Building a Spend Taxonomy for a $10B+ Procurement Portfolio
Defined a four-tier spend taxonomy and trained an ML classifier to give executives visibility into a large procurement portfolio.

Athul P M
Senior Associate Consultant at Indus Insights





From their time as

Associate Consultant II
Indus Insights • 2025 - 2025
Overview
Athul joined a supply chain project with a structural visibility problem: a procurement team managing a $10B+ portfolio had no consistent way to look at spend across categories. Each category manager was using their own taxonomy, and the classifications that worked at the analyst level did not translate into anything meaningful for executive decision-making.
The Story
Athul joined a supply chain project with a structural visibility problem: a procurement team managing a $10B+ portfolio had no consistent way to look at spend across categories. Each category manager was using their own taxonomy, and the classifications that worked at the analyst level did not translate into anything meaningful for executive decision-making.
He spent months in working sessions with individual category managers, learning what actually went into each category from the people who owned it. From those conversations, he ran frequency analyses to understand which sub-segmentations were statistically significant, which were legacy artifacts, and which mattered from a dollar standpoint.
Making the Hard Calls
The messier part was that many legacy categories were emotionally significant to the people who had built them. Removing a classification was not just a data decision; it required Athul to make a case grounded in business logic. His anchor was dollar significance: a category with a million invoices totaling less than $100K in annual spend was not worth maintaining at the executive level, regardless of its historical granularity.
He aligned with the Director of Procurement on this framing, which gave him the confidence to make calls that the data supported even when they ran against established convention. The result was a four-tier spend taxonomy that was generalized across the firm and designed for executive use.
He also owned the training-data strategy for an ML invoice classifier built on top of the taxonomy, defining the sampling logic that ensured reliable classification and cutting model-validation time per analyst by more than 20 hours through automated identification of misclassified entries.
