AI Build-vs-Buy Framework
Before I compared a single model, I found out my measuring stick was bent.
A 125M-parameter specialist beat six foundation models on a locked human-labeled benchmark. The recommendation was neither build nor buy. It was a hybrid operating model.
A five-person MBA research team. I led the scope, owned the build and codebase, and ran the model evaluation. Human labeling and adjudication were shared.
Fictional complaint fragments · aggregate agreement metrics measured. Source described only as a public sample from the CFPB Consumer Complaint Database.
Call an API, or build a specialist?
When a seven-class classification capability runs continuously at volume, should the organization call a general model API or build and fine-tune a small specialist?
The codebook, not the model, was the first thing to fix.
Accuracy stayed visible. It did not lead the decision.
Two classes represented 83.3% of the locked holdout; macro-F1 and MCC carry more information about minority-class performance.
A five-person MBA research project. Teja led scope, owned the build and codebase, and ran the head-to-head evaluation. Human labeling and adjudication were shared. Research project · no production deployment. Source described only as a public sample from the CFPB Consumer Complaint Database; complaint fragments shown anywhere in this artifact are fictional.