cd ../blog
5 min read

Fix the Data Before You Fine-Tune

DataMLTraining
Cover art for data quality before fine-tuning

Cover art for data quality before fine-tuning

Fine-tuning is expensive attention. Point it at clean, representative data or you will polish noise.

Audit label guidelines with two people on the same samples. Disagreement is free insight; silent ambiguity becomes training signal trash.

Watch for leakage across train and test — timestamps, users, documents that appear in both with near-duplicates.

Start with strong baselines: classic classifiers, good retrieval, careful prompting. Fine-tune when the gap is proven.

Version datasets like code. If you cannot rebuild last quarter’s training set, you cannot debug last quarter’s model.