5 min read
Fix the Data Before You Fine-Tune
DataMLTraining

Cover art for data quality before fine-tuning
Fine-tuning is expensive attention. Point it at clean, representative data or you will polish noise.
Audit label guidelines with two people on the same samples. Disagreement is free insight; silent ambiguity becomes training signal trash.
Watch for leakage across train and test — timestamps, users, documents that appear in both with near-duplicates.
Start with strong baselines: classic classifiers, good retrieval, careful prompting. Fine-tune when the gap is proven.
Version datasets like code. If you cannot rebuild last quarter’s training set, you cannot debug last quarter’s model.