Header image: Working title. by Rick Camacho, CC BY 2.0, via flickr via Openverse — cropped to 16:9 and colour-adjusted.
Key takeaways
- AutoData discovered entire algorithms for pre-training data selection, outperforming human-designed pipelines.
- It treats data engineering as a machine learning problem, not a manual art.
- The discovered algorithms transfer and improve performance at larger scales.
An AI agent autonomously discovered entire algorithms for selecting pre-training data. And it did it better than human-designed pipelines.
This is a fundamental reframing of data engineering as a machine learning problem, not a manual art. If the results hold—and the transferability tests in the paper suggest they do—we’re watching a significant development in data curation.
The Bottleneck No One Could Crack
The standard playbook? Heuristic engineering over per-document features—lexical statistics, categorical labels, perplexity scores.
Until AutoData.
AutoData’s Approach: Treating Data Selection as an Agentic Problem
It searches directly over executable selection algorithms. The agent explores a program space of scoring rules, stratification strategies, and stochastic selection mechanisms. It’s discovering the heuristics themselves.
The search space includes combinations of rules. The agent mixes and matches these rules, uncovering feature interactions human designers would miss.
It treats data selection as a program synthesis problem, where the goal is to discover the optimal algorithm for curating a dataset.
How AutoData Refines Its Own Algorithms
AutoData’s search process is iterative and feedback-driven. The agent iteratively refines algorithms with validation feedback from a proxy model. That performance signal guides the agent toward better algorithms.
AutoData completes its search within an overnight run. And crucially, the discovered algorithms transfer. The paper shows that a selection algorithm found using a small proxy model improves performance on a larger scale, as measured by CORE, a benchmark for evaluating pre-training data quality.
Performance: Outperforming Human-Designed Pipelines
The headline result is simple: AutoData discovered a selection algorithm that outperforms existing human-designed curation pipelines. Full stop.
The algorithm was discovered using a small proxy model, but it improved CORE when applied at larger scales.
The Implications: Data Engineering as Machine Learning
The paper’s core claim is bold: data engineering can be treated as an agentic machine learning problem.
Why?
Limitations and Open Questions
The search was conducted on a small proxy model, leaving open questions about scalability.
The Human Factor: What’s Left for Researchers?
Far from it.
How do different properties of data interact? What are the limits of agentic data engineering?
It’s liberating.
The Road Ahead: Where Does Agentic Data Engineering Go Next?
But one thing is clear: AutoData has shown that data engineering is no longer a manual art. It’s a machine learning problem.
The End of Manual Data Curation?
AutoData automates algorithmic discovery. It outperforms human-designed pipelines. It proves transferability. It’s a leap.