Work/2026
Job Ad Classification at Scale
Employers describe the same job in a hundred different ways. This system reads any job advert and decides which standard occupation it actually is — across four countries and six hundred million adverts.
- Role
- Sole ML engineer, end to end
- Client
- Pearson, via TenX
- Built with
- Databricks · PySpark · SageMaker · MLflow · FAISS · FP16
- job ads processed
- 600M
- top-1 accuracy
- 65.13%
- compute cost per full run
- $12K
- full-corpus runtime
- 4.7 days
The problem
Job adverts are written by humans, so the same role appears as "Senior Dev", "Software Engineer II" and "Full-Stack Ninja". To make labour-market data comparable, every advert has to map to one standardised occupation code.
Doing that at six hundred million adverts is where it gets hard. The obvious approach — compare each advert against every occupation — was both too inaccurate to use and too expensive to run.
What I built
A two-stage ranking pipeline. A fast bi-encoder narrows millions of possibilities down to a shortlist, then a slower, more careful cross-encoder re-ranks that shortlist to pick the winner. This is the standard retrieval-then-rerank pattern, and it is what makes the accuracy affordable.
The accuracy gain came mostly from hard-negative mining: training the model on more than 2.5 million pairs of examples that look similar but are genuinely different occupations. Easy examples teach a model very little.
Cost came down through FP16 inference and partitioning the workload so autoscaling GPU workers stayed busy instead of idling.
The outcome
Top-1 accuracy went from 12.48% to 65.13%; end-to-end pipeline accuracy from 10.47% to 62.24%.
The cost of a full 600M run fell from a projected $48,000 to roughly $12,000, and runtime from about 23 days to about 4.7 days.
Validated in production on 20.9 million records in around 12 hours, then run across the full corpus of 600 million adverts.
- 01
600M job ads
- 02
Bi-encoder
- 03
Cross-encoder
- 04
Occupation code