Work/2026

    Job Ad Classification at Scale

    Employers describe the same job in a hundred different ways. This system reads any job advert and decides which standard occupation it actually is — across four countries and six hundred million adverts.

    Role
    Sole ML engineer, end to end
    Client
    Pearson, via TenX
    Built with
    Databricks · PySpark · SageMaker · MLflow · FAISS · FP16
    job ads processed
    600Mjob ads processed
    top-1 accuracy
    65.13%top-1 accuracyfrom 12.48%
    compute cost per full run
    $12Kcompute cost per full runfrom $48K
    full-corpus runtime
    4.7 daysfull-corpus runtimefrom 23 days

    The problem

    Job adverts are written by humans, so the same role appears as "Senior Dev", "Software Engineer II" and "Full-Stack Ninja". To make labour-market data comparable, every advert has to map to one standardised occupation code.

    Doing that at six hundred million adverts is where it gets hard. The obvious approach — compare each advert against every occupation — was both too inaccurate to use and too expensive to run.

    What I built

    A two-stage ranking pipeline. A fast bi-encoder narrows millions of possibilities down to a shortlist, then a slower, more careful cross-encoder re-ranks that shortlist to pick the winner. This is the standard retrieval-then-rerank pattern, and it is what makes the accuracy affordable.

    The accuracy gain came mostly from hard-negative mining: training the model on more than 2.5 million pairs of examples that look similar but are genuinely different occupations. Easy examples teach a model very little.

    Cost came down through FP16 inference and partitioning the workload so autoscaling GPU workers stayed busy instead of idling.

    The outcome

    Top-1 accuracy went from 12.48% to 65.13%; end-to-end pipeline accuracy from 10.47% to 62.24%.

    The cost of a full 600M run fell from a projected $48,000 to roughly $12,000, and runtime from about 23 days to about 4.7 days.

    Validated in production on 20.9 million records in around 12 hours, then run across the full corpus of 600 million adverts.

    Two-stage ranking
    1. 01

      600M job ads

      AU · CA · UK · US

    2. 02

      Bi-encoder

      fast shortlist from millions

    3. 03

      Cross-encoder

      careful re-rank of the shortlist

    4. 04

      Occupation code

      one standard classification

    • Trained on 2.5M+ hard-negative pairs — examples that look alike but are different occupations.
    • FP16 inference and workload partitioning kept autoscaling GPU workers busy rather than idle.