About

ClassDex inventories 1,312 text classification benchmarks published in the ACL Anthology between 2013–2025, and organises them under a hand-built task taxonomy of 25 families and 43 leaves.

The taxonomy

Each family is defined by a discriminant description: what separates it from its neighbours (« vs Sentiment », « vs Deception ») rather than a definition in isolation. Leaves name sub-tasks wherever the literature draws a clear line. The taxonomy was built by manually annotating papers, then consolidated through disambiguation passes over neighbouring families.

Dataset availability

A benchmark counts as findable when the paper points to a dataset that can actually be retrieved. On this inventory, 31% of benchmarks (406 of 1,312) carry no reachable dataset link: they exist in the papers, but not where people look for them. The mosaic on the home page shades every family by that share.

Scope

The inventory covers single-task classification benchmarks only. Papers whose task could not be assigned to a family are kept under Unassigned rather than dropped, so the totals always add up.

Affiliation

ClassDex is developed atRALI — Recherche appliquée en linguistique informatique, atUniversité de Montréal.