August 17, 2026·Science·8 min read

From Gaokao Fairness to Ranking Systems: A Detour Through Evaluation

We began by asking how a single exam solution should be graded. Before long, the harder question was whether a fair ranking measures facts, or quietly defines what kind of person a system wants.

A few days ago, a friend and I began talking about fairness in the Gaokao. The starting point was modest: multiple-choice questions invite guessing, while open-response questions inevitably admit some human judgment.

My first thought was simple. Why not use more short-answer questions in quantitative subjects? They retain the convenience of objective grading while making blind guessing far less useful. My friend cared more about a different criterion: if sufficiently similar solutions always receive sufficiently similar scores, then the system is fair. I was not fully satisfied. A system can be remarkably consistent and consistently wrong. Low variance does not make bias disappear. Equal treatment matters, but it is no substitute for a correct judgment.

AI grading makes that disagreement less sharp. In mathematics or physics, a future system need not enumerate every acceptable solution in advance. It could inspect chains of reasoning, call symbolic algebra or theorem-verification tools, and send only unusual or uncertain cases to human reviewers. This could reduce both random grader variation and the familiar penalty imposed on a correct solution because the grader has never seen it before.

But that solves only how to grade, not what to reward. The harder question would remain: what kind of performance deserves how many points in the first place?

The conversation had quietly wandered from examination technique into social choice.

What a single total score throws away

Society plainly needs more than one kind of capable person. Some calculate quickly; some prove things unusually well; some can model a messy problem. Others may be unremarkable under time pressure and exceptional when the question is open-ended. If these abilities are represented as a vector, a conventional total score merely chooses a set of weights and compresses the vector into one number.

But why should every university and every discipline use the same weights?

An engineering program may care strongly about whether a result is correct and a model actually works. A mathematics program may care more about proof, abstraction, and nonstandard reasoning. An apparently natural design follows: a national examination system measures a multidimensional profile as reliably as possible, while universities publish their value functions in advance and use different weights to select different kinds of students. Common measurements preserve comparability without pretending that “the best student” has one universal definition.

Taken one step further, this is no longer only a Gaokao problem. Competitions, hiring systems, university rankings, and even machine-learning benchmarks can all be separated into four layers: measurement, valuation, ranking, and audit. First ask what a person or system actually demonstrated. Then state which dimensions matter. Only then produce a ranking. Finally, test whether the order survives reasonable measurement error and changes in weights.

The neatest sentence to emerge from this scheme was:

Measurement should be as objective as possible; value judgments need not be universal, but they should be explicit.

The sentence is persuasive enough to feel new. Unfortunately, one of intuition's small talents is to mistake familiarity for discovery.

Literature searches are useful ways to cool an intuition

Friedler, Scheidegger, and Venkatasubramanian had already distinguished latent constructs, observed quantities, and downstream decisions. Jacobs and Wallach explicitly brought measurement theory into the study of fairness: measuring something reliably is one matter; measuring the thing one intended is another. Bothmann and colleagues likewise emphasize that fairness cannot be separated from normative assumptions. They gave the world we think ought to exist a disarmingly candid name: the fictitious world.

The larger collision came from a field I had not initially considered at all: multicriteria decision analysis. SMAA, introduced in 1998, already allowed uncertainty in criteria and weights, and explored the weight space to ask what kinds of valuation would make each alternative preferable. Robust Ordinal Regression later distinguished necessary from possible preference. If A outranks B under every value function compatible with the stated preferences, the relation is necessary; if it holds only under some compatible functions, it is merely possible. Corrente and colleagues extended the same framework to imprecise evaluations themselves.

The question can therefore be narrowed. Instead of insisting on a handsome list of ranks, one might first ask whether the list is entitled to sound so certain. If A and B are separated by less than the measurement error, or if a small change in weights reverses them, solemnly announcing first place and second place is a little like measuring fog with a vernier caliper.

Others had, of course, walked this path as well. Singh, Kempe, and Joachims studied fair ranking when merit itself is uncertain. Statistics offers even more direct methods for constructing confidence sets for ranks. In the framework of Mogstad and colleagues, an honest conclusion may sometimes be not “this object ranks third,” but “its true rank may lie between second and sixth.” Forcing noisy estimates into a complete order with no uncertainty attached mistakes the tidiness of computer output for the sufficiency of evidence.

The four-layer scheme did not disappear; it merely lost the pleasant costume of novelty. Independently arriving at an old idea does not make it a discovery, though it often shows that the idea grows naturally from the structure of the problem. Many questions become interesting not when the first elegant thought appears, but when that thought collides with what is already known and reveals the part that remains unresolved.

A ranking is not a stationary ruler

Continuing the argument changed the problem once again.

Almost everything above treats those being evaluated as stationary. The rules are set, people submit their performance, and the system measures and ranks it. Real people, however, can read the rules—and usually read them rather carefully.

If a university announces that it strongly values a particular attribute, schools, families, and tutoring markets will train for that attribute. If a university ranking emphasizes a handful of indicators, universities will redirect resources toward those indicators. If a machine-learning benchmark becomes prestigious, model development will optimize toward that benchmark. Once the ruler is placed on the table, the thing being measured begins to grow toward the ruler.

This, too, is well-trodden ground. Strategic classification studies how people change observable features once they know the classification rule. Strategic ranking places competition and ordering directly into settings such as college admissions. Performative prediction studies the wider phenomenon in which a prediction or decision changes the world it was meant to predict. Work on ranking games has likewise examined the incentives rankings create for both ranked actors and ranking providers.

Yet at this point, the original Gaokao conversation became more interesting rather than less.

Fairness may never have been a matter of finding the best exam paper, the smartest AI grader, or the perfect weighted formula. An evaluation system both measures people and trains them. It describes what counts as excellence while helping manufacture that form of excellence. A rule may be highly accurate when examined statically and still alter its object through repeated social use, until even the meaning of the original measurement has changed.

The harder question may therefore be not “How do we rank people most fairly?” but this:

Can an evaluation system still measure what we originally cared about after everyone has learned to optimize against it?

That is a long way from multiple-choice questions, short answers, and grading variance. I rather like the detour. It left no universal ranking formula and removed several attractive answers along the way; but as the answers diminished, the question became clearer.

Sometimes the point of reading the literature is not to prove one's intuition right. It is to peel away, layer by layer, the parts other people have already understood. Whatever remains after that removal is where thinking truly begins.

References

  1. On the (im)possibility of fairness · Sorelle A. Friedler, Carlos Scheidegger, Suresh Venkatasubramanian
  2. Measurement and Fairness · Abigail Z. Jacobs, Hanna Wallach
  3. What is Fairness? On Protected Attributes and Fictitious Worlds · Ludwig Bothmann, Kristina Peters, Bernd Bischl
  4. SMAA - Stochastic multiobjective acceptability analysis · Risto Lahdelma, Joonas Hokkanen, Pekka Salminen
  5. Ordinal regression revisited: Multiple criteria ranking using a set of additive value functions · Salvatore Greco, Vincent Mousseau, Roman Słowiński
  6. Robust Ordinal Regression in case of Imprecise Evaluations · Salvatore Corrente, Salvatore Greco, Roman Słowiński
  7. Fairness in Ranking under Uncertainty · Ashudeep Singh, David Kempe, Thorsten Joachims
  8. Inference for Ranks with Applications to Mobility across Neighbourhoods and Academic Achievement across Countries · Magne Mogstad, Joseph P. Romano, Azeem M. Shaikh, Daniel Wilhelm
  9. Strategic Classification · Moritz Hardt, Nimrod Megiddo, Christos Papadimitriou, Mary Wootters
  10. Strategic ranking · Lydia T. Liu, Nikhil Garg, Christian Borgs
  11. Performative Prediction · Juan Perdomo, Tijana Zrnic, Celestine Mendler-Dünner, Moritz Hardt
  12. Ranking Games · Margit Osterloh, Bruno S. Frey