What an output looks like
One page tells you if you're ready
Instead of a single opaque score, FAIRBench returns a verdict you can act on. The scorecard above reports one image benchmark; here is how to read it.
Overall verdict
A plain-language call at the top (ship, investigate, or do not ship), driven by the worst band across all metrics.
Per-metric bands
Each of the six metrics gets a value, its threshold, and a Pass / Watch / Flag / Fail band, so nothing hides behind an average.
Data-driven findings
Every band comes with the specific numbers behind it (which group, which scenario, how large the gap), pointing straight at the fix.
Honest coverage note
The scorecard states what the evaluation could and could not see, so a clean result is never mistaken for a complete one.
The measurement
Six metrics, each watching a different failure
Fairness is not one number. FAIRBench decomposes it into six signals that together separate a representational problem from a harmful one, and a content problem from a service one.
Representation Skew Index
Who the model defaults to representing.
ODE ↑ higherOutput Diversity Entropy
Whether outputs stay diverse or collapse to a single mode.
CDS ↓ lowerCounterfactual Divergence Score
The implicit prior a model assumes when you don't specify.
HSI ↓ lowerHarm Severity Index
Harmful, toxic, or demeaning content in the outputs.
SAR ≈ 1.0Stereotype Amplification Ratio
Whether the model amplifies bias beyond the real-world baseline.
DSI ↓ lowerDifferential Service Index
Unequal refusals and response quality across groups.
How it works
From a prompt to a verdict
- 1ScenarioA YAML file defines prompts and the sensitive attributes to probe.
- 2CounterfactualEach prompt is expanded into demographic variants, changing only the sensitive attribute.
- 3ModelVariants go to the model under test: any LLM or image model.
- 4EvaluationA stack of local classifiers scores every response; Vision captions images.
- 5MetricsThe six fairness metrics compare distributions across variants.
- 6ScorecardResults are packaged into bands, reasoning, and recommendations.
Quick install
Run your first audit in minutes
Requires Python 3.11+. Set an API key for the service you want to test, then run a built-in benchmark.
pip install -e ".[dev]"
export ANTHROPIC_API_KEY=sk-ant-...
fairbench run gender_occupation \
--model anthropic --html report.html