Crash-tested, not validated.
Before NORTH reads a real organisation, it reads hundreds of synthetic ones. Here is exactly what that proves — and what it cannot.
We tell you what the diagnostic can and cannot do.
A leadership diagnostic only earns trust when its limits are visible. We publish the corpus we test against, the checks we run, and the cases the tests cannot cover.
What the crash test proves: consistency (the same answers always produce the same reading), coverage (all eight archetypes and every band are reachable), and sensitivity (injected failure patterns move the reading). What it cannot prove: accuracy. The synthetic cohort is generated from the same rulebook that scores it — a consistency check, not independent ground truth. Anyone who shows you a validation number built that way is grading their own homework. We would rather you know how ours is built.
Four gates. One pipeline.
Every release of the SIGNAL classifier must pass these thresholds before it reaches a respondent.
The corpus ships with the site: 300 deterministically generated organisations, 2,700 simulated responses, every one scored by the same production engine that reads real respondents.
Three records per organisation across a baseline and two rescore points. The corpus exercises the full scoring space rather than a few showcase profiles.
The engine is a rule system, not a model. Identical inputs always produce the identical reading — no sampling noise, no mood. If we change a rule, the version changes with it.
Every archetype and readiness band is reachable, and when a synthetic organisation carries an injected failure pattern, the reading moves. The engine responds to the failures it was built to find.
How we manufacture doubt on purpose.
Generate synthetic organisations
We generate 300 synthetic organisations with assigned archetypes, dimension profiles and trajectory types, using a seeded deterministic generator. The corpus ships publicly so anyone can inspect it.
Apply controlled perturbations
For sensitivity tests we deliberately shift one or two dimensions against the archetype fingerprint — for example, making a Governance Steward score low on Accountability. The test then checks whether the reading moves.
Run the production engine
The same code that reads real respondents reads the synthetic data. No separate test harness, no special casing — what the corpus experiences is what you experience.
Publish the limits
The synthetic cohort is generated from the same rulebook that scores it, so these tests cannot prove accuracy — and we do not claim they do. They prove consistency, coverage, and sensitivity. Real-world proof begins with the first rescored programme, and we will publish those results, including the misses.
Five dimensions. Eight ranges.
The classifier matches a respondent's dimension pattern against these fingerprints. The ranges are deliberately tight so that edge cases surface as contradictions, not confident misclassifications.
What we do not claim.
It is not a predictive test.
NORTH reads how you trade off values under pressure. It does not predict promotion, failure, or stock price.
Archetypes are context-dependent.
The same leader can surface a different archetype in a different organisational moment. The score is a photograph, not a tattoo.
Synthetic tests cannot prove accuracy.
The crash-test cohort is generated from the same rulebook that scores it. It proves consistency, coverage, and sensitivity — never how often the engine is right about real people. That proof comes only from rescored real programmes, and we will publish it, including the misses.
It does not replace judgment.
The Brief and the Coalition Map are inputs for leadership conversation. They are not instructions.
Read the full methodology.
The 8 archetypes, 5 dimensions, 4 readiness bands, and the philosophy behind the NORTH diagnostic.