Type
Website article
Status
ready-for-review
Week
4

Why Triangulating Your Assessments Never Closed the Performance Gap

A senior appointment is preceded by a thorough assessment: a personality profile, a competency assessment, and a structured multi-rater review, all of them pointing the same way. The process was not rushed and no corner was cut, and within a year the appointment is plainly wrong. The instinct afterwards is that something in the process must have been too thin. The harder possibility is that the process was rigorous, the instruments agreed for good reason, and all of them were reading the wrong layer.

Executive summary

The appointment that did everything right

There is a particular kind of hiring failure that unsettles a leadership team more than a careless one does. The careless appointment at least offers a lesson: slow down, check more, involve more people. The unsettling case is the one where all of that was already done. The candidate was assessed on personality, evaluated against a competency framework, and reviewed by people above, beside, and below them, and the readings converged on a confident recommendation. The appointment was made on the strength of that convergence, and it still did not work.

Faced with this, the natural response is to assume a gap in coverage and to close it: add another instrument, lengthen the process, bring in a further panel. That response makes sense only if the miss was caused by insufficient measurement. It is worth pausing on the fact that the measurement already in hand was both extensive and in agreement, because more of the same kind of measurement is precisely what already converged on the wrong answer.

Agreement is not the same as independent confirmation

The reason convergence is so persuasive is that it resembles triangulation. When three different readings arrive at the same conclusion, it feels like independent lines of evidence meeting at a point, and independent agreement is one of the most reliable signals we have. The strength of that signal, however, depends entirely on the word independent. Triangulation reduces error only to the extent that the sources of evidence do not share the same blind spot. If they do, agreement stops being confirmation and becomes a coordinated echo. The confidence it produces is then confidence in the echo, not in the answer.

The instruments differ in method and read the same layer

This is where the assessment stack quietly fails a test it appears to pass. The instruments are genuinely different in method. A personality inventory, a competency evaluation, and a multi-rater review are not the same tool, and combining unlike methods is exactly what proper triangulation calls for. What they share is not their method but their subject. Every one of them reads the observable layer: what a person has done, how they tend to operate, and how the people around them perceive both.

That layer is the midstream, the behaviours and traits and perceptions the whole existing landscape of assessment is built to capture, and it is real and worth measuring. Triangulating three different windows onto it gives you an unusually reliable picture of the midstream. It carries you not one step toward the layer underneath, because none of the three windows looks there. The agreement is entirely genuine, and it is genuine about the wrong altitude.

A thorough reading of typical behaviour, silent on the case that matters

There is a second reason a well-assessed appointment fails, and it sharpens the first. Assessment, by its nature, captures typical performance: how a person operates under ordinary conditions, sustained over time and mostly unobserved. Typical performance is a distinct thing from maximum performance, which is what a person produces when the stakes are high and everything is being asked of them at once. The two are known to correlate only weakly (https://experts.umn.edu/en/publications/relations-between-measures-of-typical-and-maximum-job-performance). A profile can therefore be an accurate account of someone's ordinary operation and a poor forecast of their conduct at the limit.

That distinction is not merely statistical; it has a physical basis. Ordinary conditions run largely on the brain's memory systems, the stored, well-practised responses that carry a person capably through situations resembling ones they have handled before. This is exactly what a track record and an assessment of typical behaviour describe. The situations that break an appointment are the ones that resemble nothing in the archive, where sustained pressure and genuine novelty hand control to a different set of brain systems entirely. A measurement of how someone performs typically is, by construction, silent about the systems that decide the hard case, because those systems only take the load when the typical case runs out.

A boundary of the instrument, not a failure of it

None of this makes the assessment stack defective. Each tool is valid on its own terms:

Each does what it was built to do, and each does it well. The gap is not a flaw inside any of these instruments; it is a boundary around all of them. They were built to read what a person does and reports, and capacity is not a thing a person does or reports. It is the substrate that determines what they will be able to do when demand exceeds the record. Asking a midstream instrument to read it is not asking a good tool to work harder; it is asking it to measure something outside the layer it was designed for. No number of well-chosen instruments trained on the same layer adds up to a reading of a different one.

The confidence scales faster than the coverage

The commercial danger is that the missing layer does not merely stay hidden. It hides more effectively the more thorough the assessment becomes. Each additional instrument that agrees raises the confidence attached to the decision, and none of them touches the blind direction. So a leadership team can grow steadily more certain of an appointment precisely as it accumulates more evidence that cannot speak to the thing most likely to sink it. The most consequential appointments, the ones that justify the deepest assessment (https://vanaya.co.id/leadership/), are therefore the ones most likely to be made with high confidence on the reading least able to warn. Thoroughness on the wrong layer does not correct the error. It underwrites it.

An upstream question needs an upstream measurement

The way out is not a better instrument on the same layer, but a reading of a different one. The question that decides an appointment is whether a leader has the capacity to hold when demand exceeds anything the record contains. That question is upstream of everything a personality profile, a competency framework, or a multi-rater review can see. Answering it requires measuring capacity directly, rather than inferring it from behaviour that never tested it (https://vanaya.co.id/neurometric/). That is a different kind of measurement, aimed at a different layer. It is the question the whole month has been circling: not how to assess what a leader does, which the existing stack already does well, but how to read what a leader can do before the situation that requires it arrives.