noetly
Profiles with something to say

Personality tests: what each model actually measures

A test is judged on two things: the model it uses, meaning the list of dimensions it treats as real, and the instrument, meaning how it measures them. Most tests online borrow a known model and improvise the instrument.

· 2 min read

The models you will meet

The five-factor model, known as the Big Five or OCEAN, has been the research reference since the 1980s. Its five dimensions are continuous: openness, conscientiousness, extraversion, agreeableness, emotional reactivity. They reappear across most of the languages studied, which is the strongest argument in their favour.

The MBTI cuts four of those dimensions into sixteen types, and leaves the fifth out entirely. Its page covers what that cut costs in stability.

DISC, sold in many countries as four colours, describes styles of workplace behaviour. It works as shared vocabulary in a team, while its ability to predict behaviour stays weak.

The Enneagram sorts people into nine types organised around a core fear. Its empirical validation sits far below the three models above, and Noetly does not measure it.

Alongside those four, Noetly measures more specialised instruments: honesty-humility from the HEXACO model, the ten Schwartz values, the two axes of adult attachment, Holland’s six vocational interests, McClelland’s three motives, and two political axes kept apart from the rest.

What separates an instrument from a questionnaire

Three checks settle the question in front of any test you meet.

The first is the precision on display. A test reporting “78% extraversion” without the margin of error around it leaves you unable to tell whether the figure is 78 or 62, and therefore unable to do anything with it.

The second is stability over time. When a result changes on retaking the test three weeks later, the gap belongs to the questionnaire and not to the person, who did not change in the meantime.

The third is how specific the descriptions are. A line like “you have potential you do not always use” fits everybody, and a great many free tests run on that effect, described by Bertram Forer in 1949.

How Noetly works

One set of questions feeds every layer. Each answer updates a probability distribution over your traits, the next question is chosen by what remains uncertain, and written answers are read by a language model that pulls out cues feeding the same estimate.

So you do not take an MBTI test, then a values test, then an attachment test. The readings unlock as the precision becomes good enough to carry them, and the progress bar you watch is exactly that precision.

Frequently asked

Which personality test is the most reliable?
Instruments built on the five-factor model, provided they publish a margin of error and their scoring method. Reliability belongs to the instrument more than to the model: two Big Five questionnaires can differ widely in precision.
Are free personality tests worth anything?
Some are, and three checks tell them apart: is the margin of error shown, does the result hold if you retake the test a few weeks later, and are the descriptions specific enough that they could not apply to anyone.
How many questions does a real measurement need?
It depends on the precision you want and on how the questions are chosen. A fixed questionnaire often needs around a hundred items to cover facets; an adaptive test reaches the same precision in fewer, because it only asks about what is still uncertain.