Methodology
How the Archetypes Are Built: The Methodology
Archetype's ten personality types were built by clustering 145,388 real Big Five (OCEAN) profiles from the Johnson IPIP-NEO-300 sample, then validated on roughly 853,000 more responses from Open Psychometrics. We tested whether 16 clusters held up; they didn't. Ten did. Your answers become z-scores against population norms, get residualized against a response-style axis (PC1), and are matched to the nearest archetype centroid by distance. This page tells the whole story, with the real numbers and the honest limits.
145,388
real Big Five profiles clustered
~853K
more responses it was validated on
10 of 16
candidate clusters that actually held up
50 questions
→
Score 5 traits
→
Residualize (PC1)
→
Nearest archetype
The question behind the personality archetype methodology
Most personality tests start with a theory and sort you into it. We started with a question: how many genuinely distinct kinds of people are there, really? Not how many we'd like there to be, or how many make for a tidy quiz, but how many actually show up when you look at hundreds of thousands of real personality profiles.
To answer it honestly you need two things. First, a measurement model that academic psychology actually trusts: the Big Five, also called OCEAN, which scores five broad traits, Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism (we frame the last as emotional style). The Big Five is the model decades of peer-reviewed research converged on because it replicates across cultures and predicts real outcomes. Second, you need a lot of real people, scored the same way, so that any pattern you find is in the data and not in your imagination.
That framing matters because it sets the bar. An archetype is only worth naming if real people actually cluster there. Everything below is the process of finding out which clusters were real and which were wishful thinking.
Clustering 145,388 real Big Five profiles
The foundation is the Johnson IPIP-NEO-300 dataset: 145,388 people who each completed a 300-item Big Five inventory and consented to anonymous research use. That depth of measurement matters. With 300 items per person, each of the five trait scores is estimated precisely, so the map of where people sit in personality space is sharp rather than blurry.
We represented every person as a point in five-dimensional space, one coordinate per trait. Then we used clustering, an unsupervised method that looks for natural groupings without being told in advance what to find. The algorithm doesn't know about 'Diplomats' or 'Captains'; it only sees the geometry of 145,388 points and asks where they bunch together. Each cluster it returns has a centre, a centroid, which is the average personality profile of everyone in that group.
Crucially, clustering doesn't tell you the right number of groups; you have to test candidate numbers and see which structure is stable and meaningful. So this step produced candidates, not conclusions. The real work was deciding which of those candidate clusters described a coherent kind of person and which were just an artifact of asking the algorithm for too many groups.
Why 16 clusters didn't hold up and 10 did
It is tempting to land on 16 types. The number is culturally familiar from four-letter typologies, and asking the algorithm for 16 groups will always return 16 groups, that's what it's built to do. But returning a cluster and having a real cluster are different things. When we examined a 16-cluster solution, several of the groups were thin, unstable, or sat so close to their neighbours that the boundary between them was noise rather than signal. Split a real population into too many pieces and you start carving distinctions that don't survive a second look.
Ten clusters held up. At ten, each archetype occupied a distinct, well-populated region of personality space, the centroids were far enough apart to be genuinely different profiles, and the structure reproduced rather than dissolving into rounding error. The six clusters that fell away weren't deleted to make a rounder number; they failed the test of being a stable, separable kind of person.
The honest takeaway is that ten is what the data supported, not a marketing decision. The ten that survived are The Diplomat, The Captain, The Anchor, The Organizer, The Dreamer, The Skeptic, The Caretaker, The Explorer, The Rebel, and The Loyalist. Each is the centroid of a region where real people actually concentrate.
Validating on ~853,000 more responses
A pattern found in one dataset can be a quirk of that dataset. The test of a real structure is whether it reappears in data it has never seen. So we took the ten-archetype structure derived from the 145,388-person Johnson sample and checked it against an entirely separate body of data: roughly 853,000 responses from the Open Psychometrics IPIP-50 sample, a shorter 50-item Big Five inventory completed by a different population of people.
This is the step that separates a defensible model from a horoscope. The Open Psychometrics respondents were never part of building the archetypes. If the ten centroids only described the original group, they would have matched the new data poorly. They didn't, the same ten regions of personality space were populated in the validation sample too. The structure travelled across two independent datasets and two different inventory lengths, which is exactly what you want to see before trusting it.
The 50-item inventory also happens to be the test you take in the app, about seven minutes of questions. So validation did double duty: it confirmed the ten archetypes generalise, and it confirmed they can be recovered from the shorter survey real users actually complete.
The scoring pipeline in plain language
Here is exactly what happens to your answers, with no black box. Step one: you answer 50 statements on a 1-to-5 agree-disagree scale. We average the relevant items (reversing the negatively worded ones) to get a raw score for each of the five traits.
Step two, z-scores. A raw score on its own is meaningless, scoring a 3.8 on Extraversion tells you nothing until you know how the population scores. So we convert each raw trait score into a z-score: how far above or below the population average you sit, measured in standard deviations. A z-score of 0 is dead average; positive is above, negative is below. This puts all five traits on the same comparable scale.
Step three, PC1 residualization. There's a well-known nuisance in self-report data: some people tend to agree with everything (acquiescence) or use the scale in their own characteristic way, which nudges all five trait scores up or down together. That shared response-style direction shows up as the first principal component, PC1. Before matching, we statistically remove your position along that PC1 axis, so we compare your trait shape, the contour that's distinctively you, rather than your overall tendency to agree. This is the single biggest reason the matching is robust rather than naive.
Step four, nearest centroid. We then measure the straight-line (Euclidean) distance from your residualized five-trait point to each of the ten archetype centroids, and match you to the closest one. That's it: average, standardise, remove response-style, find the nearest real cluster. No hidden weighting, no personality-quiz sleight of hand.
What this does, and does not, mean
This is the part most personality products skip, so read it carefully. Your archetype is the nearest centroid, the closest of ten real population clusters to your particular trait profile. It is a useful, evidence-based summary of where you sit. It is not a box, a destiny, or a fixed essence.
Traits are spectra, not switches. The Big Five describes you with five continuous dials, and almost everyone lands somewhere in the middle of most of them. Your archetype is shorthand for the overall shape those dials make, but two people sharing an archetype can still differ, and you may sit near a boundary between two. That's why the app shows your actual trait percentiles, not just a label, the label is the headline and the traits are the real story.
Big Five is descriptive, not prescriptive. It tells you how you tend to behave, on average, across situations; it does not predict any single moment, diagnose anything clinical, or set a ceiling on who you can become. People change, and a seven-minute self-report is a snapshot, not an X-ray. We'd rather you trust the method because we're honest about its limits than because we oversold it. No test is destiny, and anyone who tells you otherwise is selling something.
Why the model can be trusted: reproducibility
Trust in a method should come from how it was built, not from how confidently it's asserted. Three things make this model checkable rather than merely claimed. First, the inputs are public, named research datasets: the Johnson IPIP-NEO-300 sample and the Open Psychometrics IPIP-50 sample, both built on the open-source IPIP item pool. We didn't invent a proprietary questionnaire whose properties no one can inspect.
Second, the structure was cross-validated. The ten archetypes were derived on 145,388 people and confirmed on a separate ~853,000, across two different inventory lengths. A result that reappears in independent data is the opposite of a fluke; reproducibility is the entire point.
Third, the scoring pipeline is fully specified and deterministic: z-scores against fixed population norms, residualization against a fixed PC1 axis, nearest-centroid matching by Euclidean distance. The same answers always produce the same result, and every step is one you could in principle recompute by hand. The Big Five engine underneath is the one academic psychology actually uses, which is the whole wedge here, the identity payoff of 'I'm The Diplomat' resting on a foundation you're allowed to interrogate, rather than asked to take on faith.
Questions people ask first
How many personality profiles were used to build the archetypes?
The ten archetypes were clustered from 145,388 real Big Five profiles in the Johnson IPIP-NEO-300 sample, then validated on roughly 853,000 additional responses from the Open Psychometrics IPIP-50 sample. That's close to a million people across the two datasets, which is what lets us say the structure is real and not a quirk of one group.
Why are there 10 archetypes and not 16?
Because ten is what the data actually supported. A clustering algorithm will return any number of groups you ask for, so we tested candidates. At 16, several clusters were thin, unstable, or too close to their neighbours to be genuinely distinct. At 10, each archetype occupied a well-populated, separable region of personality space and the structure reproduced across datasets. The six that fell away failed the test of being a stable kind of person.
What is PC1 residualization and why does it matter?
PC1 is the first principal component of the trait scores, which captures a general response-style or acquiescence tendency, some people simply agree with more statements, nudging all five traits up or down together. Residualizing against PC1 removes that shared tendency before matching, so we compare your distinctive trait shape rather than your overall agreeableness with the scale. It's the main reason the matching is robust instead of being thrown off by how you use a rating scale.
Is the Big Five more scientific than MBTI or 16-type tests?
Yes, by the standards psychology uses. The Big Five is the model decades of peer-reviewed research converged on; it replicates across cultures and predicts real outcomes, and it treats traits as continuous spectra. MBTI-style four-letter typologies, built on Jungian theory, are widely criticised for low test-retest reliability and for forcing each trait into a binary either/or. Archetype keeps the satisfying 'I'm a type' payoff but rests it on the Big Five foundation.
Does my archetype mean my personality is fixed?
No. Your archetype is the nearest of ten real population clusters to your current trait profile, a useful summary, not a destiny or a box. Big Five traits are continuous spectra, people change over time, and a seven-minute self-report is a snapshot rather than an X-ray. The app shows your actual trait percentiles alongside the label precisely so you can see the spectrum, not just the headline.
Can I trust the results to be consistent?
The scoring pipeline is fully deterministic: the same 50 answers always yield the same z-scores, the same PC1 residualization, and the same nearest-centroid match. It's built on public, named research datasets and the open IPIP item pool, and the ten-archetype structure was confirmed on an independent ~853,000-person sample. Consistency comes from a specified, reproducible method rather than from anything hidden.
See which of the ten you are
The test is 50 questions, about seven minutes, and built on the method above. No sign-up required to see your full result, your five trait percentiles and the archetype nearest to your profile. Free, honest about its limits, and grounded in the Big Five that psychology actually uses.
Take the free test