Is the personality test real?
Yes, and it is short on purpose. Thirty-eight statements score you on five traits: your energy, how you take in information, how you decide, how much structure you want, and how steady you run under pressure.
Those five are the Big Five, the model most personality research has settled on. Goldberg laid out the five-factor structure in 1990 and Costa and McCrae published a full questionnaire for it in 1992. John and Srivastava (1999) traced where the model came from, and Soto and John (2017) built one of the measures in use today.
Source · Goldberg, 1990 · Costa and McCrae, 1992 · John and Srivastava, 1999 · Soto and John, 2017
How a score is built
The test on this site is thirty-eight statements: eight each for energy, information, decisions and structure, and six for the fifth trait. Each answer carries a weight, and the weighted sum lands you at one position between 0 and 100 on that trait. The same answers always produce the same numbers, which is arithmetic rather than proof of anything: a calculator is consistent too.
That 0 to 100 is where your answers sit between the two ends of the scale. It is not a rank against other people: a rank needs a norm sample, and we do not have one. Nor have we published reliability or validity figures for these thirty-eight statements, the kind a research inventory carries, so read the numbers as a careful reading of what you answered and not as a measured fact about you.
Did you write the questions yourselves?
Yes. We wrote all thirty-eight, aimed at the same traits the public-domain item pool measures (Goldberg and colleagues, 2006), and worded for a phone rather than a lab. The pool is the reference, not the text: no statement here is lifted from it, and none is lifted from a test you have to pay for.
Source · Goldberg and colleagues, 2006
Where do the four letters come from?
From four of your five scores, each turned into a letter. McCrae and Costa (1989) found that those four letters line up with four of the Big Five traits, and argued in the same paper that the scores tell you more read as scales than as boxes. That is exactly how we use them here. Your fifth trait, Emotional Climate, has no letter at all, so it stays a number.
So a code is a label for the numbers, and a label is the part that can be wrong: when a score sits near the middle, answering on a different day can put that letter the other way. Reviewers who examined four-letter systems went further and argued the whole idea of sorting people into them is not supported (Pittenger, 2005; Stein and Swan, 2019). We think they are right, which is why the number is the result here and the letter is only its name, and why we print the runner-up code beside your own whenever the call was close.
Source · McCrae and Costa, 1989 · Pittenger, 2005 · Stein and Swan, 2019
What a close call looks like
Say your deciding trait comes out at 52 out of 100. Anything from 40 to 60 counts as a close call here, so 52 is one: answer on a different day and that letter could go the other way. When that happens we say so on your result and name the runner-up code beside your own, instead of printing one code as if the call had been clear.
Those ten points either side of the middle are a line we drew, not a margin of error we measured. A real one would come from giving the same people the same test twice, and we have not done that, so this page gives no rate for how often a letter moves and you should not trust one from anybody who has not measured it either.
Why do you ask about gender?
Men and women answer a little differently on average, most clearly on two of the five traits (Costa, Terracciano and McCrae, 2001). The direction repeated in all twenty-six cultures that study covered, but the size of the gap did not: it was widest in Western samples, which the authors called a surprise themselves. And every one of these gaps is small next to how much people differ inside each group.
So the answer moves one trait: how steady you run under pressure, by four points out of a hundred, toward the middle of the group you named. Those four points are a judgement of ours, deliberately smaller than the close-call band, not a figure read off a norm sample we do not have. They can tip a call that was already close and cannot overturn a clear one, and the other four traits are read the same way for everyone. The question also has an answer for anyone it does not fit, and that answer is read exactly like a skip: everything against the overall average, nothing guessed.
Source · Costa, Terracciano and McCrae, 2001
Do people change?
Slowly, and mostly in the same direction: people tend to become steadier and more conscientious as they get older (Roberts and colleagues, 2006). Read that carefully, because it is easy to stretch. It tracks the average of large groups across decades. It does not say how much your own score can move between last month and today, and we have not measured that for this test either.
So if you take it again and a letter has moved, the honest first explanation is not that you changed. Look at the close calls: a letter that flips is usually one we had already told you was sitting near the middle.
Source · Roberts and colleagues, 2006
Where does a compatibility score come from?
From the four letters on each side and from how steady each of you runs under pressure: alike where it helps, different where it stays interesting.
The steadiness half is the half we usually do not have, because it is a number and not a letter. That is why a pair page shows a range instead of a single figure, and why the range narrows the moment one of you has taken the test. It is our reading of two profiles, not a measurement of your relationship. Use it to start a conversation, not to end one.
Attachment, reading a real conversation, and what a read cannot see are answered in the Charactly app, and you can join the early-access list.
Charactly is a reading, not a clinical assessment and not a diagnosis. Where a number here and the person you know yourself to be disagree, believe the person.