Over the past few months, I conducted a self-experiment.
Following the idea that “the best test is the self-test,” I worked through four different personality and strengths assessments — including tools used in assessment centers of major companies. I repeated some of them up to three times. In times of scaling and AI-supported preselection, I wanted to understand what we are actually basing decisions on.

The result made me think.

When I place my own results side by side, I am looking at two different people. The same person, partly within the same week — but with hardly any overlap. In psychometric terms: the convergent validity between the instruments was alarmingly low. And the repeated tests revealed the next problem — weak test-retest reliability: when someone takes the same test again after a few weeks, they often end up somewhere else. A measure that shows something different the second time does not measure anything stable.

What particularly bothered me was this: even the tests that do not assign types or categories ultimately provide deterministic descriptions — statements that define who I “am.” They may avoid the crude typologizing of the classic Myers-Briggs Type Indicator (MBTI), but they reproduce the same underlying assumption: that personality is a fixed state rather than a snapshot in time. Yet traits are demonstrably distributed dimensionally — not either/or, but more or less, as the psychometrically robust Five-Factor Model (Big Five) has shown for decades. On top of that comes the Barnum effect: vague formulations in which everyone can recognize themselves.

And one final point genuinely surprised me: a test does not become more valid simply because the user is asked at the end to confirm whether they were “in a quiet environment” and tested “in their preferred language.” That is not control of confounding variables — it is shifting responsibility for them onto the test taker. True measurement quality comes from sound test construction, not from a self-report checkbox.

Daraus ergeben sich für mich unbequeme Fragen:
👉 Wurden die Konstruktionsprobleme des MBTI je wirklich gelöst – oder nur eleganter verpackt?
👉 Was bedeutet das für Führung, wenn wir Teams nach „Farben“ und „Typen“ sortieren?
👉 Kann man einen Menschen am Ende doch nur über qualitative, dialogische Verfahren wirklich einschätzen?
👉 Und etwas provokant: Müssten Recruiter nicht eher einen Kurs in qualitativer Forschung belegen – Interviewführung, Triangulation, Quellenkritik – bevor sie Menschen für Stellen auswählen?

I do not have a ready-made answer to that. But I am convinced: anyone who measures people should at least know how well their measurement instrument actually measures.

How do you approach this — do you rely on assessment tools, on conversation, or on both?