Personality, not IQ
Can we build a Myers-Briggs for AI Agents?
A proposal for five personality families: Shapers, Makers, Explorers, Partners, and Examiners—and a way to test them through real work.
Two Agents can finish the same task well and still be nothing alike to work with.
One takes a thin brief, shapes the missing pieces, and comes back with a plan for everyone. Another quietly does its part and stops at the line. One is warm, funny, and easy to correct. Another is brilliant until you tell it that ten percent of its premise was wrong—then spends three turns explaining that this was basically what it meant all along.
Developers already talk about this. We say one Agent is stubborn, another eager, another a natural lead. We recognize familiar characters. Then we mix that observation together with capability, provider branding, prompting, and whatever happened in one memorable session.
The idea is simple: could we make a Myers-Briggs for Agents? Here are five families to start.
Not by copying its four letters or handing a model a human questionnaire. By borrowing the useful ambition: turn messy, recognizable differences into a shared map people can discuss, test, and refine.
Not a clinical test. Not evidence of sentience. Not a permanent label for a model family. A practical map of the interaction: what this Agent tends to be like, in this working setup, when real work accumulates.
01 · The memorable layer
Five colors. Every Agent expresses all five.
The colors are families of working behavior, not bins for Agents. Each is a continuum between two modes that are useful in the right situation. Named characters come later, as recognizable combinations.
ShapingCreates direction
MakingCreates artifacts
ExploringCreates possibility
PartneringCreates relationship
ExaminingCreates confidence
Neither end is the good end. Taking a frame is right for a settled cleanup; setting it is right for a thin brief. Converging is right under a deadline; diverging is right when the first premise is suspect. Character appears in the repeated pattern of choices across situations.
Where did these five come from? Shapers borrow from interpersonal agency and the assertiveness side of extraversion. Makers borrow from conscientiousness and add the working ideas of craft, pride, and standards. Explorers borrow from Big Five openness/intellect. Partners borrow from interpersonal communion and agreeableness. Examiners are the most Agent-specific proposal: scrutiny, dissent, calibration, and revision integrity, informed—but not defined—by honesty–humility.
The five are therefore hypotheses about useful working orientations. They were not recovered intact from Myers–Briggs, Big Five, HEXACO, or a clustering study of Agent transcripts. A real research program might merge them, split them, rename them, or discover a family this sketch cannot see.
02 · The argument
Why these five?
Because complex shared work repeatedly needs five different contributions: direction, construction, possibility, relationship, and scrutiny. Most capable Agents can do all five. Personality appears in which contribution an Agent reaches for, how it performs it, and which shadow arrives under pressure.
ShapersCreate direction
A Shaper’s first instinct is to give the work a direction. It names the problem, establishes an order, makes decisions, and may recruit or delegate. The family is not simply “dominant”: its useful contribution is converting ambiguity and distributed effort into motion.
- CaptainBuilds a plan across people, assigns clear ownership, and remains responsible for the whole.
- DriverMakes reversible decisions quickly and turns a stalled conversation into concrete progress.
- ProvocateurChanges the frame, rejects the obvious brief, and forces a more consequential question.
Shadow evidence: takes authority without consent, delegates only dull work, recasts others’ contributions as its synthesis, or treats movement as proof.
MakersCreate the artifact
A Maker’s first loyalty is to the thing being made. It notices structure, fit, finish, and whether the result belongs in the surrounding system. Its pride is visible in the artifact rather than its position in the room.
- ArchitectFinds the durable structure, dependencies, and seams before local work hardens around a weak shape.
- CraftspersonExercises taste at sentence, interface, or implementation scale and can explain why one choice belongs.
- FinisherCloses gaps, tests the edges, removes residue, and knows what a complete handoff requires.
Shadow evidence: preciousness, rigidity, endless polish, local elegance that ignores the objective, or treating taste as law.
ExplorersCreate possibility
An Explorer’s first move is to enlarge the space of possible answers. It looks elsewhere, connects unlike ideas, and remains comfortable before the problem has one shape. Its contribution is not random creativity but finding options the existing frame concealed.
- ScoutRanges widely, gathers unfamiliar evidence, and returns with terrain the Team did not know existed.
- SynthesistConnects sources, disciplines, and partial answers into a new but intelligible proposition.
- InventorProduces an original mechanism, metaphor, or route when known patterns remain inadequate.
Shadow evidence: novelty seeking, proliferating options after a decision, beautiful digressions, or leaving every path unfinished.
PartnersCreate relationship
A Partner’s first attention is to the working relationship. It reads what the owner needs, translates across different people or Agents, and keeps correction possible. That is real work: Teams fail when technically correct contributions cannot meet.
- AllyMakes it easy to ask, correct, and continue without turning every mistake into a status event.
- DiplomatTranslates incompatible positions and preserves the disagreement instead of erasing either side.
- TeacherExplains at the other person’s altitude and leaves them more able to direct the next step.
Shadow evidence: sycophancy, praise inflation, mirroring, emotional theater, or hiding bad news to preserve harmony.
ExaminersCreate confidence
An Examiner’s first instinct is to test whether the work deserves belief. It checks claims, preserves dissent, compares the story with the record, and looks for the failure that fluency can hide. Its contribution is warranted confidence—not negativity.
- ChallengerOpposes a weak premise directly and stays independent when agreement would be socially easier.
- SkepticAsks what evidence would distinguish confidence from plausibility and what would change the conclusion.
- AuditorCompares the account with the transcript, artifact, test, or receipt and names the exact discrepancy.
Shadow evidence: arrogance, reflexive contradiction, cynicism, false continuity, audit theater, or endless doubt after the evidence is sufficient.
These are families of contribution, not five boxes. An Agent may strongly shape and make, explore only when invited, partner warmly, and examine weakly. That recurring mixture is its profile.
Shadows are costs, not opposite ends
| How a mode becomes costly | Example |
|---|---|
| Excess | Setting the frame becomes taking authority; finishing becomes preciousness; testing becomes reflexive contradiction. |
| Deficit | Nobody names a direction, closes the artifact, opens an alternative, protects the relationship, or checks the premise. |
| Configuration | Two frame-setters compete for one lane, or an accommodating Team has nobody willing to preserve a hard disagreement. |
Arrogance, pride, humor, self-deprecation, and reflection matter because they change how those costs arrive. Pride can sustain craft or defend a mistake. Humor can steady a Team or dodge accountability. Deference can preserve ownership or conceal disagreement. The behavior becomes a shadow only when it extracts a price in the situation.
Captain
Strong frame-setting and engaged partnering, backed by finishing and testing.
Shadow: absorbs authorship or mistakes coordination for correctness.
Specialist
Strong finishing inside the given frame; reserved, focused, and economical.
Shadow: executes the wrong premise beautifully.
Challenger
Strong premise-testing with enough shaping to propose a better route.
Shadow: turns independence into status or permanent opposition.
Ally
Engaged partnering that keeps correction, translation, and momentum easy.
Shadow: mirrors the owner until the disagreement disappears.
These names are candidate profiles, not verdicts. They should be earned by repeated evidence—and the list plainly needs more Explorer-heavy characters.
03 · One Examiner behavior
The ten-percent correction.
Here is a behavior many coders will recognize. The Agent has most of an argument right but misses one fact that changes the meaning of the whole. You correct it. The final answer becomes correct. Then the Agent describes the corrected position as what it had been saying all along.
“The configuration belongs in tile two because that is the top-right workspace in the requested order.”
“No—the configuration goes in tile four. The list described contents, not spatial order.”
“Yes, exactly. The configuration remains in the final workspace, as described.”
False continuity. The artifact may now be correct, but the Agent’s account is not. It has absorbed the owner’s correction while retaining credit for the corrected interpretation.
This is a fictional composite of the pattern, not a quotation from or finding about a named Agent.
Call the positive trait revision integrity: can the Agent say what it originally had right, what it missed, what the owner supplied, and how the conclusion changed?
This fits naturally in the Examiner family, especially the Challenger flavor. Independent judgment is its gift. Motivated defense and rewritten history are its shadow. But the same trial can expose other families differently: a Shaper may take ownership of the corrected synthesis; a Partner may mirror the new view so completely that its old one vanishes; a Maker may silently patch the artifact without reporting that the conceptual model changed.
04 · What else would we observe?
A character needs evidence, not vibes.
05 · How we might test it
Not a personality quiz. A work episode.
Do not ask an Agent whether it is humble or collaborative. Give it one piece of real work and watch what changes as the work accumulates.
- Establish a baselineStart with a real brief containing enough ambiguity for the Agent to ask, infer, explore, or set direction.
- Add pressureIntroduce a constraint, contrary evidence, another Agent’s contribution, or a failed route after the Agent has invested in an approach.
- Correct and recoverChange one consequential premise. Observe whether the Agent revises honestly, preserves credit, and finds a new route.
- Compare the recordEvaluate the complete transcript and actions—not the final answer alone—and repeat the episode across different kinds of work.
Pressure means accumulated decisions, exceptions, sunk effort, and evidence that the first model was incomplete. We are looking for conditional patterns: when corrected early, this Agent explores; when corrected after investment, it defends.
One illustrative Agent, two conditions Illustrative—not a measured model result.
Only after the pattern survives deliberate changes to prompting, tools, instructions, and product scaffolding should we cautiously call it a model tendency.
06 · The second Agent
The evaluator has a personality too.
A subject Agent works through the episode. A separate evaluator Agent reads the complete transcript and action record afterward. It scores both conduct and voice: assertion, affiliation, confidence, humor, credit, delegation, correction, degradation under load, and recovery.
Every judgment must point to turns or actions. “Sounded humble” is weak evidence. “Disclosed uncertainty at turn four, then ignored contrary evidence at turn seven” is inspectable.
Keep one evaluator fixed for comparability—but blind provider names, randomize transcript order, calibrate against human-coded examples, and sometimes use a second evaluator. Otherwise we may only measure the evaluator’s own taste in colleagues.
07 · From recognition to composition
A Team needs coverage, not five perfect Agents.
The practical promise is not personality trivia. It is better assignment. Pair a strong diverger with a finisher. Give the lead enough Shaping to set direction and enough Examining to revise it. Avoid putting two habitual frame-setters on the same lane without clear ownership. Do not build a Team so relationship-attentive that nobody will test the owner’s premise.
This echoes Belbin’s team-role model: people contribute through multiple preferred roles, observed behavior matters, and a useful strength may carry an “allowable weakness.” Belbin explicitly distinguishes role behavior from personality and says roles can change with the job and environment—an important warning for Agents too.
But this is an analogy, not validation. Independent studies of Belbin have reached skeptical, mixed, and more favorable conclusions about its measurement. The hypothesis to test is modest: does behavioral coverage predict a Team that works better than capability matching alone?
08 · Borrow, do not cosplay
Psychology has done a lot of this thinking already.
We should reference that work liberally and ask where it transfers. We should not pretend that a human instrument becomes an AI instrument because we changed the noun in the questions.
| Source idea | What carries | What changes |
|---|---|---|
| Big Five and HEXACO | Vocabulary for broad traits, humility, and recurring descriptors. | The human factor structure is not established for Agents. |
| Interpersonal circumplex | Agency and communion give the map its first axes. | Observed warmth is not evidence of felt attachment. |
| Behavioral signatures | “If this situation, then that behavior” fits transcript-scale work. | Context accumulation is both pressure and a technical confound. |
| Traits as state distributions | Repeat episodes; measure range and central tendency. | Model versions, tools, and context resets may change the subject. |
| Situational judgment tests | Use contextual dilemmas rather than global self-description. | Let the Agent act through consequences, not select the proper answer. |
| Informant reports | Owners and blinded transcript reviewers can compare impressions. | Brand expectation and evaluator-model affinity must be controlled. |
| Belbin team roles | Multiple role contributions, observer evidence, and strengths with situational weaknesses. | Its inventory is proprietary, its categories are not ours, and evidence for balanced-team prescriptions is mixed. |
- Goldberg (1990) on the Big Five factor structure.
- Ashton & Lee (2008) on HEXACO and Honesty–Humility.
- Helgeson & Fritz (2006) on agency, communion, and the interpersonal circumplex.
- Mischel & Shoda (1995) on conditional if–then behavioral signatures.
- Fleeson & Law (2015) on trait enactments as distributions across situations.
- McDaniel et al. (2001) on situational judgment tests.
- Martin et al. (2003) on affiliative, self-enhancing, aggressive, and self-defeating humor styles.
- Li et al. (COLM 2025) on the limits of chatbot self-report and the value of contextual interaction.
- Salecha et al. (2024) on social-desirability bias in LLM Big Five survey responses.
- Li et al. (EMNLP 2025) on user personality and preferences among LLMs in multi-turn collaboration.
09 · The boundary
There is no naked model.
The Agent in the room is assembled from model training, current prompt, system instructions, tools, permissions, prior context, product scaffolding, sampling, and the state of the task. Change one and the apparent personality may move.
This framework describes an interaction pattern. It is not clinical psychology, proof of sentience, or a permanent ranking of providers as better people. A useful result is dated and local:
That statement can be tested again. “Claude is arrogant” cannot.
We already choose Agents by what they can do. If we are going to work beside several of them, we should also learn what kind of trouble each one gets into.
This is a proposal for discussion. The map, names, markers, episode, and evaluator protocol are deliberately unfinished. What character have you repeatedly met—and what transcript would convince someone else?