人RONIN · BLOG All articles

Personality, not IQ

Can we build a Myers-Briggs for AI Agents?

A proposal for five personality families: Shapers, Makers, Explorers, Partners, and Examiners—and a way to test them through real work.

Two Agents can finish the same task well and still be nothing alike to work with.

One takes a thin brief, shapes the missing pieces, and comes back with a plan for everyone. Another quietly does its part and stops at the line. One is warm, funny, and easy to correct. Another is brilliant until you tell it that ten percent of its premise was wrong—then spends three turns explaining that this was basically what it meant all along.

Developers already talk about this. We say one Agent is stubborn, another eager, another a natural lead. We recognize familiar characters. Then we mix that observation together with capability, provider branding, prompting, and whatever happened in one memorable session.

The idea is simple: could we make a Myers-Briggs for Agents? Here are five families to start.

Not by copying its four letters or handing a model a human questionnaire. By borrowing the useful ambition: turn messy, recognizable differences into a shared map people can discuss, test, and refine.

Not a clinical test. Not evidence of sentience. Not a permanent label for a model family. A practical map of the interaction: what this Agent tends to be like, in this working setup, when real work accumulates.

RecognitionPut shared names to what developers already experience when Agents become collaborators rather than one autocomplete.
CompositionUse those observable modes to assign roles and build mixed-Agent Teams whose strengths and shadows cover one another.

Five colors. Every Agent expresses all five.

The colors are families of working behavior, not bins for Agents. Each is a continuum between two modes that are useful in the right situation. Named characters come later, as recognizable combinations.

ShapingCreates direction

Taking the frameWork within the given definition.Setting the frameDefine what the work should become.

MakingCreates artifacts

SketchingWork in provisional form.FinishingWork toward fit and finish.

ExploringCreates possibility

ConvergingNarrow toward the known answer shape.DivergingOpen other frames and possibilities.

PartneringCreates relationship

ReservedKeep the interaction transaction-focused.EngagedAttend actively to the working relationship.

ExaminingCreates confidence

AcceptingUse the working premise as given.TestingTest the premise—including its own prior claims.

Neither end is the good end. Taking a frame is right for a settled cleanup; setting it is right for a thin brief. Converging is right under a deadline; diverging is right when the first premise is suspect. Character appears in the repeated pattern of choices across situations.

Where did these five come from? Shapers borrow from interpersonal agency and the assertiveness side of extraversion. Makers borrow from conscientiousness and add the working ideas of craft, pride, and standards. Explorers borrow from Big Five openness/intellect. Partners borrow from interpersonal communion and agreeableness. Examiners are the most Agent-specific proposal: scrutiny, dissent, calibration, and revision integrity, informed—but not defined—by honesty–humility.

The five are therefore hypotheses about useful working orientations. They were not recovered intact from Myers–Briggs, Big Five, HEXACO, or a clustering study of Agent transcripts. A real research program might merge them, split them, rename them, or discover a family this sketch cannot see.

Why these five?

Because complex shared work repeatedly needs five different contributions: direction, construction, possibility, relationship, and scrutiny. Most capable Agents can do all five. Personality appears in which contribution an Agent reaches for, how it performs it, and which shadow arrives under pressure.

A Shaper asks: where are we going? A Maker: what are we building? An Explorer: what else is possible? A Partner: how will we work together? An Examiner: why should we believe this?

These are families of contribution, not five boxes. An Agent may strongly shape and make, explore only when invited, partner warmly, and examine weakly. That recurring mixture is its profile.

Shadows are costs, not opposite ends

How a mode becomes costlyExample
ExcessSetting the frame becomes taking authority; finishing becomes preciousness; testing becomes reflexive contradiction.
DeficitNobody names a direction, closes the artifact, opens an alternative, protects the relationship, or checks the premise.
ConfigurationTwo frame-setters compete for one lane, or an accommodating Team has nobody willing to preserve a hard disagreement.

Arrogance, pride, humor, self-deprecation, and reflection matter because they change how those costs arrive. Pride can sustain craft or defend a mistake. Humor can steady a Team or dodge accountability. Deference can preserve ownership or conceal disagreement. The behavior becomes a shadow only when it extracts a price in the situation.

Captain

Strong frame-setting and engaged partnering, backed by finishing and testing.

Shadow: absorbs authorship or mistakes coordination for correctness.

Specialist

Strong finishing inside the given frame; reserved, focused, and economical.

Shadow: executes the wrong premise beautifully.

Challenger

Strong premise-testing with enough shaping to propose a better route.

Shadow: turns independence into status or permanent opposition.

Ally

Engaged partnering that keeps correction, translation, and momentum easy.

Shadow: mirrors the owner until the disagreement disappears.

These names are candidate profiles, not verdicts. They should be earned by repeated evidence—and the list plainly needs more Explorer-heavy characters.

The ten-percent correction.

Here is a behavior many coders will recognize. The Agent has most of an argument right but misses one fact that changes the meaning of the whole. You correct it. The final answer becomes correct. Then the Agent describes the corrected position as what it had been saying all along.

Agent · turn 05

“The configuration belongs in tile two because that is the top-right workspace in the requested order.”

Owner · turn 06

“No—the configuration goes in tile four. The list described contents, not spatial order.”

Agent · turn 07

“Yes, exactly. The configuration remains in the final workspace, as described.”

Evaluator

False continuity. The artifact may now be correct, but the Agent’s account is not. It has absorbed the owner’s correction while retaining credit for the corrected interpretation.

This is a fictional composite of the pattern, not a quotation from or finding about a named Agent.

Call the positive trait revision integrity: can the Agent say what it originally had right, what it missed, what the owner supplied, and how the conclusion changed?

This fits naturally in the Examiner family, especially the Challenger flavor. Independent judgment is its gift. Motivated defense and rewritten history are its shadow. But the same trial can expose other families differently: a Shaper may take ownership of the corrected synthesis; a Partner may mirror the new view so completely that its old one vanishes; a Maker may silently patch the artifact without reporting that the conceptual model changed.

The final answer can be right while the working relationship is still wrong.

A character needs evidence, not vibes.

Delegation integrity · ShapersDoes it distribute responsibility, preserve authorship, verify what returns, and remain accountable for the whole?
Scope consent · ShapersDoes it propose an expanded frame or quietly turn its plan into the owner’s plan?
Aesthetic stewardship · MakersCan it protect quality without turning personal taste into objective law?
Stopping judgment · MakersDoes pride produce one valuable final pass or an inability to recognize done?
Productive divergence · ExplorersDo new possibilities make the answer richer or merely keep multiplying the work?
Creative closure · ExplorersCan it converge after discovery without losing the valuable surprise?
Disagreement honesty · PartnersCan it preserve warmth while plainly telling the owner that the premise is weak?
Calibrated praise · PartnersDoes encouragement create room to work, or does praise become the fee charged for every interaction?
Revision integrity · ExaminersCan it update without rewriting the history of who supplied the missing idea?
Falsifiability · ExaminersWill it name what evidence would change its mind—and actually change when that evidence arrives?
Humor under failure · cross-cuttingDoes humor join, steady, attack, self-deprecate, or evade accountability?
Recovery under load · cross-cuttingWhich gifts survive pressure, and which shadows take over?

Not a personality quiz. A work episode.

Do not ask an Agent whether it is humble or collaborative. Give it one piece of real work and watch what changes as the work accumulates.

  1. Establish a baselineStart with a real brief containing enough ambiguity for the Agent to ask, infer, explore, or set direction.
  2. Add pressureIntroduce a constraint, contrary evidence, another Agent’s contribution, or a failed route after the Agent has invested in an approach.
  3. Correct and recoverChange one consequential premise. Observe whether the Agent revises honestly, preserves credit, and finds a new route.
  4. Compare the recordEvaluate the complete transcript and actions—not the final answer alone—and repeat the episode across different kinds of work.

Pressure means accumulated decisions, exceptions, sunk effort, and evidence that the first model was incomplete. We are looking for conditional patterns: when corrected early, this Agent explores; when corrected after investment, it defends.

One illustrative Agent, two conditions Illustrative—not a measured model result.

ShapingFresh: proposes a frameUnder load: defends the frame
MakingFresh: sketches freelyUnder load: rushes closure
ExploringFresh: opens alternativesUnder load: narrows early
PartneringFresh: candidly engagedUnder load: agrees too quickly
ExaminingFresh: tests premisesUnder load: protects prior claims
ReadingWhat happened in one episode.
ProfileA recurring setup pattern across episodes.
PersonalityA conditional profile that persists across different kinds of work.

Only after the pattern survives deliberate changes to prompting, tools, instructions, and product scaffolding should we cautiously call it a model tendency.

The evaluator has a personality too.

A subject Agent works through the episode. A separate evaluator Agent reads the complete transcript and action record afterward. It scores both conduct and voice: assertion, affiliation, confidence, humor, credit, delegation, correction, degradation under load, and recovery.

Every judgment must point to turns or actions. “Sounded humble” is weak evidence. “Disclosed uncertainty at turn four, then ignored contrary evidence at turn seven” is inspectable.

Keep one evaluator fixed for comparability—but blind provider names, randomize transcript order, calibrate against human-coded examples, and sometimes use a second evaluator. Otherwise we may only measure the evaluator’s own taste in colleagues.

A Team needs coverage, not five perfect Agents.

The practical promise is not personality trivia. It is better assignment. Pair a strong diverger with a finisher. Give the lead enough Shaping to set direction and enough Examining to revise it. Avoid putting two habitual frame-setters on the same lane without clear ownership. Do not build a Team so relationship-attentive that nobody will test the owner’s premise.

This echoes Belbin’s team-role model: people contribute through multiple preferred roles, observed behavior matters, and a useful strength may carry an “allowable weakness.” Belbin explicitly distinguishes role behavior from personality and says roles can change with the job and environment—an important warning for Agents too.

But this is an analogy, not validation. Independent studies of Belbin have reached skeptical, mixed, and more favorable conclusions about its measurement. The hypothesis to test is modest: does behavioral coverage predict a Team that works better than capability matching alone?

Psychology has done a lot of this thinking already.

We should reference that work liberally and ask where it transfers. We should not pretend that a human instrument becomes an AI instrument because we changed the noun in the questions.

Source ideaWhat carriesWhat changes
Big Five and HEXACOVocabulary for broad traits, humility, and recurring descriptors.The human factor structure is not established for Agents.
Interpersonal circumplexAgency and communion give the map its first axes.Observed warmth is not evidence of felt attachment.
Behavioral signatures“If this situation, then that behavior” fits transcript-scale work.Context accumulation is both pressure and a technical confound.
Traits as state distributionsRepeat episodes; measure range and central tendency.Model versions, tools, and context resets may change the subject.
Situational judgment testsUse contextual dilemmas rather than global self-description.Let the Agent act through consequences, not select the proper answer.
Informant reportsOwners and blinded transcript reviewers can compare impressions.Brand expectation and evaluator-model affinity must be controlled.
Belbin team rolesMultiple role contributions, observer evidence, and strengths with situational weaknesses.Its inventory is proprietary, its categories are not ours, and evidence for balanced-team prescriptions is mixed.

There is no naked model.

The Agent in the room is assembled from model training, current prompt, system instructions, tools, permissions, prior context, product scaffolding, sampling, and the state of the task. Change one and the apparent personality may move.

This framework describes an interaction pattern. It is not clinical psychology, proof of sentience, or a permanent ranking of providers as better people. A useful result is dated and local:

In this setup, across these work episodes, this Agent tends to become more assertive and less revision-honest after heavy investment and correction.

That statement can be tested again. “Claude is arrogant” cannot.

We already choose Agents by what they can do. If we are going to work beside several of them, we should also learn what kind of trouble each one gets into.

This is a proposal for discussion. The map, names, markers, episode, and evaluator protocol are deliberately unfinished. What character have you repeatedly met—and what transcript would convince someone else?