The Quad Model

Most of what we know about people's attitudes comes from asking them. That works until it doesn't. Sometimes people are unwilling to say what they think, particularly about socially sensitive topics. Sometimes they are willing but unable, because the attitude isn't fully available to introspection. And sometimes they are simply not motivated to dig for an honest answer. Indirect methods were built to work around this: rather than asking people to report an attitude, they infer it from how people perform a task.

The Implicit Association Test

The most widely used indirect method is the Implicit Association Test (IAT). It is a sorting task: people see items representing two categories — faces, words, or names — one at a time and sort them using one of two keys. The catch is in how the keys are assigned: in one arrangement of the common race-based version, White faces and good words share a key; in another, Black faces and good words share it.

The task is meant to capture the difference in performance between these arrangements, on the assumption that a stronger association will show up as better performance under one arrangement than the other. Better performance when Black faces and good words share a key, for instance, is taken to suggest an association that favors responding to Black targets more positively; better performance when White faces and good words share a key suggests the reverse.

The Implicit Association Test: the same stimuli under two key pairings Two task screens side by side. In both, the good and bad attribute labels stay on the same sides of the screen while the White and Black category labels swap. Below each screen, a bar shows typical response time: faster under the first pairing, slower under the second. The standardized difference between them is the IAT D-score. FIRST PAIRING Good White Bad Black face or word appears here press E press I SECOND PAIRING Good Black Bad White face or word appears here press E press I typical response time faster typical response time slower Good and bad stay put. The group labels swap. The standardized difference in response times between the two pairings is the IAT D-score — one number per person.
The IAT compares performance across two key pairings. The standardized difference between them is the D-score.

The difference in performance between the pairings is typically computed as the D-score: a single numeric summary of how much faster someone responded under one pairing than the other. Faster responding is taken to reflect a stronger preferential association toward one group, treating the difference as a measure of bias itself. Most of the literature on implicit intergroup bias relies on this metric.

What one number leaves out

The D-score is a singular, relative index. It suggests a person favored one group over the other, and by how much. That is genuinely useful, but it is also less information than what is offered by IAT data. Moreover, very different psychological processes may produce the same D-score.

Consider two ways a preference for one group over another could emerge. It could rest on more positive evaluations of the favored group or more negative evaluations of the other. These are not the same explanation, and a person could arrive at an identical score either way.

Two ways to arrive at the same implicit preference Two stacked bars of equal total length. In the first, negative evaluations of the group not favored make up most of the length. In the second, positive evaluations of the favored group do. Both correspond to the same observed implicit preference. Positive evaluations of the favored group Negative evaluations of the other group same implicit preference Mostly negativity Performance driven by negative evaluations of one group more than positive evaluations of the other. Mostly positivity Performance driven by positive evaluations of one group more than negative evaluations of the other. Both evaluations shape IAT performance. The D-score cannot say which one is doing the most work.
Illustrative. Two compositions a single score cannot distinguish.

Now consider a second ambiguity. Evaluative associations (e.g., "Black = good") do not always influence behavior whenever they are activated. Regulatory processes such as suppression and inhibition can gate the expression of an evaluation before it reaches behavior. With that in mind, the same intergroup bias can arise from two very different psychological stories: the evaluation was rarely activated but also rarely regulated, or it was often activated but often regulated. The outcome is the same in both. What differs is how much the attitude itself is driving that outcome, making it risky to assume a one-to-one relationship between a single metric like the D-score and evaluations.

Two ways to arrive at the same weak or absent bias Two bars showing an evaluation, each split into the portion held back by regulation and the portion expressed in behavior. In the first, almost nothing is activated. In the second, a great deal is activated but nearly all of it is held back. The portion expressed is the same in both. Expressed in behavior Held back by regulation Bar length = evaluations activated expressed in behavior Rarely activated, rarely regulated The evaluation is seldom triggered, so there is little for regulation to act on. Often activated, often regulated The evaluation is triggered readily, but regulation keeps most of it from being expressed. Absent and suppressed look identical from the outside.
Illustrative, and simplified — detection and guessing also shape responses. Absence and suppression are not distinguishable from a single score.

Neither ambiguity is a flaw in the IAT per se. It is a limitation in relying on a single metric for a process-level explanation of IAT performance, with two common misinterpretations. First, the discrimination and prejudice literatures routinely describe "anti-Black" or "anti-gay" bias for D-scores that may be as much or more about "pro-White" and "pro-straight" evaluations. Second, nonattitudinal processes are often ignored in explanations of D-score variance. Performance on any task is not process-pure; it is the joint product of several processes, and the D-score alone cannot tell you which ones produced it. But there are tools that offer a solution. Below we discuss one.

Introducing the Quad model

The quad model is a multinomial processing tree (MPT) model: a formal theory about the processing pathways that lead to particular responses on the IAT. On any given trial of the IAT, several things might happen. An evaluative association might be activated or not. The correct answer might be detected or not. If the two point to different response options, their conflict must be resolved so that one wins out over the other.

These possibilities — call them psychological states — combine into a set of paths, each running from the stimulus to one of the two outcomes the task records: a correct response or an incorrect one. Take a trial where an activated association points toward the wrong key. A correct response can still come about three ways. The association was never activated, and the correct key was detected. Or it was activated, but was overcome in favor of the detected correct response. Or neither happened, and a response bias toward the key that happens to be correct on this trial was engaged. Each path has a probability, and the probability of responding correctly on that trial is their sum.

Run that for every trial type and you get a system of equations linking a handful of unknown process probabilities to something you can actually count: how often people were right and wrong on each kind of trial. Solve the system, and the counts you observe become estimates of the processes you don't observe.

The four parameters

The conditional relationships between these parameters matters. OB is only estimable on trials where AC and D pull in opposite directions, and G only comes into play when the first two fail to influence responses. That nesting is exactly what the tree diagram shows, and it is why the model can separate processes that a single score folds together. Additionally, the psychological interpretation of these parameters is not assumed. Each has been validated against independent evidence.

Work through the model

The tool below lets you build the equations yourself. Because it is the most familiar case, the tool works through the Race IAT (Black/White, good/bad). The same structure applies to any evaluative IAT. Pick a trial type and the diagram highlights the paths available on that trial, with the corresponding equations beside it. Two things are worth watching for as you move through the eight trial types: which trials have an OB branch at all (only the ones where the association and the correct answer conflict), and how the trials that lack one become diagnostic of detection and guessing instead.

Trial Type

White + good share a response key

White + bad share a response key

ACwgDOB(1 − OB)(1 − D)(1 − ACwg)D(1 − D)G(1 − G)trialnot in this trialnot in this trial✓ correct✓ correct✓ correct✓ correct✗ incorrect

White target — White + good share a response key

Response category 1 · White-good / Black-bad specification · no OB step

Assumes White targets are evaluated positively and Black targets negatively. AC estimates activation of the White-good and Black-bad associations; OB estimates control over those associations.

P(correct)ACwg + (1 − ACwg) × D + (1 − ACwg) × (1 − D) × G
P(incorrect)(1 − ACwg) × (1 − D) × (1 − G)

The activated association points at the correct key here, so this trial is compatible under this specification. With no conflict to resolve, it is diagnostic mainly of D — detection and G — guessing.

The equations above write ACwg as a single term. That is the algebraic collapse of ACwg × D and ACwg × (1 − D): on this trial type detection does not change the response, so D factors out. The diagram shows both branches, converging on the same outcome.

Model specification

The model has to assume a direction: which group is being evaluated positively and which negatively. Toggle between the two specifications above to see how the same eight trial types map onto different underlying processes depending on which assumption is made. That assumption turns out to be far more useful than it first appears, as we’ll see in the next section.

An example of the Quad model's impact: understanding the absence of ingroup preference

Both ambiguities above are live questions in one of the more striking findings in the study of intergroup behavior.

Ingroup preference is among the most reliably documented phenomenon in the psychological sciences. People favor their own groups in cooperation, in trust, in helping, in memory, in moral judgment — even when group membership is assigned arbitrarily. And yet for some, it is relatively absent. Beginning with Kenneth and Mamie Clark's doll studies in the 1940s, and repeatedly since, members of lower-status groups have shown attenuated ingroup preference, and sometimes outright preference for the higher-status outgroup. The pattern appears far more consistently on implicit methods like the IAT.

The standard explanation is internalization: members of lower-status groups are thought to absorb the negative stereotypes their society propagates about their ingroup. It is a plausible account, and for decades it went largely untested at the process-level, because a D-score cannot test it.

First question: what is the preference made of?

If an implicit outgroup preference reflects internalized ingroup negativity, then negative evaluations of the ingroup should be doing most of the work. If instead it reflects positive evaluations of the higher-status outgroup, that is a different psychological story altogether.

Splitting the AC parameter separately for positive-outgroup and negative-ingroup evaluation makes the two accounts separable. I joined my colleagues in doing just that, this across four social domains and over 800,000 online respondents showing implicit outgroup preference.

What outgroup bias is made of: positive outgroup versus negative ingroup evaluation Difference between the positive-outgroup and negative-ingroup association parameters for respondents showing implicit outgroup bias. Lower-status groups fall consistently on the positive-outgroup side. Higher-status groups are mixed. LOWER-STATUS RESPONDENTS Gay men Lesbian women Black Older Asian (n.s.) HIGHER-STATUS RESPONDENTS Straight women Straight men Younger White (vs. Asian) White (vs. Black) −0.04 −0.02 0 0.02 0.04 0.06 ← driven more by negative evaluations of the ingroup driven more by positive → evaluations of the outgroup Preferring the outgroup is not the same thing as disliking your own group.
Adapted from Calanchini et al. (2022), confirmatory sample. Values are the difference between the positive-outgroup and negative-ingroup association parameters. The sexuality IAT was run in three stimulus versions; the version with both gay and lesbian stimuli is shown.

Outgroup bias consistently reflected positive outgroup evaluations more than negative ingroup evaluations. Higher-status groups showed no such consistency. Ingroup negativity was present, it was simply not the dominant ingredient. Preferring the higher-status outgroup, is not the same thing as disliking your own group.

Second question: is it only about associations?

Decomposing the associations still leaves the second ambiguity open. Everything above concerns what people have learned. But an absence of ingroup preference could also mean that ingroup-affirming evaluations are present and actively unexpressed.

To tackle this question, my colleagues and I relied on two specifications of the quad model, as seen in the tool above. We fit each respondent's IAT data to a version of the quad model that assumed the evaluative associations were ingroup-affirming (i.e., positive evaluations of the ingroup, negative evaluations of the outgroup), and again to a version that assumed the evaluative associations were outgroup-affirming (i.e., negative evaluations of the ingroup, positive evaluations of the outgroup). The OB parameters in these models therefore assume the overcoming of ingroup-affirming and outgroup-affirming evaluations, respectively.

We fit both specifications to each of 5.7 million respondents' data, across sixteen years and six attitude domains, keeping whichever one fit better for each person, then compared OB across the two.

Asymmetry in self-regulatory control, by respondent status and attitude domain For each of six attitude domains, the difference in the Overcoming Bias parameter between the ingroup-affirming and outgroup-affirming model specifications. Among lower-status respondents the difference is positive and sizable in five of six domains. Among higher-status respondents it is small and inconsistent in direction. Lower-status respondents Higher-status respondents Sexuality Race Skin tone Weight Age Disability −0.4 −0.2 0 +0.2 +0.4 +0.6 +0.8 ← outgroup-affirming evaluations regulated more often ingroup-affirming evaluations → regulated more often Ingroup-affirming evaluations were disproportionately regulated among lower-status groups.
Adapted from Klein et al. (under review). Values are the difference in the Overcoming Bias parameter between the ingroup-affirming and outgroup-affirming specifications. The lower-status difference was not statistically significant in the disability domain, where estimates were near zero overall.

Among lower-status respondents, ingroup-affirming evaluations were overcome more often than outgroup-affirming ones in five of six domains, and often by a wide margin. Among higher-status respondents, no consistent direction emerged at all.

And it appears even in domains where lower-status respondents still show ingroup preference on average, only attenuated, meaning the associations alone do not account for the phenomenon.

The general point

Two findings, one phenomenon, and neither one visible in a D-score. The first says the absence of ingroup preference is less about ingroup negativity than the field assumed. The second says it is not only about evaluations at all, but also about the regulatory processes governing their expression. Understanding how social hierarchies persist may require looking past the content of our biases to the machinery that governs whether they are expressed.

Apply the quad model yourself

If you'd like to figure out how to implement the quad model for your own data, I suggest you start by walking through the template I've put together. If you click the button below, you'll be taken to an OSF page with template code, sample data, and the standard quad model equations file. The template code is commented to walk you through the analysis for a 2-group between-participants estimation and analysis of quad model parameters.

Quad modeling template

I haven't added the model specification comparison to the template yet, so please see my recent papers (and their respective OSF pages) for insights on implementing that procedure.

Contact me with any questions that cannot be answered with the material I've put together already. I hope this has been helpful!

References

Thumbnail for Measuring and Modeling Implicit Cognition
Measuring and Modeling Implicit Cognition Samuel A. W. Klein & Jeffrey W. Sherman In J. Robert Thompson (Ed.), The Routledge Handbook of Philosophy and Implicit Cognition (pp. 44–55). Routledge, 2022
Thumbnail for The contributions of positive outgroup and negative ingroup evaluation to implicit bias favoring outgroups
The contributions of positive outgroup and negative ingroup evaluation to implicit bias favoring outgroups Jimmy Calanchini, Kathleen Schmidt, Jeffrey W. Sherman, & Samuel A. W. Klein Proceedings of the National Academy of Sciences, 2022
Thumbnail for A Role for Self-Regulatory Control in the Absence of Ingroup Preference
A Role for Self-Regulatory Control in the Absence of Ingroup Preference Samuel A. W. Klein, Tessa E. S. Charlesworth, Daniel W. Heck, Mahzarin R. Banaji, & Jeffrey W. Sherman Proceedings of the National Academy of Sciences, 2026