Paper 01 · Experimental economics
Take the most-replicated experiment in the economics of trust. Change exactly one thing: the person deciding whether to trust is spending someone else’s money. The people who trust behave the same. The people being trusted do not — and trust stops paying for itself.
Ola Kvaløy · Miguel Luzuriaga · “Playing the trust game with other people’s money” · Experimental Economics 17 (2014) 615–630 · doi:10.1007/s10683-013-9386-4
Received 23 November 2012 · accepted 7 December 2013 · published 18 December 2013 · University of Stavanger, Norway
A great deal of economic life is conducted by people spending money that is not theirs.
A fund manager commits your capital. A CEO decides, on the board’s behalf, whether to enter a partnership with a rival who could walk off with everything they learn. A section leader hires someone the CEO will never interview. In every one of these, somebody is extending trust — accepting a risk that cannot be enforced by any contract — and the person who bears the consequences is not the person making the call.
Economics has a good empirical handle on trust between two people who each own what they are risking. We know that people trust more than self-interest can justify, and that they are repaid more than self-interest can justify. The engine behind that is reciprocity: people reward kindness.
But reciprocity is a fragile enforcement device, because it responds to intentions, not just to outcomes. And that is exactly what delegation interferes with. If a sender is not risking her own money, how kind is her gesture? And even if she is kind, how would you repay her — when nothing you do can change what she takes home? Kvaløy and Luzuriaga put the question this way:
Does reciprocity still facilitate economic exchange when the people making the transactional decisions are not the people who reap the rewards of trust?
Everything here is built on one experiment: the trust game, also called the investment game, introduced by Berg, Dickhaut and McCabe in 1995. It has two players who never learn each other’s identity and never meet again.
Nothing obliges the receiver to return anything. There is no contract, no court, no reputation, no second round, no way to identify the other person afterwards. So x is a clean measure of trust — money handed to a stranger with no enforceable claim on it — and y is a clean measure of trustworthiness. Nothing else in the situation can produce them.
The sender is spending her own 100. Whatever comes back is hers.
Out of the 100 she owns. The experimenter triples it on the way over: the receiver sees 3x = 180.
Free choice anywhere in [0, 3x]. Nothing enforces a return — no contract, no repeat play, no way to find out who the other person was.
Exactly break-even: the money came back, the gains all stayed with the receiver.
Both bars on the same scale, 0–400.
Write πS for the sender’s final money and πR for the receiver’s. The sender starts with 100, gives away x, and gets y back. The receiver starts with 100, gains 3x, and gives away y:
Add the two payoffs and the transfers cancel:
So the efficient choice is unambiguous: x = 100, which puts 400 on the table instead of 200. Doubling the room’s wealth requires nothing but trust. Whether anyone gets to enjoy it is a separate question.
The sender ends up better off than her outside option of 100 exactly when 100 − x + y > 100, i.e. when y > x. The receiver ends up better off than his 100 exactly when y < 3x. Put the two together:
Subtract one payoff from the other:
Three landmarks, then, and it is worth carrying all three into the results: y = x is the sender’s break-even, y = 2x is the equal split, and y = 3x is the receiver giving away everything he gained. The paper’s headline statistic — the share returned, y / 3x — puts break-even at 1/3 and the equal split at 2/3. Baseline receivers averaged 0.42. OPM receivers averaged 0.31: below break-even.
Before looking at what people did, we need the benchmark they are being compared against. It is obtained by backward induction: solve the last decision first, then work up the tree, assuming each player correctly anticipates what follows.
The unique subgame-perfect Nash equilibrium is (x* = 0, y* = 0), paying (100, 100). And it is the worst available joint outcome: W = 200 against a possible 400. The equilibrium burns 200 kroner that both players would rather have had — a social dilemma in two moves.
The sender moves first and picks any x in [0, 100]. The receiver sees 3x arrive, knows exactly what the sender gave up, and picks any y in [0, 3x]. Then the game ends: one shot, anonymous, no reputation to protect. Three branches are drawn here; there are really 101.
This prediction is wrong, and known to be wrong — Berg, Dickhaut and McCabe found senders sending about half and receivers returning about a third, and thousands of replications since have agreed. The equilibrium is not a forecast. It is a measuring stick: everything above zero is social preference, and the whole literature is an argument about what that surplus is made of.
Now the treatment. Kvaløy and Luzuriaga ran two conditions with 90 subjects each.
A faithful replication of Berg et al. Two players, 100 each, sender sends x from her own endowment, receiver returns y. Payoffs: 100 − x + y and 100 + 3x − y.
A third person enters: the client. The client gives the sender 100. The sender keeps her own 100 regardless, and decides how much of the client’s money to send. Whatever the receiver returns goes to the client, not to her.
Three features of this design carry the whole paper, and each is deliberate:
Notice what has not changed. The receiver’s choice set is identical, his payoff function is identical, and the amount in front of him is identical. If a theory says the receiver optimises over the distribution of money, it cannot possibly predict a difference between these two treatments. That constraint is what makes the experiment a genuine test rather than a demonstration.
The result is much harder to forget once you have sat in the sender’s chair. Send the same amount a few times in each treatment.
Whatever comes back is yours. You are risking your own money.
Nothing yet. Try the same x in both treatments a few times — the gap is a distribution, not a single number, and one round will not show it to you.
45 sender–receiver pairs per treatment, the same number Kvaløy and Luzuriaga ran. Senders draw from the reported mean and standard deviation of x; receivers use the reported share for their quintile. Run it a few times and watch how unstable a 45-pair mean is — this is the honest way to feel what a p-value of 0.04 is claiming.
People plainly do not play the selfish equilibrium, so we need a model in which they care about something beyond their own money. There are two families, and they disagree about what that something is.
Inequity aversion(Fehr & Schmidt 1999; Bolton & Ockenfels 2000). People dislike unequal payoffs — more when they are behind, but also when they are ahead. For two players:
Take a receiver holding 100 + 3x. Above we found πR − πS = 4x − 2y, so he is ahead whenever y < 2x. In that region only the guilt term is active:
So the model gives a crisp answer. A receiver with β < ½ returns nothing. A receiver with β > ½ returns exactly y = 2x — the equal split, and not one krone more. Inequity aversion generates reciprocal behaviour without any taste for reciprocity: what looks like gratitude is just discomfort at being ahead.
Run that derivation again for OPM. The receiver’s payoff is the same; the person he is compared with holds 100 − x + y, the same as before. Every symbol is unchanged. α and β describe preferences over distributions of money and contain no slot for who acted or why.
Fehr–Schmidt explains reciprocal behaviour in both treatments and predicts no difference between them. Whatever a given receiver returns in Baseline, the same receiver returns in OPM.
The authors also check the version with three players, since OPM has one more person with money at stake:
The paper’s reading is that this normalisation means the extra player has no behavioural consequence, so Prediction A survives. (There is a wrinkle in that step worth having ready — see §11.)
Reciprocity models(Rabin 1993; Dufwenberg & Kirchsteiger 2004; Falk & Fischbacher 2006). Here people respond to the kindness behind an action, not only its result. Add two ingredients: a kindness term φj, how kind player i judges player j to have been, and a reciprocation term σi, how strongly i converts that judgement into money.
Differentiate in each treatment. In Baseline the only other player is the sender:
In OPM there are two others. The client holds 100 − x + y; the sender holds a constant 100:
Because φS > 0 ≥ φC, the marginal value of returning is strictly lower in OPM. Less money comes back under delegation, and the client’s profit falls below the Baseline sender’s.
Look for the sender in the OPM derivative. She is not there. Her payoff is a constant, so it differentiates away: a receiver who finds her admirable has no channel through which to pay her for it. Delegation does not merely reduce the kindness available to reward — it severs the wire between gratitude and money. That is why the effect lands on receivers rather than senders.
Here theory genuinely cannot call it, and the paper says so. Three forces pull in different directions. A sender who cares about her client and expects a poorer return should send less. Inequity aversion still restrains her, since sending makes the receiver rich. But betrayal aversion(Bohnet & Zeckhauser 2004) — people accept risk more readily from a dice roll than from a person who might betray them — should push the other way: it is not her trust being abused, so she should send more. And the purely selfish sender, indifferent and paid either way, may send generously out of a taste for efficiency.
So: a clear prediction for receivers, no prediction for senders. Worth stating plainly in a presentation, because it is why Result 1 is interesting rather than disappointing.
A kind sender (φS > 0) is worth rewarding; a client who has done nothing (φC ≤ 0) is not. Same person, same α and β, two different answers — this is the paper’s explanation of Result 2.
Fehr–Schmidt’s advantageous-inequality aversion. Must satisfy 0 ≤ β < 1 and β ≤ α.
Disadvantageous-inequality aversion. It only bites past the equal-split point, so it rarely changes the receiver’s answer here.
How strongly this person converts perceived kindness into money. Set it to 0 and you are back to pure Fehr–Schmidt.
She handed over her own money knowing she might get nothing back. Positive.
The client made no decision and was not even in the room. Zero at best; the paper argues it can go negative — not blame, just no earned claim.
Sets the size of the pie under discussion.
This receiver reciprocates a sender who risked her own money and gives nothing to a client who did nothing. That is Result 2 in its sharpest form.
Notice what is missing from the OPM expression: the sender. Her payoff is a flat 100 no matter what the receiver does, so she drops out of the derivative entirely. Even a receiver who finds her admirable has no way to pay her for it. Delegation does not just lower the kindness on offer — it disconnects the only lever a reciprocal person has.
Marginal value of returning one more krone, as sensitivity to kindness σ rises. With φS positive and φC negative, the two lines pull apart in opposite directions — the same trait that makes someone more generous to a sender makes them less generous to a client.
The clients are the delicate part of the design, and worth describing precisely because a good question will land here. They were passive — no decision at any point — and anonymous to senders and receivers alike. They were recruited during the baseline sessions and told only that they could earn additional money through decisions made in another experiment.
Each was given a ticket to keep for a couple of days and then to redeem at an office. The sender in OPM held the duplicate of a specific client’s ticket, and wrote on it the amount earned on that client’s behalf. So the client is real, identifiable to the mechanism, and paid according to what the receiver decides — and yet, to everyone in the room, entirely faceless.
It makes the client’s money credibly real without putting the client in the room. That is the design’s great strength and, as §11 argues, also the source of its most awkward confound.
Comparisons use the Mann–Whitney U-test throughout — a rank-based test that asks whether one sample’s values tend to sit above the other’s. It makes no assumption of normality, which matters here because the data are lumpy, censored at zero, and nothing like a bell curve.
| Baseline mean | Std | OPM mean | Std | z | p | |
|---|---|---|---|---|---|---|
| Money sent | 65.04 | 32.21 | 59.18 | 34.27 | 0.80 | 0.42 |
| Money returned | 78.27 | 70.79 | 50.91 | 59.32 | 2.02 | 0.04 |
| Share returned | 0.42 | 0.28 | 0.31 | 0.31 | 2.35 | 0.02 |
Senders who manage other people’s money do not behave significantly different from senders who manage their own money.
NOK 65.04 vs 59.18 sent; Mann–Whitney z = 0.80, p = 0.42.
Senders sent about 6 kroner less with a client’s money — nowhere near significance. Whatever moral weight people feel when spending someone else’s money, it did not change what they did with it. And note the level: 65 out of 100 is high trust by the standards of this literature.
Receivers return less money when senders send a third party’s money than when senders send their own money.
NOK 78.27 vs 50.91 returned (z = 2.02, p = 0.04); share 0.42 vs 0.31 (z = 2.35, p = 0.02).
A 35% drop in the amount returned, and a fall in the share from 0.42 to 0.31. Recall the landmarks from §3: 1/3 is where the first mover merely breaks even. Baseline receivers cleared it. OPM receivers did not.
The treatment effect is not spread evenly. In Q1–Q3 the bars are close, and Q2 even runs the wrong way. The gap opens in the top two quintiles — precisely where the sender has been most generous, and so precisely where an intention-based theory says the largest reward is owed. The share returned falls from 40% to 23% in Q4 and from 38% to 23% in Q5: almost halved, exactly where reciprocity should have been strongest.
| Mean | Std | Min | Max | vs. | p | |
|---|---|---|---|---|---|---|
| Sender (Baseline) | 113.22 | 62.62 | 0 | 300 | z = 2.22 | 0.03 |
| Client (OPM) | 91.73 | 57.69 | 0 | 300 | ||
| Receiver (Baseline) | 216.87 | 89.25 | 100 | 400 | z = -0.30 | 0.76 |
| Receiver (OPM) | 226.62 | 99.96 | 100 | 400 |
Trust is profitable when playing with one’s own money, but not when playing with other people’s money.
Sender payoff in Baseline 113.22 vs client payoff in OPM 91.73 (z = 2.22, p = 0.03) — below the 100 a client would have kept if the sender had sent nothing.
The average client would have been better off if their agent had sent nothing at all. The client’s 91.73 is below the 100 they started with. Delegated trust did not merely underperform — it destroyed value for the person paying for it, while the receivers’ average payoff was, if anything, slightly higher in OPM (226.62 vs 216.87, p = 0.76).
The authors state clearly that the gender effect was not hypothesised in advance. It is also the most striking thing in the paper.
Cell sizes are uneven and worth knowing: receivers were 23 men / 22 women in Baseline, and 14 men / 31 women in OPM. Fourteen men carry the entire male OPM estimate.
| Men: Baseline vs OPM | z = -0.47 | p = 0.64 |
| Women: Baseline vs OPM | z = 2.99 | p = 0.00 |
| Men vs women, Baseline | z = -1.22 | p = 0.22 |
| Men vs women, OPM | z = 1.88 | p = 0.06 |
Men are flat across treatments (71.96 → 75.71, p = 0.64). Women halve (84.86 → 39.71, p = 0.00). The aggregate treatment effect in Table 1 is, in its entirety, a female effect.
Women return less money when senders send a third party’s money, while men do not exhibit a significant difference.
Women 84.86 → 39.71 (z = 2.99, p = 0.00); men 71.96 → 75.71 (z = −0.47, p = 0.64).
Women return significantly less money than men when senders send a third party’s money, while there is no significant difference when senders send their own money.
OPM shares 0.25 vs 0.44 (z = 2.22, p = 0.03); Baseline shares 0.44 vs 0.41 (z = −0.05, p = 0.96).
This is a genuine reversal, not a shading. The prior literature says women are, if anything, more reciprocal than men in trust games, and more generous in dictator games. Here they are markedly less generous — but only under delegation. The model in §7 accommodates it with one parameter: if women have a higher σ, then a positive φ makes them return more and a negative φ makes them return less. High sensitivity is not the same as high generosity. Croson and Gneezy’s survey of this literature reaches the same conclusion from the other direction: women’s social behaviour is more responsive to context, which is why gender differences in these games look so inconsistent across studies.
Finally the regressions, which do two jobs: check that the treatment effect survives controls, and locate exactly where it lives.
| Regressor | (1) | (2) | (3) |
|---|---|---|---|
| OPM | −33.64** (16.51) | −27.00* (14.56) | 16.71 (28.29) |
| Money Received | — | 0.3138*** (0.0839) | 0.3218*** (0.0813) |
| Female | — | −10.88 (16.30) | 22.63 (20.59) |
| OPM × Female | — | — | −73.35** (32.21) |
| Intercept | 74.25*** (11.73) | 19.01 (14.91) | 0.7972 (18.45) |
| R² (pseudo) | 0.005 | 0.023 | 0.029 |
| F-statistic | 0.045 | 0.000 | 0.000 |
The paper’s pivot. Women in OPM return 73 kroner less than the main effects alone would predict. The aggregate treatment effect in Table 1 is entirely this cell.
The design is built so that outcome-based and intention-based social preferences make different predictions, and the data side with intentions. Inequity aversion is not refuted — it explains a great deal of what receivers do in both treatments — but it cannot explain the difference, and the difference is what the experiment was built to see. Kindness has to be in the utility function.
A run of experiments has shown that delegation blunts negativereciprocity: responders are less willing to reject an unfair offer, or to punish, when a hired agent made the decision (Fershtman & Gneezy 2001; Coffman 2011; Bartling & Fischbacher 2012). In those settings blunted reciprocity is good news for the principal — delegation raises profit by disarming the punisher.
This paper finds the same attenuation in the domain of positive reciprocity, where the sign flips. When the force being blunted is reward rather than punishment, delegation destroys profit instead of protecting it. The general lesson is neither “delegation pays” nor “delegation costs”, but: delegation weakens reciprocity, and whether that helps you depends on which way reciprocity was pointing.
And this design improves on the earlier ones in a specific way. In Fershtman and Gneezy’s setup you cannot tell whether responders spare the agent (the hostage effect) or spare the principal who is only indirectly responsible. Here the sender bears full responsibility and cannot be paid, so the reduced return has to be about the client’s lack of a claim.
Reciprocity is the enforcement mechanism for everything a contract cannot specify — relational contracts, implicit promises, the parts of a partnership no court will police. This experiment says that mechanism is weakened precisely when the decision is delegated, which is to say precisely in the settings where firms operate. Whoever sits across the table from your agent may honour the deal less than they would have honoured it with you, and the whole cost lands on you.
It has been argued that the second stage of a trust game adds nothing to a dictator game: the receiver holds money and decides how much to give away. If that were true the standard dictator finding — women give more — should hold here. Under OPM women gave significantly less. Something in the receiver’s position is doing work that a dictator game cannot capture: he is interpreting how the money arrived, which is information a dictator never has.
Anticipating these is the difference between presenting a paper and defending one. The first three are acknowledged by the authors; the rest are fair game.
The paper wants the client to differ from the Baseline sender in one respect: having taken no action. But the client also differs by being physically absent — the authors note in a footnote that this may itself have made clients seem less entitled. Presence is known to increase giving. So “did nothing” and “was not there” are confounded, and no part of the design separates them. A cleaner version would seat a silent, visible client in the room.
Not hypothesised ex ante, and resting on small, uneven cells: 23 male and 22 female receivers in Baseline, but 14 male and 31 female in OPM. Fourteen people carry the entire claim that men are unaffected. Say “this is a hypothesis for the next experiment” before someone says it for you.
Because utility is linear in y within each region, both theories predict a return of either zero or exactly the equal split. Real receivers are spread smoothly across the whole interval. The model earns its keep by getting the direction of the treatment effect right; it should not be read as a description of any individual.
Worth having in your pocket, since it is the kind of thing a sharp questioner finds. The paper argues that the 1/(n−1) normalisation in equation (2) means the extra player has no behavioural consequence. But work the derivative through for OPM, where the receiver compares himself with both the client and the flat-fee sender:
So on a literal reading, pure Fehr–Schmidt does predict a treatment effect for receivers whose β falls between ½ and 2/3 — and in the same direction the authors found. This does not rescue inequity aversion: it cannot produce the gender crossover, and the size of the effect depends entirely on the arbitrary level of the sender’s fee. But “Fehr–Schmidt predicts nothing here” is a slightly stronger claim than the algebra supports, and it is a good thing to be able to say yourself. This wrinkle is my own reading, not the paper’s.
A shape that works in about ten minutes: (1) the hook — people spend other people’s money constantly; (2) the original trust game and the selfish prediction, so the audience has the benchmark; (3) the one change, with the two payoff blocks side by side so they can see the receiver’s payoff is untouched; (4) the two theories as a genuine fork — one predicts nothing, the other predicts a drop; (5) Results 1–3, ending on the client’s 91.73 against 100; (6) the gender crossover, flagged as exploratory; (7) the takeaway — delegation weakens reciprocity, and the sign of that depends on which way reciprocity was pointing.
If you have one sentence: senders trust the same with other people’s money, but receivers stop honouring it — so delegated trust destroys the very value that trust is supposed to create.
They ran Berg, Dickhaut and McCabe’s 1995 trust game twice — once normally, and once where the person deciding whether to trust was spending an absent third party’s money for a flat fee — and compared the two.
| Endowment / multiplier | NOK 100 each · sent amount tripled |
| Subjects | 180 students, Stavanger · 90 per treatment · 45 pairs each |
| Sent (B / OPM) | 65.04 / 59.18 · p = 0.42 · no effect |
| Returned (B / OPM) | 78.27 / 50.91 · p = 0.04 |
| Share (B / OPM) | 0.42 / 0.31 · p = 0.02 |
| First mover’s payoff | 113.22 (sender, B) vs 91.73 (client, OPM) · p = 0.03 |
| Women returned | 84.86 → 39.71 · p = 0.00 |
| Men returned | 71.96 → 75.71 · p = 0.64 |
| The interaction | OPM × Female = −73.35 (se 32.21) |
| Female slope | 0.50 per krone in Baseline → 0.23 in OPM |
Every figure on this page comes from the paper itself. Tables 1–5 are reproduced as published; the per-quintile values are read from Figs. 1 and 2. The derivations in §3, §4 and §7 are worked out step by step here — the paper states most of them in compressed form — and the wrinkle in §11 is flagged as mine rather than theirs. The simulated receiver in §6 is a teaching device built on the paper’s reported statistics and should not be cited as data.
Ola Kvaløy · Miguel Luzuriaga (2014). “Playing the trust game with other people’s money”. Experimental Economics 17, 615–630. doi:10.1007/s10683-013-9386-4
Berg, J., Dickhaut, J., & McCabe, K. (1995). “Trust, reciprocity, and social history”. Games and Economic Behavior 10, 122–142.
Fehr, E., & Schmidt, K. (1999). “A theory of fairness, competition and cooperation”. QJE 114, 817–868.
Falk, A., & Fischbacher, U. (2006). “A theory of reciprocity”. Games and Economic Behavior 54, 293–315.
Dufwenberg, M., & Kirchsteiger, G. (2004). “A theory of sequential reciprocity”. Games and Economic Behavior 47, 268–298.
Fershtman, C., & Gneezy, U. (2001). “Strategic delegation: an experimental study”. JEBO 45, 371–380.
Bartling, B., & Fischbacher, U. (2012). “Shifting the blame: on delegation and responsibility”. Review of Economic Studies 79(1), 67–87.
Bohnet, I., & Zeckhauser, R. (2004). “Trust, risk and betrayal”. JEBO 55(4), 467–484.
Croson, R., & Gneezy, U. (2009). “Gender differences in preferences”. Journal of Economic Literature 47(2), 448–474.
Study notes by Sazid · The Econ Lab · back to the index