Three moving parts. Thaler's 1999 review organizes the subject around three components, and they are worth separating because each produces a different kind of mistake.
The first is how outcomes are perceived and evaluated. A gain or a loss is never scored in the abstract; it is scored against a reference point and inside a particular account. This is where the asymmetry between how a loss and an equivalent gain feel does its work, and it is why the same $500 can register as a windfall, a shortfall or nothing at all depending on what it is being compared against.
The second is categorization: the assignment of spending to accounts. Heath and Soll's account of it in 1996 describes expenses as first being booked, meaning noticed at all, and then posted to a category by similarity. Once a category has a budget, spending in it competes with itself rather than with everything else, so a large dinner out can make a theatre ticket feel unaffordable while an equally large car repair does not.
The third is grouping and frequency: how narrowly or broadly choices are bracketed, and how often the books are balanced. An account reckoned daily behaves differently from one reckoned annually, and a decision considered on its own behaves differently from the same decision considered as one of twenty. Thaler's own review notes that people can define their accounts narrowly or broadly and balance them daily, monthly or yearly, and that the choice is largely unexamined.
Transaction utility, which is the part with the most everyday consequences. Thaler's 1985 framework separates the value of what you get from the pleasure or pain of the deal itself. Acquisition utility is whether the thing is worth the money; transaction utility is whether the price feels like a good price, judged against what you expected to pay. The two come apart constantly, and that is why a bargain on something you did not need still feels like a win, and why paying an ordinary price for something you badly wanted can feel like a defeat. The classic demonstration is the same drink priced differently by the seller: people report a higher acceptable price for a soda bought at a resort hotel than for the identical soda from a run-down grocery store, even though what they consume is identical. That study replicated in 2025, with a modest effect.
⚠️ The evidence base needs stating carefully, because this is a field where the most quotable examples are not always the best-supported ones. Li and Feldman published a registered-report replication of the problems reviewed in Thaler's 1999 paper in Royal Society Open Science in 2025. Testing 17 classic problems with roughly 500 participants each, they concluded, in their own words, "a mostly successful replication: out of the 17 problems, we found empirical support for 11, mixed empirical support for three and no empirical support for three." So the construct stands, and some of the demonstrations that made it famous do not.
One failure is worth naming, because it is the illustration most likely to be reached for. The scenario in which people will spend twenty minutes to save $5 on a $15 item but not to save the same $5 on a $125 item, originally from Tversky and Kahneman in 1981, is Problem 2 in Thaler's review, and the replication reported that it "failed to find support" for the original finding. The reason is instructive rather than damning: hardly anyone in the modern sample was willing to spend twenty minutes to save $5 in either condition, which the authors attribute to inflation since the 1980s. The effect may well be real and the vignette simply worn out. Either way, quoting it as an established result is no longer safe. The replication also found the opposite of the original result on part of at least one problem, which is a reminder that effect sizes and even directions in this literature are contingent on stakes, framing and population.
🔑 The two-sidedness is the practically useful part, and it is what distinguishes mental accounting from most named biases. The label attached to a sum of money changes nothing about the arithmetic and a great deal about the behavior, which means it can be pointed in either direction.
Pointed the wrong way, it costs money. A household carrying a balance at a high credit card rate while holding cash in an account labeled for a holiday is paying for the label. A tax refund or a bonus, being coded as a different kind of money from salary, gets spent in ways the same amount of salary would not. An investment held at a loss is kept because selling would post the loss to an account that has so far only recorded a paper figure, which is the mechanism behind the tendency to sell winners and hold losers. A separate account nominally reserved for one purpose is quietly raided for another, and because the accounts were never written down, nobody notices the ledger no longer balances.
Pointed the right way, it is one of the more reliable behavioral tools available. Named savings accounts work because the label makes withdrawing money feel like taking it from a specific future purpose rather than from an undifferentiated pile. Automatic transfers work because the money is posted to its account before it can be booked as spendable. Envelope systems and category budgets work for the same reason. None of these does anything a spreadsheet could not do; what they do is make the irrational bit of the machinery pull in a chosen direction. Goals-based planning is this idea applied deliberately at the level of a whole financial plan.
Two honest limits belong on the page alongside the uses. Knowing the name of the effect is weak protection against it, which is why the fixes that work are changes to the environment rather than resolutions about willpower. And strict separation has a real cost: holding cash in one labeled account while carrying debt in another is inefficient, and the discipline the labels buy has to be worth more than the interest the separation loses.