Preferences to Rewards
Building from preferences and a minimal set of axioms to a utility function expressible as a sum of discounted rewards: the familiar framing in reinforcement learning.
By Fernando E. Rosas
- Understand the difference between preferences, utility, and reward: preferences being a primary, largely uncontroversial notion, and utility and rewards being derived notions resting on specific assumptions.
- Be able to derive the relationship between various preference structures and rationality axioms.
- Critically assess alternative notions of rationality, and the consequences of dropping various classical decision theory assumptions.
Overview
This note develops a short route from preferences over complete trajectories to expected utility, reward, and discount. We begin with preference relations on deterministic trajectories and explain how completeness and transitivity yield an ordinal utility representation. We then show how lotteries, together with the von Neumann–Morgenstern axioms, produce a cardinal utility over trajectories, and we clarify the distinction between ordinal preference utility and vNM utility. Next, following Bowling et al., we add a fifth temporal axiom that is necessary and sufficient for a recursive representation in terms of local rewards and discounting. Finally, we explain why reward is not unique: different reward functions can encode the same utility or the same preference ordering, affine changes of utility induce corresponding changes of reward, and potential-based shaping provides a canonical example of reward equivalence.
1. Introduction
What does it mean for a system to have a goal? In sequential decision making, the object of evaluation is often not a single isolated choice or prize but a whole trajectory: a complete history of states, observations, actions, and consequences unfolding through time. Before asking how goals are achieved, it is worth asking a more basic question: what does it mean to prefer some trajectories over others, and what follows from that?
Recent work in reinforcement learning has revived this fundamental question, treating preferences over histories as primitive and asking what additional assumptions are needed before one can recover scalar reward functions. The present note provides an introduction to these ideas, which proceeds in three stages.
- First, one asks for a numerical representation of how whole trajectories are ranked.
- Second, once one allows lotteries over trajectories, one asks when these lotteries can be ranked by the expectation of that trajectory utility.
- Third, one asks what additional requirements are needed in order to decompose this expected utility into rewards assigned at each time step.
The first and second steps are closely related to the classical theory developed by von Neumann and Morgenstern (von Neumann & Morgenstern 1944). The third asks what extra temporal structure is needed before utility over whole trajectories can be decomposed into stagewise rewards, following the line of work developed in modern reinforcement learning by Pitis 2019, Shakerinava & Ravanbakhsh 2022, and Bowling et al. 2023.
2. Preferences over trajectories
Let be a finite set of observations and a finite set of actions. A one-step interaction is then given by . For each , define the space of trajectories of length by
We write for the unique trajectory of length . The space of all finite trajectories is
A typical element of has the form . We will keep the notation for the set of all finite trajectories throughout.
Definition 2.1 (Preference). A preference relation on is a binary relation where
means that trajectory is judged at least as good as trajectory .
From we derive the usual companion relations:
We call `indifference', as an agent has no reason to prefer one over the other.
At this point, has no properties whatsoever. One may naturally wonder what kinds of properties it is reasonable to require of , and what follows from them — which is what we study in the next sections.
3. When are preferences problematic?
Can a preference relation be intrinsically `bad'? The relevant concern here is whether it leads to some form of self-defeat, avoidable loss, or failure of coherent behaviour. The discussion in the decision-theory literature suggests at least three grades of concern.
- Representability failures. A first and weakest concern is that a preference relation may fail to be representable in a convenient form. This does not, by itself, imply that the preference is irrational.
Representation failures may merely be inconvenient, but they become more significant when they are symptoms of deeper issues of the kinds described next (Aumann 1962; Fishburn 1970). 2. Self-defeat and avoidable loss. A more serious concern is that preferences may guide choice poorly — as judged by the agent's own interest. One important case is static self-defeat: choosing an option that is worse than another available one, or adopting a policy that is systematically improvable. A classic case of suboptimality is dominance: one option or policy dominates another when it is at least as good in every relevant respect and strictly better in some, so choosing the dominated option is a clear mistake (Kreps 1988; Mas-Colell et al. 1995). A second case is diachronic self-defeat: a plan that the agent endorses now is predictably undone later in a way that leaves the agent worse off overall. 3. Vulnerability. The most vivid coherence arguments show that a collection of individually acceptable choices can be combined into a guaranteed loss. Dutch-book arguments play this role for credences; money-pump arguments play the analogous role for preferences. The standard example is a preference cycle
If the agent is willing to pay a small fee to move each time to a strictly preferred trajectory, then an adversary can guide it around the cycle and back to where it started, poorer than before (Gustafsson 2010). This is why intransitivity is usually regarded as a particularly severe pathology: it is not merely hard to represent, but also vulnerable to exploitation under natural trading assumptions.
Coherence is not selection.
It is useful to distinguish three kinds of formal results. A representation result says that if preferences satisfy certain axioms, then they can be written in a particular mathematical form; the vNM and Savage theorems being classical examples. A coherence result is different: it links violations of some constraint to a penalty such as a Dutch book, money pump, dynamic inconsistency, or dominated choice. Complete-class and admissibility theorems belong more naturally in this second family than in the first, since they characterize undominated decision rules rather than utility representations. A selection result is different again: it adds a story about some optimization process — such as market competition, evolution, or training dynamics — and argues that systems lacking a certain property tend to be selected against. In short, representation concerns form, coherence concerns vulnerability or domination, and selection concerns survival under pressure.
Note that representation, coherence, and selection arguments are related, but not identical. In general, a coherence argument is a within-agent claim: it says that if a single agent violates some structural constraint, then the agent is vulnerable to a penalty such as a money pump, Dutch book, or dominated choice. A selection argument is different: it asks whether agents lacking that property would tend to disappear under some broader optimization pressure, such as market competition, training dynamics, or evolutionary selection. A coherence result may help motivate a selection story, yet it does not by itself show that realistic environments actually select against the offending preference pattern. Similarly, a representational failure with no plausible selection story may still be mathematically interesting while being less central for explaining the structure of real agents.
For the purposes of this note, the key point is that these notions of badness do not all coincide. A preference can fail standard representation without being exploitably incoherent, and not every departure from expected utility is thereby problematic. Nevertheless, the axioms studied below are useful because — as we will see — they rule out several undesirable properties.
4. Completeness, transitivity, and ordinal utility
Two structural conditions are especially important.
Definition 4.1 (Completeness). A preference relation on is complete if for every pair ,
Definition 4.2 (Transitivity). A preference relation on is transitive if for every ,
Completeness says that the agent can compare any two trajectories. This is a claim about comparability, not about certainty: it says that the preference relation returns a verdict on every pair, not that the agent knows everything about the consequences of those trajectories. Transitivity says that these verdicts fit together consistently across chains of comparison. For example, if and , then transitivity requires as well.
A preference relation satisfying both conditions is often called a weak order in economics and decision theory (Fishburn 1970; Kreps 1988; Mas-Colell et al. 1995). On a countable domain such as , weak orders admit an ordinal utility representation. More general representation theorems on richer domains go back to the classic work of Debreu 1954.
Proposition 4.3 (Ordinal representation on the trajectory space). A preference relation on is complete and transitive if and only if there exists a function such that
Proof
If such a function exists, then completeness and transitivity are inherited from the total order on . Conversely, if is complete and transitive, then trajectories can be grouped into indifference classes, where each class contains all trajectories tied with one another. The quotient of these classes is a countable total order. Any countable total order can be embedded in , so we may assign real numbers to the indifference classes in a way that preserves their order. Composing that assignment with the quotient map gives the desired utility function on trajectories.
This utility is ordinal: only the ranking matters. Any strictly increasing transformation of represents the same preference relation. So ordinal utility lets us encode the order of deterministic trajectories, but it does not yet tell us how to compare risky mixtures of them. The numbers themselves have no independent meaning beyond the order they induce. If represents a preference relation, then so does , or , provided the transformation remains strictly increasing. This is why ordinal utility is best thought of as a numerical labeling of ranks, not yet as a measure of how much one trajectory is preferred to another (Fishburn 1970).
It is also useful to understand what happens when either condition fails.
What if completeness or transitivity fail
If completeness fails. Then some pairs of trajectories are incomparable. This may be a feature rather than a bug: the agent may genuinely refuse to rank certain alternatives because its values are plural, under-specified, or context-sensitive. But once incomparability is allowed, no single real-valued utility function can exactly represent the relation in the sense of Proposition 4.3, because any two real numbers are themselves comparable. One must then move to a different formalism, such as partial orders, sets of utility functions, or multi-criteria representations (Aumann 1962).
If transitivity fails. Then local pairwise judgments need not assemble into a global ranking. In the simplest case one obtains a cycle
No scalar utility can represent such a cycle, since it would require
which is impossible. The logical problem is already serious: there is no single global ranking of the options. Under the additional behavioral assumption that the agent is willing to pay a small fee to move from any option to a strictly preferred one, this logical failure becomes operational through the money pump problem. Suppose an agent currently has and is willing to pay a small fee each time it moves to a strictly preferred trajectory. The cycle above licenses the sequence
with the agent paying at each step. At the end it is back where it started, but poorer by . Repeating the cycle pumps away arbitrarily much money (Gustafsson 2010).
There is also an interesting geometric view on transitivity. This paragraph is not needed for the rest of the note, so readers meeting these ideas for the first time can safely treat it as an optional aside. Fix a finite menu and draw the complete graph whose vertices are the trajectories in . Encode pairwise comparisons by an antisymmetric edge flow :
where means that is preferred to , while means the reverse. If preferences come from a utility function , then each edge weight is just a utility difference,
so is a discrete gradient field. In the language of the discrete Helmholtz–Hodge decomposition, any edge flow can be split into a gradient part, which comes from a potential, and a cyclic part, which records genuine loops. On the complete comparison graph there is no extra harmonic remainder, so inconsistency is entirely captured by the cyclic component (Jiang et al. 2011). A three-cycle corresponds exactly to a nonzero discrete curl:
Thus transitive preferences are precisely the potential part of the decomposition, with the potential given by the utility, while intransitive cycles show up as the rotational or cyclic part. If this language feels abstract, the key takeaway is simple: utility means all local comparisons come from one global ranking, whereas cycles are the leftover pattern that cannot be explained by any single scalar potential.
5. Lotteries and the von Neumann–Morgenstern axioms
Up to this point, we have just considered preferences over a countable set without considering stochasticity. To introduce uncertainty, we expand the domain from trajectories to lotteries over trajectories.
Let for denote the set of probability distributions over . An element can be expressed as
which should be read as a lottery that yields outcome with probability . The outcome is identified with the degenerate lottery . In this way, preferences over outcomes can be viewed as a special case of preferences over lotteries. Furthermore, for and , write
for the compound lottery that first flips a coin with bias , then samples from or accordingly.
The von Neumann–Morgenstern (vNM) framework studies weak preference relations over and asks when such preferences admit an expected-utility representation (von Neumann & Morgenstern 1944). The framework considers four axioms on preferences over lotteries.
Axiom 1 (Completeness). For all , either or (or both).
Axiom 2 (Transitivity). For all , if and , then .
Axiom 3 (Continuity). For all with , there exists such that
Axiom 4 (Independence). For all and ,
We already know about completeness and transitivity from the previous section; the new players are continuity and independence. Continuity says intermediate prospects admit a break-even mixture between better and worse ones. Independence says that if is preferred to , then mixing both with the same background lottery should not reverse that preference.
The natural question is therefore when a preference over lotteries can be represented by the expectation of a utility function on trajectories:
This formulation says that the value of a lottery is the probability-weighted average of the utilities of its possible trajectories. In its simplest finite-outcome form, this question is answered by the celebrated vNM theorem.
Theorem 5.1 (von Neumann–Morgenstern). Let be finite. A weak preference relation on satisfies completeness, transitivity, continuity, and independence if and only if there exists a function such that for all lotteries ,
Moreover, is unique up to positive affine transformations:
Proof
We first prove the easy direction: if preferences admit an expected-utility representation, then they satisfy the four axioms.
Completeness and transitivity follow immediately from the total order on . Continuity follows because if then , so one can choose such that
Hence
Finally, independence follows from linearity:
Therefore
We now prove the converse direction, namely that the four axioms imply an expected-utility representation.
- Pick best and worst outcomes (completeness and transitivity). Because is finite and the restriction of to degenerate lotteries is complete and transitive, either all outcomes in are indifferent or there exist such that
If all outcomes are indifferent, then the constant utility function already represents the preference relation. So we may assume . 2. Calibrate each outcome against and (continuity). For set , and for set . For , continuity gives some such that
Thus every outcome is indifferent to a lottery over the two benchmark outcomes and . 3. Compare the benchmark lotteries (independence and transitivity). We first show that for ,
For , write
We claim that if , then . Indeed, if , then , and independence applied to with mixing weight and background lottery gives
So it remains to consider the case . Define
Since , independence gives
But the left-hand side is exactly , while the right-hand side is . Hence larger values of yield strictly better benchmark lotteries. Together with Step 2 and transitivity, this proves the displayed equivalence above. 4. Reduce an arbitrary lottery to a benchmark lottery (independence). Let
By Step 2, each is indifferent to . Replacing each by the corresponding inside , one at a time, and using independence together with transitivity, yields
By ordinary probability arithmetic, the compound lottery on the right reduces to
- Conclude the expected-utility representation. Applying Step 3 to the benchmark lotteries from Step 4, we obtain
This is exactly the desired expected-utility representation.
It remains to prove affine uniqueness, which uses continuity, independence, and the representation just obtained. Suppose is another expected-utility representation of the same preference relation. Let
For any , Step 2 gives
Since also represents the same preferences, indifference implies equality of expected -value:
So is a positive affine transformation of .
Exercise 5.1 (Guided proof of the von Neumann–Morgenstern theorem). Assume is finite and that on satisfies completeness, transitivity, continuity, and independence.
(a) Show that there exist trajectories such that for every .
(b) Using continuity, show that for every there exists such that
(c) Use transitivity to show that if
then
(d) Let
Use independence repeatedly to replace each by the equivalent lottery and show that
(e) Reduce the compound lottery in part (d) to a simple lottery and prove that
(f) Show that
This gives the expected-utility representation.
(g) Prove affine uniqueness: if also represents the same preference relation in expected-utility form, then for some and .
(h) Prove the converse direction: if preferences admit an expected-utility representation, then they satisfy completeness, transitivity, continuity, and independence.
Solution
(a) Induction on . A single trajectory is its own best and worst element. Given a maximal and minimal for , completeness compares with each of them and transitivity makes the better of maximal and the worse of minimal for .
(b) If then every is indifferent to both and any constant works. Otherwise . By (a) we have , so continuity applied to the triple gives directly some with ; set . For uniqueness, independence implies the mixtures are strictly increasing in when , so two different weights cannot both be indifferent to .
(c) Suppose . Mixture monotonicity gives , and chaining the two calibrating indifferences through transitivity yields . If the same argument gives . Together: .
(d) Independence states that implies . View as a mixture in which appears with weight against the rest; replacing by its calibrated equivalent therefore leaves the whole lottery indifferent. Doing this for each in turn — finitely many steps — and chaining with transitivity gives .
(e) The right-hand side is a compound lottery whose only outcomes are and ; it awards with total probability . Identifying a compound lottery with the simple lottery it induces gives .
(f) By (e), and are indifferent to calibrated – mixtures with weights and ; by mixture monotonicity and transitivity, .
(g) Apply the -representation to the calibration of :
This is the required form , with intercept and slope , the latter strictly positive because and represents . (Note the statement's scalars are unrelated to the trajectories of part (a); the letter does double duty here, so read as the utility of the best trajectory throughout.)
(h) Let . Completeness and transitivity are inherited from the total order on . Continuity: is continuous, so upper and lower contour sets in are closed. Independence: mixing both sides with adds the same and scales the difference by , leaving the comparison unchanged.
In contrast to our previous result, vNM utility is cardinal up to positive affine transformations, not merely ordinal. A general monotone transformation would destroy the expectation formula. The independence axiom forces exactly the amount of structure needed for linear averaging.
Two notions of utility.
In economics it is standard to distinguish two different notions of utility (Kreps 1988; Mas-Colell et al. 1995).
The first is ordinal utility, also known as preference utility or `F1', which represents preferences over certain outcomes. This is the object obtained in Proposition 4.3: if preferences over deterministic trajectories are complete and transitive, then there exists a function such that
Only the ranking matters. Any strictly increasing transformation of represents the same preferences (Fishburn 1970).
The second is von Neumann–Morgenstern utility, also called `F2'. This is the function that appears inside the expectation operator when preferences over lotteries satisfy continuity and independence:
Unlike ordinal utility, is unique only up to positive affine transformations. Its numerical differences therefore carry behavioral content: they determine which mixtures an agent is willing to accept, and so encode the agent's attitudes toward risk and gambling structure (von Neumann & Morgenstern 1944; Mas-Colell et al. 1995).
Exercise 5.2 (Ordinal versus vNM utility). Let satisfy .
(a) Give two different ordinal utility functions that represent the same ranking of .
(b) Explain why these two functions are equally good as ordinal representations.
(c) Suppose in addition that
Show that not every strictly increasing transformation of a vNM utility can preserve this indifference.
Solution
(a) For instance and .
(b) An ordinal representation encodes only the ranking: represents iff , and both functions induce . Composing with any strictly increasing map preserves exactly this information, so no ordinal criterion can distinguish from .
(c) Under expected utility the indifference forces . Take the vNM utility and the strictly increasing map . Then , while the lottery has expected value . Since , the relation represented by strictly prefers the lottery to : the indifference is destroyed. Only positive affine transformations preserve all such midpoint identities, which is the uniqueness part of the vNM theorem.
If continuity or independence fails
It is useful to separate the roles of the four vNM axioms:
- As before, completeness and transitivity give us an ordinal utility on deterministic trajectories.
- Continuity and independence turn that ordinal picture into an expected utility theory over uncertain prospects.
If completeness or transitivity fail, as discussed in the previous section, we lose hope of describing the preference by a single scalar quantity. If continuity or independence fails, the preference will generally no longer admit the linear expectation form.
Failure of continuity. To build intuition, let and define
Then
so as increases from to , the lottery moves from the worse prospect to the better prospect . Continuity says, informally, that preferences vary smoothly along this path. In particular, it says that the intermediate prospect can be matched by some mixture of the better and worse prospects:
for some . Intuitively, the axiom requires no abrupt jumps in preference as the mixture probabilities vary.
When continuity fails, some distinctions can become lexicographic. For example, we could have a case where , but for all . This means that has absolute priority over : as soon as appears with any nonzero probability, the lottery becomes strictly better than . One way to interpret this is that some considerations are given lexical priority over others. For instance, avoiding catastrophe might outrank ordinary gains so completely that no finite improvement in ordinary reward compensates for even an arbitrarily small increase in catastrophic risk. Such preferences need not be contradictory; they are simply too sharp to be captured by a single real-valued utility whose expectations are taken in the usual way.
What breaks in this case is not all representations, but the real-valued vNM representation. If continuity is dropped, one is naturally led to lexicographic or non-Archimedean representations: for example, ordered pairs or vectors of utilities compared lexicographically, or utilities with infinitesimal scales (Fishburn 1971). In those models, the first coordinate records the highest-priority consideration, and lower coordinates matter only when higher ones tie.
Failure of independence.
Independence says that common lottery components should cancel. If , then mixing both with the same background lottery at the same rate should preserve the ranking:
So the relative ranking of and should depend only on how they differ, not on the common part they share.
When independence fails, the value of a prospect depends on its context. The same local substitution can be attractive in one background and unattractive in another. This is exactly what the Allais pattern shows (Allais 1953). Let denote trajectories yielding , , and million, and consider
Many people prefer to , but prefer to . Under independence this is impossible, because versus and versus differ only by a common consequence. The reversal shows that certainty is treated as psychologically special: replacing a sure outcome by a tiny risk of getting nothing matters more than expected utility allows.
This is the general lesson of independence failure. Probabilities are no longer aggregated linearly against a fixed utility function on trajectories. Common branches cannot be canceled, and the whole shape of the distribution starts to matter. Decision makers may overweight certainty, distort small probabilities, care about disappointment or regret, or evaluate gains and losses relative to a reference point rather than in absolute terms. This is the route taken by prospect theory and related non-expected-utility models (Kahneman & Tversky 1979; Machina 1982).
From the perspective of sequential choice, independence also matters because it allows one to replace a sublottery by an equivalent one without changing the value of the larger plan. If independence fails, the value of a branch may depend on the branches surrounding it, so local and global evaluations need not line up. Hammond's consequentialist argument shows that, together with dynamic consistency and suitable sequential assumptions, one is pushed back toward independence (Hammond 1988). But that only shows one route to coherent planning. One may instead keep a richer, non-linear evaluation of lotteries and give up the idea that common consequences are always behaviorally irrelevant.
So the two failures have different meanings. Failure of continuity says that some priorities are infinitely sharp, leading naturally to lexicographic or infinitesimal utility scales. Failure of independence says that uncertainty is evaluated holistically rather than by linear averaging, leading to models in which background risk, certainty, or reference dependence affect choice.
Exercise 5.3 (Failures of the vNM axioms). Here we will study the consequences of different axioms.
(a) Construct a simple incomplete preference relation on three trajectories. Why can it not be represented by a single real-valued ordinal utility?
(b) Construct a three-cycle and explain how it gives rise to a money pump.
(c) Give an example of a lexicographic preference over three outcomes and show that it violates continuity.
(d) Write down the Allais pattern from Section 5 and explain which axiom it violates.
Solution
(a) Let , , with and incomparable (neither nor ). Any satisfies or because the reals are totally ordered, so the relation it represents is complete; incomparability cannot be encoded by a single real-valued function.
(b) Take . Suppose you hold and are willing to pay some small to exchange an item for one you strictly prefer. Since you pay to swap ; since you pay to swap ; since you pay to swap . You are back where you started, poorer, and the cycle can be run forever.
(c) Give outcomes two attributes and compare lexicographically (the first attribute decides unless tied): , , , so . Extend to lotteries by comparing expected attribute vectors lexicographically. Continuity would require some with . But that mixture has attribute vector : for every it beats on the first attribute, and for it is -below . No mixing weight produces indifference.
(d) With paying , and million, the pattern of Section 5 is
with but . Both pairs differ only by a common consequence: replacing of by of turns into and into . Independence says a ranking is unchanged when the same consequence is mixed into both sides with the same weight, so the pattern violates independence (completeness, transitivity, and continuity are all consistent with it).
6. A fifth axiom: reward and discount
The vNM theorem gives an expected-utility representation over lotteries of whole trajectories, but it does not yet tell us that utility can be generated locally from stepwise rewards.
To investigate the possibility of localizing utility over time, let us denote by the set of one-step transitions. For and , write for the trajectory obtained by prepending to , and extend this operation to lotteries by
Now, using a vNM utility one can always define a continuation-dependent increment as
However, this increment may depend on the whole continuation . What is missing is a temporal condition ensuring that the effect of prepending a transition depends only on that transition itself. Following (Bowling et al. 2023), this can be explored with the following additional axiom.
Axiom 5 (Temporal -indifference). There exists a function such that for all and all ,
Under expected utility, the indifference above is equivalent to
which can be rearranged as
So the effect of prepending the same transition is to rescale the utility difference between two continuations by a factor . In that sense, measures how much the future still matters after the step has occurred. When , future differences are preserved without discount at that step. When , the future still matters but is damped. When , once has occurred the continuation no longer affects the comparison. Bowling's axiom is closely related to the transition-dependent discounting discussed by White 2017 and Pitis 2019.
Bowling et al. show that adding this axiom to the four vNM axioms is necessary and sufficient for a discounted-reward representation of preferences.
Theorem 6.1 (Markov reward representation, after Bowling et al.). A weak preference relation on satisfies completeness, transitivity, continuity, independence, and temporal -indifference if and only if there exist functions , , and such that ,
and, for all lotteries ,
Moreover, is unique up to positive scale, and is the same function appearing in the fifth axiom (Bowling et al. 2023).
This theorem sharpens the vNM result in exactly the way reinforcement learning needs. Utility is no longer an arbitrary scalar attached to a complete trajectory; it is generated recursively from a local reward and a local weight on the future. Unrolling the recursion for a trajectory gives
Thus the fifth axiom is what allows utility over whole trajectories to be decomposed into local reward and discounting. Two familiar special cases are worth highlighting.
- If is constant, then
which is the standard discounted-return objective.
- If , then
so utility is simply the additive cumulative reward.
In summary, the step from vNM utility over whole trajectories to RL-style reward requires one more axiom. That axiom simultaneously identifies both the local reward signal and the form of discounting.
Remark (MDPs, POMDPs, and locality). The locality result of this section is especially natural in fully observed MDPs, where one often expects reward to depend only on the current state transition. In the present note, however, the primitive histories are sequences of observations and actions, and the corresponding local reward has the form . In partially observed settings this can be restrictive: a goal may be Markov in the hidden state, in the agent's belief state, or in some augmented memory state, without being reducible to a function of the current observation-action pair alone. In that sense, the fifth axiom should be read as characterizing when preferences admit a reward that is local in the chosen representation of experience. If the raw observation stream is too coarse, one may need to enrich the state description before a Markov reward representation becomes available (Bowling et al. 2023).
Utility is not reward.
It is helpful to keep the following four different objects apart.
- Preference is the primitive relation . It says only which trajectories or lotteries are weakly preferred to which others.
- Preference utility (`F1') is an ordinal representation of preferences over deterministic trajectories. Under completeness and transitivity, it is any function such that
It is unique only up to strictly increasing transformations (Fishburn 1970).
- vNM utility (`F2') is the stronger, cardinal utility that appears when preferences over lotteries satisfy the vNM axioms. It is the function for which
It is unique only up to positive affine transformations (von Neumann & Morgenstern 1944). Thus F2' refines F1': it agrees with it on the ranking of certain trajectories, but adds the extra structure needed to compare lotteries.
- Reward is not either of these utilities. A reward function is a local representation introduced only after adding temporal structure. In the present framework, reward appears when the fifth axiom allows utility to be written recursively as
So reward is a way of decomposing utility over complete trajectories into stepwise contributions. It is therefore downstream of preference, and even downstream of vNM utility. This is why different reward functions can encode the same utility or the same preference ordering, a point we return to in the next section (Bowling et al. 2023).
Exercise 6.1 (Guided proof of the Markov reward representation theorem). Assume the hypotheses of Exercise 5.1 together with temporal -indifference, and let be a vNM utility representation over trajectories.
(a) Apply expected utility to temporal -indifference and show that for all lotteries and all transitions ,
(b) Rearrange part (a) to prove that
is independent of the choice of .
(c) Define
for any . Deduce that for every deterministic trajectory ,
(d) Show that this recursion extends to lotteries by linearity:
(e) Prove the converse direction: if , , and satisfy
for all and , and if extends linearly to lotteries, then temporal -indifference holds.
(f) Unroll the recursion along a finite trajectory and derive
(g) Show that when is constant, the previous formula reduces to standard discounted return, and when it reduces to additive cumulative reward.
(h) Show that if , then the corresponding reward must be
Interpret this as a source of reward non-uniqueness.
Solution
(a) Apply (in expected-utility form) to both sides of the axiom's indifference and multiply by :
(b) Rearranging, for all : the quantity does not depend on which lottery is used to compute it.
(c) By (b) the definition of is unambiguous. Taking to be the point mass on a deterministic continuation gives .
(d) is the pushforward of under prepending , and is linear on lotteries, so .
(e) With the recursion and linearity, , which is symmetric under ; hence . Dividing by shows both mixtures in the axiom have equal expected utility, and since represents , they are indifferent.
(f) Normalise (allowed by affine freedom). Induction on : , and expanding the inner term yields
(g) If , the product is and : discounted return. If every product is and : additive cumulative reward.
(h) . The same preferences thus admit a family of reward functions: even with fixed, is only determined up to a positive scaling and an additive shift modulated by .
7. Reward equivalence and shaping
The preceding theorem shows how to pass from utility over lotteries of trajectories to a local reward representation. But it also raises an important question: how unique is that reward? The answer has two parts. If one fixes a particular utility function and a particular discount function , then the reward is determined. But if one changes the numerical representative of the same preference relation, or allows shaping transformations, then many different rewards can encode the same underlying ordering.
First observe that once and are fixed, the reward is fixed as well. Indeed, if
for all and , then necessarily
and the right-hand side must be independent of the continuation . So the non-uniqueness of reward does not come from ambiguity inside a fixed recursive representation. It comes from the fact that utility itself is not unique as a numerical object.
Proposition 7.1 (Multiple rewards can encode the same preference relation). Suppose , , and satisfy
for all and , and that preferences over lotteries are represented by the expected value of . For any and , define
Then for all and , and the expected value of represents the same preference relation as the expected value of .
The proposition shows that reward non-uniqueness already appears at the level of affine reparameterizations of utility: changing the numerical representative of the same vNM preference relation generally changes the associated reward function as well. This point is especially simple under the normalization used in the previous section. Then the additive degree of freedom disappears, and one is left only with positive rescalings:
So in the normalized setting reward is unique only up to choice of units.
There is also a second, less trivial kind of non-uniqueness, in which one changes rewards while leaving the induced trajectory ordering unchanged, first studied by Ng et al. 1999.
Proposition 7.2 (Potential-based shaping changes utility only by a boundary term). Consider a state-based setting with transitions and a potential function on states.
- If and , then
for any trajectory from to . 2. If is constant and , then
The previous result shows that shaping into changes utility only by a boundary term. In particular, if all compared trajectories share the same start state and terminal potential, then the induced preference ordering is unchanged; if the boundary term vanishes on all admissible trajectories, then even the numerical utility is unchanged.
The general lesson is that the same preferences admit many different reward functions, so no single one of them is canonical. Some transformations, such as positive scaling, merely change the numerical units. Others, such as potential-based shaping, redistribute value along the trajectory while leaving the overall preference over complete trajectories unchanged. This is why reward design is often underdetermined: what matters behaviorally is not a raw reward function in isolation, but the utility and preference structure it induces.
Exercise 7.1 (Reward shaping). Assume additive reward with and define
(a) Show that along any finite trajectory the total shaped reward differs from the original total reward by the boundary term .
(b) Under what condition on the admissible trajectories does this shaping leave utility exactly unchanged?
(c) Under what weaker condition does it leave only the induced preference ordering unchanged?
Solution
(a) Writing , the shaped total is
since the middle sum telescopes.
(b) Utility is exactly unchanged iff the boundary term vanishes on every admissible trajectory, i.e. throughout — for example when all trajectories start in a fixed and takes the value on every reachable terminal state.
(c) Only the ordering is at stake if the boundary term is the same constant for all compared trajectories: then every utility shifts by and the induced ranking is untouched. A fixed start state together with constant on the reachable terminal states suffices, whatever that constant is.
8. Conclusion
The main lesson of this note is that, in sequential settings, the primitive object is not a one-step reward but a preference over complete trajectories. From that starting point one can distinguish several layers of structure. Completeness and transitivity yield an ordinal utility over deterministic trajectories. Once lotteries are introduced, continuity and independence refine that ordinal picture into a von Neumann–Morgenstern utility, which agrees with the ranking of certain trajectories but carries strictly more information because it calibrates trade-offs between risky prospects.
This separation of layers clarifies both what expected utility achieves and what it leaves out. It explains why "utility" in economics can refer either to an ordinal representation of certain preferences or to a cardinal object suitable for evaluating lotteries. It also shows what fails when particular axioms are dropped: without completeness or transitivity, scalar representation itself can fail; without continuity or independence, one may still have meaningful preferences, but no longer the linear expectation form of vNM. In that sense, expected utility is not the whole of rationality, but a particular and mathematically powerful strengthening of it.
The final step is to see that even vNM utility is not yet reward. To obtain a local reward-and-discount representation one needs additional temporal structure, captured here by Bowling et al.'s fifth axiom. Under that axiom, utility over whole trajectories admits a recursive decomposition into local reward and discounting. But even then reward is not unique: affine changes of utility induce corresponding changes of reward, and potential-based shaping can redistribute value along a trajectory while preserving the underlying ordering. Moreover, the locality of reward depends on the chosen representation of experience, which is natural in fully observed MDPs but can be restrictive in partially observed settings unless the state is suitably enriched. The reward hypothesis, understood in this way, is therefore best read not as the claim that goals are primitively rewards, but as the claim that sufficiently structured preferences over trajectories can be represented by rewards.
Open-ended questions
Exercise 8.1. Do the results have any prescriptive or descriptive implications? What kinds of agents, with what kinds of preferences should we design? By default, what kind of agents should we expect to obtain via typical training processes?
Solution
Discussion notes, not a unique answer. Prescriptively, the money-pump and dominance arguments say that an agent which cares about a fungible resource is under pressure toward completeness and transitivity — coherence is an attractor for agents that can be exploited otherwise. Descriptively, nothing guarantees that trained systems satisfy any axiom: gradient descent on episodic objectives can produce context-dependent, intransitive, or incomparable preferences, especially off-distribution. A reasonable expectation is approximate coherence where incoherence was penalised during training, and no guarantee elsewhere. For design, the trade-off runs both ways: highly coherent agents are more predictable and analysable but are exactly the ones for which instrumental-convergence arguments bite; agents with incomplete or unstable preferences may be harder to exploit into goal-directed resource acquisition, at the cost of being harder to reason about.
Exercise 8.2. Let's assume a superintelligence has preferences over trajectories. Are there weaker versions of the axioms that make it more plausible that these preferences are "safe for us" compared to preferences that lead to utilities or even rewards?
Solution
Discussion notes. The natural candidates weaken one axiom at a time. Dropping completeness is the most studied: an agent with incomplete preferences can remain undecided between continuing and being shut down, so it is not pushed by coherence arguments toward shutdown-resistance; money pumps also lose force because the agent may simply refuse trades between incomparable options. Weakening continuity permits lexicographic safety: "never cross the constraint, then optimise" cannot be represented by a single real-valued utility, which is arguably a feature. Weakening independence allows certainty-favouring (Allais-like) preferences, which dampen gambling-for-resources behaviour. Weakening temporal -indifference removes the local reward representation: goals about the shape of a trajectory as a whole (diversity, "never do X") stop being expressible as accumulated reward. The common pattern: each axiom kept buys a representation theorem and optimisation pressure; each axiom dropped blocks a coherence-based failure mode while making the agent's behaviour less analysable.
Exercise 8.3. More generally, are there interesting examples of preferences over trajectories that the students could analyse, that do or do not satisfy some of the axioms?
Solution
Discussion notes; instructive examples include: (i) lexicographic safety-first preferences — violate continuity (see Exercise 5.3); (ii) Pareto/multi-objective preferences, undominated but unaggregated — violate completeness; (iii) hyperbolic discounting — satisfies the vNM axioms at a single decision time yet violates temporal -indifference, hence admits utility but no stationary reward/discount pair, and is dynamically inconsistent; (iv) the certainty effect (Allais) — violates independence only; (v) satisficing ("anything above the threshold is equally fine") — complete and transitive, so ordinal utility exists, but indifference plateaus interact oddly with lotteries near the threshold; (vi) trajectory-shape goals such as "visit as many distinct states as possible" — can satisfy all four vNM axioms (a utility exists) while violating the temporal axiom, illustrating exactly the gap between utility and reward.
References
Maurice Allais (1953). Le comportement de l'homme rationnel devant le risque: critique des postulats et axiomes de l'ecole Americaine. Econometrica.
Robert J Aumann (1962). Utility theory without the completeness axiom. Econometrica.
Michael Bowling, John D. Martin, David Abel, and Will Dabney (2023). Settling the Reward Hypothesis. arXiv preprint arXiv:2212.10420.
Gerard Debreu (1954). Representation of a preference ordering by a numerical function. Decision Processes.
Peter C. Fishburn (1970). Utility Theory for Decision Making. Wiley.
Peter C. Fishburn (1971). A Study of Lexicographic Expected Utility. Management Science.
Johan E. Gustafsson (2010). A Money-Pump for Acyclic Intransitive Preferences. Dialectica.
Peter J. Hammond (1988). Consequentialist Foundations for Expected Utility. Theory and Decision.
Xiaoye Jiang, Lek-Heng Lim, Yuan Yao, and Yinyu Ye (2011). Statistical ranking and combinatorial Hodge theory. Mathematical Programming.
Daniel Kahneman and Amos Tversky (1979). Prospect Theory: An Analysis of Decision under Risk. Econometrica.
David M. Kreps (1988). Notes on the Theory of Choice. Westview Press.
Mark J. Machina (1982). "Expected Utility" Analysis without the Independence Axiom. Econometrica.
Andreu Mas-Colell, Michael D. Whinston, and Jerry R. Green (1995). Microeconomic Theory. Oxford University Press.
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell (1999). Policy Invariance under Reward Transformations: Theory and Application to Reward Shaping. Proceedings of the Sixteenth International Conference on Machine Learning.
Silviu Pitis (2019). Rethinking the discount factor in reinforcement learning: A decision theoretic approach. Proceedings of the AAAI Conference on Artificial Intelligence.
John von Neumann and Oskar Morgenstern (1944). Theory of Games and Economic Behavior. Princeton University Press.
Martha White (2017). Unifying Task Specification in Reinforcement Learning. Proceedings of the 34th International Conference on Machine Learning.