This is the sixth part of an effort to add control to the Hidden Markov Model RxInfer example. Eventually this work will be used in a more elaborate scheduling project making use of active inference.
In the previous parts, the agent first watched a Roomba move through a 3-room apartment, consisting of a bedroom, a living room, and a bathroom, inferring which room it was in (Parts 1 and 2) and whether that room was occupied (Part 3). In Part 4, the agent took control, with a single goal: to avoid disturbing people. In Part 5, it got a second goal, to clean, and the model of the environment grew from one state factor to five:
the room the Roomba is in (bedroom, living room or bathroom), which the agent controls,
where the person is (bedroom, living room, bathroom, or away), who moves around independently, and
for each of the three rooms, whether it is dirty (clean or dirty).
With both goals, the agent kept the apartment cleaner than a Roomba with only the goal of avoiding people, most of all the living room, at the cost of a little extra disturbance. To do so, it combined the five factors into their joint states, \(3 \cdot 4 \cdot 2 \cdot 2 \cdot 2 = 96\) in total, and computed its beliefs about them exactly.
The problem: the joint states multiply. Exact inference worked well for 96 joint states, but it does not scale. Every new factor multiplies the number of joint states by its number of values. An apartment with ten rooms, a person who can be in any of them or away, and dirt in each room would already have \(10 \cdot 11 \cdot 2^{10} = 112\,640\) joint states. A belief over all of them, updated and predicted at every tick for every candidate plan, quickly becomes too expensive. Yet the model itself stays small: each factor and each sensor depends on only a few others, which is exactly what multi-factor modeling is about.
The remedy: factorized beliefs. In this part, the agent no longer keeps a belief over the joint states. Instead, it keeps one belief per factor: where the Roomba is, where the person is, and whether each room is dirty. That is \(3 + 4 + 2 + 2 + 2 = 13\) numbers instead of 96, and in the ten-room apartment \(10 + 11 + 10 \cdot 2 = 41\) instead of 112 640. The agent treats the product of these beliefs as its belief about the whole state. This is the mean-field approximation, the standard way of handling larger models in active inference, and the one used by pymdp. Its beliefs are computed by a fixed-point iteration: at each tick, each factor’s belief is updated in turn, taking the others’ beliefs into account, until they settle. This iteration minimizes a variational free energy, so the free energy, absent in Parts 4 and 5, returns, now within each tick.
The price. A product of separate beliefs cannot represent relations between factors. In Part 5, a motion reading told the agent that the Roomba and the person were probably in the same room, a statement about two factors together. A factorized belief can only say where each of them probably is on its own. How much this matters is the main question of this part.
A controlled comparison. Everything else stays as in Part 5: the environment, the sensors, the two preferences, no motion and dirt picked up, and the planning horizon of three steps. Only the agent’s inference changes, so any difference in its beliefs or behavior comes from the approximation alone. We compare the two agents in three ways:
beliefs: an exact filter runs alongside the factorized agent, fed the same observations and actions, so the two sets of beliefs can be compared tick by tick,
decisions: at each tick, we check whether the exact agent would have chosen the same action, and
behavior: disturbance, cleanliness and runtime, against the exact agent of Part 5.
We expect the approximation to do well here, since the camera keeps the Roomba’s room nearly certain, and when one factor is certain, the relations the approximation loses hardly matter. To also see where it breaks down, we make the camera less reliable, so that the room becomes uncertain.
The structure of the code is the same as in Part 5:
create_envir()
execute(): carries out the chosen action, moving the Roomba and the person, and updating dirt
observe()
state() (true state, for evaluation only)
create_agent()
act(): returns the chosen action
future(): predicts future states and observations for the candidate actions
compute(): infers the current state, now by fixed-point iteration, and evaluates the candidate actions
slide(): moves the agent’s time window one step forward
As in the previous parts, and as in the active inference literature, \(\mathbf{A}\) holds the observation (likelihood) tensors and \(\mathbf{B}\) the state transition tensors, swapped compared to the original RxInfer example, which follows the engineering convention of using \(A\) for state transitions. Each tensor has one axis for each factor it depends on.
versioninfo() ## Julia version
Julia Version 1.12.7
Commit 6d172b025e4 (2026-08-15 08:05 UTC)
Build Info:
Official https://julialang.org release
Platform Info:
OS: Linux (x86_64-linux-gnu)
CPU: 12 × Intel(R) Core(TM) i7-8700B CPU @ 3.20GHz
WORD_SIZE: 64
LLVM: libLLVM-18.1.7 (ORCJIT, skylake)
GC: Built with stock GC
Threads: 1 default, 1 interactive, 1 GC (on 12 virtual cores)
Updating registry at `~/.julia/registries/General.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
## tuple0(x...): build a 0-based tuple (x^{[0]}, ..., x^{[N]}), mirroring the math.## - each argument becomes one entry: tuple0(v) wraps v; it never re-indexes v itself## - the index range 0:N is derived from the number of arguments## - entries may differ in shape, e.g. likelihood matrices for different modalities## Returns an OffsetVector (mutable), not a Julia Tuple.tuple0(x...) =OffsetVector(collect(x), 0:length(x)-1)## offset0(v): re-index an existing vector from 0, e.g. a trajectory over t = 0, ..., T.## Unlike tuple0, it does not wrap v; it shifts v's own indices.offset0(v) =OffsetVector(v, 0:length(v)-1)
offset0 (generic function with 1 method)
Why tuple0() rather than OffsetVector() directly?
It prevents a level mix-up. With OffsetVector, it is easy to offset the wrong level:
OffsetVector([1.0, 0.0, 0.0], 0:2) ## the one-hot vector itself, re-indexed 0:2 (wrong level)OffsetVector([[1.0, 0.0, 0.0]], 0:0) ## a tuple containing the one-hot vector (intended)tuple0([1.0, 0.0, 0.0]) ## the same as the second line, with no way to get it wrong
The first line is valid code but means something else: the one-hot vector $\hat{\boldsymbol{s}}^{\ast[0,0]}$ itself, now indexed from 0, rather than the tuple $\hat{\mathbf{s}}^{\ast(0)} = \big(\hat{\boldsymbol{s}}^{\ast[0,0]}\big)$. Nothing would complain, and later code that indexes `[0]` would get a single number instead of a vector. `tuple0` always wraps its arguments as *entries*, so this mistake cannot happen.
The index range is computed automatically. With OffsetVector, the range is written by hand and has to agree with the number of entries: adding a second modality to LAˣ while forgetting to change 0:0 to 0:1 gives an error at best. tuple0 derives the range from the number of arguments, so adding or removing an entry is a one-line change.
It reads like the math.tuple0(A0, A1) looks like \(\big(\boldsymbol{A}^{[0]}, \boldsymbol{A}^{[1]}\big)\): entries separated by commas, nothing else. OffsetVector([A0, A1], 0:1) adds square brackets and a range that have nothing to do with the math.
It states the intent.OffsetVector says “an array with shifted indices”, which could be anything. tuple0 says “a tuple of factors or modalities, as in the math”, the specific concept from (10.33). And since all tuples are built the same way, a search for tuple0( finds every one of them.
There is one place to change things. Changing how tuples work, e.g. switching to an immutable 0-based type, or adding a check that every column of every \(\boldsymbol{A}\) and \(\boldsymbol{B}\) sums to 1, means changing one line instead of every place a tuple is built.
Entries of different shapes just work.collect builds a vector of whatever is passed in, so tuple0(A0, A1) with a \(3\times 3\) and a \(2\times 3\) matrix gives a vector of matrices, which is exactly what a tuple of likelihood matrices with different numbers of observations needs.
2 DATA UNDERSTANDING
The data is simulated and, as in Parts 4 and 5, it arises from the interaction between the agent and the environment. At each tick, the agent observes, chooses an action, and the environment responds, so the Roomba’s path depends on the agent’s choices. A run covers \(T + 1 = 600\) ticks, \(t = 0, \ldots, T\), and records three kinds of one-hot encoded vectors:
Observations, which the agent receives: the observation tuple \(\hat{\mathbf{o}}^{(t)} = \big(\hat{\boldsymbol{o}}^{[t,0]}, \hat{\boldsymbol{o}}^{[t,1]}, \hat{\boldsymbol{o}}^{[t,2]}\big)\), with one entry per sensor:
\(\hat{\boldsymbol{o}}^{[t,0]}\) (Texture) over \(\{\text{hardwood}, \text{carpet}, \text{tiles}\}\)
\(\hat{\boldsymbol{o}}^{[t,1]}\) (Motion) over \(\{\text{no motion}, \text{motion}\}\)
\(\hat{\boldsymbol{o}}^{[t,2]}\) (Dirt) over \(\{\text{nothing picked up}, \text{dirt picked up}\}\)
Actions, which the agent chooses: the action tuple \(\hat{\mathbf{u}}^{(t)} = \big(\hat{\boldsymbol{u}}^{[t,0]}\big)\), whose entry \(\hat{\boldsymbol{u}}^{[t,0]}\) is over the room the Roomba heads for: \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\). Heading for the room it is already in means staying put. The action chosen at tick \(t\) takes effect at tick \(t+1\), so actions are recorded for \(t = 0, \ldots, T-1\).
True states, which the agent never sees: the state tuple \(\hat{\mathbf{s}}^{\ast(t)} = \big(\hat{\boldsymbol{s}}^{\ast[t,0]}, \ldots, \hat{\boldsymbol{s}}^{\ast[t,4]}\big)\), with one entry per state factor, as in Part 5:
\(\hat{\boldsymbol{s}}^{\ast[t,0]}\) (Room) over \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\), where the Roomba is
\(\hat{\boldsymbol{s}}^{\ast[t,1]}\) (Person) over \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}, \text{away}\}\), where the person is
\(\hat{\boldsymbol{s}}^{\ast[t,2]}\), \(\hat{\boldsymbol{s}}^{\ast[t,3]}\) and \(\hat{\boldsymbol{s}}^{\ast[t,4]}\) (Dirt in the bedroom, living room and bathroom), each over \(\{\text{clean}, \text{dirty}\}\)
In each one-hot vector, the component that is 1 marks the value that occurred: for the state factors, where the Roomba and the person actually are, and which rooms are actually dirty; for \(\hat{\boldsymbol{o}}^{[t,0]}\), \(\hat{\boldsymbol{o}}^{[t,1]}\) and \(\hat{\boldsymbol{o}}^{[t,2]}\), the texture, motion and dirt the sensors report; for \(\hat{\boldsymbol{u}}^{[t,0]}\), the room the agent chose to head for.
As in Part 5, the state tuple has five entries, which differ in size: 3 values for the room, 4 for the person, 2 for each room’s dirt. This is why the states form a tuple rather than a single vector. In this part, the agent’s beliefs follow the same structure: one belief per factor, with no belief over their combinations.
The agent works with the observations only, and with the actions it chose itself. The true states are kept for evaluation: they let us check how well the agent’s beliefs match the real situation, how often the Roomba ended up in the same room as the person, and how clean the apartment was kept.
The data serves two further purposes in this part. First, the observations and actions of the factorized agent’s run are also fed to an exact filter, as used in Part 5, so that the factorized beliefs can be compared with the exact ones on exactly the same data. Second, for the stress test, a second run is made with a less reliable camera, i.e. a different Texture likelihood in the environment and in the agent’s model; everything else about the data stays the same.
Changes in the data compared to Part 5
The data itself is unchanged: the same observations, actions and true states, with the same structure.
The agent’s beliefs are one per factor, following the structure of the state tuple, with no belief over the joint states.
Two further uses of the data: the factorized agent’s observations and actions are also fed to an exact filter, for a comparison on identical data, and a second run with a less reliable camera serves as a stress test.
3 DATA PREPARATION
As in the previous parts, we use the data from the simulation directly; no cleaning or transformation is needed. As in Parts 4 and 5, the data becomes available tick by tick: at each tick, the agent receives only the current observation tuple \(\hat{\mathbf{o}}^{(t)}\), updates its beliefs, and chooses its next action. There is therefore no separate preparation step before inference; the agent processes the data as it arrives.
Throughout this notebook, trajectories are indexed from 0, so that index \(t\) always means time \(t\). During a run, we record five trajectories, each re-indexed from 0 with offset0:
Lôː: the received observation tuples \(\hat{\mathbf{o}}^{(0:T)}\),
Lûː: the chosen action tuples \(\hat{\mathbf{u}}^{(0:T-1)}\),
Lsː: the agent’s belief tuples \(\mathbf{s}^{(0:T)}\), one belief per state factor at each tick,
Ls̃ː: the exact reference beliefs, in the same format, from an exact filter fed the same observations and actions, for evaluation only, and
Lŝˣː: the true state tuples \(\hat{\mathbf{s}}^{\ast(0:T)}\), for evaluation only.
In Part 5, the agent worked internally with the 96 joint states, i.e. all combinations of the five factors, and its belief about each factor was a marginal, obtained by adding up the joint belief. In this part, there is no joint belief: the agent’s belief consists of five separate beliefs, one per factor: \(\boldsymbol{s}^{[t,0]}\) about the Roomba’s room, \(\boldsymbol{s}^{[t,1]}\) about the person’s location, and \(\boldsymbol{s}^{[t,2]}\), \(\boldsymbol{s}^{[t,3]}\) and \(\boldsymbol{s}^{[t,4]}\) about the dirt in each room. They are what the agent computes and acts on, not a summary of something larger. Recording them is therefore straightforward, and since they have the same structure as the true state tuple, they can be compared with it factor by factor.
The exact reference beliefs \(\tilde{\boldsymbol{s}}^{[t,n]}\) are the marginals of the exact filter of Part 5, computed alongside the run but never used by the agent. They have the same structure as the agent’s beliefs, so the two can be compared factor by factor and tick by tick. Because both filters see exactly the same observations and actions, any difference between them comes from the approximation alone.
As in Parts 4 and 5, this part does not use RxInfer: the agent’s computations are written directly in Julia, so no conversion at an RxInfer boundary is needed.
Changes in the data preparation compared to Part 5
The agent’s beliefs are recorded directly: they are five separate per-factor beliefs, not marginals of a joint belief over 96 states.
A fifth trajectory,Ls̃ː, holds the beliefs of an exact filter fed the same observations and actions, written \(\tilde{\boldsymbol{s}}^{[t,n]}\). It is used only for evaluation, to measure the effect of the approximation.
4 MODELING
4.1 Narrative
In general, we will have the following symbol conventions. They are shown for states \(s\); observations \(o\) and actions \(u\) follow the same pattern.
Indices. Time-only indices are written in \((\,)\), compound or non-time indices in \([\,]\). The index therefore tells you what kind of object you are looking at: a superscript \((t)\) collects everything at time \(t\), while \([t,n]\) picks out a single state factor \(n\) at time \(t\). New in this part, the iterations within a tick get their own brackets, \(\langle\,\rangle\), written as a subscript: \(\boldsymbol{s}^{[t,n]}_{\langle k \rangle}\) is the belief about factor \(n\) at time \(t\) after iteration \(k\).
\(t\): time step, \(t = 0, \ldots, T\)
\(k\): iteration of the fixed-point iteration within a tick, \(k = 1, \ldots, K\)
\(s^{[t,n]}\): categorical random variable for state factor \(n\) at time \(t\). Its values are labels. As in Part 5, there are five state factors:
factor 0 (Room): \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\), where the Roomba is
factor 1 (Person): \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}, \text{away}\}\), where the person is
factors 2, 3 and 4 (Dirt in the bedroom, living room and bathroom): each \(\{\text{clean}, \text{dirty}\}\)
A random variable is not bold, because a single categorical random variable has no structure.
\(\mathbf{s}^{(t)} = \big(s^{[t,0]}, \ldots, s^{[t,4]}\big)\): tuple of the random variables of all state factors at time \(t\)
\(u^{[t,0]}\): categorical random variable for the action at time \(t\), i.e. the room the Roomba heads for: \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\). Heading for the room it is already in means staying put. The action at time \(t\) takes effect at time \(t+1\). Only factor 0 (Room) is controlled; the other factors change on their own.
Fonts. Bold marks anything with structure, and the style of bold tells you which kind:
upright bold (\(\mathbf{s}\), \(\mathbf{A}\)): a tuple, an ordered collection whose entries may differ in size, such as one entry per state factor or observation modality
italic bold (\(\boldsymbol{s}\), \(\boldsymbol{A}\)): a single vector, matrix or tensor, such as a one-hot vector, a probability vector, a belief, or a likelihood tensor
For example, \(\mathbf{A} = \big(\boldsymbol{A}^{[0]}, \ldots, \boldsymbol{A}^{[M]}\big)\) is the tuple of likelihood tensors, and \(\boldsymbol{A}^{[m]}\) is the tensor for observation modality \(m\). Likewise, \(\mathbf{B} = \big(\boldsymbol{B}^{[0]}, \ldots, \boldsymbol{B}^{[N]}\big)\) and \(\mathbf{D} = \big(\boldsymbol{D}^{[0]}, \ldots, \boldsymbol{D}^{[N]}\big)\) collect one array per state factor, and \(\mathbf{C} = \big(\boldsymbol{C}^{[0]}, \ldots, \boldsymbol{C}^{[M]}\big)\) collects one preference vector per observation modality. Upright bold is only used for Latin letters, because Greek letters such as \(\pi\) do not render in upright bold.
Dependencies. With several state factors, each tensor depends on only some of them. We write \(\mathcal{D}_A^{[m]}\) for the set of factors that modality \(m\) depends on, and \(\mathcal{D}_B^{[n]}\) for the set of factors whose previous values the transition of factor \(n\) depends on (always including factor \(n\) itself). Each tensor has one axis for its own variable, followed by one axis per factor it depends on, and, for the controlled factor, one axis for the action:
tensor
depends on
shape
axes
\(\boldsymbol{A}^{[0]}\) (Texture)
\(\mathcal{D}_A^{[0]} = \{0\}\)
\(3 \times 3\)
texture, room
\(\boldsymbol{A}^{[1]}\) (Motion)
\(\mathcal{D}_A^{[1]} = \{0, 1\}\)
\(2 \times 3 \times 4\)
motion, room, person
\(\boldsymbol{A}^{[2]}\) (Dirt)
\(\mathcal{D}_A^{[2]} = \{0, 2, 3, 4\}\)
\(2 \times 3 \times 2 \times 2 \times 2\)
dirt reading, room, dirt in each room
\(\boldsymbol{B}^{[0]}\) (Room)
\(\mathcal{D}_B^{[0]} = \{0\}\), and the action
\(3 \times 3 \times 3\)
next room, previous room, action
\(\boldsymbol{B}^{[1]}\) (Person)
\(\mathcal{D}_B^{[1]} = \{1\}\)
\(4 \times 4\)
next location, previous location
\(\boldsymbol{B}^{[n]}\) (Dirt, \(n = 2, 3, 4\))
\(\mathcal{D}_B^{[n]} = \{n, 0\}\)
\(2 \times 2 \times 3\)
next dirt, previous dirt, previous room
For example, \(\boldsymbol{A}^{[1]}[\bullet, r, p]\) is the distribution over the motion reading when the Roomba is in room \(r\) and the person in location \(p\). The dirt sensor depends on all three dirt factors because which room’s dirt it senses depends on where the Roomba is.
Marks. A hat marks a realized value: something that has actually happened. A bold symbol without a hat is a distribution: a probability vector over possible values.
\(\hat{\boldsymbol{s}}^{[t,n]}\): realized value of \(s^{[t,n]}\), one-hot encoded as a vector, e.g. \([0,1,0]^{\top}\) for the Roomba being in the living room
\(\hat{\mathbf{s}}^{(t)} = \big(\hat{\boldsymbol{s}}^{[t,0]}, \ldots, \hat{\boldsymbol{s}}^{[t,4]}\big)\): tuple of the realized values of all state factors at time \(t\)
\(\hat{\mathbf{s}}^{(0:T)} = \big(\hat{\mathbf{s}}^{(0)}, \ldots, \hat{\mathbf{s}}^{(T)}\big)\): trajectory (sequence over time) of realized states
\(\hat{\boldsymbol{u}}^{[t,0]}\): the action the agent actually chose at time \(t\), one-hot encoded, e.g. \([0,0,1]^{\top}\) for heading to the bathroom; \(\hat{\mathbf{u}}^{(t)} = \big(\hat{\boldsymbol{u}}^{[t,0]}\big)\) is the tuple of all actions at time \(t\)
\(\boldsymbol{s}^{[t,n]}\), without a hat: the agent’s belief about state factor \(n\) at time \(t\), a probability vector with one entry per value of \(s^{[t,n]}\)
\(\boldsymbol{B}^{\ast[0]}[\bullet, \bullet, u]\,\hat{\boldsymbol{s}}^{\ast[t-1,0]}\): a probability vector (here, over the next room, given action \(u\)), not a realized value
\(\bar{\boldsymbol{A}}^{[m]}\), with a bar: a posterior mean, the agent’s point estimate of a learned matrix after inference (used in Parts 2 and 3; in this part, the agent’s tensors are given)
\(\tilde{\boldsymbol{s}}^{[t,n]}\), with a tilde: the exact reference belief about factor \(n\) at time \(t\), computed by the exact filter of Part 5 on the same observations and actions. It is used only for evaluation; the agent never sees it.
\(\tilde{u}^{[t,0]}\): the action the exact agent would have chosen at time \(t\), from the exact reference belief. It is used only for evaluation; the agent never carries it out.
Factorized beliefs. In Part 5, the agent held a joint belief over all combinations of the five factors, \(3 \cdot 4 \cdot 2 \cdot 2 \cdot 2 = 96\) in total. In this part, it holds one belief per factor, \(\boldsymbol{s}^{[t,0]}, \ldots, \boldsymbol{s}^{[t,4]}\), and treats their product as its belief about the whole state:
This is the mean-field approximation. Written as an array with one axis per factor, the belief it implies about the combinations is the outer product \(\boldsymbol{s}^{[t,0]} \otimes \cdots \otimes \boldsymbol{s}^{[t,4]}\), whose entry for a combination is the product of the factors’ probabilities for their values in it. Such a product cannot represent relations between factors: it cannot express “the Roomba and the person are probably in the same room”, only where each of them probably is on its own. The beliefs \(\boldsymbol{s}^{[t,n]}\) are no longer marginals of something larger; they are the agent’s belief.
The exact reference keeps the Part 5 picture: its joint belief \(\tilde{\boldsymbol{s}}^{(t)}\), over the 96 combinations, and its marginals \(\tilde{\boldsymbol{s}}^{[t,n]}\), obtained by adding up the joint belief. Comparing \(\boldsymbol{s}^{[t,n]}\) with \(\tilde{\boldsymbol{s}}^{[t,n]}\) shows the effect of the approximation, factor by factor.
The agent’s beliefs are computed by a fixed-point iteration: at each tick, the agent starts from its predicted beliefs, \(\boldsymbol{s}^{[t,n]}_{\langle 0 \rangle}\), and updates the beliefs about the factors in turn, each taking the current beliefs about the other factors into account. After \(K\) iterations, it takes \(\boldsymbol{s}^{[t,n]} = \boldsymbol{s}^{[t,n]}_{\langle K \rangle}\). Each iteration lowers the variational free energy\(F^{(t)}_{\langle k \rangle}\), a measure of how well the factorized belief explains the prediction and the observations at time \(t\); \(\boldsymbol{F}^{(t)}\) collects its values over the iterations.
The agent never observes the states, so its state variables never carry a hat: it has random variables \(s^{[t,n]}\) and beliefs \(\boldsymbol{s}^{[t,n]}\). The observations are different, because the agent does receive them. Each observation therefore appears in three forms, which share the same index but are different objects:
\(o^{[t,0]}\): the categorical random variable for the Texture observation at time \(t\), with values \(\{\text{hardwood}, \text{carpet}, \text{tiles}\}\)
\(\hat{\boldsymbol{o}}^{[t,0]}\): the observation the agent actually received, one-hot encoded. If the camera reads hardwood at time \(t\), then \(\hat{\boldsymbol{o}}^{[t,0]} = [1,0,0]^{\top}\). If at the same time the motion sensor detects motion and the dirt sensor picks up nothing, then \(\hat{\boldsymbol{o}}^{[t,1]} = [0,1]^{\top}\) and \(\hat{\boldsymbol{o}}^{[t,2]} = [1,0]^{\top}\), and the tuple of all modalities is \(\hat{\mathbf{o}}^{(t)} = \big([1,0,0]^{\top}, [0,1]^{\top}, [1,0]^{\top}\big)\).
\(\boldsymbol{o}^{[t,0]}\), without a hat: an expected observation, i.e. a probability vector the agent computes from its beliefs
For a modality that depends on a single factor, the expected observation is a matrix-vector product, as before. Suppose the agent believes the Roomba is most likely in the bedroom, \(\boldsymbol{s}^{[t,0]} = [0.7, 0.2, 0.1]^{\top}\), and its texture matrix has the same values as \(\boldsymbol{A}^{\ast[0]}\). Then it expects
that is, hardwood with probability 0.645, carpet 0.22, and tiles 0.135. For a modality that depends on several factors, the tensor is contracted with the beliefs about all of them. The motion sensor detects motion with probability 0.8 if the person is in the Roomba’s room, and 0.05 otherwise. Suppose the agent believes the person is in the living room with probability 0.5, away with 0.3, and in the bedroom and bathroom with 0.1 each. Treating the two factors as independent, the probability that the person is in the same room as the Roomba is \(0.7 \cdot 0.1 + 0.2 \cdot 0.5 + 0.1 \cdot 0.1 = 0.18\), so the agent expects motion with probability \(0.8 \cdot 0.18 + 0.05 \cdot 0.82 \approx 0.185\). In Part 5, the agent used its joint belief for this, which also accounted for any relation between the two factors. In this part, treating the factors as independent is exactly what the agent does: this example is the mean-field computation.
So, for observations the agent has both hatted and unhatted forms: what it expects (\(\boldsymbol{o}\)) and what it gets (\(\hat{\boldsymbol{o}}\)). For states, it only ever has the unhatted forms. Only the environment has \(\hat{\boldsymbol{s}}^{\ast[t,n]}\): where the Roomba and the person really are, and which rooms are really dirty.
Actions, preferences and planning. With control, the agent needs three more kinds of objects:
\(\boldsymbol{B}^{[0]}[\bullet, \bullet, u]\): the transition matrix of the controlled factor under action\(u\). The action also matters for the dirt factors, but only indirectly: it determines where the Roomba goes, and that determines which room gets cleaned.
\(\boldsymbol{C}^{[m]}\): the agent’s preference over the observations of modality \(m\), a probability vector expressing which observations it prefers to receive. As in Part 5, two modalities have a real preference: Motion, for no motion, e.g. \(\boldsymbol{C}^{[1]} = [0.9, 0.1]^{\top}\), and Dirt, for dirt picked up, e.g. \(\boldsymbol{C}^{[2]} = [0.2, 0.8]^{\top}\). Texture has a flat preference.
planning symbols, following Algorithm 18 in Chapter 9:
\(H\): the planning horizon, i.e. how many steps the agent looks ahead; \(\tau = 0, \ldots, H-1\) counts the imagined steps, where step \(\tau\) predicts time \(t + \tau + 1\)
\(\boldsymbol{V}\): the matrix of candidate policies, i.e. action sequences. Row \(p\), \(\pi^{[p]} = \boldsymbol{V}[p, \bullet]\), is policy \(p\), and \(\boldsymbol{V}[p, \tau]\) is the action it prescribes at step \(\tau\).
\(\boldsymbol{s}_{\pi^{[p]}}^{[\tau,n]}\) and \(\boldsymbol{o}_{\pi^{[p]}}^{[\tau,m]}\): the predicted belief about factor \(n\) and predicted observation at step \(\tau\) under policy \(p\), unhatted because they are predictions. As the beliefs themselves, the predicted beliefs are factorized: one per factor, with no joint.
\(G^{[p,\tau]}\): the expected free energy of policy \(p\) at step \(\tau\), and \(G^{[p]} = \sum_{\tau} G^{[p,\tau]}\) its total
\(\pi^{(t)}\): the random variable for the policy chosen at time \(t\), with posterior \(Q\big(\pi^{(t)} = \pi^{[p]}\big) = \sigma\big(-\boldsymbol{G}\big)_p\), where \(\boldsymbol{G}\) collects the \(G^{[p]}\) and \(\sigma\) is the softmax
\(Q\big(u^{[t,0]}\big)\): the resulting distribution over the next action, obtained by adding up \(Q\big(\pi^{(t)} = \pi^{[p]}\big)\) over all policies whose first action is \(u\)
The variational free energy \(F^{(t)}\), which the agent minimizes in perception, and the expected free energy \(G^{[p]}\), which it minimizes in planning, are different quantities: the first measures how well the agent’s current beliefs explain what it has seen, the second how well a policy is expected to lead to preferred and informative observations.
The asterisk\({}^{\ast}\) marks what belongs to the environment (the generative process), e.g. \(\hat{\mathbf{s}}^{\ast(t)}\) and \(\boldsymbol{B}^{\ast[0]}\). Symbols without it belong to the agent (the generative model). In this part, the agent is given its tensors: \(\mathbf{A} = \mathbf{A}^{\ast}\) and \(\mathbf{B} = \mathbf{B}^{\ast}\). The preferences \(\mathbf{C}\) belong to the agent alone; the environment has none. The observation \(\hat{\mathbf{o}}^{(t)}\) has no asterisk because it is shared: the environment generates it and the agent receives it. The same holds for the action \(\hat{\mathbf{u}}^{(t)}\): the agent chooses it and the environment carries it out.
Code symbols. The code mirrors the math as closely as identifiers allow:
L prefix: upright bold (\(\mathbf{s}\), \(\mathbf{A}\)), a tuple. The L suggests verticality (upright).
_ prefix: italic bold (\(\boldsymbol{s}\), \(\boldsymbol{A}\)), a single vector, matrix or tensor that is not part of a tuple. The _ suggests the underline of bold.
hat: a Unicode circumflex on the letter, e.g. ŝ for \(\hat{s}\), ô for \(\hat{o}\), û for \(\hat{u}\)
bar: a combining macron on the letter (typed \bar then Tab), e.g. _Ā⁰ for \(\bar{\boldsymbol{A}}^{[0]}\)
tilde: a combining tilde on the letter (typed \tilde then Tab), e.g. Ls̃ₜ for the exact reference beliefs \(\big(\tilde{\boldsymbol{s}}^{[t,0]}, \ldots, \tilde{\boldsymbol{s}}^{[t,4]}\big)\). For \(\tilde{u}\), a precomposed character exists, ũ, as û does for the hat, e.g. ũₜ for \(\tilde{u}^{[t,0]}\)
asterisk \({}^{\ast}\): a superscript ˣ (typed \^x then Tab)
time index \((t)\): a subscript in the name: ₜ for \(t\), ₜ₋₁ for \(t-1\), ₀ for \(t=0\)
iteration index \(\langle k \rangle\): an ordinary loop variable k; the beliefs are updated in place, so only the final iteration is kept, apart from the free energies in _Fₜ
trajectory \((0{:}T)\): a ː suffix (triangular colon), e.g. Lôː for \(\hat{\mathbf{o}}^{(0:T)}\)
tuple index \([n]\) or \([m]\): real 0-based indexing, e.g. LAˣ[0]. In Julia, tuples over factors or modalities are built with tuple0(a, b, ...), which wraps each argument as one entry. In Python, ordinary tuples and lists are already 0-based.
trajectory index \(t\): real 0-based indexing too, e.g. Lŝˣː[t]. In Julia, a trajectory is an existing vector re-indexed from 0 with offset0(v), which shifts the vector’s own indices instead of wrapping it.
Indexing and the tear-off mnemonic. Only tuples have names; their entries are written by indexing. Read a factor or modality index ([n] or [m]) as tearing the vertical line off the L, turning it into _: LAˣ[0] reads as _Aˣ[0], i.e. italic bold \(\boldsymbol{A}^{\ast[0]}\), just as indexing into a tuple turns the upright \(\mathbf{A}^{\ast}\) into the italic \(\boldsymbol{A}^{\ast[0]}\) in the math. Three details:
A time index into a trajectory does not tear off the line: Lŝˣː[t] is still a tuple, \(\hat{\mathbf{s}}^{\ast(t)}\). Only the factor index that follows tears it off: Lŝˣː[t][1] is \(\hat{\boldsymbol{s}}^{\ast[t,1]}\), the person’s location.
The mnemonic applies wherever the brackets appear, on either side of an assignment: LAˣ[0] = … assigns to \(\boldsymbol{A}^{\ast[0]}\).
Each further index changes the font of the result as it would in the math: LAˣ[1][2, 1, 1] is a single number, an entry of the tensor \(\boldsymbol{A}^{\ast[1]}\), so it is no longer bold at all. Likewise, LBˣ[0][:, :, u] is the transition matrix of the Room factor for action u, still italic bold.
Since each tuple has exactly one name, there is nothing to keep in sync: when a tuple gets a new value, a single assignment updates it, e.g. Lŝˣₜ = tuple0(...).
As in Parts 4 and 5, this part does not use RxInfer: the agent’s computations are written directly in Julia.
chosen action tuple at \(t\), and its entry (one-hot over target rooms)
\(\hat{\mathbf{u}}^{(0:T-1)}\)
Lûː
trajectory of chosen actions
\(\boldsymbol{s}^{[t,n]}\)
Lsₜ[n]
agent’s belief about factor \(n\) at \(t\), one factor of its factorized belief
\(\boldsymbol{s}^{[0:T,\bullet]}\)
Lsː
trajectory of the agent’s beliefs
\(K\), \(k\)
K, k
number of fixed-point iterations per tick, and the current iteration
\(\boldsymbol{F}^{(t)}\)
_Fₜ
the variational free energies \(F^{(t)}_{\langle k \rangle}\) at \(t\), one per iteration
\(\tilde{\boldsymbol{s}}^{(t)}\)
_s̃ₜ
exact reference: joint belief over the 96 combinations of all factors
\(\tilde{\boldsymbol{s}}^{[t,n]}\)
Ls̃ₜ[n]
exact reference: marginal belief about factor \(n\) at \(t\)
\(\tilde{\boldsymbol{s}}^{[0:T,\bullet]}\)
Ls̃ː
trajectory of the exact reference beliefs
\(\tilde{u}^{[t,0]}\)
ũₜ
exact reference: the action the exact agent would have chosen at \(t\)
\(H\), \(\boldsymbol{V}\)
H, _V
planning horizon, matrix of candidate policies
\(\boldsymbol{G}\), \(Q\big(\pi^{(t)}\big)\)
_G, _qπ
expected free energies of the policies, and the posterior over policies
\(t\), \(T\)
t, T
time step, last time step
The code stores the agent’s beliefs, not its random variables. So the tuple Lsₜ holds the beliefs \(\big(\boldsymbol{s}^{[t,0]}, \ldots, \boldsymbol{s}^{[t,4]}\big)\), while in the math, \(\mathbf{s}^{(t)}\) denotes the tuple of random variables. Unlike in Part 5, the agent has no joint belief: the only joint belief in this part is the exact reference _s̃ₜ, a single array, not a tuple, which the font, and the prefix, make visible.
In Python, ˣ is normalized to a plain x, so LAˣ and LAx are the same name; never use LAx for anything else.
Changes in the symbol conventions compared to Part 5
A third kind of index bracket,\(\langle k \rangle\), for the iterations of the fixed-point iteration within a tick, with \(k = 1, \ldots, K\).
Factorized beliefs replace the joint belief: the agent holds one belief \(\boldsymbol{s}^{[t,n]}\) per factor and treats their product as its belief about the whole state (the mean-field approximation). The beliefs are no longer marginals; they are the belief itself.
A tilde marks the exact reference:\(\tilde{\boldsymbol{s}}^{(t)}\) is the joint belief of the exact filter of Part 5, run on the same data for evaluation, and \(\tilde{\boldsymbol{s}}^{[t,n]}\) are its marginals. In code, a combining tilde: _s̃ₜ, Ls̃ₜ, Ls̃ː.
The variational free energy\(F^{(t)}_{\langle k \rangle}\) returns, within each tick, collected in \(\boldsymbol{F}^{(t)}\) (_Fₜ). It is distinguished from the expected free energy \(G^{[p]}\) used in planning.
The motion example of treating two factors as independent is now exactly the agent’s own computation.
Predicted beliefs in planning are per factor,\(\boldsymbol{s}_{\pi^{[p]}}^{[\tau,n]}\), instead of a predicted joint belief.
The code table drops the agent’s joint belief _sₜ, and adds K, k, _Fₜ, and the reference beliefs _s̃ₜ, Ls̃ₜ and Ls̃ː.
A tilde also marks the exact agent’s would-be action,\(\tilde{u}^{[t,0]}\) (ũₜ), used for the decision agreement.
4.2 Core Elements
This section attempts to answer three important questions:
What metrics are we going to track?
What decisions do we intend to make?
What are the sources of uncertainty?
We track two groups of metrics. The first, as in Part 5, measures how well the Roomba does its job:
disturbance: the fraction of time steps at which the Roomba is in the same room as the person. Lower is better.
cleanliness: the fraction of time steps at which each room is dirty, and its average over the three rooms. Lower is better.
tracking: the accuracy of the agent’s beliefs about each factor, i.e. the fraction of time steps at which its most probable value equals the true one, for the Roomba’s room, the person’s location, and the dirt in each room. In particular, we check how well the agent knows whether the person is in its room, and whether the room it is in is dirty.
The second group is new, and measures the approximation itself, by comparing the factorized agent with the exact reference:
belief discrepancy: how far the agent’s belief about each factor, \(\boldsymbol{s}^{[t,n]}\), is from the exact reference belief \(\tilde{\boldsymbol{s}}^{[t,n]}\), computed on the same observations and actions. We use the total variation distance, half the sum of the absolute differences between the two probability vectors: 0 if they are equal, 1 if they put all their probability on different values. Lower is better.
decision agreement: the fraction of time steps at which the exact agent, given the exact reference belief, would have chosen the same action as the factorized agent. Higher is better. A belief discrepancy only matters if it changes what the agent does.
runtime: the time per tick, for the factorized agent and the exact agent. This is what the approximation is meant to save, although with only 96 joint states, the saving is expected to be modest.
Since only the agent’s inference changes, the main comparison is between two agents with the same preferences, no motion and dirt picked up: the exact agent of Part 5 and the factorized agent of this part, each running in the same environment with the same seed. The random Roomba of 4.3 remains as a reference point. Finally, a stress test repeats the comparison with a less reliable camera, to see whether the approximation breaks down when the Roomba’s room becomes uncertain.
Decisions. The decisions are the same as in Part 5. At each time step, the agent chooses where the Roomba goes next: it stays in its current room, or heads for one of the other rooms. It chooses by looking three steps ahead and scoring each candidate sequence of moves by its expected free energy, which favors moves that are expected to lead to no motion and to dirt being picked up, as well as moves that reduce the agent’s uncertainty about the state. The two preferences can pull in opposite directions: the living room gets dirty fastest, but it is also the room the person uses most. What changes is only the basis of the decision: the factorized belief instead of the joint belief, both in the current belief and in the predictions under each policy.
There are two sources of uncertainty in the environment itself. The first has to do with the fact that the state transitions are not deterministic but rather stochastic, and each factor has its own dynamics:
Room: a move does not always succeed, and since there is no door between the bedroom and the living room, the Roomba has to pass through the bathroom to get from one to the other.
Person: the person moves between the rooms and sometimes leaves the apartment, independently of the Roomba, spending most time in the living room.
Dirt: each clean room gets dirty with a probability per tick that differs per room, the living room fastest; a dirty room only gets clean when the Roomba is in it, and even then cleaning may take more than one tick.
This is captured in the tuple of state transition tensors \(\mathbf{B}^{\ast} = \big(\boldsymbol{B}^{\ast[0]}, \ldots, \boldsymbol{B}^{\ast[4]}\big)\), with one tensor per state factor, each depending only on the factors listed in 4.1.
The second relates to observations, and all three sensors contribute to it. The Roomba’s camera sometimes misidentifies the floor surface in a room. The motion sensor sometimes misses the person when they are sitting still, and occasionally detects motion when they are not in the room. The dirt sensor sometimes misses dirt, and occasionally reports dirt in a clean room. This uncertainty is captured in the tuple of observation likelihood tensors \(\mathbf{A}^{\ast} = \big(\boldsymbol{A}^{\ast[0]}, \boldsymbol{A}^{\ast[1]}, \boldsymbol{A}^{\ast[2]}\big)\), for Texture, Motion and Dirt. In the stress test, the camera’s mistakes become more frequent.
The initial state is not a source of uncertainty in the environment, because the run always starts from the same state: the Roomba in the bedroom, the person in the living room, and all rooms clean.
There is uncertainty on the agent’s side too. The agent is again given its model of the world, \(\mathbf{A} = \mathbf{A}^{\ast}\) and \(\mathbf{B} = \mathbf{B}^{\ast}\), so it is not uncertain about how the environment works. It is uncertain about the state: where the Roomba is, where the person is, and which rooms are dirty, both now and in the future. It does not know the starting state either, so it begins with a uniform prior for each factor. The dirt in rooms the Roomba is not in is never observed: the agent can only predict it, and its uncertainty grows the longer it stays away. Reducing this uncertainty is part of what the agent’s choices aim at: going back to check a room has value in itself, besides the chance of finding dirt to clean.
New in this part is a source of error on the agent’s side: the approximation. The factorized belief cannot represent relations between factors, so after an observation that involves several factors, such as a motion reading, its beliefs can differ from the exact ones. This error does not come from the environment, and it would not disappear with better sensors; it comes from the form of the belief itself. A known property of the mean-field approximation is that it tends to be overconfident: forced to describe a coupled situation with independent factors, it typically commits to one explanation rather than spreading its belief over several. The belief discrepancy and the stress test are designed to measure how large this error is, and when it matters.
Changes in the core elements compared to Part 5
A second group of metrics measures the approximation: the belief discrepancy between the agent’s beliefs and the exact reference (total variation distance), the agreement between the factorized agent’s decisions and those the exact agent would have made, and the runtime per tick.
The main comparison is between two agents with the same preferences: the exact agent of Part 5 and the factorized agent of this part. The random Roomba remains as a reference, and a stress test with a less reliable camera is added.
The decisions are unchanged, but are based on the factorized belief, both now and in the predictions under each policy.
A new source of error on the agent’s side: the approximation itself, which cannot represent relations between factors and tends to make the agent overconfident.
4.3 System-Under-Steer / Environment / Generative Process
First, before letting the agent take control of the Roomba, we describe the environment it acts in, i.e. the generative process that produces the data. The environment is the same as in Part 5: its state has five aspects, each described by its own state factor:
factor 0 (Room): \(s^{\ast[t,0]} \in \{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\), where the Roomba is
factor 1 (Person): \(s^{\ast[t,1]} \in \{\text{bedroom}, \text{livingroom}, \text{bathroom}, \text{away}\}\), where the person is
factors 2, 3 and 4 (Dirt in the bedroom, living room and bathroom): \(s^{\ast[t,n]} \in \{\text{clean}, \text{dirty}\}\), \(n = 2, 3, 4\)
At time \(t\), the true state is the realized state tuple \[
\hat{\mathbf{s}}^{\ast(t)} = \big(\hat{\boldsymbol{s}}^{\ast[t,0]}, \hat{\boldsymbol{s}}^{\ast[t,1]}, \hat{\boldsymbol{s}}^{\ast[t,2]}, \hat{\boldsymbol{s}}^{\ast[t,3]}, \hat{\boldsymbol{s}}^{\ast[t,4]}\big)
\] whose entries are the one-hot encoded values of the five factors. The entries differ in size, 3, 4 and 2 values, which is why the state is a tuple rather than a single vector.
In Parts 3 and 4, occupancy was a property of the Roomba’s room, combined with the room into a single factor. Since Part 5, the model describes where the person is instead. The Roomba’s room is then occupied exactly when the Roomba and the person are in the same room.
The initial state is given by the tuple of initial state distributions \(\mathbf{D}^{\ast} = \big(\boldsymbol{D}^{\ast[0]}, \ldots, \boldsymbol{D}^{\ast[4]}\big)\). In our simulation, the run always starts from the same state: the Roomba in the bedroom, the person in the living room, and all rooms clean, so each \(\boldsymbol{D}^{\ast[n]}\) is a one-hot vector.
Actions. As in Parts 4 and 5, at each time step the agent chooses an action, the room the Roomba heads for,
and the environment carries it out. Heading for the room it is already in means staying put. The action chosen at time \(t-1\) determines how the Roomba moves from \(t-1\) to \(t\).
Transitions, one per factor. The true transition probabilities are captured by the tuple \(\mathbf{B}^{\ast} = \big(\boldsymbol{B}^{\ast[0]}, \ldots, \boldsymbol{B}^{\ast[4]}\big)\), with one tensor per factor. Given the previous state and action, each factor changes independently of the others, and each depends only on a few factors:
Room, \(\boldsymbol{B}^{\ast[0]} \in [0,1]^{3\times 3\times 3}\) (next room, previous room, action): if the Roomba heads for the room it is in, it stays; if it heads for an adjacent room, it gets there with probability \(p_{\text{move}}\), and otherwise stays where it is. There is no door between the bedroom and the living room, so when it heads from one to the other, it first moves to the bathroom.
Person, \(\boldsymbol{B}^{\ast[1]} \in [0,1]^{4\times 4}\) (next location, previous location): the person moves between the rooms and sometimes leaves the apartment, independently of the Roomba. The living room, with its abundance of natural light, is where the person spends most time; bathroom visits are short.
Dirt in room \(n\), \(\boldsymbol{B}^{\ast[n]} \in [0,1]^{2\times 2\times 3}\) (next dirt, previous dirt, previous room of the Roomba), for \(n = 2, 3, 4\): if the Roomba was in that room, a dirty room becomes clean with probability \(p_{\text{clean}}\); otherwise, a clean room becomes dirty with probability \(p_{\text{dirty}}(n)\), which is highest for the living room, and a dirty room stays dirty.
The dirt factors are where the dependencies between factors appear in the transitions: a room’s dirt depends on its own previous value and on where the Roomba was. The action influences the dirt only indirectly, through where it takes the Roomba.
Our Roomba is equipped with three sensors. The first is a small camera that tracks the surface it is moving over, since there are
hardwood floors in the bedroom
carpet in the living room, and
tiles in the bathroom.
However, this method is not foolproof, and sometimes the Roomba will mistake one floor type for another. Don’t be too hard on the little guy, it’s just a Roomba after all.
The second is a motion sensor that reports whether it detects movement. It is the most direct evidence that the person is in the Roomba’s room: it usually detects motion when they are, but sometimes misses them when they are sitting still, and occasionally detects motion when they are not.
The third, added in Part 5, is a dirt sensor that reports whether the Roomba’s brushes are picking up dirt. It usually does so when the room it is in is dirty, but sometimes misses dirt, and occasionally reports dirt in a clean room.
At time \(t\), the received observation tuple is \(\hat{\mathbf{o}}^{(t)} = \big(\hat{\boldsymbol{o}}^{[t,0]}, \hat{\boldsymbol{o}}^{[t,1]}, \hat{\boldsymbol{o}}^{[t,2]}\big)\), whose entries are the one-hot encoded values of three observation modalities:
Observations, each depending on a few factors. The true mapping from the state to the observations is \(\mathbf{A}^{\ast} = \big(\boldsymbol{A}^{\ast[0]}, \boldsymbol{A}^{\ast[1]}, \boldsymbol{A}^{\ast[2]}\big)\), with one likelihood tensor per modality:
Texture, \(\boldsymbol{A}^{\ast[0]} \in [0,1]^{3\times 3}\), depends only on the Roomba’s room.
Motion, \(\boldsymbol{A}^{\ast[1]} \in [0,1]^{2\times 3\times 4}\), depends on the Roomba’s room and the person’s location: what matters is whether the two are the same.
Dirt, \(\boldsymbol{A}^{\ast[2]} \in [0,1]^{2\times 3\times 2\times 2\times 2}\), depends on the Roomba’s room and the dirt in all three rooms: the room tells which room’s dirt the sensor senses.
The tensors also encode the probability that a sensor gives a misleading reading. Observations do not depend on the action. This leaves us with the following generative process specification, where a realized value without bold, such as \(\hat{s}^{\ast[t-1,0]}\), is used as an index into a tensor:
Each line picks out, from a tensor, the probability vector for the current values of exactly the factors it depends on, as listed in 4.1. The specification makes the structure of the environment visible at a glance: each factor and each sensor looks at only a few others.
With actions and several state factors, this process has the structure of a factorized partially observable Markov decision process: five hidden Markov chains, one per factor, coupled through the dependencies in their transitions, with the Room chain driven by the agent’s actions, and three sensors emitting noisy observations of them. The agent never sees the true state directly. As in Part 5, it is given its model of the environment, \(\mathbf{A} = \mathbf{A}^{\ast}\) and \(\mathbf{B} = \mathbf{B}^{\ast}\), and its goal is to choose actions that keep the rooms clean while staying out of the person’s way.
A factorized process, but not factorized beliefs. Since the environment is built from separate factors, it may seem natural that the agent’s beliefs can be kept separate too. They cannot, exactly. The process is factorized: each factor’s next value is drawn independently, given the previous state. But the agent does not know the previous state, and the sensors look at several factors at once. When the motion sensor detects motion, the explanations “the Roomba is in the bedroom and the person too” and “the Roomba is in the living room and the person too” both fit, while mixed combinations do not. The exact posterior therefore couples the factors, even though the process does not. The factorized agent of this part ignores this coupling, which is exactly the approximation we evaluate.
Changes in the environment compared to Part 5
None: the factors, transitions, sensors and process specification are the same as in Part 5. Only the references to earlier parts are updated.
A closing paragraph explains why a factorized process does not give factorized beliefs: the agent does not know the state, and sensors that look at several factors couple them in its posterior. This is the coupling the factorized agent ignores.
## As in Part 5, the tuple has entries of different sizes: 3, 4 and 2 values,## which is what tuple0 is for, and why the state isn't a single vector.## factor 0 (Room): 1 bedroom 2 livingroom 3 bathroom## factor 1 (Person): 1 bedroom 2 livingroom 3 bathroom 4 away## factors 2–4 (Dirt in bedroom, livingroom, bathroom): 1 clean 2 dirtyLŝˣ₀ =tuple0( [1.0, 0.0, 0.0], ## ŝ^{*[0,0]}: Roomba in the bedroom [0.0, 1.0, 0.0, 0.0], ## ŝ^{*[0,1]}: person in the living room [1.0, 0.0], ## ŝ^{*[0,2]}: bedroom clean [1.0, 0.0], ## ŝ^{*[0,3]}: living room clean [1.0, 0.0], ## ŝ^{*[0,4]}: bathroom clean) ## ŝ^{*(0)}: the initial state tuple
5-element OffsetArray(::Vector{Vector{Float64}}, 0:4) with eltype Vector{Float64} with indices 0:4:
[1.0, 0.0, 0.0]
[0.0, 1.0, 0.0, 0.0]
[1.0, 0.0]
[1.0, 0.0]
[1.0, 0.0]
The state variables represent what we need to know. The true state at time \(t\) is given by a (compound) tuple \(\hat{\mathbf{s}}^{\ast(t)}\) whose components are the one-hot encoded state factors (here five factors):
one-hot encoded as \(\hat{\boldsymbol{s}}^{\ast[t,n]} \in \big\{[1,0]^{\top}, [0,1]^{\top}\big\}\)
The factors have different numbers of values, 3, 4 and 2, so their one-hot vectors differ in length. This is why the state is a tuple rather than a single vector. All combinations together give \(3 \cdot 4 \cdot 2 \cdot 2 \cdot 2 = 96\) possible states. At any time, the environment is in exactly one of them. The exact reference keeps a belief over all 96, while the factorized agent keeps one belief per factor, \(3 + 4 + 2 + 2 + 2 = 13\) numbers in total.
The observation variables represent what we can sense. Similarly, the (compound) observation tuple received by the Roomba at time \(t\) is given by
one-hot encoded as \(\hat{\boldsymbol{o}}^{[t,2]} \in \big\{[1,0]^{\top}, [0,1]^{\top}\big\}\)
Changes in the state and observation variables compared to Part 5
None in the variables themselves. A sentence contrasts the 96 combinations the exact reference keeps a belief over with the 13 numbers of the factorized agent.
4.3.2 Decision variables
The decision variables represent what we can control. As in Parts 4 and 5, the agent chooses an action at each time step. The action at time \(t\) is given by a (compound) tuple \(\hat{\mathbf{u}}^{(t)}\) whose components are the one-hot encoded control factors (here a single factor):
control factor 0 (Destination), acting on state factor 0 (Room):
\(u^{[t,0]} \in \{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\), the room the Roomba heads for,
one-hot encoded as \(\hat{\boldsymbol{u}}^{[t,0]} \in \big\{[1,0,0]^{\top}, [0,1,0]^{\top}, [0,0,1]^{\top}\big\}\)
Heading for the room the Roomba is already in means staying put. The action chosen at time \(t\) takes effect at time \(t+1\): it selects which of the transition matrices \(\boldsymbol{B}^{\ast[0]}[\bullet, \bullet, u]\) the environment uses for the Room factor in the next step.
Of the five state factors, the action controls only one, the Room, but its effects reach further. Through where it takes the Roomba, the action also determines which room gets cleaned: the dirt factors depend on the Roomba’s room, so the action influences them indirectly. The person’s location, on the other hand, is beyond the agent’s control: it can choose its room, but not its company.
The agent’s choice is therefore always a bet on two things at once. It heads for a room it expects to be quiet and in need of cleaning, and its sensors then tell it whether those expectations came true: the motion sensor whether the person is there, the dirt sensor whether there is dirt to pick up. The factorized agent makes the same bet, but evaluates it with separate beliefs about where the Roomba will be, where the person will be, and how dirty each room will be, combined as if they were independent.
Changes in the decision variables compared to Part 5
None in the variables themselves. A sentence notes that the factorized agent evaluates its choices with separate beliefs per factor, combined as if independent.
## actions (target room): 1 bedroom## 2 livingroom## 3 bathroom## heading for the room the Roomba is already in means staying putaction_labels = ["bedroom", "livingroom", "bathroom"]
Exogenous forces represent what changes on its own: the influences on the environment that the agent cannot control, arriving between its decisions. In an active inference model, they appear as the stochastic dynamics of the state factors, in so far as these dynamics do not depend on the agent’s actions.
Forces versus factors. A state factor and an exogenous force are not the same thing. In the sequential-decision framework, the state \(S_t\), the decision \(x_t\) and the exogenous information \(W_{t+1}\) are separate, and the next state follows from all three:
\[
S_{t+1} = S^M\big(S_t, x_t, W_{t+1}\big)
\]
The person’s location, for example, is a state; the person deciding to walk to the bathroom is the exogenous force acting on it. In an active inference model, such a force has no variable of its own: it is contained in the randomness of the transition tensors \(\mathbf{B}^{\ast}\).
Kinds of state factors. Not every factor the agent cannot control is driven by exogenous forces alone. With the factors of this part, four kinds can be distinguished:
kind of factor
example
exogenous?
controlled: the action selects its transition
Room
only the chance that a move fails
indirectly influenced: not controlled, but its transition depends on a controlled factor
Dirt (cleaning depends on where the Roomba is)
partly: dirt accumulates exogenously, but is removed by the agent
fully exogenous: its transition depends on nothing the agent does
Person
yes
static: never changes at all
the layout of the apartment
no force acts on it: it is a hidden fact, not a force
The exogenous forces in this part. As in Part 5, there are two:
The person’s movements. The person moves between the rooms and sometimes leaves the apartment, independently of the Roomba. The agent can neither influence nor directly observe this; it only notices the person through the motion sensor, when they are in the same room. These forces are contained in \(\boldsymbol{B}^{\ast[1]}\).
Dirt accumulating. Each clean room gets dirty with some probability per tick, the living room fastest. The agent cannot prevent this; it can only remove dirt by cleaning, i.e. by being in the room. These forces are contained in the growth of dirt in \(\boldsymbol{B}^{\ast[2]}\) to \(\boldsymbol{B}^{\ast[4]}\), while the removal of dirt is the agent’s doing.
Because the agent senses the person only through the motion sensor, which also depends on the Roomba’s room, the person’s movements are where the factorized beliefs of this part are most likely to differ from the exact ones.
In Parts 3 and 4, the changes in occupancy were an exogenous force as well, but they were folded into the combined RoomOccupancy factor. Since Part 5, with a separate factor for the person, they are explicit. Further exogenous forces could be added in later parts, for example doors that are closed at random.
Changes in the exogenous forces compared to Part 5
None in the forces themselves. A sentence notes that the person’s movements, sensed only together with the Roomba’s room, are where the factorized beliefs are most likely to differ from the exact ones.
4.3.4 State Transition and Observation Generation functions
Starting from \(\hat{\boldsymbol{s}}^{\ast[0,n]} \sim \mathcal{Cat}\big(\boldsymbol{D}^{\ast[n]}\big)\), \(n = 0, \ldots, 4\), with one-hot vectors for the Roomba in the bedroom, the person in the living room, and all rooms clean, the environment evolves according to
where \(\hat{u}^{[t-1,0]}\) is the action chosen at time \(t-1\). With one tensor per factor, each table below stays small: the 96 joint states never have to be written out.
State Transition model \(\boldsymbol{B}^{\ast[0]}\) (factor 0, Room)
One table per action \(u\). If the Roomba heads for the room it is in, it stays. If it heads for an adjacent room, it gets there with \(p_{\text{move}} = 0.9\), and otherwise stays where it is. There is no door between the bedroom and the living room, so when it heads from one to the other, it first moves to the bathroom. Each column is a probability distribution over the next room, given the previous room, so each column sums to 1.
Action: head for the bedroom
next room previous room
bedroom
livingroom
bathroom
bedroom
1.0
0.0
0.9
livingroom
0.0
0.1
0.0
bathroom
0.0
0.9
0.1
Action: head for the living room
next room previous room
bedroom
livingroom
bathroom
bedroom
0.1
0.0
0.0
livingroom
0.0
1.0
0.9
bathroom
0.9
0.0
0.1
Action: head for the bathroom
next room previous room
bedroom
livingroom
bathroom
bedroom
0.1
0.0
0.0
livingroom
0.0
0.1
0.0
bathroom
0.9
0.9
1.0
The bathroom is the hub of the apartment: from the bedroom, heading for the living room leads to the bathroom first, and vice versa.
State Transition model \(\boldsymbol{B}^{\ast[1]}\) (factor 1, Person)
The person moves independently of the Roomba and of the action. Each column is a probability distribution over the person’s next location, given the previous one, so each column sums to 1.
next previous
bedroom
livingroom
bathroom
away
bedroom
0.80
0.03
0.20
0.00
livingroom
0.13
0.90
0.30
0.05
bathroom
0.05
0.04
0.50
0.00
away
0.02
0.03
0.00
0.95
The living room, with its abundance of natural light, is where the person spends most time (staying with 0.90). Bathroom visits are short (0.50). When the person leaves the apartment, they mostly stay away for a while (0.95), and they come back through the living room.
State Transition model \(\boldsymbol{B}^{\ast[n]}\) (factors 2–4, Dirt)
Each dirt factor depends on its own previous value and on where the Roomba was. If the Roomba was in that room, a dirty room becomes clean with \(p_{\text{clean}} = 0.8\), and a clean room stays clean. If the Roomba was elsewhere, a clean room becomes dirty with \(p_{\text{dirty}}(n)\), and a dirty room stays dirty. Each column is a probability distribution over the next value, given the previous one, so each column sums to 1.
The Roomba was in this room (cleaning)
next previous
clean
dirty
clean
1.0
0.8
dirty
0.0
0.2
The Roomba was elsewhere (dirt accumulating)
next previous
clean
dirty
clean
\(1 - p_{\text{dirty}}(n)\)
0.0
dirty
\(p_{\text{dirty}}(n)\)
1.0
with a rate of dirt accumulation that differs per room:
bedroom
livingroom
bathroom
\(p_{\text{dirty}}\)
0.02
0.05
0.03
The living room, the room in most use, gets dirty fastest: on average after about 20 ticks, against 50 for the bedroom and about 33 for the bathroom. Since the Roomba’s room has three possible values, \(\boldsymbol{B}^{\ast[n]}\) consists of three \(2 \times 2\) matrices: the cleaning table for the value matching room \(n\), and the accumulating table for the other two.
Observation Generation model \(\boldsymbol{A}^{\ast[0]}\) (modality 0, Texture)
The texture depends only on the Roomba’s room. Each column is a probability distribution over the observed texture, so each column sums to 1.
observation \(o^{[t,0]}\) room
bedroom
livingroom
bathroom
hardwood
0.9
0.05
0.05
carpet
0.05
0.9
0.05
tiles
0.05
0.05
0.9
The off-diagonal entries are the camera’s mistakes: in each room, it misreads the floor with probability 0.1, split equally between the two other textures.
Stress test: a less reliable camera. For the stress test in 4.6, the camera misreads the floor with probability 0.4 instead of 0.1, again split equally between the two other textures:
observation \(o^{[t,0]}\) room
bedroom
livingroom
bathroom
hardwood
0.6
0.2
0.2
carpet
0.2
0.6
0.2
tiles
0.2
0.2
0.6
A reading now favors the matching room by a factor of 3 instead of 18, so a single reading no longer settles where the Roomba is. The agent still has its knowledge of where it was heading, but when a move may have failed and the camera is unsure, the Roomba’s room becomes genuinely uncertain. That is the situation in which the relations between factors matter most: a motion reading then says that the person is in the same room as the Roomba, whichever room that is, which a factorized belief cannot express. In the stress test, both the environment and the agent’s model use this camera, so the agent is not misinformed, only less informed; all other tensors stay the same.
Observation Generation model \(\boldsymbol{A}^{\ast[1]}\) (modality 1, Motion)
Motion depends on the Roomba’s room and the person’s location, but only through one question: are they in the same room? The table gives the probability of detecting motion; the probability of no motion is 1 minus each entry.
room person
bedroom
livingroom
bathroom
away
bedroom
0.8
0.05
0.05
0.05
livingroom
0.05
0.8
0.05
0.05
bathroom
0.05
0.05
0.8
0.05
The motion sensor detects the person with probability 0.8 when they are in the Roomba’s room; it misses them when they are sitting still with probability 0.2. Otherwise, it raises a false alarm with probability 0.05.
Observation Generation model \(\boldsymbol{A}^{\ast[2]}\) (modality 2, Dirt)
The dirt reading depends on the Roomba’s room and on the dirt in all three rooms, but again only through one question: is the room the Roomba is in dirty? Each column is a probability distribution over the dirt reading, so each column sums to 1.
observation \(o^{[t,2]}\) dirt in the Roomba’s room
clean
dirty
nothing picked up
0.95
0.1
dirt picked up
0.05
0.9
The dirt sensor picks up dirt in a dirty room with probability 0.9, and misses it with probability 0.1. In a clean room, it falsely reports dirt with probability 0.05. Although \(\boldsymbol{A}^{\ast[2]}\) has \(2 \times 3 \times 2 \times 2 \times 2 = 48\) entries, they are all filled from this small table: for each room, the column is selected by the dirt factor of that room.
Changes in the transition and observation functions compared to Part 5
None in the main model. All tables are the same as in Part 5.
New: a less reliable camera for the stress test, which misreads the floor with probability 0.4 instead of 0.1. It is used by both the environment and the agent in the stress test only, to make the Roomba’s room uncertain, where the factorized beliefs are expected to differ most from the exact ones.
## factor 0 (Room): 1 bedroom 2 livingroom 3 bathroom## factor 1 (Person): 1 bedroom 2 livingroom 3 bathroom 4 away## factors 2–4 (Dirt in bedroom, livingroom, bathroom): 1 clean 2 dirty## actions (target room): 1 bedroom 2 livingroom 3 bathroom## ---------- factor 0: Room (3×3×3: next room, previous room, action) ----------p_move =0.9## probability that a move succeedsnext_room(r, u) = r == u ? r : (r ==3|| u ==3) ? u :3## no door between bedroom 1 and livingroom 2_B_room =zeros(3, 3, 3)for u in1:3, r in1:3 r′ =next_room(r, u)if r′ == r _B_room[r, r, u] =1.0## heading for the current room: stayelse _B_room[r′, r, u] = p_move ## move succeeds _B_room[r, r, u] =1- p_move ## move fails: stayendend## ---------- factor 1: Person (4×4: next location, previous location) ----------_B_person = [0.800.030.200.00; ## to bedroom (columns: from bedroom, livingroom, bathroom, away)0.130.900.300.05; ## to livingroom0.050.040.500.00; ## to bathroom0.020.030.000.95] ## to away## ---------- factors 2–4: Dirt (2×2×3: next dirt, previous dirt, previous room of the Roomba) ----------p_clean =0.8## probability that a dirty room is cleaned, per tickp_dirty = [0.02, 0.05, 0.03] ## rate of dirt accumulation: bedroom, livingroom, bathroomfunctiondirt_tensor(room, p_dirty_room) _B =zeros(2, 2, 3)for r in1:3if r == room ## the Roomba was in this room: cleaning _B[:, :, r] = [1.0 p_clean;0.01- p_clean]else## the Roomba was elsewhere: dirt accumulating _B[:, :, r] = [1- p_dirty_room 0.0; p_dirty_room 1.0]endend _BendLBˣ =tuple0( _B_room, ## B^{*[0]}: Room _B_person, ## B^{*[1]}: Persondirt_tensor(1, p_dirty[1]), ## B^{*[2]}: Dirt in the bedroomdirt_tensor(2, p_dirty[2]), ## B^{*[3]}: Dirt in the living roomdirt_tensor(3, p_dirty[3]), ## B^{*[4]}: Dirt in the bathroom)## D_B^{[n]}: the factors whose previous values each transition depends on (own factor first)LdepsB =tuple0([0], [1], [2, 0], [3, 0], [4, 0])size.(LBˣ) ## (3,3,3), (4,4), (2,2,3), (2,2,3), (2,2,3)
5-element OffsetArray(::Vector{Tuple{Int64, Int64, Vararg{Int64}}}, 0:4) with eltype Tuple{Int64, Int64, Vararg{Int64}} with indices 0:4:
(3, 3, 3)
(4, 4)
(2, 2, 3)
(2, 2, 3)
(2, 2, 3)
Changes in the transition code compared to Part 5: none.
## factor 0 (Room): 1 bedroom 2 livingroom 3 bathroom## factor 1 (Person): 1 bedroom 2 livingroom 3 bathroom 4 away## factors 2–4 (Dirt in bedroom, livingroom, bathroom): 1 clean 2 dirty## ---------- modality 0: Texture (3×3: texture, room) ----------## the camera reads the right texture with probability 1 - p_misread,## and each of the two wrong textures with probability p_misread / 2texture_tensor(p_misread) = [i == j ? 1- p_misread : p_misread /2 for i in1:3, j in1:3]p_misread =0.1## the camera of Part 5p_misread_stress =0.4## the less reliable camera of the stress test_A_texture =texture_tensor(p_misread) ## rows hardwood, carpet, tiles; columns bedroom, livingroom, bathroom_A_texture_stress =texture_tensor(p_misread_stress)## ---------- modality 1: Motion (2×3×4: motion, room, person) ----------p_detect, p_false_motion =0.8, 0.05## motion if the person is in the Roomba's room / otherwise_A_motion =zeros(2, 3, 4)for r in1:3, p in1:4 p_motion = p == r ? p_detect : p_false_motion ## same room? (person 4 = away is never the same room) _A_motion[:, r, p] = [1- p_motion, p_motion] ## (no motion, motion)end## ---------- modality 2: Dirt (2×3×2×2×2: dirt reading, room, dirt in bedroom, livingroom, bathroom) ----------p_pickup, p_false_pickup =0.9, 0.05## dirt picked up in a dirty room / in a clean room_A_dirt =zeros(2, 3, 2, 2, 2)for r in1:3, d₁ in1:2, d₂ in1:2, d₃ in1:2 dirty = (d₁, d₂, d₃)[r] ==2## is the Roomba's room dirty? p_dirt = dirty ? p_pickup : p_false_pickup _A_dirt[:, r, d₁, d₂, d₃] = [1- p_dirt, p_dirt] ## (nothing picked up, dirt picked up)endLAˣ =tuple0( _A_texture, ## A^{*[0]}: Texture _A_motion, ## A^{*[1]}: Motion _A_dirt, ## A^{*[2]}: Dirt)## the stress test: the same sensors, except for a less reliable cameraLAˣ_stress =tuple0( _A_texture_stress, ## A^{*[0]}: Texture, misreading with probability 0.4 _A_motion, ## A^{*[1]}: Motion, unchanged _A_dirt, ## A^{*[2]}: Dirt, unchanged)## D_A^{[m]}: the factors each modality depends on, in the order of the tensor's axes after the firstLdepsA =tuple0([0], [0, 1], [0, 2, 3, 4])@assert LAˣ[0] ≈ [0.90.050.05; 0.050.90.05; 0.050.050.9] ## same camera as in Part 5size.(LAˣ) ## (3,3), (2,3,4), (2,3,2,2,2)
3-element OffsetArray(::Vector{Tuple{Int64, Int64, Vararg{Int64}}}, 0:2) with eltype Tuple{Int64, Int64, Vararg{Int64}} with indices 0:2:
(3, 3)
(2, 3, 4)
(2, 3, 2, 2, 2)
Changes in the observation code compared to Part 5
The camera is built from its misread probability with texture_tensor, so that the normal camera (p_misread = 0.1, the same as in Part 5) and the less reliable camera of the stress test (p_misread_stress = 0.4) come from the same function. A check confirms that the normal camera is unchanged.
New: LAˣ_stress, the tuple of likelihood tensors for the stress test, with the less reliable camera and the same Motion and Dirt tensors.
LAˣ[0] ## the Texture likelihood matrix, A^{*[0]} (3×3: texture, room)
LAˣ[2][2, :, 2, 2, 2] ## A^{*[2]}: probability of picking up dirt, by room, when all rooms are dirty
3-element Vector{Float64}:
0.9
0.9
0.9
LAˣ[2][2, :, 1, 1, 1] ## A^{*[2]}: probability of picking up dirt, by room, when all rooms are clean
3-element Vector{Float64}:
0.05
0.05
0.05
4.3.5 Objective function
The objective function in this section is the designer’s measure of performance: how well the system as a whole does, judged from the outside with knowledge of the true states \(\hat{\mathbf{s}}^{\ast(t)}\). It is not something the environment pursues, nor the quantity the agent minimizes internally; the agent’s own objectives, the variational free energy in perception and the expected free energy in planning, are described in 4.5.
As in Part 5, the agent has two goals, so there are two main measures of its behavior. The first is the disturbance: the fraction of ticks the Roomba spends in the same room as the person,
where \(\mathbb{1}[\cdot]\) is 1 if the condition holds and 0 otherwise. Disturbance is a condition on two factors: the Roomba and the person are in the same room. When the person is away, the condition never holds. Lower is better.
The second is the cleanliness: the fraction of ticks each room \(r\) is dirty, and its average over the three rooms,
where room \(r = 1, 2, 3\) (bedroom, living room, bathroom) corresponds to dirt factor \(n = r + 1\). Lower is better. The living room deserves particular attention: it gets dirty fastest, and it is the room the person uses most, so it is where the two goals conflict.
We also keep the tracking accuracies, one per state factor:
i.e. how often the most probable value under the agent’s belief \(\boldsymbol{s}^{[t,n]}\) equals the true value, for the Roomba’s room, the person’s location, and the dirt in each room. Higher is better, with a maximum of 1.
New in this part are three measures of the approximation itself. They compare the factorized agent with the exact reference, run on the same observations and actions, so that only the form of the belief differs. The first is the belief discrepancy: for each factor, the total variation distance between the agent’s belief and the exact reference belief, averaged over the run,
where \(i\) runs over the values of factor \(n\). A distance of 0 means the two beliefs are equal, 1 that they put all their probability on different values. Since an approximation can be good on average and still fail badly at a few moments, we also report the largest distance over the run, \(\max_t d^{[t,n]}\). Lower is better.
The second is the decision agreement: the fraction of ticks at which the exact agent, starting from the exact reference belief, would have chosen the same action as the factorized agent,
where \(\hat{u}^{[t,0]}\) is the action the factorized agent chose, and \(\tilde{u}^{[t,0]}\) the action the exact agent would have chosen in its place. Higher is better, with a maximum of 1. A belief discrepancy only matters if it changes what the agent does, and this measure shows whether it does.
The third is the belief entropy, which shows how the beliefs differ. The entropy of a belief measures how spread out it is: 0 for a belief that is certain, and \(\ln\) of the number of values for a uniform belief. For each factor, we compare the average entropy of the agent’s beliefs with that of the exact reference beliefs, at the same ticks:
A negative \(J_{\text{H}}(n)\) means that the agent’s beliefs about factor \(n\) are, on average, sharper than the exact ones: the agent is overconfident, as the mean-field approximation tends to be (see 4.2). A value near 0 means the approximation neither sharpens nor blurs the beliefs. Unlike the other measures, this one has no “better” direction in itself; it shows the character of the error.
Finally, the runtime per tick, for the factorized agent and the exact agent, shows what the approximation saves. With only 96 joint states, the saving is expected to be modest; the point of the approximation lies in larger models.
To judge these measures, we make two kinds of comparison:
on the same data: the belief discrepancy, the decision agreement and the belief entropy compare the factorized agent with the exact reference, fed the factorized agent’s own observations and actions. This isolates the approximation completely.
in closed loop: the disturbance, cleanliness and runtime compare the factorized agent with the exact agent of Part 5, each running on its own in the same environment with the same seed, with the random Roomba as a reference point. Here the agents’ paths diverge as soon as one of them makes a different choice, so this comparison shows the combined effect of all differences.
Both comparisons are made twice: with the camera of Part 5, and in the stress test, with the less reliable camera. We expect the approximation to do well with the normal camera, with small belief discrepancies and near-complete decision agreement, and to show larger discrepancies in the stress test, where the Roomba’s room becomes uncertain, together with a clearly negative \(J_{\text{H}}\) for the factors involved.
The agents do not optimize these measures directly: they cannot see the true states. Instead, they act on their preferences for observing no motion and dirt picked up. The measures tell us, from the outside, how well those preferences translate into the behavior we want, and how much the approximation changes that.
Changes in the objective function compared to Part 5
The behavior measures are unchanged: disturbance, cleanliness and tracking.
New: three measures of the approximation, computed on the same observations and actions as the exact reference: the belief discrepancy per factor (the total variation distance, on average and at its largest), the decision agreement, the fraction of ticks at which the exact agent would have chosen the same action, written \(\tilde{u}^{[t,0]}\), and the belief entropy difference \(J_{\text{H}}(n)\), which shows whether the agent is overconfident.
New: the runtime per tick of both agents.
Two kinds of comparison replace the comparison of three Roombas: on the same data, to isolate the approximation, and in closed loop, the factorized agent against the exact agent of Part 5 with the same seed, with the random Roomba as a reference. Both are repeated in the stress test.
4.3.6 Implementation of the System-Under-Steer / Environment / Generative Process
## returns a one-hot encoding of a random sample from a categorical distributionfunctionrand_1hot_vec(rng, distribution::Categorical) K =ncategories(distribution) sample =zeros(K) drawn_category =rand(rng, distribution) sample[drawn_category] =1.0return sampleend
rand_1hot_vec (generic function with 1 method)
functioncreate_envir(; LAˣ, LBˣ, LdepsA, LdepsB, Lŝˣ₀, seed=42) rng =MersenneTwister(seed)## the current values (1-based indices) of the factors in deps, read from a tuple of one-hot vectorsvalues_of(Lŝ, deps) =Tuple(argmax(Lŝ[n]) for n in deps)## ô^{[t,m]} ~ Cat(A^{*[m]}[•, values of the factors in D_A^{[m]}]), for every modality mgenerate_obs(Lŝ) =tuple0((rand_1hot_vec(rng, Categorical(LAˣ[m][:, values_of(Lŝ, LdepsA[m])...]))for m ineachindex(LAˣ))...)## t = 0 Lŝˣₜ = Lŝˣ₀ ## ŝ^{*(0)} Lôₜ =generate_obs(Lŝˣₜ) ## ô^{(0)} execute = (Lûₜ₋₁) ->begin## Lûₜ₋₁ = û^{(t-1)}: the action chosen at t-1, a tuple of one-hot vectors u =argmax(Lûₜ₋₁[0]) ## û^{[t-1,0]} as an index: 1 bedroom, 2 livingroom, 3 bathroom Lŝˣₜ₋₁ = Lŝˣₜ## ŝ^{*[t,n]} ~ Cat(B^{*[n]}[•, values of the factors in D_B^{[n]}, (action, for the controlled factor 0)]) Lŝˣₜ =tuple0((begin action = n ==0 ? (u,) : () ## only factor 0 (Room) has an action axis p = LBˣ[n][:, values_of(Lŝˣₜ₋₁, LdepsB[n])..., action...]rand_1hot_vec(rng, Categorical(p))endfor n ineachindex(LBˣ))...) Lôₜ =generate_obs(Lŝˣₜ) ## ô^{(t)}nothingend observe = () -> Lôₜ ## ô^{(t)} = (ô^{[t,0]}, ô^{[t,1]}, ô^{[t,2]}) state = () -> Lŝˣₜ ## ŝ^{*(t)} = (ŝ^{*[t,0]}, ..., ŝ^{*[t,4]}), true state (for evaluation only)return (execute, observe, state)end
create_envir (generic function with 1 method)
Changes in create_envir compared to Part 5: none. The stress test passes its tensors, LAˣ_stress, as the LAˣ argument.
In order to generate data to mimic the observations of the Roomba, we need to specify four things: the initial state (i.e., where the Roomba and the person start, and which rooms are dirty), the actual transition probabilities \(\mathbf{B}^{\ast}\), one tensor per state factor (i.e., how likely the Roomba is to reach the room it heads for, how the person moves, and how dirt accumulates and is cleaned), the observation distributions \(\mathbf{A}^{\ast}\), one tensor per sensor (i.e., what texture the Roomba will observe in each room, whether it will detect motion, depending on where the person is, and whether it will pick up dirt, depending on whether its room is dirty), and a way to choose the actions. We can then use these specifications to generate observations from the generative process, which has the structure of a factorized partially observable Markov decision process (POMDP).
In this section, we test the environment on its own, with a Roomba that chooses its destination at random. This also gives the reference point for the measures from 4.3.5: how often a Roomba without preferences ends up in the same room as the person, and how dirty the rooms get. Since the random Roomba ignores its sensors, a less reliable camera would not change where it goes, so this one run serves as the reference for the stress test too. In 4.6, the agents’ choices take the place of the random ones.
To generate our data, we’ll follow these steps: 1. Assume an initial state \(\hat{\mathbf{s}}^{\ast(0)}\), one value per factor: the Roomba in the bedroom, the person in the living room, and all rooms clean. 2. Determine the observations in the starting state, one per sensor \(m = 0, 1, 2\) (Texture, Motion, Dirt), by drawing from the categorical distribution that \(\boldsymbol{A}^{\ast[m]}\) gives for the current values of the factors that sensor depends on. 3. Choose an action \(\hat{u}^{[t-1,0]}\), the room the Roomba heads for. Here, it is drawn at random from the three rooms. 4. Determine the next state, one factor at a time: for each factor \(n\), draw its next value from the categorical distribution that \(\boldsymbol{B}^{\ast[n]}\) gives for the previous values of the factors it depends on, and, for the Room factor, the chosen action. The Roomba moves, the person moves, and the dirt in each room grows or is cleaned. 5. Determine the observations in the new state, as in step 2. 6. Repeat steps 3-5 for as many time steps as we want.
We will generate 600 data points (\(t = 0, \ldots, T\) with \(T = 599\)) to simulate 600 ticks of the Roomba moving through the apartment. Lôː will contain the Roomba’s observations at each tick, \(\hat{\mathbf{o}}^{(0:T)}\), where each entry is a tuple holding the texture of the floor it’s on, whether it detects motion, and whether it picks up dirt. Lûː will contain the action chosen at each tick, \(\hat{\mathbf{u}}^{(0:T-1)}\). Lŝˣː will contain the true state at each tick, \(\hat{\mathbf{s}}^{\ast(0:T)}\): where the Roomba and the person actually were, and which rooms were actually dirty.
The following code implements this process and generates our data:
Changes in the data generation compared to Part 5: none. The random Roomba ignores its sensors, so its run also serves as the reference for the stress test.
T =599## 600 ticks: t = 0, ..., T## factor 0 (Room): 1 bedroom 2 livingroom 3 bathroom## factor 1 (Person): 1 bedroom 2 livingroom 3 bathroom 4 away## factors 2–4 (Dirt in bedroom, livingroom, bathroom): 1 clean 2 dirtyLŝˣ₀ =tuple0([1.0, 0.0, 0.0], ## Roomba in the bedroom [0.0, 1.0, 0.0, 0.0], ## person in the living room [1.0, 0.0], ## bedroom clean [1.0, 0.0], ## living room clean [1.0, 0.0]) ## bathroom clean → ŝ^{*(0)}(execute_sim, observe_sim, state_sim) =create_envir(; LAˣ, LBˣ, LdepsA, LdepsB, Lŝˣ₀)## random Roomba: chooses its destination at random (reference point for 4.3.5)rng_act =MersenneTwister(7) ## separate random stream for the actionsrandom_action(rng) =tuple0(Float64.(1:3.==rand(rng, 1:3))) ## a random one-hot action tupleLŝˣː =offset0(Vector{Any}(undef, T+1)) ## ŝ^{*(0:T)}: true states, each a tuple (Room, Person, Dirt ×3)Lôː =offset0(Vector{Any}(undef, T+1)) ## ô^{(0:T)}: observations, each a tuple (Texture, Motion, Dirt)Lûː =offset0(Vector{Any}(undef, T)) ## û^{(0:T-1)}: chosen actions, each a tuple (Destination)for t =0:Tif t >0 Lûː[t-1] =random_action(rng_act) ## choose û^{(t-1)} (here at random)execute_sim(Lûː[t-1]) ## the environment carries it out: move, update person and dirt, observeend Lŝˣː[t] =state_sim() Lôː[t] =observe_sim()end
Changes in the random Roomba simulation compared to Part 5: none.
[Lŝˣ[0] for Lŝˣ in Lŝˣː] ## ŝ^{*[t,0]} for t = 0, ..., T: the Room factor (one-hot over 3 rooms)
## t, Roomba's room, person's location, dirt in bedroom, livingroom, bathroom[(t, room_labels[argmax(Lŝˣː[t][0])], person_labels[argmax(Lŝˣː[t][1])], (dirt_labels[argmax(Lŝˣː[t][n])] for n in2:4)...) for t in0:T]
textures = [argmax(Lô[0]) for Lô inparent(Lôː)] ## texture index (1, 2, 3) of ô^{[t,0]} for t = 0, ..., Tplot(title="Observations of the Random Roomba")plot!(0:T, textures, linetype=:steppost, linestyle=:solid, leg=:right, label="observed texture", xlabel="Time t", yticks=(1:3, ["Hardwood", "Carpet", "Tiles"]))
pickups = [argmax(Lô[2]) for Lô inparent(Lôː)] ## dirt index (1 = nothing, 2 = dirt picked up) of ô^{[t,2]} for t = 0, ..., Tplot(title="Dirt Readings of the Random Roomba")plot!(0:T, pickups, linetype=:steppost, linestyle=:solid, leg=:right, label="dirt reading", xlabel="Time t", yticks=(1:2, ["Nothing", "Dirt picked up"]))
motions = [argmax(Lô[1]) for Lô inparent(Lôː)] ## motion index (1 = no motion, 2 = motion) of ô^{[t,1]} for t = 0, ..., Tplot(title="Motion Readings of the Random Roomba")plot!(0:T, motions, linetype=:steppost, linestyle=:solid, leg=:right, label="observed motion", xlabel="Time t", yticks=(1:2, ["No motion", "Motion"]))
rooms = [argmax(Lŝˣ[0]) for Lŝˣ inparent(Lŝˣː)] ## 1 bedroom, 2 livingroom, 3 bathroompersons = [argmax(Lŝˣ[1]) for Lŝˣ inparent(Lŝˣː)] ## 1 bedroom, 2 livingroom, 3 bathroom, 4 awaysame_room = (rooms .== persons) .+1## 1 no, 2 yes: the person is in the Roomba's roomdirt_here = [argmax(Lŝˣː[t][rooms[t+1] +1]) for t in0:T] ## 1 clean, 2 dirty: the dirt factor of the Roomba's roomtextures = [argmax(Lô[0]) for Lô inparent(Lôː)] ## 1 hardwood, 2 carpet, 3 tilesmotions = [argmax(Lô[1]) for Lô inparent(Lôː)] ## 1 no motion, 2 motionpickups = [argmax(Lô[2]) for Lô inparent(Lôː)] ## 1 nothing, 2 dirt picked updestinations = [argmax(Lû[0]) for Lû inparent(Lûː)] ## 1 bedroom, 2 livingroom, 3 bathroom (t = 0, ..., T-1)window = (0, 150) ## zoom in so individual ticks are visible; use (0, T) for the full runpanel(ts, y, ticks, label, color) =plot(ts, y, linetype=:steppost, label=label, color=color, yticks=ticks, xlims=window, legend=:outerright)plot(panel(0:T-1, destinations, (1:3, ["Bedroom", "Living room", "Bathroom"]), "destination", :gray),panel(0:T, rooms, (1:3, ["Bedroom", "Living room", "Bathroom"]), "Roomba's room", :blue),panel(0:T, textures, (1:3, ["Hardwood", "Carpet", "Tiles"]), "texture", :steelblue),panel(0:T, persons, (1:4, ["Bedroom", "Living room", "Bathroom", "Away"]), "person's location", :red),panel(0:T, same_room, (1:2, ["No", "Yes"]), "same room", :red),panel(0:T, motions, (1:2, ["No motion", "Motion"]), "motion", :darkred),panel(0:T, dirt_here, (1:2, ["Clean", "Dirty"]), "Roomba's room dirty", :brown),panel(0:T, pickups, (1:2, ["Nothing", "Dirt"]), "dirt reading", :sienna), layout=(8, 1), size=(900, 1150), link=:x, xlabel=["""""""""""""""Time t"], plot_title="Random Roomba: actions, hidden states, and sensor readings",)
present = same_room .==2misses =count(present .& (motions .==1)) ## person in the room, no motion detectedfalse_alarms =count(.!present .& (motions .==2)) ## person elsewhere, motion detectedprintln("motion sensor misses: $misses of $(count(present)) ticks with the person present ($(round(100*misses/count(present), digits=1))%)")println("motion sensor false alarms: $false_alarms of $(count(.!present)) ticks without ($(round(100*false_alarms/count(.!present), digits=1))%)")
motion sensor misses: 16 of 95 ticks with the person present (16.8%)
motion sensor false alarms: 13 of 505 ticks without (2.6%)
## Empirical rate of dirt accumulation per room, compared with p_dirty from 4.3.4.## Counts the ticks at which a room was clean and the Roomba was elsewhere,## and how often the room was dirty at the next tick.dirty = [[argmax(Lŝˣ[n]) ==2 for Lŝˣ inparent(Lŝˣː)] for n in2:4] ## dirt in bedroom, livingroom, bathroomfor (r, room) inenumerate(["Bedroom", "Living room", "Bathroom"]) chances = [t for t in1:T if !dirty[r][t] && rooms[t] != r] ## clean, and the Roomba elsewhere (1-based positions) rate =mean(dirty[r][t+1] for t in chances)println("$room: got dirty at $(round(rate, digits=3)) of $(length(chances)) chances (p_dirty: $(p_dirty[r]))")end
Bedroom: got dirty at 0.021 of 423 chances (p_dirty: 0.02)
Living room: got dirty at 0.05 of 341 chances (p_dirty: 0.05)
Bathroom: got dirty at 0.035 of 284 chances (p_dirty: 0.03)
## ---------- random Roomba: reference values for 4.3.5 ----------rooms_random = [argmax(Lŝˣ[0]) for Lŝˣ inparent(Lŝˣː)]persons_random = [argmax(Lŝˣ[1]) for Lŝˣ inparent(Lŝˣː)]dirty_random = [[argmax(Lŝˣ[n]) ==2 for Lŝˣ inparent(Lŝˣː)] for n in2:4] ## dirt in bedroom, livingroom, bathroomdisturbance_random =mean(rooms_random .== persons_random) ## J_disturb: fraction of ticks in the person's roomdirtiness_random = [mean(dirty_random[r]) for r in1:3] ## J_dirty(r): fraction of ticks each room is dirtycoverage_random = [mean(rooms_random .== r) for r in1:3] ## fraction of ticks in each room, for referenceprintln("random Roomba, disturbance: $(round(100*disturbance_random, digits=1))% of ticks in the person's room")println("random Roomba, dirty: ",join(["$(r): $(round(100*dirtiness_random[i], digits=1))%" for (i, r) inenumerate(["bedroom", "livingroom", "bathroom"])], ", ")," (average $(round(100*mean(dirtiness_random), digits=1))%)")println("random Roomba, coverage: ",join(["$(r): $(round(100*coverage_random[i], digits=1))%" for (i, r) inenumerate(["bedroom", "livingroom", "bathroom"])], ", "))
random Roomba, disturbance: 15.8% of ticks in the person's room
random Roomba, dirty: bedroom: 8.7%, livingroom: 20.7%, bathroom: 3.2% (average 10.8%)
random Roomba, coverage: bedroom: 22.5%, livingroom: 25.8%, bathroom: 51.7%
4.4 Uncertainty Model
The uncertainty in the environment is captured in the tuple of state transition tensors \(\mathbf{B}^{\ast} = \big(\boldsymbol{B}^{\ast[0]}, \ldots, \boldsymbol{B}^{\ast[4]}\big)\), one per state factor, and the tuple of observation likelihood tensors \(\mathbf{A}^{\ast} = \big(\boldsymbol{A}^{\ast[0]}, \boldsymbol{A}^{\ast[1]}, \boldsymbol{A}^{\ast[2]}\big)\), for Texture, Motion and Dirt. Their entries are the probabilities of the Roomba’s moves, including moves that fail, of the person’s movements, of dirt accumulating and being cleaned, and of the readings of the camera, the motion sensor and the dirt sensor. With several factors, the uncertainty is spread over several small tensors, each describing one aspect of the environment. The actions themselves add no uncertainty: the agent knows which action it chose, only not exactly what will come of it. The initial state adds no uncertainty either, because the run always starts from the same state: the Roomba in the bedroom, the person in the living room, and all rooms clean. In the stress test, the camera’s tensor \(\boldsymbol{A}^{\ast[0]}\) is replaced by a less reliable one, which adds uncertainty about the Roomba’s room; everything else stays the same.
On the agent’s side, there is no uncertainty about how the environment works: as in Part 5, the agent is given \(\mathbf{A} = \mathbf{A}^{\ast}\) and \(\mathbf{B} = \mathbf{B}^{\ast}\). What remains is uncertainty about the state: where the Roomba is, where the person is, and which rooms are dirty, and, when planning, where its actions will lead, whether it will meet the person there, and whether there will be dirt to clean. The uncertainty differs per factor. The Roomba’s room is well tracked, since the agent knows where it was heading and the camera confirms it. The person is only noticed when they are in the Roomba’s room. And the dirt in other rooms is never observed at all: the agent can only predict it, so its uncertainty about those rooms grows the longer the Roomba stays away. The agent handles all this uncertainty with its beliefs, and takes it into account when choosing actions.
New in this part is how the agent represents its uncertainty. Its factorized belief can express how uncertain it is about each factor on its own, but not how the uncertainties about different factors are related. In Part 5, the joint belief could hold, for example, “the Roomba is in the bedroom or the bathroom, and the person is in the same room as the Roomba”, an uncertainty about two factors that is tied together. The factorized belief can only hold “the Roomba is in the bedroom or the bathroom” and “the person is in the bedroom or the bathroom”, which also allows the two to be in different rooms. When one factor is nearly certain, this makes little difference: if the agent knows the Roomba is in the bedroom, “the person is in the same room” simply means “the person is in the bedroom”. In Part 5, the room was tracked correctly 94.7% of the time, so we expect the factorized belief to do well with the normal camera. With the less reliable camera of the stress test, the room is uncertain more often, and the relations the factorized belief cannot hold should start to matter.
Changes in the uncertainty model compared to Part 5
The environment’s uncertainty is unchanged, except in the stress test, where a less reliable camera adds uncertainty about the Roomba’s room.
The agent represents its uncertainty per factor: it can express how uncertain it is about each factor, but not how these uncertainties are related. This matters little when one of the factors is nearly certain, as the Roomba’s room usually is, and more when the room becomes uncertain, as in the stress test.
4.5 Agent / Generative Model
4.5.1 State and Observation variables
On the agent’s side, the state of the environment at time \(t\) is not known: the agent represents it by a (compound) tuple of categorical random variables, one per state factor (here five factors):
factor 0 (Room): \(s^{[t,0]} \in \{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\), where the Roomba is
factor 1 (Person): \(s^{[t,1]} \in \{\text{bedroom}, \text{livingroom}, \text{bathroom}, \text{away}\}\), where the person is
factors 2, 3 and 4 (Dirt in the bedroom, living room and bathroom): \(s^{[t,n]} \in \{\text{clean}, \text{dirty}\}\), \(n = 2, 3, 4\)
The agent never observes these variables directly. What it holds instead is a belief. In Part 5, this was a joint belief over all 96 combinations of the five factors. In this part, it is a factorized belief: one probability vector per factor,
13 numbers in total. For example, \(\boldsymbol{s}^{[t,1]}\) is the agent’s belief about where the person is. The agent treats the product of these vectors as its belief about the whole state, \(Q\big(\mathbf{s}^{(t)}\big) = \prod_{n} \mathcal{Cat}\big(s^{[t,n]};\, \boldsymbol{s}^{[t,n]}\big)\), as introduced in 4.1. This is an approximate posterior: the exact posterior, which the reference filter computes, is in general not a product.
What the agent gives up is exactly what Part 5 needed the joint belief for. Its sensors tie the factors together. When the motion sensor detects motion, the exact posterior learns little about the Roomba’s room or the person’s location on their own, but a lot about the two together: they are probably in the same room. Likewise, when the dirt sensor picks up dirt, it learns that whichever room the Roomba is in is dirty. The factorized belief has to express such information through the separate factors: “the person is probably where I think the Roomba is”. As long as the agent is fairly sure where the Roomba is, this loses little. When it is not, the factorized belief has to settle for one explanation, while the exact posterior keeps several.
where \(\boldsymbol{D}^{[n]}\) parameterizes the categorical distribution of \(s^{[0,n]}\). Unlike the environment, the agent does not know the starting state, so each factor gets a uniform prior: \(\boldsymbol{D}^{[0]} = [\tfrac{1}{3}, \tfrac{1}{3}, \tfrac{1}{3}]^{\top}\) for the room, \(\boldsymbol{D}^{[1]} = [\tfrac{1}{4}, \tfrac{1}{4}, \tfrac{1}{4}, \tfrac{1}{4}]^{\top}\) for the person, and \(\boldsymbol{D}^{[n]} = [\tfrac{1}{2}, \tfrac{1}{2}]^{\top}\) for the dirt in each room, \(n = 2, 3, 4\). The prior is already a product of one distribution per factor, so before any observation arrives, the factorized belief is exact. From the first observations on, it is an approximation.
The (compound) observation tuple received by the agent at time \(t\) is the same one the environment generates:
one-hot encoded as \(\hat{\boldsymbol{o}}^{[t,2]} \in \big\{[1,0]^{\top}, [0,1]^{\top}\big\}\)
As in Parts 4 and 5, the agent also chooses actions, and it prefers some observations over others. These are described in 4.5.2 and 4.5.3.
Changes in the agent’s state and observation variables compared to Part 5
The belief is factorized: one probability vector per factor, 13 numbers in total, instead of a joint belief over 96 combinations. Their product is the agent’s approximate posterior.
What is given up: the relations between factors that the motion and dirt sensors reveal. The factorized belief expresses them through the separate factors, which works well when the Roomba’s room is nearly certain, and forces a single explanation when it is not.
The prior is unchanged, and being a product, it is the one point at which the factorized belief is exact.
4.5.2 Decision variables
As in Parts 4 and 5, the agent makes decisions. At each time step, it chooses an action: the room the Roomba heads for. The agent represents its action at time \(t\) by a (compound) tuple of categorical random variables, one per control factor (here a single factor):
\[
\mathbf{u}^{(t)} = \big(u^{[t,0]}\big)
\]
where the control factors are:
control factor 0 (Destination), acting on state factor 0 (Room):
Heading for the room the Roomba is already in means staying put. Unlike the state, the agent always knows which action it has taken: once chosen, the action is a realized value \(\hat{\boldsymbol{u}}^{[t,0]}\), which the agent hands to the environment and also uses itself, to predict how the state will change. What the agent does not know is the outcome: whether the move succeeds, whether the person is in the room it reaches, and whether there is dirt there to clean.
The action selects the transition matrix of the Room factor only. Its effect on the other factors runs through their dependencies: the dirt factors depend on where the Roomba is, so in its predictions the agent’s choice of room also determines which room gets cleaned. The person’s movements, on the other hand, are predicted the same way whatever the agent does. With a factorized belief, the effect on the dirt is predicted from the agent’s belief about where the Roomba is, rather than from the combinations of room and dirt: each room’s predicted dirt is averaged over where the agent thinks the Roomba will be. As long as that belief is sharp, this is close to exact.
Policies. To choose an action, the agent looks \(H\) steps ahead. It considers candidate policies, i.e. sequences of \(H\) actions, collected as the rows of a matrix \(\boldsymbol{V}\): policy \(p\) is \(\pi^{[p]} = \boldsymbol{V}[p, \bullet]\), and \(\boldsymbol{V}[p, \tau]\) is the action it prescribes at step \(\tau = 0, \ldots, H-1\). With three possible actions per step, there are \(3^H\) policies. As in Part 5, \(H = 3\), which gives 27 policies. Three steps are needed for cleaning to pay off: from the bedroom, reaching the living room takes two moves, through the bathroom, and the dirt is picked up and removed only once the Roomba is there. A policy such as head for the bathroom, then the living room, then stay lets the agent see that pay-off. The agent scores each policy by its expected free energy (see 4.5.6), turns the scores into a probability for each policy, and from those derives a probability for each possible next action. Only the first action of the chosen policy is carried out; at the next time step, the agent observes the result and plans again.
The factorized belief makes each policy cheaper to evaluate, but it does not reduce the number of policies. That number grows exponentially with the horizon and with the number of actions, so in larger models planning needs its own remedy, such as pruning unpromising policies or formulating planning as inference. That is a separate problem from the one this part addresses.
Changes in the agent’s decision variables compared to Part 5
The decisions are unchanged: the same action, horizon \(H = 3\), and 27 policies.
The action’s effect on the dirt is predicted from the belief about the Roomba’s room, averaged over where the agent thinks the Roomba will be, rather than per combination of room and dirt.
The factorized belief does not reduce the number of policies, which grows exponentially with the horizon; this needs a separate remedy in larger models.
4.5.3 Preference variables / Setpoints
The agent’s goals are expressed as preferences over observations: the agent prefers to receive the observations it wants to be true. There is one preference vector per modality, collected in \(\mathbf{C} = \big(\boldsymbol{C}^{[0]}, \boldsymbol{C}^{[1]}, \boldsymbol{C}^{[2]}\big)\). Each \(\boldsymbol{C}^{[m]}\) is a probability vector over the observations of modality \(m\), playing the role of a setpoint: the distribution of observations the agent would like to see.
\(\boldsymbol{C}^{[0]}\) (Texture) is flat, \(\boldsymbol{C}^{[0]} = [\tfrac{1}{3}, \tfrac{1}{3}, \tfrac{1}{3}]^{\top}\): the agent has no preference for any floor type.
\(\boldsymbol{C}^{[1]}\) (Motion) prefers no motion, e.g. \(\boldsymbol{C}^{[1]} = [0.9, 0.1]^{\top}\): the agent would like to observe no motion 90% of the time. This is the goal introduced in Part 4: not disturbing anyone.
\(\boldsymbol{C}^{[2]}\) (Dirt) prefers dirt picked up, e.g. \(\boldsymbol{C}^{[2]} = [0.2, 0.8]^{\top}\): the agent would like its brushes to pick up dirt 80% of the time. This is the goal added in Part 5: cleaning.
The agent cannot prefer unoccupied rooms or clean rooms directly, because it never observes the person’s location or the dirt. Instead, it prefers the observations that go with what it wants, and chooses actions it expects to produce them. This is the usual way of expressing goals in active inference: preferences are stated over what the agent can sense, not over the hidden state.
For cleaning, the choice of observation needs care. It may seem natural to prefer observing a clean floor, i.e. nothing picked up, but that preference would backfire: the agent would be drawn to rooms that are already clean, where it has nothing to do. Preferring dirt picked up draws it to rooms that are dirty, which are exactly the rooms where its presence makes a difference. Once a room is clean, the preference is no longer satisfied there, and it pulls the agent on to wherever dirt is most likely to have built up: the living room, which gets dirty fastest, and rooms it has not visited for a while.
How strongly the agent holds each preference is set by how far its \(\boldsymbol{C}^{[m]}\) is from flat. When scoring a policy, the agent compares the observations it expects under that policy with its preferences (see 4.5.6), and \(\log \boldsymbol{C}^{[m]}\) determines how much each expected reading counts:
with \(\boldsymbol{C}^{[1]} = [0.9, 0.1]^{\top}\), observing motion is \(\log(0.9/0.1) \approx 2.2\) worse than observing no motion, and
with \(\boldsymbol{C}^{[2]} = [0.2, 0.8]^{\top}\), observing dirt picked up is \(\log(0.8/0.2) \approx 1.4\) better than observing nothing.
The ratio of the two sets the agent’s balance between its goals. Here, avoiding the person weighs somewhat more than cleaning, so the agent should not simply go wherever dirt is likely, but clean the living room when it expects it to be quiet. More extreme preferences make the agent’s choices more driven by its goals, and less by reducing its uncertainty about the state.
In this part, the factorized agent and the exact agent have the same preferences. Preferences are stated over observations, which are the same for both, so the approximation does not change what the agent wants. It only changes the expected observations against which the preferences are compared, since these are computed from the factorized belief. Any difference in behavior between the two agents therefore comes from how they predict, not from what they prefer.
The flat preference for Texture does not mean the camera is useless: it still helps the agent infer where the Roomba is, which it needs in order to predict where its actions will lead, and which room’s dirt the dirt sensor is sensing. In this part, it has one more role: by keeping the Roomba’s room nearly certain, it is what keeps the factorized belief close to the exact one. The stress test weakens exactly this role.
Changes in the preferences compared to Part 5
The preferences are unchanged, and the factorized agent and the exact agent share them.
The approximation affects only the expected observations against which the preferences are compared, not the preferences themselves, so behavioral differences come from prediction alone.
The comparison with the Part 4 agent (a flat Dirt preference) is dropped; it belonged to Part 5.
The camera has an extra role: keeping the Roomba’s room nearly certain, which keeps the factorized belief close to the exact one. The stress test weakens this role.
4.5.4 State Transition and Observation Generation functions
The state transition functions, one per state factor, are given by
where, as in 4.3, the values of the random variables are used as indices into the tensors. Each line picks out the column of \(\boldsymbol{B}^{[n]}\) for the previous values of exactly the factors in \(\mathcal{D}_B^{[n]}\), and, for the Room factor, the action chosen at \(t-1\). Given the previous state and action, the factors change independently of each other, so the transition of the whole state is the product
where each line picks out the column of \(\boldsymbol{A}^{[m]}\) for the current values of exactly the factors in \(\mathcal{D}_A^{[m]}\). It parameterizes the categorical distribution of \(o^{[t,m]}\): over textures for \(m = 0\), over motion readings for \(m = 1\), and over dirt readings for \(m = 2\). Observations do not depend on the action.
As in Part 5, the agent is given its tensors, \(\boldsymbol{A}^{[m]} = \boldsymbol{A}^{\ast[m]}\) and \(\boldsymbol{B}^{[n]} = \boldsymbol{B}^{\ast[n]}\), and it knows their dependencies, which are the same as the environment’s. The tensors are known parameters, not random variables, so they appear as subscripts. The model is the same as in Part 5; what changes is how the agent uses it with a factorized belief.
An entry of \(\boldsymbol{A}^{[m]}\) is the probability of a specific observation given specific values of the factors it depends on, and each column, obtained by fixing all axes except the first, is a categorical distribution over observations. The same holds for each \(\boldsymbol{B}^{[n]}\), with the next value of factor \(n\) in place of the observation. If the state were known, the right column could simply be looked up. The agent, however, only has its beliefs about the separate factors, so it averages over the values of the factors a tensor depends on, weighting each combination of their values by the product of the factors’ probabilities. To write this down, let \(j_d\) denote a value of factor \(d\), and \(j_{\mathcal{D}}\) the values of the factors in a dependency set \(\mathcal{D}\), in order. For example, \(j_{\mathcal{D}_A^{[1]}} = (j_0, j_1)\), the Roomba’s room and the person’s location, whose combination gets the weight \(\boldsymbol{s}^{[t,0]}[j_0]\,\boldsymbol{s}^{[t,1]}[j_1]\). Only the factors a tensor depends on are involved: the 96 combinations of all factors never appear.
Using the functions for inference. Perception has two steps, as in Part 5: predict, then update. Both now work factor by factor.
Prediction. After carrying out action \(\hat{u}^{[t-1,0]}\), the agent predicts each factor from its previous beliefs about the factors that factor’s transition depends on:
where the action index appears only for the Room factor. For the person, this is the familiar \(\boldsymbol{B}^{[1]}\,\boldsymbol{s}^{[t-1,1]}\). For the dirt in room \(n\), it averages the cleaning and accumulating tables over where the agent thought the Roomba was. Each predicted belief is exactly what the exact filter would predict for that factor, if its previous belief had been a product. What is lost is the relation between the predicted factors: in the exact prediction, “the Roomba stayed in the living room” goes together with “the living room was cleaned”, whereas the factorized prediction keeps only how likely each is on its own. At \(t = 0\), the prediction is replaced by the uniform priors \(\boldsymbol{D}^{[n]}\) from 4.5.1.
Update by fixed-point iteration. The agent now combines the predicted beliefs with the observations. With a factorized belief, this cannot be done in one step, because each sensor ties several factors together. Instead, the agent updates the beliefs about the factors in turn, each time holding the beliefs about the other factors fixed:
where \(\hat{o}^{[t,m]}\), the received observation without bold, is used as an index, the result is normalized to sum to 1, and the beliefs \(\boldsymbol{s}^{[t,d]}\) about the other factors are their current values. In words: the new belief about factor \(n\) is its prediction, reweighted by the evidence of every sensor that depends on it, where the evidence for value \(j\) is the log-likelihood of the received reading, averaged over the agent’s current beliefs about the sensor’s other factors. One pass over all five factors is one iteration \(k\); after \(K\) iterations, or once the beliefs no longer change, the agent takes \(\boldsymbol{s}^{[t,n]} = \boldsymbol{s}^{[t,n]}_{\langle K \rangle}\).
Each sensor reaches only the factors it depends on:
Texture depends only on the room. Its evidence enters the Room belief directly, as in a single-factor model.
Motion depends on the room and the person. Its evidence about the person is averaged over the current belief about the room, and vice versa. This is where the iteration matters: a motion reading pulls the person’s belief towards the room the agent currently believes the Roomba is in. If that room belief is sharp, one pass suffices; if it is not, the two beliefs adjust to each other over several iterations, and settle on the most consistent single explanation.
Dirt depends on the room and the three dirt factors. With a sharp room belief, its evidence lands almost entirely on the dirt factor of the Roomba’s room; the dirt in the other rooms is, as before, only predicted.
The averaging inside the logarithm’s exponent is where the approximation lies. The exact filter averages the likelihoods themselves over the joint belief; the factorized agent averages their logarithms over a product of beliefs. When the averaged-over beliefs are nearly certain, the two agree; when they are spread out, the factorized update is more decisive than the exact one, which is the overconfidence mentioned in 4.2.
The variational free energy. The iteration minimizes the variational free energy of the factorized belief at time \(t\),
The first term, the complexity, measures how far the beliefs have moved from their predictions. The second, the accuracy with its sign reversed, measures how poorly the beliefs explain the received observations. Each update of a single factor is the exact minimizer of \(F^{(t)}\) over that factor’s belief, with the others held fixed, so every update lowers \(F^{(t)}\), and the values \(F^{(t)}_{\langle k \rangle}\) after each iteration decrease until the beliefs settle. This in-turn scheme is one variant of the fixed-point iteration; pymdp by default updates all factors at once from the previous iteration, which is usually close but not guaranteed to decrease \(F^{(t)}\) at every step.
The exact reference. For comparison, the exact filter of Part 5 runs alongside, on the same observations and actions. With \(\mathbf{j} = (j_0, \ldots, j_4)\) a combination of all factor values, it computes
over all 96 combinations, and its marginals \(\tilde{\boldsymbol{s}}^{[t,n]}\) by adding up. The agent never uses it.
Using the functions for planning. To evaluate policy \(p\), the agent applies the factorized prediction to the future, step by step, starting from its current beliefs:
Here \(\boldsymbol{s}_{\pi^{[p]}}^{[\tau,n]}\) is the predicted belief about factor \(n\) at time \(t + \tau + 1\) if the agent follows policy \(p\), and \(\boldsymbol{o}_{\pi^{[p]}}^{[\tau,m]}\) is the observation of sensor \(m\) it then expects. This is the computation of the example in 4.1, which treated the room and the person as independent. No real observations are available for the future, so there is no update and no iteration in planning: the predictions are what the agent compares with its preferences when it scores policy \(p\) (see 4.5.6).
The two preferences make the corresponding predictions the ones that matter most:
\(\boldsymbol{o}_{\pi^{[p]}}^{[\tau,1]}\) tells the agent how likely it is to observe motion if it follows policy \(p\). Since the person’s movements do not depend on the action, this follows from where the policy takes the Roomba and where the person is predicted to be, which, without new evidence, drifts towards the living room. With factorized predictions, the two are combined as if independent.
\(\boldsymbol{o}_{\pi^{[p]}}^{[\tau,2]}\) tells the agent how likely it is to pick up dirt. In the predictions, dirt grows in every room the Roomba is not in, and is removed in the room it is in, averaged over where the agent expects the Roomba to be. So a policy that heads for a room the agent has not visited for a while, especially the living room, is expected to pick up dirt, while a policy that stays in a room it has just cleaned is not.
The cost. Every formula above sums only over the values of the factors a single tensor depends on: at most \(3 \cdot 2 \cdot 2 \cdot 2 = 24\) combinations, for the dirt sensor, instead of the 96 combinations of the exact filter, and instead of \(96 \times 96\) transitions in its prediction. In general, the cost of the factorized agent grows with the size of the largest tensor, not with the number of joint states, which is what makes larger models feasible.
Changes in the agent’s transition and observation functions compared to Part 5
The model is unchanged; what changes is how the agent uses it with a factorized belief.
Prediction is per factor: each factor is predicted from the beliefs about the factors its transition depends on. Each prediction is what the exact filter would give for a product belief, but the relations between the predicted factors are lost.
The update is a fixed-point iteration: the factors are updated in turn, each reweighting its prediction by the evidence of the sensors that depend on it, with log-likelihoods averaged over the current beliefs about the sensor’s other factors. Averaging log-likelihoods over a product, rather than likelihoods over the joint, is where the approximation lies, and why it tends to be overconfident.
New: the variational free energy\(F^{(t)}\), complexity minus accuracy, which every in-turn update lowers, unlike pymdp’s default all-at-once update.
The exact reference is the exact filter of Part 5, run alongside on the same data.
Planning uses factorized predictions, with expected observations computed as if the factors were independent, as in the example of 4.1. There is no iteration in planning.
The cost grows with the size of the largest tensor, at most 24 combinations here, instead of with the 96 joint states.
4.5.5 Implementation of the Agent / Generative Model / Internal Model
We start by specifying a probabilistic model for the agent that describes the agent’s internal beliefs about the external dynamics of the environment. As in Parts 4 and 5, this model has two parts. Up to the current time \(t\), it describes what has happened: the states the environment went through, the observations the Roomba received, and the actions \(\hat{u}^{[0:t-1,0]}\) the agent actually took. Beyond \(t\), it describes what could happen over the next \(H\) steps, depending on which policy \(\pi^{(t)}\) the agent chooses. With several factors and sensors, we write the model in terms of the tuples \(\mathbf{s}^{(t')}\) and \(\mathbf{o}^{(t')}\). The generative model at time \(t\) is the same as in Part 5:
where \(P_D\) is the agent’s uniform initial state prior from 4.5.1, and \(P_{\mathbf{A}}\) and \(P_{\mathbf{B}}\) combine the functions from 4.5.4, one per sensor and one per factor:
with, as in 4.5.4, the dots standing for exactly the factors in \(\mathcal{D}_A^{[m]}\) and \(\mathcal{D}_B^{[n]}\), and, for the Room factor, the action \(u\). The tensors \(\mathbf{A} = \mathbf{A}^{\ast}\) and \(\mathbf{B} = \mathbf{B}^{\ast}\) and the prior \(\mathbf{D}\) are known parameters, so they appear as subscripts. As in Parts 4 and 5, there are no parameter priors: the agent learns nothing about its tensors in this part.
Written out, the model is a product of many small pieces, each involving only a few variables: one prior per factor, one observation function per sensor and time step, and one transition function per factor and time step. This is what a factorized model means, and it is why each tensor can stay small. The action appears in only one of these pieces, the transition of the Room factor; its influence on the dirt runs through the dirt transitions’ dependence on the Roomba’s previous room.
The two transition products differ only in where the action comes from. In the past, the action is the one the agent actually took, \(\hat{u}^{[t'-1,0]}\), a known value. In the future, it is the action \(\boldsymbol{V}[\pi^{(t)}, \tau]\) that the policy prescribes at step \(\tau\); since the policy is not yet chosen, it is a random variable. Step \(\tau\) predicts time \(t + \tau + 1\), as in 4.5.4.
The policy prior\(P\big(\pi^{(t)}\big)\) is where the agent’s preferences enter the model. In active inference, it is not chosen freely: policies are a priori more probable the lower their expected free energy,
where \(\boldsymbol{G}\) collects the expected free energies \(G^{[p]}\) of the 27 policies, which depend on both preferences in \(\mathbf{C}\) from 4.5.3 (see 4.5.6), and \(\sigma\) is the softmax. The agent thus expects to act in ways that lead to its preferred observations: no motion, and dirt picked up.
The product inside \(P_{\mathbf{A}}\) is where the three sensors are combined. Each sensor depends on its own subset of the factors, and given the whole state, the three readings are independent. For the future, these observations are not yet available: the agent uses their expected values, \(\boldsymbol{o}_{\pi^{[p]}}^{[\tau,m]}\) from 4.5.4, to evaluate each policy.
The approximate posterior. The factorization is a property of the model, not of the exact posterior: once observations come in, the exact posterior over the state no longer factorizes, as explained in 4.5.1. In Part 5, the agent therefore kept the exact posterior, as a joint belief over all 96 combinations. In this part, it replaces the exact posterior at each tick by an approximate posterior of a simpler form, a product of one distribution per factor,
and chooses its beliefs \(\boldsymbol{s}^{[t,n]}\) to make this product as close as possible to the exact posterior, in the sense of the variational free energy \(F^{(t)}\) of 4.5.4. The model is exactly the same as for the exact agent; only the form of the belief is restricted.
This happens tick by tick. The agent does not revisit the past: at each tick, everything it has learned so far is summarized in its previous beliefs \(\boldsymbol{s}^{[t-1,n]}\), from which it predicts the new beliefs \(\boldsymbol{s}^{[t,n]}_{\langle 0 \rangle}\), and the fixed-point iteration then takes the new observations into account. So there are two approximations, applied at every tick: the prediction is computed from a product belief, which loses the relations between factors that the exact prediction would carry forward, and the update is restricted to a product belief, which cannot hold the relations that the sensors create. Both are evaluated together by the comparison with the exact reference.
Changes in the generative model compared to Part 5
The generative model is unchanged.
The posterior is approximated: at each tick, the agent replaces the exact posterior by a product of one distribution per factor, chosen to minimize the variational free energy.
The approximation enters twice per tick: in the prediction, which is computed from a product belief, and in the update, which is restricted to a product belief. The comparison with the exact reference measures both together.
Now it is time to build our agent. As in Parts 4 and 5, the agent does not learn anything: it is given its model of the world, i.e. the observation tensors \(\mathbf{A} = \mathbf{A}^{\ast}\) with their dependencies \(\mathcal{D}_A^{[m]}\), the transition tensors \(\mathbf{B} = \mathbf{B}^{\ast}\) with their dependencies \(\mathcal{D}_B^{[n]}\), the uniform initial priors \(\mathbf{D}\) from 4.5.1, and its preferences \(\mathbf{C}\) from 4.5.3. What it has to do is use this model, at every tick, in three steps.
1. Update its beliefs (perception). After the environment has carried out the previous action \(\hat{u}^{[t-1,0]}\) and returned the new observations \(\hat{\mathbf{o}}^{(t)}\), the agent first predicts each factor, then corrects these predictions with what its sensors report, as in 4.5.4:
the prediction of factor \(n\) contracts its transition tensor \(\boldsymbol{B}^{[n]}\) with the previous beliefs about the factors it depends on, and, for the Room factor, selects the slice of the action taken. For the person, this is a matrix-vector product; for the dirt in each room, an average over where the agent thought the Roomba was. At \(t = 0\), the predictions are the uniform priors \(\boldsymbol{D}^{[n]}\).
the fixed-point iteration then updates the factors in turn. For factor \(n\), each sensor that depends on it contributes the logarithm of the slice \(\boldsymbol{A}^{[m]}[\hat{o}^{[t,m]}, \ldots]\) for the received observation, contracted with the current beliefs about the sensor’s other factors, leaving a vector over the values of factor \(n\). These log-evidence vectors are added to the logarithm of the prediction, and the result is exponentiated and normalized. One pass over all five factors is one iteration; the agent stops once the beliefs no longer change, or after at most \(K = 10\) iterations.
After each iteration, the agent computes the variational free energy \(F^{(t)}_{\langle k \rangle}\), which should decrease and level off. This is a check that the iteration works, and it shows how many iterations are actually needed. Both steps use only the tensors and the five belief vectors: there is no joint belief anywhere in the agent.
2. Evaluate the policies (planning). For each of the \(3^H = 27\) policies \(\pi^{[p]}\), the agent predicts its factorized beliefs and expected observations over the next \(H = 3\) steps, as in 4.5.4, and scores the policy by its expected free energy \(G^{[p]}\) (see 4.5.6). The score favors policies that are expected to lead to preferred observations, i.e. no motion and dirt picked up, and policies that are expected to reduce the agent’s uncertainty about the state, e.g. by going back to check on a room it has not seen for a while. The predictions use the same per-factor contraction as in step 1, without an update, since no observations are available for the future.
3. Choose an action (decision). The scores are turned into a posterior over policies, \(Q\big(\pi^{(t)} = \pi^{[p]}\big) = \sigma\big(-\boldsymbol{G}\big)_p\), and from it into a distribution over the next action, \(Q\big(u^{[t,0]}\big)\), by adding up the probabilities of all policies that start with that action. The agent then carries out the most probable action, \(\hat{u}^{[t,0]}\), and the cycle repeats at the next tick. This step is the same as in Part 5.
The exact reference. Alongside the agent, we run the exact filter and planner of Part 5, unchanged, on the factorized agent’s observations and actions. At each tick, it computes the exact joint belief \(\tilde{\boldsymbol{s}}^{(t)}\) and its marginals \(\tilde{\boldsymbol{s}}^{[t,n]}\), and scores the policies from them, which gives the action \(\tilde{u}^{[t,0]}\) the exact agent would have chosen. The reference never influences the factorized agent: it only watches, so that its beliefs and would-be decisions can be compared with the agent’s on exactly the same data.
All of this is computed directly in Julia, without RxInfer. As in Part 5, the code is generic over factors and modalities: the dependency tuples LdepsA and LdepsB tell it which beliefs each tensor is contracted with, so a single contraction function serves the prediction, the fixed-point iteration, the free energy and the expected observations. This keeps every quantity of Algorithm 18 visible in the code. RxInfer also supports active inference, including factorized beliefs, as message passing on a factor graph, and planning formulated as inference; this becomes attractive for larger models, where the policies can no longer be enumerated either, and is a possible direction for a later part. Let’s build the agent!
Changes in the agent’s three steps compared to Part 5
Perception is factorized: each factor is predicted by contracting its transition tensor with the previous beliefs about its dependencies, and the beliefs are then updated in turn by the fixed-point iteration, for at most \(K = 10\) iterations. The variational free energy is computed after each iteration, as a check.
Planning predicts factorized beliefs, using the same contraction as in perception.
The decision step is unchanged.
The exact reference runs the exact filter and planner of Part 5 alongside, on the same data, to provide \(\tilde{\boldsymbol{s}}^{[t,n]}\) and \(\tilde{u}^{[t,0]}\) for the evaluation. It never influences the agent.
A single contraction function serves the prediction, the update, the free energy and the expected observations.
## ---------- the agent's model: given, not learned ----------LA = LAˣ ## A^{[m]} = A^{*[m]}, m = 0, 1, 2 (dependencies D_A^{[m]} in LdepsA)LB = LBˣ ## B^{[n]} = B^{*[n]}, n = 0, ..., 4 (dependencies D_B^{[n]} in LdepsB)LD =tuple0(fill(1/3, 3), ## D^{[0]}: Room, uniformfill(1/4, 4), ## D^{[1]}: Person, uniformfill(1/2, 2), fill(1/2, 2), fill(1/2, 2)) ## D^{[2]}, D^{[3]}, D^{[4]}: Dirt, uniformLC =tuple0(fill(1/3, 3), ## C^{[0]}: Texture, flat [0.9, 0.1], ## C^{[1]}: Motion, prefers no motion [0.2, 0.8]) ## C^{[2]}: Dirt, prefers dirt picked up## ---------- policies: all action sequences of length H ----------H =3## planning horizon_V = [p[τ] for p invec(collect(Iterators.product(ntuple(_ ->1:3, H)...))), τ in1:H] ## P×H; column τ+1 is step τP =size(_V, 1) ## number of policies, 3^H = 27## ---------- fixed-point iteration ----------K =10## at most K iterations per ticktol_fpi =1e-8## stop early once no belief changes by more than this## ---------- helpers ----------normalize(v) = v ./sum(v)log_stable(x) =log.(max.(x, 1e-16)) ## avoids log(0)softmax(x) =normalize(exp.(x .-maximum(x)))## H[A^{[m]}]: entropy of each column of A^{[m]}, an array over the factors in D_A^{[m]}entropies(LA) =tuple0((dropdims(-sum(LA[m] .*log_stable(LA[m]), dims=1), dims=1) for m ineachindex(LA))...)## the agent's sensor model: the likelihood tensors and their column entropies## (the stress test gives both the exact and the factorized agent a less reliable camera)model(LA) = (; LA, LH =entropies(LA))model_normal =model(LAˣ)model_stress =model(LAˣ_stress)## ---------- from policy scores to an action (shared by both agents) ----------functionplan_from(_G) _qπ =softmax(-_G) ## Q(π^{(t)} = π^{[p]}) = σ(-G)_p _qu = [sum(_qπ[_V[:, 1] .== u]) for u in1:3] ## Q(u^{[t,0]}): add up policies by their first action (_G, _qπ, _qu)endchoose(_qu) =argmax(_qu) ## the most probable action (1 bedroom, 2 livingroom, 3 bathroom)## ==================================================================================================## THE FACTORIZED AGENT: one belief per factor, Ls[n] = s^{[t,n]}## ==================================================================================================## contracts an array whose axes correspond to the vectors in vs with all of these vectors,## except the one at position keep (0: contract all, giving a number; otherwise a vector over that axis)functiondot_except(arr, vs, keep=0) a = arrfor k inlength(vs):-1:1## from the last axis down, so earlier axes keep their positions k == keep &&continue shape =ntuple(i -> i == k ? length(vs[k]) :1, ndims(a)) a =dropdims(sum(a .*reshape(vs[k], shape), dims=k), dims=k)endndims(a) ==0 ? a[] : aend## contracts all axes of T but the first, giving a vector over the first axis (the next value, or the observation)dot_tail(T, vs) = [dot_except(selectdim(T, 1, i), vs) for i inaxes(T, 1)]## ---------- prediction, one factor at a time ----------## s^{[t,n]}_{⟨0⟩}[j] = Σ_{j_D} B^{[n]}[j, j_D, (u)] ∏_{d ∈ D_B^{[n]}} s^{[t-1,d]}[j_d]predict_factor(Ls, n, u) =dot_tail(n ==0 ? LB[0][:, :, u] : LB[n], [Ls[d] for d in LdepsB[n]])predict_fact(Ls, u) =tuple0((predict_factor(Ls, n, u) for n ineachindex(LB))...)## ---------- the variational free energy of factorized beliefs Ls, with predictions Lprior ----------## F = Σ_n KL(s^{[t,n]} ‖ s^{[t,n]}_{⟨0⟩}) - Σ_m E_Q[ log A^{[m]}[ô^{[t,m]}, ...] ]## (complexity) (accuracy)vfe(Ls, Lprior, LlogA) =sum(sum(Ls[n] .* (log_stable(Ls[n]) .-log_stable(Lprior[n]))) for n ineachindex(Ls)) -sum(dot_except(LlogA[m], [Ls[d] for d in LdepsA[m]]) for m ineachindex(LlogA))## ---------- step 1: perception (prediction, then fixed-point iteration) ----------## s^{[t,n]}[j] ∝ s^{[t,n]}_{⟨0⟩}[j] · exp( Σ_{m: n ∈ D_A^{[m]}} E_{other factors}[ log A^{[m]}[ô^{[t,m]}, ...] ] )functionperceive_fact(M, Lsₜ₋₁, Lôₜ, ûₜ₋₁) Lprior = Lsₜ₋₁ ===nothing ? LD :predict_fact(Lsₜ₋₁, ûₜ₋₁) ## s^{[t,n]}_{⟨0⟩}; at t = 0 the priors D^{[n]} LlogA =tuple0((log_stable(selectdim(M.LA[m], 1, argmax(Lôₜ[m]))) for m ineachindex(M.LA))...) ## log A^{[m]}[ô^{[t,m]}, ...] Lsₜ =tuple0((copy(Lprior[n]) for n ineachindex(LB))...) ## start the iteration at the predictions _Fₜ =Float64[] ## F^{(t)}_{⟨k⟩}, k = 1, 2, ...for k in1:K Lold =tuple0((copy(Lsₜ[n]) for n ineachindex(LB))...)for n ineachindex(LB) ## update the factors in turn logev =log_stable(Lprior[n])for m ineachindex(M.LA) pos =findfirst(==(n), LdepsA[m]) ## does sensor m depend on factor n? pos ===nothing&&continue logev = logev .+dot_except(LlogA[m], [Lsₜ[d] for d in LdepsA[m]], pos) ## averaged over the othersend Lsₜ[n] =softmax(logev)endpush!(_Fₜ, vfe(Lsₜ, Lprior, LlogA))maximum(maximum(abs.(Lsₜ[n] .- Lold[n])) for n ineachindex(LB)) < tol_fpi &&breakend (Lsₜ, _Fₜ)end## ---------- step 2: expected free energy of policy p, with factorized predictions ----------## G^{[p]} = Σ_τ Σ_m [ o_π^{[τ,m]}·(log o_π^{[τ,m]} - log C^{[m]}) + E_Q[ H[A^{[m]}] ] ]functionefe_fact(M, Lsₜ, p, LC) Ls = Lsₜ G =0.0for τ in1:H ## code τ = math τ + 1 Ls =predict_fact(Ls, _V[p, τ]) ## s_π^{[τ,n]}: predicted beliefs, one per factorfor m ineachindex(M.LA) vs = [Ls[d] for d in LdepsA[m]] ## predicted beliefs about the factors sensor m depends on o =dot_tail(M.LA[m], vs) ## o_π^{[τ,m]}: expected observation, factors treated as independent G += o'* (log_stable(o) -log_stable(LC[m])) +dot_except(M.LH[m], vs)endend Gend## ---------- step 3: from policy scores to an action ----------plan_fact(M, Lsₜ, LC) =plan_from([efe_fact(M, Lsₜ, p, LC) for p in1:P])## ==================================================================================================## THE EXACT REFERENCE: the agent of Part 5, a joint belief over all 96 combinations, _s̃## ==================================================================================================sizes =Tuple(size(LB[n], 1) for n ineachindex(LB)) ## number of values per factor: (3, 4, 2, 2, 2), 96 combinations## the marginal of a joint belief over the factors in deps, with its axes in the order of depsfunctionmarginal(_s, deps) others =Tuple(k for k in1:ndims(_s) if (k-1) ∉ deps) ## axes to add up over (axis k is factor k-1) m =dropdims(sum(_s, dims=others), dims=others) ## axes now in increasing factor orderpermutedims(m, invperm(sortperm(deps))) ## reorder to the order of depsendmarginals(_s) =tuple0((marginal(_s, [n]) for n ineachindex(LB))...) ## (s̃^{[t,0]}, ..., s̃^{[t,4]})## an array over the factors in deps (axes in the order of deps), spread out over the joint state's other axesfunctionspread(arr, deps) a =permutedims(arr, sortperm(deps)) ## axes in increasing factor orderreshape(a, ntuple(k -> (k-1) in deps ? sizes[k] :1, length(sizes))) ## size 1 on the other axes: broadcastsend_D =reduce((a, b) -> a .* b, (spread(LD[n], [n]) for n ineachindex(LD))) ## joint prior: ∏_n D^{[n]}, 1/96 each## a factor is moved before the factors it depends on, so that it uses their previous valuesupdate_order = [4, 3, 2, 1, 0]@assertall(findfirst(==(n), update_order) <findfirst(==(d), update_order)for n ineachindex(LB) for d in LdepsB[n] if d != n)## moves the probability of each combination along the axis of factor n, using B^{[n]}functiontransition(_s, n, u) _s′ =zero(_s) action = n ==0 ? (u,) : () ## only factor 0 (Room) has an action axisfor I inCartesianIndices(_s) p = _s[I] p ==0&&continue i =Tuple(I) ## the combination (1-based values of all factors) col =view(LB[n], :, (i[d+1] for d in LdepsB[n])..., action...) ## B^{[n]}[•, values of D_B^{[n]}, (u)]for (j, b) inenumerate(col) _s′[Base.setindex(i, j, n+1)...] += b * p ## factor n takes value j, the others keep theirsendend _s′endfunctionpredict_joint(_s, u)for n in update_order _s =transition(_s, n, u)end _send## exact Bayesian filtering over the joint states, as in Part 5functionperceive_exact(M, _s̃ₜ₋₁, Lôₜ, ûₜ₋₁) prediction = _s̃ₜ₋₁ ===nothing ? _D :predict_joint(_s̃ₜ₋₁, ûₜ₋₁) likelihood =reduce((a, b) -> a .* b, (spread(selectdim(M.LA[m], 1, argmax(Lôₜ[m])), LdepsA[m]) for m ineachindex(M.LA)))normalize(likelihood .* prediction) ## _s̃ₜ = s̃^{(t)}end## the expected free energy of policy p, with joint predictions, as in Part 5functionefe_exact(M, _s̃ₜ, p, LC) _s = _s̃ₜ G =0.0for τ in1:H _s =predict_joint(_s, _V[p, τ])for m ineachindex(M.LA) q =marginal(_s, LdepsA[m]) o =reshape(M.LA[m], size(M.LA[m], 1), :) *vec(q) G += o'* (log_stable(o) -log_stable(LC[m])) +sum(M.LH[m] .* q)endend Gendplan_exact(M, _s̃ₜ, LC) =plan_from([efe_exact(M, _s̃ₜ, p, LC) for p in1:P])
plan_exact (generic function with 1 method)
Changes in the agent code compared to Part 5
The sensor model is passed as an argument,M = model(LA), holding the likelihood tensors and their column entropies, so that the stress test can give both agents the less reliable camera (model_stress) without changing any global.
New: the factorized agent.dot_except contracts a tensor with the beliefs about all its factors but one; on it are built the per-factor prediction (predict_fact), the fixed-point iteration (perceive_fact), the variational free energy (vfe), and planning with factorized predictions (efe_fact, plan_fact).
The exact reference is the code of Part 5, with its functions renamed (predict_joint, perceive_exact, efe_exact, plan_exact) and the sensor model as an argument.
Both agents share plan_from and choose, so differences in their decisions come only from their scores.
LC_motion is removed, and the fixed-point iteration is controlled by K = 10 and the tolerance tol_fpi.
Notes on the agent code
Names: the code mirrors the agent’s math.
LA, LB, LC and LD are tuples, upright bold \(\mathbf{A}\), \(\mathbf{B}\), \(\mathbf{C}\) and \(\mathbf{D}\), built with tuple0 as in the environment. By the tear-off mnemonic, LA[m] is the italic-bold likelihood tensor \(\boldsymbol{A}^{[m]}\), LB[n] the transition tensor \(\boldsymbol{B}^{[n]}\), LC[m] the preference vector \(\boldsymbol{C}^{[m]}\), and LD[n] the initial prior \(\boldsymbol{D}^{[n]}\) of factor \(n\). Since the agent is given its model, LA and LB are simply the environment’s LAˣ and LBˣ, without the asterisk, and their dependencies are the environment’s LdepsA and LdepsB.
model(LA) bundles the likelihood tensors with LH, the entropy of each column of each \(\boldsymbol{A}^{[m]}\), i.e. how uncertain the sensor’s reading is for each combination of the factors it depends on. Every function that uses the sensors takes this bundle as its first argument, M. model_normal holds the camera of Part 5, model_stress the less reliable camera of the stress test; nothing else differs between them.
Lsₜ and Lsₜ₋₁ are the factorized agent’s beliefs \(\big(\boldsymbol{s}^{[t,0]}, \ldots, \boldsymbol{s}^{[t,4]}\big)\) at \(t\) and \(t-1\): a tuple of five probability vectors, with no joint belief behind them. Inside perceive_fact, Lprior holds the predictions \(\boldsymbol{s}^{[t,n]}_{\langle 0 \rangle}\), and _Fₜ the free energies \(F^{(t)}_{\langle k \rangle}\) after each iteration. Inside efe_fact, Ls holds the predicted beliefs \(\boldsymbol{s}_{\pi^{[p]}}^{[\tau,n]}\).
_s̃ₜ and _s̃ₜ₋₁ are the exact reference’s joint beliefs \(\tilde{\boldsymbol{s}}^{(t)}\) and \(\tilde{\boldsymbol{s}}^{(t-1)}\): single arrays of size \(3 \times 4 \times 2 \times 2 \times 2\) (sizes), as in Part 5. marginals(_s̃ₜ) returns their marginals \(\big(\tilde{\boldsymbol{s}}^{[t,0]}, \ldots, \tilde{\boldsymbol{s}}^{[t,4]}\big)\), which have the same format as Lsₜ, so the two can be compared directly.
_V, _G, _qπ and _qu are single italic-bold arrays: the policy matrix \(\boldsymbol{V}\), the expected free energies \(\boldsymbol{G}\), the posterior over policies \(Q\big(\pi^{(t)}\big)\) and the distribution over the next action \(Q\big(u^{[t,0]}\big)\).
K and tol_fpi control the fixed-point iteration: at most \(K = 10\) iterations, stopping early once no belief changes by more than tol_fpi.
The factorized agent: one contraction for everything. Every computation of the factorized agent combines a tensor with the beliefs about the factors it depends on, as in 4.5.4. A single function does this:
dot_except(arr, vs, keep) multiplies an array, whose axes correspond to the vectors in vs, with all of these vectors and adds up, except along the axis at position keep. With keep = 0, everything is added up and the result is a number, an expectation under the product of the beliefs; with keep set to a factor’s position, the result is a vector over that factor’s values.
dot_tail(T, vs) does the same for a tensor with a leading axis of its own, the next value of a factor or the observation, which is never contracted. It gives the vector over that leading axis.
On these two functions, the agent is built in four pieces:
predict_factor and predict_fact compute the prediction \(\boldsymbol{s}^{[t,n]}_{\langle 0 \rangle}\) of each factor, by contracting \(\boldsymbol{B}^{[n]}\) with the previous beliefs about the factors in \(\mathcal{D}_B^{[n]}\), using dot_tail. For the Room factor, the slice of the action is selected first. Since every prediction is computed from the previous beliefs, no update order is needed, unlike in the exact prediction.
perceive_fact is step 1, the perception: it predicts, then runs the fixed-point iteration, starting from the predictions. For factor \(n\), each sensor that depends on it contributes dot_except of the log-likelihood slice for the received observation, with keep at factor \(n\)’s position: its log-evidence, averaged over the current beliefs about the sensor’s other factors. The factors are updated in turn, each with the latest beliefs about the others, so every update lowers the free energy. It returns the beliefs and _Fₜ.
vfe computes the variational free energy: the complexity, a sum of KL divergences between the beliefs and their predictions, minus the accuracy, the expected log-likelihood of the received observations, using dot_except with keep = 0.
efe_fact is step 2, the evaluation of one policy: it rolls the factorized beliefs forward through the policy’s actions with predict_fact, and sums, over the steps and the three sensors, the risk and the ambiguity, with the expected observation from dot_tail and the expected entropy from dot_except. This follows line 16 of Algorithm 18. plan_fact scores all policies with it.
The exact reference: the agent of Part 5. The reference is the code of Part 5, with its functions renamed and the sensor model as an argument: predict_joint (with transition and update_order), perceive_exact, efe_exact and plan_exact, using the helpers marginal and spread between the joint belief and the tensors. Its logic is unchanged, so its beliefs are exact, and the decisions it would make are those of the Part 5 agent.
Shared: the decision.plan_from turns a vector of policy scores into the posterior over policies and the distribution over the next action, and choose picks the most probable action. This follows lines 19 and 20 of Algorithm 18. Both agents use them, so any difference in their decisions comes from their scores, i.e. from their beliefs and predictions alone.
Indexing. Tuples are 0-based as everywhere else, e.g. LA[0] to LA[2], LB[0] to LB[4] and Lsₜ[0] to Lsₜ[4], and so are the factor numbers in LdepsA and LdepsB, which is why [Ls[d] for d in LdepsA[m]] picks out the beliefs about the factors sensor \(m\) depends on, in the order of the tensor’s axes. Inside the italic-bold arrays, Julia’s ordinary 1-based indices are used: in the joint belief of the reference, factor \(n\) is axis \(n + 1\), and the values of each factor, the three actions, and the planning steps are numbered from 1. The math’s step \(\tau = 0, \ldots, H-1\) is column \(\tau + 1\) of _V. The position keep in dot_except is 1-based too: it is a position in the list of dependencies, found with findfirst.
As in Parts 4 and 5, this part does not use RxInfer: all computations are written directly in Julia, which keeps every quantity of Algorithm 18, and of the fixed-point iteration, visible.
Next, we define the agent and the time-stepping procedure.
## M: the sensor model, model_normal or model_stress## LC: the agent's preferences## exact: false for the factorized agent, true for the exact agent of Part 5 (also used as the reference)functioncreate_agent(; M, LC, exact=false) bₜ₋₁ =nothing## the previous belief: factorized beliefs Ls (a tuple), or the exact joint belief _s̃ (an array) ûₜ₋₁ =nothing## û^{[t-1,0]}: the previous action, as an index 1..3 (none before t = 0) bₜ =nothing## the current belief, in the same form ûₜ =nothing## û^{[t,0]}: the action this agent prefers at t _Fₜ =nothing## F^{(t)}_{⟨k⟩}: the free energies of the fixed-point iteration (factorized agent only) _G = _qπ = _qu =nothing## the current plan: G, Q(π^{(t)}), Q(u^{[t,0]})## step 1 and 2: update the belief with the new observations, then score all policies compute = (Lôₜ) ->beginif exact bₜ =perceive_exact(M, bₜ₋₁, Lôₜ, ûₜ₋₁) ## exact filtering over the 96 joint states _G, _qπ, _qu =plan_exact(M, bₜ, LC)else bₜ, _Fₜ =perceive_fact(M, bₜ₋₁, Lôₜ, ûₜ₋₁) ## prediction and fixed-point iteration, per factor _G, _qπ, _qu =plan_fact(M, bₜ, LC)endnothingend## step 3: choose the next action, returned as a one-hot action tuple û^{(t)}## (for the reference, this is the action the exact agent *would* have chosen, ũ^{[t,0]}) act = () ->begin ûₜ =choose(_qu)tuple0(Float64.(1:3.== ûₜ))end## predicted beliefs under policy p, τ = 0, ..., H-1 (for inspection):## each entry is a tuple of five beliefs, one per factor future = (p) ->begin b = bₜ Lsπː =Any[]for τ in1:Hif exact b =predict_joint(b, _V[p, τ]) ## s̃_π^{(τ)}: predicted joint beliefpush!(Lsπː, marginals(b))else b =predict_fact(b, _V[p, τ]) ## s_π^{[τ,n]}: predicted factorized beliefspush!(Lsπː, b)endendoffset0(Lsπː)end## move one tick forward: the current belief becomes the previous one, and so does the action;## by default this agent's own action, or, if given, the action actually taken (for the reference) slide = (Lû =nothing) ->begin bₜ₋₁ = bₜ ûₜ₋₁ = Lû ===nothing ? ûₜ :argmax(Lû[0])nothingend belief = () -> exact ? marginals(bₜ) : bₜ ## (s^{[t,0]}, ..., s^{[t,4]}), in the same form for both agents scores = () -> (_G, _qπ, _qu) ## the current plan, for inspection free_energy = () -> _Fₜ ## F^{(t)}_{⟨k⟩}, k = 1, 2, ... (nothing for the exact agent)return (compute, act, future, slide, belief, scores, free_energy)end
create_agent (generic function with 1 method)
Changes in create_agent compared to Part 5
One function, two agents:exact=false builds the factorized agent, exact=true the exact agent of Part 5. The sensor model is an argument, M, so that the stress test can give either agent the less reliable camera.
The reference can watch:slide optionally takes the action actually carried out, so an exact agent can move forward with the factorized agent’s actions while still computing, with act, the action it would have chosen itself.
Beliefs are returned in the same form for both agents: a tuple of five beliefs, one per factor, the marginals in the case of the exact agent.
New: free_energy returns the free energies of the last fixed-point iteration, for the factorized agent.
4.6 Agent Evaluation
Now it’s time to hand the Roomba over to the factorized agent. Does it still manage to clean the apartment while staying out of people’s way, and how far do its beliefs and decisions differ from those of the exact agent?
As in Parts 4 and 5, there is no separate inference step on finished data. The agent and the environment run together, tick by tick, for \(T + 1 = 600\) ticks. In the main run, the exact reference watches alongside. At each tick:
the environment reports the observations \(\hat{\mathbf{o}}^{(t)}\) from the three sensors: texture, motion and dirt,
the factorized agent predicts its beliefs and updates them by fixed-point iteration, one factor at a time (compute, step 1), and the reference updates its exact joint belief from the same observations,
the factorized agent scores all 27 policies by their expected free energy and chooses the next destination \(\hat{\boldsymbol{u}}^{[t,0]}\) (compute, step 2, and act); the reference does the same from its own belief, which gives the action \(\tilde{u}^{[t,0]}\) the exact agent would have chosen, but this action is never carried out,
the environment carries out the factorized agent’s action, moving the Roomba, moving the person, and letting dirt accumulate or be cleaned (execute), and
both move one tick forward (slide), the reference with the factorized agent’s action, so that it keeps seeing exactly the same data.
During the run, we record the true states \(\hat{\mathbf{s}}^{\ast(0:T)}\), the observations \(\hat{\mathbf{o}}^{(0:T)}\), the actions \(\hat{\mathbf{u}}^{(0:T-1)}\), the agent’s beliefs \(\boldsymbol{s}^{[0:T,\bullet]}\), the reference beliefs \(\tilde{\boldsymbol{s}}^{[0:T,\bullet]}\), the would-be actions \(\tilde{u}^{[0:T-1,0]}\), the free energies of the fixed-point iteration at each tick, and the runtime of both. From these, we compute the measures from 4.3.5: the belief discrepancy and the decision agreement, which isolate the approximation, and the disturbance, cleanliness and tracking, which show how the factorized agent does its job.
To see the combined effect of the approximation in closed loop, we also run the exact agent of Part 5 on its own, in the same environment with the same seed. Its path diverges from the factorized agent’s as soon as one of them makes a different choice, so comparing their disturbance and cleanliness shows how much the approximation changes the outcome, not just the beliefs. The random Roomba from 4.3 remains as a reference point.
Both comparisons are then repeated in the stress test, with the less reliable camera, model_stress, in the environment and in both agents. This gives four runs in total:
run
camera
agent acting
reference watching
main
normal
factorized
exact
exact
normal
exact
none
stress, main
less reliable
factorized
exact
stress, exact
less reliable
exact
none
Since the agents are given their model of the world, there is nothing to learn. The exact agent’s beliefs follow from exact Bayesian filtering over the 96 joint states. The factorized agent’s beliefs follow from the fixed-point iteration, which minimizes a variational free energy at each tick: unlike in Parts 4 and 5, there is a free energy curve to check, one per tick. Both agents’ choices follow from the expected free energy of their policies. Let’s start the runs!
Changes in the evaluation setup compared to Part 5
The acting agent is factorized, and the exact agent of Part 5 watches alongside as the reference, computing its beliefs and would-be actions on the same data without influencing the run.
The recorded data adds the reference beliefs, the would-be actions, the free energies of the fixed-point iteration, and the runtimes.
Four runs replace the comparison of the cleaning goal: the factorized agent with the watching reference, and the exact agent on its own, each with the normal camera and with the less reliable camera of the stress test.
The free energy returns, as one curve per tick from the fixed-point iteration.
4.6.1 Evaluate with simulated data
4.6.1.1 Naive approach {N/A}
4.6.1.2 Active inference approach
### Simulation parameters## Time steps t = 0, ..., T (600 ticks)T =599## Initial state: the Roomba in the bedroom, the person in the living room, all rooms clean## factor 0 (Room): 1 bedroom 2 livingroom 3 bathroom## factor 1 (Person): 1 bedroom 2 livingroom 3 bathroom 4 away## factors 2–4 (Dirt in bedroom, livingroom, bathroom): 1 clean 2 dirtyLŝˣ₀ =tuple0([1.0, 0.0, 0.0], ## ŝ^{*[0,0]}: Roomba in the bedroom [0.0, 1.0, 0.0, 0.0], ## ŝ^{*[0,1]}: person in the living room [1.0, 0.0], ## ŝ^{*[0,2]}: bedroom clean [1.0, 0.0], ## ŝ^{*[0,3]}: living room clean [1.0, 0.0]) ## ŝ^{*[0,4]}: bathroom clean → ŝ^{*(0)}
5-element OffsetArray(::Vector{Vector{Float64}}, 0:4) with eltype Vector{Float64} with indices 0:4:
[1.0, 0.0, 0.0]
[0.0, 1.0, 0.0, 0.0]
[1.0, 0.0]
[1.0, 0.0]
[1.0, 0.0]
Changes in the simulation parameters compared to Part 5: none. All four runs start from this state.
### OVERRIDES
Let’s take a look at our results. We start with a check: the exact agent, running on its own with the normal camera, uses the same code, preferences and seed as the agent with both preferences in Part 5, so it should reproduce Part 5’s results exactly: a disturbance of 11.0% and an average dirtiness of 8.3%. If it does, any difference that follows comes from the approximation.
With the normal camera, we expect the approximation to do well. The camera keeps the Roomba’s room nearly certain, and when the room is certain, the factorized update is exact, as 4.5.4 showed. So we expect:
small belief discrepancies for most factors at most ticks, with occasional larger ones at moments when the room is uncertain: after a camera misread, or after a move that may have failed,
the largest discrepancies for the person, since the person is sensed only through the motion sensor, which ties the person to the room,
a near-complete decision agreement: most of the time, the exact agent would have chosen the same action,
similar behavior: disturbance and cleanliness close to those of the exact agent, although the two runs can drift apart after a single different choice, and
a fixed-point iteration that settles quickly, in a few iterations per tick, with a free energy that decreases at every iteration.
In the stress test, the room becomes uncertain much more often, so we expect larger and more frequent belief discrepancies, a lower decision agreement, and, as discussed in 4.2, a factorized belief that is overconfident: sharper than the exact belief, committing to one explanation where the exact belief keeps several. Both agents will track the room less well with the less reliable camera, but the question is whether the factorized agent loses more than the exact agent does.
Finally, the runtime. With only 96 joint states, the exact agent is cheap, and the fixed-point iteration does several passes per tick, so the factorized agent need not be faster here. Its advantage lies in how its cost grows with the size of the model, which this small model cannot show.
Let’s see if it worked.
Changes in the expected results compared to Part 5
A reproduction check: the exact agent should reproduce Part 5’s results exactly, confirming that any difference comes from the approximation.
The expectations concern the approximation: small belief discrepancies with the normal camera, largest for the person, near-complete decision agreement and similar behavior; larger discrepancies, lower agreement and overconfidence in the stress test.
The runtime is not expected to favor the factorized agent at this size.
The behavioral expectations of Part 5 are dropped, since both agents share the same goals.
## one run in a fresh environment, with the same seed as the random Roomba (4.3)## M: the sensor model, for the environment and the agents (model_normal or model_stress)## exact: false for the factorized agent, true for the exact agent of Part 5## watch: true to let the exact reference watch alongside, on the same datafunctionrun_agent(; M, exact=false, watch=false, seed=42) (execute_ai, observe_ai, state_ai) =create_envir(; LAˣ=M.LA, LBˣ, LdepsA, LdepsB, Lŝˣ₀, seed) agent =create_agent(; M, LC, exact) (compute_ai, act_ai, future_ai, slide_ai, belief_ai, scores_ai, free_energy_ai) = agentif watch (compute_ref, act_ref, _, slide_ref, belief_ref, _, _) =create_agent(; M, LC, exact=true)end Lŝˣː =offset0(Vector{Any}(undef, T+1)) ## ŝ^{*(0:T)}: true states, each a tuple (Room, Person, Dirt ×3) Lôː =offset0(Vector{Any}(undef, T+1)) ## ô^{(0:T)}: observations, each a tuple (Texture, Motion, Dirt) Lûː =offset0(Vector{Any}(undef, T)) ## û^{(0:T-1)}: chosen actions, each a tuple (Destination) Lsː =offset0(Vector{Any}(undef, T+1)) ## s^{[0:T,•]}: the agent's beliefs, each a tuple of five Lquː =offset0(Vector{Any}(undef, T+1)) ## Q(u^{[t,0]}): the agent's action probabilities at each tick _Fː =offset0(Vector{Any}(undef, T+1)) ## F^{(t)}_{⟨k⟩}: free energies of the fixed-point iteration (factorized agent) runtimeː =offset0(zeros(T+1)) ## seconds spent in compute at each tick Ls̃ː =offset0(Vector{Any}(undef, T+1)) ## s̃^{[0:T,•]}: the reference beliefs (if watching) ũː =offset0(zeros(Int, T)) ## ũ^{[0:T-1,0]}: the actions the exact agent would have chosen (if watching) runtime_refː =offset0(zeros(T+1)) ## seconds spent in the reference's compute (if watching)for t =0:T t >0&&execute_ai(Lûː[t-1]) ## 4. the environment carries out the previous action Lŝˣː[t] =state_ai() Lôː[t] =observe_ai() ## 1. the environment reports the observations runtimeː[t] =@elapsedcompute_ai(Lôː[t]) ## 2.–3. the agent updates its beliefs and scores the policies Lsː[t] =belief_ai() Lquː[t] =scores_ai()[3] _Fː[t] =free_energy_ai()if watch runtime_refː[t] =@elapsedcompute_ref(Lôː[t]) ## the reference: same observations, exact belief Ls̃ː[t] =belief_ref()endif t < T Lûː[t] =act_ai() ## 3. the agent chooses the next destinationslide_ai() ## 5. the agent moves one tick forwardif watch ũː[t] =argmax(act_ref()[0]) ## the action the exact agent would have chosenslide_ref(Lûː[t]) ## ... but it moves forward with the action actually takenendendendreturn (; Lŝˣː, Lôː, Lûː, Lsː, Lquː, _Fː, runtimeː, Ls̃ː, ũː, runtime_refː, agent)end## the four runs of 4.6@time run_main =run_agent(; M=model_normal, watch=true) ## factorized agent, reference watching@time run_exact =run_agent(; M=model_normal, exact=true) ## exact agent on its own (should reproduce Part 5)@time run_stress =run_agent(; M=model_stress, watch=true) ## stress test: factorized agent, reference watching@time run_stress_exact =run_agent(; M=model_stress, exact=true) ## stress test: exact agent on its own## from here on, the recorded data and the agent functions are those of the main run(; Lŝˣː, Lôː, Lûː, Lsː, Lquː, _Fː, Ls̃ː, ũː) = run_main(compute_ai, act_ai, future_ai, slide_ai, belief_ai, scores_ai, free_energy_ai) = run_main.agent
74.481889 seconds (728.92 M allocations: 20.921 GiB, 6.11% gc time, 16.36% compilation time)
56.631703 seconds (682.98 M allocations: 19.124 GiB, 6.75% gc time, 0.04% compilation time)
62.531370 seconds (715.00 M allocations: 20.232 GiB, 6.77% gc time)
59.359194 seconds (682.97 M allocations: 19.124 GiB, 7.19% gc time)
run_agent takes the sensor model and the kind of agent,** and can let the exact reference watch: it updates from the same observations, records its beliefs and the action it would have chosen, and moves forward with the action actually taken.
New recordings: the free energies of the fixed-point iteration, the runtime of each agent per tick, and, when watching, the reference beliefs and would-be actions.
Four runs: the factorized agent with the reference watching, and the exact agent on its own, with the normal camera and in the stress test.
The familiar names refer to the main run, so the inspection cells that follow work as before.
## t, Roomba's room, person's location, dirt in bedroom, livingroom, bathroom: the true states of the main run[(t, room_labels[argmax(Lŝˣː[t][0])], person_labels[argmax(Lŝˣː[t][1])], (dirt_labels[argmax(Lŝˣː[t][n])] for n in2:4)...) for t in0:T]
[action_labels[argmax(Lû[0])] for Lû in Lûː] ## destination chosen by the factorized agent at t = 0, ..., T-1, by name
599-element OffsetArray(::Vector{String}, 0:598) with eltype String with indices 0:598:
"livingroom"
"livingroom"
"bedroom"
"bedroom"
"livingroom"
"bedroom"
"livingroom"
"bedroom"
"livingroom"
"bedroom"
⋮
"livingroom"
"livingroom"
"bedroom"
"bedroom"
"livingroom"
"livingroom"
"bedroom"
"bedroom"
"livingroom"
New: the ticks at which the exact agent, watching the same run, would have chosen a different action than the factorized agent.
## the ticks at which the exact agent would have chosen differently: (t, factorized agent's action, exact agent's action)[(t, action_labels[argmax(Lûː[t][0])], action_labels[ũː[t]]) for t in0:T-1 if argmax(Lûː[t][0]) != ũː[t]]
## t, the agent's belief about the Roomba's room, the exact reference belief, and their total variation distancetv(a, b) =sum(abs.(a .- b)) /2[(t, round.(Lsː[t][0], digits=2), round.(Ls̃ː[t][0], digits=2), round(tv(Lsː[t][0], Ls̃ː[t][0]), digits=3)) for t in0:T]
## the same for the person's location[(t, round.(Lsː[t][1], digits=2), round.(Ls̃ː[t][1], digits=2), round(tv(Lsː[t][1], Ls̃ː[t][1]), digits=3)) for t in0:T]
## the same for the dirt in the living room[(t, round.(Lsː[t][3], digits=2), round.(Ls̃ː[t][3], digits=2), round(tv(Lsː[t][3], Ls̃ː[t][3]), digits=3)) for t in0:T]
Each listing shows the agent’s belief next to the exact reference belief, with their total variation distance \(d^{[t,n]}\) from 4.3.5, for the Roomba’s room, the person’s location and the dirt in the living room.
## t, and the factorized agent's most probable value of each factor: room, person, dirt in bedroom, livingroom, bathroom[(t, room_labels[argmax(Lsː[t][0])], person_labels[argmax(Lsː[t][1])], (dirt_labels[argmax(Lsː[t][n])] for n in2:4)...) for t in0:T]
New: the ticks and factors at which the agent’s most probable value differs from the exact reference’s.
## the ticks and factors at which the agent's most probable value differs from the exact reference's:## (t, factor, the agent's most probable value, the reference's most probable value)labels = [room_labels, person_labels, dirt_labels, dirt_labels, dirt_labels]names = ["room", "person", "dirt bedroom", "dirt living room", "dirt bathroom"][(t, names[n+1], labels[n+1][argmax(Lsː[t][n])], labels[n+1][argmax(Ls̃ː[t][n])]) for t in0:T for n in0:4 if argmax(Lsː[t][n]) !=argmax(Ls̃ː[t][n])]
## t, true (room, person, dirt here), the factorized agent's most probable (room, person, dirt here),## the exact reference's most probable (room, person, dirt here), the action taken, and the exact agent's would-be action## "dirt here" is the dirt factor of the room the Roomba is actually in: factor r + 1 for room rhere(t) =argmax(Lŝˣː[t][0]) +1triple(Ls, t) = (room_labels[argmax(Ls[0])], person_labels[argmax(Ls[1])], dirt_labels[argmax(Ls[here(t)])])[(t,triple(Lŝˣː[t], t),triple(Lsː[t], t),triple(Ls̃ː[t], t), t < T ? action_labels[argmax(Lûː[t][0])] :"-", t < T ? action_labels[ũː[t]] :"-") for t in0:T]
Changes in the combined listing compared to Part 5
The exact reference is added: its most probable (room, person, dirt here), and the action the exact agent would have chosen, next to the factorized agent’s, so that each row shows where the two agree and where they differ.
[Lô[0] for Lô in Lôː] ## ô^{[t,0]} for t = 0, ..., T: the received Texture observations (main run)
_G, _qπ, _qu =scores_ai() ## the plan of the factorized agent at the last tick, t = T## the 27 policies, from lowest expected free energy (most probable) to highest[(join(action_labels[_V[p, :]], " → "), round(_G[p], digits=2), round(_qπ[p], digits=3)) for p insortperm(_G)]
New: the ticks at which the fixed-point iteration needed the most iterations, with the free energy after each iteration.
## the 5 ticks with the most fixed-point iterations: (t, number of iterations, F^{(t)}_{⟨k⟩} for k = 1, 2, ...)n_iter = [length(_Fː[t]) for t in0:T][(t, n_iter[t+1], round.(_Fː[t], digits=4)) for t in (0:T)[sortperm(n_iter, rev=true)[1:5]]]
Finally, we can check how well the approximation works. The evaluation has four parts, each computed for the normal camera and for the stress test.
1. The approximation on the same data. From the main runs, where the reference watched, we compute for each factor the belief discrepancy: the total variation distance \(d^{[t,n]}\) between the agent’s belief \(\boldsymbol{s}^{[t,n]}\) and the exact reference belief \(\tilde{\boldsymbol{s}}^{[t,n]}\), on average and at its largest. We also compute the decision agreement, the fraction of ticks at which the exact agent would have chosen the same action. To test whether the factorized beliefs are overconfident, we compare their entropy, a measure of how spread out a belief is, with that of the exact beliefs at the same ticks: if the factorized beliefs are systematically sharper, their average entropy is lower. A plot of \(d^{[t,n]}\) over time shows when the discrepancies occur, and whether they coincide with camera misreads, when the room is uncertain.
2. The fixed-point iteration. For each tick, we check that the free energy \(F^{(t)}_{\langle k \rangle}\) decreased at every iteration, and count how many iterations were needed before the beliefs settled. The free energy curves of a few ticks, including those that needed the most iterations, show the iteration at work.
3. The behavior in closed loop. From the true states, we compute the disturbance and the cleanliness of the factorized agent and the exact agent, each in its own run, and of the random Roomba from 4.3 as a reference point. The exact agent’s results with the normal camera should match those of Part 5 exactly, which checks the whole setup. We also compute the tracking accuracies of both agents per factor: how often each agent’s most probable value equals the true one. With the less reliable camera, both agents will track the room less well; the question is whether the factorized agent loses more than the exact one.
4. The runtime. We compare the median time per tick of the factorized agent and the exact reference, which ran on exactly the same data. The median is used because the first ticks include Julia’s compilation time.
The evaluation does not repeat the analysis of Part 5, i.e. whether the agent can overrule a misleading sensor reading: with the same model and the same goals, those questions were answered there. Here, the question is only how much of that behavior survives the approximation.
Changes in the evaluation measures compared to Part 5
The evaluation focuses on the approximation: belief discrepancy, decision agreement and belief entropy on the same data; the behavior of the factorized and the exact agent in closed loop; and the runtime per tick, each for the normal camera and the stress test.
The free energy of the fixed-point iteration is checked for a decrease at every iteration, with the number of iterations per tick.
New: belief entropy, to test whether the factorized beliefs are overconfident.
The analysis of overruled sensor readings is not repeated, since the model and the goals are those of Part 5.
## ---------- 1. the approximation on the same data (main runs, reference watching) ----------tv(a, b) =sum(abs.(a .- b)) /2## total variation distanceentropy(p) =-sum(x >0 ? x *log(x) :0.0for x in p) ## H[p], in natspct(x) =round.(100.* x, digits=1)factor_names = ["Roomba's room", "person's location", "dirt in bedroom", "dirt in living room", "dirt in bathroom"]functionapprox_measures(run) Ls, Ls̃ =parent(run.Lsː), parent(run.Ls̃ː) ## 1-based over ticks; factors stay 0-based d = [tv(Ls[i][n], Ls̃[i][n]) for i in1:T+1, n in0:4] ## d^{[t,n]}, (T+1)×5 dH = [entropy(Ls[i][n]) -entropy(Ls̃[i][n]) for i in1:T+1, n in0:4] ## H[s] - H[s̃] (; d, J_TV =vec(mean(d, dims=1)), ## J_TV(n): average distance d_max =vec(maximum(d, dims=1)), ## max_t d^{[t,n]} J_H =vec(mean(dH, dims=1)), ## J_H(n): negative means overconfident agree =mean(argmax(run.Lûː[t][0]) == run.ũː[t] for t in0:T-1)) ## J_agreeenda_main =approx_measures(run_main)a_stress =approx_measures(run_stress)for (name, a) in [("normal camera", a_main), ("stress test", a_stress)]println("---------- $name: decision agreement $(pct(a.agree))% ----------")println(rpad("factor", 22), rpad("avg distance", 14), rpad("max distance", 14), "entropy difference")for n in0:4println(rpad(factor_names[n+1], 22), rpad(round(a.J_TV[n+1], digits=4), 14),rpad(round(a.d_max[n+1], digits=3), 14), round(a.J_H[n+1], digits=4))endend## ---------- when do the discrepancies occur? ----------misreads(run) = [argmax(run.Lôː[t][0]) !=argmax(run.Lŝˣː[t][0]) for t in0:T] ## camera misread the floorfunctiondiscrepancy_plot(a, run, title) p =plot(0:T, a.d[:, 1], label="room", color=:blue, title=title, ylabel="distance d", legend=:outerright)plot!(p, 0:T, a.d[:, 2], label="person", color=:red)plot!(p, 0:T, vec(maximum(a.d[:, 3:5], dims=2)), label="dirt (largest)", color=:brown) m =misreads(run)scatter!(p, (0:T)[m], fill(-0.02, count(m)), markershape=:vline, markersize=4, color=:gray, label="camera misread") pendplot(discrepancy_plot(a_main, run_main, "Belief discrepancy, normal camera"),discrepancy_plot(a_stress, run_stress, "Belief discrepancy, stress test"), layout=@layout([a; b]), size=(900, 650), link=:all, xlabel=["""Time t"])
---------- normal camera: decision agreement 97.7% ----------
factor avg distance max distance entropy difference
Roomba's room 0.0116 0.408 -0.0269
person's location 0.0243 0.358 -0.0266
dirt in bedroom 0.0052 0.16 -0.0091
dirt in living room 0.0195 0.317 -0.0111
dirt in bathroom 0.0041 0.162 -0.0033
---------- stress test: decision agreement 91.5% ----------
factor avg distance max distance entropy difference
Roomba's room 0.0505 0.622 -0.091
person's location 0.0605 0.467 -0.0819
dirt in bedroom 0.0134 0.177 -0.0361
dirt in living room 0.0496 0.481 -0.0691
dirt in bathroom 0.0092 0.264 -0.0201
## ---------- 2. the fixed-point iteration ----------for (name, run) in [("normal camera", run_main), ("stress test", run_stress)] Fs =parent(run._Fː) ## one vector F^{(t)}_{⟨k⟩} per tick iters =length.(Fs) mono =all(all(diff(F) .<=1e-10) for F in Fs) ## decreased at every iteration?println("$name: free energy decreased at every iteration: $mono; ","iterations per tick: mean $(round(mean(iters), digits=2)), max $(maximum(iters)), ","stopped at K = $K at $(count(==(K), iters)) ticks")end## the free energy curves of the 5 ticks that needed the most iterations in the stress test,## shown as the distance to their final value (log scale), so that the convergence is visibleFs =parent(run_stress._Fː)top =sortperm(length.(Fs), rev=true)[1:5]p =plot(title="Fixed-point iteration, stress test: F(k) - F(final)", xlabel="iteration k", yscale=:log10, legend=:outerright)for i in top F = Fs[i]length(F) >1&&plot!(p, 1:length(F)-1, F[1:end-1] .- F[end] .+1e-12, marker=:circle, label="t = $(i-1)")endp
normal camera: free energy decreased at every iteration: true; iterations per tick: mean 4.99, max 10, stopped at K = 10 at 28 ticks
stress test: free energy decreased at every iteration: true; iterations per tick: mean 6.46, max 10, stopped at K = 10 at 77 ticks
## ---------- 3. the behavior in closed loop ----------## measures(Lŝˣː): disturbance, cleanliness and coverage of one run, as in Part 5functionmeasures(Lŝˣː) rooms = [argmax(Lŝˣ[0]) for Lŝˣ inparent(Lŝˣː)] persons = [argmax(Lŝˣ[1]) for Lŝˣ inparent(Lŝˣː)] dirty = [[argmax(Lŝˣ[n]) ==2 for Lŝˣ inparent(Lŝˣː)] for n in2:4] ## dirt in bedroom, livingroom, bathroom (; disturbance =mean(rooms .== persons), ## J_disturb dirtiness = [mean(dirty[r]) for r in1:3], ## J_dirty(r) coverage = [mean(rooms .== r) for r in1:3]) ## backgroundend## J_track(n): how often each agent's most probable value equals the true onetracking(run) = [mean(argmax(run.Lsː[t][n]) ==argmax(run.Lŝˣː[t][n]) for t in0:T) for n in0:4]runs = [("factorized, normal", run_main), ("exact, normal", run_exact), ("factorized, stress", run_stress), ("exact, stress", run_stress_exact)]m =Dict(name =>measures(run.Lŝˣː) for (name, run) in runs)println(rpad("", 22), rpad("disturbance", 13), rpad("dirty (bed, living, bath)", 28), rpad("avg dirty", 11), "coverage")println(rpad("random Roomba", 22), rpad("$(pct(disturbance_random))%", 13), rpad("$(pct(dirtiness_random))%", 28),rpad("$(pct(mean(dirtiness_random)))%", 11), "$(pct(coverage_random))%")for (name, _) in runs r = m[name]println(rpad(name, 22), rpad("$(pct(r.disturbance))%", 13), rpad("$(pct(r.dirtiness))%", 28),rpad("$(pct(mean(r.dirtiness)))%", 11), "$(pct(r.coverage))%")endprintln("reproduction check (Part 5: 11.0% disturbance, 8.3% average dirty): exact, normal gives ","$(pct(m["exact, normal"].disturbance))% and $(pct(mean(m["exact, normal"].dirtiness)))%")println()println(rpad("tracking", 22), join(rpad.(factor_names, 22)))for (name, run) in runsprintln(rpad(name, 22), join(rpad.(string.(pct(tracking(run))) .*"%", 22)))end## cleanliness per room, normal camera and stress testx =1:3functiondirt_plot(f, e, title) p =bar(x .-0.27, pct(dirtiness_random), bar_width=0.27, label="random Roomba", color=:gray)bar!(p, x, pct(m[e].dirtiness), bar_width=0.27, label="exact agent", color=:steelblue)bar!(p, x .+0.27, pct(m[f].dirtiness), bar_width=0.27, label="factorized agent", color=:green, xticks=(x, ["Bedroom", "Living room", "Bathroom"]), ylabel="% of ticks dirty", legend=:outerright, title="$title — disturbance: exact $(pct(m[e].disturbance))%, factorized $(pct(m[f].disturbance))%", titlefontsize=10)endplot(dirt_plot("factorized, normal", "exact, normal", "Normal camera"),dirt_plot("factorized, stress", "exact, stress", "Stress test"), layout=@layout([a; b]), size=(900, 650), link=:y)
disturbance dirty (bed, living, bath) avg dirty coverage
random Roomba 15.8% [8.7, 20.7, 3.2]% 10.8% [22.5, 25.8, 51.7]%
factorized, normal 11.2% [5.8, 19.3, 2.5]% 9.2% [32.2, 17.7, 50.2]%
exact, normal 11.0% [4.3, 17.8, 2.7]% 8.3% [32.2, 17.8, 50.0]%
factorized, stress 12.0% [3.3, 15.7, 3.3]% 7.4% [35.3, 19.2, 45.5]%
exact, stress 11.5% [3.8, 25.2, 3.0]% 10.7% [35.5, 19.8, 44.7]%
reproduction check (Part 5: 11.0% disturbance, 8.3% average dirty): exact, normal gives 11.0% and 8.3%
tracking Roomba's room person's location dirt in bedroom dirt in living room dirt in bathroom
factorized, normal 94.5% 71.5% 95.8% 83.7% 97.5%
exact, normal 94.7% 70.3% 97.3% 84.8% 98.0%
factorized, stress 85.0% 74.3% 97.0% 87.8% 96.3%
exact, stress 87.2% 73.3% 96.0% 79.0% 97.2%
## ---------- 4. the runtime per tick (median, since the first ticks include compilation) ----------for (name, run) in [("normal camera", run_main), ("stress test", run_stress)] t_fact =median(parent(run.runtimeː)) t_exact =median(parent(run.runtime_refː))println("$name: factorized agent $(round(1000* t_fact, digits=2)) ms per tick, ","exact reference $(round(1000* t_exact, digits=2)) ms per tick ","(ratio exact / factorized: $(round(t_exact / t_fact, digits=2)))")end
normal camera: factorized agent 8.12 ms per tick, exact reference 89.19 ms per tick (ratio exact / factorized: 10.98)
stress test: factorized agent 8.37 ms per tick, exact reference 88.91 ms per tick (ratio exact / factorized: 10.62)
Changes in the evaluation code compared to Part 5
Four cells instead of one, following the four parts of the evaluation.
New: the approximation on the same data: belief discrepancy (average and largest), entropy difference and decision agreement per factor, and a plot of the discrepancies over time with the camera misreads marked.
New: the fixed-point iteration: a check that the free energy decreased at every iteration, the number of iterations per tick, and the free energy curves of the hardest ticks.
The behavior measures are computed for the factorized and the exact agent with both cameras, with the random Roomba as a reference point, and with a check that the exact agent reproduces Part 5. Tracking is computed for every run.
New: the runtime per tick of both agents on the same data.
Removed: the comparison with the Part 4 agent, and the analysis of overruled sensor readings.
4.7 Conclusion and outlook
In this part, the agent got a second goal. Besides staying out of people’s way, as in Part 4, it now had to clean. To express this, the model of the environment grew from a single state factor to five: the Roomba’s room, the person’s location, and the dirt in each of the three rooms. Four ingredients made this possible:
Several state factors with dependencies. Each factor’s transitions, and each sensor’s readings, depend on only a few factors: a room’s dirt on itself and on where the Roomba was, the motion sensor on where the Roomba and the person are. These dependencies, \(\mathcal{D}_A^{[m]}\) and \(\mathcal{D}_B^{[n]}\), are the core of multi-factor modeling. They keep every tensor small: the 96 joint states never had to be written out.
A joint belief, with marginals for evaluation. Although the model factorizes, the agent’s beliefs do not: a motion reading says something about the Roomba and the person together. The agent therefore filtered exactly over the 96 joint states, stored as an array with one axis per factor, and the beliefs about single factors followed by adding up.
Two preferences. The agent preferred to observe no motion and dirt picked up. The latter, rather than a clean floor, is what draws the agent to rooms where its presence makes a difference.
A longer horizon. With \(H = 3\), the agent could see the pay-off of going to the living room, two moves away from the bedroom.
As in Part 4, all of this was written directly in Julia, with code that is generic over factors and sensors: the dependency lists tell each function which axes of the joint belief a tensor acts on.
The cleaning goal works, more modestly than expected. The three Roombas compare as follows:
disturbance
dirty: bedroom
dirty: living room
dirty: bathroom
dirty: average
random Roomba
15.8%
8.7%
20.7%
3.2%
10.8%
Part 4 agent (no motion only)
8.8%
6.5%
26.2%
2.8%
11.8%
both preferences
11.0%
4.3%
17.8%
2.7%
8.3%
The Part 4 agent avoided the person best, but it confirms the problem that motivated this part: it left the living room dirty 26.2% of the time, more than the random Roomba. Adding the cleaning goal brought this down to 17.8%, and the average over the three rooms from 11.8% to 8.3%, the cleanest of the three Roombas. The price was 2.2 percentage points of extra disturbance (11.0% against 8.8%), still about 30% less than the random Roomba.
How it achieved this is visible in the coverage. The agent with both preferences spent only slightly more time in the living room than the Part 4 agent (17.8% against 15.3%), yet left it dirty much less often. It did not clean by spending more time in the living room, but by going there at the right moments: when it expected dirt to have built up. The dirt plot shows the same pattern: periods with a dirty room under the Roomba are short, because the agent stays while it picks up dirt and moves on once the room is clean.
Still, the living room remains the dirtiest room by far, and the improvement over the random Roomba is modest (17.8% against 20.7%). The two preferences pull in opposite directions there, and with the motion preference weighing more than the dirt preference, avoiding the person often wins. While the Roomba is elsewhere, the agent cannot see the person, so its belief drifts towards the living room, where the person usually is: the agent expects motion there more often than it would find it.
Tracking: the person now carries over. The agent tracked the Roomba’s room correctly 94.7% of the time, and whether the person was in the Roomba’s room 96.0% of the time. The latter is a real improvement on Part 4. There, the agent was wrong at every one of the 38 ticks at which the motion sensor was wrong: its belief simply followed the latest reading. Here, the motion sensor was wrong at 31 ticks, and the agent still got it right at 51.6% of those. The reason is in the model: the person now moves independently of the Roomba and tends to stay put, so earlier observations carry over, whereas in Part 4 occupancy was drawn fresh each time the Roomba entered a room. The dirt sensor was wrong at 27 ticks, and the agent overruled it at 77.8% of those, since dirt persists until it is cleaned and the agent knows where it has recently cleaned. The camera, finally, misread the floor at 53 ticks, at 50.9% of which the agent still inferred the right room.
The tracking of the other factors needs more care. The person’s location was tracked correctly only 70.3% of the time: the agent senses the person only when they are in the Roomba’s room, and otherwise has to rely on its predictions. The dirt in the bedroom (97.3%) and the bathroom (98.0%) seems well tracked, but these rooms were rarely dirty, so always guessing clean would already have been right 95.7% and 97.3% of the time. The living room, which gets dirty fastest, is the real test, and there the agent was right 84.8% of the time, against 82.2% for always guessing clean. The agent knows the dirt where it is; elsewhere, it knows little more than the base rates.
These results come from a single run of 600 ticks per Roomba, so the exact percentages should not be over-interpreted. The differences in disturbance between the two agents, in particular, are small. The patterns, though, follow from the model: a Roomba that avoids people neglects the busiest room, a cleaning goal brings it back at well-chosen moments, and the person’s location persists in the agent’s belief now that it is a factor of its own.
Possible next steps:
Tuning the balance. The relative strength of \(\boldsymbol{C}^{[1]}\) and \(\boldsymbol{C}^{[2]}\) sets the trade-off between disturbance and cleanliness. Running the agents for a range of preference strengths, averaged over several random seeds, would trace out this trade-off, and show how much cleaner the living room can get for each extra percent of disturbance. A precision parameter for how decisively the agent picks the best policy is worth including.
Factorized beliefs. Exact filtering over the joint states works for 96 combinations, but their number grows multiplicatively with every factor. Approximating the belief by a product of per-factor beliefs, as is common in larger active inference models, makes much larger models feasible, at the cost of losing exactly the relations between factors that the joint belief captures here.
Valuing information. The agent’s weak knowledge of the person’s location and the living room’s dirt suggests that checking on rooms could be worth more than it currently is. A longer planning horizon, or a stronger weight on reducing uncertainty, may make the agent value this.
Learning while acting. Here, the agent was given its tensors. Letting it learn them while it acts, e.g. how fast each room gets dirty, or where the person tends to be, brings in the trade-off between exploring to learn and exploiting what it already knows.
Richer environments. More exogenous forces, such as doors that are closed at random, a person whose routine depends on the time of day, or more than one person, bring the model closer to the scheduling project this series is building up to.
Planning as inference with RxInfer. For larger models, where neither the joint states nor the policies can be enumerated, formulating planning as inference in RxInfer is the natural next tool.
With this part, the Roomba both cleans and stays out of the way. The balance between the two can still be improved, and the model is now built from factors that can be added one at a time.
4.7 Conclusion and outlook
In this part, the model stayed exactly as in Part 5: five state factors, three sensors, and the two goals of cleaning and not disturbing anyone. What changed was the agent’s belief. Instead of an exact joint belief over all 96 combinations of the factors, the agent kept one belief per factor, 13 numbers in total, and treated their product as its belief about the whole state: the mean-field approximation. Four ingredients made this possible:
Prediction per factor. Each factor’s belief was predicted from the beliefs about the factors its transition depends on, by contracting its transition tensor with them. For a product belief, this gives exactly the exact filter’s prediction for each factor; what it loses are the relations between the predicted factors.
A fixed-point iteration. Since each sensor ties several factors together, the beliefs were updated in turn, each against the evidence of its sensors averaged over the current beliefs about the sensors’ other factors, until they settled. This iteration minimizes a variational free energy, which brought back the free energy curve that Parts 4 and 5 lacked.
Planning with factorized predictions. Expected observations were computed as if the factors were independent, exactly as in the example of 4.1.
An exact reference that watches. The exact agent of Part 5 ran alongside the factorized agent, on the same observations and actions, computing its beliefs and the actions it would have chosen, without influencing the run. This isolated the approximation completely.
To see where the approximation breaks down, every comparison was repeated in a stress test with a less reliable camera, which misreads the floor with probability 0.4 instead of 0.1.
The approximation is good with the normal camera, and clearly worse in the stress test. On the same data, the factorized and exact beliefs compare as follows, with the average and largest total variation distance \(d\) and the average entropy difference \(J_{\text{H}}\):
normal: avg \(d\)
normal: max \(d\)
normal: \(J_{\text{H}}\)
stress: avg \(d\)
stress: max \(d\)
stress: \(J_{\text{H}}\)
Roomba’s room
0.012
0.41
−0.027
0.051
0.62
−0.091
person’s location
0.024
0.36
−0.027
0.061
0.47
−0.082
dirt in bedroom
0.005
0.16
−0.009
0.013
0.18
−0.036
dirt in living room
0.020
0.32
−0.011
0.050
0.48
−0.069
dirt in bathroom
0.004
0.16
−0.003
0.009
0.26
−0.020
decision agreement
97.7%
91.5%
With the normal camera, the beliefs are close on average, and the exact agent would have chosen the same action at 97.7% of the ticks, disagreeing at about 14 of 599. In the stress test, the average distances grow by a factor of 2 to 4, the largest rises to 0.62, and the agreement drops to 91.5%, about 51 disagreements. The discrepancy plot shows when the differences occur: with the normal camera, the spikes coincide with camera misreads, the moments when the Roomba’s room is uncertain; in the stress test, misreads are frequent, and the distance rarely returns to zero. The largest discrepancies of all occur in the first few dozen ticks of the stress test, when every belief still starts from a uniform prior and the noisy camera has not yet settled where the Roomba is. This confirms the main expectation: mean-field is safe when the factor that ties the others together, here the Roomba’s room, is nearly certain, and degrades as that factor becomes uncertain.
The approximation is overconfident, as expected. The entropy difference is negative for every factor, under both cameras: the factorized beliefs are consistently sharper than the exact ones, and three to four times more so in the stress test. Forced to describe a coupled situation with independent factors, the factorized belief commits to one explanation where the exact belief keeps several.
Two findings differ from what we expected. First, the room is affected as much as the person. The person has the largest average distance, as expected, but the room has the largest single distance under both cameras, and in the stress test its average is close to the person’s. The reason is that the factorized update works in both directions: a motion or dirt reading also pushes the belief about the room, and by committing to one explanation, it can push it further than the exact posterior would. Second, the dirt in the living room is the third trouble spot, well above the other two rooms. Its dirt depends on where the Roomba was, and the living room is where the Roomba’s comings and goings matter most; whenever the room is uncertain, the factorized prediction loses the link between “the Roomba was there” and “the room was cleaned”.
The fixed-point iteration works as the theory says. The free energy decreased at every iteration of every tick, under both cameras. The iteration settled after 5.0 iterations on average with the normal camera, and 6.5 in the stress test, where the factors are harder to reconcile. It reached the limit of \(K = 10\) at 28 and 77 ticks, but harmlessly: the free energy curves fall along a straight line on a log scale, and after nine iterations the remaining gap is \(10^{-8}\) to \(10^{-11}\), far below anything that could change a decision.
In closed loop, the outcome is close with the normal camera, and inconclusive in the stress test. First a check: running on its own, the exact agent reproduced Part 5 exactly, with a disturbance of 11.0% and an average dirtiness of 8.3%, so every difference comes from the approximation. The two agents then compare as follows, with the random Roomba as a reference:
disturbance
dirty: bedroom
dirty: living room
dirty: bathroom
dirty: average
random Roomba
15.8%
8.7%
20.7%
3.2%
10.8%
exact, normal
11.0%
4.3%
17.8%
2.7%
8.3%
factorized, normal
11.2%
5.8%
19.3%
2.5%
9.2%
exact, stress
11.5%
3.8%
25.2%
3.0%
10.7%
factorized, stress
12.0%
3.3%
15.7%
3.3%
7.4%
With the normal camera, the factorized agent behaves almost like the exact one: slightly more disturbance and slightly more dirt, with nearly identical coverage, and with tracking accuracies within 1.5 percentage points for every factor. It goes to the same rooms as often, just slightly less well-timed. In the stress test, the factorized agent kept the apartment cleaner than the exact agent in this run, mostly in the living room. This should not be read as the approximation improving behavior. The two runs diverge after the first different choice, and with about 51 such choices, they end up in quite different situations. Both agents spent almost the same time in the living room (19.2% and 19.8%), so the difference looks more like timing and chance in a single run than like a systematic effect. For the same reason, the tracking accuracies of the two stress runs are hard to compare, since the rooms were dirty at different rates. Only the room tracking compares fairly, and there the factorized agent lost a little more than the exact one (85.0% against 87.2%), as expected.
The runtime saving is larger than expected. On the same data, the factorized agent needed about 5 ms per tick, the exact reference about 80 ms: a factor of almost 16, under both cameras. We expected little or no saving with only 96 joint states, and that expectation was wrong. Most of the time goes into planning: every tick, each of the 27 policies is predicted three steps ahead, 81 predictions in all. Each exact prediction moves the probability of all 96 combinations through every factor’s transition, while each factorized prediction sums over at most a few combinations per factor. Part of the factor of 16 reflects how the two are implemented, but the direction is fundamental: the factorized cost grows with the size of the largest tensor, the exact cost with the number of joint states.
These results come from a single run of 600 ticks per agent and camera, so the exact percentages should not be over-interpreted, the closed-loop differences least of all. The comparisons on the same data are more robust, since they compare the two beliefs tick by tick on identical observations.
A note on rerunning the notebook. All results except the runtimes are reproducible: the environment draws its randomness from its own generator with a fixed seed, the random Roomba from another, and the agents use no randomness at all, since perception and planning are deterministic and the action is chosen with argmax. Rerunning the notebook therefore gives exactly the same observations, beliefs, actions and measures, as the reproduction check of the exact agent against Part 5 also shows. Only the runtimes change from run to run, since they measure wall-clock time; the conclusions above rely on their rounded values and on the ratio between them, not on their exact values. This holds for the same versions of Julia and its packages: a newer version of Julia or of Distributions.jl may draw samples differently, which would change the runs and the numbers reported here. The notebook’s Manifest.toml records the exact versions used.
Possible next steps:
Several seeds. Repeating the closed-loop runs over several random seeds would show whether the differences in behavior, in particular in the stress test, exceed the variation between runs.
A larger model. The real case for factorized beliefs is a model in which the joint states can no longer be enumerated: more rooms, more people, more factors. There, the exact reference is no longer available, so the approximation would have to be judged by behavior alone, which this part has shown to be possible with care.
Structured approximations. Between the exact joint belief and a full product lies a middle ground: keeping a joint belief over the factors that are most tightly coupled, here the Roomba’s room and the person’s location, and factorizing the rest. This would target exactly the discrepancies found here.
Tuning the balance. As in Part 5, the relative strength of the two preferences sets the trade-off between disturbance and cleanliness, and is worth exploring systematically, together with a precision parameter for how decisively the agent picks the best policy.
Learning while acting. Letting the agent learn its tensors while it acts, e.g. how fast each room gets dirty, or where the person tends to be, brings in the trade-off between exploring to learn and exploiting what it already knows.
Message passing with RxInfer. The fixed-point iteration of this part is a form of variational message passing on a factor graph, which is what RxInfer does natively. For larger models, where the policies can no longer be enumerated either, formulating both perception and planning as inference in RxInfer is the natural next tool.
With this part, the Roomba’s agent has become cheaper to run without changing what it wants: it keeps one belief per factor, and loses little while it knows where it is. When it does not, its beliefs become overconfident, and that is the price of the approximation.