This is the fifth part of an effort to add control to the Hidden Markov Model RxInfer example. Eventually this work will be used in a more elaborate scheduling project making use of active inference.
In the previous parts, the agent first watched a Roomba move through a 3-room apartment, consisting of a bedroom, a living room, and a bathroom, inferring which room it was in (Parts 1 and 2) and whether that room was occupied (Part 3). In Part 4, the agent took control, with a single goal: to avoid disturbing people. It succeeded, with about a third less disturbance than a Roomba moving at random, but it did so by largely avoiding the living room, the room most often occupied. A Roomba that does not clean the busiest room is not much of a cleaning robot.
In this part, the agent gets a second goal: to clean. It now has to balance two goals that can conflict, cleaning where it is needed and not disturbing anyone, and the most interesting case is the living room: it gets dirty fastest, but it is also the room most often in use. We expect the agent to clean it when it is quiet, rather than avoiding it altogether.
Several state factors. To express this, the model of the environment grows from one state factor to five, each describing one aspect of the world:
the room the Roomba is in (bedroom, living room or bathroom), which the agent controls,
where the person is (bedroom, living room, bathroom, or away), who moves around independently, and
for each of the three rooms, whether it is dirty (clean or dirty). Rooms get dirty over time, the living room fastest, and a room only gets clean when the Roomba is in it.
In Parts 3 and 4, room and occupancy were combined into a single state factor. Splitting the state into factors is the standard way of building larger active inference models, and gaining experience with it is a main purpose of this part. Each factor’s transitions depend only on a few factors, e.g. a room’s dirt depends on itself and on where the Roomba is, and each sensor depends only on a few factors, e.g. the motion sensor depends on where the Roomba is and where the person is. Specifying these dependencies is the core of multi-factor modeling.
The Roomba has three sensors:
a camera that recognizes the floor texture (hardwood in the bedroom, carpet in the living room, tiles in the bathroom),
a motion sensor that detects movement, the most direct evidence that the person is in the same room as the Roomba, and
new in this part, a dirt sensor that reports whether the Roomba’s brushes are picking up dirt, i.e. whether the room it is in is dirty.
The light sensor from the previous parts added little and is dropped, to keep the model focused.
Two preferences. As before, goals are expressed as preferences over observations. The agent prefers to observe no motion, as in Part 4, and now also prefers to observe dirt being picked up, which draws it to dirty rooms. A preference for observing a clean floor would backfire: the agent would prefer to sit in rooms that are already clean. How strongly the agent holds each preference determines its balance between cleaning and avoiding people.
Planning. As in Part 4, the agent looks a few steps ahead, scores each candidate sequence of moves by its expected free energy, and carries out the first move of the best sequences. Since getting from the bedroom to the living room takes two moves, and cleaning takes at least one more, the agent now looks three steps ahead.
To keep the focus on multi-factor modeling and on balancing two goals, the agent is again given its model of the world rather than learning it. Internally, it combines the five factors into their joint states, \(3 \cdot 4 \cdot 2 \cdot 2 \cdot 2 = 96\) in total, so that perception and planning can still be computed exactly. A later part can replace this with a factorized approximation, which is what makes much larger models feasible.
The structure of the code is the same as in Part 4:
create_envir()
execute(): carries out the chosen action, moving the Roomba and the person, and updating dirt
observe()
state() (true state, for evaluation only)
create_agent()
act(): returns the chosen action
future(): predicts future states and observations for the candidate actions
compute(): infers the current state and evaluates the candidate actions
slide(): moves the agent’s time window one step forward
As in the previous parts, and as in the active inference literature, \(\mathbf{A}\) holds the observation (likelihood) matrices and \(\mathbf{B}\) the state transition matrices, swapped compared to the original RxInfer example, which follows the engineering convention of using \(A\) for state transitions. With several factors, these matrices become tensors, with one axis for each factor they depend on.
versioninfo() ## Julia version
Julia Version 1.12.7
Commit 6d172b025e4 (2026-08-15 08:05 UTC)
Build Info:
Official https://julialang.org release
Platform Info:
OS: Linux (x86_64-linux-gnu)
CPU: 12 × Intel(R) Core(TM) i7-8700B CPU @ 3.20GHz
WORD_SIZE: 64
LLVM: libLLVM-18.1.7 (ORCJIT, skylake)
GC: Built with stock GC
Threads: 1 default, 1 interactive, 1 GC (on 12 virtual cores)
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
Project No packages added to or removed from `~/.julia/environments/v1.12/Project.toml`
Manifest No packages added to or removed from `~/.julia/environments/v1.12/Manifest.toml`
## tuple0(x...): build a 0-based tuple (x^{[0]}, ..., x^{[N]}), mirroring the math.## - each argument becomes one entry: tuple0(v) wraps v; it never re-indexes v itself## - the index range 0:N is derived from the number of arguments## - entries may differ in shape, e.g. likelihood matrices for different modalities## Returns an OffsetVector (mutable), not a Julia Tuple.tuple0(x...) =OffsetVector(collect(x), 0:length(x)-1)## offset0(v): re-index an existing vector from 0, e.g. a trajectory over t = 0, ..., T.## Unlike tuple0, it does not wrap v; it shifts v's own indices.offset0(v) =OffsetVector(v, 0:length(v)-1)
offset0 (generic function with 1 method)
Why tuple0() rather than OffsetVector() directly?
It prevents a level mix-up. With OffsetVector, it is easy to offset the wrong level:
OffsetVector([1.0, 0.0, 0.0], 0:2) ## the one-hot vector itself, re-indexed 0:2 (wrong level)OffsetVector([[1.0, 0.0, 0.0]], 0:0) ## a tuple containing the one-hot vector (intended)tuple0([1.0, 0.0, 0.0]) ## the same as the second line, with no way to get it wrong
The first line is valid code but means something else: the one-hot vector $\hat{\boldsymbol{s}}^{\ast[0,0]}$ itself, now indexed from 0, rather than the tuple $\hat{\mathbf{s}}^{\ast(0)} = \big(\hat{\boldsymbol{s}}^{\ast[0,0]}\big)$. Nothing would complain, and later code that indexes `[0]` would get a single number instead of a vector. `tuple0` always wraps its arguments as *entries*, so this mistake cannot happen.
The index range is computed automatically. With OffsetVector, the range is written by hand and has to agree with the number of entries: adding a second modality to LAˣ while forgetting to change 0:0 to 0:1 gives an error at best. tuple0 derives the range from the number of arguments, so adding or removing an entry is a one-line change.
It reads like the math.tuple0(A0, A1) looks like \(\big(\boldsymbol{A}^{[0]}, \boldsymbol{A}^{[1]}\big)\): entries separated by commas, nothing else. OffsetVector([A0, A1], 0:1) adds square brackets and a range that have nothing to do with the math.
It states the intent.OffsetVector says “an array with shifted indices”, which could be anything. tuple0 says “a tuple of factors or modalities, as in the math”, the specific concept from (10.33). And since all tuples are built the same way, a search for tuple0( finds every one of them.
There is one place to change things. Changing how tuples work, e.g. switching to an immutable 0-based type, or adding a check that every column of every \(\boldsymbol{A}\) and \(\boldsymbol{B}\) sums to 1, means changing one line instead of every place a tuple is built.
Entries of different shapes just work.collect builds a vector of whatever is passed in, so tuple0(A0, A1) with a \(3\times 3\) and a \(2\times 3\) matrix gives a vector of matrices, which is exactly what a tuple of likelihood matrices with different numbers of observations needs.
2 DATA UNDERSTANDING
The data is simulated and, as in Part 4, it arises from the interaction between the agent and the environment. At each tick, the agent observes, chooses an action, and the environment responds, so the Roomba’s path depends on the agent’s choices. A run covers \(T + 1 = 600\) ticks, \(t = 0, \ldots, T\), and records three kinds of one-hot encoded vectors:
Observations, which the agent receives: the observation tuple \(\hat{\mathbf{o}}^{(t)} = \big(\hat{\boldsymbol{o}}^{[t,0]}, \hat{\boldsymbol{o}}^{[t,1]}, \hat{\boldsymbol{o}}^{[t,2]}\big)\), with one entry per sensor:
\(\hat{\boldsymbol{o}}^{[t,0]}\) (Texture) over \(\{\text{hardwood}, \text{carpet}, \text{tiles}\}\)
\(\hat{\boldsymbol{o}}^{[t,1]}\) (Motion) over \(\{\text{no motion}, \text{motion}\}\)
\(\hat{\boldsymbol{o}}^{[t,2]}\) (Dirt) over \(\{\text{nothing picked up}, \text{dirt picked up}\}\)
Actions, which the agent chooses: the action tuple \(\hat{\mathbf{u}}^{(t)} = \big(\hat{\boldsymbol{u}}^{[t,0]}\big)\), whose entry \(\hat{\boldsymbol{u}}^{[t,0]}\) is over the room the Roomba heads for: \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\). Heading for the room it is already in means staying put. The action chosen at tick \(t\) takes effect at tick \(t+1\), so actions are recorded for \(t = 0, \ldots, T-1\).
True states, which the agent never sees: the state tuple \(\hat{\mathbf{s}}^{\ast(t)} = \big(\hat{\boldsymbol{s}}^{\ast[t,0]}, \ldots, \hat{\boldsymbol{s}}^{\ast[t,4]}\big)\), now with one entry per state factor:
\(\hat{\boldsymbol{s}}^{\ast[t,0]}\) (Room) over \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\), where the Roomba is
\(\hat{\boldsymbol{s}}^{\ast[t,1]}\) (Person) over \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}, \text{away}\}\), where the person is
\(\hat{\boldsymbol{s}}^{\ast[t,2]}\), \(\hat{\boldsymbol{s}}^{\ast[t,3]}\) and \(\hat{\boldsymbol{s}}^{\ast[t,4]}\) (Dirt in the bedroom, living room and bathroom), each over \(\{\text{clean}, \text{dirty}\}\)
In each one-hot vector, the component that is 1 marks the value that occurred: for the state factors, where the Roomba and the person actually are, and which rooms are actually dirty; for \(\hat{\boldsymbol{o}}^{[t,0]}\), \(\hat{\boldsymbol{o}}^{[t,1]}\) and \(\hat{\boldsymbol{o}}^{[t,2]}\), the texture, motion and dirt the sensors report; for \(\hat{\boldsymbol{u}}^{[t,0]}\), the room the agent chose to head for.
Unlike in Parts 3 and 4, the state tuple has more than one entry, and the entries differ in size: 3 values for the room, 4 for the person, 2 for each room’s dirt. This is exactly why the states form a tuple rather than a single vector.
The agent works with the observations only, and with the actions it chose itself. The true states are kept for evaluation: they let us check how well the agent’s beliefs match the real situation, how often the Roomba ended up in the same room as the person, and, new in this part, how clean the apartment was kept.
3 DATA PREPARATION
As in the previous parts, we use the data from the simulation directly; no cleaning or transformation is needed. As in Part 4, the data becomes available tick by tick: at each tick, the agent receives only the current observation tuple \(\hat{\mathbf{o}}^{(t)}\), updates its beliefs, and chooses its next action. There is therefore no separate preparation step before inference; the agent processes the data as it arrives.
Throughout this notebook, trajectories are indexed from 0, so that index \(t\) always means time \(t\). During a run, we record four trajectories, each re-indexed from 0 with offset0:
Lôː: the received observation tuples \(\hat{\mathbf{o}}^{(0:T)}\),
Lûː: the chosen action tuples \(\hat{\mathbf{u}}^{(0:T-1)}\),
Lsː: the agent’s belief tuples \(\mathbf{s}^{(0:T)}\), one belief per state factor at each tick, and
Lŝˣː: the true state tuples \(\hat{\mathbf{s}}^{\ast(0:T)}\), for evaluation only.
With several state factors, the agent’s belief needs a word of explanation. Internally, the agent works with the 96 joint states, i.e. all combinations of the five factors, so that its inference is exact. For recording and evaluation, this joint belief is summarized as a tuple of five marginal beliefs, one per factor: the belief \(\boldsymbol{s}^{[t,0]}\) about the Roomba’s room, \(\boldsymbol{s}^{[t,1]}\) about the person’s location, and \(\boldsymbol{s}^{[t,2]}\), \(\boldsymbol{s}^{[t,3]}\) and \(\boldsymbol{s}^{[t,4]}\) about the dirt in each room. Each marginal follows from the joint belief by adding up the entries of all joint states with the same value for that factor. The marginals have the same structure as the true state tuple, so they can be compared factor by factor.
As in Part 4, this part does not use RxInfer: the agent’s computations are written directly in Julia, so no conversion at an RxInfer boundary is needed.
4 MODELING
4.1 Narrative
In general, we will have the following symbol conventions. They are shown for states \(s\); observations \(o\) and actions \(u\) follow the same pattern.
Indices. Time-only indices are written in \((\,)\), compound or non-time indices in \([\,]\). The index therefore tells you what kind of object you are looking at: a superscript \((t)\) collects everything at time \(t\), while \([t,n]\) picks out a single state factor \(n\) at time \(t\).
\(t\): time step, \(t = 0, \ldots, T\)
\(s^{[t,n]}\): categorical random variable for state factor \(n\) at time \(t\). Its values are labels. In this part there are five state factors:
factor 0 (Room): \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\), where the Roomba is
factor 1 (Person): \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}, \text{away}\}\), where the person is
factors 2, 3 and 4 (Dirt in the bedroom, living room and bathroom): each \(\{\text{clean}, \text{dirty}\}\)
A random variable is not bold, because a single categorical random variable has no structure.
\(\mathbf{s}^{(t)} = \big(s^{[t,0]}, \ldots, s^{[t,4]}\big)\): tuple of the random variables of all state factors at time \(t\)
\(u^{[t,0]}\): categorical random variable for the action at time \(t\), i.e. the room the Roomba heads for: \(\{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\). Heading for the room it is already in means staying put. The action at time \(t\) takes effect at time \(t+1\). Only factor 0 (Room) is controlled; the other factors change on their own.
Fonts. Bold marks anything with structure, and the style of bold tells you which kind:
upright bold (\(\mathbf{s}\), \(\mathbf{A}\)): a tuple, an ordered collection whose entries may differ in size, such as one entry per state factor or observation modality
italic bold (\(\boldsymbol{s}\), \(\boldsymbol{A}\)): a single vector, matrix or tensor, such as a one-hot vector, a probability vector, a belief, or a likelihood tensor
For example, \(\mathbf{A} = \big(\boldsymbol{A}^{[0]}, \ldots, \boldsymbol{A}^{[M]}\big)\) is the tuple of likelihood tensors, and \(\boldsymbol{A}^{[m]}\) is the tensor for observation modality \(m\). Likewise, \(\mathbf{B} = \big(\boldsymbol{B}^{[0]}, \ldots, \boldsymbol{B}^{[N]}\big)\) and \(\mathbf{D} = \big(\boldsymbol{D}^{[0]}, \ldots, \boldsymbol{D}^{[N]}\big)\) collect one array per state factor, and \(\mathbf{C} = \big(\boldsymbol{C}^{[0]}, \ldots, \boldsymbol{C}^{[M]}\big)\) collects one preference vector per observation modality. Upright bold is only used for Latin letters, because Greek letters such as \(\pi\) do not render in upright bold.
Dependencies. With several state factors, each tensor depends on only some of them. We write \(\mathcal{D}_A^{[m]}\) for the set of factors that modality \(m\) depends on, and \(\mathcal{D}_B^{[n]}\) for the set of factors whose previous values the transition of factor \(n\) depends on (always including factor \(n\) itself). Each tensor has one axis for its own variable, followed by one axis per factor it depends on, and, for the controlled factor, one axis for the action:
tensor
depends on
shape
axes
\(\boldsymbol{A}^{[0]}\) (Texture)
\(\mathcal{D}_A^{[0]} = \{0\}\)
\(3 \times 3\)
texture, room
\(\boldsymbol{A}^{[1]}\) (Motion)
\(\mathcal{D}_A^{[1]} = \{0, 1\}\)
\(2 \times 3 \times 4\)
motion, room, person
\(\boldsymbol{A}^{[2]}\) (Dirt)
\(\mathcal{D}_A^{[2]} = \{0, 2, 3, 4\}\)
\(2 \times 3 \times 2 \times 2 \times 2\)
dirt reading, room, dirt in each room
\(\boldsymbol{B}^{[0]}\) (Room)
\(\mathcal{D}_B^{[0]} = \{0\}\), and the action
\(3 \times 3 \times 3\)
next room, previous room, action
\(\boldsymbol{B}^{[1]}\) (Person)
\(\mathcal{D}_B^{[1]} = \{1\}\)
\(4 \times 4\)
next location, previous location
\(\boldsymbol{B}^{[n]}\) (Dirt, \(n = 2, 3, 4\))
\(\mathcal{D}_B^{[n]} = \{n, 0\}\)
\(2 \times 2 \times 3\)
next dirt, previous dirt, previous room
For example, \(\boldsymbol{A}^{[1]}[\bullet, r, p]\) is the distribution over the motion reading when the Roomba is in room \(r\) and the person in location \(p\). The dirt sensor depends on all three dirt factors because which room’s dirt it senses depends on where the Roomba is.
Marks. A hat marks a realized value: something that has actually happened. A bold symbol without a hat is a distribution: a probability vector over possible values.
\(\hat{\boldsymbol{s}}^{[t,n]}\): realized value of \(s^{[t,n]}\), one-hot encoded as a vector, e.g. \([0,1,0]^{\top}\) for the Roomba being in the living room
\(\hat{\mathbf{s}}^{(t)} = \big(\hat{\boldsymbol{s}}^{[t,0]}, \ldots, \hat{\boldsymbol{s}}^{[t,4]}\big)\): tuple of the realized values of all state factors at time \(t\)
\(\hat{\mathbf{s}}^{(0:T)} = \big(\hat{\mathbf{s}}^{(0)}, \ldots, \hat{\mathbf{s}}^{(T)}\big)\): trajectory (sequence over time) of realized states
\(\hat{\boldsymbol{u}}^{[t,0]}\): the action the agent actually chose at time \(t\), one-hot encoded, e.g. \([0,0,1]^{\top}\) for heading to the bathroom; \(\hat{\mathbf{u}}^{(t)} = \big(\hat{\boldsymbol{u}}^{[t,0]}\big)\) is the tuple of all actions at time \(t\)
\(\boldsymbol{s}^{[t,n]}\), without a hat: the agent’s belief about state factor \(n\) at time \(t\), a probability vector with one entry per value of \(s^{[t,n]}\)
\(\boldsymbol{B}^{\ast[0]}[\bullet, \bullet, u]\,\hat{\boldsymbol{s}}^{\ast[t-1,0]}\): a probability vector (here, over the next room, given action \(u\)), not a realized value
\(\bar{\boldsymbol{A}}^{[m]}\), with a bar: a posterior mean, the agent’s point estimate of a learned matrix after inference (used in Parts 2 and 3; in this part, the agent’s tensors are given)
Joint and marginal beliefs. Internally, the agent works with the joint states: all combinations of the five factors, \(3 \cdot 4 \cdot 2 \cdot 2 \cdot 2 = 96\) in total. Its joint belief \(\boldsymbol{s}^{(t)}\) is a single probability vector over these 96 combinations, italic bold and with a time-only index, since it describes everything at time \(t\) at once. From it, the belief about a single factor, its marginal\(\boldsymbol{s}^{[t,n]}\), follows by adding up the entries of all combinations with the same value for that factor. The tuple of marginals, \(\big(\boldsymbol{s}^{[t,0]}, \ldots, \boldsymbol{s}^{[t,4]}\big)\), is what we record and evaluate. The marginals lose one thing: how the factors are related. For example, the joint belief can express “the Roomba and the person are probably in the same room”, while the two marginals on their own cannot.
The agent never observes the states, so its state variables never carry a hat: it has random variables \(s^{[t,n]}\) and beliefs \(\boldsymbol{s}^{[t,n]}\) and \(\boldsymbol{s}^{(t)}\). The observations are different, because the agent does receive them. Each observation therefore appears in three forms, which share the same index but are different objects:
\(o^{[t,0]}\): the categorical random variable for the Texture observation at time \(t\), with values \(\{\text{hardwood}, \text{carpet}, \text{tiles}\}\)
\(\hat{\boldsymbol{o}}^{[t,0]}\): the observation the agent actually received, one-hot encoded. If the camera reads hardwood at time \(t\), then \(\hat{\boldsymbol{o}}^{[t,0]} = [1,0,0]^{\top}\). If at the same time the motion sensor detects motion and the dirt sensor picks up nothing, then \(\hat{\boldsymbol{o}}^{[t,1]} = [0,1]^{\top}\) and \(\hat{\boldsymbol{o}}^{[t,2]} = [1,0]^{\top}\), and the tuple of all modalities is \(\hat{\mathbf{o}}^{(t)} = \big([1,0,0]^{\top}, [0,1]^{\top}, [1,0]^{\top}\big)\).
\(\boldsymbol{o}^{[t,0]}\), without a hat: an expected observation, i.e. a probability vector the agent computes from its beliefs
For a modality that depends on a single factor, the expected observation is a matrix-vector product, as before. Suppose the agent believes the Roomba is most likely in the bedroom, \(\boldsymbol{s}^{[t,0]} = [0.7, 0.2, 0.1]^{\top}\), and its texture matrix has the same values as \(\boldsymbol{A}^{\ast[0]}\). Then it expects
that is, hardwood with probability 0.645, carpet 0.22, and tiles 0.135. For a modality that depends on several factors, the tensor is contracted with the beliefs about all of them. The motion sensor detects motion with probability 0.8 if the person is in the Roomba’s room, and 0.05 otherwise. Suppose the agent believes the person is in the living room with probability 0.5, away with 0.3, and in the bedroom and bathroom with 0.1 each. Treating the two factors as independent, the probability that the person is in the same room as the Roomba is \(0.7 \cdot 0.1 + 0.2 \cdot 0.5 + 0.1 \cdot 0.1 = 0.18\), so the agent expects motion with probability \(0.8 \cdot 0.18 + 0.05 \cdot 0.82 \approx 0.185\). The agent itself uses its joint belief for this, which also accounts for any relation between the two factors.
So, for observations the agent has both hatted and unhatted forms: what it expects (\(\boldsymbol{o}\)) and what it gets (\(\hat{\boldsymbol{o}}\)). For states, it only ever has the unhatted forms. Only the environment has \(\hat{\boldsymbol{s}}^{\ast[t,n]}\): where the Roomba and the person really are, and which rooms are really dirty.
Actions, preferences and planning. With control, the agent needs three more kinds of objects:
\(\boldsymbol{B}^{[0]}[\bullet, \bullet, u]\): the transition matrix of the controlled factor under action\(u\). The action also matters for the dirt factors, but only indirectly: it determines where the Roomba goes, and that determines which room gets cleaned.
\(\boldsymbol{C}^{[m]}\): the agent’s preference over the observations of modality \(m\), a probability vector expressing which observations it prefers to receive. In this part, two modalities have a real preference: Motion, for no motion, e.g. \(\boldsymbol{C}^{[1]} = [0.9, 0.1]^{\top}\), and Dirt, for dirt picked up, e.g. \(\boldsymbol{C}^{[2]} = [0.2, 0.8]^{\top}\). Texture has a flat preference.
planning symbols, following Algorithm 18 in Chapter 9:
\(H\): the planning horizon, i.e. how many steps the agent looks ahead; \(\tau = 0, \ldots, H-1\) counts the imagined steps, where step \(\tau\) predicts time \(t + \tau + 1\)
\(\boldsymbol{V}\): the matrix of candidate policies, i.e. action sequences. Row \(p\), \(\pi^{[p]} = \boldsymbol{V}[p, \bullet]\), is policy \(p\), and \(\boldsymbol{V}[p, \tau]\) is the action it prescribes at step \(\tau\).
\(\boldsymbol{s}_{\pi^{[p]}}^{(\tau)}\) and \(\boldsymbol{o}_{\pi^{[p]}}^{[\tau,m]}\): the predicted joint belief and predicted observation at step \(\tau\) under policy \(p\), unhatted because they are predictions
\(G^{[p,\tau]}\): the expected free energy of policy \(p\) at step \(\tau\), and \(G^{[p]} = \sum_{\tau} G^{[p,\tau]}\) its total
\(\pi^{(t)}\): the random variable for the policy chosen at time \(t\), with posterior \(Q\big(\pi^{(t)} = \pi^{[p]}\big) = \sigma\big(-\boldsymbol{G}\big)_p\), where \(\boldsymbol{G}\) collects the \(G^{[p]}\) and \(\sigma\) is the softmax
\(Q\big(u^{[t,0]}\big)\): the resulting distribution over the next action, obtained by adding up \(Q\big(\pi^{(t)} = \pi^{[p]}\big)\) over all policies whose first action is \(u\)
The asterisk\({}^{\ast}\) marks what belongs to the environment (the generative process), e.g. \(\hat{\mathbf{s}}^{\ast(t)}\) and \(\boldsymbol{B}^{\ast[0]}\). Symbols without it belong to the agent (the generative model). In this part, the agent is given its tensors: \(\mathbf{A} = \mathbf{A}^{\ast}\) and \(\mathbf{B} = \mathbf{B}^{\ast}\). The preferences \(\mathbf{C}\) belong to the agent alone; the environment has none. The observation \(\hat{\mathbf{o}}^{(t)}\) has no asterisk because it is shared: the environment generates it and the agent receives it. The same holds for the action \(\hat{\mathbf{u}}^{(t)}\): the agent chooses it and the environment carries it out.
Code symbols. The code mirrors the math as closely as identifiers allow:
L prefix: upright bold (\(\mathbf{s}\), \(\mathbf{A}\)), a tuple. The L suggests verticality (upright).
_ prefix: italic bold (\(\boldsymbol{s}\), \(\boldsymbol{A}\)), a single vector, matrix or tensor that is not part of a tuple. The _ suggests the underline of bold.
hat: a Unicode circumflex on the letter, e.g. ŝ for \(\hat{s}\), ô for \(\hat{o}\), û for \(\hat{u}\)
bar: a combining macron on the letter (typed \bar then Tab), e.g. _Ā⁰ for \(\bar{\boldsymbol{A}}^{[0]}\)
asterisk \({}^{\ast}\): a superscript ˣ (typed \^x then Tab)
time index \((t)\): a subscript in the name: ₜ for \(t\), ₜ₋₁ for \(t-1\), ₀ for \(t=0\)
trajectory \((0{:}T)\): a ː suffix (triangular colon), e.g. Lôː for \(\hat{\mathbf{o}}^{(0:T)}\)
tuple index \([n]\) or \([m]\): real 0-based indexing, e.g. LAˣ[0]. In Julia, tuples over factors or modalities are built with tuple0(a, b, ...), which wraps each argument as one entry. In Python, ordinary tuples and lists are already 0-based.
trajectory index \(t\): real 0-based indexing too, e.g. Lŝˣː[t]. In Julia, a trajectory is an existing vector re-indexed from 0 with offset0(v), which shifts the vector’s own indices instead of wrapping it.
Indexing and the tear-off mnemonic. Only tuples have names; their entries are written by indexing. Read a factor or modality index ([n] or [m]) as tearing the vertical line off the L, turning it into _: LAˣ[0] reads as _Aˣ[0], i.e. italic bold \(\boldsymbol{A}^{\ast[0]}\), just as indexing into a tuple turns the upright \(\mathbf{A}^{\ast}\) into the italic \(\boldsymbol{A}^{\ast[0]}\) in the math. Three details:
A time index into a trajectory does not tear off the line: Lŝˣː[t] is still a tuple, \(\hat{\mathbf{s}}^{\ast(t)}\). Only the factor index that follows tears it off: Lŝˣː[t][1] is \(\hat{\boldsymbol{s}}^{\ast[t,1]}\), the person’s location.
The mnemonic applies wherever the brackets appear, on either side of an assignment: LAˣ[0] = … assigns to \(\boldsymbol{A}^{\ast[0]}\).
Each further index changes the font of the result as it would in the math: LAˣ[1][2, 1, 1] is a single number, an entry of the tensor \(\boldsymbol{A}^{\ast[1]}\), so it is no longer bold at all. Likewise, LBˣ[0][:, :, u] is the transition matrix of the Room factor for action u, still italic bold.
Since each tuple has exactly one name, there is nothing to keep in sync: when a tuple gets a new value, a single assignment updates it, e.g. Lŝˣₜ = tuple0(...).
As in Part 4, this part does not use RxInfer: the agent’s computations are written directly in Julia.
chosen action tuple at \(t\), and its entry (one-hot over target rooms)
\(\hat{\mathbf{u}}^{(0:T-1)}\)
Lûː
trajectory of chosen actions
\(\boldsymbol{s}^{(t)}\)
_sₜ
agent’s joint belief over the 96 combinations of all factors
\(\boldsymbol{s}^{[t,n]}\)
Lsₜ[n]
agent’s marginal belief about factor \(n\) at \(t\)
\(\boldsymbol{s}^{[0:T,\bullet]}\)
Lsː
trajectory of the agent’s marginal beliefs
\(H\), \(\boldsymbol{V}\)
H, _V
planning horizon, matrix of candidate policies
\(\boldsymbol{G}\), \(Q\big(\pi^{(t)}\big)\)
_G, _qπ
expected free energies of the policies, and the posterior over policies
\(t\), \(T\)
t, T
time step, last time step
The code stores the agent’s beliefs, not its random variables. So the tuple Lsₜ holds the marginal beliefs \(\big(\boldsymbol{s}^{[t,0]}, \ldots, \boldsymbol{s}^{[t,4]}\big)\), while in the math, \(\mathbf{s}^{(t)}\) denotes the tuple of random variables. The joint belief has its own name, _sₜ: it is a single vector, not a tuple, which the font, and the prefix, make visible.
In Python, ˣ is normalized to a plain x, so LAˣ and LAx are the same name; never use LAx for anything else.
Changes in the symbol conventions compared to Part 4
Five state factors instead of one: Room, Person, and Dirt in each of the three rooms. Only Room is controlled by the agent’s actions; the other factors change on their own.
Dependencies. The sets \(\mathcal{D}_A^{[m]}\) and \(\mathcal{D}_B^{[n]}\) list which factors each likelihood tensor and each transition tensor depends on, with a table of every tensor’s dependencies, shape and axes. This is the core of multi-factor modeling, and corresponds to A_dependencies and B_dependencies in pymdp. The calligraphic \(\mathcal{D}\) for dependency sets is distinct from the upright \(\mathbf{D}\) for the initial-state tuple.
Tensors instead of matrices, wherever an array can have more than two axes.
Joint and marginal beliefs. The agent’s joint belief over all 96 combinations of the five factors is \(\boldsymbol{s}^{(t)}\): italic bold, because it is a single vector, and with a time-only index, because it describes everything at time \(t\). This follows directly from the existing index and font rules, so no new symbol is needed. The per-factor beliefs \(\boldsymbol{s}^{[t,n]}\) are its marginals. In code, the joint belief is _sₜ and the tuple of marginals is Lsₜ.
Expected observations with several factors. For a modality that depends on one factor, the expected observation is still a matrix-vector product. For a modality that depends on several factors, such as Motion, the tensor is contracted with the beliefs about all of them.
New modality numbering: Texture (0), Motion (1) and Dirt (2). The light sensor of the previous parts is dropped.
Two preferences:\(\boldsymbol{C}^{[1]}\) for no motion and \(\boldsymbol{C}^{[2]}\) for dirt picked up. Their relative strength sets the agent’s balance between avoiding people and cleaning.
Predicted beliefs in planning are joint beliefs,\(\boldsymbol{s}_{\pi^{[p]}}^{(\tau)}\), over all factors at once.
The RxInfer boundary is reduced to a single remark, since this part, like Part 4, does not use RxInfer.
The code table covers the per-factor tensors, the dependency tuples LdepsA and LdepsB, both preferences, and the joint and marginal beliefs. The rows that were specific to the light sensor and to the combined RoomOccupancy factor are removed.
4.2 Core Elements
This section attempts to answer three important questions:
What metrics are we going to track?
What decisions do we intend to make?
What are the sources of uncertainty?
We track three metrics:
disturbance: the fraction of time steps at which the Roomba is in the same room as the person. Lower is better.
cleanliness: the fraction of time steps at which each room is dirty, and its average over the three rooms. Lower is better. This replaces the coverage of Part 4: what matters is not how evenly the Roomba spreads its time, but whether the rooms are kept clean.
tracking: the accuracy of the agent’s beliefs about each factor, i.e. the fraction of time steps at which its most probable value equals the true one, for the Roomba’s room, the person’s location, and the dirt in each room. In particular, we check how well the agent knows whether the person is in its room, and whether the room it is in is dirty.
Since the agent now has two goals that can conflict, we compare three Roombas in the same environment: one that moves at random, one with only the Part 4 preference for no motion, and one with both preferences. This shows what the cleaning goal adds, and what it costs in disturbance.
Decisions. At each time step, the agent chooses where the Roomba goes next: it stays in its current room, or heads for one of the other rooms. It chooses by looking three steps ahead and scoring each candidate sequence of moves by its expected free energy, which now favors moves that are expected to lead to no motion and to dirt being picked up, as well as moves that reduce the agent’s uncertainty about the state. The two preferences can pull in opposite directions: the living room gets dirty fastest, but it is also the room the person uses most.
There are two sources of uncertainty in the environment itself. The first has to do with the fact that the state transitions are not deterministic but rather stochastic, and each factor has its own dynamics:
Room: a move does not always succeed, and since there is no door between the bedroom and the living room, the Roomba has to pass through the bathroom to get from one to the other.
Person: the person moves between the rooms and sometimes leaves the apartment, independently of the Roomba, spending most time in the living room.
Dirt: each clean room gets dirty with a probability per tick that differs per room, the living room fastest; a dirty room only gets clean when the Roomba is in it, and even then cleaning may take more than one tick.
This is captured in the tuple of state transition tensors \(\mathbf{B}^{\ast} = \big(\boldsymbol{B}^{\ast[0]}, \ldots, \boldsymbol{B}^{\ast[4]}\big)\), with one tensor per state factor, each depending only on the factors listed in 4.1.
The second relates to observations, and all three sensors contribute to it. The Roomba’s camera sometimes misidentifies the floor surface in a room. The motion sensor sometimes misses the person when they are sitting still, and occasionally detects motion when they are not in the room. The dirt sensor sometimes misses dirt, and occasionally reports dirt in a clean room. This uncertainty is captured in the tuple of observation likelihood tensors \(\mathbf{A}^{\ast} = \big(\boldsymbol{A}^{\ast[0]}, \boldsymbol{A}^{\ast[1]}, \boldsymbol{A}^{\ast[2]}\big)\), for Texture, Motion and Dirt.
The initial state is not a source of uncertainty in the environment, because the run always starts from the same state: the Roomba in the bedroom, the person in the living room, and all rooms clean.
Finally, there is uncertainty on the agent’s side. The agent is again given its model of the world, \(\mathbf{A} = \mathbf{A}^{\ast}\) and \(\mathbf{B} = \mathbf{B}^{\ast}\), so it is not uncertain about how the environment works. It is uncertain about the state: where the Roomba is, where the person is, and which rooms are dirty, both now and in the future. It does not know the starting state either, so it begins with a uniform prior over all combinations. The dirt in rooms the Roomba is not in is never observed: the agent can only predict it, and its uncertainty grows the longer it stays away. Reducing this uncertainty is part of what the agent’s choices aim at: going back to check a room has value in itself, besides the chance of finding dirt to clean.
Changes in the core elements compared to Part 4
Cleanliness replaces coverage as a metric: what matters is not how evenly the Roomba spreads its time, but whether the rooms are kept clean. Tracking now covers all five state factors.
Three Roombas are compared in the same environment: one that moves at random, one with only the Part 4 preference for no motion, and one with both preferences. This shows what the cleaning goal adds, and what it costs in disturbance.
The decisions are guided by two preferences, no motion and dirt picked up, with a planning horizon of three steps. The preferences can conflict, most of all in the living room, which gets dirty fastest but is also the room the person uses most.
The transition uncertainty is described per state factor: failed moves for the Room, independent movements for the Person, and growing dirt that is only removed by cleaning for each room. \(\mathbf{B}^{\ast}\) is now a tuple of five tensors.
The observation uncertainty includes the new dirt sensor, which can miss dirt or report dirt in a clean room. The light sensor is dropped, and the motion sensor now detects the person in the Roomba’s room.
The initial state is fixed: the Roomba in the bedroom, the person in the living room, and all rooms clean.
The agent’s uncertainty covers all five factors. Dirt in rooms the Roomba is not in is never observed directly, so the agent’s uncertainty about it grows over time, which gives going back to check a room value in itself.
4.3 System-Under-Steer / Environment / Generative Process
First, before letting the agent take control of the Roomba, we describe the environment it acts in, i.e. the generative process that produces the data. In this part, the environment’s state has five aspects, each described by its own state factor:
factor 0 (Room): \(s^{\ast[t,0]} \in \{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\), where the Roomba is
factor 1 (Person): \(s^{\ast[t,1]} \in \{\text{bedroom}, \text{livingroom}, \text{bathroom}, \text{away}\}\), where the person is
factors 2, 3 and 4 (Dirt in the bedroom, living room and bathroom): \(s^{\ast[t,n]} \in \{\text{clean}, \text{dirty}\}\), \(n = 2, 3, 4\)
At time \(t\), the true state is the realized state tuple \[
\hat{\mathbf{s}}^{\ast(t)} = \big(\hat{\boldsymbol{s}}^{\ast[t,0]}, \hat{\boldsymbol{s}}^{\ast[t,1]}, \hat{\boldsymbol{s}}^{\ast[t,2]}, \hat{\boldsymbol{s}}^{\ast[t,3]}, \hat{\boldsymbol{s}}^{\ast[t,4]}\big)
\] whose entries are the one-hot encoded values of the five factors. The entries differ in size, 3, 4 and 2 values, which is why the state is a tuple rather than a single vector.
In Parts 3 and 4, occupancy was a property of the Roomba’s room, combined with the room into a single factor. With several factors, it is more natural to describe where the person is. The Roomba’s room is then occupied exactly when the Roomba and the person are in the same room.
The initial state is given by the tuple of initial state distributions \(\mathbf{D}^{\ast} = \big(\boldsymbol{D}^{\ast[0]}, \ldots, \boldsymbol{D}^{\ast[4]}\big)\). In our simulation, the run always starts from the same state: the Roomba in the bedroom, the person in the living room, and all rooms clean, so each \(\boldsymbol{D}^{\ast[n]}\) is a one-hot vector.
Actions. As in Part 4, at each time step the agent chooses an action, the room the Roomba heads for,
and the environment carries it out. Heading for the room it is already in means staying put. The action chosen at time \(t-1\) determines how the Roomba moves from \(t-1\) to \(t\).
Transitions, one per factor. The true transition probabilities are captured by the tuple \(\mathbf{B}^{\ast} = \big(\boldsymbol{B}^{\ast[0]}, \ldots, \boldsymbol{B}^{\ast[4]}\big)\), with one tensor per factor. Given the previous state and action, each factor changes independently of the others, and each depends only on a few factors:
Room, \(\boldsymbol{B}^{\ast[0]} \in [0,1]^{3\times 3\times 3}\) (next room, previous room, action), as in Part 4: if the Roomba heads for the room it is in, it stays; if it heads for an adjacent room, it gets there with probability \(p_{\text{move}}\), and otherwise stays where it is. There is no door between the bedroom and the living room, so when it heads from one to the other, it first moves to the bathroom.
Person, \(\boldsymbol{B}^{\ast[1]} \in [0,1]^{4\times 4}\) (next location, previous location): the person moves between the rooms and sometimes leaves the apartment, independently of the Roomba. The living room, with its abundance of natural light, is where the person spends most time; bathroom visits are short.
Dirt in room \(n\), \(\boldsymbol{B}^{\ast[n]} \in [0,1]^{2\times 2\times 3}\) (next dirt, previous dirt, previous room of the Roomba), for \(n = 2, 3, 4\): if the Roomba was in that room, a dirty room becomes clean with probability \(p_{\text{clean}}\); otherwise, a clean room becomes dirty with probability \(p_{\text{dirty}}(n)\), which is highest for the living room, and a dirty room stays dirty.
The dirt factors are where the dependencies between factors appear in the transitions: a room’s dirt depends on its own previous value and on where the Roomba was. The action influences the dirt only indirectly, through where it takes the Roomba.
Our Roomba is equipped with three sensors. The first is a small camera that tracks the surface it is moving over, since there are
hardwood floors in the bedroom
carpet in the living room, and
tiles in the bathroom.
However, this method is not foolproof, and sometimes the Roomba will mistake one floor type for another. Don’t be too hard on the little guy, it’s just a Roomba after all.
The second is a motion sensor that reports whether it detects movement. It is the most direct evidence that the person is in the Roomba’s room: it usually detects motion when they are, but sometimes misses them when they are sitting still, and occasionally detects motion when they are not.
The third, new sensor is a dirt sensor that reports whether the Roomba’s brushes are picking up dirt. It usually does so when the room it is in is dirty, but sometimes misses dirt, and occasionally reports dirt in a clean room. The light sensor of the previous parts is dropped.
At time \(t\), the received observation tuple is \(\hat{\mathbf{o}}^{(t)} = \big(\hat{\boldsymbol{o}}^{[t,0]}, \hat{\boldsymbol{o}}^{[t,1]}, \hat{\boldsymbol{o}}^{[t,2]}\big)\), whose entries are the one-hot encoded values of three observation modalities:
Observations, each depending on a few factors. The true mapping from the state to the observations is \(\mathbf{A}^{\ast} = \big(\boldsymbol{A}^{\ast[0]}, \boldsymbol{A}^{\ast[1]}, \boldsymbol{A}^{\ast[2]}\big)\), with one likelihood tensor per modality:
Texture, \(\boldsymbol{A}^{\ast[0]} \in [0,1]^{3\times 3}\), depends only on the Roomba’s room.
Motion, \(\boldsymbol{A}^{\ast[1]} \in [0,1]^{2\times 3\times 4}\), depends on the Roomba’s room and the person’s location: what matters is whether the two are the same.
Dirt, \(\boldsymbol{A}^{\ast[2]} \in [0,1]^{2\times 3\times 2\times 2\times 2}\), depends on the Roomba’s room and the dirt in all three rooms: the room tells which room’s dirt the sensor senses.
The tensors also encode the probability that a sensor gives a misleading reading. Observations do not depend on the action. This leaves us with the following generative process specification, where a realized value without bold, such as \(\hat{s}^{\ast[t-1,0]}\), is used as an index into a tensor:
Each line picks out, from a tensor, the probability vector for the current values of exactly the factors it depends on, as listed in 4.1. The specification makes the structure of the environment visible at a glance: each factor and each sensor looks at only a few others.
With actions and several state factors, this process has the structure of a factorized partially observable Markov decision process: five hidden Markov chains, one per factor, coupled through the dependencies in their transitions, with the Room chain driven by the agent’s actions, and three sensors emitting noisy observations of them. The agent never sees the true state directly. In this part, it is given its model of the environment, \(\mathbf{A} = \mathbf{A}^{\ast}\) and \(\mathbf{B} = \mathbf{B}^{\ast}\), and its goal is to choose actions that keep the rooms clean while staying out of the person’s way.
Changes in the environment compared to Part 4
Five state factors replace the combined RoomOccupancy factor: Room, Person, and Dirt in each of the three rooms. Occupancy is no longer a property of the Roomba’s room, but follows from where the person is: the Roomba’s room is occupied exactly when the Roomba and the person are in the same room.
The initial state distributions form a tuple \(\mathbf{D}^{\ast}\) of one-hot vectors: the run always starts with the Roomba in the bedroom, the person in the living room, and all rooms clean.
The transitions are given per factor. Room moves as in Part 4, driven by the agent’s actions. Person moves independently of the Roomba. The dirt in each room depends on its own previous value and on where the Roomba was, which is where dependencies between factors appear in the transitions.
The sensors: the light sensor is dropped, the motion sensor now detects the person in the Roomba’s room, and a new dirt sensor reports whether the Roomba’s brushes are picking up dirt.
The observation tensors are listed with their shapes and dependencies: Texture depends on the room only, Motion on the room and the person, and the dirt sensor on the room and the dirt in all three rooms.
The process specification has one line per factor and per sensor, each picking out, from its tensor, the probability vector for the current values of exactly the factors it depends on.
The process is now a factorized POMDP: five hidden Markov chains, coupled through their dependencies. The agent’s goal is to keep the rooms clean while staying out of the person’s way.
## This is the first time the tuple has entries of different sizes: 3, 4 and 2 values. ## That's exactly what tuple0 is for, and why the state isn't a single vector.## factor 0 (Room): 1 bedroom 2 livingroom 3 bathroom## factor 1 (Person): 1 bedroom 2 livingroom 3 bathroom 4 away## factors 2–4 (Dirt in bedroom, livingroom, bathroom): 1 clean 2 dirtyLŝˣ₀ =tuple0( [1.0, 0.0, 0.0], ## ŝ^{*[0,0]}: Roomba in the bedroom [0.0, 1.0, 0.0, 0.0], ## ŝ^{*[0,1]}: person in the living room [1.0, 0.0], ## ŝ^{*[0,2]}: bedroom clean [1.0, 0.0], ## ŝ^{*[0,3]}: living room clean [1.0, 0.0], ## ŝ^{*[0,4]}: bathroom clean) ## ŝ^{*(0)}: the initial state tuple
5-element OffsetArray(::Vector{Vector{Float64}}, 0:4) with eltype Vector{Float64} with indices 0:4:
[1.0, 0.0, 0.0]
[0.0, 1.0, 0.0, 0.0]
[1.0, 0.0]
[1.0, 0.0]
[1.0, 0.0]
The state variables represent what we need to know. The true state at time \(t\) is given by a (compound) tuple \(\hat{\mathbf{s}}^{\ast(t)}\) whose components are the one-hot encoded state factors (here five factors):
one-hot encoded as \(\hat{\boldsymbol{s}}^{\ast[t,n]} \in \big\{[1,0]^{\top}, [0,1]^{\top}\big\}\)
The factors have different numbers of values, 3, 4 and 2, so their one-hot vectors differ in length. This is why the state is a tuple rather than a single vector. All combinations together give \(3 \cdot 4 \cdot 2 \cdot 2 \cdot 2 = 96\) possible states.
The observation variables represent what we can sense. Similarly, the (compound) observation tuple received by the Roomba at time \(t\) is given by
one-hot encoded as \(\hat{\boldsymbol{o}}^{[t,2]} \in \big\{[1,0]^{\top}, [0,1]^{\top}\big\}\)
Changes in the state and observation variables compared to Part 4
Five state factors replace the combined RoomOccupancy factor: Room (3 values), Person (4 values), and Dirt in each of the three rooms (2 values each). The three dirt factors are listed together, since they have the same values.
The state is now a tuple whose entries differ in length: the factors have different numbers of values, so their one-hot vectors have different lengths. All combinations together give \(3 \cdot 4 \cdot 2 \cdot 2 \cdot 2 = 96\) possible states.
The modalities are renumbered: the light sensor is dropped, Motion is now modality 1, and the new Dirt sensor is modality 2, with values nothing picked up and dirt picked up.
4.3.2 Decision variables
The decision variables represent what we can control. As in Part 4, the agent chooses an action at each time step. The action at time \(t\) is given by a (compound) tuple \(\hat{\mathbf{u}}^{(t)}\) whose components are the one-hot encoded control factors (here a single factor):
control factor 0 (Destination), acting on state factor 0 (Room):
\(u^{[t,0]} \in \{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\), the room the Roomba heads for,
one-hot encoded as \(\hat{\boldsymbol{u}}^{[t,0]} \in \big\{[1,0,0]^{\top}, [0,1,0]^{\top}, [0,0,1]^{\top}\big\}\)
Heading for the room the Roomba is already in means staying put. The action chosen at time \(t\) takes effect at time \(t+1\): it selects which of the transition matrices \(\boldsymbol{B}^{\ast[0]}[\bullet, \bullet, u]\) the environment uses for the Room factor in the next step.
Of the five state factors, the action controls only one, the Room, but its effects reach further. Through where it takes the Roomba, the action also determines which room gets cleaned: the dirt factors depend on the Roomba’s room, so the action influences them indirectly. The person’s location, on the other hand, is beyond the agent’s control: it can choose its room, but not its company.
The agent’s choice is therefore always a bet on two things at once. It heads for a room it expects to be quiet and in need of cleaning, and its sensors then tell it whether those expectations came true: the motion sensor whether the person is there, the dirt sensor whether there is dirt to pick up.
Changes in the decision variables compared to Part 4
The control factor acts on state factor 0 (Room), instead of the combined RoomOccupancy factor, and selects the transition matrix of the Room factor only.
The action’s reach is described explicitly: it directly controls only the Room factor, but through where it takes the Roomba, it also determines which room gets cleaned, since the dirt factors depend on the Roomba’s room. The person’s location remains beyond the agent’s control. This shows how a single control factor can affect several state factors through their dependencies.
The agent’s choice is a bet on two things at once: heading for a room that is quiet and in need of cleaning, with the motion sensor and the dirt sensor telling it whether each expectation came true.
## actions (target room): 1 bedroom## 2 livingroom## 3 bathroom## heading for the room the Roomba is already in means staying putaction_labels = ["bedroom", "livingroom", "bathroom"]
Exogenous forces represent what changes on its own: the influences on the environment that the agent cannot control, arriving between its decisions. In an active inference model, they appear as the stochastic dynamics of the state factors, in so far as these dynamics do not depend on the agent’s actions.
Forces versus factors. A state factor and an exogenous force are not the same thing. In the sequential-decision framework, the state \(S_t\), the decision \(x_t\) and the exogenous information \(W_{t+1}\) are separate, and the next state follows from all three:
\[
S_{t+1} = S^M\big(S_t, x_t, W_{t+1}\big)
\]
The person’s location, for example, is a state; the person deciding to walk to the bathroom is the exogenous force acting on it. In an active inference model, such a force has no variable of its own: it is contained in the randomness of the transition tensors \(\mathbf{B}^{\ast}\).
Kinds of state factors. Not every factor the agent cannot control is driven by exogenous forces alone. With the factors of this part, four kinds can be distinguished:
kind of factor
example
exogenous?
controlled: the action selects its transition
Room
only the chance that a move fails
indirectly influenced: not controlled, but its transition depends on a controlled factor
Dirt (cleaning depends on where the Roomba is)
partly: dirt accumulates exogenously, but is removed by the agent
fully exogenous: its transition depends on nothing the agent does
Person
yes
static: never changes at all
the layout of the apartment
no force acts on it: it is a hidden fact, not a force
The exogenous forces in this part. There are two:
The person’s movements. The person moves between the rooms and sometimes leaves the apartment, independently of the Roomba. The agent can neither influence nor directly observe this; it only notices the person through the motion sensor, when they are in the same room. These forces are contained in \(\boldsymbol{B}^{\ast[1]}\).
Dirt accumulating. Each clean room gets dirty with some probability per tick, the living room fastest. The agent cannot prevent this; it can only remove dirt by cleaning, i.e. by being in the room. These forces are contained in the growth of dirt in \(\boldsymbol{B}^{\ast[2]}\) to \(\boldsymbol{B}^{\ast[4]}\), while the removal of dirt is the agent’s doing.
In the previous parts, the changes in occupancy were an exogenous force as well, but they were folded into the combined RoomOccupancy factor. With a separate factor for the person, they are now explicit. Further exogenous forces could be added in later parts, for example doors that are closed at random.
Changes in the exogenous forces compared to Part 4
This section has content for the first time. In the previous parts, it stated that there were no exogenous force variables (yet). In this part, two exogenous forces are made explicit: the person’s movements and the accumulation of dirt.
Forces are distinguished from factors: an exogenous force is the influence that changes a factor from outside the agent’s control, contained in the randomness of the transition tensors, not the factor itself.
Four kinds of state factors are distinguished: controlled, indirectly influenced, fully exogenous, and static, depending on how much of their dynamics the agent’s actions can affect.
The occupancy changes of the previous parts are recognized as an exogenous force that was folded into the combined RoomOccupancy factor, and is now explicit through the separate Person factor.
4.3.4 State Transition and Observation Generation functions
Starting from \(\hat{\boldsymbol{s}}^{\ast[0,n]} \sim \mathcal{Cat}\big(\boldsymbol{D}^{\ast[n]}\big)\), \(n = 0, \ldots, 4\), with one-hot vectors for the Roomba in the bedroom, the person in the living room, and all rooms clean, the environment evolves according to
where \(\hat{u}^{[t-1,0]}\) is the action chosen at time \(t-1\). With one tensor per factor, each table below stays small: the 96 joint states never have to be written out.
State Transition model \(\boldsymbol{B}^{\ast[0]}\) (factor 0, Room)
One table per action \(u\). If the Roomba heads for the room it is in, it stays. If it heads for an adjacent room, it gets there with \(p_{\text{move}} = 0.9\), and otherwise stays where it is. There is no door between the bedroom and the living room, so when it heads from one to the other, it first moves to the bathroom. Each column is a probability distribution over the next room, given the previous room, so each column sums to 1.
Action: head for the bedroom
next room previous room
bedroom
livingroom
bathroom
bedroom
1.0
0.0
0.9
livingroom
0.0
0.1
0.0
bathroom
0.0
0.9
0.1
Action: head for the living room
next room previous room
bedroom
livingroom
bathroom
bedroom
0.1
0.0
0.0
livingroom
0.0
1.0
0.9
bathroom
0.9
0.0
0.1
Action: head for the bathroom
next room previous room
bedroom
livingroom
bathroom
bedroom
0.1
0.0
0.0
livingroom
0.0
0.1
0.0
bathroom
0.9
0.9
1.0
The bathroom is the hub of the apartment: from the bedroom, heading for the living room leads to the bathroom first, and vice versa.
State Transition model \(\boldsymbol{B}^{\ast[1]}\) (factor 1, Person)
The person moves independently of the Roomba and of the action. Each column is a probability distribution over the person’s next location, given the previous one, so each column sums to 1.
next previous
bedroom
livingroom
bathroom
away
bedroom
0.80
0.03
0.20
0.00
livingroom
0.13
0.90
0.30
0.05
bathroom
0.05
0.04
0.50
0.00
away
0.02
0.03
0.00
0.95
The living room, with its abundance of natural light, is where the person spends most time (staying with 0.90). Bathroom visits are short (0.50). When the person leaves the apartment, they mostly stay away for a while (0.95), and they come back through the living room.
State Transition model \(\boldsymbol{B}^{\ast[n]}\) (factors 2–4, Dirt)
Each dirt factor depends on its own previous value and on where the Roomba was. If the Roomba was in that room, a dirty room becomes clean with \(p_{\text{clean}} = 0.8\), and a clean room stays clean. If the Roomba was elsewhere, a clean room becomes dirty with \(p_{\text{dirty}}(n)\), and a dirty room stays dirty. Each column is a probability distribution over the next value, given the previous one, so each column sums to 1.
The Roomba was in this room (cleaning)
next previous
clean
dirty
clean
1.0
0.8
dirty
0.0
0.2
The Roomba was elsewhere (dirt accumulating)
next previous
clean
dirty
clean
\(1 - p_{\text{dirty}}(n)\)
0.0
dirty
\(p_{\text{dirty}}(n)\)
1.0
with a rate of dirt accumulation that differs per room:
bedroom
livingroom
bathroom
\(p_{\text{dirty}}\)
0.02
0.05
0.03
The living room, the room in most use, gets dirty fastest: on average after about 20 ticks, against 50 for the bedroom and about 33 for the bathroom. Since the Roomba’s room has three possible values, \(\boldsymbol{B}^{\ast[n]}\) consists of three \(2 \times 2\) matrices: the cleaning table for the value matching room \(n\), and the accumulating table for the other two.
Observation Generation model \(\boldsymbol{A}^{\ast[0]}\) (modality 0, Texture)
The texture depends only on the Roomba’s room. Each column is a probability distribution over the observed texture, so each column sums to 1.
observation \(o^{[t,0]}\) room
bedroom
livingroom
bathroom
hardwood
0.9
0.05
0.05
carpet
0.05
0.9
0.05
tiles
0.05
0.05
0.9
The off-diagonal entries are the camera’s mistakes: in each room, it misreads the floor with probability 0.1, split equally between the two other textures.
Observation Generation model \(\boldsymbol{A}^{\ast[1]}\) (modality 1, Motion)
Motion depends on the Roomba’s room and the person’s location, but only through one question: are they in the same room? The table gives the probability of detecting motion; the probability of no motion is 1 minus each entry.
room person
bedroom
livingroom
bathroom
away
bedroom
0.8
0.05
0.05
0.05
livingroom
0.05
0.8
0.05
0.05
bathroom
0.05
0.05
0.8
0.05
The motion sensor detects the person with probability 0.8 when they are in the Roomba’s room; it misses them when they are sitting still with probability 0.2. Otherwise, it raises a false alarm with probability 0.05.
Observation Generation model \(\boldsymbol{A}^{\ast[2]}\) (modality 2, Dirt)
The dirt reading depends on the Roomba’s room and on the dirt in all three rooms, but again only through one question: is the room the Roomba is in dirty? Each column is a probability distribution over the dirt reading, so each column sums to 1.
observation \(o^{[t,2]}\) dirt in the Roomba’s room
clean
dirty
nothing picked up
0.95
0.1
dirt picked up
0.05
0.9
The dirt sensor picks up dirt in a dirty room with probability 0.9, and misses it with probability 0.1. In a clean room, it falsely reports dirt with probability 0.05. Although \(\boldsymbol{A}^{\ast[2]}\) has \(2 \times 3 \times 2 \times 2 \times 2 = 48\) entries, they are all filled from this small table: for each room, the column is selected by the dirt factor of that room.
Changes in the transition and observation functions compared to Part 4
The equations have one line per state factor and per sensor, each indexing its tensor only by the factors it depends on. The run starts from a fixed state: the Roomba in the bedroom, the person in the living room, and all rooms clean.
Each factor has its own small table instead of the combined \(6 \times 6\) matrix per action. The 96 joint states never have to be written out, which is one of the main practical advantages of splitting the state into factors.
\(\boldsymbol{B}^{\ast[0]}\) keeps the three room tables from Part 4, now for the Room factor only.
New: \(\boldsymbol{B}^{\ast[1]}\) for the Person, a \(4 \times 4\) matrix in which the person spends most time in the living room, makes short bathroom visits, and sometimes leaves the apartment.
New: \(\boldsymbol{B}^{\ast[2]}\) to \(\boldsymbol{B}^{\ast[4]}\) for the Dirt, each built from a cleaning table, used when the Roomba was in that room, with \(p_{\text{clean}} = 0.8\), and an accumulating table, used otherwise, with a rate \(p_{\text{dirty}}\) of 0.02 for the bedroom, 0.05 for the living room and 0.03 for the bathroom.
The occupancy rules (\(p_{\text{stay}}\), \(p_{\text{occ}}\)) and the light sensor are dropped.
\(\boldsymbol{A}^{\ast[1]}\) (Motion) is a \(3 \times 4\) table of detection probabilities by the Roomba’s room and the person’s location: 0.8 if they are in the same room, 0.05 otherwise.
New: \(\boldsymbol{A}^{\ast[2]}\) (Dirt). Although it has 48 entries, all are filled from one small table, selected by the dirt factor of the Roomba’s room: 0.9 to pick up dirt in a dirty room, 0.05 to falsely report dirt in a clean one.
## factor 0 (Room): 1 bedroom 2 livingroom 3 bathroom## factor 1 (Person): 1 bedroom 2 livingroom 3 bathroom 4 away## factors 2–4 (Dirt in bedroom, livingroom, bathroom): 1 clean 2 dirty## actions (target room): 1 bedroom 2 livingroom 3 bathroom## ---------- factor 0: Room (3×3×3: next room, previous room, action) ----------p_move =0.9## probability that a move succeedsnext_room(r, u) = r == u ? r : (r ==3|| u ==3) ? u :3## no door between bedroom 1 and livingroom 2_B_room =zeros(3, 3, 3)for u in1:3, r in1:3 r′ =next_room(r, u)if r′ == r _B_room[r, r, u] =1.0## heading for the current room: stayelse _B_room[r′, r, u] = p_move ## move succeeds _B_room[r, r, u] =1- p_move ## move fails: stayendend## ---------- factor 1: Person (4×4: next location, previous location) ----------_B_person = [0.800.030.200.00; ## to bedroom (columns: from bedroom, livingroom, bathroom, away)0.130.900.300.05; ## to livingroom0.050.040.500.00; ## to bathroom0.020.030.000.95] ## to away## ---------- factors 2–4: Dirt (2×2×3: next dirt, previous dirt, previous room of the Roomba) ----------p_clean =0.8## probability that a dirty room is cleaned, per tickp_dirty = [0.02, 0.05, 0.03] ## rate of dirt accumulation: bedroom, livingroom, bathroomfunctiondirt_tensor(room, p_dirty_room) _B =zeros(2, 2, 3)for r in1:3if r == room ## the Roomba was in this room: cleaning _B[:, :, r] = [1.0 p_clean;0.01- p_clean]else## the Roomba was elsewhere: dirt accumulating _B[:, :, r] = [1- p_dirty_room 0.0; p_dirty_room 1.0]endend _BendLBˣ =tuple0( _B_room, ## B^{*[0]}: Room _B_person, ## B^{*[1]}: Persondirt_tensor(1, p_dirty[1]), ## B^{*[2]}: Dirt in the bedroomdirt_tensor(2, p_dirty[2]), ## B^{*[3]}: Dirt in the living roomdirt_tensor(3, p_dirty[3]), ## B^{*[4]}: Dirt in the bathroom)## D_B^{[n]}: the factors whose previous values each transition depends on (own factor first)LdepsB =tuple0([0], [1], [2, 0], [3, 0], [4, 0])size.(LBˣ) ## (3,3,3), (4,4), (2,2,3), (2,2,3), (2,2,3)
5-element OffsetArray(::Vector{Tuple{Int64, Int64, Vararg{Int64}}}, 0:4) with eltype Tuple{Int64, Int64, Vararg{Int64}} with indices 0:4:
(3, 3, 3)
(4, 4)
(2, 2, 3)
(2, 2, 3)
(2, 2, 3)
Changes in the transition code compared to Part 4
No joint states: the helper idx, the occupancy rules p_occ and p_stay, and the combined \(6 \times 6 \times 3\) array _B_joint are removed. Each factor gets its own tensor, built directly from the tables in 4.3.4.
_B_room is unchanged from Part 4, now as the tensor of the Room factor only.
_B_person, the \(4 \times 4\) transition matrix of the Person factor, is entered directly from its table.
dirt_tensor builds the \(2 \times 2 \times 3\) tensor for one room’s dirt: the cleaning matrix in the slice where the Roomba was in that room, and the accumulating matrix in the other two slices.
LBˣ is a tuple of five tensors, one per state factor, with different shapes: \(3 \times 3 \times 3\), \(4 \times 4\), and three times \(2 \times 2 \times 3\).
LdepsB records the dependencies\(\mathcal{D}_B^{[n]}\) of each transition tensor, using 0-based factor indices. Each list starts with the factor itself, and the order of the list matches the order of the tensor’s axes. The action axis of the Room factor is implicit, since it is the only controlled factor.
## factor 0 (Room): 1 bedroom 2 livingroom 3 bathroom## factor 1 (Person): 1 bedroom 2 livingroom 3 bathroom 4 away## factors 2–4 (Dirt in bedroom, livingroom, bathroom): 1 clean 2 dirty## ---------- modality 0: Texture (3×3: texture, room) ----------_A_texture = [0.90.050.05; ## hardwood (columns: room bedroom, livingroom, bathroom)0.050.90.05; ## carpet0.050.050.9 ] ## tiles## ---------- modality 1: Motion (2×3×4: motion, room, person) ----------p_detect, p_false_motion =0.8, 0.05## motion if the person is in the Roomba's room / otherwise_A_motion =zeros(2, 3, 4)for r in1:3, p in1:4 p_motion = p == r ? p_detect : p_false_motion ## same room? (person 4 = away is never the same room) _A_motion[:, r, p] = [1- p_motion, p_motion] ## (no motion, motion)end## ---------- modality 2: Dirt (2×3×2×2×2: dirt reading, room, dirt in bedroom, livingroom, bathroom) ----------p_pickup, p_false_pickup =0.9, 0.05## dirt picked up in a dirty room / in a clean room_A_dirt =zeros(2, 3, 2, 2, 2)for r in1:3, d₁ in1:2, d₂ in1:2, d₃ in1:2 dirty = (d₁, d₂, d₃)[r] ==2## is the Roomba's room dirty? p_dirt = dirty ? p_pickup : p_false_pickup _A_dirt[:, r, d₁, d₂, d₃] = [1- p_dirt, p_dirt] ## (nothing picked up, dirt picked up)endLAˣ =tuple0( _A_texture, ## A^{*[0]}: Texture _A_motion, ## A^{*[1]}: Motion _A_dirt, ## A^{*[2]}: Dirt)## D_A^{[m]}: the factors each modality depends on, in the order of the tensor's axes after the firstLdepsA =tuple0([0], [0, 1], [0, 2, 3, 4])size.(LAˣ) ## (3,3), (2,3,4), (2,3,2,2,2)
3-element OffsetArray(::Vector{Tuple{Int64, Int64, Vararg{Int64}}}, 0:2) with eltype Tuple{Int64, Int64, Vararg{Int64}} with indices 0:2:
(3, 3)
(2, 3, 4)
(2, 3, 2, 2, 2)
Changes in the observation code compared to Part 4
Tensors instead of repeated columns: each likelihood tensor now has one axis per factor it depends on, so the repeat constructions that spread each room’s column over the six joint states are no longer needed.
The light sensor is dropped, so Motion becomes modality 1, and the new Dirt sensor is modality 2.
_A_texture, the \(3 \times 3\) table of the Texture modality, is entered directly from its table.
_A_motion is a \(2 \times 3 \times 4\) tensor over the Roomba’s room and the person’s location, filled by a loop: motion is detected with probability 0.8 when the two are in the same room, and falsely reported with 0.05 otherwise.
_A_dirt is a \(2 \times 3 \times 2 \times 2 \times 2\) tensor over the Roomba’s room and the dirt in all three rooms, filled by a loop that picks out the dirt factor of the Roomba’s room, (d₁, d₂, d₃)[r], to decide whether dirt is picked up.
LdepsA records the dependencies\(\mathcal{D}_A^{[m]}\) of each likelihood tensor, using 0-based factor indices, in the order of the tensor’s axes after the observation axis.
LAˣ[0] ## the Texture likelihood matrix, A^{*[0]} (3×3: texture, room)
LAˣ[2][2, :, 2, 2, 2] ## A^{*[2]}: probability of picking up dirt, by room, when all rooms are dirty
3-element Vector{Float64}:
0.9
0.9
0.9
LAˣ[2][2, :, 1, 1, 1] ## A^{*[2]}: probability of picking up dirt, by room, when all rooms are clean
3-element Vector{Float64}:
0.05
0.05
0.05
4.3.5 Objective function
The objective function in this section is the designer’s measure of performance: how well the system as a whole does, judged from the outside with knowledge of the true states \(\hat{\mathbf{s}}^{\ast(t)}\). It is not something the environment pursues, nor the quantity the agent minimizes internally; the agent’s own objective, the expected free energy, is described in 4.5.6.
In this part, the agent has two goals, so there are two main measures. The first is the disturbance: the fraction of ticks the Roomba spends in the same room as the person,
where \(\mathbb{1}[\cdot]\) is 1 if the condition holds and 0 otherwise. With separate factors for the Roomba’s room and the person’s location, disturbance is now a condition on two factors: they are in the same room. When the person is away, the condition never holds. Lower is better.
The second is the cleanliness: the fraction of ticks each room \(r\) is dirty, and its average over the three rooms,
where room \(r = 1, 2, 3\) (bedroom, living room, bathroom) corresponds to dirt factor \(n = r + 1\). Lower is better. This replaces the coverage of Part 4: what matters is not how the Roomba spends its time, but whether the rooms are kept clean. The living room deserves particular attention: it gets dirty fastest, and it is the room the person uses most, so it is where the two goals conflict.
Finally, we keep the tracking accuracies, now one per state factor:
i.e. how often the most probable value under the agent’s marginal belief \(\boldsymbol{s}^{[t,n]}\) equals the true value, for the Roomba’s room, the person’s location, and the dirt in each room. Higher is better, with a maximum of 1. The agent can only avoid the person and find dirty rooms if it knows where they are.
To judge these measures, we compare three Roombas in the same environment:
one that chooses its destination at random, as a reference point,
one with only the Part 4 preference for no motion, which shows what avoiding the person alone does to cleanliness, and
one with both preferences, the agent of this part.
If the cleaning goal works, the agent with both preferences should keep the rooms, and in particular the living room, cleaner than the agent with only the motion preference, at the cost of some extra disturbance. How much extra is the price of cleaning.
The agent does not optimize these measures directly: it cannot see the true states. Instead, it acts on its preferences for observing no motion and dirt picked up. The measures tell us, from the outside, how well those preferences translate into the behavior we want.
Changes in the objective function compared to Part 4
Two main measures for two goals: disturbance, for avoiding the person, and cleanliness, for cleaning.
Disturbance is a condition on two factors: the Roomba’s room and the person’s location are the same. When the person is away, the Roomba cannot disturb them.
Cleanliness replaces coverage: the fraction of ticks each room is dirty, and its average over the three rooms. What matters is not how the Roomba spends its time, but whether the rooms are kept clean. The living room deserves particular attention, since it is where the two goals conflict.
Tracking has one accuracy per state factor, computed from the agent’s marginal beliefs: for the Roomba’s room, the person’s location, and the dirt in each room.
Three Roombas are compared: one moving at random, one with only the Part 4 preference for no motion, and one with both preferences. The difference in disturbance between the last two is the price of cleaning.
The agent acts on two preferences,no motion and dirt picked up, rather than on the measures themselves, which it cannot compute since it never sees the true states.
4.3.6 Implementation of the System-Under-Steer / Environment / Generative Process
## returns a one-hot encoding of a random sample from a categorical distributionfunctionrand_1hot_vec(rng, distribution::Categorical) K =ncategories(distribution) sample =zeros(K) drawn_category =rand(rng, distribution) sample[drawn_category] =1.0return sampleend
rand_1hot_vec (generic function with 1 method)
functioncreate_envir(; LAˣ, LBˣ, LdepsA, LdepsB, Lŝˣ₀, seed=42) rng =MersenneTwister(seed)## the current values (1-based indices) of the factors in deps, read from a tuple of one-hot vectorsvalues_of(Lŝ, deps) =Tuple(argmax(Lŝ[n]) for n in deps)## ô^{[t,m]} ~ Cat(A^{*[m]}[•, values of the factors in D_A^{[m]}]), for every modality mgenerate_obs(Lŝ) =tuple0((rand_1hot_vec(rng, Categorical(LAˣ[m][:, values_of(Lŝ, LdepsA[m])...]))for m ineachindex(LAˣ))...)## t = 0 Lŝˣₜ = Lŝˣ₀ ## ŝ^{*(0)} Lôₜ =generate_obs(Lŝˣₜ) ## ô^{(0)} execute = (Lûₜ₋₁) ->begin## Lûₜ₋₁ = û^{(t-1)}: the action chosen at t-1, a tuple of one-hot vectors u =argmax(Lûₜ₋₁[0]) ## û^{[t-1,0]} as an index: 1 bedroom, 2 livingroom, 3 bathroom Lŝˣₜ₋₁ = Lŝˣₜ## ŝ^{*[t,n]} ~ Cat(B^{*[n]}[•, values of the factors in D_B^{[n]}, (action, for the controlled factor 0)]) Lŝˣₜ =tuple0((begin action = n ==0 ? (u,) : () ## only factor 0 (Room) has an action axis p = LBˣ[n][:, values_of(Lŝˣₜ₋₁, LdepsB[n])..., action...]rand_1hot_vec(rng, Categorical(p))endfor n ineachindex(LBˣ))...) Lôₜ =generate_obs(Lŝˣₜ) ## ô^{(t)}nothingend observe = () -> Lôₜ ## ô^{(t)} = (ô^{[t,0]}, ô^{[t,1]}, ô^{[t,2]}) state = () -> Lŝˣₜ ## ŝ^{*(t)} = (ŝ^{*[t,0]}, ..., ŝ^{*[t,4]}), true state (for evaluation only)return (execute, observe, state)end
create_envir (generic function with 1 method)
Changes in create_envir compared to Part 4
Generic over factors and modalities: the single-factor matrix products are replaced by a lookup in each tensor, using the current values of exactly the factors that tensor depends on. This follows the process specification in 4.3, line by line.
All factors are updated from the previous state: each factor’s next value is drawn from its own tensor, using only the previous state tuple and, for the Room factor, the action. This is what “given the previous state, each factor changes independently” means in code.
values_of turns the one-hot vectors of the factors a tensor depends on into indices, in the order of the tensor’s axes as listed in the dependency tuples.
The action is used only for factor 0 (Room), as an extra index after its dependencies, matching the action axis of its tensor.
generate_obs draws one observation per modality in the same way, both at \(t = 0\) and after each step. The intermediate probability vectors of Part 4 are no longer stored.
LdepsA and LdepsB are new arguments, passed in from the cells that define the likelihood and transition tensors. The dependency lists must match the axis order of the tensors exactly: own factor first for the transitions, then the other dependencies in the listed order.
In order to generate data to mimic the observations of the Roomba, we need to specify four things: the initial state (i.e., where the Roomba and the person start, and which rooms are dirty), the actual transition probabilities \(\mathbf{B}^{\ast}\), one tensor per state factor (i.e., how likely the Roomba is to reach the room it heads for, how the person moves, and how dirt accumulates and is cleaned), the observation distributions \(\mathbf{A}^{\ast}\), one tensor per sensor (i.e., what texture the Roomba will observe in each room, whether it will detect motion, depending on where the person is, and whether it will pick up dirt, depending on whether its room is dirty), and a way to choose the actions. We can then use these specifications to generate observations from the generative process, which has the structure of a factorized partially observable Markov decision process (POMDP).
In this section, we test the environment on its own, with a Roomba that chooses its destination at random. This also gives the reference point for the measures from 4.3.5: how often a Roomba without preferences ends up in the same room as the person, and how dirty the rooms get. In 4.6, the agent’s choices take the place of the random ones.
To generate our data, we’ll follow these steps: 1. Assume an initial state \(\hat{\mathbf{s}}^{\ast(0)}\), one value per factor: the Roomba in the bedroom, the person in the living room, and all rooms clean. 2. Determine the observations in the starting state, one per sensor \(m = 0, 1, 2\) (Texture, Motion, Dirt), by drawing from the categorical distribution that \(\boldsymbol{A}^{\ast[m]}\) gives for the current values of the factors that sensor depends on. 3. Choose an action \(\hat{u}^{[t-1,0]}\), the room the Roomba heads for. Here, it is drawn at random from the three rooms. 4. Determine the next state, one factor at a time: for each factor \(n\), draw its next value from the categorical distribution that \(\boldsymbol{B}^{\ast[n]}\) gives for the previous values of the factors it depends on, and, for the Room factor, the chosen action. The Roomba moves, the person moves, and the dirt in each room grows or is cleaned. 5. Determine the observations in the new state, as in step 2. 6. Repeat steps 3-5 for as many time steps as we want.
We will generate 600 data points (\(t = 0, \ldots, T\) with \(T = 599\)) to simulate 600 ticks of the Roomba moving through the apartment. Lôː will contain the Roomba’s observations at each tick, \(\hat{\mathbf{o}}^{(0:T)}\), where each entry is a tuple holding the texture of the floor it’s on, whether it detects motion, and whether it picks up dirt. Lûː will contain the action chosen at each tick, \(\hat{\mathbf{u}}^{(0:T-1)}\). Lŝˣː will contain the true state at each tick, \(\hat{\mathbf{s}}^{\ast(0:T)}\): where the Roomba and the person actually were, and which rooms were actually dirty.
The following code implements this process and generates our data:
Changes in the data generation compared to Part 4
The four specifications are tuples: the initial state, the transitions \(\mathbf{B}^{\ast}\) and the observations \(\mathbf{A}^{\ast}\) now have one entry per state factor or sensor. The transitions cover the Roomba’s moves, the person’s movements, and the dirt, and the light sensor is replaced by the dirt sensor.
The random Roomba is the reference point for both measures: disturbance and cleanliness.
The initial state has one value per factor: the Roomba in the bedroom, the person in the living room, and all rooms clean.
Each step draws one value per factor: the next state is generated factor by factor, each from its own tensor, using the previous values of the factors it depends on and, for the Room factor, the chosen action.
The observations are drawn per sensor from the tensor values for the current values of the factors that sensor depends on.
The recorded data: the observations hold texture, motion and dirt, and the true states hold where the Roomba and the person were and which rooms were dirty.
T =599## 600 ticks: t = 0, ..., T## factor 0 (Room): 1 bedroom 2 livingroom 3 bathroom## factor 1 (Person): 1 bedroom 2 livingroom 3 bathroom 4 away## factors 2–4 (Dirt in bedroom, livingroom, bathroom): 1 clean 2 dirtyLŝˣ₀ =tuple0([1.0, 0.0, 0.0], ## Roomba in the bedroom [0.0, 1.0, 0.0, 0.0], ## person in the living room [1.0, 0.0], ## bedroom clean [1.0, 0.0], ## living room clean [1.0, 0.0]) ## bathroom clean → ŝ^{*(0)}(execute_sim, observe_sim, state_sim) =create_envir(; LAˣ, LBˣ, LdepsA, LdepsB, Lŝˣ₀)## random Roomba: chooses its destination at random (reference point for 4.3.5)rng_act =MersenneTwister(7) ## separate random stream for the actionsrandom_action(rng) =tuple0(Float64.(1:3.==rand(rng, 1:3))) ## a random one-hot action tupleLŝˣː =offset0(Vector{Any}(undef, T+1)) ## ŝ^{*(0:T)}: true states, each a tuple (Room, Person, Dirt ×3)Lôː =offset0(Vector{Any}(undef, T+1)) ## ô^{(0:T)}: observations, each a tuple (Texture, Motion, Dirt)Lûː =offset0(Vector{Any}(undef, T)) ## û^{(0:T-1)}: chosen actions, each a tuple (Destination)for t =0:Tif t >0 Lûː[t-1] =random_action(rng_act) ## choose û^{(t-1)} (here at random)execute_sim(Lûː[t-1]) ## the environment carries it out: move, update person and dirt, observeend Lŝˣː[t] =state_sim() Lôː[t] =observe_sim()end
Changes in the random Roomba simulation compared to Part 4
The initial state has five factors: the Roomba in the bedroom, the person in the living room, and all rooms clean.
create_envir receives the dependency tuplesLdepsA and LdepsB, which tell the environment which factors each tensor depends on.
The recorded tuples have new contents: each true state holds five factors (Room, Person, and Dirt in each room), and each observation holds Texture, Motion and Dirt.
The loop itself is unchanged: the random Roomba chooses its destination and the environment carries it out, as in Part 4. Each step now also moves the person and updates the dirt in each room.
[Lŝˣ[0] for Lŝˣ in Lŝˣː] ## ŝ^{*[t,0]} for t = 0, ..., T: the Room factor (one-hot over 3 rooms)
## t, Roomba's room, person's location, dirt in bedroom, livingroom, bathroom[(t, room_labels[argmax(Lŝˣː[t][0])], person_labels[argmax(Lŝˣː[t][1])], (dirt_labels[argmax(Lŝˣː[t][n])] for n in2:4)...) for t in0:T]
Each factor is read directly: the Roomba’s room, the person’s location and the dirt in each room come straight from their own factors, so the index arithmetic that decoded the six combined RoomOccupancy states is no longer needed.
Occupancy becomes same_room: the Roomba’s room equals the person’s location, which is the disturbance condition from 4.3.5.
Three panels instead of two: the Roomba’s room with the chosen destinations, the person’s location with markers where the Roomba is in the same room, and the dirt in each room, marking the ticks at which that room is dirty.
The dirt panel shows the dirt dynamics: dirt builds up in a room while the Roomba is away, and disappears when it cleans there.
textures = [argmax(Lô[0]) for Lô inparent(Lôː)] ## texture index (1, 2, 3) of ô^{[t,0]} for t = 0, ..., Tplot(title="Observations of the Random Roomba")plot!(0:T, textures, linetype=:steppost, linestyle=:solid, leg=:right, label="observed texture", xlabel="Time t", yticks=(1:3, ["Hardwood", "Carpet", "Tiles"]))
pickups = [argmax(Lô[2]) for Lô inparent(Lôː)] ## dirt index (1 = nothing, 2 = dirt picked up) of ô^{[t,2]} for t = 0, ..., Tplot(title="Dirt Readings of the Random Roomba")plot!(0:T, pickups, linetype=:steppost, linestyle=:solid, leg=:right, label="dirt reading", xlabel="Time t", yticks=(1:2, ["Nothing", "Dirt picked up"]))
motions = [argmax(Lô[1]) for Lô inparent(Lôː)] ## motion index (1 = no motion, 2 = motion) of ô^{[t,1]} for t = 0, ..., Tplot(title="Motion Readings of the Random Roomba")plot!(0:T, motions, linetype=:steppost, linestyle=:solid, leg=:right, label="observed motion", xlabel="Time t", yticks=(1:2, ["No motion", "Motion"]))
rooms = [argmax(Lŝˣ[0]) for Lŝˣ inparent(Lŝˣː)] ## 1 bedroom, 2 livingroom, 3 bathroompersons = [argmax(Lŝˣ[1]) for Lŝˣ inparent(Lŝˣː)] ## 1 bedroom, 2 livingroom, 3 bathroom, 4 awaysame_room = (rooms .== persons) .+1## 1 no, 2 yes: the person is in the Roomba's roomdirt_here = [argmax(Lŝˣː[t][rooms[t+1] +1]) for t in0:T] ## 1 clean, 2 dirty: the dirt factor of the Roomba's roomtextures = [argmax(Lô[0]) for Lô inparent(Lôː)] ## 1 hardwood, 2 carpet, 3 tilesmotions = [argmax(Lô[1]) for Lô inparent(Lôː)] ## 1 no motion, 2 motionpickups = [argmax(Lô[2]) for Lô inparent(Lôː)] ## 1 nothing, 2 dirt picked updestinations = [argmax(Lû[0]) for Lû inparent(Lûː)] ## 1 bedroom, 2 livingroom, 3 bathroom (t = 0, ..., T-1)window = (0, 150) ## zoom in so individual ticks are visible; use (0, T) for the full runpanel(ts, y, ticks, label, color) =plot(ts, y, linetype=:steppost, label=label, color=color, yticks=ticks, xlims=window, legend=:outerright)plot(panel(0:T-1, destinations, (1:3, ["Bedroom", "Living room", "Bathroom"]), "destination", :gray),panel(0:T, rooms, (1:3, ["Bedroom", "Living room", "Bathroom"]), "Roomba's room", :blue),panel(0:T, textures, (1:3, ["Hardwood", "Carpet", "Tiles"]), "texture", :steelblue),panel(0:T, persons, (1:4, ["Bedroom", "Living room", "Bathroom", "Away"]), "person's location", :red),panel(0:T, same_room, (1:2, ["No", "Yes"]), "same room", :red),panel(0:T, motions, (1:2, ["No motion", "Motion"]), "motion", :darkred),panel(0:T, dirt_here, (1:2, ["Clean", "Dirty"]), "Roomba's room dirty", :brown),panel(0:T, pickups, (1:2, ["Nothing", "Dirt"]), "dirt reading", :sienna), layout=(8, 1), size=(900, 1150), link=:x, xlabel=["""""""""""""""Time t"], plot_title="Random Roomba: actions, hidden states, and sensor readings",)
present = same_room .==2misses =count(present .& (motions .==1)) ## person in the room, no motion detectedfalse_alarms =count(.!present .& (motions .==2)) ## person elsewhere, motion detectedprintln("motion sensor misses: $misses of $(count(present)) ticks with the person present ($(round(100*misses/count(present), digits=1))%)")println("motion sensor false alarms: $false_alarms of $(count(.!present)) ticks without ($(round(100*false_alarms/count(.!present), digits=1))%)")
motion sensor misses: 16 of 95 ticks with the person present (16.8%)
motion sensor false alarms: 13 of 505 ticks without (2.6%)
## Empirical rate of dirt accumulation per room, compared with p_dirty from 4.3.4.## Counts the ticks at which a room was clean and the Roomba was elsewhere,## and how often the room was dirty at the next tick.dirty = [[argmax(Lŝˣ[n]) ==2 for Lŝˣ inparent(Lŝˣː)] for n in2:4] ## dirt in bedroom, livingroom, bathroomfor (r, room) inenumerate(["Bedroom", "Living room", "Bathroom"]) chances = [t for t in1:T if !dirty[r][t] && rooms[t] != r] ## clean, and the Roomba elsewhere (1-based positions) rate =mean(dirty[r][t+1] for t in chances)println("$room: got dirty at $(round(rate, digits=3)) of $(length(chances)) chances (p_dirty: $(p_dirty[r]))")end
Bedroom: got dirty at 0.021 of 423 chances (p_dirty: 0.02)
Living room: got dirty at 0.05 of 341 chances (p_dirty: 0.05)
Bathroom: got dirty at 0.035 of 284 chances (p_dirty: 0.03)
## ---------- random Roomba: reference values for 4.3.5 ----------rooms_random = [argmax(Lŝˣ[0]) for Lŝˣ inparent(Lŝˣː)]persons_random = [argmax(Lŝˣ[1]) for Lŝˣ inparent(Lŝˣː)]dirty_random = [[argmax(Lŝˣ[n]) ==2 for Lŝˣ inparent(Lŝˣː)] for n in2:4] ## dirt in bedroom, livingroom, bathroomdisturbance_random =mean(rooms_random .== persons_random) ## J_disturb: fraction of ticks in the person's roomdirtiness_random = [mean(dirty_random[r]) for r in1:3] ## J_dirty(r): fraction of ticks each room is dirtycoverage_random = [mean(rooms_random .== r) for r in1:3] ## fraction of ticks in each room, for referenceprintln("random Roomba, disturbance: $(round(100*disturbance_random, digits=1))% of ticks in the person's room")println("random Roomba, dirty: ",join(["$(r): $(round(100*dirtiness_random[i], digits=1))%" for (i, r) inenumerate(["bedroom", "livingroom", "bathroom"])], ", ")," (average $(round(100*mean(dirtiness_random), digits=1))%)")println("random Roomba, coverage: ",join(["$(r): $(round(100*coverage_random[i], digits=1))%" for (i, r) inenumerate(["bedroom", "livingroom", "bathroom"])], ", "))
random Roomba, disturbance: 15.8% of ticks in the person's room
random Roomba, dirty: bedroom: 8.7%, livingroom: 20.7%, bathroom: 3.2% (average 10.8%)
random Roomba, coverage: bedroom: 22.5%, livingroom: 25.8%, bathroom: 51.7%
Changes in the random Roomba’s reference values compared to Part 4
Disturbance is now the fraction of ticks the Roomba spends in the person’s room, computed from the Room and Person factors.
New: cleanliness, the fraction of ticks each room is dirty, and its average over the three rooms, the second main measure of this part.
Coverage is kept as background: it is no longer a main measure, but helps explain the cleanliness, since a room that is rarely visited is often dirty.
The values are saved under their own names,disturbance_random, dirtiness_random and coverage_random, so that they survive the agent’s run in 4.6, which overwrites the recorded states. This cell replaces the separate reference-value cell of Part 4.
4.4 Uncertainty Model
The uncertainty in the environment is captured in the tuple of state transition tensors \(\mathbf{B}^{\ast} = \big(\boldsymbol{B}^{\ast[0]}, \ldots, \boldsymbol{B}^{\ast[4]}\big)\), one per state factor, and the tuple of observation likelihood tensors \(\mathbf{A}^{\ast} = \big(\boldsymbol{A}^{\ast[0]}, \boldsymbol{A}^{\ast[1]}, \boldsymbol{A}^{\ast[2]}\big)\), for Texture, Motion and Dirt. Their entries are the probabilities of the Roomba’s moves, including moves that fail, of the person’s movements, of dirt accumulating and being cleaned, and of the readings of the camera, the motion sensor and the dirt sensor. With several factors, the uncertainty is spread over several small tensors, each describing one aspect of the environment. The actions themselves add no uncertainty: the agent knows which action it chose, only not exactly what will come of it. The initial state adds no uncertainty either, because the run always starts from the same state: the Roomba in the bedroom, the person in the living room, and all rooms clean.
On the agent’s side, there is no uncertainty about how the environment works: in this part, the agent is given \(\mathbf{A} = \mathbf{A}^{\ast}\) and \(\mathbf{B} = \mathbf{B}^{\ast}\). What remains is uncertainty about the state: where the Roomba is, where the person is, and which rooms are dirty, and, when planning, where its actions will lead, whether it will meet the person there, and whether there will be dirt to clean. The uncertainty differs per factor. The Roomba’s room is well tracked, since the agent knows where it was heading and the camera confirms it. The person is only noticed when they are in the Roomba’s room. And the dirt in other rooms is never observed at all: the agent can only predict it, so its uncertainty about those rooms grows the longer the Roomba stays away. The agent handles all this uncertainty with its beliefs, and takes it into account when choosing actions.
Changes in the uncertainty model compared to Part 4
The uncertainty is spread over several small tensors:\(\mathbf{B}^{\ast}\) holds one transition tensor per state factor, covering the Roomba’s moves, the person’s movements, and the dirt, and \(\mathbf{A}^{\ast}\) holds the tensors for Texture, Motion and Dirt.
The initial state is the fixed five-factor state: the Roomba in the bedroom, the person in the living room, and all rooms clean.
The agent’s uncertainty differs per factor: the Roomba’s room is well tracked, the person is only noticed when they are in the Roomba’s room, and the dirt in other rooms is never observed, so the agent’s uncertainty about it grows the longer the Roomba stays away.
4.5 Agent / Generative Model
4.5.1 State and Observation variables
On the agent’s side, the state of the environment at time \(t\) is not known: the agent represents it by a (compound) tuple of categorical random variables, one per state factor (here five factors):
factor 0 (Room): \(s^{[t,0]} \in \{\text{bedroom}, \text{livingroom}, \text{bathroom}\}\), where the Roomba is
factor 1 (Person): \(s^{[t,1]} \in \{\text{bedroom}, \text{livingroom}, \text{bathroom}, \text{away}\}\), where the person is
factors 2, 3 and 4 (Dirt in the bedroom, living room and bathroom): \(s^{[t,n]} \in \{\text{clean}, \text{dirty}\}\), \(n = 2, 3, 4\)
The agent never observes these variables directly. What it holds instead is a belief. Internally, this is the joint belief \(\boldsymbol{s}^{(t)} \in [0,1]^{96}\), with one probability per combination of the five factors. It is convenient to arrange it as an array with one axis per factor, of size \(3 \times 4 \times 2 \times 2 \times 2\), in the factor order used for the axes of the tensors in 4.1. From it, the agent’s belief about a single factor, the marginal \(\boldsymbol{s}^{[t,n]}\), follows by adding up over all other axes. For example, \(\boldsymbol{s}^{[t,1]} \in [0,1]^{4}\) is the agent’s belief about where the person is.
The agent needs the joint belief, rather than five separate marginals, because its sensors tie the factors together. When the motion sensor detects motion, the agent learns little about the Roomba’s room or the person’s location on their own, but a lot about the two together: they are probably in the same room. Likewise, when the dirt sensor picks up dirt, the agent learns that whichever room the Roomba is in is dirty. Only the joint belief can hold such statements; the marginals lose them, as noted in 4.1.
where \(\boldsymbol{D}^{[n]}\) parameterizes the categorical distribution of \(s^{[0,n]}\). Unlike the environment, the agent does not know the starting state, so each factor gets a uniform prior: \(\boldsymbol{D}^{[0]} = [\tfrac{1}{3}, \tfrac{1}{3}, \tfrac{1}{3}]^{\top}\) for the room, \(\boldsymbol{D}^{[1]} = [\tfrac{1}{4}, \tfrac{1}{4}, \tfrac{1}{4}, \tfrac{1}{4}]^{\top}\) for the person, and \(\boldsymbol{D}^{[n]} = [\tfrac{1}{2}, \tfrac{1}{2}]^{\top}\) for the dirt in each room, \(n = 2, 3, 4\). Together, they give the uniform joint prior over all 96 combinations, \(\tfrac{1}{96}\) each. The prior is the only point at which the factors are independent: once the first observations arrive, the joint belief in general no longer splits into a product of its marginals.
The (compound) observation tuple received by the agent at time \(t\) is the same one the environment generates:
one-hot encoded as \(\hat{\boldsymbol{o}}^{[t,2]} \in \big\{[1,0]^{\top}, [0,1]^{\top}\big\}\)
As in Part 4, the agent also chooses actions, and it prefers some observations over others. These are described in 4.5.2 and 4.5.3.
Changes in the agent’s state and observation variables compared to Part 4
Five state factors replace the combined RoomOccupancy factor: Room, Person, and Dirt in each of the three rooms, with the same values as in the environment.
The belief is joint, with marginals derived from it: the agent holds a belief over all 96 combinations, arranged as a \(3 \times 4 \times 2 \times 2 \times 2\) array with one axis per factor, and the belief about each factor follows by adding up over the other axes.
Why joint: the motion and dirt sensors each depend on more than one factor, so their readings tell the agent how factors are related, e.g. that the Roomba and the person are probably in the same room. Separate marginals cannot represent this.
The initial prior is a tuple\(\mathbf{D} = \big(\boldsymbol{D}^{[0]}, \ldots, \boldsymbol{D}^{[4]}\big)\) of uniform vectors, one per factor. Their product is the uniform prior over the 96 joint states.
The modalities are renumbered: Texture (0), Motion (1) and Dirt (2). The light sensor is dropped.
4.5.2 Decision variables
As in Part 4, the agent makes decisions. At each time step, it chooses an action: the room the Roomba heads for. The agent represents its action at time \(t\) by a (compound) tuple of categorical random variables, one per control factor (here a single factor):
\[
\mathbf{u}^{(t)} = \big(u^{[t,0]}\big)
\]
where the control factors are:
control factor 0 (Destination), acting on state factor 0 (Room):
Heading for the room the Roomba is already in means staying put. Unlike the state, the agent always knows which action it has taken: once chosen, the action is a realized value \(\hat{\boldsymbol{u}}^{[t,0]}\), which the agent hands to the environment and also uses itself, to predict how the state will change. What the agent does not know is the outcome: whether the move succeeds, whether the person is in the room it reaches, and whether there is dirt there to clean.
The action selects the transition matrix of the Room factor only. Its effect on the other factors runs through their dependencies: the dirt factors depend on where the Roomba is, so in its predictions the agent’s choice of room also determines which room gets cleaned. The person’s movements, on the other hand, are predicted the same way whatever the agent does.
Policies. To choose an action, the agent looks \(H\) steps ahead. It considers candidate policies, i.e. sequences of \(H\) actions, collected as the rows of a matrix \(\boldsymbol{V}\): policy \(p\) is \(\pi^{[p]} = \boldsymbol{V}[p, \bullet]\), and \(\boldsymbol{V}[p, \tau]\) is the action it prescribes at step \(\tau = 0, \ldots, H-1\). With three possible actions per step, there are \(3^H\) policies. In this part, \(H = 3\), which gives 27 policies. Three steps are needed for cleaning to pay off: from the bedroom, reaching the living room takes two moves, through the bathroom, and the dirt is picked up and removed only once the Roomba is there. A policy such as head for the bathroom, then the living room, then stay lets the agent see that pay-off. The agent scores each policy by its expected free energy (see 4.5.6), turns the scores into a probability for each policy, and from those derives a probability for each possible next action. Only the first action of the chosen policy is carried out; at the next time step, the agent observes the result and plans again.
Changes in the agent’s decision variables compared to Part 4
The control factor acts on state factor 0 (Room), instead of the combined RoomOccupancy factor.
The unknown outcome of an action now includes whether there is dirt to clean, besides whether the move succeeds and whether the person is there.
The action’s effect on other factors is explicit: it directly selects only the Room transition, but through the dependencies of the dirt factors on the Roomba’s room, it also determines, in the agent’s predictions, which room gets cleaned. The person’s predicted movements are the same whatever the agent does.
The planning horizon is \(H = 3\), giving 27 policies instead of 9. Reaching the living room from the bedroom takes two moves, and the pay-off from cleaning comes only once the Roomba is there, so a shorter horizon would not let the agent see the benefit.
4.5.3 Preference variables / Setpoints
The agent’s goals are expressed as preferences over observations: the agent prefers to receive the observations it wants to be true. There is one preference vector per modality, collected in \(\mathbf{C} = \big(\boldsymbol{C}^{[0]}, \boldsymbol{C}^{[1]}, \boldsymbol{C}^{[2]}\big)\). Each \(\boldsymbol{C}^{[m]}\) is a probability vector over the observations of modality \(m\), playing the role of a setpoint: the distribution of observations the agent would like to see.
\(\boldsymbol{C}^{[0]}\) (Texture) is flat, \(\boldsymbol{C}^{[0]} = [\tfrac{1}{3}, \tfrac{1}{3}, \tfrac{1}{3}]^{\top}\): the agent has no preference for any floor type.
\(\boldsymbol{C}^{[1]}\) (Motion) prefers no motion, e.g. \(\boldsymbol{C}^{[1]} = [0.9, 0.1]^{\top}\): the agent would like to observe no motion 90% of the time. This is the goal of Part 4: not disturbing anyone.
\(\boldsymbol{C}^{[2]}\) (Dirt) prefers dirt picked up, e.g. \(\boldsymbol{C}^{[2]} = [0.2, 0.8]^{\top}\): the agent would like its brushes to pick up dirt 80% of the time. This is the new goal: cleaning.
The agent cannot prefer unoccupied rooms or clean rooms directly, because it never observes the person’s location or the dirt. Instead, it prefers the observations that go with what it wants, and chooses actions it expects to produce them. This is the usual way of expressing goals in active inference: preferences are stated over what the agent can sense, not over the hidden state.
For cleaning, the choice of observation needs care. It may seem natural to prefer observing a clean floor, i.e. nothing picked up, but that preference would backfire: the agent would be drawn to rooms that are already clean, where it has nothing to do. Preferring dirt picked up draws it to rooms that are dirty, which are exactly the rooms where its presence makes a difference. Once a room is clean, the preference is no longer satisfied there, and it pulls the agent on to wherever dirt is most likely to have built up: the living room, which gets dirty fastest, and rooms it has not visited for a while.
How strongly the agent holds each preference is set by how far its \(\boldsymbol{C}^{[m]}\) is from flat. When scoring a policy, the agent compares the observations it expects under that policy with its preferences (see 4.5.6), and \(\log \boldsymbol{C}^{[m]}\) determines how much each expected reading counts:
with \(\boldsymbol{C}^{[1]} = [0.9, 0.1]^{\top}\), observing motion is \(\log(0.9/0.1) \approx 2.2\) worse than observing no motion, and
with \(\boldsymbol{C}^{[2]} = [0.2, 0.8]^{\top}\), observing dirt picked up is \(\log(0.8/0.2) \approx 1.4\) better than observing nothing.
The ratio of the two sets the agent’s balance between its goals. Here, avoiding the person weighs somewhat more than cleaning, so the agent should not simply go wherever dirt is likely, but clean the living room when it expects it to be quiet. A flat \(\boldsymbol{C}^{[2]}\) gives back the agent of Part 4, with only the preference for no motion; this is the second of the three Roombas we compare (see 4.3.5). More extreme preferences make the agent’s choices more driven by its goals, and less by reducing its uncertainty about the state.
The flat preference for Texture does not mean the camera is useless: it still helps the agent infer where the Roomba is, which it needs in order to predict where its actions will lead, and which room’s dirt the dirt sensor is sensing. It just does not make any room more attractive than another.
Changes in the preferences compared to Part 4
Two preferences instead of one:\(\boldsymbol{C}^{[1]}\) for no motion, the goal of Part 4, and the new \(\boldsymbol{C}^{[2]}\) for dirt picked up, the cleaning goal. Only Texture has a flat preference; the light sensor is dropped.
The cleaning preference is stated over dirt being picked up, not over a clean floor. A preference for a clean floor would draw the agent to rooms that are already clean. Preferring dirt picked up draws it to dirty rooms, and keeps pulling it to wherever dirt is likely to have built up.
The balance between the two goals is set by the relative strength of the preferences: observing motion counts about 2.2 against the agent, observing dirt picked up about 1.4 in its favor. A flat \(\boldsymbol{C}^{[2]}\) gives back the Part 4 agent, the second of the three Roombas compared.
The camera has an extra role: knowing the Roomba’s room also tells the agent which room’s dirt the dirt sensor is reading.
4.5.4 State Transition and Observation Generation functions
The state transition functions, one per state factor, are given by
where, as in 4.3, the values of the random variables are used as indices into the tensors. Each line picks out the column of \(\boldsymbol{B}^{[n]}\) for the previous values of exactly the factors in \(\mathcal{D}_B^{[n]}\), and, for the Room factor, the action chosen at \(t-1\). Given the previous state and action, the factors change independently of each other, so the transition of the whole state is the product
where each line picks out the column of \(\boldsymbol{A}^{[m]}\) for the current values of exactly the factors in \(\mathcal{D}_A^{[m]}\). It parameterizes the categorical distribution of \(o^{[t,m]}\): over textures for \(m = 0\), over motion readings for \(m = 1\), and over dirt readings for \(m = 2\). Observations do not depend on the action.
In this part, the agent is given its tensors, \(\boldsymbol{A}^{[m]} = \boldsymbol{A}^{\ast[m]}\) and \(\boldsymbol{B}^{[n]} = \boldsymbol{B}^{\ast[n]}\), and it knows their dependencies, which are the same as the environment’s. The tensors are known parameters, not random variables, so they appear as subscripts.
An entry of \(\boldsymbol{A}^{[m]}\) is the probability of a specific observation given specific values of the factors it depends on, and each column, obtained by fixing all axes except the first, is a categorical distribution over observations. The same holds for each \(\boldsymbol{B}^{[n]}\), with the next value of factor \(n\) in place of the observation. If the state were known, the right column could simply be looked up. The agent, however, only has its joint belief \(\boldsymbol{s}^{(t)}\), so it takes the weighted average of the columns over all 96 combinations, each weighted by its probability under the belief. To write this down, let \(\mathbf{j} = (j_0, \ldots, j_4)\) denote one combination of factor values, so that \(\boldsymbol{s}^{(t)}[\mathbf{j}]\) is its probability, and let \(j_{\mathcal{D}}\) denote the values of the factors in a dependency set \(\mathcal{D}\), in order. For example, \(j_{\mathcal{D}_A^{[1]}} = (j_0, j_1)\), the Roomba’s room and the person’s location.
Using the functions for inference. After carrying out action \(\hat{u}^{[t-1,0]}\), the agent first predicts the new state, by moving the probability of each combination \(\mathbf{i}\) in its previous belief to the combinations \(\mathbf{j}\) it can lead to:
Then it updates this prediction with the observations it receives. Given the state, the three sensors are independent of each other, so it combines them by multiplying their likelihoods:
where \(\hat{o}^{[t,m]}\), the received observation without bold, is used as an index, and the result is normalized to sum to 1 over all 96 combinations. At \(t = 0\), the prediction is replaced by the uniform initial prior from 4.5.1. A combination that explains all readings well gains belief, while one that explains only some of them loses some. Each sensor contributes differently:
Texture depends only on the room, so it tells the rooms apart.
Motion depends on the room and the person, and distinguishes combinations in which they are in the same room from those in which they are not. It says little about where the person is when they are elsewhere.
Dirt depends on the room and the three dirt factors, and tells the agent whether the room the Roomba is in is dirty. It says nothing about the dirt in the other rooms: that the agent can only predict, from how dirt accumulates.
This is exact Bayesian filtering over the 96 joint states, which is feasible because the number of combinations is small.
Using the functions for planning. To evaluate policy \(p\), the agent applies the same functions to the future, step by step, starting from its current belief. Its predicted joint beliefs are
Here \(\boldsymbol{s}_{\pi^{[p]}}^{(\tau)}\) is the predicted joint belief about time \(t + \tau + 1\) if the agent follows policy \(p\), and \(\boldsymbol{o}_{\pi^{[p]}}^{[\tau,m]}\) is the observation of sensor \(m\) it then expects. Unlike the example in 4.1, which treated the room and the person as independent, the expected observation is computed from the joint belief, so it also accounts for any relation between the factors a sensor depends on. No real observations are available for the future, so these predictions are never updated; they are what the agent compares with its preferences when it scores policy \(p\) (see 4.5.6).
The two preferences make the corresponding predictions the ones that matter most:
\(\boldsymbol{o}_{\pi^{[p]}}^{[\tau,1]}\) tells the agent how likely it is to observe motion if it follows policy \(p\). Since the person’s movements do not depend on the action, this follows from where the policy takes the Roomba and where the person is predicted to be, which, without new evidence, drifts towards the living room.
\(\boldsymbol{o}_{\pi^{[p]}}^{[\tau,2]}\) tells the agent how likely it is to pick up dirt. In the predictions, dirt grows in every room the Roomba is not in, and is removed in the room it is in. So a policy that heads for a room the agent has not visited for a while, especially the living room, is expected to pick up dirt, while a policy that stays in a room it has just cleaned is not.
Each policy is predicted over all 96 joint states, but with 27 policies and three steps this remains a small computation.
Changes in the agent’s transition and observation functions compared to Part 4
One transition function per state factor, each depending only on the factors in \(\mathcal{D}_B^{[n]}\) and, for the Room factor, the action. Their product is the transition of the whole state, \(P_{\mathbf{B}}\), since the factors change independently given the previous state.
One observation function per sensor, each depending only on the factors in \(\mathcal{D}_A^{[m]}\): Texture on the room, Motion on the room and the person, and Dirt on the room and the dirt in all three rooms. The light sensor is dropped.
Indexing the joint belief:\(\mathbf{j} = (j_0, \ldots, j_4)\) denotes one of the 96 combinations, \(\boldsymbol{s}^{(t)}[\mathbf{j}]\) its probability, and \(j_{\mathcal{D}}\) the values of the factors in a dependency set. Matrix-vector products are replaced by sums over combinations.
Inference is exact filtering over the joint states: predict with \(P_{\mathbf{B}}\), then multiply by the likelihoods of the three received observations and normalize. Texture tells the rooms apart, Motion whether the person is in the Roomba’s room, and Dirt whether the Roomba’s room is dirty. Dirt elsewhere is only predicted.
Planning predicts joint beliefs\(\boldsymbol{s}_{\pi^{[p]}}^{(\tau)}\), and expected observations are computed from them, so relations between factors are taken into account.
Two predictions drive the agent’s choices: expected motion, which depends on where the policy takes the Roomba and where the person is predicted to be, and expected dirt pick-up, which favors rooms the Roomba has not visited for a while.
4.5.5 Implementation of the Agent / Generative Model / Internal Model
We start by specifying a probabilistic model for the agent that describes the agent’s internal beliefs about the external dynamics of the environment. As in Part 4, this model has two parts. Up to the current time \(t\), it describes what has happened: the states the environment went through, the observations the Roomba received, and the actions \(\hat{u}^{[0:t-1,0]}\) the agent actually took. Beyond \(t\), it describes what could happen over the next \(H\) steps, depending on which policy \(\pi^{(t)}\) the agent chooses. With several factors and sensors, we write the model in terms of the tuples \(\mathbf{s}^{(t')}\) and \(\mathbf{o}^{(t')}\). The generative model at time \(t\) is defined as follows:
where \(P_D\) is the agent’s uniform initial state prior from 4.5.1, and \(P_{\mathbf{A}}\) and \(P_{\mathbf{B}}\) combine the functions from 4.5.4, one per sensor and one per factor:
with, as in 4.5.4, the dots standing for exactly the factors in \(\mathcal{D}_A^{[m]}\) and \(\mathcal{D}_B^{[n]}\), and, for the Room factor, the action \(u\). The tensors \(\mathbf{A} = \mathbf{A}^{\ast}\) and \(\mathbf{B} = \mathbf{B}^{\ast}\) and the prior \(\mathbf{D}\) are known parameters, so they appear as subscripts. As in Part 4, there are no parameter priors: the agent learns nothing about its tensors in this part.
Written out, the model is a product of many small pieces, each involving only a few variables: one prior per factor, one observation function per sensor and time step, and one transition function per factor and time step. This is what a factorized model means, and it is why each tensor can stay small. The action appears in only one of these pieces, the transition of the Room factor; its influence on the dirt runs through the dirt transitions’ dependence on the Roomba’s previous room.
The two transition products differ only in where the action comes from. In the past, the action is the one the agent actually took, \(\hat{u}^{[t'-1,0]}\), a known value. In the future, it is the action \(\boldsymbol{V}[\pi^{(t)}, \tau]\) that the policy prescribes at step \(\tau\); since the policy is not yet chosen, it is a random variable. Step \(\tau\) predicts time \(t + \tau + 1\), as in 4.5.4.
The policy prior\(P\big(\pi^{(t)}\big)\) is where the agent’s preferences enter the model. In active inference, it is not chosen freely: policies are a priori more probable the lower their expected free energy,
where \(\boldsymbol{G}\) collects the expected free energies \(G^{[p]}\) of the 27 policies, which depend on both preferences in \(\mathbf{C}\) from 4.5.3 (see 4.5.6), and \(\sigma\) is the softmax. The agent thus expects to act in ways that lead to its preferred observations: no motion, and dirt picked up.
The product inside \(P_{\mathbf{A}}\) is where the three sensors are combined. Each sensor depends on its own subset of the factors, and given the whole state, the three readings are independent. For the future, these observations are not yet available: the agent uses their expected values, \(\boldsymbol{o}_{\pi^{[p]}}^{[\tau,m]}\) from 4.5.4, to evaluate each policy.
Note that the factorization is a property of the model, not of the agent’s beliefs. The model is a product of small pieces, but once observations come in, the posterior over the state no longer factorizes, as explained in 4.5.1. This is why the agent works with the joint belief over all 96 combinations. In larger models, where the joint states become too many, the belief itself is approximated by a product of per-factor beliefs. That approximation is left for a later part.
Changes in the generative model compared to Part 4
The model is written in terms of tuples: the state \(\mathbf{s}^{(t')}\) of five factors and the observation \(\mathbf{o}^{(t')}\) of three sensors, with \(P_{\mathbf{A}}\) and \(P_{\mathbf{B}}\) as products of the per-sensor and per-factor functions from 4.5.4, and \(P_D\) as the product of the five uniform per-factor priors.
The model is factorized: a product of small pieces, each involving only the factors it depends on. The action enters only the transition of the Room factor, and reaches the dirt through the dirt transitions.
The policy prior covers 27 policies, scored with both preferences: no motion, and dirt picked up.
Each sensor depends on its own subset of factors, and the three readings are independent given the whole state.
The model factorizes, but the beliefs do not: once observations arrive, the posterior over the state couples the factors, which is why the agent works with the 96 joint states. A factorized approximation of the beliefs is left for a later part.
Now it is time to build our agent. As in Part 4, the agent does not learn anything: it is given its model of the world, i.e. the observation tensors \(\mathbf{A} = \mathbf{A}^{\ast}\) with their dependencies \(\mathcal{D}_A^{[m]}\), the transition tensors \(\mathbf{B} = \mathbf{B}^{\ast}\) with their dependencies \(\mathcal{D}_B^{[n]}\), the uniform initial priors \(\mathbf{D}\) from 4.5.1, and its preferences \(\mathbf{C}\) from 4.5.3. What it has to do is use this model, at every tick, in three steps.
1. Update its belief (perception). After the environment has carried out the previous action \(\hat{u}^{[t-1,0]}\) and returned the new observations \(\hat{\mathbf{o}}^{(t)}\), the agent first predicts the new joint state, then corrects that prediction with what its sensors report, as in 4.5.4:
With the joint belief stored as a \(3 \times 4 \times 2 \times 2 \times 2\) array, both parts are array operations:
the prediction moves probability between combinations, each factor according to its own tensor, using the previous values of the factors it depends on. The 96 × 96 transition matrix of the joint states is never written out.
the likelihood of each sensor is the slice \(\boldsymbol{A}^{[m]}[\hat{o}^{[t,m]}, \ldots]\) for the received observation: an array over only the factors that sensor depends on, e.g. \(3 \times 4\) for Motion. It is spread out over the axes of the other factors, which it does not depend on, and the three slices are multiplied entry by entry with the prediction.
The result is normalized to sum to 1. At \(t = 0\), the prediction is replaced by the uniform joint prior, the product of the \(\boldsymbol{D}^{[n]}\). With known tensors and only 96 joint states, this is exact Bayesian filtering: no variational approximation is needed. For recording and evaluation, the agent also computes the marginals \(\boldsymbol{s}^{[t,n]}\) from the joint belief, by adding up over the other axes.
2. Evaluate the policies (planning). For each of the \(3^H = 27\) policies \(\pi^{[p]}\), the agent predicts its joint beliefs and expected observations over the next \(H = 3\) steps, as in 4.5.4, and scores the policy by its expected free energy \(G^{[p]}\) (see 4.5.6). The score favors policies that are expected to lead to preferred observations, i.e. no motion and dirt picked up, and policies that are expected to reduce the agent’s uncertainty about the state, e.g. by going back to check on a room it has not seen for a while.
3. Choose an action (decision). The scores are turned into a posterior over policies, \(Q\big(\pi^{(t)} = \pi^{[p]}\big) = \sigma\big(-\boldsymbol{G}\big)_p\), and from it into a distribution over the next action, \(Q\big(u^{[t,0]}\big)\), by adding up the probabilities of all policies that start with that action. The agent then carries out the most probable action, \(\hat{u}^{[t,0]}\), and the cycle repeats at the next tick.
The model is larger than in Part 4, with 96 joint states instead of six, and 27 policies instead of 9, but still small enough for all three steps to be computed directly in Julia, without RxInfer. As in create_envir, the code is generic over factors and modalities: the dependency tuples LdepsA and LdepsB tell it which axes of the joint belief each tensor acts on, so the same few functions handle every factor and sensor. This keeps every quantity of Algorithm 18 visible in the code. RxInfer also supports active inference, as planning formulated as inference, which becomes attractive for larger models, where neither the joint states nor the policies can be enumerated; this is a possible direction for a later part. Let’s build the agent!
Changes in the agent’s three steps compared to Part 4
Perception works on the joint belief: the prediction moves probability between the 96 combinations factor by factor, without building a 96 × 96 matrix, and each sensor’s likelihood is a slice of its tensor over only the factors it depends on, spread out over the other axes. The agent also computes the marginals of each factor, for recording.
The initial prior is the product of the five uniform per-factor priors \(\boldsymbol{D}^{[n]}\).
Planning evaluates 27 policies over \(H = 3\) steps, scored with both preferences, no motion and dirt picked up, plus the value of reducing uncertainty, e.g. checking on a room the Roomba has not visited for a while.
The decision step is unchanged.
The code is generic over factors and modalities, like create_envir: the dependency tuples tell each function which axes of the joint belief a tensor acts on.
## ---------- the agent's model: given, not learned ----------LA = LAˣ ## A^{[m]} = A^{*[m]}, m = 0, 1, 2 (dependencies D_A^{[m]} in LdepsA)LB = LBˣ ## B^{[n]} = B^{*[n]}, n = 0, ..., 4 (dependencies D_B^{[n]} in LdepsB)LD =tuple0(fill(1/3, 3), ## D^{[0]}: Room, uniformfill(1/4, 4), ## D^{[1]}: Person, uniformfill(1/2, 2), fill(1/2, 2), fill(1/2, 2)) ## D^{[2]}, D^{[3]}, D^{[4]}: Dirt, uniformLC =tuple0(fill(1/3, 3), ## C^{[0]}: Texture, flat [0.9, 0.1], ## C^{[1]}: Motion, prefers no motion [0.2, 0.8]) ## C^{[2]}: Dirt, prefers dirt picked upLC_motion =tuple0(fill(1/3, 3), [0.9, 0.1], fill(1/2, 2)) ## the Part 4 agent: no cleaning goal (for comparison in 4.6)## ---------- the joint state: an array with one axis per factor ----------sizes =Tuple(size(LB[n], 1) for n ineachindex(LB)) ## number of values per factor: (3, 4, 2, 2, 2), 96 combinations## ---------- policies: all action sequences of length H ----------H =3## planning horizon_V = [p[τ] for p invec(collect(Iterators.product(ntuple(_ ->1:3, H)...))), τ in1:H] ## P×H; column τ+1 is step τP =size(_V, 1) ## number of policies, 3^H = 27## ---------- helpers ----------normalize(v) = v ./sum(v)log_stable(x) =log.(max.(x, 1e-16)) ## avoids log(0)softmax(x) =normalize(exp.(x .-maximum(x)))## the marginal of the joint belief over the factors in deps, with its axes in the order of depsfunctionmarginal(_s, deps) others =Tuple(k for k in1:ndims(_s) if (k-1) ∉ deps) ## axes to add up over (axis k is factor k-1) m =dropdims(sum(_s, dims=others), dims=others) ## axes now in increasing factor orderpermutedims(m, invperm(sortperm(deps))) ## reorder to the order of depsendmarginals(_s) =tuple0((marginal(_s, [n]) for n ineachindex(LB))...) ## (s^{[t,0]}, ..., s^{[t,4]}), for recording## an array over the factors in deps (axes in the order of deps), spread out over the joint state's other axesfunctionspread(arr, deps) a =permutedims(arr, sortperm(deps)) ## axes in increasing factor orderreshape(a, ntuple(k -> (k-1) in deps ? sizes[k] :1, length(sizes))) ## size 1 on the other axes: broadcastsend_D =reduce((a, b) -> a .* b, (spread(LD[n], [n]) for n ineachindex(LD))) ## joint prior: ∏_n D^{[n]}, 1/96 each## H[A^{[m]}]: entropy of each column of A^{[m]}, an array over the factors in D_A^{[m]}LH =tuple0((dropdims(-sum(LA[m] .*log_stable(LA[m]), dims=1), dims=1) for m ineachindex(LA))...)## ---------- prediction: Σ_i P_B(j | i, u) s[i], one factor at a time ----------## each factor must use the *previous* values of the factors it depends on, so a factor is updated## before the factors it depends on: the dirt factors (which depend on the Roomba's previous room)## before the Room factorupdate_order = [4, 3, 2, 1, 0]@assertall(findfirst(==(n), update_order) <findfirst(==(d), update_order)for n ineachindex(LB) for d in LdepsB[n] if d != n)## moves the probability of each combination along the axis of factor n, using B^{[n]}functiontransition(_s, n, u) _s′ =zero(_s) action = n ==0 ? (u,) : () ## only factor 0 (Room) has an action axisfor I inCartesianIndices(_s) p = _s[I] p ==0&&continue i =Tuple(I) ## the combination (1-based values of all factors) col =view(LB[n], :, (i[d+1] for d in LdepsB[n])..., action...) ## B^{[n]}[•, values of D_B^{[n]}, (u)]for (j, b) inenumerate(col) _s′[Base.setindex(i, j, n+1)...] += b * p ## factor n takes value j, the others keep theirsendend _s′endfunctionpredict(_s, u)for n in update_order _s =transition(_s, n, u)end _send## ---------- step 1: perception (exact Bayesian filtering over the joint states) ----------## s^{(t)}[j] ∝ (∏_m A^{[m]}[ô^{[t,m]}, j_{D_A^{[m]}}]) · Σ_i P_B(j | i, û^{[t-1,0]}) s^{(t-1)}[i]; at t = 0 the prediction is the joint priorfunctionperceive(_sₜ₋₁, Lôₜ, ûₜ₋₁) prediction = _sₜ₋₁ ===nothing ? _D :predict(_sₜ₋₁, ûₜ₋₁) likelihood =reduce((a, b) -> a .* b, (spread(selectdim(LA[m], 1, argmax(Lôₜ[m])), LdepsA[m]) for m ineachindex(LA))) ## the slices A^{[m]}[ô^{[t,m]}, ...], spread outnormalize(likelihood .* prediction) ## _sₜ = s^{(t)}end## ---------- step 2: expected free energy of policy p ----------## G^{[p]} = Σ_τ Σ_m [ o_π^{[τ,m]}·(log o_π^{[τ,m]} - log C^{[m]}) + Σ_j H[A^{[m]}][j_{D_A^{[m]}}] s_π^{(τ)}[j] ]## (risk: expected vs. preferred observations) (ambiguity: expected sensor noise)functionefe(_sₜ, p, LC) _s = _sₜ G =0.0for τ in1:H ## code τ = math τ + 1 _s =predict(_s, _V[p, τ]) ## s_π^{(τ)}: predicted joint belieffor m ineachindex(LA) q =marginal(_s, LdepsA[m]) ## predicted belief over the factors sensor m depends on o =reshape(LA[m], size(LA[m], 1), :) *vec(q) ## o_π^{[τ,m]}: expected observation G += o'* (log_stable(o) -log_stable(LC[m])) +sum(LH[m] .* q)endend Gend## ---------- step 3: from policy scores to an action ----------functionplan(_sₜ, LC) _G = [efe(_sₜ, p, LC) for p in1:P] ## G^{[p]}, one per policy _qπ =softmax(-_G) ## Q(π^{(t)} = π^{[p]}) = σ(-G)_p _qu = [sum(_qπ[_V[:, 1] .== u]) for u in1:3] ## Q(u^{[t,0]}): add up policies by their first action (_G, _qπ, _qu)endchoose(_qu) =argmax(_qu) ## û^{[t,0]}: the most probable action (1 bedroom, 2 livingroom, 3 bathroom)
choose (generic function with 1 method)
Changes in the agent code compared to Part 4
The joint belief is an array with one axis per factor, of size \(3 \times 4 \times 2 \times 2 \times 2\). Two helpers connect it to the tensors: marginal adds it up to the factors a tensor depends on, and spread broadcasts an array over a few factors across the full joint shape.
The prediction is computed one factor at a time (transition, predict), without building the 96 × 96 transition matrix. A factor is updated before the factors it depends on, so that it uses their previous values: the dirt factors before the Room factor. An @assert checks this order against LdepsB.
Perception multiplies the slices of the likelihood tensors for the received observations, each spread out over the factors it does not depend on, with the prediction, and normalizes. At \(t = 0\), the prediction is the uniform joint prior _D, the product of the per-factor priors in LD.
The expected free energy rolls the joint belief forward under each policy, and computes the expected observation and the ambiguity of each sensor from the marginal over the factors that sensor depends on.
Two preference tuples:LC with both preferences, and LC_motion for the Part 4 agent with only the preference for no motion. efe and plan take the preferences as an argument, so the same code serves both agents.
The planning horizon is \(H = 3\), giving 27 policies.
Notes on the agent code
Names: the code mirrors the agent’s math.
LA, LB, LC and LD are tuples, upright bold \(\mathbf{A}\), \(\mathbf{B}\), \(\mathbf{C}\) and \(\mathbf{D}\), built with tuple0 as in the environment. By the tear-off mnemonic, LA[m] is the italic-bold likelihood tensor \(\boldsymbol{A}^{[m]}\), LB[n] the transition tensor \(\boldsymbol{B}^{[n]}\), LC[m] the preference vector \(\boldsymbol{C}^{[m]}\), and LD[n] the initial prior \(\boldsymbol{D}^{[n]}\) of factor \(n\). Since the agent is given its model, LA and LB are simply the environment’s LAˣ and LBˣ, without the asterisk, and their dependencies are the environment’s LdepsA and LdepsB.
LC_motion is a second preference tuple, with a flat preference for Dirt: the preferences of the Part 4 agent, used for the comparison in 4.6.
_sₜ and _sₜ₋₁ are the agent’s joint beliefs \(\boldsymbol{s}^{(t)}\) and \(\boldsymbol{s}^{(t-1)}\): single italic-bold arrays, of size \(3 \times 4 \times 2 \times 2 \times 2\) (sizes), so that _sₜ[j₀+1, j₁+1, j₂+1, j₃+1, j₄+1] is the probability of one combination \(\mathbf{j}\). Inside efe, _s is the predicted joint belief \(\boldsymbol{s}_{\pi^{[p]}}^{(\tau)}\). _D is the uniform joint prior, the product of the per-factor priors in LD.
marginals(_sₜ) returns the tuple of marginals \(\big(\boldsymbol{s}^{[t,0]}, \ldots, \boldsymbol{s}^{[t,4]}\big)\), i.e. Lsₜ, which is what we record and evaluate.
_V, _G, _qπ and _qu are single italic-bold arrays: the policy matrix \(\boldsymbol{V}\), the expected free energies \(\boldsymbol{G}\), the posterior over policies \(Q\big(\pi^{(t)}\big)\) and the distribution over the next action \(Q\big(u^{[t,0]}\big)\).
LH holds, per modality, the entropy of each column of \(\boldsymbol{A}^{[m]}\): how uncertain the sensor’s reading is for each combination of the factors it depends on. It is an array over those factors, e.g. \(3 \times 4\) for Motion, and it is computed once, since the tensors do not change.
Helpers: connecting the joint belief to the tensors. Each tensor depends on only a few factors, while the joint belief has an axis for every factor. Two helpers translate between them, using the dependency lists:
marginal(_s, deps) adds up the joint belief over all factors not in deps, leaving the belief over exactly the factors a tensor depends on, with its axes in the order of deps, i.e. the order of the tensor’s axes. It is used whenever a tensor is applied to a belief.
spread(arr, deps) does the opposite: it gives an array over the factors in deps the shape of the joint belief, with size 1 on all other axes, so that Julia’s broadcasting repeats it along them. It is used to multiply the likelihoods into the joint belief.
Structure: one function per step.
transition and predict compute the prediction \(\sum_{\mathbf{i}} P_{\mathbf{B}}(\mathbf{j} \mid \mathbf{i}, u)\,\boldsymbol{s}[\mathbf{i}]\) one factor at a time: transition moves the probability of each combination along the axis of factor \(n\), using the column of \(\boldsymbol{B}^{[n]}\) for the values of the factors it depends on, and predict does this for all five factors. Each factor must use the previous values of its dependencies, so a factor is updated before the factors it depends on: the dirt factors, which depend on the Roomba’s previous room, before the Room factor. This is the order in update_order, and the @assert checks it against LdepsB. In the wrong order, the dirt would be cleaned in the room the Roomba is moving to, rather than the room it was in.
perceive is step 1, the perception: it predicts the new joint state with the previous action, multiplies by the likelihoods of the three received observations, each the slice of \(\boldsymbol{A}^{[m]}\) for the received observation, spread out over the joint state, and normalizes. At \(t = 0\), it starts from the joint prior _D instead of a prediction. This is exact Bayesian filtering over the 96 joint states.
efe is step 2, the evaluation of one policy: it rolls the joint belief forward through the policy’s actions and sums, over the steps and the three sensors, the risk (how far the expected observations are from the preferences) and the ambiguity (how uncertain the sensors’ readings are expected to be). Both are computed from the marginal over the factors the sensor depends on. This follows line 16 of Algorithm 18.
plan and choose are step 3, the decision: plan turns the scores into the posterior over policies and the distribution over the next action, and choose picks the most probable action. This follows lines 19 and 20 of Algorithm 18. efe and plan take the preferences as an argument, so the same code serves both the agent with two preferences (LC) and the Part 4 agent (LC_motion).
Indexing. Tuples are 0-based as everywhere else, e.g. LA[0] to LA[2] and LB[0] to LB[4], and so are the factor numbers in LdepsA and LdepsB. Inside the italic-bold arrays, Julia’s ordinary 1-based indices are used: factor \(n\) is axis \(n + 1\) of the joint belief, which is why the code reads i[d+1] for the value of factor d, and the values of each factor, the three actions, and the planning steps are numbered from 1. The math’s step \(\tau = 0, \ldots, H-1\) is column \(\tau + 1\) of _V.
This part does not use RxInfer: with known tensors and only 96 joint states, all three steps take a few dozen lines of Julia, which keeps every quantity of Algorithm 18 visible.
Next, we define the agent and the time-stepping procedure.
## LC: the agent's preferences, e.g. LC (both goals) or LC_motion (the Part 4 agent)functioncreate_agent(; LC) _sₜ₋₁ =nothing## s^{(t-1)}: the previous joint belief (none before t = 0) ûₜ₋₁ =nothing## û^{[t-1,0]}: the previous action, as an index 1..3 (none before t = 0) _sₜ =nothing## s^{(t)}: the current joint belief, a 3×4×2×2×2 array ûₜ =nothing## û^{[t,0]}: the current action _G = _qπ = _qu =nothing## the current plan: G, Q(π^{(t)}), Q(u^{[t,0]})## step 1 and 2: update the joint belief with the new observations, then score all policies compute = (Lôₜ) ->begin _sₜ =perceive(_sₜ₋₁, Lôₜ, ûₜ₋₁) _G, _qπ, _qu =plan(_sₜ, LC)nothingend## step 3: choose the next action, returned as a one-hot action tuple û^{(t)} act = () ->begin ûₜ =choose(_qu)tuple0(Float64.(1:3.== ûₜ))end## predicted beliefs under policy p, τ = 0, ..., H-1 (for inspection):## each entry is the tuple of marginals of the predicted joint belief s_π^{(τ)} future = (p) ->begin _s = _sₜ Lsπː =Any[]for τ in1:H _s =predict(_s, _V[p, τ]) ## s_π^{(τ)}: predicted joint beliefpush!(Lsπː, marginals(_s))endoffset0(Lsπː)end## move one tick forward: the current belief and action become the previous ones slide = () ->begin _sₜ₋₁ = _sₜ ûₜ₋₁ = ûₜnothingend belief = () ->marginals(_sₜ) ## (s^{[t,0]}, ..., s^{[t,4]}): the marginals, for recording scores = () -> (_G, _qπ, _qu) ## the current plan, for inspectionreturn (compute, act, future, slide, belief, scores)end
create_agent (generic function with 1 method)
Changes in create_agent compared to Part 4
The preferences are an argument,create_agent(; LC), so the same code builds both the agent with two preferences and the Part 4 agent with only the preference for no motion.
The agent keeps the joint belief over the 96 combinations, _sₜ, rather than a tuple with a single belief. The joint belief is only used inside the agent.
belief returns the marginals, one per factor, which is what is recorded and evaluated.
future predicts joint beliefs with the factor-by-factor prediction, and returns their marginals at each step, so the predicted room, person location and dirt under a policy can be inspected.
4.6 Agent Evaluation
Now it’s time to hand the Roomba over to the agent. Does it manage to clean the apartment while staying out of people’s way?
As in Part 4, there is no separate inference step on finished data. The agent and the environment run together, tick by tick, for \(T + 1 = 600\) ticks. At each tick:
the environment reports the observations \(\hat{\mathbf{o}}^{(t)}\) from the three sensors: texture, motion and dirt,
the agent updates its joint belief \(\boldsymbol{s}^{(t)}\) about the room, the person and the dirt in each room (compute, step 1),
the agent scores all 27 policies by their expected free energy and chooses the next destination \(\hat{\boldsymbol{u}}^{[t,0]}\) (compute, step 2, and act),
the environment carries out the action, moving the Roomba, moving the person, and letting dirt accumulate or be cleaned (execute), and
the agent moves one tick forward (slide).
During the run, we record the true states \(\hat{\mathbf{s}}^{\ast(0:T)}\), the observations \(\hat{\mathbf{o}}^{(0:T)}\), the actions \(\hat{\mathbf{u}}^{(0:T-1)}\) and the agent’s marginal beliefs \(\boldsymbol{s}^{[0:T,\bullet]}\), one per factor at each tick. From these, we compute the measures from 4.3.5: the disturbance (how often the Roomba ended up in the same room as the person), the cleanliness (how often each room was dirty), and the tracking accuracies (how well the agent knew where the Roomba was, where the person was, and which rooms were dirty).
To see what the cleaning goal adds, we run two agents, as announced in 4.3.5:
the Part 4 agent, with only the preference for no motion (LC_motion), and
the agent of this part, with both preferences, no motion and dirt picked up (LC).
Together with the random Roomba from 4.3, this gives three Roombas to compare. Each run uses the same environment with the same random seed, so the comparison is as fair as possible: only the way the actions are chosen differs. We expect the agent with both preferences to keep the rooms, and in particular the living room, cleaner than the Part 4 agent, at the cost of some extra disturbance.
Since the agents are given their model of the world, there is nothing to learn and no free energy to minimize over iterations: their beliefs follow from exact Bayesian filtering over the 96 joint states, and their choices from the expected free energy of their policies. Let’s start the runs!
Changes in the evaluation setup compared to Part 4
Each tick now involves three sensors (texture, motion, dirt), a joint belief over room, person and dirt, 27 policies, and an environment step that moves the Roomba and the person and updates the dirt.
The recorded beliefs are the agent’s marginals, one per factor at each tick.
The measures are disturbance, cleanliness and tracking per factor; coverage is no longer a main measure.
Two agents are run in the same environment with the same seed: the Part 4 agent with only the preference for no motion, and the agent with both preferences. Together with the random Roomba, this gives the three-way comparison from 4.3.5.
4.6.1 Evaluate with simulated data
4.6.1.1 Naive approach {N/A}
4.6.1.2 Active inference approach
### Simulation parameters## Time steps t = 0, ..., T (600 ticks)T =599## Initial state: the Roomba in the bedroom, the person in the living room, all rooms clean## factor 0 (Room): 1 bedroom 2 livingroom 3 bathroom## factor 1 (Person): 1 bedroom 2 livingroom 3 bathroom 4 away## factors 2–4 (Dirt in bedroom, livingroom, bathroom): 1 clean 2 dirtyLŝˣ₀ =tuple0([1.0, 0.0, 0.0], ## ŝ^{*[0,0]}: Roomba in the bedroom [0.0, 1.0, 0.0, 0.0], ## ŝ^{*[0,1]}: person in the living room [1.0, 0.0], ## ŝ^{*[0,2]}: bedroom clean [1.0, 0.0], ## ŝ^{*[0,3]}: living room clean [1.0, 0.0]) ## ŝ^{*[0,4]}: bathroom clean → ŝ^{*(0)}
5-element OffsetArray(::Vector{Vector{Float64}}, 0:4) with eltype Vector{Float64} with indices 0:4:
[1.0, 0.0, 0.0]
[0.0, 1.0, 0.0, 0.0]
[1.0, 0.0]
[1.0, 0.0]
[1.0, 0.0]
Changes in the simulation parameters compared to Part 4
The initial state has one entry per factor: the Roomba in the bedroom, the person in the living room, and all rooms clean, the same state as for the random Roomba, so all three Roombas start alike.
### OVERRIDES
Let’s take a look at our results. We compare three Roombas: the random Roomba from 4.3, the Part 4 agent with only the preference for no motion, and the agent with both preferences. If the preferences work, both agents should spend clearly less of their time in the person’s room than the random Roomba (a lower disturbance). The two agents should differ in cleanliness: the Part 4 agent is expected to neglect the living room, as in Part 4, so that room stays dirty much of the time, while the agent with both preferences should keep all rooms, and the living room in particular, cleaner. The price is likely some extra disturbance, since the living room is also where the person spends most time. Behind this behavior, the agents should keep good track of the state (high tracking accuracies), but not equally well for every factor: the Roomba’s room should be tracked well, the dirt in the Roomba’s own room too, but the person’s location and the dirt in other rooms less so, since the agent only senses them when it is in the same room.
It is also worth looking at how the agent with both preferences behaves, not just how well. We expect it to:
leave a room soon after detecting motion there, as in Part 4,
stay in a dirty room while it picks up dirt, and move on once the dirt sensor goes quiet,
visit the living room in short bursts, when it expects it to be quiet, rather than avoiding it altogether, and
return to rooms it has not visited for a while, since it expects dirt to have built up there, and checking also reduces its uncertainty about them.
Let’s see if it worked.
Changes in the expected results compared to Part 4
Three Roombas are compared: the random Roomba, the Part 4 agent with only the preference for no motion, and the agent with both preferences.
Cleanliness replaces coverage: the Part 4 agent is expected to neglect the living room, and the agent with both preferences to keep it clean, at the cost of some extra disturbance.
Tracking is expected to differ per factor: good for the Roomba’s room and the dirt in its own room, weaker for the person and the dirt in the other rooms.
The expected behavior now includes staying in a dirty room until it is clean, short visits to the living room when it is expected to be quiet, and returning to rooms not visited for a while.
## one run of an agent with preferences LC, in a fresh environment with the same seed as the random Roomba (4.3)functionrun_agent(LC; seed=42) (execute_ai, observe_ai, state_ai) =create_envir(; LAˣ, LBˣ, LdepsA, LdepsB, Lŝˣ₀, seed) agent =create_agent(; LC) (compute_ai, act_ai, future_ai, slide_ai, belief_ai, scores_ai) = agent Lŝˣː =offset0(Vector{Any}(undef, T+1)) ## ŝ^{*(0:T)}: true states, each a tuple (Room, Person, Dirt ×3) Lôː =offset0(Vector{Any}(undef, T+1)) ## ô^{(0:T)}: observations, each a tuple (Texture, Motion, Dirt) Lûː =offset0(Vector{Any}(undef, T)) ## û^{(0:T-1)}: chosen actions, each a tuple (Destination) Lsː =offset0(Vector{Any}(undef, T+1)) ## s^{[0:T,•]}: the agent's marginal beliefs, each a tuple of five Lquː =offset0(Vector{Any}(undef, T+1)) ## Q(u^{[t,0]}): the agent's action probabilities at each tickfor t =0:T t >0&&execute_ai(Lûː[t-1]) ## 4. the environment carries out the previous action Lŝˣː[t] =state_ai() Lôː[t] =observe_ai() ## 1. the environment reports the observationscompute_ai(Lôː[t]) ## 2.–3. the agent updates its joint belief and scores the policies Lsː[t] =belief_ai() ## the marginals of the joint belief Lquː[t] =scores_ai()[3]if t < T Lûː[t] =act_ai() ## 3. the agent chooses the next destinationslide_ai() ## 5. the agent moves one tick forwardendendreturn (; Lŝˣː, Lôː, Lûː, Lsː, Lquː, agent)end@time run_motion =run_agent(LC_motion) ## the Part 4 agent: only the preference for no motion@time run_both =run_agent(LC) ## the agent of this part: no motion and dirt picked up## from here on, the recorded data and the agent functions are those of the agent with both preferences(; Lŝˣː, Lôː, Lûː, Lsː, Lquː) = run_both(compute_ai, act_ai, future_ai, slide_ai, belief_ai, scores_ai) = run_both.agent
59.640369 seconds (690.16 M allocations: 19.473 GiB, 6.62% gc time, 8.89% compilation time)
53.818729 seconds (683.17 M allocations: 19.129 GiB, 7.16% gc time)
Two runs: the loop is wrapped in a function, run_agent, which runs one agent with given preferences in a fresh environment. It is called once for the Part 4 agent (LC_motion) and once for the agent with both preferences (LC), both with the same seed as the random Roomba.
The environment receives the dependency tuplesLdepsA and LdepsB, and the agent its preferences.
The recorded data holds five-factor true states, Texture, Motion and Dirt observations, and the agent’s five marginal beliefs at each tick.
The familiar names refer to the agent with both preferences, so the inspection cells that follow work unchanged; the Part 4 agent’s data is kept in run_motion for the comparison.
Each run is timed, since the larger model makes it slower than in Part 4.
## t, Roomba's room, person's location, dirt in bedroom, livingroom, bathroom: the true states of the agent's run[(t, room_labels[argmax(Lŝˣː[t][0])], person_labels[argmax(Lŝˣː[t][1])], (dirt_labels[argmax(Lŝˣː[t][n])] for n in2:4)...) for t in0:T]
One listing per factor of interest: the recorded beliefs are now marginals, one per factor, so the belief about the Roomba’s room, the person’s location and the dirt in the living room are listed separately, instead of a single belief over six joint states.
## t, and the agent's most probable value of each factor: room, person, dirt in bedroom, livingroom, bathroom[(t, room_labels[argmax(Lsː[t][0])], person_labels[argmax(Lsː[t][1])], (dirt_labels[argmax(Lsː[t][n])] for n in2:4)...) for t in0:T]
Changes in the agent’s named state listing compared to Part 4
One name per factor: instead of the most probable of six joint states, the listing shows, for each factor, the most probable value under the agent’s marginal belief, in the same format as the true-state table.
## t, true (room, person, dirt here), agent's most probable (room, person, dirt here), destination chosen (none at t = T)## "dirt here" is the dirt factor of the room the Roomba is actually in: factor r + 1 for room rhere(t) =argmax(Lŝˣː[t][0]) +1[(t, (room_labels[argmax(Lŝˣː[t][0])], person_labels[argmax(Lŝˣː[t][1])], dirt_labels[argmax(Lŝˣː[t][here(t)])]), (room_labels[argmax(Lsː[t][0])], person_labels[argmax(Lsː[t][1])], dirt_labels[argmax(Lsː[t][here(t)])]), t < T ? action_labels[argmax(Lûː[t][0])] :"-") for t in0:T]
Changes in the combined listing compared to Part 4
True and believed states are compact triples: the Roomba’s room, the person’s location, and the dirt in the room the Roomba is actually in, instead of the six joint-state names. The dirt in the other rooms is left out to keep the lines readable; it can be found in the per-factor listings.
[Lô[0] for Lô in Lôː] ## ô^{[t,0]} for t = 0, ..., T: the received Texture observations (agent with both preferences)
Changes in the recorded beliefs compared to Part 4
Each entry is a tuple of five marginals, one per factor, instead of a tuple with a single belief over six joint states. The agent’s joint belief is not recorded; only its marginals are.
_G, _qπ, _qu =scores_ai() ## the plan of the agent with both preferences at the last tick, t = T## the 27 policies, from lowest expected free energy (most probable) to highest[(join(action_labels[_V[p, :]], " → "), round(_G[p], digits=2), round(_qπ[p], digits=3)) for p insortperm(_G)]
The 27 policies are sorted from lowest to highest expected free energy, so the most probable ones come first.
Finally, we can check whether the agent achieved what it was built for: keeping the Roomba out of people’s way while still cleaning the apartment. From the true states \(\hat{\mathbf{s}}^{\ast(t)}\), we compute, for each of the three Roombas, the disturbance, the fraction of ticks the Roomba spent in the same room as the person, and the cleanliness, the fraction of ticks each room was dirty. The coverage, the fraction of ticks spent in each room, is shown as background, since it helps explain the cleanliness. The comparison between the Part 4 agent and the agent with both preferences shows what the cleaning goal achieves, and what it costs in disturbance.
Behind this behavior lies the agent’s tracking. From its marginal beliefs \(\boldsymbol{s}^{[t,n]}\), we derive its most probable value of each factor at each tick, and compare it with the true value. We expect the tracking to differ per factor:
the Roomba’s room should be tracked about as well as in Part 4, since the camera still does most of that work, and the agent knows where it was heading.
the dirt in the Roomba’s room should be tracked well too, since the dirt sensor reports on it at every tick. Here a natural question is whether the agent can do better than the dirt sensor alone: dirt persists until it is cleaned, so a single false reading could be overruled by what the agent expects.
whether the person is in the Roomba’s room is the question behind the disturbance. In Part 4, the agent’s belief simply followed the latest motion reading, because occupancy was drawn fresh each time the Roomba entered a room. Now the person moves independently of the Roomba and tends to stay put, so the agent’s earlier observations carry over, and it may do better than the motion sensor alone.
the person’s location and the dirt in the other rooms are only predicted, not sensed, while the Roomba is elsewhere, so we expect them to be tracked less well.
Tracking matters for both goals: when the agent thinks the person is elsewhere while they are not, it stays where it causes disturbance, and when it thinks a room is clean while it is not, it stays away from where it is needed.
There is no free energy curve to check this time: the agent’s beliefs come from exact Bayesian filtering at each tick, not from an iterative approximation.
Changes in the evaluation measures compared to Part 4
Disturbance and cleanliness are computed for all three Roombas. Coverage is kept as background only.
Tracking is evaluated per factor from the agent’s marginal beliefs: well for the Roomba’s room and the dirt in its room, less well for the person’s location and the dirt in other rooms, which are only predicted while the Roomba is elsewhere.
Two questions about overruling a sensor: whether the agent can correct false dirt readings, since dirt persists until cleaned, and whether it can do better than the motion sensor, since the person now persists independently of the Roomba, unlike the occupancy of Part 4.
Lqsː = Lsː ## the agent's marginal beliefs s^{[t,n]}, t = 0, ..., T (0-based; recorded during the run)## ---------- behavior: disturbance, cleanliness and coverage (4.3.5), for one run ----------functionmeasures(Lŝˣː) rooms = [argmax(Lŝˣ[0]) for Lŝˣ inparent(Lŝˣː)] persons = [argmax(Lŝˣ[1]) for Lŝˣ inparent(Lŝˣː)] dirty = [[argmax(Lŝˣ[n]) ==2 for Lŝˣ inparent(Lŝˣː)] for n in2:4] ## dirt in bedroom, livingroom, bathroom (; disturbance =mean(rooms .== persons), ## J_disturb dirtiness = [mean(dirty[r]) for r in1:3], ## J_dirty(r) coverage = [mean(rooms .== r) for r in1:3]) ## backgroundendm_motion =measures(run_motion.Lŝˣː) ## the Part 4 agentm_both =measures(Lŝˣː) ## the agent with both preferencespct(x) =round.(100.* x, digits=1)println(" disturbance dirty (bedroom, livingroom, bathroom) average dirty coverage")println("random Roomba: $(pct(disturbance_random))% $(pct(dirtiness_random))% $(pct(mean(dirtiness_random)))% $(pct(coverage_random))%")println("Part 4 agent: $(pct(m_motion.disturbance))% $(pct(m_motion.dirtiness))% $(pct(mean(m_motion.dirtiness)))% $(pct(m_motion.coverage))%")println("both preferences: $(pct(m_both.disturbance))% $(pct(m_both.dirtiness))% $(pct(mean(m_both.dirtiness)))% $(pct(m_both.coverage))%")## ---------- tracking (agent with both preferences) ----------factor_names = ["Roomba's room", "person's location", "dirt in bedroom", "dirt in living room", "dirt in bathroom"]accuracy = [mean(argmax(Lqsː[t][n]) ==argmax(Lŝˣː[t][n]) for t in0:T) for n in0:4] ## J_track(n)for n in0:4println("tracking, $(factor_names[n+1]): $(pct(accuracy[n+1]))%")endrooms = [argmax(Lŝˣ[0]) for Lŝˣ inparent(Lŝˣː)] ## true room: 1 bedroom, 2 livingroom, 3 bathroompersons = [argmax(Lŝˣ[1]) for Lŝˣ inparent(Lŝˣː)] ## true person locationinferred_rooms = [argmax(Ls[0]) for Ls inparent(Lqsː)] ## most probable room## the person in the Roomba's (true) room: the person marginal's entry for that roomperson_here_true = persons .== roomsp_person_here = [Lqsː[t][1][rooms[t+1]] for t in0:T]inferred_person_here = p_person_here .>0.5accuracy_person_here =mean(inferred_person_here .== person_here_true)## the dirt in the Roomba's (true) room: the dirt factor r + 1 for room rdirty_here_true = [argmax(Lŝˣː[t][rooms[t+1]+1]) ==2 for t in0:T]p_dirty_here = [Lqsː[t][rooms[t+1]+1][2] for t in0:T]inferred_dirty_here = p_dirty_here .>0.5accuracy_dirty_here =mean(inferred_dirty_here .== dirty_here_true)## can the agent overrule its sensors?textures = [argmax(Lô[0]) for Lô inparent(Lôː)]camera_wrong = textures .!= roomsaccuracy_cam_wrong =mean(inferred_rooms[camera_wrong] .== rooms[camera_wrong])println("camera misread the floor at $(count(camera_wrong)) of $(T+1) ticks; ","agent still inferred the right room at $(pct(accuracy_cam_wrong))% of those")motion_detected = [argmax(Lô[1]) for Lô inparent(Lôː)] .==2## modality 1: Motionmotion_wrong = motion_detected .!= person_here_trueaccuracy_mot_wrong =mean(inferred_person_here[motion_wrong] .== person_here_true[motion_wrong])println("motion sensor was wrong at $(count(motion_wrong)) of $(T+1) ticks; ","agent still inferred correctly whether the person was in its room at $(pct(accuracy_mot_wrong))% of those")dirt_picked_up = [argmax(Lô[2]) for Lô inparent(Lôː)] .==2## modality 2: Dirtdirt_wrong = dirt_picked_up .!= dirty_here_trueaccuracy_dirt_wrong =mean(inferred_dirty_here[dirt_wrong] .== dirty_here_true[dirt_wrong])println("dirt sensor was wrong at $(count(dirt_wrong)) of $(T+1) ticks; ","agent still inferred correctly whether its room was dirty at $(pct(accuracy_dirt_wrong))% of those")## ---------- plots ----------p1 =scatter(0:T, rooms, title="Room: inferred correctly $(pct(accuracy[1]))% of the time", label="True room", legend=:outerright, yticks=(1:3, ["Bedroom", "Living room", "Bathroom"]), markercolor=:cyan, markersize=6, markerstrokewidth=0, alpha=.5)scatter!(p1, 0:T, inferred_rooms, label="Inferred room", markercolor=:green, markersize=3, markerstrokewidth=0, alpha=.5)p2 =plot(0:T, person_here_true, linetype=:steppost, color=:red, label="Person in the room", title="Person in the Roomba's room: inferred correctly $(pct(accuracy_person_here))% of the time", yticks=([0, 0.5, 1], ["No", "0.5", "Yes"]), ylims=(-0.05, 1.05), legend=:outerright)plot!(p2, 0:T, p_person_here, color=:black, alpha=0.7, label="P(person here)")p3 =plot(0:T, dirty_here_true, linetype=:steppost, color=:brown, label="Room dirty", title="Dirt in the Roomba's room: inferred correctly $(pct(accuracy_dirty_here))% of the time", yticks=([0, 0.5, 1], ["Clean", "0.5", "Dirty"]), ylims=(-0.05, 1.05), legend=:outerright)plot!(p3, 0:T, p_dirty_here, color=:black, alpha=0.7, label="P(dirty here)", xlabel="Time t")x =1:3p4 =bar(x .-0.27, pct(dirtiness_random), bar_width=0.27, label="random Roomba", color=:gray)bar!(p4, x, pct(m_motion.dirtiness), bar_width=0.27, label="Part 4 agent", color=:orange)bar!(p4, x .+0.27, pct(m_both.dirtiness), bar_width=0.27, label="both preferences", color=:green, xticks=(x, ["Bedroom", "Living room", "Bathroom"]), ylabel="% of ticks dirty", legend=:outerright, title="Dirty — disturbance: random $(pct(disturbance_random))%, Part 4 $(pct(m_motion.disturbance))%, both $(pct(m_both.disturbance))%", titlefontsize=10)plot(p1, p2, p3, p4, layout=@layout([a; b; c; d]), size=(900, 1100))
disturbance dirty (bedroom, livingroom, bathroom) average dirty coverage
random Roomba: 15.8% [8.7, 20.7, 3.2]% 10.8% [22.5, 25.8, 51.7]%
Part 4 agent: 8.8% [6.5, 26.2, 2.8]% 11.8% [34.3, 15.3, 50.3]%
both preferences: 11.0% [4.3, 17.8, 2.7]% 8.3% [32.2, 17.8, 50.0]%
tracking, Roomba's room: 94.7%
tracking, person's location: 70.3%
tracking, dirt in bedroom: 97.3%
tracking, dirt in living room: 84.8%
tracking, dirt in bathroom: 98.0%
camera misread the floor at 53 of 600 ticks; agent still inferred the right room at 50.9% of those
motion sensor was wrong at 31 of 600 ticks; agent still inferred correctly whether the person was in its room at 51.6% of those
dirt sensor was wrong at 27 of 600 ticks; agent still inferred correctly whether its room was dirty at 77.8% of those
Changes in the evaluation code compared to Part 4
The behavior measures are computed for all three Roombas with a helper, measures: disturbance (the Roomba in the person’s room), the fraction of ticks each room is dirty, and coverage as background. They are printed side by side.
Tracking is computed for every factor from the agent’s marginal beliefs, replacing the derivation of room and occupancy from six joint states.
The person and the dirt in the Roomba’s room are evaluated at the Roomba’s true room, which only needs a single entry of the person or dirt marginal, so no information from the joint belief is needed.
Three sensors are checked for errors the agent overrules: the camera, the motion sensor, and the new dirt sensor.
The plots show the room, the probability that the person is in the Roomba’s room, the probability that its room is dirty, and the cleanliness of each room for the three Roombas, with their disturbance in the title.
4.7 Conclusion and outlook
In this part, the agent got a second goal. Besides staying out of people’s way, as in Part 4, it now had to clean. To express this, the model of the environment grew from a single state factor to five: the Roomba’s room, the person’s location, and the dirt in each of the three rooms. Four ingredients made this possible:
Several state factors with dependencies. Each factor’s transitions, and each sensor’s readings, depend on only a few factors: a room’s dirt on itself and on where the Roomba was, the motion sensor on where the Roomba and the person are. These dependencies, \(\mathcal{D}_A^{[m]}\) and \(\mathcal{D}_B^{[n]}\), are the core of multi-factor modeling. They keep every tensor small: the 96 joint states never had to be written out.
A joint belief, with marginals for evaluation. Although the model factorizes, the agent’s beliefs do not: a motion reading says something about the Roomba and the person together. The agent therefore filtered exactly over the 96 joint states, stored as an array with one axis per factor, and the beliefs about single factors followed by adding up.
Two preferences. The agent preferred to observe no motion and dirt picked up. The latter, rather than a clean floor, is what draws the agent to rooms where its presence makes a difference.
A longer horizon. With \(H = 3\), the agent could see the pay-off of going to the living room, two moves away from the bedroom.
As in Part 4, all of this was written directly in Julia, with code that is generic over factors and sensors: the dependency lists tell each function which axes of the joint belief a tensor acts on.
The cleaning goal works, more modestly than expected. The three Roombas compare as follows:
disturbance
dirty: bedroom
dirty: living room
dirty: bathroom
dirty: average
random Roomba
15.8%
8.7%
20.7%
3.2%
10.8%
Part 4 agent (no motion only)
8.8%
6.5%
26.2%
2.8%
11.8%
both preferences
11.0%
4.3%
17.8%
2.7%
8.3%
The Part 4 agent avoided the person best, but it confirms the problem that motivated this part: it left the living room dirty 26.2% of the time, more than the random Roomba. Adding the cleaning goal brought this down to 17.8%, and the average over the three rooms from 11.8% to 8.3%, the cleanest of the three Roombas. The price was 2.2 percentage points of extra disturbance (11.0% against 8.8%), still about 30% less than the random Roomba.
How it achieved this is visible in the coverage. The agent with both preferences spent only slightly more time in the living room than the Part 4 agent (17.8% against 15.3%), yet left it dirty much less often. It did not clean by spending more time in the living room, but by going there at the right moments: when it expected dirt to have built up. The dirt plot shows the same pattern: periods with a dirty room under the Roomba are short, because the agent stays while it picks up dirt and moves on once the room is clean.
Still, the living room remains the dirtiest room by far, and the improvement over the random Roomba is modest (17.8% against 20.7%). The two preferences pull in opposite directions there, and with the motion preference weighing more than the dirt preference, avoiding the person often wins. While the Roomba is elsewhere, the agent cannot see the person, so its belief drifts towards the living room, where the person usually is: the agent expects motion there more often than it would find it.
Tracking: the person now carries over. The agent tracked the Roomba’s room correctly 94.7% of the time, and whether the person was in the Roomba’s room 96.0% of the time. The latter is a real improvement on Part 4. There, the agent was wrong at every one of the 38 ticks at which the motion sensor was wrong: its belief simply followed the latest reading. Here, the motion sensor was wrong at 31 ticks, and the agent still got it right at 51.6% of those. The reason is in the model: the person now moves independently of the Roomba and tends to stay put, so earlier observations carry over, whereas in Part 4 occupancy was drawn fresh each time the Roomba entered a room. The dirt sensor was wrong at 27 ticks, and the agent overruled it at 77.8% of those, since dirt persists until it is cleaned and the agent knows where it has recently cleaned. The camera, finally, misread the floor at 53 ticks, at 50.9% of which the agent still inferred the right room.
The tracking of the other factors needs more care. The person’s location was tracked correctly only 70.3% of the time: the agent senses the person only when they are in the Roomba’s room, and otherwise has to rely on its predictions. The dirt in the bedroom (97.3%) and the bathroom (98.0%) seems well tracked, but these rooms were rarely dirty, so always guessing clean would already have been right 95.7% and 97.3% of the time. The living room, which gets dirty fastest, is the real test, and there the agent was right 84.8% of the time, against 82.2% for always guessing clean. The agent knows the dirt where it is; elsewhere, it knows little more than the base rates.
These results come from a single run of 600 ticks per Roomba, so the exact percentages should not be over-interpreted. The differences in disturbance between the two agents, in particular, are small. The patterns, though, follow from the model: a Roomba that avoids people neglects the busiest room, a cleaning goal brings it back at well-chosen moments, and the person’s location persists in the agent’s belief now that it is a factor of its own.
Possible next steps:
Tuning the balance. The relative strength of \(\boldsymbol{C}^{[1]}\) and \(\boldsymbol{C}^{[2]}\) sets the trade-off between disturbance and cleanliness. Running the agents for a range of preference strengths, averaged over several random seeds, would trace out this trade-off, and show how much cleaner the living room can get for each extra percent of disturbance. A precision parameter for how decisively the agent picks the best policy is worth including.
Factorized beliefs. Exact filtering over the joint states works for 96 combinations, but their number grows multiplicatively with every factor. Approximating the belief by a product of per-factor beliefs, as is common in larger active inference models, makes much larger models feasible, at the cost of losing exactly the relations between factors that the joint belief captures here.
Valuing information. The agent’s weak knowledge of the person’s location and the living room’s dirt suggests that checking on rooms could be worth more than it currently is. A longer planning horizon, or a stronger weight on reducing uncertainty, may make the agent value this.
Learning while acting. Here, the agent was given its tensors. Letting it learn them while it acts, e.g. how fast each room gets dirty, or where the person tends to be, brings in the trade-off between exploring to learn and exploiting what it already knows.
Richer environments. More exogenous forces, such as doors that are closed at random, a person whose routine depends on the time of day, or more than one person, bring the model closer to the scheduling project this series is building up to.
Planning as inference with RxInfer. For larger models, where neither the joint states nor the policies can be enumerated, formulating planning as inference in RxInfer is the natural next tool.
With this part, the Roomba both cleans and stays out of the way. The balance between the two can still be improved, and the model is now built from factors that can be added one at a time.