Learning.IsAlgEnvSeq.hasLaw_feedback_zero_cond
From the authors
Conditionally on the event (O 0, A 0) = p, the first feedback has law env.ฮฝ0 p.
-
๐ : Type u_1m๐ : MeasurableSpace ๐A measurable space is a space equipped with a ฯ-algebra.MeasurableSingletonClass ๐A typeclass mixin forMeasurableSpaces such that each singleton is measurable. -
๐ : Type u_2m๐ : MeasurableSpace ๐MeasurableSingletonClass ๐ -
๐จ : Type u_3m๐จ : MeasurableSpace ๐จ -
ฮฉ : Type u_4mฮฉ : MeasurableSpace ฮฉ
-
O : โ โ ฮฉ โ ๐ -
A : โ โ ฮฉ โ ๐ -
Y : โ โ ฮฉ โ ๐จ -
alg : Algorithm ๐ ๐ ๐จA stochastic, sequential algorithm. -
env : Environment ๐ ๐ ๐จA stochastic environment. -
P : MeasureTheory.Measure ฮฉA measure is defined to be an outer measure that is countably additive on measurable sets, with the additional assumption that the outer measure is the canonical extension of the restricted measure.MeasureTheory.IsFiniteMeasure PA measureฮผis called finite ifฮผ univ < โ. -
p : ๐ ร ๐
-
h : IsAlgEnvSeq O A Y alg env PAn algorithm-environment sequence: a sequence of observations, actions and feedbacks generated by an algorithm interacting with an environment. -
hP : P ((fun ฯ => (O 0 ฯ, A 0 ฯ)) โปยน' {p}) โ 0
ProbabilityTheory.HasLaw (Y 0) (env.ฮฝ0 p) P[|(fun ฯ => (O 0 ฯ, A 0 ฯ)) โปยน' {p}]The predicate HasLaw X ฮผ P registers the fact that the random variable X has law ฮผ under the measure P, in other words that P.map X = ฮผ.MeasurableSpace : Type u_6 โ Type u_6A measurable space is a space equipped with a ฯ-algebra.
MeasurableSingletonClass : (ฮฑ : Type u_6) โ [MeasurableSpace ฮฑ] โ PropA typeclass mixin for `MeasurableSpace`s such that each singleton is measurable.
Nat : TypeThe natural numbers, starting at zero. This type is special-cased by both the kernel and the compiler, and overridden with an efficient implementation. Both use a fast arbitrary-precision arithmetic library (usually [GMP](https://gmplib.org/)); at runtime, `Nat` values that are sufficiently small are unboxed.
Learning.Algorithm : (๐ : Type u_5) โ
(๐ : Type u_6) โ
(๐จ : Type u_7) โ [MeasurableSpace ๐] โ [MeasurableSpace ๐] โ [MeasurableSpace ๐จ] โ Type (max (max u_5 u_6) u_7)A stochastic, sequential algorithm. At each round, it sees an observation in `๐`, then takes an action in `๐`, and finally receives feedback in `๐จ`. The action is a random function of the past rounds and the current observation.Go to its page
Learning.Environment : (๐ : Type u_5) โ
(๐ : Type u_6) โ
(๐จ : Type u_7) โ [MeasurableSpace ๐] โ [MeasurableSpace ๐] โ [MeasurableSpace ๐จ] โ Type (max (max u_5 u_6) u_7)A stochastic environment. At each round, an observation is drawn prior to the algorithm taking an action. Then the environment provides feedback based on the observation and the action.Go to its page
MeasureTheory.IsFiniteMeasure : {ฮฑ : Type u_1} โ {m0 : MeasurableSpace ฮฑ} โ MeasureTheory.Measure ฮฑ โ PropA measure `ฮผ` is called finite if `ฮผ univ < โ`.
MeasureTheory.Measure : (ฮฑ : Type u_5) โ [MeasurableSpace ฮฑ] โ Type u_5A measure is defined to be an outer measure that is countably additive on measurable sets, with the additional assumption that the outer measure is the canonical extension of the restricted measure. The measure of a set `s`, denoted `ฮผ s`, is an extended nonnegative real. The real-valued version is written `ฮผ.real s`.
Prod : Type u โ Type v โ Type (max u v)The product type, usually written `ฮฑ ร ฮฒ`. Product types are also called pair or tuple types. Elements of this type are pairs in which the first element is an `ฮฑ` and the second element is a `ฮฒ`. Products nest to the right, so `(x, y, z) : ฮฑ ร ฮฒ ร ฮณ` is equivalent to `(x, (y, z)) : ฮฑ ร (ฮฒ ร ฮณ)`. Conventions for notations in identifiers: * The recommended spelling of `ร` in identifiers is `Prod`.
Learning.IsAlgEnvSeq : {๐ : Type u_1} โ
{๐ : Type u_2} โ
{๐จ : Type u_3} โ
{ฮฉ : Type u_4} โ
{m๐ : MeasurableSpace ๐} โ
{m๐ : MeasurableSpace ๐} โ
{m๐จ : MeasurableSpace ๐จ} โ
{mฮฉ : MeasurableSpace ฮฉ} โ
(โ โ ฮฉ โ ๐) โ
(โ โ ฮฉ โ ๐) โ
(โ โ ฮฉ โ ๐จ) โ
Learning.Algorithm ๐ ๐ ๐จ โโฆAn algorithm-environment sequence: a sequence of observations, actions and feedbacks generated by an algorithm interacting with an environment.Go to its page
Ne : {ฮฑ : Sort u} โ ฮฑ โ ฮฑ โ Prop`a โ b`, or `Ne a b` is defined as `ยฌ (a = b)` or `a = b โ False`, and asserts that `a` and `b` are not equal. Conventions for notations in identifiers: * The recommended spelling of `โ ` in identifiers is `ne`.
Prod.mk : {ฮฑ : Type u} โ {ฮฒ : Type v} โ ฮฑ โ ฮฒ โ ฮฑ ร ฮฒConstructs a pair. This is usually written `(x, y)` instead of `Prod.mk x y`. Conventions for notations in identifiers: * The recommended spelling of `(a, b)` in identifiers is `mk`.
Set.preimage : {ฮฑ : Type u} โ {ฮฒ : Type v} โ (ฮฑ โ ฮฒ) โ Set ฮฒ โ Set ฮฑThe preimage of `s : Set ฮฒ` by `f : ฮฑ โ ฮฒ`, written `f โปยน' s`, is the set of `x : ฮฑ` such that `f x โ s`.
Singleton.singleton : {ฮฑ : outParam (Type u)} โ {ฮฒ : Type v} โ [self : Singleton ฮฑ ฮฒ] โ ฮฑ โ ฮฒ`singleton x` is a collection with the single element `x` (notation: `{x}`).
Conventions for notations in identifiers:
* The recommended spelling of `{x}` in identifiers is `singleton`.ProbabilityTheory.HasLaw : {ฮฉ : Type u_1} โ
{๐ง : Type u_2} โ
{mฮฉ : MeasurableSpace ฮฉ} โ
{m๐ง : MeasurableSpace ๐ง} โ
(ฮฉ โ ๐ง) โ MeasureTheory.Measure ๐ง โ autoParam (MeasureTheory.Measure ฮฉ) ProbabilityTheory.HasLaw._auto_1 โ PropThe predicate `HasLaw X ฮผ P` registers the fact that the random variable `X` has law `ฮผ` under the measure `P`, in other words that `P.map X = ฮผ`. We also require `X` to be `AEMeasurable`, to allow for nice interactions with operations on the codomain of `X`. See for instance `HasLaw.comp`, `IndepFun.hasLaw_mul` and `IndepFun.hasLaw_add`.
Learning.Environment.ฮฝ0 : {๐ : Type u_1} โ
{๐ : Type u_2} โ
{๐จ : Type u_3} โ
{m๐ : MeasurableSpace ๐} โ
{m๐ : MeasurableSpace ๐} โ
{m๐จ : MeasurableSpace ๐จ} โ Learning.Environment ๐ ๐ ๐จ โ ProbabilityTheory.Kernel (๐ ร ๐) ๐จDistribution of the first feedback given the first observation and action: the feedback kernel at time `0` applied to the empty history.Go to its page
ProbabilityTheory.cond : {ฮฉ : Type u_1} โ {m : MeasurableSpace ฮฉ} โ MeasureTheory.Measure ฮฉ โ Set ฮฉ โ MeasureTheory.Measure ฮฉThe conditional probability measure of measure `ฮผ` on set `s` is `ฮผ` restricted to `s` and scaled by the inverse of `ฮผ s` (to make it a probability measure): `(ฮผ s)โปยน โข ฮผ.restrict s`.
Code
lemma IsAlgEnvSeq.hasLaw_feedback_zero_cond [MeasurableSingletonClass ๐]
[MeasurableSingletonClass ๐] (h : IsAlgEnvSeq O A Y alg env P) {p : ๐ ร ๐}
(hP : P ((fun ฯ โฆ (O 0 ฯ, A 0 ฯ)) โปยน' {p}) โ 0) :
HasLaw (Y 0) (env.ฮฝ0 p) P[|(fun ฯ โฆ (O 0 ฯ, A 0 ฯ)) โปยน' {p}]Proof
h.hasCondDistrib_feedback_zero.hasLaw_cond (h.measurable_feedback 0)
(measurableSet_singleton p) (fun a ha โฆ by rw [Set.mem_singleton_iff.1 ha]) hPMeaning last changed in v4.34.0-rc2-76-g565f652 (2026-09-10), the 3th recorded change.
Self-contained, with its dependencies inlined and proofs replaced by sorry: download the raw file ยท open it in the Lean web editor.
Dependency graph
Audit surface: 7 project declarations, 45 external constants
โ Proved: no sorry anywhere in its closure
This is the tool's own reading of one build's recorded axioms, and it is not robust against an author who wants it to pass. Checking meant to be relied on should go through Comparator, which replays the proof through the kernel from an export against an explicit list of permitted axioms.