
Jargon
| Info theory |
ML-land |
| Source |
A data distribution: $p(x)$ |
| Message |
A sample from $p(x)$ |
| Channel |
Optional, almost always a random transformation that yields $p(y |
| Decoder |
Same usage in ML - $p(x |
Intro

In this first half, we study (from a high level) how to approach:
- Measuring information content
- Compressing data
- Perfect communication over imperfect communication channels
Probability Setup
Forward and Backward Probability
Communication Channels
A Toy Example
A First Inference Problem*
Asides
Stochastic Processes
Entropy & Information