๐ŸŽฒ Maximum A Posteriori

Machine Learning / Optimization

Given historical data , we want to estimate the parameters . For maximum a posteriori (MAP), we already have some pre-existing hypothesis about our probabilities, termed a prior , and use our new data to update our hypothesis to get the posterior, .. In other words, we find thatโ€™s most likely explained by as well as our prior:

The prior has an influence on the final probability density, and in practice biases our model toward a smoother, simpler distribution.

Example

We'll illustrate this concept with a coin-flip example. Let be a set of coin-flip results, with heads and tails, and let be the probability of the coin landing heads. We'll find that maximizes the probability of given that occurred. For conjugacy, let

follow the same family of distributions as the likelihood and posterior. Then, we can maximize the posterior with ๐Ÿช™ Bayes' Theorem.

Content by William Liang, written in Obsidian.
Thank you to all the educators who made these notes possible.