Probability for Data Science
eBook  ›  Chapter 8 · Estimation
Chapter 8

Summary

In this chapter, we have discussed the basic principles of parameter estimation. The three building blocks are:

The three building blocks give us several strategies to estimate the parameters:

As discussed in this chapter, no single estimation strategy is universally “better” because one needs to specify the optimality criterion. If the goal is to minimize the mean squared error, then the MMSE estimator is the optimal strategy. If the goal is to maximize the likelihood without assuming any prior knowledge, the ML estimator would be the optimal strategy. It may appear that if we knew the ground truth parameter \(\vtheta^*\) we could simply minimize the distance between the estimated parameter \(\vtheta\) and the true value \(\vtheta^*\) — but “distance” is not a single, canonical notion, and different choices lead to different optimal estimators. For example, if one cares about the mean absolute error (MAE) rather than the mean squared error, the optimal estimator is the median of the posterior distribution instead of its mean. Therefore, it is the end user's responsibility to specify the optimality criterion.

Whenever we consider parameter estimation, we tend to think that it is about estimating the model parameters, such as the mean of a Gaussian PDF. While in many statistics problems this is indeed the case, parameter estimation can be much broader if we link it with regression. Specifically, a regularized linear regression problem can be formulated as a MAP estimation

$$\vtheta^* = \argmin{\vtheta} \;\; \underset{\textcolor{blue}{-\log f_{\mX|\mTheta}(\vx|\vtheta)}}{\underbrace{\|\mX\vtheta - \vy\|^2}} + \underset{\textcolor{blue}{-\log f_{\mTheta}(\vtheta)}}{\underbrace{ \lambda R(\vtheta) }},$$

for some regularization \(R(\vtheta)\), which is also the negative log of the prior. Expressed in this way, we recognize that the MAP estimation can be used to recover signals. For example, we can model \(\mX\) as a linear degradation process of certain imaging systems. Then solving the MAP estimation is equivalent to finding the best signal explaining the degraded observation using the posterior as the criterion. There is rich literature dealing with solving MAP estimation problems similar to these in subjects such as computational imaging, communication systems, remote sensing, radar engineering, and recommendation systems, to name a few.