损失函数 Loss Function
交叉熵损失函数
假设有一个图像分类任务,我们希望根据图片中动物的轮廓、颜色等特征来预测动物的类别。现在有3种可预测类别:猫、狗、猪,并且当前我们已经获得了两个预测模型。
模型1:
| 预测 | 真实 | 是否正确 |
|---|---|---|
| 0.3 0.3 0.4 | 0 0 1 (猪) | 正确 |
| 0.3 0.4 0.3 | 0 1 0 (狗) | 正确 |
| 0.1 0.2 0.7 | 1 0 0 (猫) | 错误 |
模型2:
| 预测 | 真实 | 是否正确 |
|---|---|---|
| 0.1 0.2 0.7 | 0 0 1 (猪) | 正确 |
| 0.1 0.7 0.2 | 0 1 0 (狗) | 正确 |
| 0.3 0.4 0.3 | 1 0 0 (猫) | 错误 |
从上述结果中我们可以看出,模型2对于样本1和样本2的判断更为准确,同时对样本3的判断也没有像样本1那样错得离谱。
1 Classification Error
最为直接的一种损失函数可以被定义为:
$$\text{classification error} = \frac{\text { count of error items }}{\text { count of all items }}$$
- 模型1:$\text{classification error} = \frac{1}{3}$
- 模型2:$\text{classification error} = \frac{1}{3}$
显然,$\text{classification error}$并不能很好地体现模型1和2之间的差距。
2 Mean Squared Error
均方误差的定义为:
$$\text{MSE}=\frac{1}{n} \sum_{i}^{n}\left(\hat{y}{i}-y{i}\right)^{2}$$
- 模型1:
$$\begin{array}{l}\text { sample } 1 \text { loss }=(0.3-0)^{2}+(0.3-0)^{2}+(0.4-1)^{2}=0.54 \ \text { sample } 2 \text { loss }=(0.3-0)^{2}+(0.4-1)^{2}+(0.3-0)^{2}=0.54 \ \text { sample } 3 \text { loss }=(0.1-1)^{2}+(0.2-0)^{2}+(0.7-0)^{2}=1.34\end{array}$$
对所有loss求平均:
$$\text{MSE}=\frac{0.54+0.54+1.34}{3}=0.81$$ - 模型2:
$$\begin{array}{l}\text { sample } 1 \text { loss }=(0.1-0)^{2}+(0.2-0)^{2}+(0.7-1)^{2}=0.14 \ \text { sample } 2 \text { loss }=(0.1-0)^{2}+(0.7-1)^{2}+(0.2-0)^{2}=0.14 \ \text { sample } 3 \text { loss }=(0.3-1)^{2}+(0.4-0)^{2}+(0.3-0)^{2}=0.74\end{array}$$
对所有loss求平均:
$$\text{MSE}=\frac{0.14+0.14+0.74}{3}=0.34$$
3 Cross Entropy Loss Function
香农关于熵的定义:无损编码事件信息的最小平均编码长度。
$$\text{Entropy} = - \sum_{i}P(i) \log_2P(i)$$
罕见信息比常见信息有着更多的信息量,因为罕见信息能够排除更多其他的可能性,传递一个更为确切的信息。以编码天气信息为例,
| Fine(50%) | Cloudy(25%) | Rainy(12.5%) | Snow(12.5%) |
|---|---|---|---|
| 0 | 10 | 110 | 111 |
| -log(0.5) = 1 | -log(0.25) = 2 | -log(0.125) = 3 | -log(0.125) = 3 |
当接收到Rainy信息时,我们能够减少87.5%的不确定性(Fine,Cloudy,Snow);而如果接收到Fine信息时,我们仅能减少50%的不确定性。
对于连续变量$x$的概率分布$P(x)$,熵的公式可以表示为:
$$\text{Entropy}=-\int P(x) \log {2}{P}(x) d x$$
使用$x \sim P$表示使用概率分布$P$来计算期望,熵的公式可以简写为:
$$H(P)=\text{Entropy}=\mathbb{E}{x \sim P}[-\log P(x)]$$
$$\text{EstimatedEntropy}=\mathbb{E}{x \sim Q}[-\log Q(x)]$$
$$\text{CrossEntropy}=\mathbb{E}{x \sim P}[-\log Q(x)]$$
- P:真实概率分布
- Q:预估概率分布
二分类交叉熵表达式:
$$L=\frac{1}{N} \sum_{i} L_{i}=\frac{1}{N} \sum_{i}-\left[y_{i} \cdot \log \left(p_{i}\right)+\left(1-y_{i}\right) \cdot \log \left(1-p_{i}\right)\right]$$
其中:
- $y_i$ — 表示样本$i$的label,正类为1,负类为0
- $p_i$ — 表示样本$i$预测为正类的概率
多分类交叉熵表达式:
- Title: 损失函数 Loss Function
- Author: Kaleido
- Created at : 2024-03-29 14:44:06
- Updated at : 2024-04-11 00:16:25
- Link: https://redefine.ohevan.com/2024/03/29/ML-损失函数/
- License: This work is licensed under CC BY-NC-SA 4.0.