A model training method for improving the quality of the predicted probability distribution of a semantic segmentation network
By generating masks in a fully convolutional semantic segmentation network and using softmax function to calculate the predicted probability distribution, and training with cross entropy loss function, the problem of low quality of the prediction probability distribution in the full convolutional network is solved, the prediction confidence distinction ability of the model is improved, and the security and robustness of the model are enhanced.
Patent Information
- Application Number
- CN202211086940.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-06
AI Technical Summary
The existing fully convolutional semantic segmentation network has shortcomings in predicting the quality of probability distribution, which leads to the inability to effectively distinguish correct and misclassified samples, especially in applications with high security requirements.
By generating the mask and mask functions that meet the conditions act on the output of the semantic segmentation network of the fully convolutional image, the predicted probability distribution is calculated using the Bernoulli distribution and softmax functions, and the cross entropy loss function is used for model training, ensuring that the end-to-end training process of the model does not introduce additional computational costs.
Without affecting the model segmentation performance, the predicted probability distribution quality of the model is significantly improved, and the robustness and security of the model are improved.
Smart Images

Figure CN115546225B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a model training method for improving the quality of the predicted probability distribution of a semantic segmentation network. Background Art
[0002] The purpose of image semantic segmentation is to assign a semantic class label to each pixel in an image, which belongs to the pixel-level dense classification task. Generally speaking, semantic segmentation is one of the basic tasks that pave the way for achieving comprehensive scene understanding. More and more applications also obtain knowledge from image data, including autonomous driving, human-computer interaction, indoor navigation, image editing, augmented reality, and virtual reality, etc.
[0003] Image semantic segmentation methods can be divided into two categories: one is the traditional method, such as threshold-based segmentation, edge-based segmentation, region-based segmentation, graph theory-based segmentation, energy functional-based segmentation, etc.; the other is the deep learning-based method. In recent years, with the development of deep neural networks, deep learning has shown increasing advantages in the field of computer vision. Deep convolutional networks are particularly effective for image data and can be used to efficiently extract pixel features in images, overcoming the limitation that traditional methods rely heavily on manually selected features and obtaining better segmentation effects.
[0004] Jonathan Long et al. proposed using a fully convolutional network (FCN) for semantic segmentation in the article "Fully Convolutional Networks for Semantic Segmentation", which has greatly promoted the development of deep learning-based semantic segmentation technology in recent years. Various models based on FCN have significantly improved the accuracy of semantic segmentation, but there is a problem of low quality of the predicted probability distribution. Specifically, the model gives extremely high prediction confidence for different samples, resulting in the inability to effectively distinguish misclassified samples through the prediction confidence. In applications with high safety requirements, there are great hidden dangers, seriously hindering the application of the FCN model in actual scenarios. Ideally, the model should give high confidence to correctly classified samples and low confidence to misclassified samples to improve the robustness of the entire system. Therefore, in practical applications, it is necessary to improve the quality of the predicted probability distribution of the model. Summary of the Invention
[0005] In order to improve the quality of the predicted probability distribution of the fully convolutional semantic segmentation network, that is, to give higher confidence to correctly classified samples and lower confidence to misclassified samples, the present invention provides a model training method for improving the quality of the predicted probability distribution of a semantic segmentation network.
[0006] The object of the present invention is achieved only through one of the following technical solutions.
[0007] A model training method for improving the quality of the predicted probability distribution of a semantic segmentation network, comprising the following steps:
[0008] S1. Select any fully convolutional image semantic segmentation network for supervised training, and obtain the output generated by the selected network for the input sample;
[0009] S2. Generate a mask and a mask function that meet the conditions, and apply the mask to the network output obtained in step S1 through the mask function;
[0010] S3. Based on the network output after the mask is applied, use the softmax function to calculate the predicted probability distribution of the input sample, and use the cross-entropy loss function to supervise the model training until convergence.
[0011] Further, in step S1, the output of the last layer of the selected fully convolutional image semantic segmentation network is used as the output of the entire fully convolutional image semantic segmentation network.
[0012] Further, step S2 includes the following steps:
[0013] S2.1. Generate a mask M using the Bernoulli distribution, K is the number of semantic segmentation categories in the selected fully convolutional image semantic segmentation network, and k is the category index of the input pixel sample, specifically as follows:
[0014]
[0015] where, represents the Bernoulli distribution, δ is an adjustable hyperparameter, and mk represents the mask acting on the prediction score of the k-th category;
[0016] S2.2. Define the mask function Apply the mask M to the output L of the selected fully convolutional image semantic segmentation network through the mask function specifically as follows:
[0017]
[0018] where, l k represents the prediction score of the model for the input sample belonging to category k, L′ is the output after masking, l′ k represents the prediction score of the masked input sample belonging to category k, represents element-wise multiplication;
[0019] S2.3. The mathematical expectation of the network output before and after masking remains unchanged, specifically as follows:
[0020]
[0021] Among them, represents the mathematical expectation.
[0022] Furthermore, step S3 includes the following steps:
[0023] S3.1. Based on the network output L' after the action of the mask, use the softmax function to calculate the predicted probability distribution;
[0024] S3.2. Input the predicted probability distribution and the corresponding semantic segmentation annotation, and use the cross-entropy loss function to calculate the sample loss;
[0025] S3.3. Use the gradient descent method to train the fully convolutional image semantic segmentation network selected for segmentation until convergence.
[0026] Compared with the existing methods, the present invention has the following advantages and effects:
[0027] The present invention does not introduce any additional sub-models or design new loss functions, is simple and easy to expand, and the computational cost brought during training can be ignored. In addition, the present invention ensures the end-to-end training of the model, greatly simplifying the model training process. Description of the Drawings
[0028] Figure 1 is a schematic flowchart of a model training method for improving the quality of the predicted probability distribution of a semantic segmentation network in an embodiment of the present invention.
[0029] Figure 2 is a schematic flowchart of a naive model training method. Detailed Embodiments
[0030] In order to make the technical solutions and advantages of the present invention clearer, the following further details the specific implementation of the present invention in combination with the drawings and embodiments, but the implementation and protection of the present invention are not limited thereto.
[0031] In the following description, the technical solutions are elaborated in combination with specific drawings for a full understanding of the present invention application. However, the present invention application can be implemented in many other ways different from those described herein. Similar extended embodiments made by those of ordinary skill in the art without creative efforts all fall within the scope of protection of the present invention.
[0032] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to limit the description. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms, unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0033] Embodiment 1:
[0034] A model training method for improving the quality of the predicted probability distribution of a semantic segmentation network, comprising the following steps:
[0035] S1. Select any fully convolutional image semantic segmentation network for supervised training to obtain the output generated by the selected network for the input samples;
[0036] In this embodiment, the method described in "Fully Convolutional Networks for Semantic Segmentation" is selected, and an 18-layer residual network (ResNet) is used as the backbone network, denoted as FCN-R18, and the last layer of FCN-R18 is used as the output.
[0037] S2. Generate a qualified mask and a mask function, and apply the mask to the network output obtained in step S1 through the mask function, including the following steps:
[0038] S2.1. Generate a mask M using the Bernoulli distribution, K is the number of semantic segmentation categories in the selected fully convolutional image semantic segmentation network, and k is the category index of the input pixel sample, specifically as follows:
[0039]
[0040] Among them, represents the Bernoulli distribution, δ is an adjustable hyperparameter, and m k represents the mask acting on the k-th class prediction score. In this embodiment, δ is set to 0.9;
[0041] S2.2. Define the mask function Apply the mask M to the output L of the selected fully convolutional image semantic segmentation network through the mask function specifically as follows:
[0042]
[0043] Among them, l krepresents the prediction score of the model for the input sample belonging to class k, and L′ is the output after masking. l′ k represents the prediction score of the masked input sample belonging to class k. represents element-wise multiplication;
[0044] S2.3. The mathematical expectation of the network output before and after masking remains unchanged, specifically as follows:
[0045]
[0046] where represents the mathematical expectation.
[0047] S3. Based on the network output after the masking effect, use the softmax function to calculate the prediction probability distribution of the input sample, and use the cross-entropy loss function to supervise the model training until convergence, including the following steps:
[0048] S3.1. Based on the network output L′ after the masking effect, use the softmax function to calculate the prediction probability distribution;
[0049] S3.2. Input the prediction probability distribution and the corresponding semantic segmentation annotation, and use the cross-entropy loss function to calculate the sample loss;
[0050] S3.3. Use the gradient descent method to train the fully convolutional image semantic segmentation network selected for segmentation until convergence.
[0051] In this embodiment, the area under the receiver operating characteristic curve (Area Under Receiver Operating Characteristic, AUC) is used as the evaluation criterion for the quality of the prediction probability distribution. On the publicly available CamVid dataset, the AUC score of the model trained by the training method of the present invention is 83.54%, Figure 2 and the AUC score of the model trained by the shown naive training method is 61.53%. Without affecting the segmentation performance of the model, the present invention effectively improves the quality of the model's prediction probability distribution.
[0052] Embodiment 2:
[0053] A model training method for improving the quality of the prediction probability distribution of a semantic segmentation network, including the following steps:
[0054] S1. Select any fully convolutional image semantic segmentation network for supervised training to obtain the output generated by the selected network for the input sample;
[0055] Select the method described in "Rethinking atrous convolution for semantic image segmentation", and use the 101-layer Residual Network (ResNet) as the backbone network, denoted as DeepLabv3-R101, and use the last layer of DeepLabv3-R101 as the output.
[0056] S2. Generate a qualified mask and a mask function, and apply the mask to the network output obtained in step S1 through the mask function, including the following steps:
[0057] S2.1. Generate a mask M using the Bernoulli distribution. K is the number of semantic segmentation categories in the selected fully convolutional image semantic segmentation network, and k is the category index of the input pixel sample, specifically as follows:
[0058]
[0059] Among them, represents the Bernoulli distribution, δ is an adjustable hyperparameter, mk represents the mask applied to the prediction score of the k-th category, and δ is set to 0.9 in this embodiment;
[0060] S2.2. Define the mask function Apply the mask M to the output L of the selected fully convolutional image semantic segmentation network through the mask function Specifically as follows:
[0061]
[0062] Among them, l k represents the prediction score of the model for the input sample belonging to category k, and L′ is the output after masking. l′ k represents the prediction score of the masked input sample belonging to category k. represents element-wise multiplication;
[0063] S2.3. The mathematical expectation of the network output before and after masking remains unchanged, specifically as follows:
[0064]
[0065] Among them, represents the mathematical expectation.
[0066] S3. Based on the network output after the mask is applied, use the softmax function to calculate the prediction probability distribution of the input sample, and use the cross-entropy loss function to supervise the model training until convergence, including the following steps:
[0067] S3.1. Calculate the predicted probability distribution using the softmax function based on the network output L' after the mask effect;
[0068] S3.2. Input the predicted probability distribution and the corresponding semantic segmentation annotation, and calculate the sample loss using the cross-entropy loss function;
[0069] S3.3. Use the gradient descent method to train the fully convolutional image semantic segmentation network selected for segmentation until convergence.
[0070] In this embodiment, on the publicly available Cityscapes dataset, the AUC score of the model trained by the training method of the present invention is 73.57%, Figure 2 The AUC score of the model trained by the shown naive training method is 54.35%.
[0071] Embodiment 3:
[0072] A model training method for improving the quality of the predicted probability distribution of a semantic segmentation network, including the following steps:
[0073] S1. Select any fully convolutional image semantic segmentation network for supervised training, and obtain the output generated by the selected network for the input sample;
[0074] Select the method described in "Alignseg: Feature-aligned segmentation networks", and use a 101-layer residual network (ResNet) as the backbone network, denoted as AlignSeg-R101, and use the last layer of AlignSeg-R101 as the output.
[0075] S2. Generate a mask and a mask function that meet the conditions, and apply the mask to the network output obtained in step S1 through the mask function, including the following steps:
[0076] S2.1. Generate a mask M using the Bernoulli distribution, K is the number of semantic segmentation categories in the selected fully convolutional image semantic segmentation network, and k is the category index of the input pixel sample, specifically as follows:
[0077]
[0078] Among them, represents the Bernoulli distribution, δ is an adjustable hyperparameter, and m k represents the mask acting on the k-th class prediction score. In this embodiment, δ is set to 0.9;
[0079] S2.2. Define the mask function Apply the mask M through the mask function Act on the output L of the selected fully convolutional image semantic segmentation network, specifically as follows:
[0080]
[0081] Among them, l k represents the prediction score of the model for the input sample belonging to class k, and L' is the output after masking. l' k represents the prediction score of the masked input sample belonging to class k. represents element-wise multiplication;
[0082] S2.3. The mathematical expectation of the network output before and after masking remains unchanged, specifically as follows:
[0083]
[0084] Among them, represents the mathematical expectation.
[0085] S3. Based on the network output after the masking effect, use the softmax function to calculate the prediction probability distribution of the input sample, and use the cross-entropy loss function to supervise the model training until convergence, including the following steps:
[0086] S3.1. Based on the network output L' after the masking effect, use the softmax function to calculate the prediction probability distribution;
[0087] S3.2. Input the prediction probability distribution and the corresponding semantic segmentation annotation, and use the cross-entropy loss function to calculate the sample loss;
[0088] S3.3. Use the gradient descent method to train the selected fully convolutional image semantic segmentation network until convergence.
[0089] In this embodiment, on the publicly available Cityscapes dataset, the AUC score of the model trained by the training method of the present invention is 77.71%, Figure 2 and the AUC score of the model trained by the shown naive training method is 55.16%.
[0090] It should be noted that for the embodiments of the model training method for improving the quality of the prediction probability distribution of the semantic segmentation network in the embodiments, for the sake of simplicity of description, they are all expressed as a series of steps or combinations of operations. However, those skilled in the art should know that the present application is not limited by the described order of actions, because according to the present application, certain steps or operations can be performed in other orders or simultaneously.
[0091] The preferred embodiments of the present application disclosed above are only used to help understand the present invention and its core idea. For those of ordinary skill in the art, according to the idea of the present invention, there will be changes in specific application scenarios and implementation operations. This specification should not be construed as a limitation to the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A model training method for improving the quality of the predicted probability distribution of a semantic segmentation network, characterized in that It includes the following steps: S1. Select any fully convolutional image semantic segmentation network for supervised training to obtain the output generated by the selected network for the input samples; S2. Generate a mask and a mask function that meet the conditions, and apply the mask to the network output obtained in step S1 through the mask function; specifically, it includes the following steps: S2.
1. Generate a mask M using the Bernoulli distribution. K is the number of semantic segmentation categories in the selected fully convolutional image semantic segmentation network, and k is the category index of the input pixel sample, specifically as follows: Among them represents the Bernoulli distribution, δ is an adjustable hyperparameter, and m k represents the mask acting on the prediction score of the k-th class; S2.
2. Define the masking function Apply the mask M to the output L of the selected fully convolutional image semantic segmentation network through the masking function Specifically as follows: Among them, l k represents the predicted score of the model for the input sample belonging to class k, and L' is the output after masking, l' k represents the predicted score of the masked input sample belonging to class k, represents element-wise multiplication; S2.
3. The mathematical expectation of the network output before and after the mask remains unchanged; S3. Based on the network output after the mask is applied, calculate the predicted probability distribution of the input samples, and supervise the model training until convergence.
2. The model training method for improving the quality of the predicted probability distribution of a semantic segmentation network according to claim 1, wherein In step S1, the output of the last layer of the selected fully convolutional image semantic segmentation network is used as the output of the entire fully convolutional image semantic segmentation network.
3. A model training method for improving the quality of the predicted probability distribution of a semantic segmentation network, characterized in that In step S2.3, specifically as follows: Among them, represents the mathematical expectation.
4. A model training method for improving the quality of the predicted probability distribution of a semantic segmentation network, according to any one of claims 1 to 3, characterized in that, Step S3 includes the following steps: S3.
1. Based on the network output L' after the mask is applied, calculate the predicted probability distribution; S3.
2. Input the predicted probability distribution and the corresponding semantic segmentation annotation, and calculate the sample loss; S3.
3. Train the selected fully convolutional image semantic segmentation network for segmentation until convergence.
5. A model training method for improving the quality of the predicted probability distribution of a semantic segmentation network, characterized in that, In step S3.1, the softmax function is used to calculate the predicted probability distribution.
6. The model training method for improving the quality of the predicted probability distribution of the semantic segmentation network according to claim 4, characterized in that In step S3.2, the cross-entropy loss function is used to calculate the sample loss.
7. A model training method for improving the quality of the predicted probability distribution of a semantic segmentation network, characterized in that, In step S3.3, the gradient descent method is used to train the selected fully convolutional image semantic segmentation network for segmentation until convergence.
Citation Information
Patent Citations
Named entity recognition model training method and device, equipment and medium
CN112966517A
Method for improving robust performance of convolutional neural network
CN113255768A