A novel method for analyzing CT images of COVID-19 based on semi-supervised deep learning
By adopting semi-supervised deep learning method and attention mechanism in the CT image analysis of COVID-19, the problems of high data labeling cost and model overfitting are solved, the diagnostic efficiency and accuracy are improved, and higher generalization and performance are achieved.
Patent Information
- Application Number
- CN202210151409.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-02-18
AI Technical Summary
The prior art faces the problems of high data labeling cost, insufficient data volume and overfitting of models in the CT image analysis of COVID-19, resulting in low diagnostic efficiency and accuracy.
The CT image analysis method of COVID-19 based on semi-supervised deep learning is adopted. By introducing attention mechanisms into the classification model and using data augmentation and pseudo-label generation technology, the training signals with supervised data are gradually released to improve the generalization ability of the model.
It improves the diagnostic efficiency and accuracy of CT image analysis of COVID-19, and can achieve higher generalization and performance under the conditions of a small amount of labeled data and a large amount of labeled data.
Smart Images

Figure CN114549452B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for analyzing CT images of COVID-19, and in particular to a method for analyzing CT images of COVID-19 based on semi-supervised deep learning. Background Art
[0002] One of the most critical steps in preventing and fighting COVID-19 is to effectively screen suspected infected patients. In the early stages of the epidemic, reverse transcription polymerase chain reaction (RT-PCR) was usually used to determine whether a patient had COVID-19. However, due to the rapid outbreak of the epidemic, many countries lacked sufficient test kits to test suspected patients. Moreover, RT-PCR testing takes several days to produce results, and excessive testing time will lead to delays in epidemic control and treatment. In addition, RT-PCR testing has low sensitivity, and one test may not be able to make an accurate judgment, so multiple tests are required to make a final judgment. In clinical practice, researchers have found that chest computed tomography (CT) images of COVID-19 patients all show imaging features such as ground-glass shadows and multifocal patchy consolidation. Moreover, compared with RT-PCR testing, doctors can get chest CT scans and corresponding diagnostic results faster. Moreover, CT scanning equipment is very popular in modern healthcare systems, so CT has become another effective way to screen and diagnose COVID-19 in the early stage.
[0003] In recent years, as deep learning has made breakthrough progress in computer vision, it has been widely used in image classification, image positioning and detection, and medical image segmentation, which has greatly reduced the burden on doctors caused by massive medical image data. Currently, most commonly used medical image diagnosis methods are based on supervised learning and require a large amount of labeled data. However, in many practical work, there may be only a few labeled samples available because the cost of labeling data is very high. The CT acquisition and labeling of COVID-19 requires a lot of time and energy from professional doctors, which is even more serious during the epidemic. Training deep learning models requires a large amount of labeled data to achieve clinical standard performance. Insufficient data will lead to overfitting of the model, resulting in poor performance of the model. Secondly, because medical image data involves patient privacy issues, many CT image datasets are not public. Models trained with these non-public datasets cannot be used in other hospitals. Summary of the invention
[0004] The purpose of the present invention is to provide a method for analyzing COVID-19 CT images based on semi-supervised learning and attention mechanism.
[0005] The present invention adopts the following technical solutions:
[0006] A method for analyzing COVID-19 CT images based on semi-supervised deep learning, characterized in that it comprises the following steps:
[0007] (1) Establish the relationship between images and category labels, that is, the classification model. The classification model is based on the residual neural network in deep learning and adds an attention module;
[0008] (2) Perform two different data augmentations on each unlabeled training sample to obtain two new images;
[0009] (3) classifying the image obtained after data enhancement by the classification model trained in step (1) to obtain its classification result;
[0010] (4) Perform entropy minimization on the classification results of the image after unlabeled sample enhancement, and regard the processed results as its pseudo-label;
[0011] (5) Perform data augmentation on each labeled training sample to obtain its enhanced image;
[0012] (6) Mixing the unlabeled training samples after the pseudo-label data enhancement and the labeled training samples after the data enhancement to obtain new training samples;
[0013] (7) Substitute the new training samples and corresponding labels into the classification model for training and update the network parameter information;
[0014] (8) During the training process, as the number of unlabeled training samples increases, labeled training samples are gradually removed, gradually releasing the training signal of supervised data;
[0015] (9) Perform weighted summation on the weights of the fully connected layer and feature map in the classification model to generate an attention map, highlighting the important areas that are closely related to the prediction results;
[0016] (10) Use the trained classification model to analyze the sample images in the test set and generate a visualization of the analysis results.
[0017] Furthermore, the novel coronavirus pneumonia CT image analysis method based on semi-supervised deep learning of the present invention also has the following characteristics: the specific process of step (1) is as follows: the classification model adopts the residual neural network model in the classic image classification model, and adds a module of the attention mechanism to the model. First, the feature map extracted by the residual neural network is subjected to the global maximum pooling and global average pooling based on width and height respectively to obtain two feature maps after the convolution layer. Then, the output features of the convolution layer are added. The final attention map A is generated by the sigmoid function:
[0018] A = Sigmoid (Conv (Avgpool (F)) + Conv (Maxpool (F))), where F is the feature extracted by the residual neural network, AvgPool is the average pooling function, MaxPool is the maximum pooling function, and Conv is the convolution function.
[0019] Furthermore, the novel coronavirus pneumonia CT image analysis method based on semi-supervised deep learning of the present invention also has the following characteristics: the specific process of step (2) is: performing data enhancement twice on the unlabeled image, and the enhancement methods include: standardization, geometric transformation, random adjustment of brightness, and random adjustment of contrast.
[0020] Furthermore, the novel coronavirus pneumonia CT image analysis method based on semi-supervised deep learning of the present invention also has the following characteristics: the specific process of step (4) is: the classification results are processed to minimize entropy, forcing the classifier to make low entropy predictions for unlabeled training samples, and using the sharpening function to minimize the entropy of unlabeled data, which is in the following form: Where p is the probability category, T is the temperature parameter used to adjust the classification entropy, i is the number of samples, j represents from 1 to the number of categories, and L is the total number of categories.
[0021] Furthermore, the novel coronavirus pneumonia CT image analysis method based on semi-supervised deep learning of the present invention also has the following characteristics: the specific process of step (6) is: Mixup the unlabeled training samples after the pseudo-label data enhancement and the labeled training samples after the data enhancement to obtain new training samples and enhance the robustness of the classification model. The formula of mixup is as follows:
[0022] x′=μ′x 1 +(1-μ′)x 2
[0023] p′=μ′p 1 +(1-μ′)p 2
[0024] μ~Beta(α,α)
[0025] μ′=max(μ,1-μ)
[0026] where x 1 , p 1 are images of labeled training samples and their corresponding labels, x 2 , p 2 are the images and corresponding labels of unlabeled training samples. α represents the distribution parameter of Beta, and μ represents the sample mixing weight.
[0027] Furthermore, the novel coronavirus pneumonia CT image analysis method based on semi-supervised deep learning of the present invention also has the following characteristics: the specific process of step (8) is as follows: at the training time t, a threshold ηt is set, and 1 / K≤η t ≤1, where K is the number of categories. When the probability of the correct category P of a labeled example is higher than the threshold η t , the model removes this example from the loss function and only trains other labeled examples in this minibatch.
[0028] The present invention also provides a COVID-19 CT image analysis system based on semi-supervised deep learning, which is characterized by comprising:
[0029] The deep learning module establishes the relationship between images and category labels to form a classification model module. The classification model module is based on the residual neural network in deep learning and adds an attention module.
[0030] The deep learning module performs two different data augmentations on each unlabeled training sample to obtain two new images;
[0031] The classification model module classifies the image obtained after data enhancement to obtain its classification result;
[0032] The classification model module performs entropy minimization on the classification results of the image after unlabeled sample enhancement, and regards the processed results as its pseudo-label;
[0033] The data enhancement module performs data enhancement on each labeled training sample to obtain its enhanced image;
[0034] The data mixing module mixes the unlabeled training samples after the pseudo-label data enhancement and the labeled training samples after the data enhancement to obtain new training samples;
[0035] The classification model module trains new training samples and corresponding labels and updates network parameter information. During the training process, as the number of unlabeled training samples increases, labeled training samples are gradually removed, gradually releasing the training signal of supervised data.
[0036] The attention map generation module performs weighted summation on the weights of the fully connected layer and feature map in the classification model module to generate an attention map, highlighting important areas that are closely related to the prediction results.
[0037] Beneficial effects of the invention: The present invention proposes a general semi-supervised deep learning method, which can introduce unlabeled samples and use the hidden distribution information learned by the model to promote the classifier to move in the correct decision direction, thereby achieving higher generalization and accuracy. In the field of natural image recognition, semi-supervised learning can use a small amount of labeled data and a large amount of unlabeled data to alleviate the problem of insufficient data. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a flowchart of the COVID-19 CT image diagnosis method based on semi-supervised learning and attention mechanism provided by the present invention.
[0039] Figure 2 This is an image example from the open-source COVID-19 CT image dataset.
[0040] Figure 3 It is a structural diagram of the classification model proposed in the present invention.
[0041] Figure 4(a) is a histogram of the classification indicators of ablation learning.
[0042] Figure 4(b) is the classification indicator confusion matrix of ResNet50 ablation learning.
[0043] Figure 4(c) is the classification indicator confusion matrix of ResNet50+Attention+SSL ablation learning.
[0044] Figure 4(d) is the ROC curve of the classification indicator of ablation learning.
[0045] Figure 5 To better understand the model’s decision, we visualize the attention map of the lesion area. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. The process of the present invention is as follows Figure 1 and Figure 3 The steps shown include:
[0047] (1) Establish the relationship between images and category labels, that is, the classification model. The classification model is based on the residual neural network in deep learning, and an attention module is added on its basis.
[0048] The classification model adopts the residual neural network model in the classic image classification model, and adds the module of the attention mechanism to the model to enhance the model's ability to extract features. First, the feature map extracted by the residual neural network is subjected to global maximum pooling and global average pooling based on width and height respectively to obtain two feature maps after the convolution layer. Then, the output features of the convolution layer are added. Finally, the attention map A is generated by the sigmoid function: A = Sigmoid (Conv (Avgpool (F)) + Conv (Maxpool (F))), where F is the feature extracted by the residual neural network, AvgPool is the average pooling function, MaxPool is the maximum pooling function, and Conv is the convolution function.
[0049] (2) Perform two different data augmentations on each unlabeled training sample to obtain two new images. Commonly used data augmentation methods in images generally include: standardization, geometric transformation (translation, flipping, rotation), random adjustment of brightness, random adjustment of contrast, etc.
[0050] (3) The image obtained after data enhancement is classified using the classification model trained in the previous stage to obtain its classification result.
[0051] (4) Perform entropy minimization on the classification results of the image after unlabeled sample enhancement, and regard the processed results as their pseudo labels. Perform entropy minimization on the classification results to force the classifier to make low entropy predictions for unlabeled training samples. The sharpening function is used in the present invention to minimize the entropy of unlabeled data, which is as follows: Where p is the probability category, T is the temperature parameter used to adjust the classification entropy, i is the number of samples, j represents from 1 to the number of categories, and L is the total number of categories.
[0052] (5) Perform data augmentation on each labeled training sample to obtain its enhanced image.
[0053] (6) Mixup the unlabeled training samples after the pseudo-label data enhancement and the labeled training samples after the data enhancement to obtain new training samples. Enhance the robustness of the classification model. The formula for mixup is as follows:
[0054] x′=μ′x 1 +(1-μ′)x 2
[0055] p′=μ′p 1 +(1-μ′)p 2
[0056] μ~Beta(α,α)
[0057] μ′=max(μ,1-μ)
[0058] where x 1 , p 1 are images of labeled training samples and their corresponding labels, x 2 , p 2 are the images and corresponding labels of unlabeled training samples. α represents the distribution parameter of Beta, and μ represents the sample mixing weight.
[0059] (7) Substitute the new training samples and corresponding labels into the classification model for training and update the network parameter information.
[0060] (8) During the training process, as the number of unlabeled training samples increases, labeled training samples are gradually removed, and the training signal of supervised data is gradually released. At the training time t, a threshold ηt is set, and 1 / K≤η t ≤1, where K is the number of categories. When the probability of the correct category P of a labeled example is higher than the threshold η t When , the model removes this example from the loss function and only trains other labeled examples under this minibatch. t Used to prevent the model from overfitting to the labeled data. t As it approaches 1, the model can only slowly get supervision from the labeled instances, which greatly alleviates the overfitting problem.
[0061] (9) The weights of the fully connected layer and feature map in the classification model are weighted summed to generate an attention map, which highlights the important areas that are closely related to the prediction results.
[0062] (10) Use the trained classification model to analyze the sample images in the test set and generate a visualization of the analysis results.
[0063] The embodiment adopts the method for COVID-19 CT diagnosis based on semi-supervised learning and attention mechanism provided by the present invention, and is verified by a public dataset of COVID-19 CT images.
[0064] A labeled CT dataset and an unlabeled CT dataset are used to evaluate the proposed method in the diagnosis of COVID-19 CT images. The labeled CT dataset is a COVID-19 public dataset containing 349 positive and 397 negative CT scans. Figure 3 An example of COVID-19 CT images is shown. The division of the dataset is shown in Table 1.
[0065] Table 1: Division of COVID-19 CT image dataset
[0066] Class Training set Validation set Test Set Coronavirus disease 191 60 98 normal 234 58 105 total 425 118 203
[0067] The positive samples are 760 COVID-19 preprints from medRxiv and bioRxiv, and the negative samples are CT scans of normal people or other types of diseases. The unlabeled sample dataset comes from the LUNA dataset, which consists of low-dose lung CT images and is designed for the detection and segmentation of lung nodules in patients. We selected 500 of them as unlabeled samples and added them to the training set, and used these images as unlabeled images for semi-supervised learning.
[0068] In order to fully understand the effect of each part of the method provided by the present invention, the following ablation studies were conducted, which are divided into four cases: (1) using residual neural network alone; (2) using residual neural network plus attention mechanism; (3) using residual neural network plus semi-supervised learning; (4) using residual neural network plus attention mechanism plus semi-supervised learning.
[0069] First from Figure 4(a) to Figure 4(d) It can be seen that the model with the attention module performs better than the model without the attention module. This shows that the proposed attention module can ensure that the decision of the model mainly depends on the infected area and suppress the contribution of irrelevant parts of the image, thereby improving the performance of the model. Secondly, it is obvious from this table that the performance of the model is improved when semi-supervised learning is used. This shows that semi-supervised learning can improve the generalization ability of the model by expanding the dataset and reducing the risk of overfitting the model on a smaller dataset. When the attention module and semi-supervised learning are used together, both models achieve the best performance. Third, when semi-supervised learning and the attention module are used alone, the specificity of the model is reduced, which shows that some advantages of the model may be sacrificed in order to meet specific needs.
[0070] Figure 5 The visualization of the classification results of the baseline and the model is shown. The first column represents the original COVID-19 CT image. The second and third columns in the figure show the results of using the residual neural network alone. The colors from dark red to dark blue correspond to the values of the class significance of the pixels from large to small. The fourth and fifth columns show the results of the proposed model. By comparing the baseline results with the results of the present invention, we observe that the proposed model can capture almost all the salient regions of the prediction.
[0071] This embodiment also provides a COVID-19 CT image analysis system based on semi-supervised deep learning, including:
[0072] The deep learning module establishes the relationship between images and category labels to form a classification model module. The classification model module is based on the residual neural network in deep learning and adds an attention module.
[0073] The deep learning module performs two different data augmentations on each unlabeled training sample to obtain two new images;
[0074] The classification model module classifies the image obtained after data enhancement to obtain its classification result;
[0075] The classification model module performs entropy minimization on the classification results of the image after unlabeled sample enhancement, and regards the processed results as its pseudo-label;
[0076] The data enhancement module performs data enhancement on each labeled training sample to obtain its enhanced image;
[0077] The data mixing module mixes the unlabeled training samples after the pseudo-label data enhancement and the labeled training samples after the data enhancement to obtain new training samples;
[0078] The classification model module trains new training samples and corresponding labels and updates network parameter information. During the training process, as the number of unlabeled training samples increases, labeled training samples are gradually removed, gradually releasing the training signal of supervised data.
[0079] The attention map generation module performs weighted summation on the weights of the fully connected layer and feature map in the classification model module to generate an attention map, highlighting important areas that are closely related to the prediction results.
[0080] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for analyzing COVID-19 CT images based on semi-supervised deep learning, characterized in that: The following steps are involved: (1) Establish the relationship between images and category labels, that is, the classification model. The classification model is based on the residual neural network in deep learning and adds an attention module; (2) Perform two different data augmentations on each unlabeled training sample to obtain two new images; (3) classifying the image obtained after data enhancement by the classification model trained in step (1) to obtain its classification result; (4) Perform entropy minimization on the classification results of the image after unlabeled sample enhancement, and regard the processed results as its pseudo-label; (5) Perform data augmentation on each labeled training sample to obtain its enhanced image; (6) Mixing the unlabeled training samples after the pseudo-label data enhancement and the labeled training samples after the data enhancement to obtain new training samples; (7) Substitute the new training samples and corresponding labels into the classification model for training and update the network parameter information; (8) During the training process, as the number of unlabeled training samples increases, labeled training samples are gradually removed, gradually releasing the training signal of supervised data; (9) Perform weighted summation on the weights of the fully connected layer and feature map in the classification model to generate an attention map, highlighting the important areas that are closely related to the prediction results; (10) Use the trained classification model to analyze the sample images in the test set and generate a visualization of the analysis results.
2. The method for analyzing COVID-19 CT images based on semi-supervised deep learning according to claim 1 is characterized in that: The specific process of step (1) is as follows: the classification model adopts the residual neural network model in the classic image classification model, and a module of the attention mechanism is added to the model. First, the feature map extracted by the residual neural network is subjected to the global maximum pooling and global average pooling based on width and height respectively to obtain two feature maps after the convolution layer. Then, the output features of the convolution layer are added, and the final attention map A is generated by the sigmoid function: A=Sigmoid(Conv(Avgpool(F))+Conv(Maxpool(F))), wherein F is the feature extracted by the residual neural network, AvgPool is the average pooling function, MaxPool is the maximum pooling function, and Conv is the convolution function.
3. The method for analyzing COVID-19 CT images based on semi-supervised deep learning according to claim 1 is characterized in that: The specific process of step (2) is: performing two data enhancements on the unlabeled image, and the enhancement methods include: standardization, geometric transformation, random brightness adjustment, and random contrast adjustment.
4. The method for analyzing COVID-19 CT images based on semi-supervised deep learning according to claim 1, characterized in that: The specific process of step (4) is: perform entropy minimization processing on the classification results, force the classifier to make low entropy predictions for unlabeled training samples, and use the sharpening function to minimize the entropy of unlabeled data, in the following form: Where p is the probability category, T is the temperature parameter used to adjust the classification entropy, i is the number of samples, j represents from 1 to the number of categories, and L is the total number of categories.
5. The method for analyzing COVID-19 CT images based on semi-supervised deep learning according to claim 1, characterized in that: The specific process of step (6) is: Mixup the unlabeled training samples after the pseudo-label data enhancement and the labeled training samples after the data enhancement to obtain new training samples and enhance the robustness of the classification model, wherein the formula of mixup is as follows: x′=μ′x1+(1-μ′)x2 p′=μ′p1+(1-μ′)p2 μ~Beta(α,α) μ′=max(μ,1-μ) Among them, x1, p1 are the images of labeled training samples and the corresponding labels, x2, p2 are the images of unlabeled training samples and the corresponding labels, α represents the distribution parameter of Beta, and μ represents the sample mixing weight.
6. The method for analyzing COVID-19 CT images based on semi-supervised deep learning according to claim 1, characterized in that: The specific process of step (8) is as follows: at the training time t, a threshold η is set t , and 1 / K≤η t ≤1, where K is the number of categories. When the probability of the correct category P of a labeled example is higher than the threshold η t , the model removes this example from the loss function and only trains other labeled examples in this minibatch.
7. A COVID-19 CT image analysis system based on semi-supervised deep learning, characterized in that: include: The deep learning module establishes the relationship between images and category labels to form a classification model module. The classification model module is based on the residual neural network in deep learning and adds an attention module. The deep learning module performs two different data augmentations on each unlabeled training sample to obtain two new images; The classification model module classifies the image obtained after data enhancement to obtain its classification result; The classification model module performs entropy minimization on the classification results of the image after unlabeled sample enhancement, and regards the processed results as its pseudo-label; The data enhancement module performs data enhancement on each labeled training sample to obtain its enhanced image; The data mixing module mixes the unlabeled training samples after the pseudo-label data enhancement and the labeled training samples after the data enhancement to obtain new training samples; The classification model module trains new training samples and corresponding labels and updates network parameter information; During the training process, as the number of unlabeled training samples increases, labeled training samples are gradually removed, gradually releasing the training signal of supervised data; The attention map generation module performs weighted summation on the weights of the fully connected layer and feature map in the classification model module to generate an attention map, highlighting important areas that are closely related to the prediction results.
Citation Information
Patent Citations
Image classification method based on semi-supervised self-paced learning cross-task deep network
CN108764281A
An unsupervised / semi-supervised CT image reconstruction depth network train method
CN109035169A