Method for Semantic Segmentation of COVID-19 CT Images Based on the Segmentation Network Ref-Net
By designing the Ref-Net network, combined with the Res2Net backbone network and edge attention, attention positioning and context exploration modules, the problems of inaccurate detection results and scarce data sets in the segmentation of COVID-19 were solved, and higher segmentation accuracy and effective screening of infected people were achieved.
Patent Information
- Application Number
- CN202210903236.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The prior art has problems of inaccurate detection results and scarce data sets in the segmentation of CT images of COVID-19, especially because the lung nodules are small and similar to the background, making the accuracy of the detection results difficult to guarantee.
A regional segmentation network for COVID-19 infection was designed, Ref-Net, using Res2Net as the backbone network, and on the basis of it, the edge attention module (EAM), attention positioning module (APM) and context exploration module (CEM) were added to improve the accuracy of segmentation results.
Through the Ref-Net network, the accuracy of the CT image segmentation results of COVID-19 has been improved, the average absolute error (MAE) has been increased by more than 26.8%, and the Dice similarity coefficient, sensitivity, specificity rate and structural indicators have been improved, which has significantly promoted the effective screening of infected people.
Smart Images

Figure CN115359255B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a method for semantic segmentation of COVID-19 CT images based on a segmentation network Ref-Net. Background Art
[0002] Corona Virus Disease 2019 (COVID-19), abbreviated as "COVID-19", is an acute respiratory infectious disease caused by the novel coronavirus; currently, the global epidemic situation is still very serious, and a key step in fighting pneumonia is the effective screening of infected people; at present, the screening of COVID-19 mainly involves ribonucleic acid testing of respiratory specimens of subjects through reverse transcription polymerase chain reaction (RT-PCR), which is a complex and time-consuming manual process with a highly variable positive rate and the positive rate decreasing over time after the onset of symptoms; therefore, the application research of X-ray and computed tomography (CT) radiography technology in COVID-19 detection is of great significance.
[0003] "Fang Y, Zhang H, Xie J, et al. Sensitivity of Chest CT for COVID-19: Comparison to RT-PCR[J]. Radiology, 2020:200432" compared the sensitivities of chest CT and RT-PCR for COVID-19. The survey results showed that the detection rate of the first RT-PCR was 71% in 51 samples, and the confidence interval was 56% - 83%; while the detection rate of the first CT reached 98% (50 / 51), and the confidence interval was 90% - 100%; it can be seen that the sensitivity of chest CT is higher than that of reverse transcription polymerase chain reaction. "Ye, Z., et al. "Chest CT manifestations of new coronavirus disease 2019 (COVID-19): a pictorial review." European Radiology 30.8(2020)" pointed out that chest computed tomography (CT) can be used as an effective supplementary detection method for reverse transcription polymerase chain reaction (RT-PCR). "Wang, L., and A. Wong. "COVID-Net: A Tailored Deep Convolutional Neural Network Design for Detection of COVID-19 Cases from Chest X-Ray Images." (2020)" proposed the COVID-Net deep convolutional neural network, which can detect COVID-19 cases from chest X-ray images. In addition, the COVID-19 dataset COVIDx was also provided: it consists of 5941 chest X-ray images provided by 2839 pneumonia patients; compared with X-rays, CT images can perform three-dimensional imaging of the lungs, so screening for COVID-19 through CT images is more widely accepted."Xu, X., et al. "Deep Learning System to Screen Coronavirus Disease 2019 Pneumonia." arXiv (2020)" trained on a CT image dataset using a deep learning model, segmented candidate infection regions, and used a position attention classification model to classify the infection regions, which can be used as an auxiliary diagnostic method for front-line clinicians. At the same time, combining "Zheng, Chuansheng, et al. "Deep Learning-based Detection for COVID-19 from Chest CT using Weak Label." (2020)" and "Kaheel, Hussein, A. Hussein, and A. Chehab. "AI-Based Image Processing for COVID-19 Detection in Chest CT Scan Images." (2021)", it can be found that there are not many studies on the CT image segmentation of COVID-19 at present. The main reasons are as follows: (1) The pulmonary nodule regions are very small and the background and infection regions in CT images are extremely similar, often resulting in inaccurate detection results; (2) The dataset is scarce (especially the dataset with infection region annotations), and it is time-consuming and laborious to obtain a high-quality dataset, making it difficult to train a high-quality deep learning network. Summary of the Invention
[0004] The purpose of the present invention is to improve the accuracy of detection results and thus effectively screen infected persons.
[0005] To achieve the above object, the present invention uses Res2Net as the backbone network and designs a COVID-19 infection region segmentation network Ref-Net. Based on this segmentation network Ref-Net, semantic segmentation of COVID-19 CT images is performed. The network structure framework diagram of the segmentation network Ref-Net is as Figure 1 shown. Res2Net increases the receptive field range in each layer of the network and can represent multi-scale features with finer granularity. Therefore, multi-layer features are extracted using Res2Net as the backbone network , , , , ; First, since the underlying features retain the edge information of the target object, an edge attention module (EAM) is added after to generate an edge feature map , and the standard binary cross-entropy loss function is used to calculate its true edge feature map generated with the ground truth image Differences; secondly, since the high layer is rich in the global positioning semantic information of the target object, The input is fed into the Attention Positioning Module (APM) composed of channel attention and spatial attention to obtain a blurred prediction map ; finally, in order to eliminate the false predictions (false negatives and false positives) in the blurred prediction map, , , The Context Exploration Module (CEM) is added to fuse the current feature and the upper layer feature respectively, where The feature obtained after passing through the attention positioning module is used as the high layer feature of the first CEM, and after being refined by 3 Context Exploration Modules (CEM), a clear and accurate prediction map (Prediction Map) is finally obtained;
[0006] Edge Attention Module
[0007] There are differences and complementarities between the high layer features and the low layer features of the convolutional neural network; specifically, the deep high layer features retain semantic information to obtain an abstract description, but they have less image detail information, and the shallow low layer features focus on spatial information to construct object boundaries. Since the high layer prediction map is clear and rich in global semantic information, most networks attach importance to the high layer features while ignoring the low layer edge result information, resulting in blurred predicted edges; therefore, the present invention feeds the low layer features with a more appropriate resolution into the edge attention module to focus on the edge features; specifically, the edge attention module is composed of a convolutional layer with only one convolutional kernel, and an edge feature map is obtained after passing through this convolutional layer ;
[0008] In order to compare the differences between the edge feature map and the edge feature map of the ground truth image a standard binary cross-entropy loss function is introduced; the calculation of the difference loss function is shown in Equations (1) and (2):
[0009] (1)
[0010] , (2)
[0011] where w and h respectively represent the width and height of the feature map, is the ground truth edge feature map derived from the ground truth map GT, is the edge feature map predicted by passing through the edge attention module; where, represents the weight, which is set to 1 in the present invention; two parts of and Global supervision and local supervision are provided respectively, and the segmentation result is more accurate;
[0012] Attention localization module
[0013] The attention mechanism can selectively focus on important regions in the image. Channel attention focuses on which channel features are more meaningful; spatial attention focuses on which regional features are more meaningful. Combining the two attention mechanisms enables the network to focus on more meaningful features and more accurately locate the target in both the channel and spatial dimensions;
[0014] The structural diagram of the attention localization module consists of channel attention and spatial attention, aiming to further enhance the high-level global semantic information to obtain a preliminary segmentation result; as Figure 2 shown, the specific structure of the attention localization module is: the output of the deepest layer is used as the input feature F of the attention localization module. The number of channels, height, and width of the feature F are denoted as C, H, and W respectively; first, the input feature F is reshaped to obtain the query (Q), key (K), and value (V) respectively, where , N refers to the number of pixels. Then, after transposing Q and performing matrix multiplication with K and passing through the Softmax layer, the channel attention map (Channel Attention Map) is obtained . The influence of the j-th channel on the i-th channel in the channel attention map X is specifically expressed as shown in Equation (3):
[0015] (3)
[0016] where represents the Q th i row of the matrix represents the K th j row of the matrix
[0017] X and V After performing matrix multiplication between the transposes of, the resulting feature shape size is converted to . In addition, the attention localization module also introduces a scaling parameter , which is learnable, with an initial value of 1 and continuously learns and updates the weights. Finally, the final channel attention output feature is obtained through a skip connection , and any row of it can be expressed by the formula as shown in Equation (4):
[0018] (4)
[0019] where Represents the output feature of channel attention The i th row, is the proportional parameter, represents the value V The j th row of the matrix, Represents the i th row of the input feature of channel attention;
[0020] The specific process of spatial attention is similar to that of channel attention. As the input of spatial attention, it passes through 3 convolutions and changes the shape of the convolution result to obtain new query ( ), key ( ), and value ( ) respectively; It should be noted that , while . Similarly, The transpose of and perform matrix multiplication and normalization to obtain the Spatial Attention Map. j In spatial attention, the influence of the i th position on the th position is calculated as shown in Equation (5):
[0021] (5)
[0022] Where, represents the th column of the query i , represents the th column of the key j . The subsequent process is similar to that of channel attention. and perform matrix multiplication after the transpose of and then introduce the proportional parameter and finally obtain the final output through skip connection.
[0023] Context Exploration Module
[0024] The COVID-19 infection area is small and highly similar to the background, so there will be false positive predictions and false negative predictions in the predicted segmentation results, resulting in inaccurate segmentation results. The present invention discovers that humans will use context reasoning to infer ambiguous problems and make final decisions during observation. In the image segmentation task, context information has extremely strong spatial constraints. Therefore, in the image segmentation task, by reasoning about context feature information, some ambiguous regions in the prediction results can be removed. The present invention refines the initial segmentation results through a context exploration module, discovers false positive regions and false negative regions and eliminates them to obtain a more accurate segmentation result. The structure diagram of the context exploration module is as shown in Figure 3 shown. The high-level prediction map is upsampled and normalized and multiplied by the current level features to obtain the foreground attention features . At the same time, the upsampled and normalized high-level prediction map is inverted and multiplied by the previous level features to obtain the background attention features . and are respectively fed into two parallel exploration modules (Explore Block, EB) to detect the false positive regions and false negative regions in the prediction results;
[0025] After detecting the false positive regions and false negative regions , the present invention eliminates false positives and false negatives in the following manner. This process is represented by formulas as shown in Eqs. (6-8):
[0026] (6)
[0027] (7)
[0028] (8)
[0029] where is the upper-level input feature, is the output feature after eliminating false positives and false negatives, C , B , R respectively represent convolution, normalization, and ReLU, U represents upsampling, and are both learnable scaling parameters; and subtract element-wise to eliminate false positives, and add element-wise to eliminate false negatives;
[0030] As mentioned above, there are two parallel exploration modules, and each exploration module consists of 4 branches with similar structures. The convolution of is used for channel reduction. The convolution of is used for local feature extraction ( , , , ). The dilation convolution with a kernel size of and a dilation rate of is used for context awareness ( , , , ). After each convolution, normalization (BN) and ReLU non-linear operations are performed. After channel reduction, local feature extraction, and context awareness are performed on each branch, the output is input to the next branch for further processing. Finally, the output results of the 4 branches are stacked and fused in the channel dimension. The structure of the exploration module can enrich the context exploration ability to discover false positives and false negatives in the prediction results.
[0031] 1.4 Loss Function
[0032] The present invention also supervises the mapping graph output by the attention localization module and the outputs of the three context exploration modules , , . The loss function of each image is expressed as shown in Equation (9):
[0033] (9)
[0034] where GT is the ground truth map of the infected area segmentation, represents upsampling the image to the same size as the ground truth map GT;
[0035] In addition, the present invention uses to supervise the output of the edge attention module. Therefore, the total loss function is expressed by the formula as shown in Equation (10):
[0036] (10).
[0037] Beneficial Effects Obtained by the Present Invention:
[0038] The segmentation network Ref-Net disclosed by the present invention uses Res2Net as the backbone network, and adds an edge attention module (EAM), an attention localization module (APM), and a context exploration module (CEM) to refine the segmentation results. Without using any auxiliary optimization means, it performs superiorly compared with the latest networks in various metrics, and the mean absolute error ( MAE), which is increased by more than 26.8%, and the Dice similarity coefficient, sensitivity (Sen), specificity (Spec), structure index ( ), and enhancement-alignment index ( ) are all improved simultaneously, which overall improves the accuracy of the detection results and promotes the effective screening of infected persons. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is the structural diagram of the Ref-Net network;
[0040] Figure 2 is the structural diagram of the attention localization module;
[0041] Figure 3 is the structural diagram of the context exploration module;
[0042] Figure 4 is the visual comparison diagram of the segmentation results of the infected area;
[0043] Figure 5 is the PR curve (left) and the weighted F-measure (right). DETAILED DESCRIPTION OF THE INVENTION
[0044] The present invention will be further described in detail below in conjunction with specific embodiments. Embodiment
[0045] Experimental settings This model is implemented under the PyTorch framework, and the hardware configuration is: CPU, Intel i7-4790; GPU, NVIDIA GTX TITAN X (video memory, 12G); memory, 32G; the input image size during training is , and the optimizer Adam is used for training, and the learning rate is set to ; the batch size during training is 2, and the number of training iterations is 100;
[0046] Dataset The dataset used includes 100 labeled CT images from the COVID-19 CT segmentation dataset. All CT images are from more than 40 COVID-19 patients, collected by the Italian Society of Medical and Interventional Radiology and segmented and labeled by radiologists; this is the first open-access COVID-19 dataset for lung infection segmentation, but the scale is small; we divide the 100 labeled CT images into training, validation, and test datasets, including 45 randomly selected CT images as training samples, 5 CT images for validation, and the remaining 50 images for testing; the resolutions of the CT slices are not consistent, and before training, we uniformly adjust the resolutions of all CT slices to 352×352;
[0047] Evaluation metrics: Three widely used metrics in the field of medical images are used to evaluate the segmentation results: Dice similarity coefficient, sensitivity (Sen), and specificity (Spec). In addition, three commonly used metrics in the field of object detection are introduced: structural metric ( ), enhanced alignment metric ( ), and mean absolute error ( MAE ); during evaluation, the upsampling of generated after passing through three context exploration modules is used as the final prediction map , and its similarity is compared with the ground truth map GT ( );
[0048] Dice similarity coefficient (DSC): It is a similarity evaluation metric used to calculate the similarity between two samples. Its value is 1 when the segmentation result is the best and 0 when it is the worst. Its calculation method is shown in Equation (11):
[0049] (11)
[0050] Among them, TP represents the part that is predicted as the infected area and is actually the infected area; TN represents the area that is predicted as the background and is actually the background; FP represents the part that is predicted as the infected area but is actually the non-infected area; FN represents the area that is predicted as the background but is actually the infected area; the same applies hereinafter;
[0051] Sensitivity (Sen): It is used to calculate the proportion of the predicted infected area in all infected areas. Its calculation method is expressed by the formula as shown in Equation (12):
[0052] (12)
[0053] Specificity (Spec): It is used to calculate the proportion of the identified background in all backgrounds. Its calculation method is expressed by the formula as shown in Equation (13):
[0054] (13)
[0055] Structural metric ( ): It is used to measure the structural similarity between the prediction map ( ) and the ground truth map ( ). Its calculation method is expressed by the formula as shown in Equation (14):
[0056] (14)
[0057] Among them, The balance coefficient is default set to 0.5 in the present invention, represents the object-level similarity, represents the region-level similarity;
[0058] Enhanced alignment metric ( ) is used to simultaneously measure the local and global similarities between and , and its formula is shown as Equation (15):
[0059] (15)
[0060] Among them, w and h represent the width and height of the ground truth map respectively, represents the pixel coordinates, represents the enhanced alignment matrix, and the mean of all in the experiment is reported, denoted as .
[0061] Mean absolute error ( MAE ) is used to measure the error between and at the pixel level, and its calculation method is shown as Equation (16):
[0062] (16).
[0063] Comparative experiments
[0064] In the experiment on the COVID-19 infection area, the present invention compared some classic models in the medical field: U-Net, U-Net++, Attention U-Net, Gated-UNet, Dense-UNet and the latest model: Inf-Net; they were evaluated on 6 gold evaluation metrics, where the smaller the value of the mean absolute error ( MAE ), the better the effect, and the larger the values of other metrics, the better the effect; the COVID-19 infection area segmentation results are shown in Table 1. It can be seen that the network (Ref-Net) proposed by the present invention performs excellently. The MAE of the current latest COVID-19 segmentation model Inf-Net is stable at about 0.09 - 0.08. Compared with the segmentation model Inf-Net, the mean absolute error ( MAE ) of the present invention has increased by more than 26.8%. Other evaluation metrics include Dice similarity coefficient, sensitivity (Sen), specificity (Spec), structure metric ( ), enhanced-alignment metric ( They have all been improved simultaneously.
[0065] Table 1
[0066]
[0067] To more intuitively demonstrate the superiority of the present invention, 10 pictures were extracted from the segmentation test results of the present invention for comparison with the segmentation results of the classical networks U-Net, U-Net++, and the latest network Inf-Net, showing a visual comparison diagram of the segmentation results of the lung infection area, as Figure 1 shown. It can be seen from Figure 1 that the segmentation result of the present invention is clearer than that of Inf-Net, and it can be seen that the segmentation result of the present invention removes the ambiguous areas in the prediction result of Inf-Net to a certain extent; that is to say, in the segmentation result of Inf-Net, some uninfected areas are segmented out (false positives) while some infected areas are not detected (false negatives), while the present invention eliminates these false positive areas and false negative areas to a certain extent, and the segmentation result is more accurate. The index of the mean absolute error (MAE) has been significantly improved. This is because the present invention pays attention to edge information and eliminates false positives and false negatives through 3 context exploration modules to explore the context from the high level to the low level. Therefore, it has a clearer boundary and is significantly superior to other network models. Among them, the clear boundary is because the low-level edge attention module constrains the boundary, and the accurate segmentation result is attributed to the attention localization module to locate the target from the global scope, and the context exploration module gradually removes the false negatives and false positives in the prediction result.
[0068] Figure 2 (Left) is the PR (Precision-Recall) curve of the present invention and U-Net, U-Net++, and Inf-Net. In the PR curve, P refers to precision and R refers to recall. The PR curve represents the relationship between precision and recall. After plotting the PR curves of the data of different networks, if the PR curve of a network X encloses the PR curve of another network Y, it can be asserted that the performance of network X is better than Y. It can be seen from Figure 2 (Left) that the curve of the present invention encloses other algorithms, so it can be judged that the present invention is superior to other algorithms. Figure 2 (Right) is the evaluation result of another evaluation index, the weighted F-measure method. The weighted F-measure method comprehensively considers by taking the weighted harmonic mean of precision and recall. This is because precision and recall sometimes conflict, and the most commonly used evaluation method at this time is the weighted F-measure method. Therefore, the present invention Figure 2 simultaneously plots the PR curve and the F-measure curve to more comprehensively prove the superiority of the performance of the present invention.
[0069] Ablation experiment
[0070] Based on the original model (Ref-Net) of the present invention, an attempt is made to add the attention localization module of the feature layer to the feature layer , , respectively, to seek the best position of the attention localization module; in this group of ablation experiments, taking the addition of the attention localization module in as an example, the specific modification details are as follows: without changing the position and number of other modules, delete the attention localization module after the feature layer , and use the feature layer and the prediction map generated by it as the input of the first context exploration module, replacing the feature layer and prediction after passing through the attention localization module in the original model ; next, input into the attention localization module, and use the obtained output as the low-level input of the first context exploration; according to the above similar process, add the attention localization module to , respectively; the results evaluated by 6 indicators are shown in Table 2; it can be seen from Table 2 that adding the attention localization module (APM) at a higher layer has the best effect, because the higher layer of the convolutional neural network is rich in the global localization semantic information of the target object, and can locate the target object more comprehensively and accurately.
[0071] Table 2
[0072]
[0073] The algorithm of the present invention has good portability and is applicable to various current popular backbone networks. Replace Inf-Net and the algorithm Ref-Net of the present invention with other backbone networks respectively, and train with the same data set. The results are shown in Table 3; it can be seen from Table 3 that using VGG-Net, ResNet, and Res2Net as the backbone networks to apply the network structure models Inf-Net and Ref-Net respectively, by comparing the results of 6 common evaluation indicators, it can be concluded that even if different backbone networks are used, the algorithm of the present invention still performs superiorly compared to Inf-Net.
[0074] Table 3
[0075]
[0076] In this article, specific examples are used to elaborate on the principles and implementation modes of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be noted that for those of ordinary skill in the art, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for semantic segmentation of COVID-19 CT images based on the segmentation network Ref-Net, characterized in that, The segmentation network Ref-Net uses Res2Net as the backbone network to extract multi-level features , , , , , retain the edge information of the target object based on the underlying features, and add an edge attention module after to generate an edge feature map , and use the standard binary cross-entropy loss function to calculate its difference from the ground-truth edge feature map generated by the ground-truth image ; Based on the global positioning semantic information rich in target objects at the high level, the input is fed into the attention localization module composed of channel attention and spatial attention to obtain a blurred prediction map ; In , , the context exploration module CEM is added to fuse the current feature and the upper-layer feature respectively to eliminate the wrong predictions in the blurred prediction map, where the feature obtained after passing through the attention localization module is used as the high-level feature of the first context exploration module CEM, and the clear and accurate prediction map is finally obtained after being refined by 3 context exploration modules CEM; among them, The edge attention module consists of a convolutional layer with only one convolutional kernel, and an edge feature map is obtained after passing through this convolutional layer , in order to compare the edge feature map of the edge feature map and the ground truth image The difference between them, the standard binary cross-entropy loss function is introduced; the calculation of the difference loss function is shown in Equations (1) and (2): (1) (2) where w and h represent the width and height of the feature map respectively, is the ground truth edge feature map derived from the ground truth map GT, is the edge feature map predicted by the edge attention module; where represents the weight, which is set to 1 in the present invention; the two parts of and provide global supervision and local supervision respectively, and the segmentation result is more accurate; The attention localization module consists of channel attention and spatial attention, and its specific structure is as follows: the output of the deepest layer is used as the input feature F of the attention localization module. The number of channels, height, and width of the feature F are denoted as C, H, and W respectively. First, the input feature F is reshaped to obtain the query (Q), key (K), and value (V) respectively, where , N refers to the number of pixels. Then, after using matrix multiplication between the transposed Q and K and passing through the Softmax layer, the channel attention map is obtained. The influence of the j-th channel on the i-th channel in the channel attention map X is specifically expressed as shown in Equation (3): (3) Among them, represents the Q th i row of the matrix ; K represents the j th row of the matrix X and V After performing matrix multiplication between the transpose of, the resulting feature shape size is converted to , the attention localization module also introduces a scale parameter , The initial value of is 1 and it continuously learns and updates the weights, and finally the final channel attention output feature is obtained through a skip connection , and any row of it can be expressed by a formula, as shown in Equation (4): (4) Among them, represents the i-th row of the channel attention output feature , represents the value V of the j row of the matrix, and represents the i-th row of the channel attention input feature; The specific process of spatial attention is similar to that of channel attention. As the input of spatial attention, it goes through 3 convolutions and changes the shape of the convolution result to obtain new query ( ), key ( ), and value ( ), , while ; Similarly, transpose of and perform matrix multiplication and normalization to obtain the spatial attention map j ; The influence of the i th position on the th position in spatial attention is calculated as shown in Equation (5): (5) Among them, represents the query 's i column, represents the key 's j column; The following process is similar to channel attention. and After performing matrix multiplication on the transpose of and then introducing a proportionality parameter ; The specific structure of the Context Exploration Module (CEM) is as follows: The high-level prediction map is upsampled and normalized, and then multiplied by the features at the current level to obtain the foreground attention features. Meanwhile, the upsampled and normalized high-level prediction map is inverted and multiplied by the features at the previous level to obtain the background attention features. , and are respectively fed into two parallel exploration modules to detect the false positive regions and false negative regions in the prediction results; The exploration module consists of 4 branches, and the structure of each branch is similar. Convolutions with are used for channel reduction, Convolutions with are used for local feature extraction, and dilated convolutions with a kernel size of and a dilation rate of are used for context awareness. After each convolution, normalization and ReLU non-linear operations are performed; After each branch performs channel reduction, local feature extraction, and context awareness, the output is input into the next branch for further processing. Finally, the output results of the 4 branches are stacked and fused in the channel dimension. False positive regions are detected and false negative regions After that, false positives and false negatives are eliminated by the following method, and this process is expressed by a formula as shown in Equation (6-8): (6) (7) (8) Among them, is the upper-level input feature, is the output feature after false positive and false negative elimination, C , B , R represent convolution, normalization, and ReLU respectively, U represents upsampling, and are both learnable scale parameters, and subtract element-wise to eliminate false positives, and add element-wise to eliminate false negatives.
2. The method according to claim 1, characterized in that: The mapping graph output by the attention localization module and the output graphs of the three Context Exploration Modules (CEMs) , , have a loss function as shown in Equation (9): (9) Among them, GT is the ground truth map for infected area segmentation, representing the upsampling of the image to the same size as the ground truth map GT.
3. The method according to claim 2, wherein: Adopt Supervise the output of the supervised edge attention module, and the total loss function is expressed by the formula as shown in Equation (10): (10)。
Citation Information
Patent Citations
New coronal pneumonia CT image infected area segmentation method and system based on improved CE-Net
CN114332133A
Remote sensing image sea-land segmentation method
CN114663439A