An Intestinal Polyp Detection Method Based on Deep Supervision and Progressive Learning
By adopting deep supervision and step-by-step learning methods in medical image detection, combined with the characteristics of Transformer and convolutional layer, the problem of insufficient image feature capture and model generalization capabilities in the prior art is solved, and more accurate and accurate intestinal polyp detection is achieved.
Patent Information
- Application Number
- CN202211007876.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-08-22
AI Technical Summary
Existing medical image detection methods have shortcomings in capturing image features and improving model generalization capabilities, especially when identifying inconspicuous polyps, lack global information capture capabilities and boundary accuracy.
Using a method based on deep supervision and step-by-step learning, features are extracted using the PVT variant in Transformer, and multi-scale detailed information is captured through convolutional layers, deep supervised learning is carried out layer by layer, and features of each layer are fused to obtain accurate detection results.
Through deep supervision and step-by-step learning methods, the accuracy and boundary accuracy of intestinal polyps detection are significantly improved, the generalization ability of the model is enhanced, and the insignificant polyps can be more effectively identified.
Smart Images

Figure CN115331024B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to a method for detecting intestinal polyps based on deep supervision and step-by-step learning. Background Art
[0002] Medical image detection is an important part of artificial intelligence-assisted diagnosis, which can provide doctors with some detailed information to assist in diagnosis. For the common cancer colon cancer, early detection and removal of polyps are effective means to prevent cancer attacks. Detecting polyps in colonoscopy-captured images is of great significance for preventing colon cancer. Recently, great progress has been made in image detection of natural images. In contrast, the detection problem in medical images still faces huge challenges. Since the dataset of medical images is generally small and the shapes of detection targets vary greatly, it is difficult to directly apply the detection methods of natural images to medical image detection. Therefore, how to accurately capture image features and improve the generalization ability of the model is crucial for further exploration of medical image detection.
[0003] Recently, medical image detection methods based on convolutional neural networks (CNNs) have achieved good performance in many datasets. The most representative method among them is U-Net, which captures context information well through skip connections. However, due to the top-down modeling method of the CNN model and the variability of polyp morphology, these models lack the ability to capture global information and generalization ability, and often fail to recognize some unobvious polyps. Xie et al. proposed SegFormer in 2021, applying Transformer to the field of image detection and proposing a multi-stage feature aggregation multi-branch decoder, which predicts features of different scales and depths by simple upsampling and then parallel fusion. CaraNet proposed by Ange et al. uses reverse attention to extract the detailed information of small objects and then models the global relationship through Transformer. CaraNet is very accurate in detecting small objects and has set a new record in medical image detection tasks. These Transformer-based methods grasp the main body of detection well, but there is still a lack in the processing of low-level texture information, resulting in inaccurate boundaries of detection results. Summary of the Invention
[0004] The present invention aims to overcome the shortcomings of the prior art and provides a method for detecting intestinal polyps based on deep supervision and step-by-step learning. Features are extracted through the variant PVT in Transformer, multi-scale detailed information is captured by convolutional layers, and learning is carried out layer by layer in a deep supervision manner, and the features of each layer are gradually fused to obtain accurate detection results.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A method for detecting intestinal polyps based on deep supervision and step-by-step learning, comprising:
[0007] Input an intestinal polyp image of 352×352×3 to be detected, and use PVT_V2 to extract features from the colonoscopy image, extracting four-scale features, and the four scales are 88×88×64, 44×44×128, 22×22×320, and 11×11×512 respectively;
[0008] Input the features of the four scales extracted into the detail enhancement module, and output the first to fourth enhanced features after detail enhancement and channel compression to 64;
[0009] Input the first, second, and third enhanced features after detail enhancement and the second, third, and fourth enhanced features in pairs into the guidance fusion module, and output the first to third fusion features after fusion;
[0010] Input the first to third fusion features and the fourth enhanced feature into the first to fourth layer multi-branch decoders respectively. The first to fourth layer multi-branch decoders are connected in sequence, and the output of the latter layer multi-branch decoder is used as the input of its previous layer multi-branch decoder at the same time, and the first to fourth decoded features after decoding by the multi-branch decoder are obtained;
[0011] Pass the first to fourth decoded features through a 3×3 convolution respectively to obtain four detection results with 1 channel number, and use the detection result corresponding to the first decoded feature as the final detection result.
[0012] Further, the detail enhancement module performs the following operations:
[0013] S21. Take any scale feature extracted Through a layer of 1×1 convolution, compress it to 64 channels and maintain the original spatial scale, remove redundant channel information in the detection task, and output a scale of H i ×W i ×64, H i 、W i Are the height and width of the feature respectively;
[0014] S22. Pass the result of S21 through 4 convolution kernels of 1×1, 3×3, 5×5, and 7×7 respectively to obtain four features that capture different scale information The scale is all H i ×W i ×64;
[0015] S23. Concatenate the results of S22 in the channel dimension to obtain a scale of Hi ×W i ×256 fused features
[0016] S24. The obtained features Through two layers of 3×3 convolutions, fuse the features that capture information at different scales to generate enhanced features Its scale is H i ×W i ×64
[0017] Furthermore, the guidance fusion module performs the following operations:
[0018] S31. For the four enhanced features extracted Input them into the guidance fusion module in the corresponding relationship of ;
[0019] S32. Upsample using bilinear interpolation to obtain features with the same spatial dimension as ;
[0020] S33. Process the upsampled features through spatial attention to obtain the attention weight smap i+1 The formula is as follows:
[0021]
[0022] where SA(·) is spatial attention;
[0023] S34. Let the feature and smap i+1 perform element-wise multiplication to highlight the features in the significant regions. The formula is as follows:
[0024]
[0025] where is element-wise multiplication;
[0026] S35. Connect and with a residual connection to retain the information of the low-level features and improve the training stability. The formula is as follows:
[0027]
[0028] S36. Concatenate and fuse and in the channel dimension to obtain a result with a scale of H i ×W i ×128
[0029] S37. The obtained features are fused with the features capturing different-scale information through a 3×3 convolution layer, and the fused features are output with a scale of H i ×W i ×64.
[0030] Furthermore, the fourth-layer multi-branch decoder performs the following operations:
[0031] S411. The fourth enhanced feature is input into a 1×1 convolution to further learn the information on different channels, and a result with a scale of 11×11×64 is obtained;
[0032] S412. The result of S41 is respectively passed through 4 convolutional kernels of 1×1, 3×3, and 5×5 to obtain 3 features capturing different-scale information The scales of the three features are all H i ×W i ×64;
[0033] S413. The three results of S42 are concatenated in the channel dimension to obtain a fused feature with a scale of H i ×W i ×192
[0034] S414. The obtained features are fused with the features capturing different-scale information through two 3×3 convolution layers to generate decoded features with a scale of H i ×W i ×64;
[0035] The decoding processes of the first to third-layer multi-branch decoders are as follows:
[0036] S421. The fused feature and the decoded feature output by the previous multi-branch decoder are concatenated in the channel dimension to obtain a fused feature with a scale of H i ×W i ×64
[0037] S422. is input into a 1×1 convolution, and the result of fusing the features of this layer and the upper layer is obtained with a scale of H i ×W i ×64
[0038] S423. Three features capturing different-scale information are obtained respectively through three convolutional kernels of 1×1, 3×3, and 5×5 The scales of the three features are all H i ×W i ×64;
[0039] S424. Concatenate the features on the channel dimension to obtain a fused feature with a scale of H i ×W i ×192
[0040] S425. Pass the obtained features through two layers of 3×3 convolutions to fuse the features capturing different-scale information and generate decoded features with a scale of H i ×W i ×64
[0041] The intestinal polyp detection method based on deep supervision and step-by-step learning provided by this application uses deep supervision to perform layer-by-layer learning on the features extracted by PVT_V2. Capture detailed information and remove redundant channel information through detail enhancement, use the guidance fusion module to gradually fuse high-semantic information and low-semantic information, and let the high-level learning results guide the low-level learning. And perform detection through a multi-branch decoder to obtain more accurate intestinal polyp detection results Brief Description of the Drawings
[0042] Figure 1 is a flowchart of the intestinal polyp detection method based on deep supervision and step-by-step learning of this application;
[0043] Figure 2 is the overall architecture diagram of the network model of this application;
[0044] Figure 3 is a schematic diagram of the structure of the detail enhancement module in the embodiment of this application;
[0045] Figure 4 is a schematic diagram of the structure of the guidance fusion module of this application;
[0046] Figure 5 is a schematic diagram of the structure of the multi-branch decoder module of this application;
[0047] Figure 6 is a schematic diagram of the structure of the spatial attention SA module of this application Detailed Embodiments
[0048] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following further details this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0049] In one embodiment, a method for detecting intestinal polyps based on deep supervision and step-by-step learning is provided, which fully utilizes the global dependence capture ability of the Transformer and the detail capture ability of the CNN to achieve accurate detection of intestinal polyp images.
[0050] Specifically, as Figure 1 shown, the method for detecting intestinal polyps based on deep supervision and step-by-step learning in this embodiment includes:
[0051] Step S1: Input an intestinal polyp image of 352×352×3 to be detected, and use PVT_V2 to extract features from the colonoscopy image, extracting four-scale features, and the four scales are 88×88×64, 44×44×128, 22×22×320, and 11×11×512 respectively.
[0052] First, obtain the intestinal polyp image to be detected, and then scale it to 352×352×3 as the input image for subsequent processing.
[0053] In this example, in order to better utilize the self-attention mechanism of the Transformer to better capture the global dependence relationship in the image, the PVT_V2 backbone network is used to extract features from the image. Using PVT_V2 to extract features from the input intestinal polyp image of 353×352×3 aims to extract features of different scales. The receptive field of the high-level network is relatively large, and its semantic information representation ability is strong, which can accurately locate the target position; the receptive field of the low-level network is relatively small, and its geometric detail information representation ability is strong, which helps to complement the boundary detail information.
[0054] After feature extraction by PVT_V2, the four-scale features obtained are 88×88×64, 44×44×128, 22×22×320, and 11×11×512 respectively, corresponding to the outputs of PVT1, PVT2, PVT3, and PVT4 in Figure 2 respectively.
[0055] Step S2: Input the four-scale features extracted into the detail enhancement module, and output the first to fourth enhanced features after detail enhancement and channel compression to 64.
[0056] In this example, as Figure 2 shown, for the four different-scale feature outputs First, channel compression is performed to remove redundant channel information and improve the model's calculation speed. Then, detail features of different scales are extracted through four convolutional kernels of different sizes. Next, these features of different scales are concatenated in the channel dimension, and the information of each scale is fused through two layers of 3×3 convolutional kernels to reduce the number of channels.
[0057] The features of four scales are obtained from the colonoscopy images through feature extraction by PVT_V2. Their scales are 88×88×64, 44×44×128, 22×22×320, and 11×11×512 respectively, and they are respectively input into the detail enhancement module. In this embodiment, the detail enhancement module is as Figure 3 shown, and the process is as follows:
[0058] S21: Any feature f i o obtained through extraction is compressed to 64 channels through a layer of 1×1 convolution while maintaining the original spatial scale, removing redundant channel information in the detection task, and the output scale is H i ×W i ×64, where H i , H i are the height and width of the feature f i o respectively.
[0059] S22: The result of S21 is respectively passed through four convolutional kernels of 1×1, 3×3, 5×5, and 7×7 to obtain four features capturing different scale information with the scale of H i ×W i ×64.
[0060] S23: The results of S22 are concatenated in the channel dimension to obtain a fused feature with the scale of H i ×W i ×256
[0061] S24: The obtained feature is passed through two layers of 3×3 convolution to fuse the features capturing different scale information and generate an enhanced feature with the scale of H i ×W i ×64.
[0062] Step S3: The first, second, and third enhanced features after detail enhancement and the second, third, and fourth enhanced features are paired and input into the guidance fusion module, and the first to third fusion features after fusion are output.
[0063] In this example, as Figure 4 shown, for the input features and After upsampling, a spatial attention map smap is generated through the SA module i+1 Multiply the spatial attention map element - by - element with the low - level features to obtain features highlighting significant regions And make a skip connection. Concatenate the result with and fuse them using a 3×3 convolutional layer to obtain the output
[0064] In this embodiment, the guidance fusion module process is as follows:
[0065] S31. For the four enhanced features extracted Input them into the guidance fusion module in the corresponding relationship of
[0066] S32. Upsample using bilinear interpolation to obtain features with the same spatial dimension as
[0067] S33. Process the upsampled features through spatial attention to obtain attention weights, denoted by smap i+1 The calculation formula is as follows:
[0068]
[0069] where SA(·) is spatial attention, and the SA module structure is as Figure 6 shown
[0070] S34. Let the feature and smap i+1 perform element - by - element multiplication to highlight the features of significant regions. The calculation formula is as follows:
[0071]
[0072] where is element - by - element multiplication
[0073] S35. Perform a residual connection between and to retain the information of low - level features and improve training stability. The calculation formula is as follows:
[0074]
[0075] S36. Concatenate with Concatenate and fuse in the channel dimension to obtain a result with dimensions H i ×W i ×128
[0076] S37. Feed the obtained features through a 3×3 convolution layer to fuse the features capturing information at different scales and output the fused features with dimensions H i ×W i ×64
[0077] Step S4. Feed the first to third fused features and the fourth enhanced feature into the first to fourth layer multi-branch decoders respectively. The multi-branch decoders of each layer are connected in sequence, and the output of the latter layer multi-branch decoder is used as the input of its previous layer multi-branch decoder at the same time, to obtain the first to fourth decoded features after decoding by the multi-branch decoders
[0078] In this example, the first to third fused features and the fourth enhanced feature are fed into their respective corresponding multi-branch decoders. The fourth enhanced feature is fed into the fourth layer multi-branch decoder, and the first to third fused features are fed into the first to third layer multi-branch decoders in sequence
[0079] As Figure 5 shown, for the input features, the multi-branch decoder first uses a 1×1 convolution to further learn the information on different channels, then extracts the information at different scales through three different convolution branches, and finally concatenates and fuses to obtain the final result
[0080] In this embodiment, the fourth enhanced feature with dimensions 11×11×64 is fed into the fourth layer multi-branch decoder, and the decoding process is as follows
[0081] S411. Feed the fourth enhanced feature into a 1×1 convolution to further learn the information on different channels and obtain a result with dimensions 11×11×64
[0082] S412. Pass the result of S41 through 4 convolutional kernels of 1×1, 3×3, and 5×5 respectively to obtain 3 features capturing information at different scales The dimensions of the three features are all H i ×W i ×64
[0083] S413. Concatenate the three results of S42 in the channel dimension to obtain a fused feature with dimensions H i ×W i ×192
[0084] S414. The obtained features Through two layers of 3×3 convolutions, the features capturing different scale information are fused to generate the decoded features whose scale is H i ×W i ×64.
[0085] In this embodiment, for the first to third layer multi-branch decoders, the input features are the fused features and the decoded features output by the previous multi-branch decoder First, they are concatenated in the channel dimension and then fused into Then, different scale information is extracted through three different convolution branches, and they are concatenated and fused again to obtain the final result
[0086] In this embodiment, for the first to third layer multi-branch decoders, the decoding process is as follows:
[0087] S421. The fused features and the decoded features output by the previous multi-branch decoder are concatenated in the channel dimension to obtain the fused features with a scale of H i ×W i ×64
[0088] In this embodiment, the output of the previous multi-branch decoder is upsampled by bilinear interpolation to obtain features with the same spatial dimension as Then and and are concatenated in the channel dimension to obtain the fused features with a scale of H i ×W i ×64
[0089] S422. Input into a 1×1 convolution, and the result of fusing the features of this layer and the upper layer features is obtained with a scale of H i ×W i ×64
[0090] S423. Pass through 3 convolutional kernels of 1×1, 3×3, and 5×5 respectively to obtain 3 features capturing different scale information The scales of the three features are all H i ×W i ×64.
[0091] S424. The features Concatenate on the channel dimension to obtain a fused feature with a scale of H i ×W i ×192
[0092] S425. The obtained features Pass through two layers of 3×3 convolutions to fuse features capturing different scale information and generate decoded features with a scale of H i ×W i ×64.
[0093] Step S5. Pass the first to fourth decoded features through a 3×3 convolution respectively to obtain four detection results with 1 channel number, and use the detection result corresponding to the first decoded feature as the final detection result.
[0094] In this step, the decoded features are passed through a 3×3 convolution respectively to obtain four detection results with 1 channel number.
[0095] During training, the detection results are also upsampled to the size of the original image by interpolation method, calculate the loss function and perform backpropagation to complete the training of the entire network model. After training the network model, use the trained network model to detect the input intestinal polyp image and output the detection result.
[0096] In this example, BCE loss and IOU loss are used to calculate the loss between the final salient object detection result and the true label.
[0097] In this example, binary cross-entropy (BCE) is used to calculate the gap between the true label and the detection result. BCE is a widely used loss in classification, and the calculation formula is as follows:
[0098]
[0099] IOU loss is mainly used to measure the overall similarity between two images, and the calculation formula is as follows:
[0100]
[0101] where g(x,y)∈[0,1] is the true label of the detection image, and p(x,y)∈[0,1] is the detection result of the model for the detection image.
[0102] When using the trained model, only the output result of the multi-branch decoder in the first layer is used. The number of channels is reduced to 1 by using a 3×3 convolution to obtain the probability value that each pixel is a polyp target. Pixels with a probability value greater than or equal to 0.5 are labeled as white pixels of polyp targets, and pixels with a probability value less than or equal to 0.5 are labeled as black pixels that are not polyp targets, obtaining the final detection result, that is, a black-and-white image with polyp targets labeled by white pixels.
[0103] In this example, the interactive encoder is used to fuse the main body features and edge features, and then fed back to the main body encoder and edge encoder for secondary iteration. The output of the secondary iteration will have clearer edge features and be more in line with the actual labels.
[0104] This embodiment uses a multi-branch fusion network to separately perform multi-scale extraction and fusion of features for the main body and the edge, which is beneficial to the edge characterization of significant targets. The method of label decoupling is introduced in the example. This method decouples the labels of intestinal polyp images. The distance transformation method is used to decouple the original labels into main body labels and edge labels. The decoupled labels are beneficial to the supervision and evaluation of the model.
[0105] This embodiment designs a detail enhancement module, a guidance fusion module, and a multi-branch decoding module. On the basis of using the Transformer backbone network to extract features, a convolutional neural network is used for local information enhancement and feature fusion. Deep supervision enables the feature fusion results of each layer to be learned, and the final result is gradually fused to be clear and accurate. On the basis of the self-attention mechanism of the Transformer accurately locating the detection area, a convolutional neural network is used to capture detail information and fuse it, making full use of the advantages of both to obtain a clear and accurate result.
[0106] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for detecting intestinal polyps based on deep supervision and step-by-step learning, characterized in that, The intestinal polyp detection method based on deep supervision and step-by-step learning includes: Input the intestinal polyp image of 352×352×3 to be detected, and use PVT_V2 to extract features from the colonoscopy images, extracting four-scale features, and the four scales are 88×88×64, 44×44×128, 22×22×320, and 11×11×512 respectively; Input the features of the four scales extracted into the detail enhancement module, and output the first to fourth enhanced features after detail enhancement and channel compression to 64; Pairwise input the first, second, and third enhanced features after detail enhancement and the second, third, and fourth enhanced features into the guidance fusion module, and output the first to third fusion features after fusion; Input the first to third fusion features and the fourth enhanced feature into the first to fourth layer multi-branch decoders respectively. The first to fourth layer multi-branch decoders are connected in sequence, and the output of the subsequent layer multi-branch decoder is used as the input of its previous layer multi-branch decoder at the same time, and obtain the first to fourth decoded features after decoding by the multi-branch decoder; Respectively pass the first to fourth decoded features through a 3×3 convolution to obtain four detection results with the number of channels being 1, and use the detection result corresponding to the first decoded feature as the final detection result; Among them, the detail enhancement module performs the following operations: S21. Take any scale feature f obtained by extraction i o Through a 1×1 convolution layer, compress it to 64 channels while maintaining the original spatial scale, remove redundant channel information in the detection task, and output a scale of H i ×W i ×64, where H i and W i are the height and width of the feature f i o respectively; S22. Respectively pass the result of S21 through four convolution kernels of 1×1, 3×3, 5×5, and 7×7 to obtain four features that capture different scale information The scale is all H i ×W i ×64; S23. Concatenate the result of S22 along the channel dimension to obtain a fused feature f with a scale of H i ×W i ×256 i decat ; S24. Obtain the feature f i decat Through two layers of 3×3 convolutions, fuse the features that capture information at different scales to generate the enhanced feature f i de , whose scale is H i ×W i ×64; The guidance fusion module performs the following operations: S31. For the four enhanced features extracted input them into the guidance fusion module in the correspondence of . S32. Upsample using bilinear interpolation to obtain a feature with the same spatial dimension as f i de S33. The upsampled features are processed through spatial attention to obtain the attention weight smap i+1 which is expressed by the following calculation formula: Among them, SA(·) is spatial attention; S34. Let the feature f i de and smap i+1 perform element-wise multiplication to highlight the features of significant regions. The calculation formula is as follows: Among them, is element multiplication; S35. Connect f i de to f i sa with a residual connection to retain the information of low-level features and improve the training stability. The calculation formula is as follows: f l gf = f i sa + f i de ; S36. Concatenate and fuse f l gf with in the channel dimension to obtain the result f with a scale of H i ×W i ×128 i gf ; S37. Obtain the feature f i gf Through a 3×3 convolution layer, fuse the features capturing information at different scales and output the fused feature f i gfout with a scale of H i ×W i ×64.
2. The intestinal polyp detection method based on deep supervision and step-by-step learning according to claim 1, characterized in that The fourth layer multi-branch decoder performs the following operations: S411. Input the fourth enhanced feature into a 1×1 convolution to further learn the information on different channels and obtain a result with a scale of 11×11×64; S412. Use four convolution kernels of 1×1, 3×3, and 5×5 respectively on the result of S41 to obtain three features that capture different scale information. The scales of the three features are all H i ×W i ×64; S413. Concatenate the three results of S42 in the channel dimension to obtain a fused feature with a scale of H i ×W i ×192 S414. Obtain the features Through two layers of 3×3 convolutions, fuse the features that capture information at different scales to generate decoded features whose scale is H i ×W i ×64; The decoding process of the first to third layer multi-branch decoders is as follows: S421. Concatenate the fused feature f i gfout and the decoded feature output by the previous multi-branch decoder in the channel dimension to obtain a fused feature f with a scale of H i ×W i ×64 i bdin ; S422. Input f i bdin into a 1×1 convolution, and the result of fusing the features of this layer and the upper layer is a result with a scale of H i ×W i ×64, namely f i bdpre ; S423. Obtain three features \(f\) i bdpre respectively through three convolutional kernels of \(1\times1\), \(3\times3\), and \(5\times5\), obtaining three features \(f\) i bd1 , \(f\) i bd2 , \(f\) i bd3 , and the scales of the three features are all \(H\) i × \(W\) i × 64; S424. Concatenate the features f i bd1 , f i bd2 , f i bd3 along the channel dimension to obtain the fused feature f with a scale of H i ×W i ×192; i bdcat ; S425. Obtain the feature f i bdcat Through two layers of 3×3 convolutions, fuse the features that capture information at different scales to generate the decoded feature f i bd with a scale of H i ×W i ×64.