A polyp image segmentation method and system, electronic device, and storage medium
By combining and fusion of multiple target image features, combined with Fourier transform and multi-scale feature denoising decoder, the problem of low segmentation accuracy of traditional polyp images is solved, and the precise extraction of detailed information such as polyp edges is achieved, which improves the accuracy of medical diagnosis.
Patent Information
- Application Number
- CN202411740060.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-29
AI Technical Summary
The traditional polyp image segmentation method has low segmentation accuracy and is difficult to effectively extract detailed information, which limits its application in medical diagnosis.
By combining and fusion of multiple target image features, the Hader code product integrates different types of feature information, generates segmented prediction maps, combines the Fourier transform to extract the key features of polyps, and builds a multi-scale feature denoising decoder to reduce the impact of background noise.
The accuracy of polyp image segmentation is improved, especially the accurate extraction of detailed information such as polyp edges, providing a reliable basis for medical diagnosis.
Smart Images

Figure CN119600295B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of image segmentation, and more specifically, relates to a polyp image segmentation method and system, an electronic device, and a storage medium. Background Art
[0002] In the field of medical image processing, the segmentation of polyp images is a crucial task. As an abnormal proliferation in human tissues, the accurate segmentation of polyps is of great significance for the early detection and treatment of diseases. However, traditional polyp image segmentation methods often have problems such as low segmentation accuracy and insufficient extraction of detailed information, which greatly limits their application in medical diagnosis. Summary of the Invention
[0003] The purpose of the present disclosure is to provide a polyp image segmentation method and system, an electronic device, and a storage medium to improve the accuracy of image segmentation.
[0004] In the first aspect of the embodiments of the present disclosure, a polyp image segmentation method is provided, including:
[0005] Combining multiple target image features pairwise to obtain a first feature, a second feature, and a third feature, where the multiple target image features include a first target image feature, a second target image feature, and a third target image feature, and the multiple target image features are features obtained from multiple image features;
[0006] Fusing the first target image feature and the first feature to obtain a first global feature;
[0007] Performing a Hadamard product on the second feature and the third feature to obtain a fourth feature;
[0008] Fusing the first global feature and the fourth feature to obtain a second global feature;
[0009] Generating a segmentation prediction map based on the second global feature.
[0010] In the second aspect of the embodiments of the present disclosure, a polyp image segmentation device is provided, including:
[0011] A feature combination module for combining multiple target image features pairwise to obtain a first feature, a second feature, and a third feature, where the multiple target image features include a first target image feature, a second target image feature, and a third target image feature, and the multiple target image features are features obtained from multiple image features;
[0012] A first fusion module for fusing the first target image feature and the first feature to obtain a first global feature;
[0013] A feature product module for performing a Hadamard product on the second feature and the third feature to obtain a fourth feature;
[0014] A second fusion module for fusing the first global feature and the fourth feature to obtain a second global feature;
[0015] A segmentation module for generating a segmentation prediction map based on the second global feature.
[0016] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned polyp image segmentation method are implemented.
[0017] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned polyp image segmentation method are implemented.
[0018] The beneficial effects of the polyp image segmentation method, system, electronic device, and storage medium provided by the embodiments of the present disclosure are as follows: By combining multiple target image features pairwise, the embodiments of the present disclosure can fully explore the relationships between features at different levels, thereby obtaining more representative and discriminative first, second, and third features, avoiding the limitations of single features. Secondly, fusing the first target image feature with the first feature to obtain the first global feature, performing a Hadamard product on the second feature and the third feature to obtain the fourth feature, and further fusing to obtain the second global feature can effectively integrate different types of feature information, making the features more rich and comprehensive, and facilitating highlighting the key features of polyps. Finally, generating a segmentation map based on the second global feature helps to improve the segmentation accuracy, especially the accurate extraction of detailed information such as polyp edges, providing a reliable basis for medical diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic flowchart of a polyp image segmentation method provided by an embodiment of the present disclosure;
[0021] Figure 2 It is a schematic diagram for extracting Fourier transform image boundary information provided by an embodiment of the present disclosure;
[0022] Figure 3 Schematic flowchart of another polyp image segmentation method provided by an embodiment of the present disclosure;
[0023] Figure 4 Schematic diagram of the denoising unit structure provided by an embodiment of the present disclosure;
[0024] Figure 5 Schematic diagram of the multi-scale feature denoising decoder structure provided by an embodiment of the present disclosure;
[0025] Figure 6 Schematic diagram of the colorectal polyp image plane provided by an embodiment of the present disclosure;
[0026] Figure 7 Result diagram of medical image segmentation provided by an embodiment of the present disclosure;
[0027] Figure 8 Visualization result diagram of ablation experiment provided by an embodiment of the present disclosure;
[0028] Figure 9 Structure block diagram of a polyp image segmentation device provided by an embodiment of the present disclosure;
[0029] Figure 10 Schematic block diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0030] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0031] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the accompanying drawings.
[0032] Please refer to Figure 1 , Figure 1 Schematic flowchart of a polyp image segmentation method provided by an embodiment of the present disclosure, the method includes:
[0033] S101: Combine multiple target image features in pairs to obtain a first feature, a second feature, and a third feature. The multiple target image features include a first target image feature, a second target image feature, and a third target image feature, and the multiple target image features are features obtained from multiple image features.
[0034] In this embodiment, feature extraction is performed on the public polyp images to obtain multiple image features.
[0035] The public polyp data can be obtained from the datasets Kvasir, CVC-ClinicDB, CVC-300, CVC-ColonDB, and ETIS.
[0036] A convolutional neural network (CNN) is constructed. The CNN in this embodiment can include a CNN encoder composed of five layers of ConvNeXt convolutional encoding networks;
[0037] The obtained public polyp images are subjected to feature extraction through the CNN encoder to obtain five layers of features (i.e., multiple image features). As Figure 3 shown, the multiple image features are respectively the first-layer shallow features in the CNN encoder 、the second-layer shallow features 、the third-layer deep features 、the fourth-layer shallow features 、the fifth-layer deep features .
[0038] Among them, the first-layer shallow features 、the second-layer shallow features 、the third-layer deep features 、the fourth-layer shallow features 、the fifth-layer deep features are respectively target image features.
[0039] As Figure 3 shown, the fifth-layer deep features are the first target image features, the fourth-layer shallow features are the second target image features, and the third-layer deep features are the third target image features.
[0040] In this embodiment, referring to Figure 5 , Figure 5 is a schematic flowchart of a multi-scale feature denoising decoder provided by an embodiment of the present disclosure. The third-layer deep features 、the fourth-layer shallow features 、and the fifth-layer deep features extracted by the CNN encoder are combined in pairs.
[0041] It can be 、 respectively sent into three denoising units. Referring to Figure 4 , Figure 4The flowchart of the denoising unit provided by an embodiment of the present disclosure. After pairwise combination of multiple target image features and entering the denoising unit, the deep features and shallow features are respectively learned through convolutions with convolution kernel sizes of , , ; perform an element-wise subtraction operation on the deep features and shallow features of the corresponding sizes, and then add the three features after subtraction; finally, obtain the first feature, the second feature, and the third feature.
[0042] The first feature, the second feature, and the third feature can be expressed as:
[0043]
[0044] where represents the deep feature, represents the shallow feature.
[0045] S102: Fuse the first target image feature and the first feature to obtain the first global feature.
[0046] In this embodiment, first, the first target image feature ( ) is learned through convolution with a convolution kernel size of to capture more global information. Then, perform an upsampling operation to restore the image size of the first target image feature.
[0047] In this embodiment, the first target image feature ( ) is subjected to channel-level feature fusion with the first feature obtained after denoising processing to increase the richness of the first target image feature, thereby obtaining the first global feature.
[0048] S103: Perform a Hadamard product on the second feature and the third feature to obtain the fourth feature.
[0049] In this embodiment, for the shallow subtracted features (i.e., the second feature and the third feature obtained after denoising processing), the Hadamard product method is used to perform element-wise multiplication on each pixel in the feature, highlighting the features useful for the segmentation task and reducing the influence of background features. After the Hadamard product processing, the fourth feature is obtained.
[0050] S104: Fuse the first global feature and the fourth feature to obtain the second global feature.
[0051] In this embodiment, the first global feature is passed through a convolution with a convolution kernel size of After convolutional learning, more global information is captured; then, an upsampling operation is performed to restore the image size of the first global feature; finally, the first global feature and the fourth feature are fused at the channel level to increase the richness of the features, thereby obtaining the second global feature.
[0052] S105: Generate a segmentation prediction map based on the second global feature.
[0053] In this embodiment, the eigenvalue of the second global feature is repositioned between 0 and 1 through a Sigmoid operation to generate a high-resolution segmentation prediction map.
[0054] In this embodiment, the generated segmentation prediction map can be deeply supervised through a Mask label, which can better guide the model to extract edge information in medical images.
[0055] It can be concluded from the above that in this embodiment, by combining multiple target image features pairwise, the relationship between features at different levels can be fully explored, thereby obtaining more representative and discriminative first, second, and third features, avoiding the limitations of single features. Secondly, fusing the first target image feature with the first feature to obtain the first global feature, and performing a Hadamard product on the second and third features to obtain the fourth feature, and further fusing to obtain the second global feature can effectively integrate different types of feature information, making the features more rich and comprehensive, which is conducive to highlighting the key features of polyps. Finally, generating a segmentation map based on the second global feature helps to improve the accuracy of segmentation, especially the accurate extraction of details such as the edges of polyps, providing a reliable basis for medical diagnosis.
[0056] In an embodiment of the present disclosure, it further includes:
[0057] Perform a Fourier transform on the public polyp image to obtain polyp image features;
[0058] Extract features from the public polyp image to obtain multiple image features;
[0059] Fuse the polyp image features and the first target image features to obtain the first polyp image feature;
[0060] Process the multiple image features to obtain multiple second polyp image features;
[0061] Fuse the first polyp image feature and the multiple second polyp image features to generate a segmentation map.
[0062] In this embodiment, the Fourier transform is performed on the public polyp image to obtain polyp image features. The most primitive public polyp image features can be enhanced in terms of semantic information of the polyp through the spatial domain branch by two identical Fourier Transform Modules (FTMs). Through the frequency domain branch, the image is transformed from the spatial domain to the frequency domain information, and key contents such as edge texture details contained in the high-frequency information are extracted. In medical images, the edges of polyps and their unique texture details can help depict the specific morphology and boundary conditions of polyps. After such Fourier transform processing, polyp image features are obtained.
[0063] In this embodiment, feature extraction is performed on the public polyp image to obtain multiple image features. The CNN encoder can be used to extract features from the obtained public polyp image, resulting in five layers of features (i.e., multiple image features). As Figure 3 shown, the multiple image features are respectively the first-layer shallow features in the CNN encoder , the second-layer shallow features , the third-layer deep features , the fourth-layer shallow features , and the fifth-layer deep features . Among them, the first-layer shallow features , the second-layer shallow features , the third-layer deep features , the fourth-layer shallow features , and the fifth-layer deep features are respectively target image features.
[0064] Each layer of features contains information at different levels and dimensions of the image. The shallow features can more reflect the local and basic appearance features of the image, while the deep features can capture more abstract and semantically stronger information. The multi-layer features together constitute multiple image features, providing a rich data basis for subsequent further processing.
[0065] In this embodiment, the first target image feature ( ) as a deep feature contains more advanced semantic information and can form a good complement with the polyp image features extracted by the Fourier transform that focus on details. The polyp image features and the first target image feature are fused to obtain the first polyp image feature. The polyp image features obtained by the Fourier transform can be fused with the first target image feature ( ).
[0066] The polyp image features obtained by the Fourier transform and the first target image feature ( are fused so that the fused first polyp image features not only possess the detailed advantages such as edge textures obtained from Fourier transform but also incorporate the high-level semantic information contained in the deep features, thereby more comprehensively depicting the characteristics of the polyp.
[0067] In this embodiment, as Figure 3 shown, multiple image features obtained from the CNN encoder (i.e., the first-layer shallow features , the second-layer shallow features , the third-layer deep features , the fourth-layer shallow features ) can be respectively and sequentially subjected to convolution, batch normalization, and activation function (Relu) processing to obtain multiple second polyp image features. Through convolution, the feature information is further mined, batch normalization is used to stabilize the data distribution and accelerate training, and Relu is used to introduce non-linear factors to enhance its expression ability, putting it in a more suitable state.
[0068] In this embodiment, the first polyp image features and multiple second polyp image features are fused to generate a segmentation map.
[0069] First, the first polyp image features are processed by sequentially performing convolution, batch normalization, and activation function (Relu). A multi-layer feature fusion decoder is constructed. The decoder includes four layers of convolution operations, and the features from each layer of the encoder in the upstream task (i.e., the first-layer shallow features , the second-layer shallow features , the third-layer deep features , the fourth-layer shallow features ) are subjected to multi-scale feature fusion with the features of the previous layer. Since the features of different layers themselves have differences in scale, for example, the scale of shallow features is relatively closer to the original image, with rich detail information but low semantic abstraction degree, while deep features have undergone multiple downsampling operations, with rich semantics but possible loss of details. Through the multi-scale feature fusion method, the features of different layers can complement each other, thereby making up for the differences in feature scales and integrating a more comprehensive and discriminative feature representation.
[0070] After the above preprocessing of the first polyp image features and the generation and processing of multiple second polyp image features, the two are fused. This fusion combines the feature advantages of both aspects after careful processing, including both the unique information carried by the first polyp image features after Fourier transform and related processing and the multi-scale feature content extracted and optimized from different layers of the CNN encoder for multiple second polyp image features.
[0071] In this embodiment, in this embodiment, the generated segmentation prediction map can be deeply supervised by the Mask label, and the segmentation prediction map can guide the CNN encoder to learn better, so as to generate a more accurate segmentation map and better extract the edge information in the medical image.
[0072] It can be concluded from the above that in this embodiment, the polyp image features obtained by Fourier transform are fused with the deep features of the fifth layer of the CNN encoder, combining detail and semantic information to make the first polyp image features more comprehensive. Secondly, the second polyp image features are obtained by processing some features of the CNN encoder, mining and optimizing these features. Furthermore, the multi-layer feature fusion decoder performs multi-scale fusion on features of different layers, compensates for scale differences, and integrates strongly discriminative features. Finally, the segmentation map is generated by fusion, which can accurately depict the shape and boundary of the polyp, improve the segmentation accuracy, provide a more reliable basis for medical diagnosis, and assist doctors in accurately judging the condition of the polyp.
[0073] As Figure 2 shown, in an embodiment of the present disclosure, performing Fourier transform on a public polyp image to obtain polyp image features includes:
[0074] Splitting the public polyp image to obtain local polyp image information and global polyp image information;
[0075] Performing spatial information extraction on the local polyp image information and the global polyp image information to obtain target local features;
[0076] Performing frequency domain information extraction on the local polyp image information and the global polyp image information to obtain target global features;
[0077] Fusing the target local features and the target global features to obtain polyp image features.
[0078] In this embodiment, the public polyp image is divided into local polyp image information and global polyp image information . The local information focuses on reflecting the detailed performance of the polyp in the local area of the image, such as the texture on the surface of the polyp, the local edge contour, etc.; while the global information reflects the relationship, position layout, etc. between the polyp and the surrounding tissues and the entire image scene from an overall perspective. Through such splitting, the characteristics of the polyp in the image can be grasped more carefully and comprehensively.
[0079] Feeding the local polyp image information and the global polyp image information into the frequency domain branch and the spatial domain branch together for spatial and frequency domain information extraction.
[0080] In the spatial domain, for local polyp image information, local detail features can be further enhanced through spatial information extraction, making it more capable of highlighting the fine texture of polyps, the shape characteristics within small regions, etc.; for global polyp image information, features at the macroscopic level such as its layout and relative position in the overall image space can be refined. After spatial information extraction, the two are fused to obtain the target local features, integrating the key features of the local and global in the spatial dimension.
[0081] In the frequency domain, the image is converted into a representation of different frequency components. Through frequency domain information extraction, high-frequency key details such as polyp edge texture can be mined from local polyp image information, and low-frequency features reflecting the overall structure can be sorted out from global polyp image information. The two are fused to form the target global features, synthesizing the important information of the local and global in the frequency domain.
[0082] In this embodiment, there are two identical Fourier transform modules. The above-mentioned target local features and target global features are input into another Fourier transform module to obtain the final target local features and target global features.
[0083] Finally, the final target local features and target global features are fused, gathering all the key features of the polyp image mined in the spatial domain and frequency domain. It includes both the local and global morphology and layout characteristics observed from the spatial perspective, and the details and overall structure information reflected by different frequency components analyzed from the frequency perspective, enabling the finally obtained polyp image features to comprehensively and multi-levelly depict the specific situation of the polyp in the image.
[0084] It can be concluded from the above that in this embodiment, the common polyp image is first shunted to distinguish local and global information, which is conducive to comprehensively and carefully grasping the characteristics of polyps. Secondly, information is extracted and fused into target local and global features in the spatial domain and frequency domain respectively, which can strengthen the detail and macroscopic features. Furthermore, two-level Fourier transform modules are set to further optimize the features, making the features more accurate. Finally, the final target local and global features are fused to obtain the polyp image features, comprehensively and multi-levelly depicting the situation of polyps, providing a rich and high-quality feature basis for subsequent polyp image analysis, segmentation and other operations, and improving the processing accuracy.
[0085] In an embodiment of the present disclosure, spatial information extraction is performed on local polyp image information and global polyp image information to obtain target local features, including:
[0086] Feature extraction is performed on global polyp image information to obtain the first global polyp image information feature;
[0087] Feature extraction is performed on local polyp image information to obtain the first local polyp image information feature;
[0088] Perform a Hadamard product on the first global polyp image information feature and the first local polyp image information feature to obtain a local feature;
[0089] Obtain the target local feature based on the local feature.
[0090] In this embodiment, when performing spatial information extraction, the local polyp image information and the global polyp image information can be fed into the spatial branch. The local polyp image information extracts features through a convolutional kernel, highlighting the details of the polyp in the local area, such as the texture on the surface of the polyp, fine details such as local edge contours, etc., and then obtaining the first local polyp image information feature; the global polyp image information extracts features through a convolutional kernel, mining the features contained therein that can reflect the layout, relative position, etc. at the macroscopic level of the polyp in the overall image scene, thereby obtaining the first global polyp image information feature.
[0091] Use Hadamard convolution to perform pixel-level multiplication on the first global polyp image information feature and the first local polyp image information feature, that is, retain the features useful for the segmentation task and reduce the impact of background noise on the segmentation. For example, if a certain pixel position is related to the key performance of the polyp in both the local feature and the global feature (such as jointly reflecting a certain detailed feature of the polyp edge), this feature will be highlighted after multiplication; while those parts that belong more to background noise or have little relevance to the segmentation task will be relatively weakened after such multiplication, thus achieving the effect of reducing the impact of background noise on the segmentation, and finally obtaining the local feature.
[0092] Finally, the local feature is processed through convolution, normalization, and the ReLu activation function to obtain the target local feature . The convolution operation can further mine the potential information in the local feature, making its feature expression more rich and accurate; normalization helps to stabilize the data distribution, making the feature data more standardized in subsequent processing, fusion and other operations; and the ReLu activation function can limit each element in the local feature to be between [0, 1]. On the one hand, this introduces non-linearity and enhances the expression ability of the model, and on the other hand, it also makes the feature values within a reasonable range, and finally obtains the final target local feature.
[0093] It can be concluded from the above that in this embodiment, by separately extracting the global and local image information features, then highlighting the useful features and weakening the background noise through Hadamard product, and finally optimizing using convolution, normalization, and the ReLu activation function, high-quality target local features can be accurately mined, providing a reliable basis for polyp image segmentation and improving the segmentation accuracy.
[0094] In an embodiment of the present disclosure, frequency domain information extraction is performed on local polyp image information and global polyp image information to obtain target global features, including:
[0095] Feature extraction is performed on the global polyp image information to obtain second global polyp image information features;
[0096] Frequency domain Fourier transform is performed on the local polyp image information to obtain second local polyp image information features;
[0097] Based on the second global polyp image information features and the second local polyp image information features, global features are obtained;
[0098] Based on the global features, target global features are obtained.
[0099] In this embodiment, when performing frequency domain information extraction, the local polyp image information Performing frequency domain Fourier transform can transform the local image information from the spatial domain to the frequency domain. In the frequency domain, the image is represented by different frequency components. Through this transformation, high-frequency key details such as polyp edge textures can be extracted, and then second local polyp image information features are obtained. For the global polyp image information , feature learning is performed through convolution operations. The convolution kernel slides on the global image information, and can extract features related to the layout and macroscopic morphology of the polyp in the overall image at different positions and scales, thereby obtaining second global polyp image information features.
[0100] The second local polyp image information features and the second global polyp image information features are added together to obtain global features; the detailed information reflected by the local features in the frequency domain and the macroscopic layout information learned from the global features through spatial convolution complement each other. After addition, the advantages of these two aspects can be integrated, and the features of the polyp from different dimensions and perspectives can be integrated, so as to obtain more comprehensive global features. After obtaining the global features, the global features are processed through convolution, normalization, and ReLu activation function to obtain target global features (i.e., global edge information).
[0101] It can be concluded from the above that in this embodiment, by performing frequency domain Fourier transform on the local polyp image information, high-frequency details such as polyp edge textures can be extracted, and by using convolution on the global polyp image information to extract macroscopic morphology and other features, features are mined from different dimensions. Secondly, the corresponding features of the two are added and fused to integrate the advantages of multiple perspectives and make the global features more comprehensive. Finally, through convolution, normalization, and ReLu activation function processing, potential information is further mined, data is stabilized, and expression ability is enhanced to obtain high-quality target global features, providing accurate and reliable basis for subsequent tasks such as polyp image segmentation and improving processing accuracy.
[0102] In one embodiment of the present disclosure, the local polyp image information is subjected to frequency domain Fourier transform to obtain the second local polyp image information feature, including:
[0103] Segmenting local polyp image information to obtain a first local feature and a second local feature;
[0104] Perform feature extraction on the second local feature to obtain a third local feature;
[0105] Based on the first local feature and the third local feature, a second local polyp image information feature is obtained.
[0106] In this embodiment, local polyp image information Preprocessing is first performed on the local polyp image information Perform convolution, normalization, and ReLu activation function processing. The convolution operation can mine local feature patterns in the image and enhance the image's detail information; normalization can stabilize data distribution, making the subsequent processed data more standardized; the ReLu activation function introduces nonlinear factors to enhance the model's ability to express complex image feature relationships.
[0107] Then, the preprocessed image information is sliced according to the feature and The ratio of local polyp image information Divided into the first local feature (proportion ) and the second local feature (proportion ). For example, when When the proportion of (the second local feature) is 0.4, The first local feature accounts for 0.6. We mine valuable information for the final feature from different parts and perform differentiated processing on features of different proportions in a targeted manner.
[0108] The second local feature ( ) to extract edge features. In polyp images, edge features contain key detail information, such as the boundary contour of the polyp and the boundary with the surrounding tissue. By extracting edge features, the morphological characteristics of the polyp in the local area can be captured more accurately, and then the third local feature can be obtained.
[0109] Then the first local feature (1- ) is added to the third local feature, the features obtained from different emphasis angles are merged, and the local information about the polyp contained in the two parts of the features is integrated to complement each other and jointly improve the description of the local characteristics of the polyp.
[0110] Finally, the added features are convolved, normalized, and passed through the ReLu activation function to obtain the second local polyp image information features, enriching the expression of polyp semantic information and enabling it to more comprehensively and accurately reflect various characteristics of the polyp in the local area.
[0111] As can be seen from the above, in this embodiment, by preprocessing the local polyp image information, the data features are optimized, making it more conducive to subsequent operations. Secondly, the feature slices are divided proportionally, enabling targeted extraction of valuable information from different parts and achieving differential processing. Furthermore, by extracting edge features, the local morphological characteristics of the polyp can be accurately captured, and by fusing features from different perspectives, the local feature description is improved. Finally, the second local polyp image information features obtained through reprocessing enrich the semantic information, enabling it to more comprehensively and accurately reflect the polyp characteristics, providing strong support for subsequent tasks such as polyp image analysis.
[0112] In an embodiment of the present disclosure, feature extraction is performed on the second local features to obtain third local features, including:
[0113] Edge feature extraction is performed on the second local features to obtain the boundary information of the second local features, and the boundary information is used as the third local features.
[0114] In this embodiment, the second local features ( ) are input into the edge feature extraction module to extract the boundary information in the polyp image.
[0115] First, a fast Fourier transform is performed on the second local features to obtain frequency domain features.
[0116] The second local features can be regarded as feature X. Feature X is first transformed from the spatial domain to the frequency domain through a fast Fourier transform. The specific operation can be expressed as:
[0117]
[0118] where u and v respectively represent the real part parameter and the imaginary part parameter after conversion to the frequency domain space, and H and W respectively represent the height and width of feature X. represents the ordinate of feature X, represents the abscissa of feature X.
[0119] Secondly, the low-frequency information in the frequency domain features is moved to the middle of the image.
[0120] After the fast Fourier transform, all the low-frequency information including the contour and texture is moved to the middle of the image. This process can be expressed as:
[0121]
[0122] where, Represents the feature that moves the low-frequency components in the frequency-domain features to the middle of the image.
[0123] In the frequency domain, low-frequency information is related to the overall structural features of the image such as contours and textures. Concentrating it in the middle position of the image facilitates subsequent unified processing and screening, and enables more targeted operations on the low-frequency information.
[0124] Again, remove the low-frequency information in the frequency-domain features to obtain high-frequency information.
[0125] By multiplying the low-frequency information within a certain range in the frequency-domain features by a zero matrix of the same size to remove the low-frequency information, this process can be expressed as:
[0126]
[0127] Among them, Represents the high-frequency information after occluding the low-frequency components, Represents the range for occluding the low-frequency information, Represents the operation of occluding the low-frequency information in the middle, specifically multiplying the middle low-frequency information by a zero matrix of this size.
[0128] and Multiplying the two can set the low-frequency information within the specified range to zero, thereby extracting the high-frequency information. In polyp images, high-frequency information usually contains detailed features such as edges. By removing the low-frequency information and retaining the high-frequency information, it is more conducive to focusing on extracting edge details and highlighting the information useful for depicting the boundary contour of the polyp.
[0129] Again, perform a Stack operation on the high-frequency information.
[0130] By performing a Stack operation on the features after extracting the high-frequency information in the frequency domain, adding the real part parameter and the imaginary part parameter can integrate the features of different dimensions (real part and imaginary part) of the high-frequency information in the frequency domain, making its representation more concise and facilitating subsequent convolutional operations for learning, further enhancing the extraction effect of the edge features.
[0131] Again, perform convolutional processing on the features after the Stack operation.
[0132] Use convolutions with three different receptive fields for feature learning, with sizes of 、 、 .
[0133] Convolution kernels with different receptive fields can capture feature information at different scales, Convolution can be used to adjust the number of channels, fuse features, etc.; and Convolution can extract features in a larger local range. By learning the high-frequency information after stacking through these three convolution operations, edge features contained in the high-frequency information can be mined from multiple angles and scales, making the extraction of edge features more comprehensive and accurate.
[0134] Finally, perform an inverse fast Fourier transform on the features after convolution processing.
[0135] Because subsequent operations (such as fusing with other spatial domain features, participating in generating the final local polyp image information features, etc.) are mostly carried out in the spatial domain, the high-frequency edge detail information extracted in the frequency domain is transformed back into the spatial domain, converting the key edge information extracted in the frequency domain back into the spatial domain, which is convenient for working with other spatial domain features, facilitating the further integration and utilization of polyp image features in the entire process, and thus providing strong support for accurately depicting the morphological characteristics of polyps in the local area, etc.
[0136] It can be concluded from the above that in this embodiment, the edge features are extracted by performing a fast Fourier transform on the second local feature and subsequent series of operations, mining information from the frequency domain perspective, and being able to accurately focus on key details such as the boundary contour of the polyp. Secondly, operations such as shifting and removing low-frequency information are beneficial to highlighting high-frequency detail features and strengthening the depiction of the edge. Thirdly, the Stack operation makes the feature representation more concise and conducive to convolution learning, and convolutions with different receptive fields comprehensively and accurately extract edge features. Finally, the inverse fast Fourier transform is convenient for collaborating with spatial domain features, providing strong support for accurately depicting the polyp morphology and improving the quality of polyp image analysis.
[0137] Exemplarily, as Figures 6 - 8 shown, the network framework of the embodiment is implemented in PyTorch and trained on an NVIDIA RTX3090 GPU with 24GB of memory. The input image is adjusted to 224×224, the batch size is 32, and the CNN encoder is initialized with pre-trained weights of ConvNeXt on ImageNet. Our model is optimized using the SGD optimizer, with an initial learning rate of 0.01, a momentum of 0.9, and a weight decay of 1e-4. During the training process, there is no overlap between the training data and the test data, and data augmentation methods such as flipping and rotation are also used to enhance the image quality.
[0138] In five specific embodiments, the public polyp training data includes Kvasir, CVC-ClinicDB, CVC-300, CVC-ColonDB, and ETIS. Among them, the training dataset includes 900 datasets from Kvasir and 550 datasets from CVC-ClinicDB, and the remaining datasets are used for testing.
[0139] To verify the effectiveness of each module of the model, in the 5th specific embodiment of this example, experiments were conducted with 15 different models using the above dataset to compare the segmentation results, and ablation experiments were carried out to prove the effectiveness of the model;
[0140] Table 2 shows the comparison of results with other methods in the first Kvasir and the second CVC-ClinicDB embodiments. These methods include UNet, UNet++, Swin-UNet, UTNet, UNeXt, PraNet, FANet, TGANet, MSnet, TMF-Net, CMUNet, CMUNeXt, PRCNet, GCtx-UNet, and CFANet. It can be seen from the table that the proposed method is superior to other methods in the comprehensive indicators of Dice similarity score and MAE.
[0141] Table 2 Comparison of results with other methods in the first Kvasir and the second CVC-ClinicDB embodiments
[0142]
[0143] Table 3 shows the comparison of results with other methods in the second embodiment of this example. These methods include UNet, UNet++, Swin-UNet, UTNet, UNeXt, PraNet, FANet, TGANet, MSnet, TMF-Net, CMUNet, CMUNeXt, PRCNet, GCtx-UNet, and CFANet. It can be seen from the table that the proposed method is superior to other methods in the IOU similarity score and MAE indicators.
[0144] Table 3 Comparison of results with other methods in the third CVC-300, the fourth CVC-ColonDB, and the fifth ETIS embodiments
[0145]
[0146] The purpose of this embodiment is to obtain the global receptive field from the multi-scale feature representation and the edge information contained in the high-frequency information in the frequency domain. Therefore, when this embodiment extracts important edge information, it extracts the image boundary features from the high-frequency domain according to the Fourier transform of the input features, and combines them with the spatial domain features to reduce background noise and irrelevant information. In addition, according to the problem of insufficient anti-interference ability of background noise and other information in the original decoder, this embodiment proposes a multi-scale feature denoising decoder, which uses multi-scale subtraction to subtract the features in the deep and shallow layer information element by element. On the one hand, it reduces the noise impact in different scale information, and on the other hand, it alleviates the problem of information loss in the upsampled high layer through the Hadamard product, further improving the accuracy of polyp segmentation. A large number of experiments conducted on the CVC-300, CVC-ClinicDB, CVC-ColonDB, Kvasir, and ETIS datasets show that this embodiment improves the accuracy of polyp segmentation.
[0147] Through the above polyp image segmentation method, this embodiment first uses the ConvNeXt convolutional layer to construct an encoder to extract the semantic information of local features and global features at different scales. Then, it constructs an image boundary information extraction module based on the Fourier transform to obtain the global receptive field in the multi-scale feature representation and the edge information contained in the high-frequency information in the frequency domain. This embodiment constructs a global receptive field in the frequency domain in the FTM, so as to extract the image boundary features from the high-frequency domain according to the Fourier transform of the input features, and combines them with the spatial domain features to reduce background noise and irrelevant information. Finally, this embodiment constructs a multi-scale feature denoising decoder, and designs a denoising module in the cross-scale multi-scale fusion decoder, which effectively reduces the impact of background noise on the segmentation task and retains the useful parts of the features for segmentation. On the one hand, it reduces the noise impact in different scale information, and on the other hand, it alleviates the problem of information loss in the upsampled high layer through the Hadamard product, further improving the accuracy of polyp segmentation. A large number of experiments conducted on the CVC-300, CVC-ClinicDB, CVC-ColonDB, Kvasir, and ETIS datasets show that this embodiment improves the accuracy of multi-organ segmentation and heart segmentation.
[0148] A polyp image segmentation method corresponding to the above embodiment Figure 9 is a structural block diagram of a polyp image segmentation device provided by an embodiment of the present disclosure. For the sake of illustration, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 9 and this polyp image segmentation device 20 includes: a feature combination module 21, a first fusion module 22, a feature product module 23, a second fusion module 24, and a segmentation module 25.
[0149] Among them, the feature combination module 21 is used to combine multiple target image features pairwise to obtain a first feature, a second feature, and a third feature. The multiple target image features include a first target image feature, a second target image feature, and a third target image feature, and the multiple target image features are features obtained from multiple image features.
[0150] The first fusion module 22 is used to fuse the first target image feature and the first feature to obtain a first global feature.
[0151] The feature product module 23 is used to perform a Hadamard product on the second feature and the third feature to obtain a fourth feature.
[0152] The second fusion module 24 is used to fuse the first global feature and the fourth feature to obtain a second global feature.
[0153] The segmentation module 25 is used to generate a segmentation prediction map based on the second global feature.
[0154] In an embodiment of the present disclosure, a polyp image segmentation device 20 further includes: an image segmentation module; specifically, the image segmentation module is used for:
[0155] Performing a Fourier transform on a public polyp image to obtain polyp image features;
[0156] Performing feature extraction on a public polyp image to obtain multiple image features;
[0157] Fusing the polyp image features and the first target image feature to obtain a first polyp image feature;
[0158] Processing the multiple image features to obtain multiple second polyp image features;
[0159] Fusing the first polyp image feature and the multiple second polyp image features to generate a segmentation map.
[0160] In an embodiment of the present disclosure, the image segmentation module is specifically further used for:
[0161] Shunting a public polyp image to obtain local polyp image information and global polyp image information;
[0162] Performing spatial information extraction on the local polyp image information and the global polyp image information to obtain target local features;
[0163] Performing frequency domain information extraction on the local polyp image information and the global polyp image information to obtain target global features;
[0164] Fusing the target local features and the target global features to obtain polyp image features.
[0165] In one embodiment of the present disclosure, the image segmentation module is further specifically configured to:
[0166] Extract features from the global polyp image information to obtain the first global polyp image information feature;
[0167] Extract features from the local polyp image information to obtain the first local polyp image information feature;
[0168] Perform a Hadamard product on the first global polyp image information feature and the first local polyp image information feature to obtain a local feature;
[0169] Obtain the target local feature based on the local feature.
[0170] In one embodiment of the present disclosure, the image segmentation module is further specifically configured to:
[0171] Extract features from the global polyp image information to obtain the second global polyp image information feature;
[0172] Perform a frequency-domain Fourier transform on the local polyp image information to obtain the second local polyp image information feature;
[0173] Obtain the global feature based on the second global polyp image information feature and the second local polyp image information feature;
[0174] Obtain the target global feature based on the global feature.
[0175] In one embodiment of the present disclosure, the image segmentation module is further specifically configured to:
[0176] Segment the local polyp image information to obtain the first local feature and the second local feature;
[0177] Extract features from the second local feature to obtain the third local feature;
[0178] Obtain the second local polyp image information feature based on the first local feature and the third local feature.
[0179] In one embodiment of the present disclosure, the image segmentation module is further specifically configured to:
[0180] Extract the edge features of the second local feature to obtain the boundary information of the second local feature, and use the boundary information as the third local feature.
[0181] See Figure 10 , Figure 10 is a schematic block diagram of an electronic device provided in an embodiment of the present disclosure. As Figure 10The electronic device 300 in the present embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module in the above-mentioned device embodiments, for example Figure 9 the functions of the modules 21 to 25 shown.
[0182] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0183] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.
[0184] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0185] In specific implementation, the processors 301, input devices 302, and output devices 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first embodiment and the second embodiment of a polyp image segmentation method provided by the embodiments of the present disclosure, and may also implement the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.
[0186] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing related hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0187] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0188] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0189] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0190] In several embodiments provided by this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, and can also be electrical, mechanical or other forms of connection.
[0191] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.
[0192] In addition, in each embodiment of the present disclosure, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0193] The above are only the specific implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A polyp image segmentation method, characterized in that, Including: Performing Fourier transform on the public polyp image to obtain polyp image features; including: Shunting the public polyp image to obtain local polyp image information and global polyp image information; Performing spatial information extraction on the local polyp image information and the global polyp image information to obtain target local features; Performing frequency domain information extraction on the local polyp image information and the global polyp image information to obtain target global features; Fusing the target local features and the target global features to obtain polyp image features; Performing feature extraction on the public polyp image to obtain multiple image features; including: Use a CNN encoder to extract features from public polyp images, obtaining multiple image features, where the multiple image features include: the first-layer shallow features , the second-layer shallow features , the third-layer deep features , the fourth-layer shallow features , the fifth-layer deep features ; Combining multiple target image features in pairs to obtain a first feature, a second feature and a third feature, wherein the multiple target image features include the first target image feature, the second target image feature and the third target image feature, and the multiple target image features are features obtained from multiple image features; the first target image feature is a fifth-layer deep feature , the second target image feature is the fourth shallow feature , the third target image feature is a third layer deep feature ; The first feature is the third layer deep feature With the fourth layer shallow features The combined features ; The second feature is the third layer deep feature With the fifth layer of deep features The combined features ; The third feature is the fourth shallow feature With the fifth layer of deep features The combined features The said combining of multiple target image features in pairs comprises: , They are respectively sent into three denoising units, and the deep features and shallow features are respectively transmitted through the convolution kernel sizes of , , The convolution is used for learning; the deep features and shallow features of corresponding sizes are subtracted element by element, and then the three layers of features after subtraction are added; Fusing the first target image feature and the first feature to obtain a first global feature; Performing Hadamard product on the second feature and the third feature to obtain a fourth feature; Fusing the first global feature and the fourth feature to obtain a second global feature; Generating a segmentation prediction map based on the second global feature and performing deep supervision on the generated segmentation prediction map through a Mask label; Fusing the polyp image features and the first target image feature to obtain a first polyp image feature; Processing the multiple image features to obtain multiple second polyp image features; Fusing the first polyp image feature and the multiple second polyp image features to generate a segmentation map.
2. The polyp image segmentation method according to claim 1, wherein The performing spatial information extraction on the local polyp image information and the global polyp image information to obtain target local features includes: Performing feature extraction on the global polyp image information to obtain a first global polyp image information feature; Performing feature extraction on the local polyp image information to obtain a first local polyp image information feature; Performing Hadamard product on the first global polyp image information feature and the first local polyp image information feature to obtain local features; Obtaining target local features based on the local features.
3. The polyp image segmentation method according to claim 1, wherein, The performing frequency domain information extraction on the local polyp image information and the global polyp image information to obtain target global features includes: Performing feature extraction on the global polyp image information to obtain a second global polyp image information feature; Performing frequency domain Fourier transform on the local polyp image information to obtain a second local polyp image information feature; Obtaining global features based on the second global polyp image information feature and the second local polyp image information feature; Obtaining target global features based on the global features.
4. The polyp image segmentation method according to claim 3, wherein, The performing frequency domain Fourier transform on the local polyp image information to obtain a second local polyp image information feature includes: Segmenting the local polyp image information to obtain a first local feature and a second local feature; Performing feature extraction on the second local feature to obtain a third local feature; Obtaining the second local polyp image information feature based on the first local feature and the third local feature.
5. The polyp image segmentation method according to claim 4, characterized in that, The performing feature extraction on the second local feature to obtain a third local feature includes: Performing edge feature extraction on the second local feature to obtain boundary information of the second local feature, and using the boundary information as the third local feature.
6. A polyp image segmentation device, characterized in that, Including: The image segmentation module is specifically used for: shunting the public polyp image to obtain local polyp image information and global polyp image information; extracting spatial information from the local polyp image information and the global polyp image information to obtain target local features; extracting frequency-domain information from the local polyp image information and the global polyp image information to obtain target global features; fusing the target local features and the target global features to obtain polyp image features; The image segmentation module is further specifically configured to: use a CNN encoder to extract features from the public polyp image to obtain multiple image features, where the multiple image features include: the first-layer shallow features , the second-layer shallow features , the third-layer deep features , the fourth-layer shallow features , the fifth-layer deep features ; A feature combination module is used to combine multiple target image features in pairs to obtain a first feature, a second feature and a third feature, wherein the multiple target image features include the first target image feature, the second target image feature and the third target image feature, and the multiple target image features are features obtained from multiple image features; the first target image feature is a fifth-layer deep feature , the second target image feature is the fourth shallow feature , the third target image feature is a third layer deep feature ; The first feature is the third layer deep feature With the fourth layer shallow features The combined features ; The second feature is the third layer deep feature With the fifth layer of deep features The combined features ; The third feature is the fourth shallow feature With the fifth layer of deep features The combined features The said combining of multiple target image features in pairs comprises: , They are respectively sent into three denoising units, and the deep features and shallow features are respectively transmitted through the convolution kernel sizes of , , The convolution is used for learning; the deep features and shallow features of corresponding sizes are subtracted element by element, and then the three layers of features after subtraction are added; The first fusion module is used for fusing the first target image feature and the first feature to obtain a first global feature; The feature multiplication module is used for performing a Hadamard product on the second feature and the third feature to obtain a fourth feature; The second fusion module is used for fusing the first global feature and the fourth feature to obtain a second global feature; The segmentation module is used for generating a segmentation prediction map based on the second global feature and performing deep supervision on the generated segmentation prediction map through a Mask label; The image segmentation module is specifically further used for: fusing the polyp image features and the first target image feature to obtain first polyp image features; processing multiple image features to obtain multiple second polyp image features; fusing the first polyp image features and the multiple second polyp image features to generate a segmentation map.
7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Colon polyp segmentation method fusing multi-class feature processing strategies
CN117173410A