A medical image segmentation method based on a fully convolutional neural network

Through deep learning algorithms and attention mechanisms, spatial pyramid pooling modules, deconvolution operations and adversarial learning strategies, the problem of insufficient feature extraction in medical image segmentation is solved, and medical image segmentation with higher accuracy is achieved.

CN119206237BActive Publication Date: 2025-07-08CHANGSHU INSTITUTE OF TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411686975.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-07-08
Estimated Expiration
2044-11-25

AI Technical Summary

Technical Problem

Traditional fully convolutional neural networks are difficult to capture sufficiently rich and fine-grained feature information in medical image segmentation, resulting in the anatomy and lesion areas being ignored, affecting the accuracy of segmentation results.

Method used

Image preprocessing is adopted for deep learning algorithms, attention mechanisms and spatial pyramid pooling modules are introduced for feature extraction, combined with deconvolution operations and optimized upsampling of jump connection structures, conditional random field models are introduced for post-processing, and network generalization capabilities are improved through adversarial learning strategies.

Benefits of technology

It significantly improves the accuracy and robustness of medical image segmentation, especially in the identification of complex anatomical structures and lesion areas, providing reliable technical support for auxiliary diagnostic treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206237B_ABST
    Figure CN119206237B_ABST
Patent Text Reader

Abstract

The present invention discloses a medical image segmentation method based on a fully convolutional neural network, comprising: aiming at the complex anatomical structures and pathological features of medical images, using a deep learning algorithm to preprocess the medical images, enhancing the key features in the images, and providing a richer information basis for subsequent feature extraction; in the feature extraction stage of the fully convolutional neural network, introducing an attention mechanism module and a spatial pyramid pooling module to learn the importance weights of different regions in the image, adaptively adjusting the weight distribution of the convolutional kernels, and extracting more fine-grained and discriminative features. Optimizing the upsampling process through deconvolution operations and skip connection structures, introducing a conditional random field model for post-processing optimization, constructing a multi-scale feature fusion module, and adopting an adversarial learning strategy to improve the generalization ability of the network. The present invention effectively improves the accuracy and robustness of medical image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to a medical image segmentation method based on a fully convolutional neural network. Background Art

[0002] In the field of medical image segmentation, the fully convolutional neural network faces a key technical challenge. Since medical images usually have complex anatomical structures and pathological features, traditional fully convolutional neural networks often have difficulty capturing sufficient rich and fine-grained feature information during the feature extraction stage. This results in the network being unable to accurately restore the spatial details of the original image during the subsequent upsampling and pixel-level classification processes, thereby affecting the accuracy of the segmentation results. Specifically, when the network extracts features from medical images, the successive operations of the convolutional layer and the pooling layer will inevitably lose a part of the spatial information. Although this information loss can be accepted in general image segmentation tasks, it may cause key anatomical structures or lesion areas to be ignored in medical image segmentation. At the same time, the deconvolution or transposed convolution operations used in the upsampling process are also difficult to completely restore the lost spatial details, further affecting the accuracy of pixel-level classification. This technical challenge highlights the uniqueness and complexity of the medical image segmentation task, and requires researchers to explore more refined and intelligent feature extraction, upsampling, and classification strategies in order to obtain more accurate and reliable segmentation results, providing strong support for subsequent diagnosis and treatment. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides a medical image segmentation method based on a fully convolutional neural network, including:

[0004] Preprocessing the medical image through a deep learning algorithm to enhance the key features of the image;

[0005] During the process of feature extraction based on the fully convolutional neural network, an attention mechanism module is introduced, and by learning the importance weights of different regions in the image, the weight distribution of the convolutional kernel is adaptively adjusted to extract fine-grained and discriminative image features;

[0006] Through the spatial pyramid pooling module embedded in the fully convolutional neural network, multi-scale feature fusion and context information capture are performed;

[0007] The fused feature map is obtained through multi-scale feature fusion, and the fused feature map is gradually restored to the same spatial size as the original image through upsampling operations;

[0008] Based on the feature map obtained by upsampling, each pixel is classified through a convolutional operation to obtain a pixel-level segmentation result;

[0009] An adversarial learning strategy is introduced. By constructing a generator and a discriminator network, the segmentation result of the fully convolutional neural network is adversarially trained with the true annotation to obtain the final segmentation result.

[0010] Preferably, the process of preprocessing medical images through a deep learning algorithm to enhance the key features of the images includes:

[0011] Obtain medical image data, and analyze and identify complex anatomical structures and pathological features in the images using a deep learning algorithm;

[0012] Preprocess the medical images, adjust the gray-scale distribution of the images through the histogram equalization algorithm, and enhance the contrast of the images;

[0013] According to the characteristics of the images, adopt an adaptive contrast adjustment algorithm to specifically enhance the contrast of local regions of the images and highlight the pathological features.

[0014] Preferably, the process of introducing an attention mechanism module to extract fine-grained and discriminative image features includes:

[0015] Construct a fully convolutional neural network model, introduce an attention mechanism module in the feature extraction stage. In the attention mechanism module, by calculating the importance weights of different regions in the medical images, generate an attention weight matrix, and perform element-wise multiplication of the attention weight matrix and the weight matrix of the convolution kernel to obtain an adjusted convolution kernel weight distribution;

[0016] Use the adjusted convolution kernel to perform convolution operations on the medical images, extract the features of key anatomical structures and lesion regions in the images, and obtain fine-grained and discriminative feature maps and corresponding image features.

[0017] Preferably, the process of performing multi-scale feature fusion and context information capture through a spatial pyramid pooling module embedded in the fully convolutional neural network includes:

[0018] Based on the spatial pyramid pooling module, perform multi-scale pooling operations on the fine-grained and discriminative image features to obtain feature representations at different scales;

[0019] Based on the feature representations at different scales, perform multi-scale feature fusion, and at the same time capture the context information in the feature maps to enhance the semantic representation ability of the features.

[0020] Preferably, the process of classifying each pixel through convolution operations based on the feature map obtained by upsampling to obtain a pixel-level segmentation result includes:

[0021] The transposed convolution operation is replaced by a deconvolution operation for upsampling, and the convolution kernel parameters during the upsampling process are learned. At the same time, a skip connection structure is introduced to fuse the high-resolution features of the shallow network and the high-level semantic features of the deep network to obtain the upsampling result.

[0022] Then, pixel-level classification is performed. Through the conditional random field model, the spatial relationship and class dependence relationship between pixels are modeled, and the segmentation result output by the fully convolutional neural network is post-processed and optimized to smooth the segmentation boundary and eliminate isolated misclassified pixels.

[0023] A multi-scale feature fusion module is constructed. Feature maps of multiple scales are extracted at different levels of the fully convolutional neural network, and through feature fusion operations, the feature information of different scales is integrated to obtain a more comprehensive and robust feature representation.

[0024] Preferably, the process of using the deconvolution operation to replace the traditional transposed convolution operation for upsampling includes:

[0025] The input low-resolution image is used to extract deep semantic features by a pre-trained convolutional neural network.

[0026] According to the preset upsampling multiple, the deep semantic features are upsampled by the deconvolution operation to obtain a first upsampled feature map.

[0027] The high-resolution features of the shallow layer corresponding to the input image are obtained, and the high-resolution features of the shallow layer are fused with the first upsampled feature map through the skip connection to obtain a second upsampled feature map.

[0028] According to the second upsampled feature map, the resolution is further improved by the deconvolution operation to obtain a third upsampled feature map.

[0029] According to the third upsampled feature map, convolution operations are used for detail enhancement to obtain a fourth upsampled feature map.

[0030] According to the fourth upsampled feature map, it is converted into a high-resolution output image through a pixel rearrangement operation.

[0031] According to the high-resolution output image and the original high-resolution image, the reconstruction error is calculated through a loss function, and the convolution kernel parameters of the deconvolution and convolution operations are optimized using the error backpropagation algorithm.

[0032] Preferably, the process of performing pixel-level classification includes:

[0033] The pixel-level segmentation result output by the fully convolutional neural network is used as the input of the conditional random field model.

[0034] Construct a conditional random field model, define pixel nodes and edge feature functions, and model the spatial relationship and class dependence relationship between pixels;

[0035] The spatial relationship feature function considers the relative position of pixels in the image, and the class dependence feature function considers the consistency of class labels of adjacent pixels;

[0036] In the conditional random field model, the optimized pixel labels are obtained through maximum a posteriori probability inference;

[0037] The inference process considers the local features and global context information of pixels, smooths the segmentation boundary, and eliminates isolated misclassified pixels.

[0038] Preferably, the inference process further includes post-processing the inference result of the conditional random field model;

[0039] The post-processing process includes:

[0040] Using connected component analysis and morphological operations to remove small isolated areas and fill holes in the segmentation result;

[0041] According to the spatial position of pixels and the optimized class labels, reconstruct the boundary contour of the segmentation result;

[0042] Through edge detection and contour tracking algorithms, obtain the boundary of the segmentation object, evaluate the optimized segmentation result, and calculate pixel-level classification accuracy and intersection over union metrics;

[0043] Compare the optimized segmentation result with the original output of the fully convolutional neural network to quantify the improvement effect of the conditional random field model;

[0044] Iteratively optimize the parameters of the conditional random field model, and find the optimal feature function weights and inference algorithm parameters through cross-validation and grid search.

[0045] Preferably, construct a multi-scale feature fusion module, extract multi-scale feature maps at different levels of the fully convolutional neural network, and the process of integrating different-scale feature information through feature fusion operations includes:

[0046] According to different levels of the fully convolutional neural network, obtain several scales of feature maps, where the feature maps contain different-scale feature information;

[0047] Perform a concatenation operation on the several scales of feature maps, concatenate the different-scale feature maps in the channel dimension to obtain a concatenated multi-scale feature map;

[0048] For the spliced multi-scale feature map, a weighting operation is performed. Through the learned weight parameters, different-scale features in the multi-scale feature map are weighted and fused to obtain a fused multi-scale feature representation.

[0049] Preferably, an adversarial learning strategy is introduced. The process of obtaining the final segmentation result by constructing a generator and a discriminator network and performing adversarial training on the segmentation result of the fully convolutional neural network and the true annotation includes:

[0050] Construct a fully convolutional neural network as the generator network. Extract multi-scale features of the medical image through convolutional layers and pooling layers, and restore the spatial resolution of the feature map through an upsampling layer to obtain a pixel-level segmentation result;

[0051] Construct a discriminator network. Extract the features of the segmentation result generated by the generator network and the true annotation through convolutional layers, and determine whether the segmentation result is a true annotation through a fully connected layer to obtain a binary classification discriminant result;

[0052] During the training process, the generator network and the discriminator network are alternately trained. The goal of the generator network is to generate a realistic segmentation result to deceive the discriminator, and the goal of the discriminator network is to determine whether the segmentation result is a true annotation;

[0053] During the adversarial training process, a gradient penalty term is introduced to limit the gradient norm of the discriminator network. A data augmentation technique is used to expand the training samples. By performing rotation, translation, and scaling transformations on the medical image, diverse training samples are generated; Through adversarial training, the generator network and the discriminator network play against each other, and finally reach a Nash equilibrium, making the segmentation result generated by the generator network indistinguishable from the true annotation, and it is also difficult for the discriminator network to distinguish the true from the false;

[0054] In the test phase, the trained generator network is used to perform segmentation prediction on unknown medical images, and post-processing techniques are used to optimize the segmentation result to obtain the final medical image segmentation result.

[0055] Compared with the prior art, the present invention has the following advantages and technical effects:

[0056] The present invention discloses a medical image segmentation method based on deep learning. Aiming at the complex anatomical structures and pathological features in medical images, the present invention uses image enhancement technology to preprocess medical images, and introduces an attention mechanism and a spatial pyramid pooling module into the fully convolutional neural network to extract more fine-grained and discriminative features. The upsampling process is optimized through deconvolution operations and skip connection structures, a conditional random field model is introduced for post-processing optimization, a multi-scale feature fusion module is constructed, and an adversarial learning strategy is adopted to improve the generalization ability of the network. This method can effectively improve the accuracy and robustness of medical image segmentation, especially showing excellent performance in the recognition of complex anatomical structures and lesion areas, providing reliable technical support for subsequent auxiliary diagnosis and treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0058] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will describe this application in detail with reference to the drawings and in conjunction with the embodiments.

[0060] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0061] As Figure 1 shown, an embodiment of the present invention provides a medical image segmentation method based on a fully convolutional neural network, including:

[0062] Preprocess the medical image through a deep learning algorithm to enhance the key features of the image;

[0063] During the process of feature extraction based on the fully convolutional neural network, introduce an attention mechanism module, adaptively adjust the weight distribution of the convolutional kernel by learning the importance weights of different regions in the image, and extract fine-grained and discriminative image features;

[0064] Through the spatial pyramid pooling module embedded in the fully convolutional neural network, perform multi-scale feature fusion and context information capture on the fine-grained and discriminative image features;

[0065] The fused feature map is obtained through multi-scale feature fusion, and the fused feature map is gradually restored to the same spatial size as the original image through upsampling operations;

[0066] Based on the feature map obtained by upsampling, each pixel is classified through convolution operations to obtain a pixel-level segmentation result;

[0067] An adversarial learning strategy is introduced. By constructing a generator and a discriminator network, the segmentation result of the fully convolutional neural network is adversarially trained with the ground truth annotation to obtain the final segmentation result.

[0068] Furthermore, the process of enhancing the key features of medical images through deep learning algorithms includes:

[0069] Obtain medical image data. For the complex anatomical structures and pathological features in the images, deep learning algorithms are used for analysis and recognition;

[0070] Preprocess the medical images. Adjust the gray-scale distribution of the images through the histogram equalization algorithm to enhance the contrast of the images;

[0071] According to the characteristics of the images, an adaptive contrast adjustment algorithm is adopted to specifically enhance the contrast of local regions of the images and highlight the pathological features.

[0072] In this embodiment, for the medical images processed by the image enhancement technology, the key anatomical structures and pathological features therein are significantly enhanced.

[0073] More specifically, in order to obtain medical image data, in this embodiment, CT, MRI and other medical images are extracted from the hospital's imaging storage system, or obtained from public medical image datasets such as the LIDC-IDRI lung nodule dataset. After obtaining the medical images, first preprocess the images, and use the histogram equalization algorithm to adjust the gray-scale distribution of the images. Histogram equalization enhances the contrast of the images by mapping the gray-scale values of the images into a uniformly distributed interval. For example, for an 8-bit gray-scale image with a gray-scale value range of 0 - 255, histogram equalization converts the gray-scale distribution of the original image into a uniform distribution, making the number of pixels at each gray level as equal as possible. On this basis, according to the characteristics of medical images, an adaptive contrast adjustment algorithm is adopted. According to the gray-scale statistical information of local regions of the images, the intensity of contrast enhancement is dynamically adjusted. The adaptive contrast adjustment algorithm in this embodiment includes adaptive histogram equalization (AHE) and contrast-limited adaptive histogram equalization (CLAHE). The CLAHE algorithm avoids over-enhancing noise and artifacts by restricting the contrast amplification factor of the local histogram. After the image enhancement process, the key anatomical structures and pathological features in the medical images become more obvious and prominent.

[0074] Furthermore, the process of introducing an attention mechanism module to extract fine-grained and discriminative image features includes:

[0075] Construct a fully convolutional neural network model. Introduce an attention mechanism module in the feature extraction stage. In the attention mechanism module, by calculating the importance weights of different regions in the medical image, generate an attention weight matrix, and perform element-wise multiplication of the attention weight matrix and the weight matrix of the convolution kernel to obtain an adjusted convolution kernel weight distribution;

[0076] Use the adjusted convolution kernel to perform convolution operations on the medical image, extract the features of key anatomical structures and lesion regions in the image, and obtain a fine-grained and discriminative feature map and corresponding image features.

[0077] More specifically, when constructing a fully convolutional neural network model, first obtain a medical image dataset, normalize these images, scale the pixel values to between 0 and 1, and at the same time apply a Gaussian filtering algorithm for denoising to reduce random noise in the image. Then, input the preprocessed image into the fully convolutional neural network, which includes multiple convolutional layers and pooling layers for extracting the basic features of the image. In the feature extraction stage, introduce an attention mechanism module, which calculates the importance weights of each region through the Softmax function, and these weights form an attention weight matrix. For example, for a 128x128 feature map, an attention mechanism outputs a weight matrix of the same size, where each element represents the importance of the corresponding position in the original feature map. Then, perform element-wise multiplication of this attention weight matrix and the weight matrix of the convolution kernel to obtain an adjusted convolution kernel weight distribution. In this way, when the network performs convolution operations, it can focus more on key regions in the image, such as lesions or abnormal structures, and extract more discriminative features.

[0078] In this embodiment, an attention mechanism is introduced during the feature extraction process to focus on key regions in the image. Commonly used attention mechanisms include channel attention and spatial attention. By learning the importance weights of different channels and spatial positions in the feature map, the feature representation is adaptively adjusted to highlight important regions and features. Finally, the extracted image features are input into a classifier or detector for disease diagnosis and lesion localization. For image classification tasks, a softmax classifier can be used to map the image features to different categories. For image localization tasks, a Region Proposal Network (RPN) can be used to generate candidate regions, and a convolutional neural network is used to classify and regress the candidate regions to obtain the accurate target position and size. Through the analysis and recognition of deep learning algorithms, high-precision medical image diagnosis and lesion detection can be achieved, providing auxiliary decision-making support for clinicians.

[0079] Furthermore, the process of multi-scale feature fusion and context information capture through the spatial pyramid pooling module embedded in the fully convolutional neural network includes:

[0080] Based on the spatial pyramid pooling module, by performing multi-scale pooling operations on fine-grained and discriminative image features, feature representations at different scales are obtained;

[0081] Based on the feature representations at different scales, multi-scale feature fusion is performed while capturing the context information in the feature map, enhancing the semantic representation ability of the features. While retaining the high-level semantic information, the fused features can restore the spatial detail information of the image as much as possible.

[0082] More specifically, to solve the problem of spatial information loss during the feature extraction process, a spatial pyramid pooling module is embedded after the third pooling layer of the fully convolutional neural network. This module performs multi-scale pooling operations on the input 64x64 feature map, using pooling windows of 2x2, 4x4, and 8x8 respectively, to obtain feature representations at three different scales, with sizes of 32x32, 16x16, and 8x8 respectively. Then, the features at these three scales are concatenated to obtain a 56x56 fused feature map, while capturing the context information in the feature map and enhancing the semantic representation ability of the features. Next, through three upsampling operations, the fused feature map is gradually restored to the same 512x512 spatial size as the original image. During the upsampling process, the bilinear interpolation algorithm is used to magnify the feature map, and the features are refined through convolutional operations. Finally, on the 512x512 feature map obtained by upsampling, a 1x1 convolutional kernel is used to classify each pixel to obtain a pixel-level segmentation result. By embedding the spatial pyramid pooling module and multi-scale feature fusion, this method can restore the spatial details of the image as much as possible while retaining the high-level semantic information, improving the accuracy of medical image segmentation.

[0083] Based on the feature map obtained by upsampling, the process of classifying each pixel through convolutional operations to obtain a pixel-level segmentation result includes:

[0084] Based on the feature map obtained by upsampling, a convolutional neural network is used to classify each pixel in the feature map. By setting the size and number of convolutional kernels, the local features of the pixels are extracted, and the class to which the pixel belongs is determined. According to the pixel classification results obtained from the convolutional operation, pixel-level segmentation of the image is performed. The softmax function is used to normalize the feature vector output by the convolution to obtain the probability distribution of each pixel belonging to each class. According to the pixel classification probability, the class to which each pixel belongs is determined to obtain the segmentation result of the image. Through the upsampling operation, the size of the feature map is increased to make the segmentation result match the resolution of the original image. The pixel classes in the segmentation result are mapped to corresponding colors or labels to obtain the final segmented image.

[0085] More specifically, in the fully convolutional neural network, by setting the convolutional kernel size to 3x3 and the number to 64, convolutional operations are performed on the upsampled feature map to extract the local features of each pixel. This step is to ensure that the detailed information around each pixel in the image can be accurately captured, so as to perform pixel classification more accurately. Then, the softmax function is used to normalize the output of the convolutional layer, converting the obtained feature vector into a probability distribution. In this process, each element in the output vector of each pixel point represents the probability that the pixel belongs to a certain class. According to the probability, the class of each pixel is determined, thus realizing the accurate segmentation of the image. Finally, by converting each pixel class in the classification result into the corresponding color or label, the final image segmentation result is generated. In this process, if the resolution of the original image is 1920x1080, the upsampling operation needs to adjust the size of the feature map to this resolution to ensure that the segmentation result has the same size as the original image, so that the segmentation result can be directly compared with the original image or further analyzed and processed.

[0086] Furthermore, based on the feature map obtained by upsampling, the process of classifying each pixel through convolutional operations to obtain the pixel-level segmentation result includes:

[0087] Use deconvolution operations to replace the traditional transposed convolution operations for upsampling and learn the convolutional kernel parameters during the upsampling process; at the same time, introduce a skip connection structure to fuse the high-resolution features of the shallow network with the high-level semantic features of the deep network to obtain the upsampling result;

[0088] Then perform pixel-level classification. Through the conditional random field model, model the spatial relationship and class dependence relationship between pixels, and post-process and optimize the segmentation result output by the fully convolutional neural network to smooth the segmentation boundary and eliminate isolated misclassified pixels.

[0089] Construct a multi-scale feature fusion module to extract multi-scale feature maps at different levels of the fully convolutional neural network. Through feature fusion operations, integrate feature information at different scales to obtain a more comprehensive and robust feature representation.

[0090] Furthermore, the process of using deconvolution operations to replace traditional transposed convolution operations for upsampling includes:

[0091] For the input low-resolution image, use a pre-trained convolutional neural network to extract deep semantic features;

[0092] According to the preset upsampling factor, perform upsampling on the deep semantic features through deconvolution operations to obtain the first upsampled feature map;

[0093] Obtain the shallow high-resolution features corresponding to the input image, and fuse the shallow high-resolution features with the first upsampled feature map through skip connections to obtain the second upsampled feature map;

[0094] According to the second upsampled feature map, further improve the resolution through deconvolution operations to obtain the third upsampled feature map;

[0095] According to the third upsampled feature map, perform detail enhancement through convolution operations to obtain the fourth upsampled feature map;

[0096] According to the fourth upsampled feature map, convert it to a high-resolution output image through pixel rearrangement operations;

[0097] According to the high-resolution output image and the original high-resolution image, calculate the reconstruction error through a loss function, and use the error backpropagation algorithm to optimize the convolution kernel parameters of the deconvolution and convolution operations.

[0098] More specifically, given a low-resolution input image with a size of 256×256, a pre-trained VGG-19 convolutional neural network is used to extract features from it, obtaining a set of 512-dimensional deep semantic feature vectors. According to the preset upsampling factor of 4, the deep semantic features are upsampled to a spatial size of 1024×1024 through deconvolution operations, obtaining the first upsampled feature map. At the same time, the shallow convolutional layer of the VGG-19 network is used to extract the high-resolution features of the input image, and through skip connections, it is element-wise added and fused with the first upsampled feature map to obtain the fused second upsampled feature map. Based on the second upsampled feature map, through a deconvolution operation with a stride of 2, the resolution of the feature map is further increased to 2048×2048, obtaining the third upsampled feature map. For the third upsampled feature map, 3 convolutional layers with a size of 3×3 are used to enhance the details, and the number of convolutional kernels is 256, 128, and 3 respectively. Finally, the fourth upsampled feature map with the same size as the target high-resolution image is obtained. Using the pixel rearrangement operation, each pixel in the fourth upsampled feature map is rearranged according to the preset mapping relationship into a high-resolution output image of 2048×2048. Through the mean square error loss function, the reconstruction error between the high-resolution output image and the original high-resolution image is measured, and using the Adam optimization algorithm, the convolutional kernel parameters in the deconvolution and convolution operations are iteratively optimized with a learning rate of 0001 to minimize the reconstruction error. After 50 epochs of training, the super-resolution reconstruction network can well learn the mapping relationship from low resolution to high resolution, effectively improving the visual quality of the image.

[0099] Furthermore, the process of performing pixel-level classification includes:

[0100] Taking the pixel-level segmentation result output by the fully convolutional neural network as the input of the conditional random field model;

[0101] Constructing a conditional random field model, defining pixel nodes and edge feature functions, and modeling the spatial relationship and class dependency relationship between pixels;

[0102] The spatial relationship feature function considers the relative position of pixels in the image, and the class dependency feature function considers the consistency of the class labels of adjacent pixels;

[0103] In the conditional random field model, the optimized pixel labels are obtained through maximum a posteriori probability inference;

[0104] The inference process considers the local features and global context information of pixels, smooths the segmentation boundary, and eliminates isolated misclassified pixels.

[0105] Furthermore, the inference process also includes post-processing the inference result of the conditional random field model; the post-processing process includes:

[0106] Using connected component analysis and morphological operations, small isolated areas are removed and holes in the segmentation results are filled;

[0107] According to the spatial positions of the pixels and the optimized class labels, the boundary contours of the segmentation results are reconstructed;

[0108] Through edge detection and contour tracking algorithms, the boundaries of the segmentation objects are obtained, and the optimized segmentation results are evaluated, and pixel-level classification accuracy and intersection over union metrics are calculated;

[0109] The optimized segmentation results are compared with the original outputs of the fully convolutional neural network to quantify the improvement effect of the conditional random field model;

[0110] The parameters of the conditional random field model are iteratively optimized, and the optimal feature function weights and inference algorithm parameters are found through cross-validation and grid search.

[0111] More specifically, pixel-level segmentation results are obtained from the fully convolutional neural network as the input of the conditional random field model. When constructing the conditional random field model, pixel node feature functions are defined, considering the RGB color values, gradient magnitudes, etc. of the pixels, and edge feature functions consider the Euclidean distance and color differences between pixels, etc. The optimized pixel labels are obtained through maximum a posteriori inference, and the mean field algorithm is used for approximate inference, with the number of iterations set to 10. During the inference process, the local feature weight of the pixels is set to 6, and the global context information weight is set to 4 to balance local consistency and global information. When post-processing the inference results, connected component analysis is used, with the area threshold set to 50 pixels, and isolated areas smaller than this threshold are removed; morphological closing operations are used, with the structuring element size of 3x3, to fill small holes. The boundaries of the segmentation objects are extracted through the Canny edge detection algorithm, with the standard deviation of the Gaussian filter being 5, the low threshold being 50, and the high threshold being 150. When evaluating the segmentation results, pixel accuracy and mean IoU are used as metrics, and the optimal parameter combination is selected through 5-fold cross-validation. During the iterative optimization process, the feature function weights are determined through grid search, with the search range from 1 to 0 and the step size being 1; the number of iterations and convergence threshold of the inference algorithm are automatically adjusted through Bayesian optimization to improve the segmentation performance.

[0112] Furthermore, a multi-scale feature fusion module is constructed, and multi-scale feature maps are extracted at different levels of the fully convolutional neural network. The process of integrating feature information of different scales through feature fusion operations includes:

[0113] According to different levels of the fully convolutional neural network, several scales of feature maps are obtained, where the feature maps contain feature information of different scales;

[0114] Perform a concatenation operation on feature maps of several scales, concatenate the feature maps of different scales in the channel dimension to obtain a concatenated multi-scale feature map;

[0115] For the concatenated multi-scale feature map, perform a weighting operation. Through the learned weight parameters, weight and fuse different-scale features in the multi-scale feature map to obtain a fused multi-scale feature representation.

[0116] According to the fused multi-scale feature representation, use the feature extraction and classifier of the convolutional neural network to determine whether there is a lesion area in the image. If there is a lesion area, determine its location. According to the determined location of the lesion area, extract the feature vector at the corresponding position from the fused multi-scale feature representation as the feature representation of the lesion area. Adopt an attention mechanism to adaptively adjust the weights of features at each scale in the multi-scale feature representation according to the feature representation of the lesion area, highlight the features of the lesion area, and suppress the features of the non-lesion area. Input the adjusted multi-scale feature representation into the subsequent network layer, and through further feature transformation and classifier, obtain the final lesion area segmentation result.

[0117] More specifically, in a fully convolutional neural network, by extracting feature maps at different levels, feature information at multiple scales is obtained. For example, in the VGG-16 network, the outputs of the 3rd, 4th, and 5th convolutional blocks can be selected as multi-scale feature maps, with scales of 1 / 8, 1 / 16, and 1 / 32 of the original image respectively. These feature maps at different scales are concatenated in the channel dimension to obtain a feature map containing multi-scale information, with the number of channels being 1024. For the concatenated multi-scale feature map, a weighted fusion method is adopted, and through the learned weight parameters, the importance of features at different scales is adaptively adjusted. For example, an attention mechanism can be used to dynamically generate the weight coefficients for each scale according to the feature representation of the lesion area, and the multi-scale feature map is weighted and summed to obtain the fused feature representation. The fused multi-scale feature representation contains semantic information and detail information at different scales and can better represent the features of the lesion area. Based on the fused feature representation, through the feature extraction and classifier of the convolutional neural network, it is judged whether there is a lesion area in the image. For example, two fully connected layers and a Softmax activation function can be used to classify the feature vector and output the probability of the lesion area. If a lesion area is detected, according to the position coordinates output by the classifier, the feature vector at the corresponding position is extracted from the fused multi-scale feature representation as the feature representation of the lesion area for subsequent segmentation tasks. To further highlight the features of the lesion area and suppress the interference of non-lesion areas, an attention mechanism can be adopted to adaptively adjust the weights of features at each scale in the multi-scale feature representation according to the feature representation of the lesion area. For example, through a fully connected layer and a Sigmoid activation function, the attention weights for each scale are generated, and the multi-scale feature map is weighted to obtain the adjusted feature representation. Finally, the adjusted multi-scale feature representation is input into the subsequent network layers, and after several convolutional layers and upsampling layers, a segmentation result with the same size as the input image is obtained. Through this method of multi-scale feature fusion and attention mechanism, different levels of semantic information and spatial details can be effectively combined to improve the segmentation accuracy of the lesion area.

[0118] Furthermore, introducing an adversarial learning strategy, the process of obtaining the final segmentation result through adversarial training of the segmentation result of the fully convolutional neural network with the true annotation by constructing a generator and a discriminator network includes:

[0119] Construct a fully convolutional neural network as the generator network, extract multi-scale features of the medical image through convolutional layers and pooling layers, and restore the spatial resolution of the feature map through an upsampling layer to obtain a pixel-level segmentation result;

[0120] Construct a discriminator network, extract the features of the segmentation result generated by the generator network and the true annotation through convolutional layers, and judge whether the segmentation result is the true annotation through a fully connected layer to obtain a binary classification discriminant result;

[0121] During the training process, the generator network and the discriminator network are alternately trained. The goal of the generator network is to generate realistic segmentation results to deceive the discriminator, and the goal of the discriminator network is to determine whether the segmentation result is a true annotation;

[0122] During the adversarial training process, a gradient penalty term is introduced to limit the gradient norm of the discriminator network. The data augmentation technique is adopted to expand the training samples. By performing rotation, translation, and scaling transformations on the medical images, diverse training samples are generated; through adversarial training, the generator network and the discriminator network play against each other, ultimately reaching a Nash equilibrium, making the segmentation results generated by the generator network indistinguishable from the true annotations, and it is also difficult for the discriminator network to distinguish between true and false;

[0123] In the testing phase, the trained generator network is used to perform segmentation prediction on unknown medical images, and post-processing techniques are adopted to optimize the segmentation results to obtain the final medical image segmentation results.

[0124] More specifically, in the generator network of the fully convolutional neural network, convolutional layers with convolutional kernel sizes of 3x3, 5x5, and 7x7 are set to extract features of the medical images at different scales. After the convolutional layers, a max-pooling operation is performed with a pooling kernel size of 2x2 and a stride of 2 to achieve downsampling of the feature map size. In the upsampling layer, a transposed convolution operation is adopted with a convolutional kernel size of 4x4 and a stride of 2 to restore the spatial resolution of the feature map to the same as that of the input medical image. The discriminator network uses 5 convolutional layers to extract features, with a convolutional kernel size of 4x4 and a stride of 2 for each layer. The LeakyReLU activation function is applied after each convolution. Finally, a binary classification result is output through a fully connected layer and a Sigmoid activation function. In each iteration, the discriminator network is trained first, and then the generator network is trained. Their learning rates are set to 0002 and 0001 respectively. At the same time, a gradient penalty term is introduced into the loss function of the discriminator network, and the penalty factor is set to 10 to limit the gradient norm of the discriminator network not to exceed 1. When performing data augmentation on the medical images, the random rotation angle range is from -10° to 10°, the random translation distance range is from -10% to 10% of the width and height of the image, and the random scaling factor range is from 8 to 2. In the post-processing stage, small regions with an area less than 100 pixels are removed through connected component analysis, and morphological closing operations are used to fill small holes in the segmentation results. The structuring element is set as a 3x3 rectangle.

[0125] The above is only a preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A medical image segmentation method based on a fully convolutional neural network, characterized in that, Including: Preprocessing medical images through a deep learning algorithm to enhance the key features of the images; During the process of feature extraction based on a fully convolutional neural network, introducing an attention mechanism module, and by learning the importance weights of different regions in the image, adaptively adjusting the weight distribution of the convolutional kernel to extract fine-grained and discriminative image features; Through a spatial pyramid pooling module embedded in the fully convolutional neural network, performing multi-scale feature fusion and capturing context information; Obtaining a fused feature map through multi-scale feature fusion, and gradually restoring the fused feature map to the same spatial size as the original image through an upsampling operation; Based on the feature map obtained by upsampling, classifying each pixel through a convolutional operation to obtain a pixel-level segmentation result; Introducing an adversarial learning strategy, and by constructing a generator and a discriminator network, performing adversarial training on the segmentation result of the fully convolutional neural network and the true annotation to obtain the final segmentation result; The process of classifying each pixel through a convolutional operation based on the feature map obtained by upsampling to obtain a pixel-level segmentation result includes: Using a transposed convolution operation to replace the traditional transposed convolution operation for upsampling and learning the convolutional kernel parameters during the upsampling process; at the same time, introducing a skip connection structure to fuse the high-resolution features of the shallow network and the high-level semantic features of the deep network to obtain an upsampling result; Then performing pixel-level classification, and through a conditional random field model, modeling the spatial relationship and class dependency relationship between pixels to post-process and optimize the segmentation result output by the fully convolutional neural network, smoothing the segmentation boundary and eliminating isolated misclassified pixels; Constructing a multi-scale feature fusion module, extracting multi-scale feature maps at different levels of the fully convolutional neural network, and integrating different-scale feature information through a feature fusion operation to obtain a more comprehensive and robust feature representation.

2. The medical image segmentation method based on a fully convolutional neural network according to claim 1, wherein The process of preprocessing medical images through a deep learning algorithm to enhance the key features of the images includes: Obtaining medical image data, and analyzing and identifying the complex anatomical structures and pathological features in the images using a deep learning algorithm; Preprocessing the medical images, adjusting the gray-scale distribution of the images through a histogram equalization algorithm to enhance the contrast of the images; According to the characteristics of the images, using an adaptive contrast adjustment algorithm to specifically enhance the contrast of local regions of the images and highlight the pathological features.

3. The medical image segmentation method based on a fully convolutional neural network according to claim 1, characterized in that The process of introducing an attention mechanism module to extract fine-grained and discriminative image features includes: Constructing a fully convolutional neural network model, introducing an attention mechanism module in the feature extraction stage. In the attention mechanism module, by calculating the importance weights of different regions in the medical image, generating an attention weight matrix, and performing element-wise multiplication of the attention weight matrix and the weight matrix of the convolutional kernel to obtain an adjusted convolutional kernel weight distribution; Using the adjusted convolutional kernel to perform a convolutional operation on the medical image to extract the features of the key anatomical structures and lesion regions in the image, obtaining a fine-grained and discriminative feature map and the corresponding image features.

4. The medical image segmentation method based on a fully convolutional neural network according to claim 1, characterized in that The process of multi-scale feature fusion and context information capture through the spatial pyramid pooling module embedded in the fully convolutional neural network includes: Based on the spatial pyramid pooling module, by performing multi-scale pooling operations on fine-grained and discriminative image features, feature representations at different scales are obtained; Based on the feature representations at different scales, multi-scale feature fusion is performed, while capturing the context information in the feature map to enhance the semantic representation ability of the features.

5. The medical image segmentation method based on a fully convolutional neural network according to claim 1, wherein The process of using deconvolution operations to replace traditional transposed convolution operations for upsampling includes: For the input low-resolution image, deep semantic features are extracted using a pre-trained convolutional neural network; According to the preset upsampling multiple, the deep semantic features are upsampled through deconvolution operations to obtain a first upsampled feature map; The shallow high-resolution features corresponding to the input image are obtained, and the shallow high-resolution features are fused with the first upsampled feature map through skip connections to obtain a second upsampled feature map; According to the second upsampled feature map, the resolution is further improved through deconvolution operations to obtain a third upsampled feature map; According to the third upsampled feature map, convolution operations are used for detail enhancement to obtain a fourth upsampled feature map; According to the fourth upsampled feature map, it is converted into a high-resolution output image through pixel rearrangement operations; According to the high-resolution output image and the original high-resolution image, the reconstruction error is calculated through a loss function, and the convolution kernel parameters of the deconvolution and convolution operations are optimized using the error backpropagation algorithm.

6. The medical image segmentation method based on a fully convolutional neural network according to claim 1, wherein The process of performing pixel-level classification includes: Taking the pixel-level segmentation result output by the fully convolutional neural network as the input of the conditional random field model; Constructing a conditional random field model, defining pixel nodes and edge feature functions, and modeling the spatial relationship and class dependence relationship between pixels; The spatial relationship feature function considers the relative position of pixels in the image, and the class dependence feature function considers the consistency of the class labels of adjacent pixels; In the conditional random field model, the optimized pixel labels are obtained through maximum a posteriori probability inference; The inference process considers the local features and global context information of pixels, smooths the segmentation boundary, and eliminates isolated misclassified pixels.

7. The medical image segmentation method based on a fully convolutional neural network according to claim 6, wherein The inference process also includes post-processing the inference result of the conditional random field model; The post-processing process includes: Using connected component analysis and morphological operations to remove small isolated areas and fill the holes in the segmentation result; According to the spatial position of the pixels and the optimized class labels, the boundary contour of the segmentation result is reconstructed; Through edge detection and contour tracking algorithms, the boundary of the segmented object is obtained, and the optimized segmentation result is evaluated, and pixel-level classification accuracy and intersection over union metrics are calculated; The optimized segmentation result is compared with the original output of the fully convolutional neural network to quantify the improvement effect of the conditional random field model; Iteratively optimize the parameters of the conditional random field model, and find the optimal feature function weights and inference algorithm parameters through cross-validation and grid search.

8. The medical image segmentation method based on a fully convolutional neural network according to claim 1, characterized in that The process of constructing a multi-scale feature fusion module, extracting multi-scale feature maps at different levels of the fully convolutional neural network, and integrating different-scale feature information through feature fusion operations includes: According to different levels of the fully convolutional neural network, several feature maps of different scales are obtained, where the feature maps contain feature information of different scales; Perform a concatenation operation on the feature maps of the several scales, concatenate the feature maps of different scales in the channel dimension to obtain a concatenated multi-scale feature map; For the concatenated multi-scale feature map, perform a weighting operation, and through the learned weight parameters, weight and fuse the features of different scales in the multi-scale feature map to obtain a fused multi-scale feature representation.

9. The medical image segmentation method based on a fully convolutional neural network according to claim 1, wherein Introduce an adversarial learning strategy. The process of obtaining the final segmentation result by constructing a generator and a discriminator network and performing adversarial training on the segmentation result of the fully convolutional neural network and the true annotation includes: Construct a fully convolutional neural network as the generator network, extract multi-scale features of the medical image through convolutional layers and pooling layers, and restore the spatial resolution of the feature map through an upsampling layer to obtain a pixel-level segmentation result; Construct a discriminator network, extract the features of the segmentation result generated by the generator network and the true annotation through convolutional layers, and judge whether the segmentation result is the true annotation through a fully connected layer to obtain a binary classification discriminant result; During the training process, alternately train the generator network and the discriminator network. The goal of the generator network is to generate a realistic segmentation result to deceive the discriminator, and the goal of the discriminator network is to judge whether the segmentation result is the true annotation; Introduce a gradient penalty term during the adversarial training process to limit the gradient norm of the discriminator network. Adopt data augmentation techniques to expand the training samples. Generate diverse training samples by rotating, translating, and scaling the medical images. Through adversarial training, enable the generator network and the discriminator network to play against each other, and finally reach a Nash equilibrium, making the segmentation result generated by the generator network indistinguishable from the true annotation, and the discriminator network also difficult to distinguish the true from the false; In the testing stage, use the trained generator network to perform segmentation prediction on unknown medical images, and adopt post-processing techniques to optimize the segmentation result to obtain the final medical image segmentation result.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation convolutional neural network method fusing diffusion semantic features

    CN118570614A