A medical image segmentation method based on long-short distance features

By designing a medical image segmentation network based on long and short distance features, using the MEtransformer module and the global local feature fusion module, the shortcomings of the existing medical image segmentation methods in feature modeling and computing resources are solved, and efficient and stable medical image segmentation effect is achieved.

CN114463341BActive Publication Date: 2025-05-06WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210026011.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-11
Publication Date
2025-05-06
Estimated Expiration
2042-01-11

AI Technical Summary

Technical Problem

The existing medical image segmentation method has room for improvement in the feature modeling of lesion regions, and the network model occupies a large amount of space and calculation, making it difficult to meet the stability and immediacy requirements of medical-assisted diagnostic systems.

Method used

Using a medical image segmentation network based on long and short distance features, a network model of encoder and decoder based on Transformer and convolutional network is designed to realize the segmentation of medical images through efficient Transformer (MEtransformer) module with mask, convolutional feature extraction module, global local feature fusion module (TCFuse) and deconvolution decoder module.

Benefits of technology

This method shows good segmentation performance on the lesion image datasets in many different regions, with stable segmentation effect and clear edges. By optimizing the network structure and loss function, the parameter quantity and calculation quantity are reduced, meeting the stability and immediacy requirements of the medical-assisted diagnostic system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463341B_ABST
    Figure CN114463341B_ABST
Patent Text Reader

Abstract

The present invention relates to a medical image segmentation method based on long-short distance features. A medical image segmentation network based on long-short distance features is constructed using the Pytorch deep learning framework, and a design mode of an encoder and a decoder based on a transformer and a convolutional network is adopted. Through the processing of four parts, namely, a ME transformer module, a convolutional feature extraction module, a global-local feature fusion module, and a deconvolution decoder module, suspicious lesion areas are segmented from input medical images. The present invention has good segmentation performance for lesion image data sets in various different regions, with stable segmentation effects and clearer edges.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image segmentation, and in particular relates to a medical image segmentation method based on long-short distance features. Background Art

[0002] With the advancement of medical technology, various new medical devices have been introduced. The internal images of the human body collected by these devices (common ones include CT, MRI, etc.) have greatly facilitated the medical diagnosis process. In the past, image analysis was mainly completed by professional physicians. However, due to the scarcity of medical experts, a computer software-based auxiliary diagnosis system is needed clinically to automatically divide the lesion area in the image. How to improve the segmentation accuracy and speed of the auxiliary system is one of the current research hotspots.

[0003] Deep learning is currently the most popular image processing method. It performs supervised learning through correctly labeled image samples to complete tasks such as classification, segmentation, and detection. Compared with natural image segmentation, the main difference in medical image segmentation lies in the data aspect: first, medical image annotation requires medical expertise, and it is more difficult and expensive to obtain data sets than natural images, so the sample size of public data sets is relatively small; second, medical images are more similar in morphology and color, and the lesion areas are more concealed; finally, medical image segmentation technology is generally used in medical auxiliary systems, which have higher requirements for stability and immediacy.

[0004] At present, scholars have done a lot of research in the field of medical image segmentation. Ronneberger et al. proposed a U-type network (Unet) based on encoder and decoder, which added jump connections on the basis of encoding and decoding, and added the underlying edge features to the decoder by superposition operation to generate segmentation results, thereby improving the segmentation accuracy. Oktar et al. proposed a Unet network structure with an attention mechanism, which supervised the extraction process of shallow shape features through deep semantic features to achieve feature optimization, and used the output of the decoder as the result of the segmentation task on the basis of optimization. Chen et al. introduced the Transformer structure based on multi-head attention in natural language processing into Unet to better extract global context information and perform more effective feature modeling to generate output results, thereby further improving the segmentation performance. However, the methods proposed by these authors still need to be improved in segmentation performance, and lack consideration for the space occupied by the network model and the amount of calculation. With the development of deep learning technology and supercomputing chips, a medical image segmentation technology that takes into account both accuracy and immediacy is needed to overcome the above-mentioned shortcomings.

[0005] In summary, the existing medical image segmentation methods still have room for improvement in the modeling of lesion area features. A network model architecture that more comprehensively and effectively integrates global and local features is needed to adapt to the complex and changeable medical detection images in clinical applications. At the same time, in order to meet the stability and immediacy of the medical auxiliary diagnosis system, the proposed network model must meet the requirements of smaller parameters and computational complexity in order to be deployed in the clinical application server. In addition, the network model also has a certain flexibility in inference time and accuracy, and can achieve a balance between accuracy and running time on different clinical devices. Summary of the invention

[0006] In view of the deficiencies of the prior art, the present invention provides a medical image segmentation method based on long-short distance features. The Pytorch deep learning framework is used to build a medical image segmentation network based on long-short distance features, and the design mode of the encoder and decoder based on the transformer and convolutional network is adopted. Through the processing of the four parts of the efficient transformer (MEtransformer) module with mask, the convolutional feature extraction module, the global-local feature fusion module (TCFuse) and the deconvolution decoder module, the suspicious lesion area is segmented from the input medical image. The present invention has good segmentation performance for lesion image data sets in various different regions, the segmentation effect is stable, and the edges are clearer.

[0007] In order to achieve the above object, the technical solution provided by the present invention is a medical image segmentation method based on long-short distance features, comprising the following steps:

[0008] Step 1: First, crop the images in the training set to a size of C×H×W, randomly flip the cropped training set images up and down and horizontally to achieve data expansion, and then divide them into training set and test set;

[0009] Step 2, construct a medical image segmentation network based on long and short distance features;

[0010] Step 3, using the training set images to train the medical image segmentation network based on long and short distance features;

[0011] Step 4: Use the segmentation network trained in step 3 to perform medical image segmentation.

[0012] Moreover, the medical image segmentation network based on long and short distance features in step 2 adopts the design mode of encoder and decoder based on transformer and convolution network, and the encoder module is composed of convolution feature extraction module, MEtransformer module and global local feature fusion module. First, the image tensor is subjected to preliminary feature extraction through two convolution feature extraction modules, and then the long-distance global features are extracted by MEtransformer module in parallel, and the short-distance local features are extracted by convolution feature extraction module. Then, the long and short distance feature vectors are fused by global local feature fusion module. On the basis of the fused feature map, the aforementioned parallel feature extraction and feature vector fusion operations are performed again to further model the long and short distance features in the data, and finally the segmentation result output of the network is obtained through three deconvolution decoder modules. The medical image segmentation network model iteratively learns the parameters through the back propagation of multiple loss functions to realize the automatic optimization of the network, and medical image segmentation can be realized after multiple rounds of training.

[0013] The convolutional feature extraction module consists of multiple blocks based on convolutional layers. Each block consists of three convolutional layers with convolution kernel sizes of 1×1, 3×3, and 1×1, as well as a normalization layer. After each normalization, it will pass through the relu activation function to ensure the distribution of feature activation. The feature map generated by each block will be saved as a skip connection of the low-order feature and sent to the subsequent deconvolution decoding operation. The final output of this module is the short-distance feature map.

[0014] The MEtransformer module includes an axial multi-head attention module and a mask module. The input of the axial multi-head attention module is an image tensor (B, C, H, W). First, the image is divided into 2×2 patches. Each patch is mapped to a vector of length C through the fully connected layer, and these N vectors are combined in the original order into a vector of size The feature map is transformed into and Two vectors, then divide channel C into multiple equal parts, and turn the two vectors into and On this basis, matrices Q, K and V are obtained respectively through matrix multiplication; the mask module performs Gaussian weighted calculation of the attention mechanism on the query set generated by the product of the Q and K matrices to guide the module to focus on long-distance information and thus reduce the weight of close-range features. Finally, the weighted query set is matrix multiplied with V to obtain the output of the module.

[0015] The input of the global-local feature fusion module is the long-distance feature map and the short-distance feature map, both of which are (B, C, H, W) in scale. First, the long-distance feature is globally averaged pooled to obtain a channel-based attention map, and then it is element-wise multiplied with the short-distance feature map and passed through a convolutional layer for preliminary feature fusion. The fused feature is then stacked with the original long-distance feature and passed through the convolutional layer again to achieve further fusion and feature dimensionality reduction, thereby generating the subsequent required high-order feature blocks.

[0016] The deconvolution decoder module consists of multiple decoder blocks based on convolution and bilinear interpolation. The features output by the second global-local feature fusion module are bilinearly upsampled and interpolated, and H and W are doubled. Then, they are stacked with skip features and then go through two layers of convolution for multi-level fusion.

[0017] Moreover, in step 3, the image tensor B×C×H×W is combined with batch size B as the training input of the network, and the network model hyperparameters are set to learn_rate=0.01, momentum=0.9, weight_decay=0.0001. After 150 epochs of iterative optimization, the loss function used in training is as follows:

[0018] Loss=0.5×Cross Entropy Loss+0.5×Dice Loss (1)

[0019] In the formula, Cross Entropy Loss represents the cross entropy loss function value, and Dice Loss is the set similarity measurement function, which is used to calculate the similarity between two samples.

[0020] Moreover, the medical image segmentation in step 4 includes an encoding stage and a decoding stage.

[0021] Encoding stage: Initially set the batch size to B, the original B×C×H×W image tensor enters two serially connected convolutional feature extraction modules to perform preliminary low-dimensional feature extraction and generate two skip features; then the low-dimensional features are sent to the MEtransformer module and another convolutional feature extraction module respectively, and the sizes of both are Long-distance global features and short-distance local features will also generate a skip feature at this time, a total of 3 skip features. The MEtransformer module converts the features into axial attention based on the transformer's multi-head attention to reduce the amount of calculation. The H and W sizes are both 32×32, and the mask of the feature map is added. The mask size is 32×32. The weight of the short-distance channel is reduced to ensure the extraction of long-distance features; then the long-distance and short-distance feature maps are sent to the global and local feature fusion module, and the long-distance features are used as input to generate C2×1×1 channel weights on the channel multiplied by the short-distance features, and then convolution compression and stacking are performed to obtain preliminary fusion features. The size is still Further long and short distance feature extraction and fusion are performed again to further compress the features into At this point, the task of image data feature compression is completed.

[0022] Decoding stage: The input is The compressed features generated by the encoding structure are first upsampled to reduce the number of channels and increase the length and width, and then stacked with the third skip generated in the encoder stage and sent to two convolutional layers. They pass through the decoder module three times in succession, corresponding to the three skip features respectively, and finally get The feature map is then passed through a convolution layer and an upsampling layer whose output channel is the number of categories to obtain the final segmentation result.

[0023] Compared with the prior art, the present invention has the following advantages:

[0024] 1) In order to solve the problem of insufficient feature extraction in medical image segmentation, a mask-based axial attention feature extraction module MEtransformer is designed. Based on transform, this module transforms different token mappings into H and W directions, calculates correlations in both directions, and designs a mask module based on Gaussian function on the correlation importance graph, which increases the weight of long-distance features and reduces the weight of short-distance features. The mask module is designed to guide the multi-head attention structure to effectively model long-distance features, and complement the advantages of the parallel convolutional feature extraction module used to model short-distance features, thereby further improving the image feature modeling capability of the entire structure and making full use of the spatial and positional information of the original image.

[0025] 2) In order to solve the problem of fusing long-distance and short-distance features, a long-short-distance feature fusion module is proposed. Long-distance features have better global information representation and play an important guiding role in the fusion process. Short-distance features have stronger expression ability for detail information and play a role in supplementing details. First, the long-distance features are used to generate a channel attention vector 1×1×C, which is multiplied with the short-distance features in the channel dimension, and then stacked. While H and W remain unchanged, the C dimension is expanded. Finally, a convolution layer with a stride of 1 is introduced to perform channel dimensionality reduction to fully fuse the information of both.

[0026] 3) In order to address the problem of large device memory and large amount of computation required by deep networks, the network structure was adjusted. On the one hand, axial attention was introduced to avoid ultra-large-scale matrix calculations and reduce the amount of computation to a reasonable level. On the other hand, the encoder and decoder hyperparameters of the network model were adjusted to remove most of the redundant parameters without affecting the accuracy, saving storage space and improving the efficiency of device use. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a schematic diagram of the medical image segmentation network based on long and short distance features of the present invention.

[0028] Figure 2 It is a structural diagram of the MEtransforme module of the present invention.

[0029] Figure 3 It is a structural diagram of the global-local feature fusion module of the present invention.

[0030] Figure 4 This is a structural diagram of the convolutional feature extraction module of the present invention.

[0031] Figure 5 It is a structural diagram of the deconvolution decoder module of the present invention.

[0032] Figure 6 This is the effect diagram of the image segmentation of the gastric polyp lesion area of ​​the present invention, wherein Figure 6 (a) is the image of the gastric polyp lesion area. Figure 6 (b) is the image segmentation effect of the gastric polyp lesion area.

[0033] Figure 7 This is the effect diagram of cell image pathology recognition and segmentation of the present invention, where Figure 7 (a) is the cell image. Figure 7 (b) is the effect diagram of cell image pathology recognition and segmentation. DETAILED DESCRIPTION

[0034] The present invention provides a medical image segmentation method based on long-short distance features. The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0035] like Figure 1 As shown, the process of the embodiment of the present invention includes the following steps:

[0036] Step 1: First, crop the images in the training set to a size of 3×512×512, randomly flip the cropped training set images up and down and horizontally to achieve data expansion, and then divide the training set and test set into a ratio of 8:2, that is, use 80% of the data for training and 20% of the data for testing the training results.

[0037] Step 2: Construct a medical image segmentation network based on long and short distance features.

[0038] The design pattern of encoder and decoder based on transformer and convolutional network is adopted. The encoder module consists of convolutional feature extraction module, MEtransformer module and global-local feature fusion module. First, the image tensor is subjected to preliminary feature extraction through two convolutional feature extraction modules, and then the MEtransformer module is used in parallel to extract long-distance global features, and the convolutional feature extraction module is used to extract short-distance local features. Then, the global-local feature fusion module is used to fuse the long-distance and short-distance feature vectors. On the basis of the fused feature map, the aforementioned parallel feature extraction and feature vector fusion operations are performed again to further model the long-distance and short-distance features in the data. Finally, the segmentation result output of the network is obtained through three deconvolution decoder modules. The medical image segmentation network model iteratively learns the parameters through back propagation of multiple loss functions to achieve automatic optimization of the network. Medical image segmentation can be achieved after multiple rounds of training.

[0039] Convolutional feature extraction module: This module consists of multiple blocks based on convolutional layers. Each block consists of three convolutional layers with convolution kernel sizes of 1×1, 3×3, and 1×1, as well as a normalization layer. After each normalization, it will pass through the relu activation function to ensure the distribution of feature activation. The feature map generated by each block will be saved as a skip connection of the low-order feature and sent to the subsequent deconvolution decoding operation. The final output of this module is the short-distance feature map.

[0040] MEtransformer module: This module includes an axial multi-head attention module and a mask module. The input of the axial multi-head attention module is an image tensor (B, C, H, W). First, the image is divided into 2×2 patches. Each patch is mapped to a vector of length C through the fully connected layer, and these N vectors are combined in the original order into a vector of size The feature map is converted into the feature map by taking the average value in the H and W directions respectively. and Two vectors, then divide channel C into multiple equal parts, and turn the two vectors into and On this basis, the matrices Q, K, and V are obtained respectively through matrix multiplication. The mask module performs Gaussian weighted calculation of the attention mechanism on the query set generated by the product of the Q and K matrices to guide the module to focus on long-distance information and reduce the weight of short-distance features. Finally, the weighted query set is matrix multiplied with V to obtain the output of the module.

[0041] Global local feature fusion module (TCFuse): The input of this module is the long-distance feature map and the short-distance feature map, both of which are (B, C, H, W) in scale. First, the long-distance feature is globally averaged pooled to obtain a channel-based attention map, and then it is element-wise multiplied with the short-distance feature map and passed through a convolutional layer for preliminary feature fusion. Then, the fused feature is stacked with the original long-distance feature and passed through the convolutional layer again to achieve further fusion and feature dimensionality reduction, thereby generating the subsequent required high-order feature blocks.

[0042] Deconvolution decoder module: This module consists of multiple decoder blocks based on convolution and bilinear interpolation. After bilinear upsampling and interpolation of the features output by the second global-local feature fusion module, H and W are doubled, and then stacked with skip features, and then subjected to two layers of convolution for multi-level fusion.

[0043] Step 3: Use the training set images to train the medical image segmentation network based on long and short distance features.

[0044] The batch size is 4, and a 4×3×512×512 image tensor is formed as the training input of the network. The network model hyperparameters are set to learn_rate=0.01, momentum=0.9, and weight_decay=0.0001. After 150 epochs of iterative optimization, the loss function used in training is as follows:

[0045] Loss=0.5×Cross Entropy Loss+0.5×Dice Loss (1)

[0046] In the formula, Cross Entropy Loss represents the cross entropy loss function value, and Dice Loss is the set similarity measurement function, which is used to calculate the similarity between two samples.

[0047] Step 4: Use the segmentation network trained in step 3 to perform medical image segmentation.

[0048] The segmentation network for medical image segmentation includes two stages: encoding and decoding.

[0049] Encoding stage: The batch size is initially set to 4, and the original 4×3×512×512 image tensor enters two serially connected convolutional feature extraction modules to perform preliminary low-dimensional feature extraction and generate two skip features, such as Figure 4 As shown. Then the low-dimensional features are sent to the MEtransformer module and another convolutional feature extraction module respectively, and the long-distance global features and short-distance local features of size 4×512×64×64 are extracted respectively (a skip feature is also generated at this time, a total of 3 skip features). The MEtransformer module converts the features into axial attention based on the multi-head attention of the transformer to reduce the amount of calculation. The H and W sizes are both 32×32, and the mask mask (32×32) of the feature map is added to ensure the extraction of long-distance features by reducing the weight of the short-distance channel. Then the long-distance and short-distance feature maps are sent to the global and local feature fusion module. The long-distance features are used as input to generate 512×1×1 channel weights on the channel and multiply them to the short-distance features. Then, convolution compression and stacking are performed to obtain preliminary fusion features (size is still 4×512×64×64), and further long-distance and short-distance feature extraction and fusion are performed again, and the features are further compressed to 4×1024×32×32. At this point, the task of image data feature compression is completed.

[0050] Decoding stage: The input is the compressed features generated by the 4×1024×32×32 encoding structure. It is first upsampled to reduce the number of channels and increase the length and width. It is then stacked with the third skip generated in the encoder stage and fed into two convolutional layers, such as Figure 5 As shown in the figure, after three consecutive decoder modules (corresponding to the three skip features respectively), a 4×127×128×128 feature map is finally obtained, and then this feature map is passed through a convolution layer and an upsampling layer with an output channel equal to the number of categories to obtain the final segmentation result.

[0051] During specific implementation, the above process can be automatically operated using computer software technology.

[0052] The specific embodiments described herein are merely examples of the spirit of the present invention. Those skilled in the art may make various modifications or additions to the specific embodiments described or replace them in similar ways, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A medical image segmentation method based on long-short distance features, characterized in that: The following steps are involved: Step 1: First, crop the images in the training set to a size of C×H×W, randomly flip the cropped training set images up and down and horizontally to achieve data expansion, and then divide them into training set and test set; Step 2, construct a medical image segmentation network based on long and short distance features; The design mode of encoder and decoder based on transformer and convolutional network is adopted. The encoder module consists of convolutional feature extraction module, MEtransformer module and global-local feature fusion module. First, the image tensor is subjected to preliminary feature extraction through two convolutional feature extraction modules, and then the MEtransformer module is used in parallel to extract long-distance global features, and the convolutional feature extraction module is used to extract short-distance local features. Then, the global-local feature fusion module is used to fuse the long-distance and short-distance feature vectors. On the basis of the fused feature map, the aforementioned parallel feature extraction and feature vector fusion operations are performed again to further model the long-distance and short-distance features in the data. Finally, the segmentation result output of the network is obtained through three deconvolution decoder modules. The MEtransformer module includes an axial multi-head attention module and a mask module. The input of the axial multi-head attention module is the image tensor (B, C, H, W). First, the image is divided into 2×2 patches. Each patch is mapped to a vector of length C through the fully connected layer, and these N vectors are combined in the original order into a vector of size The feature map is transformed into and Two vectors, then divide channel C into multiple equal parts, and turn the two vectors into and On this basis, matrices Q, K and V are obtained respectively through matrix multiplication; The mask module performs Gaussian weighted calculation of the attention mechanism on the query set generated by the product of the Q and K matrices, and finally performs matrix multiplication of the weighted query set with V to obtain the output of the module; The input of the global-local feature fusion module is the long-distance feature map and the short-distance feature map, both of which are (B, C, H, W) in scale. First, the long-distance feature is globally averaged pooled to obtain a channel-based attention map, and then it is element-wise multiplied with the short-distance feature map and passed through a convolutional layer for preliminary feature fusion. Then, the fused feature is stacked with the original long-distance feature and passed through the convolutional layer again to generate the subsequent high-order feature blocks required. Step 3, using the training set images to train the medical image segmentation network based on long and short distance features; Step 4: Use the segmentation network trained in step 3 to perform medical image segmentation.

2. A medical image segmentation method based on long-short distance features as claimed in claim 1, characterized in that: The convolutional feature extraction module in step 2 is composed of multiple blocks based on convolutional layers. Each block consists of three convolutional layers with convolution kernel sizes of 1×1, 3×3, and 1×1, and a normalization layer. After each normalization, it will pass through the relu activation function to ensure the distribution of feature activation. The feature map generated by each block will be saved as a skip connection of the low-order feature and sent to the subsequent deconvolution decoding operation. The final output of this module is the short-distance feature map.

3. The medical image segmentation method based on long-short distance features as claimed in claim 1, characterized in that: The deconvolution decoder module in step 2 is composed of multiple decoder blocks based on convolution and bilinear interpolation. The features output by the second global-local feature fusion module are bilinearly upsampled and interpolated, and H and W are doubled. Then, they are stacked with skip features and then go through two layers of convolution for multi-level fusion.

4. The medical image segmentation method based on long-short distance features as claimed in claim 1, characterized in that: In step 3, the image tensor B×C×H×W is combined with batch size B as the training input of the network. The network model hyperparameters are set to learn_rate=0.01, momentum=0.9, weight_decay=0.0001. After 150 epochs of iterative optimization, the loss function used in training is as follows: Loss=0.5×Cross Entropy Loss+0.5×Dice Loss (1) In the formula, Cross Entropy Loss represents the cross entropy loss function value, and Dice Loss is the set similarity measurement function, which is used to calculate the similarity between two samples.

5. The medical image segmentation method based on long-short distance features as claimed in claim 1, characterized in that: In step 4, the medical image segmentation includes the encoding stage and the decoding stage. In the encoding stage, the batch size is initially set to B, and the original B×C×H×W image tensor enters two serially connected convolutional feature extraction modules to perform preliminary low-dimensional feature extraction and generate two skip features; then the low-dimensional features are respectively sent to the MEtransformer module and another convolutional feature extraction module to obtain a size of Long-distance global features and short-distance local features will also generate a skip feature at this time, a total of 3 skip features. The MEtransformer module converts the features into axial attention based on the transformer's multi-head attention to reduce the amount of calculation. The H and W sizes are both 32×32, and the mask of the feature map is added. The mask size is 32×32. The weight of the short-distance channel is reduced to ensure the extraction of long-distance features; Then the long and short distance feature maps are sent to the global local feature fusion module. The long distance feature is used as input to generate C2×1×1 channel weights on the channel and multiply them by the short distance feature. Then convolution compression and stacking are performed to obtain the preliminary fusion feature. The size is still Further long and short distance feature extraction and fusion are performed again to further compress the features into At this point, the task of image data feature compression is completed.

6. A medical image segmentation method based on long-short distance features as claimed in claim 5, characterized in that: In step 4, the decoding phase: the input is The compressed features generated by the encoding structure are first upsampled to reduce the number of channels and increase the length and width, and then stacked with the third skip generated in the encoder stage and sent to two convolutional layers. They pass through the decoder module three times in succession, corresponding to the three skip features respectively, and finally get The feature map is then passed through a convolution layer and an upsampling layer whose output channel is the number of categories to obtain the final segmentation result.