Lung CT image segmentation method based on hybrid deep convolution and state space model
By using a mixed depth convolution and state space model method in lung CT image segmentation, the problem of insufficient accuracy and high computational volume in the prior art is solved, and efficient lung CT image segmentation is achieved.
Patent Information
- Application Number
- CN202510040725.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-10
AI Technical Summary
The existing lung image segmentation method is not accurate enough due to the restricted local receptive field, and the calculation is large and the efficiency is low when processing high-resolution images.
The lung CT image segmentation method based on mixed depth convolution and state space model is adopted, and the Patch Embedding layer, encoder layer, decoder layer and Linear Projection layer are connected through the MCMamba network, and local and global features are extracted using the MC-SSM module and MixConv module to reduce the calculation cost.
The balance between capturing remote information and reducing computational costs is achieved, and the accuracy and efficiency of lung CT image segmentation is improved.
Smart Images

Figure CN119941757A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image assisted interpretation, and in particular to a lung CT image segmentation method based on hybrid deep convolution and state space model. Background Art
[0002] Lung image segmentation is an important research topic in medical image processing. Traditional lung disease diagnosis mainly relies on the experience of professional doctors and naked eye observation, which may lead to misdiagnosis and other problems. Introducing deep learning models into medical image assisted interpretation can effectively assist doctors in making judgments, save time and labor costs, and optimize the allocation of medical resources.
[0003] Currently, CNN-based models or Transformer-based models are commonly used in lung image segmentation tasks. However, both CNN-based models and Transformer-based models have inherent limitations: CNN-based models can effectively extract local features, but due to the limitation of local receptive fields, their ability to capture remote information is limited; Transformer-based models can effectively capture long-distance dependencies, but because the self-attention mechanism it uses scales quadratically with the input size, its high computational cost poses a challenge.
[0004] Therefore, there is an urgent need for a medical segmentation model that can effectively capture remote information and has low computational cost. Summary of the invention
[0005] In order to solve the problem that the existing lung image segmentation methods have inaccurate segmentation results due to limited local receptive fields and high computational complexity and low efficiency when processing high-resolution images, the present invention designs an efficient medical segmentation model by fusing the state space model and hybrid deep convolution.
[0006] The lung CT image segmentation method based on hybrid deep convolution and state space model of the present invention includes:
[0007] Acquire a lung CT image to be processed; input the lung CT image to be processed into a trained lung CT image segmentation model to obtain a segmentation result; wherein the lung CT image segmentation model adopts an MCMamba network; the MCMamba network is composed of a Patch Embedding layer, an encoder layer, a decoder layer, and a Linear Projection layer connected in sequence; the encoder layer is composed of a MC-SSM module and a Patch Merging module; the MC-SSM module includes a MixConv module, an SSM module, a fusion layer, two 1×1 convolutional layers, and a normalization layer.
[0008] The beneficial effects of the present invention include:
[0009] By combining the state-space model and hybrid deep convolution, the present invention can enable the model to accurately extract local features and establish long-distance dependencies, further effectively capture global information, and improve segmentation accuracy; it provides a design idea for model research on medical image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a structural diagram of the MCMamba network described in the present invention;
[0011] Figure 2 It is a structural diagram of the MC-SSM module of the present invention;
[0012] Figure 3 It is a structural diagram of the MixConv module described in the present invention;
[0013] Figure 4 It is a structural diagram of the SSM module of the present invention;
[0014] Figure 5 It is a test chart of the embodiment of the present invention under different indicators. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solution, characteristics and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0016] Obtain a lung CT image to be processed; input the lung CT image to be processed into a trained lung CT image segmentation model to obtain a segmentation result.
[0017] Specifically, the two lung CT images used for training in this embodiment are CXR lung contour images and Convd19 pneumonia segmentation images caused by the new crown; the image size of the selected data set is 256×256 pixels, and the images are randomly flipped and randomly rotated and then normalized to obtain a training data set; training the lung CT image segmentation model includes:
[0018] S1. Input the lung CT image in the training data into the Patch Embedding layer for division and mapping to obtain an embedded image; and normalize the embedded image.
[0019] The division and mapping are processed in the Patch Embedding layer. The Patch Embedding layer divides the lung CT image into 4*4 non-overlapping blocks, each of which is 64*64 pixels in size, maps the non-overlapping blocks to a fixed dimension, defines a mapping dimension C, and C is preferably 96, to obtain an embedded image X′. The embedded image is normalized and operated using Layer Normalization.
[0020] S2: Send the normalized embedded image to the encoder layer, and process it through the MC-SSM module and the Patch Embedding module to obtain the encoded image.
[0021] The encoder layer is composed of an MC-SSM module and a Patching Merging module; the MC-SSM module is used to extract local and global features of the image, and the Patching Merging module is used to reduce the height and width of the feature map and increase the number of channels; the encoder layer and the decoder layer are jump-connected, and the data processing process includes four stages, in which the jump connection uses an addition operation, and each stage performs different degrees of processing and information abstraction on the input feature map, including:
[0022] The first stage: mainly shallow feature extraction, retaining more spatial information. The input of the encoder layer passes through two connected MC-SSM modules and a Patch Merging module in turn to obtain features with a channel number of C Figure 1 .
[0023] The second stage: focusing on extracting more structural information, features Figure 1 After two connected MC-SSM modules and a Patch Merging module, a feature with 2C channels is obtained. Figure 2 .
[0024] The third stage: further capture global semantic information, features Figure 2 After two connected MC-SSM modules and a Patch Merging module, a feature with 4C channels is obtained. Figure 3 .
[0025] Stage 4: Focus on abstract expression of high-level features, features Figure 3 After passing through two connected MC-SSM modules in sequence, the feature with a channel number of 8C is obtained Figure 4 .
[0026] The information processing process of the decoder layer includes: Stage 1: Features Figure 4 After passing through two connected MC-SSM modules in sequence, a feature map with 8C channels is obtained, which is combined with the feature map Figure 3 Add up the features Figure 5 ; Phase 2: Features Figure 5 After two connected MC-SSM modules and a Patch Merging module, a feature map with 4C channels is obtained. Figure 2The feature map 6 is obtained by adding them together; the third stage: the feature map 6 passes through two connected MC-SSM modules and a Patch Merging module in turn to obtain a feature map with a channel number of 2C. Figure 1 The feature map 7 is obtained by adding them together; the fourth stage: the feature map 7 is connected to a MC-SSM module and a Patch Merging module to obtain a feature map with a channel number C, and the feature map is added to the input of the encoder layer to obtain the feature map 8, which is used as the output of the decoder layer.
[0027] Furthermore, the MC-SSM module includes a MixConv module and an SSM module. The structure of the MC-SSM module is as follows: Figure 2 As shown in Figure 1, the feature maps obtained by the MC-SSM module are input into the MixConv module and the SSM module respectively.
[0028] The MixConv module has a module structure diagram as shown below: Figure 3 As shown in the figure, the feature maps obtained by the MixConv module are input into convolution kernels of sizes 3*3, 5*5, and 7*7 respectively, and the input feature maps are grouped; different convolution kernels are used for convolution and normalization for each group; the output of each group is spliced; in order to enhance the feature expression ability, the input of the MixConv module is added to the spliced vector to obtain the MixConv output.
[0029] Specifically, the input channel tensor is divided into g groups, and all the divided tensors X have the same spatial height h and spatial width w, and the final output tensor is The concatenation of all the partitioned tensors:
[0030]
[0031] c=c1+c2+..+c g
[0032] z0=z1+..+z g =m·c
[0033] Where c is the channel size, X (h,w,c) Represents an input tensor of shape (h, w, c), z0 represents the final output channel width, and m represents the channel multiplier, which is used to control the output channel size.
[0034] The SSM module has a module structure diagram as shown in Figure 4As shown in the figure, the input features are divided into two paths when passing through the horizontal normalization layer (LayerNorm, Layer Normalization); in branch 1, the input passes through a linear layer (LinearLayer) and then an activation function to obtain the output of branch 1; in branch 2, the input is processed by the linear layer, the activation function and the depth-wise separable convolution (DWConv) and enters the SS2D (2D-Selective-Scan, two-dimensional selective scanning) module for feature extraction. Furthermore, the SS2D module also includes: scan extension, S6 block (S6 Block, selective scanning mechanism), and scan merging.
[0035] Scanning expansion includes: expanding the input image into a sequence along four different directions: upper left to lower right, lower right to upper left, upper right to lower left, and lower left to upper right.
[0036] The S6 block includes: feature extraction of the expanded sequence to ensure that information in all directions is thoroughly scanned, thereby capturing diverse features; further, the S6 block is developed from the Mamba model, and a selection mechanism for adjusting SSM parameters based on input is added on the basis of the original S4, so that the model can filter out irrelevant information while distinguishing and retaining relevant information; the operations performed by the S6 block are expressed as:
[0037] Δ,B,C=Linear(x),Linear(x),Linear(x)
[0038]
[0039] y t =Ch t +Dx t
[0040] y=[y1,y2,···,y t ,···,y L ]
[0041] Among them, x represents the input of the feature of S6 block shape [B,L,D], B, L, D represent the batch size, batch length, and batch dimension respectively. Three linear transformations are performed on x to obtain Δ, B, and C respectively, and the hidden state h is updated t , Indicates the zero-order hold (ZOH) of A. is a discretization parameter, Represents the zero-order hold of B, using the hidden state h of the previous moment t-1 and the current input x t , calculate the output y at the current moment t , combined with the hidden state h at the current moment tand input x t , the output y at all moments t Concatenate into the final output y, where y represents a feature with a shape of [B, L, D].
[0042] Scan merging involves summing the sequences from four directions and restoring the image to the same size as the input.
[0043] The features are normalized using a horizontal normalization layer and multiplied element-by-element with the output of branch 1. The result of the multiplication passes through a linear layer, and the feature map after the linear layer is added to the input of the SSM module to obtain the SSM output; preferably, SiLU is used as the activation function of the SSM module.
[0044] The output feature map of the MixConv module is merged with the output feature map of the SSM module to obtain a merged feature map.
[0045] Furthermore, the output shape of Mixconv is (B,C MixConv ,H,W) feature map Y MixConv , the SSM output shape is (B,C SSM ,H,W) is the feature map of Y SSM , where B represents the channel size, C represents the number of channels, H represents the height of the feature map, W represents the width of the feature map, and Y is connected in series MixConv With Y SSM , and the shape is (B,C MixConv +C SSM ,H,W) feature map Y concat .
[0046] A 1×1 convolutional layer is used to extract deep features from the merged feature map to obtain a deep feature map.
[0047] Furthermore, the original input features are convolved with a 1×1 convolution. The function of this convolution layer is to perform weighted fusion on each channel through the convolution kernel, so that the feature maps from different sources are better integrated. The output shape is (B, C out ,H,W) feature map Y fused ; Perform batch normalization operation.
[0048] The feature map obtained by the MC-SSM module is convolved with 1×1 and batch normalized to obtain a shallow feature map.
[0049] The deep feature map is added to the shallow feature map to obtain the fused feature map, which is used as the output of the MC-SSM module.
[0050] S3: Send the encoded image to the decoder layer, and process it through the MC-SSM module and the Patch Embedding module to obtain the decoded image.
[0051] S4. Send the decoded image to the Linear Projection layer to restore the number of channels.
[0052] Specifically, the projection layer compresses the number of channels of the feature map to the number of channels of the target segmentation map, so that the feature map output by the decoder matches the target segmentation map in both spatial size and number of channels.
[0053] Figure 5 This is a test chart of this embodiment under different indicators. The figure records different indicator charts when different models are tested on two data sets, indicating the test results of different models.
[0054] S5. Calculate the loss function of the model according to the matched segmentation target, adjust the model parameters, and complete the model training when the loss function converges.
[0055] Specifically, the loss function of the model uses two loss functions, one for binary segmentation and the other for multi-class segmentation tasks.
[0056] Furthermore, for binary segmentation, the BceDic loss function is selected, which is composed of binary cross entropy (Bce) and Sorenson-Dice loss function (Dice Loss). The BceDice loss function can measure the difference between the probability distribution predicted by the model and the true label;
[0057] For multi-class segmentation, the CeDice loss function is used, which consists of the cross entropy (Ce) and the Sorenson-Dice loss function;
[0058] The formula of BceDic loss function L BceDice And the formula of CeDice loss function L CeDice It is expressed as:
[0059]
[0060] L BceDice =λ1L Bce +λ2L Dice
[0061] L CeDice =λ1L Ce +λ2L Dice
[0062] Among them, L Bce represents the binary cross entropy function, L Ce represents the cross entropy function, L Dice represents the Sorenson-Dyess loss function, N represents the total number of samples, C represents the total number of categories, and y i and Represent the true label and predicted label respectively, y i,c is an indicator that if sample i belongs to category c, then y i,c is equal to 1, otherwise it is 0. is the probability that the model predicts that sample i belongs to category c, |X| and |Y| represent the true value and predicted value respectively, and λ1 and λ2 represent L Bce The weight and L Dice The default values of λ1 and λ2 are both 1.
[0063] Furthermore, the entire network uses AdamW as the optimizer, combined with the Cosine AnnealingLR scheduler to dynamically adjust the learning rate; preferably, the initial learning rate is set to 0.001 and the minimum learning rate is set to 1e-5.
[0064] Finally, it should be noted that the above only describes one embodiment of the present invention. For those skilled in the art, it is conceivable that various changes, modifications, substitutions and deformations may be made to these embodiments without departing from the principle and spirit of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents, and the above-mentioned actions should be covered within the scope of protection of the present invention.
Claims
1. A lung CT image segmentation method based on hybrid deep convolution and state-space model, characterized in that: include: Acquire a lung CT image to be processed; input the lung CT image to be processed into a trained lung CT image segmentation model to obtain a segmentation result; The lung CT image segmentation model adopts the MCMamba network; the MCMamba network is composed of a PatchEmbedding layer, an encoder layer, a decoder layer, and a Linear Projection layer connected in sequence; the encoder layer is composed of a MC-SSM module and a Patch Merging module; the MC-SSM module includes a MixConv module, an SSM module, a fusion layer, two 1×1 convolutional layers and a normalization layer.
2. The lung CT image segmentation method based on hybrid deep convolution and state space model according to claim 1, characterized in that: Training the lung CT image segmentation model includes: S1. Input the lung CT image in the training data into the Patch Embedding layer to obtain an embedded image; normalize the embedded image; S2, the normalized embedded image is sent to the encoder layer, processed by the MC-SSM module and the Patch Embedding module to obtain the encoded image; S3, sending the encoded image to the decoder layer, processed by the MC-SSM module and the Patch Embedding module to obtain the decoded image; S4, send the decoded image to the Linear Projection layer to restore the number of channels; S5. Calculate the loss function of the model according to the matched segmentation target, adjust the model parameters, and complete the model training when the loss function converges.
3. The lung CT image segmentation method based on hybrid deep convolution and state space model according to claim 2, characterized in that: The Patch Embedding layer processes the lung CT image by dividing the lung CT image into non-overlapping blocks, mapping the non-overlapping blocks to a fixed dimension, and obtaining an embedded image.
4. The lung CT image segmentation method based on hybrid deep convolution and state space model according to claim 2, characterized in that: The encoder layer and the decoder layer are jump-connected, and the data processing process includes four stages, among which the jump connection adopts the addition operation.
5. The lung CT image segmentation method based on hybrid deep convolution and state space model according to claim 4, characterized in that: The encoder layer processes the data including: The first stage: the input is passed through two connected MC-SSM modules and a Patch Merging module in sequence to perform shallow feature extraction and obtain a feature map 1 with a channel number of C; The second stage: Feature map 1 passes through two connected MC-SSM modules and a Patch Merging module in turn to obtain feature map 2 with 2C channels; The third stage: Feature map 2 passes through two connected MC-SSM modules and a Patch Merging module in turn to obtain feature map 3 with 4C channels; Stage 4: Feature image 3 passes through two connected MC-SSM modules in sequence to obtain feature image 4 with 8C channels; Where C is the dimension of the mapping.
6. The lung CT image segmentation method based on hybrid deep convolution and state space model according to claim 5, characterized in that: The decoder layer processes data in four stages: Phase 1: Feature map 4 passes through two connected MC-SSM modules in sequence to obtain a feature map with a channel number of 8C, which is added to feature map 3 to obtain feature map 5; The second stage: Feature map 5 passes through two connected MC-SSM modules and a Patch Expanding module in turn to obtain a feature map with a channel number of 4C. This feature map is added to feature map 2 to obtain feature map 6. The third stage: Feature map 6 passes through two connected MC-SSM modules and a Patch Expanding module in turn to obtain a feature map with a channel number of 2C. This feature map is added to feature map 1 to obtain feature map 7; Stage 4: Feature map 7 is connected to an MC-SSM module and a Patch Expanding module to obtain a feature map with C channels. This feature map is added to the input of the encoder layer to obtain feature map 8, which is used as the output of the decoder layer. Where C is the dimension of the mapping.
7. The lung CT image segmentation method based on hybrid deep convolution and state space model according to claim 2, characterized in that: The MC-SSM module processes data in the following steps: Step 1. Input the input feature map into the MixConv module and the SSM module respectively; Step 2. Merge the output feature map of the MixConv module with the output feature map of the SSM module to obtain a merged feature map; Step 3. Use a 1×1 convolutional layer to extract deep features from the merged feature map to obtain a deep feature map; Step 4. Perform 1×1 convolution on the input feature map and perform batch normalization to obtain a shallow feature map; Step 5. Add the deep feature map and the shallow feature map to obtain the fused feature map.
8. The lung CT image segmentation method based on hybrid deep convolution and state space model according to claim 7, characterized in that: The MixConv module includes: 3 convolution kernels of size 3*3, 5*5, 7*7, 3 normalization modules, and a residual splicing module; the MixConv module processes data in the following steps: Step 1. Input the input feature maps to the convolution kernels of sizes 3*3, 5*5, and 7*7 respectively, and group the input feature maps; Step 2. Use different convolution kernels to convolve each group and normalize it; Step 3. Concatenate the output of each group; Step 4. Add the input feature map to the concatenated vector.
9. The lung CT image segmentation method based on hybrid deep convolution and state space model according to claim 7, characterized in that: The SSM module includes: 3 linear layers, 2 lateral normalization layers, a depth-wise separable convolutional layer, and a SS2D layer. The linear layer uses SiLU as the activation function.
10. The lung CT image segmentation method based on hybrid deep convolution and state space model according to claim 2, characterized in that: The operations performed by the Linear Projection layer include: compressing the number of channels of the output feature map of the decoder layer so that the number of channels of the output feature map is the same as the number of channels of the target segmentation map.
Citation Information
Patent Citations
Image super-resolution reconstruction model construction method and device, equipment and storage medium
CN114926342A
Training method and device of image super-resolution network
CN116385265A
Intestinal polyp segmentation network method based on residual double convolution and hybrid convolution of U-Net
CN117036381A
Knee joint MRI image semi-supervised segmentation method based on CPS
CN117671257A
Method and system of liver segmentation
US20110052028A1