Lung cancer CT (Computed Tomography) image data segmentation method, device, equipment, medium and product
By building a multi-scale global optimization network, combining multi-scale dynamic fusion module and global attention fusion module, the problem of multi-scale feature capture difficulty in lung cancer CT image segmentation is solved, segmentation accuracy and model generalization ability are improved, and computational complexity is reduced.
Patent Information
- Application Number
- CN202510329091.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
The existing CT image segmentation method of lung cancer is difficult to effectively capture multi-scale tumor characteristics, resulting in poor segmentation effect, large calculation volume and poor model generalization ability.
A multi-scale global optimization network is built, combining multi-scale dynamic fusion modules and global attention fusion modules, using parallel multi-branch hollow convolution and dynamic convolution kernels, combining channel and spatial attention mechanisms, enhancing the multi-scale feature capture ability of tumor areas, and smoothing the feature map through the bottleneck layer to capture global dependencies.
It improves tumor segmentation accuracy and generalization ability of the model, reduces calculation complexity, enhances the ability to capture multi-scale feature in the tumor area, and improves the segmentation effect.
Smart Images

Figure CN120259656A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image segmentation, and particularly to a method, device, equipment, medium and product for segmenting lung cancer CT image data. Background Art
[0002] Lung cancer is one of the malignant tumors with the highest incidence and mortality rates globally. Early diagnosis and treatment are crucial for improving the survival rate of patients. Computed Tomography (CT) is the main imaging tool for lung cancer screening and diagnosis, with the characteristics of high sensitivity and high visualization, and can provide high-resolution lung images.
[0003] However, characteristics such as the multi-scale distribution of the tumor area in lung cancer CT images, the blurred boundary with normal tissues, and the complex lesion morphology make it difficult for traditional image segmentation methods to achieve high-precision tumor area segmentation.
[0004] In recent years, deep learning technology has made remarkable progress in the field of medical image segmentation. Deep learning methods represented by Convolutional Neural Network (CNN), due to their powerful feature extraction capabilities, have gradually become the mainstream technology in medical image segmentation. The Fully Convolutional Network (FCN) proposed by Long et al. realizes the end-to-end training of the segmentation network without changing the image size by using downsampling operations during feature extraction and interpolating upsampling when generating the segmentation result. However, this method will cause a certain degree of loss of image details, and the accuracy needs to be further improved.
[0005] Considering that medical images have rich spatial information (such as complex texture structures), and the network downsampling process is prone to losing spatial information, the network structure based on Encoder-Decoder has begun to emerge. Ronneberger et al. proposed the U-Net structure, which establishes feature fusion channels of different scales between the symmetric encoder and decoder through skip connections, enabling the network to better utilize the global and local features of the image. With the in-depth study of neural networks, the attention mechanism has gradually been widely applied. The core of the attention mechanism is to achieve re-weighting of features by calculating the attention map. Following this idea, different attention modules can be designed according to specific medical image segmentation tasks to achieve the purpose of strengthening effective features and suppressing ineffective features. According to the different application positions, it can be divided into spatial attention and channel attention. Jun et al. paralleled these two attention modules and proposed a more effective dual-channel attention module for scene segmentation of natural images.
[0006] CNN architectures are good at detecting local features but often struggle to effectively capture global features. In contrast, Transformers can capture long-range features but may lose local feature details and result in poor segmentation accuracy for small organs.
[0007] To overcome these limitations, researchers have explored hybrid methods that combine the CNN and Transformer frameworks. Chen et al. proposed TransUNet, a network architecture that combines Transformer and U-Net, using Transformer to build a more powerful encoder that can capture richer features for multi-organ segmentation of the head and neck. Oktay et al. combined the attention mechanism with the U-Net network and proposed Attention-UNet, which achieved significant results in lesion segmentation of pancreatic CT images.
[0008] It can be seen that the current improvement in the performance of lung cancer CT image segmentation mainly benefits from the advantages of the network model in image representation learning ability. However, the current segmentation algorithms still do not meet the requirements of medical applications. The main problems are as follows:
[0009] In lung cancer CT images, the boundary between the tumor region and normal tissue is blurred, and the tumor morphology is complex and of different sizes. This leads to the problems of small inter-class differences and large intra-class differences. Existing methods such as U-Net and Attention-UNet often struggle to simultaneously consider the detailed information of small and large tumors when dealing with multi-scale tumor features, resulting in poor segmentation effects for tumors. Therefore, how to optimize the network structure, enhance the ability to capture multi-scale features of the tumor region, and improve the segmentation accuracy is a key issue that needs further research.
[0010] Traditional convolutional networks increase the receptive field by stacking multiple convolutional layers, but this method leads to a significant increase in computational complexity and may cause information dilution. Moreover, as the network depth increases, the number of network parameters increases, requiring high computational resources and more data for training. Otherwise, it is easy to cause overfitting, resulting in poor model generalization ability. Summary of the Invention
[0011] The purpose of this application is to provide a method, device, equipment, medium, and product for segmenting lung cancer CT image data to solve the problems of poor tumor segmentation effect and poor model generalization ability caused by large computational complexity.
[0012] To achieve the above object, this application provides the following solutions:
[0013] In the first aspect, this application provides a method for segmenting lung cancer CT image data, including:
[0014] Construct and train a multi-scale global optimization network to determine the optimal multi-scale global optimization network; the multi-scale global optimization network includes a multi-scale dynamic fusion module, a bottleneck layer, and a global attention fusion module; the multi-scale dynamic fusion module uses parallel multi-branch dilated convolutions and dynamic convolutional kernels, combines channel attention mechanism and spatial attention mechanism, enhances the ability to capture multi-scale features of the tumor region, and outputs an enhanced feature map; the bottleneck layer is used to smooth the enhanced feature map; the global attention fusion module combines channel attention mechanism, channel shuffling mechanism, and spatial attention mechanism to capture global dependencies in the smoothed feature map;
[0015] Input the lung cancer CT image data to be measured into the optimal multi-scale global optimization network for tumor region segmentation to determine the tumor region.
[0016] In a second aspect, the present application provides a lung cancer CT image data segmentation device, including:
[0017] A multi-scale global optimization network construction module for constructing and training a multi-scale global optimization network to determine the optimal multi-scale global optimization network; the multi-scale global optimization network includes a multi-scale dynamic fusion module, a bottleneck layer, and a global attention fusion module; the multi-scale dynamic fusion module uses parallel multi-branch dilated convolutions and dynamic convolutional kernels, combines channel attention mechanism and spatial attention mechanism, enhances the ability to capture multi-scale features of the tumor region, and outputs an enhanced feature map; the bottleneck layer is used to smooth the enhanced feature map; the global attention fusion module combines channel attention mechanism, channel shuffling mechanism, and spatial attention mechanism to capture global dependencies in the smoothed feature map;
[0018] A tumor region determination module for inputting the lung cancer CT image data to be measured into the optimal multi-scale global optimization network for tumor region segmentation to determine the tumor region.
[0019] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the lung cancer CT image data segmentation method described in any one of the above.
[0020] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the lung cancer CT image data segmentation method described in any one of the above.
[0021] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the lung cancer CT image data segmentation method described in any one of the above.
[0022] According to the specific embodiments provided in the present application, the following technical effects are disclosed in the present application:
[0023] The present application constructs a multi-scale global optimization network, which includes a multi-scale dynamic fusion module (MDFM) and a global attention fusion module (GAFM). By means of parallel multi-branch dilated convolutions, the receptive field is enlarged and spatial information in different ranges is captured, enhancing the model's ability to understand the overall layout. At the same time, combined with the attention mechanism, which includes a channel attention mechanism and a spatial attention mechanism, MDFM can capture the details and extensive context information in the image simultaneously, enhancing the multi-scale feature capture ability for tumor regions, improving the segmentation accuracy and tumor segmentation effect; GAFM combines the channel attention mechanism, channel shuffling mechanism and spatial attention mechanism to capture the global dependencies in the smoothed feature map, enhancing the diversity and expression ability of features, better fusing low-level and high-level features, while reducing the computational complexity and improving the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 It is a schematic diagram of the flow of the lung cancer CT image data segmentation method provided by the present application;
[0026] Figure 2 It is a schematic diagram of the training process of the multi-scale global optimization network provided by the present application;
[0027] Figure 3 It is a schematic diagram of the overall architecture of the multi-scale global optimization network provided by the present application;
[0028] Figure 4 It is a schematic diagram of the multi-scale dynamic fusion module provided by the present application;
[0029] Figure 5 It is a dilated convolution diagram provided by the present application;
[0030] Figure 6 It is a schematic diagram of the global attention fusion module provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0032] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] The embodiment of the present application provides a method for segmenting lung cancer CT image data. This method is executed by a computer device, specifically, it can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiment of the present application, as Figure 1 shown, this method includes the following steps.
[0034] S1: Construct and train a multi-scale global optimization network to determine the optimal multi-scale global optimization network; the multi-scale global optimization network includes a multi-scale dynamic fusion module, a bottleneck layer, and a global attention fusion module; the multi-scale dynamic fusion module uses parallel multi-branch dilated convolutions and dynamic convolution kernels, combines a channel attention mechanism and a spatial attention mechanism, enhances the multi-scale feature capture ability for the tumor region, and outputs an enhanced feature map; the bottleneck layer is used to smooth the enhanced feature map; the global attention fusion module combines a channel attention mechanism, a channel shuffle mechanism, and a spatial attention mechanism to capture the global dependence relationship in the smoothed feature map.
[0035] S2: Input the lung cancer CT image data to be measured into the optimal multi-scale global optimization network for tumor region segmentation to determine the tumor region.
[0036] In an exemplary embodiment, before S2, it further includes:
[0037] Preprocess the lung cancer CT image data to be measured, specifically including:
[0038] Perform format conversion on the lung cancer CT image data to be measured to determine the CT image data after format conversion.
[0039] Perform normalization processing on the CT image data after format conversion to determine the CT image data after normalization processing.
[0040] Perform image enhancement processing on the CT image data after normalization processing to determine the enhanced CT image data.
[0041] In an exemplary embodiment, asFigure 2 As shown, train the multi-scale global optimization network according to the historical lung cancer CT image data until the optimal multi-scale global optimization network, that is, the optimal model, is determined.
[0042] In practical applications, such as Figure 3 As shown, after inputting the historical lung cancer CT image, first preprocess the image. First, perform format conversion, convert the nii format of the CT image to the jpg format, then perform normalization processing, scale the pixel values to the range of [0, 1], and perform necessary image enhancement operations (such as random rotation, flipping, etc.) to improve the robustness of the model. And divide the dataset into a training set and a validation set according to 8:2.
[0043] The preprocessed image is input into the multi-scale global optimization network (Multi-scale Global Optimization Network, MGONet) for tumor region segmentation. The specific process is as follows: First, send it into the encoder, and extract multi-scale features through the multi-scale dynamic fusion module. The MDFM module uses parallel multi-branch dilated convolutions (different dilation rates) and dynamic convolution kernels to generate, combined with the channel-space attention mechanism, to enhance the ability to capture multi-scale features of the tumor region, especially the recognition of tumors with blurred boundaries and small tumors. Then, after the feature map undergoes a max-pooling operation with a pooling window of 2×2, the size of the output feature map is reduced to half of the original. After passing through four encoders, it is input into the Bottleneck layer, and through two 3×3 convolutions (with a Rectified Linear Unit (ReLU) activation function), the feature map is smoothed through convolution operations, reducing noise and redundant information, and then input into the decoder.
[0044] Enter the decoder. The decoder restores the resolution of the feature map through upsampling operations and performs skip connections with the feature maps of the corresponding levels of the encoder to fuse multi-scale information. After passing through the last 1×1 convolution and the Sigmoid activation function, a probability map is generated, indicating the probability that each pixel belongs to the tumor region. Through threshold processing (threshold = 0.5, can be set), the final segmentation mask is generated, where 1 represents the tumor region and 0 represents the background, completing the tumor segmentation of the CT image.
[0045] After each training cycle ends, use the validation set to evaluate the model performance, and record the Dice coefficient and IoU value. When the model reaches the optimal performance on the validation set, save the model parameters (including network weights, optimizer state) as a pre-trained model file. In subsequent training, load the pre-trained model parameters as the initial state to accelerate model convergence and improve performance.
[0046] Experimental evaluation metrics:
[0047] 1. Dice Coefficient
[0048] Definition: The Dice coefficient is a metric used to measure the similarity between two sets (such as the predicted segmentation result and the ground truth label), and is widely used in medical image segmentation tasks. Its value range is [0, 1], and the closer the value is to 1, the higher the overlap degree between the segmentation result and the ground truth label.
[0049] Formula:
[0050] X: The predicted segmentation result (the binary mask output by the model).
[0051] Y: The ground truth label.
[0052] |X ∩ Y|: The number of intersection pixels between the prediction result and the ground truth label.
[0053] |X| and |Y|: The number of pixels in the prediction result and the ground truth label.
[0054] Significance: The Dice coefficient focuses on the overlapping region between the prediction result and the ground truth label, and can effectively reflect the segmentation accuracy. It is insensitive to the size of the segmented region and is suitable for evaluating the segmentation effect of small tumors.
[0055] 2. Intersection over Union (IoU)
[0056] Definition: IoU is a metric for measuring the similarity between two sets, which calculates the ratio of the intersection to the union of the prediction result and the ground truth label. Its value range is also [0, 1], and the closer the value is to 1, the better the segmentation effect.
[0057] Formula:
[0058] X ∪ Y: The number of union pixels between the prediction result and the ground truth label.
[0059] Significance: IoU not only focuses on the overlapping region, but also considers the overall distribution of the prediction result and the ground truth label. It is more sensitive to the boundary error of the segmentation result and can better reflect the boundary fitting degree of the segmentation.
[0060] In an exemplary embodiment, S2 can be replaced by the following steps.
[0061] S21: Input the lung cancer CT image data to be measured into the multi-scale dynamic fusion module, and output an enhanced feature map.
[0062] S22: Input the enhanced feature map into the bottleneck layer, and output a smoothed feature map.
[0063] S23: Input the smoothed feature map into the global attention fusion module to output the final feature map.
[0064] S24: Identify the tumor region in the lung cancer CT image data to be measured based on the final feature map.
[0065] In an exemplary embodiment, the multi-scale dynamic fusion module aims to enhance feature expression by utilizing multiple dilation rates and integrating channel and spatial attention mechanisms. This design meets the requirement of simultaneously capturing detailed information and broad context information in the image, which is crucial for complex lung cancer tumor segmentation. The module is divided into two parts. The first part is the multi-scale dilated convolution part, and the second part is the fusion of channel and spatial attention. S21 can be replaced by the following steps.
[0066] As Figure 4 shown, the multi-scale dynamic fusion module includes: parallel multi-branch dilated convolutional layers, fully connected layers, channel attention mechanism layers, spatial attention mechanism layers, and fusion layers.
[0067] S211: Input the lung cancer CT image data to be measured into the parallel multi-branch dilated convolutional layers respectively to extract feature maps of different scales; different branches are configured with different dilation rates.
[0068] S212: Concatenate the feature maps of different scales in the channel dimension.
[0069] S213: Input the concatenated feature map into the fully connected layer to generate a dynamic convolution kernel, and perform dynamic convolution on the concatenated feature map using the dynamic convolution kernel to output a comprehensive feature map.
[0070] S214: Input the comprehensive feature map into the channel attention mechanism layer and the spatial attention mechanism layer respectively, and use the channel attention mechanism and the spatial attention mechanism to perform channel-level recalibration and spatial-level recalibration on the comprehensive feature map respectively to determine the channel feature map and the spatial feature map.
[0071] S215: Input the channel feature map and the spatial feature map into the fusion layer for fusion to determine the enhanced feature map.
[0072] In Figure 4 is Concat, feature concatenation, and the feature map is concatenated through channels; is the result F of concatenating the feature maps of multiple branches ∈R cat ∈R B×5C×H×W ; is the fully connected layer network, R B×5C×H×Wis the shape attribute of the feature map, B is the batch size, C is the number of channels, H is the height of the image, and W is the width of the image.
[0073] An example of the dilated convolution concept is as follows Figure 5 As shown, the relevant parameters of this convolution example diagram are kernel_size = 3, dilated_ratio = 2, stride = 1, and padding = 0.
[0074] Among them, kernel_size = 3: The size of the convolution kernel is 3x3.
[0075] dilated_ratio = 2: The dilation rate is 2, indicating that 1 interval is inserted between the elements of the convolution kernel, and the actual receptive field is enlarged.
[0076] stride = 1: The convolution stride is 1, indicating that the convolution kernel moves 1 pixel each time.
[0077] padding = 0: No padding is used, and the size of the output feature map will be reduced.
[0078] In practical applications, the MDFM process is as follows:
[0079] The first part implements multi-scale dilated convolution. The MDFM architecture realizes feature extraction at different scales through five parallel convolutional branches. Each branch is configured with a different dilation rate to expand the receptive field and capture spatial information in different ranges.
[0080] The first branch uses a 1x1 convolution kernel, which directly extracts features without changing the spatial scale. The second branch uses a 3x3 convolution kernel with a dilation rate of 6 to moderately expand the receptive field. The third branch uses a 3x3 convolution kernel with a dilation rate of 12 to further expand the receptive field to capture more extensive context information. The fourth branch uses a 3x3 convolution kernel with a dilation rate of 18 to provide the widest receptive field. The fifth branch uses an additional global average pooling branch to extract global context features and enhance the model's understanding ability of the overall layout.
[0081] Extract features of different scales from the five different branches, splice them in the channel dimension, then use a fully connected layer network to generate a dynamic convolution kernel, and then perform dynamic convolution on the spliced feature map with the dynamic convolution kernel to synthesize a comprehensive feature map.
[0082] The second part implements the merging and calibration of channel and spatial features. The merged feature map passes through the calibration of two attention mechanisms in parallel: the channel attention mechanism above and the spatial attention mechanism below.
[0083] Channel Attention Mechanism: First, perform global average pooling on the comprehensive feature map to obtain the global features of each channel. Then, learn the importance weights of each channel through two fully connected layers (using ReLU and Sigmoid activation functions respectively). Finally, multiply these weights with the original feature map channel by channel to achieve channel weighting.
[0084] Spatial Attention Mechanism: Perform global pooling on the comprehensive feature map in the channel dimension to obtain the spatial feature map. Learn the importance weights of each spatial position through a 1x1 convolution and Sigmoid activation function. Multiply these weights with the original feature map element by element to achieve spatial weighting.
[0085] Finally, the output feature maps of the channel attention and spatial attention mechanisms are fused through an element-wise addition (Add) operation to finally obtain the enhanced feature map. At this time, the feature maps calibrated by channels and spaces respectively are element-wise added to the original merged feature map to integrate and enhance relevant features. The enhanced feature map is finally reduced in dimension and integrated through a 1x1 convolutional layer to generate the final output feature map.
[0086] The input and output of MDFM are as follows:
[0087] 1. Input and output of the multi-scale dilated convolution branch.
[0088] Input: Feature map X ∈ R B×C×H×W
[0089] Output: The shape of the output feature map of each branch is R B×C×H×W , and the result F after concatenating the feature maps of multiple branches cat ∈ R B×5C×H×W . The number of channels of the concatenated feature map is 5C.
[0090] 2. Input and output of the dynamic convolution.
[0091] Step 1: Generation of the dynamic convolution kernel.
[0092] Input: The concatenated feature map F cat ∈ R B×5C×H×W .
[0093] Operation: Perform global average pooling on the input feature map of each branch to extract global information. Use a fully connected layer network to generate the weights of the dynamic convolution kernel.
[0094] Output: The weights W of the dynamic convolution kernel dynamic ∈ R B×Cout×Cin×K×K , where K is the height and width of the convolution kernel and can be set according to the actual situation.
[0095] Step 2: Dynamic convolution.
[0096] Input: The stitched feature map F cat and the weight W of the dynamic convolution kernel dynamic .
[0097] Operation: Apply the dynamic convolution kernel to the stitched feature map. Obtain the feature map F after dynamic convolution dynamic ∈R B ×Cout×H×W , C out is the number of channels of the feature map output after dynamic convolution.
[0098] Output: The feature map F after dynamic convolution dynamic .
[0099] 3. Input and output of channel and spatial feature merging and calibration.
[0100] Input: The dynamic convolution feature map F dynamic .
[0101] Operation: Use the channel attention mechanism to weight the comprehensive feature map. Use the spatial attention mechanism to weight the feature map after channel weighting.
[0102] Output: The calibrated feature map F calibrated ∈R B×Cout×H×W .
[0103] 4. Final convolution.
[0104] Input: The calibrated feature map F calibrated .
[0105] Operation: Use a 1x1 convolution to reduce the number of channels from C out ×N to C out .
[0106] Output: The final output feature map F out ∈R B×Cout×H×W .
[0107] This application proposes a multi-scale dynamic fusion module. At the encoding end, through multiple parallel dilated convolution branches, after channel layer stitching, a fully connected layer network is used to calculate the dynamic convolution weights for dynamic convolution, realizing the extraction and fusion of multi-scale features. Establishing global correlations of features from both spatial and channel dimensions enhances the multi-scale capture ability of tumor regions. Improves the segmentation fineness of small tumor margins.
[0108] Aiming at the problems of blurred tumor boundaries, complex shapes, and difficulty in capturing multi-scale features, this application solves the problem of difficulty in capturing multi-scale features by introducing MDFM. MDFM adopts a parallel multi-branch dilated convolution structure, and each branch is configured with different dilation rates (such as 1, 6, 12, 18) to expand the receptive field and capture tumor features at different scales. At the same time, MDFM combines channel attention and spatial attention mechanisms to dynamically adjust feature weights, enhance the attention to the tumor area, and suppress irrelevant background noise. In addition, global context information is extracted through the global average pooling branch to further optimize the ability to capture details of small and large tumors.
[0109] In an exemplary embodiment, GAFM can enhance the expressive ability of the input feature map. This module combines channel attention, channel shuffling, and spatial attention mechanisms to capture global dependencies in the feature map. S23 can be replaced by the following steps.
[0110] As Figure 6 shown, the global attention fusion module includes: a channel attention sub-module, a channel shuffling sub-module, and a spatial attention sub-module.
[0111] S231: Input the smoothed feature map into the channel attention sub-module to generate a channel attention map.
[0112] S232: Multiply the smoothed feature map and the channel attention map element-wise to determine the channel-enhanced feature map.
[0113] S233: Input the channel-enhanced feature map into the channel shuffling sub-module to apply the channel shuffling operation to determine the shuffled feature map.
[0114] S234: Input the shuffled feature map into the spatial attention sub-module to generate a spatial attention map.
[0115] S235: Multiply the shuffled feature map and the spatial attention map element-wise to output the final feature map.
[0116] In an exemplary embodiment, S231 can be replaced by the following steps.
[0117] S2311: In the channel attention sub-module, perform dimensional permutation on the smoothed feature map.
[0118] S2312: Use a multi-layer perceptron to capture the inter-channel dependencies of the permuted feature map to determine the feature map with restored dimensions.
[0119] S2313: Perform inverse permutation on the feature map with restored dimensions and use an activation function to generate a channel attention map.
[0120] In practical applications, the GAFM process is as follows:
[0121] The input smoothed feature map is first fed into the channel attention sub-module and then into the spatial attention sub-module.
[0122] The initial feature map contains multiple channels, and the spatial size of each channel is H×W. In the channel attention sub-module, the input feature map first undergoes a dimension permutation, changing from C×H×W to W×H×C. Then, the channel dependencies are captured through two layers of a Multi-Layer Perceptron (MLP). The first layer of the MLP reduces the number of channels to 1 / 5 of the original, and then introduces non-linearity through the ReLU activation function. Subsequently, the second layer of the MLP restores the number of channels to the original dimension. Finally, an inverse permutation is performed to restore to C×H×W, and a channel attention map is generated through the Sigmoid activation function. The input feature map and the channel attention map are multiplied element-wise to obtain an enhanced feature map.
[0123] To further mix and share information, a channel shuffle operation is applied. The enhanced feature map is divided into 5 groups, with each group containing C / 5 channels. A dimension permutation operation is performed on the grouped feature maps to shuffle the channel order within each group. Subsequently, the shuffled feature map is restored to the original shape C×H×W, i.e., an inverse permutation operation. This way can better mix the feature information and enhance the feature expression ability.
[0124] Finally, the mixed feature map (i.e., the shuffled feature map) is fed into the spatial attention sub-module. The input feature map passes through a 7x7 convolutional layer, and the number of channels is reduced to 1 / 5 of the original. Then, batch normalization and the ReLU activation function are used for non-linear transformation. Next, the number of channels is restored to the original dimension C through a second 7x7 convolutional layer, followed by a batch normalization layer. Finally, a spatial attention map is generated through the Sigmoid activation function. The shuffled feature map and the spatial attention map are multiplied element-wise to obtain the final output feature map. The final output feature map contains enhanced features after channel attention, channel shuffle, and spatial attention.
[0125] The GAFM input and output are as follows:
[0126] 1. Input to the decoder GAFM module.
[0127] Input feature map: Feature map from the decoder upsampling layer and skip connection of the corresponding layer encoder, X∈R B ×C×H×W .
[0128] 2. Input and output of the channel attention sub-module.
[0129] Input: Input feature map \(X\in\mathbb{R}\) B×C×H×W .
[0130] Operation:
[0131] 1) Dimension permutation: Permute the dimensions of the input feature map from \(C\times H\times W\) to \(W\times H\times C\)
[0132] 2) MLP: Use a two-layer MLP to capture the dependencies between channels.
[0133] The first layer of MLP: Reduce the number of channels to 1 / 5 of the original.
[0134] The second layer of MLP: Restore the number of channels to the original dimension \(C\).
[0135] 3) Inverse permutation: Restore the dimensions of the feature map from \(W\times H\times C\) to \(C\times H\times W\).
[0136] 4) Sigmoid activation function: Generate the channel attention map \(A\in\mathbb{R}\) channel \(\in\mathbb{R}\) B×C×1×1
[0137] 5) Channel weighting: Multiply the channel attention map \(A\) with the input feature map \(X\) channel by channel to obtain the channel-weighted feature map \(F\in\mathbb{R}\). channel and the input feature map \(X\) channel by channel to obtain the channel-weighted feature map \(F\) channel \(\in\mathbb{R}\) B×C×H×W .
[0138] Output: Channel-weighted feature map \(F\) channel .
[0139] 3. Channel Shuffle.
[0140] Input: Channel-weighted feature map \(F\in\mathbb{R}\) channel \(\in\mathbb{R}\) B×C×H×W .
[0141] Operation:
[0142] 1) Grouping: Divide the feature map into 5 groups, each group containing \(C / 5\) channels.
[0143] 2) Dimension permutation operation: Perform a dimension permutation operation on the channels within each group to shuffle the channel order.
[0144] 3) Restore shape: Restore the shuffled feature map to the original shape \(C\times H\times W\).
[0145] Output: Feature map \(F\) after channel shuffle shuffle \(\in\mathbb{R}\) B×C×H×W .
[0146] 4. Spatial attention sub-module.
[0147] Input: Feature map \(F\) after channel shuffleshuffle ∈R B×C×H×W .
[0148] operate:
[0149] 1) 7x7 convolution: Use 7x7 convolution to reduce the number of channels to 1 / 5 of the original number.
[0150] 2) Batch Normalization and ReLU: Perform batch normalization and ReLU activation function on the convolution output.
[0151] 3) 7x7 convolution: Use 7x7 convolution to restore the number of channels to the original dimension C.
[0152] 4) Batch Normalization: Batch normalization is performed on the convolution output.
[0153] 5) Sigmoid activation function: Generate spatial attention map A spatial ∈R B×1×H×W .
[0154] 6) Spatial weighting: The spatial attention map A spatial The feature map F after channel shuffle shuffle Multiply element by element to get the final output feature map Fout.
[0155] Output: Final output feature map F out ∈R B×C×H×W .
[0156] This application proposes a global attention fusion module, which strengthens the feature expression of the tumor area through global channel attention and spatial attention mechanisms at the decoding end. The channel shuffle operation is introduced to enhance the diversity and expression of features and better integrate low-level and high-level features. By dynamically weighting effective features, redundant calculations are reduced, thereby reducing the number of model parameters and memory usage. In addition, the channel shuffle operation in GAFM further enhances the diversity and expression of features and avoids the problem of feature solidification.
[0157] In response to the problems of large computational complexity, information dilution and poor generalization ability of traditional convolutional networks, this application effectively reduces the computational complexity and improves the generalization ability of the model through lightweight design and attention mechanism. First, the design of 2D convolution combined with multi-scale dilated convolution significantly reduces the amount of computation while retaining multi-scale features. Secondly, GAFM is introduced to reduce redundant calculations by dynamically weighting effective features, thereby reducing the number of model parameters and memory usage. In addition, the channel shuffling operation in GAFM further enhances the diversity and expressiveness of features, avoiding the problem of feature solidification.
[0158] Based on the same inventive concept, an embodiment of the present application further provides a lung cancer CT image data segmentation device for implementing the above-mentioned lung cancer CT image data segmentation method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following lung cancer CT image data segmentation device can refer to the limitations on the lung cancer CT image data segmentation method in the above text, and will not be repeated here.
[0159] In an exemplary embodiment, a lung cancer CT image data segmentation device is provided, including:
[0160] A multi-scale global optimization network construction module, configured to construct and train a multi-scale global optimization network to determine the optimal multi-scale global optimization network; the multi-scale global optimization network includes a multi-scale dynamic fusion module, a bottleneck layer, and a global attention fusion module; the multi-scale dynamic fusion module uses parallel multi-branch dilated convolutions and dynamic convolution kernels, combines channel attention mechanism and spatial attention mechanism, enhances the multi-scale feature capture ability for tumor regions, and outputs an enhanced feature map; the bottleneck layer is used to smooth the enhanced feature map; the global attention fusion module combines channel attention mechanism, channel shuffling mechanism, and spatial attention mechanism to capture the global dependence relationship in the smoothed feature map.
[0161] A tumor region determination module, configured to input the to-be-detected lung cancer CT image data into the optimal multi-scale global optimization network for tumor region segmentation to determine the tumor region.
[0162] In an exemplary embodiment, a computer device is provided. This computer device can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of this computer device is used to store lung cancer CT image data segmentation data. The input / output interface of this computer device is used to exchange information between the processor and external devices. The communication interface of this computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a lung cancer CT image data segmentation method.
[0163] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the above method is implemented.
[0164] In an exemplary embodiment, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the above method is implemented.
[0165] In an exemplary embodiment, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the above method is implemented.
[0166] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0167] In this application, all actions of obtaining signals, information, or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining authorization from the owner of the corresponding device.
[0168] In each of the embodiments provided in the present application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., without limitation. In each of the embodiments provided in the present application, the processor involved may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without limitation.
[0169] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0170] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for segmenting CT image data of lung cancer, characterized in that, The lung cancer CT image data segmentation method includes: Construct and train a multi-scale global optimization network to determine the optimal multi-scale global optimization network; the multi-scale global optimization network includes a multi-scale dynamic fusion module, a bottleneck layer, and a global attention fusion module; the multi-scale dynamic fusion module uses parallel multi-branch dilated convolutions and dynamic convolutional kernels, combines channel attention mechanism and spatial attention mechanism, enhances the multi-scale feature capture ability of the tumor region, and outputs an enhanced feature map; the bottleneck layer is used to smooth the enhanced feature map; the global attention fusion module combines channel attention mechanism, channel shuffling mechanism, and spatial attention mechanism to capture the global dependence relationship in the smoothed feature map; Input the lung cancer CT image data to be measured into the optimal multi-scale global optimization network for tumor region segmentation to determine the tumor region.
2. The lung cancer CT image data segmentation method according to claim 1, characterized in that Input the lung cancer CT image data to be measured into the optimal multi-scale global optimization network for tumor region segmentation to determine the tumor region, specifically including: Input the lung cancer CT image data to be measured into the multi-scale dynamic fusion module to output an enhanced feature map; Input the enhanced feature map into the bottleneck layer to output a smoothed feature map; Input the smoothed feature map into the global attention fusion module to output the final feature map; Identify the tumor region in the lung cancer CT image data to be measured according to the final feature map.
3. The lung cancer CT image data segmentation method according to claim 2, wherein Input the lung cancer CT image data to be measured into the multi-scale dynamic fusion module to output an enhanced feature map, specifically including: The multi-scale dynamic fusion module includes: parallel multi-branch dilated convolutional layers, fully connected layers, channel attention mechanism layers, spatial attention mechanism layers, and fusion layers; Input the lung cancer CT image data to be measured into the parallel multi-branch dilated convolutional layers respectively to extract feature maps of different scales; different branches are configured with different dilation rates; Stitch the feature maps of different scales in the channel dimension; Input the stitched feature map into the fully connected layer to generate a dynamic convolutional kernel, and perform dynamic convolution on the stitched feature map using the dynamic convolutional kernel to output a comprehensive feature map; Input the comprehensive feature map into the channel attention mechanism layer and the spatial attention mechanism layer respectively, and use the channel attention mechanism and the spatial attention mechanism to perform channel-level recalibration and spatial-level recalibration on the comprehensive feature map respectively to determine the channel feature map and the spatial feature map; Input the channel feature map and the spatial feature map into the fusion layer for fusion to determine the enhanced feature map.
4. The lung cancer CT image data segmentation method according to claim 2, wherein Input the smoothed feature map into the global attention fusion module to output the final feature map, specifically including: The global attention fusion module includes: a channel attention sub-module, a channel shuffling sub-module, and a spatial attention sub-module; Input the smoothed feature map into the channel attention sub-module to generate a channel attention map; Multiply the smoothed feature map and the channel attention map element by element to determine the channel-enhanced feature map; Input the feature map after channel enhancement into the channel shuffling sub-module to apply channel shuffling operation to determine the shuffled feature map; Input the shuffled feature map into the spatial attention sub-module to generate a spatial attention map; Multiply the shuffled feature map and the spatial attention map element by element and output the final feature map.
5. The lung cancer CT image data segmentation method according to claim 4, wherein, Input the smoothed feature map into the channel attention sub-module to generate a channel attention map, specifically including: In the channel attention sub-module, perform dimensional permutation on the smoothed feature map; Use a multi-layer perceptron to capture the inter-channel dependencies of the permuted feature map to determine the feature map with restored dimensions; Perform inverse permutation on the feature map with restored dimensions and use an activation function to generate a channel attention map.
6. The lung cancer CT image data segmentation method according to claim 1, characterized in that Before inputting the lung cancer CT image data to be measured into the optimal multi-scale global optimization network for tumor region segmentation to determine the tumor region, it also includes: Preprocess the lung cancer CT image data to be measured, specifically including: Perform format conversion on the lung cancer CT image data to be measured to determine the CT image data after format conversion; Perform normalization processing on the CT image data after format conversion to determine the CT image data after normalization processing; Perform image enhancement processing on the CT image data after normalization processing to determine the enhanced CT image data.
7. A lung cancer CT image data segmentation device, characterized in that, The lung cancer CT image data segmentation device includes: A multi-scale global optimization network construction module for constructing and training a multi-scale global optimization network to determine the optimal multi-scale global optimization network; the multi-scale global optimization network includes a multi-scale dynamic fusion module, a bottleneck layer, and a global attention fusion module; the multi-scale dynamic fusion module uses parallel multi-branch dilated convolutions and dynamic convolutional kernels, combines the channel attention mechanism and the spatial attention mechanism to enhance the multi-scale feature capture ability for the tumor region and outputs the enhanced feature map; the bottleneck layer is used to smooth the enhanced feature map; the global attention fusion module combines the channel attention mechanism, the channel shuffling mechanism, and the spatial attention mechanism to capture the global dependencies in the smoothed feature map; A tumor region determination module for inputting the lung cancer CT image data to be measured into the optimal multi-scale global optimization network for tumor region segmentation to determine the tumor region.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the lung cancer CT image data segmentation method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the lung cancer CT image data segmentation method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the lung cancer CT image data segmentation method according to any one of claims 1-6.
Citation Information
Cited By
Quantum heuristic progressive focusing plant cell microtubule image segmentation method and system
CN121904080A
Quantum-inspired progressive focusing plant cell microtubule image segmentation method and system
CN121904080B