A multi-type pulmonary nodule accurate segmentation method and model based on hybrid Transformer
By fusing the multi-scale convolution module and the channel attention convolution module in the lung nodule segmentation model, and improving the Transformer structure and feature bidirectional adaptive fusion module, the shortcomings of the existing models in feature extraction and generalization are solved, and the accuracy of lung nodule segmentation and model generalization ability are achieved.
Patent Information
- Application Number
- CN202211662984.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-12-23
AI Technical Summary
The existing lung nodule segmentation model has shortcomings in feature extraction ability and generalization, and it is difficult to fully learn the complex morphology and diversity of lung nodules.
The multi-type lung nodule precise segmentation method based on hybrid Transformer is adopted. By fusing the multi-scale convolution module and the channel attention convolution module in the model encoder, and improving the Transformer structure, pooling operations are combined with the self-attention layer, and feature bidirectional adaptive fusion module is added to improve feature interaction.
The feature extraction and generalization ability of the lung nodule segmentation model was improved, and the accuracy and robustness of the model in segmenting multiple types of lung nodules was enhanced.
Smart Images

Figure CN115731387B_ABST
Abstract
Description
Technical Field
[0001] The present invention provides a hybrid Transformer-based multi-type pulmonary nodule accurate segmentation method and model, belonging to the technical field of pulmonary nodule segmentation in lung CT images. Background Art
[0002] Lung cancer is the cancer with the highest morbidity and mortality rate in the world. Early detection, diagnosis and treatment of lung cancer can effectively reduce the mortality rate of patients. However, the high workload of thoracic surgeons in manually reviewing CT images can easily lead to missed diagnosis and misdiagnosis of lung nodules, causing patients to miss the best time to treat lung cancer. In order to relieve the pressure on doctors and increase the detection rate of early lung cancer, the industry is applying artificial intelligence technology to lung CT images, using AI technology to assist doctors in screening and diagnosing lung nodules and early lung lesions, and warning of early lung cancer.
[0003] Accurate segmentation of lung nodules on lung CT is the key to AI-assisted screening, diagnosis and qualitative analysis of lung nodules. The challenge of accurate segmentation of lung nodules lies in their specificity, complex and diverse shapes and sizes, and visual features similar to surrounding tissues, making it difficult for current segmentation models to fully learn all the features of lung nodules. Ronneberger et al. proposed U-Net, which consists of a symmetrical encoder-decoder network with horizontal jump connections from the encoder to the decoder, which can pass information from the shallow layer of the model to the deep layer of the model. DB-ResNet proposed by Cao et al. is a dual-branch residual network developed based on ResNet. It extracts multi-view features and multi-scale features of lung nodules through dual branches. Chen et al. proposed TransUNet, which directly inserts multiple Transformers between the encoder and decoder, making full use of the advantages of Transformer's ability to capture global information and CNN's ability to extract local information. Zhang et al. proposed TransFuse, which is a dual-branch parallel CNN and Transformer, to effectively capture global image dependencies and shallow spatial details in a simple way. Gao et al. proposed UTNet, which adopts the encoder-decoder structure of the U-type network and adds a self-attention layer after each encoder-decoder module to extract spatial and semantic features simultaneously.
[0004] The above prior art also has the following problems:
[0005] 1. The morphological features of lung nodules are complex, and feature extraction networks are required to fully extract local and global features to perform segmentation tasks. The CNN commonly used in existing technologies is good at extracting local features, but its ability to extract global features is insufficient, which can easily lead to incorrect segmentation of the edges of lung nodules. Transformer is good at representing global features and can make up for the shortcomings of CNN, but the simple combination of the two cannot improve the feature extraction ability of the model.
[0006] 2. Pulmonary nodules have various morphological types, which requires the pulmonary nodule segmentation model to be generalized in order to accurately segment multiple types of pulmonary nodules at the same time. Existing technologies often improve training strategies and use data enhancement to improve the generalization ability of the model, which can only alleviate the dilemma of insufficient generalization of the model. Model innovation can fundamentally solve this problem, but most of them only improve the depth and width of the model, and the resulting sharp increase in the number of parameters is not conducive to the actual deployment of the model. Summary of the invention
[0007] In view of the uneven internal density and complex and diverse external morphology of lung nodules, and in order to solve the problems of insufficient feature extraction capability and poor model generalization of existing lung nodule segmentation models, this paper proposes a multi-type lung nodule accurate segmentation method and model based on a hybrid Transformer.
[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is: a multi-type pulmonary nodule accurate segmentation method based on hybrid Transformer, comprising the following steps:
[0009] Step S1: Constructing a data set: Obtain lung CT images, complete the pixel-by-pixel annotation of the lung nodule lesion area through Dicom Viewer software, and perform data enhancement and data augmentation on the image;
[0010] Construct a lung nodule segmentation model, including:
[0011] Step S2: Construct the convolution module in the model encoder: According to the relationship between image resolution and the number of feature channels, design the multi-scale convolution module (Module_Multi) and the channel attention convolution module (Module_SE);
[0012] Step S3: Construct the Transformer module (Attention_Pool) in the model encoder: Improve the traditional Transformer structure and combine the pooling operation with the self-attention layer in the Transformer;
[0013] Step S4: Fuse the convolution module constructed in S2 and the Transformer module constructed in S3 in the model encoder;
[0014] Step S5: construct an upsampling module in the model decoder: group the feature maps after the dilated convolution and rearrange the channels;
[0015] Step S6: construct a feature bidirectional adaptive fusion module: by adding a learnable convolution module in the middle of the model encoder and decoder, the bidirectional fusion process of feature information is controlled;
[0016] Step S7: Use the data set processed by S1 to perform n-fold cross-training on the lung nodule segmentation model, and use the test set to output the lung nodule segmentation image and calculate the loss value;
[0017] Step S8: Adjust the model parameters according to the size of the loss value and the performance of the training phase, generate and save the trained lung nodule segmentation model, and use the evaluation index to evaluate the segmentation effect of the lung nodules.
[0018] The step S1 specifically includes the following steps:
[0019] Step S1-1: First, obtain lung CT image data and remove private information, retaining only image information, and then annotate the lesion area in the original image pixel by pixel using Dicom Viewer software, and finally obtain a label map;
[0020] Step S1-2: preprocessing the original image using random brightness and random noise image enhancement methods;
[0021] Step S1-3: perform data augmentation by rotating and cropping the image.
[0022] The specific steps of constructing the convolution module in the model encoder in step S2 are as follows:
[0023] Step S2-1: Construct a multi-scale convolution module (Module_Multi): Construct a multi-scale convolution module in the shallow layer of the encoder to extract information in the local neighborhood and outside the neighborhood, including the following sub-steps:
[0024] Step S2-1-1: Use three types of convolutions with convolution kernels of 1, 3, and 5 to form a multi-scale convolution unit, and use batch normalization and ReLU activation function after each convolution;
[0025] Step S2-1-2: Use convolution with a convolution kernel of 3, batch normalization and ReLU activation function to form a normal convolution unit;
[0026] Step S2-1-3: Use one multi-scale convolution unit and two ordinary convolution units to form a multi-scale convolution module;
[0027] Step S2-2: Construct a channel attention convolution module (Module_SE): Construct a channel attention convolution module in the deep layer of the encoder to judge the importance of each channel and effectively fuse the channels, including the following sub-steps:
[0028] Step S2-2-1: Use a global average pooling, two convolutions with a kernel of 1, a ReLU activation function, and a Sigmoid activation function to form a channel attention unit;
[0029] Step S2-2-2: Use convolution with a convolution kernel of 3, batch normalization and ReLU activation function to form a common convolution unit;
[0030] Step S2-2-3: Use one channel attention unit and three ordinary convolution units to form a channel attention convolution module.
[0031] The Transformer module (Attention_Pool) in the model encoder constructed in step S3 is achieved by adding a pooling operation to the self-attention layer of the Transformer and combining the pooling with the feature extraction process of the Transformer.
[0032] In step S4, the convolution module constructed by S2 and the Transformer module constructed by S3 are integrated in the model encoder, including three fusion architectures: a simple fusion architecture, a dual-branch parallel architecture, and a cross-fusion architecture. The three architectures can be switched according to actual conditions, including the following sub-steps:
[0033] Step S4-1: Build a simple fusion architecture: The convolution part of the encoder consists of two Module_Multi modules and two Module_SE modules, and four Attention_Pool modules are added directly after the convolution module;
[0034] Step S4-2: Build a dual-branch parallel architecture: one branch consists of two Module_Multi modules and two Module_SE modules, and the other branch consists of four Attention_Pool modules, which are finally fused through a simple convolution module;
[0035] Step S4-3: Build a cross-fusion architecture: The convolution part of the encoder consists of two Module_Multi modules and two Module_SE modules, and an Attention_Pool module is added after each convolution module.
[0036] The upsampling module in the model decoder constructed in step S5 includes the following sub-steps:
[0037] Step S5-1: Use dilated convolution to expand the receptive field without reducing the resolution of the feature map, capture multi-scale information during image resolution restoration, and accurately locate the position of the segmentation target;
[0038] Step S5-2: Grouping the feature graphs to increase the speed of encoding features by the segmentation algorithm;
[0039] Step S5-3: Rearrange the channels of the feature map to enhance the exchange of feature information between channels.
[0040] The feature bidirectional adaptive fusion module constructed in step S6 adjusts the hyperparameters in the module through back propagation of convolution to learn the ratio of feature bidirectional fusion and control the feature fusion process.
[0041] The step S7 comprises the following sub-steps:
[0042] Step S7-1: randomly divide the data set in S1 into a training set, a validation set, and a test set according to a ratio of 7:1:2;
[0043] Step S7-2: Divide the training set and the validation set by six-fold cross validation;
[0044] Step S7-3: Use the constructed segmentation model to train the data set of S7-2;
[0045] Step S7-4: Use the test set to output the lung nodule segmentation image and calculate the loss value.
[0046] A hybrid Transformer-based multi-type lung nodule accurate segmentation model, including:
[0047] Encoder: includes convolution module and Attention_Pool module and their fusion; the convolution module includes two types: multi-scale convolution module (Module_Mult) and channel attention convolution module (Module_SE); the Attention_Pool module is formed by adding pooling operation to the self-attention layer of Transformer; the fusion of the convolution module and Attention_Pool module includes three fusion architectures that can be switched according to the application situation, namely simple fusion architecture, dual-branch parallel architecture, and cross fusion architecture;
[0048] Feature bidirectional adaptive fusion module: It is set between the encoder and the decoder to realize the bidirectional flow of feature information. The hyperparameters in the module are adjusted through the back propagation of convolution to learn the proportion of bidirectional fusion of features and control the feature fusion process.
[0049] Decoder: includes multiple upsampling modules, which group and rearrange channels of feature maps after dilated convolution.
[0050] The multi-scale convolution module (Module_Mult) consists of a multi-scale convolution unit and two ordinary convolution units, wherein the multi-scale convolution unit is composed of three types of convolutions with convolution kernels of 1, 3, and 5, and batch normalization and ReLU activation function are used after each convolution, and the ordinary convolution unit is composed of convolution with a convolution kernel of 3, batch normalization and ReLU activation function;
[0051] The channel attention convolution module (Module_SE) consists of a channel attention unit and three ordinary convolution units, wherein the channel attention unit consists of a global average pooling, two convolutions with a convolution kernel of 1, a ReLU activation function, and a Sigmoid activation function.
[0052] The beneficial effects of the present invention compared with the prior art are as follows: in the multi-type lung nodule precise segmentation method and model based on hybrid Transformer provided by the present invention, the improvement and effective combination of CNN and Transformer can fully extract the local features and global features of lung nodules, thereby improving the feature extraction capability of the model; in view of the problem of feature differences between the shallow and deep layers of the lung nodule segmentation model, the present invention designs a feature bidirectional adaptive fusion module, which increases the bidirectional flow of shallow and deep feature information without increasing the depth and width of the model, so that the shallow spatial features guide the deep module segmentation, and the deep semantic features guide the shallow module training, thereby further improving the segmentation accuracy of lung nodules and improving the generalization capability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The present invention will be further described below in conjunction with the accompanying drawings:
[0054] Figure 1 is a flow chart of the segmentation method of the present invention;
[0055] Figure 2 Three fusion architecture diagrams of the convolution module and the Attention_Pool module of the present invention, where (a) is a simple fusion architecture, (b) is a dual-branch parallel architecture, and (c) is a cross-fusion architecture;
[0056] Figure 3 is a schematic diagram of an upsampling module of the present invention;
[0057] Figure 4 Schematic diagram of the feature bidirectional adaptive fusion module of the present invention, where σ and β are hyperparameters, and X and Y are feature information flows;
[0058] Figure 5 This is a schematic diagram of the structure of a multi-type lung nodule accurate segmentation model based on a hybrid Transformer, taking the cross-fusion architecture as an example. DETAILED DESCRIPTION
[0059] like Figures 1 to 5 As shown, the present invention provides a method for accurate segmentation of multiple types of lung nodules based on a hybrid Transformer, which has four main improvements: (1) The designed multi-scale convolution module (Module_Multi) and channel attention convolution module (Module_SE) can fully extract the local features of lung nodules. (2) The designed Attention_Pool module can fully extract the global features of lung nodules. The combination of the Transformer module and the pooling operation not only reduces the amount of training parameters, but also provides guidance for the image downsampling process. (3) The designed feature bidirectional adaptive fusion module can realize the bidirectional flow of feature information, thereby guiding the segmentation of deep modules and the training of shallow modules, and improving the model's ability to accurately segment multiple types of lung nodules at the same time. (4) The fusion architecture of the designed three convolution modules and the Attention_Pool module can be switched according to actual conditions during the clinical application of the segmentation model.
[0060] The flowchart of the multi-type pulmonary nodule accurate segmentation method based on hybrid Transformer proposed in this invention is as follows: Figure 1 As shown, the specific steps are as follows:
[0061] Step S1: Obtain lung CT images, complete pixel-by-pixel annotation of lung nodule lesion areas through Dicom Viewer software, and perform data enhancement and data augmentation on the images;
[0062] Step S2: Construct the convolution module in the model encoder: According to the relationship between image resolution and the number of feature channels, design the multi-scale convolution module (Module_Multi) and the channel attention convolution module (Module_SE);
[0063] Step S3: Construct the Transformer module (Attention_Pool) in the model encoder: Improve the traditional Transformer structure and combine the pooling operation with the self-attention layer in the Transformer to reduce the number of training parameters of the Transformer and retain important features during the feature extraction process.
[0064] Step S4: In the model encoder, the convolution module constructed in S2 and the Transformer module constructed in S3 are integrated to fully extract the local features and global features of the lung nodule image;
[0065] Step S5: construct an upsampling module in the model decoder: group and rearrange the feature maps after the dilated convolution;
[0066] Step S6: construct a feature bidirectional adaptive fusion module: by adding a learnable convolution module in the middle of the model encoder and decoder, the bidirectional fusion process of feature information is controlled;
[0067] Step S7: Use the data set processed by S1 to perform n-fold cross-training on the lung nodule segmentation model, and use the test set to output the lung nodule segmentation image and calculate the loss value;
[0068] Step S8: Adjust the model parameters according to the size of the loss value and the performance of the training phase, generate and save the trained lung nodule segmentation model, and use the evaluation index to evaluate the segmentation effect of the lung nodules.
[0069] The step S1 specifically includes the following steps:
[0070] Step S1-1: First, obtain lung CT image data and remove private information, retaining only image information, and then annotate the lesion area in the original image pixel by pixel using Dicom Viewer software, and finally obtain a label map;
[0071] Step S1-2: Since the original lung CT images have the characteristics of high noise, low contrast, and variable segmentation target shapes, it is necessary to use image enhancement methods such as random brightness and random noise to preprocess the original images to enhance the robustness of the segmentation model.
[0072] Step S1-3: Since CT image data is limited, it is necessary to augment the original data. The main methods are image rotation and cropping to avoid overfitting of the segmentation model.
[0073] In step S2, a convolution module in the model encoder is constructed. In the encoder of the common U-shaped pulmonary nodule segmentation model, as the model deepens, the image resolution becomes smaller and smaller, and the number of feature channels increases. The image resolution and the number of feature channels are generally inversely proportional. According to this rule, the present invention designs two convolution modules: a multi-scale convolution module (Module_Multi) and a channel attention convolution module (Module_SE). It includes the following sub-steps:
[0074] Step S2-1: Construct a multi-scale convolution module (Module_Multi): In the shallow layer of the encoder, the input resolution of the module is relatively large, and ordinary convolution can only extract local neighborhood information, which is not conducive to the extraction of detailed features of the boundary of lung nodules. Multi-scale convolution can extract information outside the neighborhood and is more suitable for large-resolution images. It includes the following sub-steps:
[0075] Step S2-1-1: Use three types of convolutions with convolution kernels of 1, 3, and 5 to form a multi-scale convolution unit, and use batch normalization and ReLU activation function after each convolution;
[0076] Step S2-1-2: Use convolution with a convolution kernel of 3, batch normalization and ReLU activation function to form a normal convolution unit;
[0077] Step S2-1-3: Use one multi-scale convolution unit and two ordinary convolution units to form a multi-scale convolution module;
[0078] Step S2-2: Construct channel attention convolution module (Module_SE): In the deep layer of the encoder, the module has many feature channels and lacks the identification index of the importance of each channel, which leads to the loss of many important features when channels are fused. Channel attention can judge the importance of each channel so that each channel can be effectively fused. It includes the following sub-steps:
[0079] Step S2-2-1: Use a global average pooling, two convolutions with a kernel of 1, a ReLU activation function, and a Sigmoid activation function to form a channel attention unit;
[0080] Step S2-2-2: Use convolution with a convolution kernel of 3, batch normalization and ReLU activation function to form a common convolution unit;
[0081] Step S2-2-3: Use one channel attention unit and three ordinary convolution units to form a channel attention convolution module.
[0082] The Transformer module (Attention_Pool) in the model encoder constructed in the step S3. CNN is good at extracting local features, but lacks the ability to represent global features. This problem can be solved by introducing the Transformer structure. However, the Transformer requires a large amount of data for training, while the available data in medicine is very scarce. In order to resolve this contradiction, the present invention improves the traditional Transformer structure. By adding a pooling operation to the self-attention layer of the Transformer, the resolution of the image can be reduced, thereby reducing the training parameters and reducing the training pressure of the Transformer module. In addition, the present invention does not add a pooling operation before and after the self-attention layer, but combines pooling into the feature extraction process of the Transformer, so that the downsampling of the image can retain important features in the feature extraction process because of the participation of the self-attention layer.
[0083] In step S4, the convolution module constructed by S2 and the Transformer module constructed by S3 are fused in the model encoder. The present invention designs three fusion architectures: simple fusion architecture, dual-branch parallel architecture, and cross-fusion architecture to adapt to different pulmonary nodule segmentation scenarios, such as Figure 2In actual use, you only need to select one of the fusion architectures for a specific scenario. The specific steps include the following:
[0084] Step S4-1: Build a simple fusion architecture: The convolution part of the encoder consists of two multi-scale convolution modules (Module_Multi) and two channel attention convolution modules (Module_SE), and four Attention_Pool modules are added directly after the convolution module;
[0085] Step S4-2: Build a dual-branch parallel architecture: one branch consists of two Module_Multi modules and two Module_SE modules, and the other branch consists of four Attention_Pool modules, which are finally fused through a simple convolution module;
[0086] Step S4-3: Build a cross-fusion architecture: The convolution part of the encoder consists of two Module_Multi modules and two Module_SE modules, and an Attention_Pool module is added after each convolution module.
[0087] In step S5, an upsampling module in the model decoder is constructed, such as Figure 3 As shown in Figure 2. The upsampling process of the feature map in the segmentation task is very important, which determines whether the features extracted by the encoder can be fully represented. It includes the following sub-steps:
[0088] Step S5-1: Use dilated convolution to expand the receptive field without reducing the resolution of the feature map, capture multi-scale information during image resolution restoration, and accurately locate the position of the segmentation target;
[0089] Step S5-2: grouping the feature maps to increase the speed of encoding features by the segmentation algorithm;
[0090] Step S5-3: Rearrange the channels of the feature map to enhance the exchange of feature information between channels.
[0091] In step S6, a feature bidirectional adaptive fusion module is constructed, such as Figure 4As shown. The shallow layer of the lung nodule segmentation model mainly extracts spatial detail features, and the deep layer mainly extracts semantic features. The bidirectional flow of feature information is realized. The introduction of spatial detail information in the deep module can make the representation of the segmentation boundary more accurate, and the introduction of semantic information in the shallow module can guide the training of the module. However, the uncontrolled fusion of shallow and deep features can easily lead to confusion and loss of features extracted by the module itself. The present invention adds a learnable convolution module in the middle of the model codec, controls the bidirectional fusion process of feature information through the ratio of bidirectional fusion of convolution learning features, thereby improving the generalization ability of the model, so that it can accurately segment multiple types of lung nodules at the same time.
[0092] The step S7 comprises the following sub-steps:
[0093] Step S7-1: randomly divide the data set in S1 into a training set, a validation set, and a test set according to a ratio of 7:1:2;
[0094] Step S7-2: Divide the training set and the validation set by six-fold cross validation;
[0095] Step S7-3: Use the constructed segmentation model to train the data set of S7-2;
[0096] Step S7-4: Use the test set to output the lung nodule segmentation image and calculate the loss value.
[0097] In step S8, the model parameters are adjusted according to the size of the loss value and the performance of the training phase to improve the segmentation accuracy during model testing. The best model is saved, and the segmentation effect of the pulmonary nodules is evaluated by indicators such as the Dice coefficient, sensitivity, and Hausdorff distance.
[0098] The present invention also proposes a segmentation model. The schematic diagram of the segmentation model structure taking the cross-fusion architecture as an example is as follows: Figure 5 As shown, including:
[0099] Encoder: includes convolution module and Attention_Pool module and their fusion; the convolution module includes two types: multi-scale convolution module (Module_Mult) and channel attention convolution module (Module_SE); the Attention_Pool module is formed by adding pooling operation to the self-attention layer of Transformer; the fusion of the convolution module and Attention_Pool module adopts a cross-fusion architecture;
[0100] Feature bidirectional adaptive fusion module: It is set between the encoder and the decoder to realize the bidirectional flow of feature information. The hyperparameters in the module are adjusted through the back propagation of convolution to learn the proportion of bidirectional fusion of features and control the feature fusion process;
[0101] Decoder: includes multiple upsampling modules, which group and rearrange channels of feature maps after dilated convolution.
[0102] The main advantages of the present invention are: (1) The improvement and effective combination of the convolution module and Transformer can extract local features and global features at the same time, improving the feature extraction ability of the model. (2) The bidirectional adaptive fusion of features increases the interaction between the deep and shallow layers of the model. The shallow spatial information guides the deep module segmentation, and the deep semantic information guides the shallow module training, thereby further improving the segmentation accuracy of lung nodules and improving the generalization ability of the model.
[0103] What needs to be explained about the specific structure of the present invention is that the connection relationship between the various component modules adopted in the present invention is definite and feasible. Except for the special instructions in the embodiments, the specific connection relationship can bring about the corresponding technical effect, and solve the technical problem proposed by the present invention on the premise of not relying on the execution of the corresponding software program. The components, modules, models of specific components and the connection methods between each other appearing in the present invention, as well as the conventional use methods and expected technical effects brought about by the above-mentioned technical features, except for the specific instructions, all belong to the disclosed contents in patents, journal articles, technical manuals, technical dictionaries, and textbooks that can be obtained by technical personnel in this field before the application date, or belong to the existing technologies such as conventional technology and common knowledge in this field, and there is no need to elaborate, so that the technical solution provided in this case is clear, complete and feasible, and the corresponding physical products can be reproduced or obtained according to the technical means.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-type pulmonary nodule accurate segmentation method based on hybrid Transformer, characterized by: The steps include: Step S1: Constructing a data set: Obtain lung CT images, complete the pixel-by-pixel annotation of the lung nodule lesion area through Dicom Viewer software, and perform data enhancement and data augmentation on the image; Construct a lung nodule segmentation model, including: Step S2: Construct the convolution module in the model encoder: According to the relationship between image resolution and the number of feature channels, design the multi-scale convolution module (Module_Multi) and the channel attention convolution module (Module_SE); The specific steps of constructing the convolution module in the model encoder in step S2 are as follows: Step S2-1: Construct a multi-scale convolution module (Module_Multi): Construct a multi-scale convolution module in the shallow layer of the encoder to extract information in the local neighborhood and outside the neighborhood, including the following sub-steps: Step S2-1-1: Use three types of convolutions with convolution kernels of 1, 3, and 5 to form a multi-scale convolution unit, and use batch normalization and ReLU activation function after each convolution; Step S2-1-2: Use convolution with a convolution kernel of 3, batch normalization and ReLU activation function to form a normal convolution unit; Step S2-1-3: Use one multi-scale convolution unit and two ordinary convolution units to form a multi-scale convolution module; Step S2-2: Construct a channel attention convolution module (Module_SE): Construct a channel attention convolution module in the deep layer of the encoder to judge the importance of each channel so that each channel can be effectively integrated, including the following sub-steps: Step S2-2-1: Use a global average pooling, two convolutions with a kernel of 1, a ReLU activation function, and a Sigmoid activation function to form a channel attention unit; Step S2-2-2: Use convolution with a convolution kernel of 3, batch normalization and ReLU activation function to form a common convolution unit; Step S2-2-3: Use one channel attention unit and three ordinary convolution units to form a channel attention convolution module; Step S3: Construct the Transformer module (Attention_Pool) in the model encoder: Improve the traditional Transformer structure and combine the pooling operation with the self-attention layer in the Transformer; Step S4: Fuse the convolution module constructed in S2 and the Transformer module constructed in S3 in the model encoder; Step S5: construct an upsampling module in the model decoder: group the feature maps after the dilated convolution and rearrange the channels; Step S6: construct a feature bidirectional adaptive fusion module: by adding a learnable convolution module in the middle of the model encoder and decoder, the bidirectional fusion process of feature information is controlled; Step S7: Use the data set processed by S1 to perform n-fold cross-training on the lung nodule segmentation model, and use the test set to output the lung nodule segmentation image and calculate the loss value; Step S8: Adjust the model parameters according to the size of the loss value and the performance of the training phase, generate and save the trained lung nodule segmentation model, and use the evaluation index to evaluate the segmentation effect of the lung nodules.
2. According to claim 1, a method for accurate segmentation of multi-type pulmonary nodules based on hybrid Transformer, characterized in that: The step S1 specifically includes the following steps: Step S1-1: First, obtain lung CT image data and remove private information, retaining only image information, and then annotate the lesion area in the original image pixel by pixel using Dicom Viewer software, and finally obtain a label map; Step S1-2: preprocessing the original image using random brightness and random noise image enhancement methods; Step S1-3: perform data augmentation by rotating and cropping the image.
3. According to claim 1, a method for accurate segmentation of multi-type pulmonary nodules based on hybrid Transformer, characterized in that: The Transformer module (Attention_Pool) in the model encoder constructed in step S3 is achieved by adding a pooling operation to the self-attention layer of the Transformer and combining the pooling with the feature extraction process of the Transformer.
4. According to claim 3, a method for accurate segmentation of multi-type pulmonary nodules based on hybrid Transformer, characterized in that: In step S4, the convolution module constructed by S2 and the Transformer module constructed by S3 are integrated in the model encoder, including three fusion architectures: a simple fusion architecture, a dual-branch parallel architecture, and a cross-fusion architecture. The three architectures can be switched according to actual conditions, including the following sub-steps: Step S4-1: Build a simple fusion architecture: The convolution part of the encoder consists of two Module_Multi modules and two Module_SE modules, and four Attention_Pool modules are added directly after the convolution module; Step S4-2: Build a dual-branch parallel architecture: one branch consists of two Module_Multi modules and two Module_SE modules, and the other branch consists of four Attention_Pool modules, which are finally fused through a simple convolution module; Step S4-3: Build a cross-fusion architecture: The convolution part of the encoder consists of two Module_Multi modules and two Module_SE modules, and an Attention_Pool module is added after each convolution module.
5. The method for accurate segmentation of multi-type pulmonary nodules based on hybrid Transformer according to claim 1, characterized in that: The upsampling module in the model decoder constructed in step S5 includes the following sub-steps: Step S5-1: Use dilated convolution to expand the receptive field without reducing the resolution of the feature map, capture multi-scale information during image resolution restoration, and accurately locate the position of the segmentation target; Step S5-2: Grouping the feature graphs to increase the speed of encoding features by the segmentation algorithm; Step S5-3: Rearrange the channels of the feature map to enhance the exchange of feature information between channels.
6. The method for accurate segmentation of multi-type pulmonary nodules based on hybrid Transformer according to claim 1, characterized in that: The feature bidirectional adaptive fusion module constructed in step S6 adjusts the hyperparameters in the module through back propagation of convolution to learn the ratio of feature bidirectional fusion and control the feature fusion process.
7. The method for accurate segmentation of multi-type pulmonary nodules based on hybrid Transformer according to claim 1, characterized in that: The step S7 comprises the following sub-steps: Step S7-1: randomly divide the data set in S1 into a training set, a validation set, and a test set according to a ratio of 7:1:2; Step S7-2: Perform six-fold cross-validation division on the training set and the validation set; Step S7-3: Use the constructed segmentation model to train the data set of S7-2; Step S7-4: Use the test set to output the lung nodule segmentation image and calculate the loss value.
8. A multi-type lung nodule accurate segmentation model based on hybrid Transformer, characterized by: include: Encoder: includes convolution module and Attention_Pool module and their fusion; the convolution module includes two types: multi-scale convolution module (Module_Mult) and channel attention convolution module (Module_SE); the Attention_Pool module is formed by adding pooling operation to the self-attention layer of Transformer; the fusion of the convolution module and Attention_Pool module includes three fusion architectures that can be switched according to the application situation, namely simple fusion architecture, dual-branch parallel architecture, and cross fusion architecture; Feature bidirectional adaptive fusion module: It is set between the encoder and the decoder to realize the bidirectional flow of feature information. The hyperparameters in the module are adjusted through the back propagation of convolution to learn the proportion of bidirectional fusion of features and control the feature fusion process. Decoder: includes multiple upsampling modules, which group and rearrange channels of feature maps after dilated convolution; The multi-scale convolution module (Module_Mult) consists of a multi-scale convolution unit and two ordinary convolution units, wherein the multi-scale convolution unit is composed of three types of convolutions with convolution kernels of 1, 3, and 5, and batch normalization and ReLU activation function are used after each convolution, and the ordinary convolution unit is composed of convolution with a convolution kernel of 3, batch normalization and ReLU activation function; The channel attention convolution module (Module_SE) consists of a channel attention unit and three ordinary convolution units, wherein the channel attention unit consists of a global average pooling, two convolutions with a convolution kernel of 1, a ReLU activation function, and a Sigmoid activation function.
Citation Information
Patent Citations
Lung CT image segmentation method based on transfer learning and attention mechanism
CN115457049A
Deep network lung texture recogniton method combined with multi-scale attention
US20210390338A1