CNN and Transform-based pulmonary tuberculosis CT image segmentation method
Through the parallel dual-branch structure and cross-enhanced fusion module, combined with the characteristics of CNN and Transformer, the problem of difficult to capture the local details and global relationships of pulmonary tuberculosis CT image segmentation in the prior art is solved, and a higher precision lesion segmentation is achieved, and more reliable auxiliary diagnostic tools are provided.
Patent Information
- Application Number
- CN202510964188.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-08-12
AI Technical Summary
In the existing CT image segmentation method of tuberculosis, traditional single-branch CNNs are difficult to capture local texture and global spatial relationships at the same time, while the pure Transformer model will lead to local details loss, and the existing feature fusion mechanism is rigid, resulting in inaccurate segmentation results.
Using a parallel dual-branch structure, combining CNN branches and Transformer branches, the fusion module and multi-scale context information extraction module are cross-enhanced, and the lesion boundary sensitivity is enhanced, and the model training is optimized by weight loss function, to achieve dynamic weight allocation and fusion of features.
It improves the accuracy and efficiency of CT image segmentation of tuberculosis, can accurately capture local details and long-range dependencies, and provides more reliable auxiliary diagnostic support.
Smart Images

Figure CN120471948A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical data prediction, and in particular relates to a pulmonary tuberculosis CT image segmentation method based on CNN and Transformer. Background Art
[0002] CT image segmentation of pulmonary tuberculosis is a key technology for auxiliary diagnosis. Pulmonary tuberculosis manifests as multi-scale lesions (miliary nodules 3-5 mm, cavities >10 mm, fibrous foci, etc.). Traditional single-branch CNNs struggle to simultaneously capture both local texture and global spatial relationships. Early infiltrative lesions have a small CT value difference from normal lung parenchyma (<100 HU), and conventional preprocessing tends to enhance non-target areas such as the mediastinum and chest wall. Caseous necrosis lesions have unclear boundaries with surrounding inflammatory exudates, resulting in fragmented edge segmentation results.
[0003] Existing pure CNN models (such as U-Net) have limited convolutional receptive fields and cannot model long-range dependencies of pulmonary metastatic lesions (such as the spatial distribution of miliary nodules). Pure Transformer models can lead to the loss of local details (such as the edges of calcifications) and increase the computational complexity of high-resolution CT. Based on the above, the feature fusion mechanism also has rigidity problems. Traditional splicing / additive fusion ignores the difference in contribution between CNN features (local edges) and Transformer features (global semantics), resulting in: the strong edge features of calcification foci are overwhelmed by global semantics; bronchial vascular bundles are misidentified as cavitary tuberculosis; single-scale dilated convolution will lead to insufficient coverage of small nodules or over-smoothing of large solid boundary due to the fixed expansion rate; post-processing boundary optimization will lead to the separation of segmentation and boundary optimization, resulting in poor real-time performance.
[0004] When designing the loss function, it is necessary to consider the adaptation to medical characteristics and the limitation of window width setting in preprocessing. Summary of the Invention
[0005] To solve the above problems in the prior art, the present invention provides a pulmonary tuberculosis CT image segmentation method based on CNN and Transformer, comprising: S1: Acquire a CT image, and preprocess the CT image by performing windowing and contrast-limited adaptive histogram equalization; S2: Extract local texture, edge features, and global long-range dependency features of lung lesions in preprocessed CT images through a parallel dual-branch structure; S3: The extracted local texture, edge features and global long-range dependency features are input into the cross-enhancement fusion module, and the features are complementary fused through dynamic weight allocation to obtain the fused features; S4: inputting the fused features into a multi-scale context information extraction module, and enhancing the sensitivity of lesion boundaries through dilated convolutions with different dilation rates; S5: Fuse the encoder features with the decoder features through skip connections, and output the segmentation result after restoring the resolution based on upsampling; S6: Use weighted loss function to optimize model training.
[0006] Specifically, the parallel dual-branch structure described in S2 includes a CNN branch and a Transformer branch; the CNN branch uses an improved residual convolution layer to replace the traditional U-Net convolution module to extract the local texture and edge features of lung lesions; the Transformer branch uses a sliding window Transformer structure to capture the global long-range dependency features of the image.
[0007] Specifically, the execution operation of the cross enhancement fusion module in S3 is: A bidirectional attention mechanism is established between the dual-branch features, and the feature contribution weights of the CNN branch and the Transformer branch are dynamically adjusted through feature channel weighting and spatial recalibration.
[0008] Specifically, the multi-scale context information extraction module in S4 includes three parallel dilated convolutional layers, and the output features are channel-wise concatenated and then dimensionally reduced by 1×1 convolution.
[0009] Specifically, the windowing process in S1 includes: The original grayscale values of the CT image were mapped to the lung window range, where the upper threshold of the lung window range was the dividing point between the lung parenchyma and the mediastinum tissue, and the upper threshold of the lung window range was the lower limit of the air density. The mapped pixel values were linearly normalized to the interval [0, 1]. Windowing processing was used as a pre-step for contrast-limited adaptive histogram equalization preprocessing to ensure that contrast enhancement only acted on lung tissue by limiting the grayscale distribution of the region of interest.
[0010] Specifically, the weighted loss function in S6 is: , Among them, Loss is the multi-task combination loss, BoundaryLoss is the boundary weighted function based on distance transformation, DiceLoss is the loss generated by the similarity of the overlapping area between the segmentation result and the true label, FocalLoss is the sample imbalance loss, α, β, γ are weight coefficients, and α+β+γ=1.
[0011] Furthermore, the residual convolution layer adopts a pre-activation structure, the batch normalization layer and the linear rectification activation function are placed before the convolution layer, the input features are first bifurcated to convolution paths of different scales, and the output features are fused with the residual connection after matrix splicing.
[0012] Furthermore, the sliding window Transformer branch captures long-range dependencies in the following ways: The input image is divided into non-overlapping image blocks; feature mapping is performed on the image blocks through a linear embedding layer; and feature extraction is performed alternately using a window multi-head self-attention mechanism and a sliding window multi-head self-attention mechanism. The window multi-head self-attention mechanism calculates self-attention within a local non-overlapping window to reduce computational complexity; the sliding window multi-head self-attention mechanism achieves global interaction through cross-window connections; and a hierarchical feature map is constructed through staged downsampling to achieve multi-scale feature fusion.
[0013] Furthermore, the bidirectional attention mechanism implements dynamic weight allocation through channel-spatial collaborative attention, specifically including: Global average pooling is performed on the CNN branch features and the Transformer branch features to generate a channel description vector. The two-channel vectors are concatenated and input into two fully connected layers, where they are activated by an activation function to generate channel attention weights. Based on the channel attention weights, channel weighting is performed on the CNN branch features and the Transformer branch features to enhance the characteristic response of the lesion-related channels. The weighted dual-branch features are concatenated along the channel dimension and a 3×3 convolution is performed to generate a spatial attention map. Normalizing the spatial attention map to generate a spatial weight mask; The spatial weight mask is multiplied with the CNN branch features and the Transformer branch features respectively to enhance the feature activation of the lesion boundary area.
[0014] Furthermore, the expansion rates of the three parallel dilated convolutional layers are 2, 4, and 6 respectively, and after the output features are spliced through channels, the method for enhancing the sensitivity of the lesion boundary is as follows: Input the hole convolution splicing features into the boundary-sensitive convolution layer to generate the initial boundary probability map; spatially modulating the original splicing features using the initial boundary probability map; Perform 1×1 convolution dimensionality reduction on the modulated features to generate multi-scale fusion features; splice the multi-scale fusion features with the initial boundary probability map along the channel dimension, and generate the final boundary enhancement features through 1×1 convolution.
[0015] The present invention's pulmonary tuberculosis CT image segmentation method based on CNN and Transformer effectively solves many problems existing in the prior art. Through a parallel dual-branch structure, the CNN branch and the Transformer branch respectively extract local texture, edge features, and global long-range dependency features, which makes up for the shortcomings of traditional single-branch CNN and pure Transformer models, and can capture local details and model long-range dependencies. The dynamic weight allocation mechanism of the cross-enhancement fusion module makes feature fusion more reasonable, avoids the rigidity of traditional fusion methods, and gives full play to the advantages of CNN features and Transformer features. The multi-scale context information extraction module uses dilated convolution with different expansion rates to enhance the sensitivity of lesion boundaries and solve the limitations of single-scale dilated convolution. Medical characteristics are taken into account in the loss function design, and model training is optimized by weighted loss function to improve the accuracy and stability of segmentation. At the same time, the preprocessing method of the present invention effectively avoids the problem of conventional preprocessing enhancing non-target areas. In summary, the method of the present invention can improve the accuracy and efficiency of pulmonary tuberculosis CT image segmentation, providing more reliable technical support for auxiliary diagnosis of pulmonary tuberculosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0017] Figure 1 This is a flow chart of a pulmonary tuberculosis CT image segmentation method based on CNN and Transformer of the present invention. DETAILED DESCRIPTION
[0018] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0019] See also Figure 1 , a pulmonary tuberculosis CT image segmentation method based on CNN and Transformer, including: S1: Acquire a CT image, and preprocess the CT image by performing windowing and contrast-limited adaptive histogram equalization; S2: Extract local texture, edge features, and global long-range dependency features of lung lesions in preprocessed CT images through a parallel dual-branch structure; S3: The extracted local texture, edge features and global long-range dependency features are input into the cross-enhancement fusion module, and the features are complementary fused through dynamic weight allocation to obtain the fused features; S4: inputting the fused features into a multi-scale context information extraction module, and enhancing the sensitivity of lesion boundaries through dilated convolutions with different dilation rates; S5: Fuse the encoder features with the decoder features through skip connections, and output the segmentation result after restoring the resolution based on upsampling; S6: Use weighted loss function to optimize model training.
[0020] In this embodiment, the skip connection is implemented as follows: the features of the i-th layer of the encoder are fused with the features of the (n-i)-th layer of the decoder (n is the total number of layers); the encoder features are adjusted for the number of channels through 1×1 convolution, and are concatenated with the upsampled decoder features, and information transmission is enhanced through a dense Swin Transformer block (DSTB); the decoder uses transposed convolution to gradually upsample (×2 times) to restore to the original resolution, and the output layer is 1×1 convolution + Sigmoid to generate a binary segmentation mask.
[0021] Specifically, the parallel dual-branch structure described in S2 includes a CNN branch and a Transformer branch; the CNN branch uses an improved residual convolution layer to replace the traditional U-Net convolution module to extract the local texture and edge features of lung lesions; the Transformer branch uses a sliding window Transformer structure to capture the global long-range dependency features of the image.
[0022] Specifically, the execution operation of the cross enhancement fusion module in S3 is: A bidirectional attention mechanism is established between the dual-branch features, and the feature contribution weights of the CNN branch and the Transformer branch are dynamically adjusted through feature channel weighting and spatial recalibration.
[0023] Specifically, the multi-scale context information extraction module in S4 includes three parallel dilated convolutional layers, and the output features are channel-wise concatenated and then dimensionally reduced by 1×1 convolution.
[0024] Specifically, the windowing process in S1 includes: The original grayscale values of the CT image were mapped to the lung window range, where the upper threshold of the lung window range was the dividing point between the lung parenchyma and the mediastinum tissue, and the upper threshold of the lung window range was the lower limit of the air density. The mapped pixel values were linearly normalized to the interval [0, 1]. Windowing processing was used as a pre-step for contrast-limited adaptive histogram equalization preprocessing to ensure that contrast enhancement only acted on lung tissue by limiting the grayscale distribution of the region of interest.
[0025] In this embodiment, the grayscale values of the original DICOM image are converted to Hounsfield units (HU). The HU values are limited to the lung window range through a linear mapping function: , Among them, I window is the pixel value after windowing, I HU is the HU value corresponding to the original DICOM image. This operation effectively reduces the grayscale range of the image, focusing on the lung tissue information. Next, the windowed pixel values are linearly normalized to the interval [0, 1], ensuring a more stable numerical range for the image data during subsequent processing. Contrast-limited adaptive histogram equalization further enhances the contrast within the lung tissue, making the difference between lesions and normal tissue more pronounced while avoiding enhancement of non-target areas such as the mediastinum and chest wall.
[0026] Specifically, the weighted loss function in S6 is: , Among them, Loss is the multi-task combination loss, BoundaryLoss is the boundary weighted function based on distance transformation, DiceLoss is the loss generated by the similarity of the overlapping area between the segmentation result and the true label, FocalLoss is the sample imbalance loss, α, β, γ are weight coefficients, and α+β+γ=1.
[0027] Furthermore, the residual convolution layer adopts a pre-activation structure, the batch normalization layer and the linear rectification activation function are placed before the convolution layer, the input features are first bifurcated to convolution paths of different scales, and the output features are fused with the residual connection after matrix splicing.
[0028] Furthermore, the sliding window Transformer branch captures long-range dependencies in the following ways: The input image is divided into non-overlapping image blocks; feature mapping is performed on the image blocks through a linear embedding layer; and feature extraction is performed alternately using a window multi-head self-attention mechanism and a sliding window multi-head self-attention mechanism. The window multi-head self-attention mechanism calculates self-attention within a local non-overlapping window to reduce computational complexity; the sliding window multi-head self-attention mechanism achieves global interaction through cross-window connections; and a hierarchical feature map is constructed through staged downsampling to achieve multi-scale feature fusion.
[0029] Furthermore, the bidirectional attention mechanism implements dynamic weight allocation through channel-spatial collaborative attention, specifically including: Global average pooling is performed on the CNN branch features and the Transformer branch features to generate a channel description vector. The two-channel vectors are concatenated and input into two fully connected layers, where they are activated by an activation function to generate channel attention weights. Based on the channel attention weights, channel weighting is performed on the CNN branch features and the Transformer branch features to enhance the characteristic response of the lesion-related channels. The weighted dual-branch features are concatenated along the channel dimension and a 3×3 convolution is performed to generate a spatial attention map. Normalizing the spatial attention map to generate a spatial weight mask; The spatial weight mask is multiplied with the CNN branch features and the Transformer branch features respectively to enhance the feature activation of the lesion boundary area.
[0030] In this embodiment, the CNN feature F c and Transformer feature F t Perform global average pooling to generate channel description vector V c , V t ; Splicing vector [V c , V t ]Input two layers of fully connected layers (FC→ReLU→FC), output channel weight W c , W t ; Weighted feature F c ˆ=W c ⊗F c , F t ˆ=W t ⊗F t , enhancing lesion-related channel responses; Splicing weighted features [F c ˆ, F t ˆ]Generate spatial attention map M through 3×3 convolution s ; Sigmoid normalization to obtain spatial mask M s ˆ; Multiply with the original feature: , CNN edge features are enhanced in the lesion boundary area, and Transformer global features are strengthened in the homogeneous area.
[0031] Furthermore, the expansion rates of the three parallel dilated convolutional layers are 2, 4, and 6 respectively, and after the output features are spliced through channels, the method for enhancing the sensitivity of the lesion boundary is as follows: Input the hole convolution splicing features into the boundary-sensitive convolution layer to generate the initial boundary probability map; spatially modulating the original splicing features using the initial boundary probability map; Perform 1×1 convolution dimensionality reduction on the modulated features to generate multi-scale fusion features; splice the multi-scale fusion features with the initial boundary probability map along the channel dimension, and generate the final boundary enhancement features through 1×1 convolution.
[0032] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments using the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A pulmonary tuberculosis CT image segmentation method based on CNN and Transformer, characterized in that: include: S1: Acquire a CT image, and preprocess the CT image by performing windowing and contrast-limited adaptive histogram equalization; S2: Extract local texture, edge features, and global long-range dependency features of lung lesions in preprocessed CT images through a parallel dual-branch structure; S3: The extracted local texture, edge features and global long-range dependency features are input into the cross-enhancement fusion module, and the features are complementary fused through dynamic weight allocation to obtain the fused features; S4: inputting the fused features into a multi-scale context information extraction module, and enhancing the sensitivity of lesion boundaries through dilated convolutions with different dilation rates; S5: Fuse the encoder features with the decoder features through skip connections, and output the segmentation result after restoring the resolution based on upsampling; S6: Use weighted loss function to optimize model training.
2. The method according to claim 1, characterized in that The parallel dual-branch structure described in S2 includes a CNN branch and a Transformer branch; the CNN branch uses an improved residual convolution layer to replace the traditional U-Net convolution module to extract local texture and edge features of lung lesions; the Transformer branch uses a sliding window Transformer structure to capture the global long-range dependency features of the image.
3. The method according to claim 1, characterized in that The execution operation of the cross enhancement fusion module in S3 is: A bidirectional attention mechanism is established between the dual-branch features, and the feature contribution weights of the CNN branch and the Transformer branch are dynamically adjusted through feature channel weighting and spatial recalibration.
4. The method according to claim 1, wherein The multi-scale context information extraction module in S4 includes three parallel dilated convolutional layers, and the output features are channel-wise concatenated and then dimensionally reduced by 1×1 convolution.
5. The method according to claim 1, wherein The windowing process in S1 specifically includes: The original grayscale values of the CT image were mapped to the lung window range, where the upper threshold of the lung window range was the dividing point between the lung parenchyma and the mediastinum tissue, and the upper threshold of the lung window range was the lower limit of the air density. The mapped pixel values were linearly normalized to the interval [0, 1]. Windowing processing was used as a pre-step for contrast-limited adaptive histogram equalization preprocessing to ensure that contrast enhancement only acted on lung tissue by limiting the grayscale distribution of the region of interest.
6. The method according to claim 1, characterized in that The weighted loss function in S6 is: , Among them, Loss is the multi-task combination loss, BoundaryLoss is the boundary weighted function based on distance transformation, DiceLoss is the loss generated by the similarity of the overlapping area between the segmentation result and the true label, FocalLoss is the sample imbalance loss, α, β, γ are weight coefficients, and α+β+γ=1.
7. The method according to claim 2, characterized in that The residual convolution layer adopts a pre-activation structure, with a batch normalization layer and a linear rectification activation function placed before the convolution layer. The input features are first bifurcated into convolution paths of different scales, and the output features are fused with the residual connection after matrix splicing.
8. The method according to claim 2, characterized in that The sliding window Transformer branch captures long-range dependencies in the following way: The input image is divided into non-overlapping image blocks; feature mapping is performed on the image blocks through a linear embedding layer; and feature extraction is performed alternately using a window multi-head self-attention mechanism and a sliding window multi-head self-attention mechanism. The window multi-head self-attention mechanism calculates self-attention within a local non-overlapping window to reduce computational complexity; the sliding window multi-head self-attention mechanism achieves global interaction through cross-window connections; and a hierarchical feature map is constructed through staged downsampling to achieve multi-scale feature fusion.
9. The method according to claim 3, characterized in that The bidirectional attention mechanism achieves dynamic weight allocation through channel-space collaborative attention, specifically including: performing global average pooling on CNN branch features and Transformer branch features respectively to generate channel description vectors; concatenating the two channel vectors and inputting them into two fully connected layers, activating them through an activation function to generate channel attention weights; and performing channel weighting on CNN branch features and Transformer branch features based on the channel attention weights to enhance the feature response of lesion-related channels. The weighted dual-branch features are concatenated along the channel dimension and a 3×3 convolution is performed to generate a spatial attention map. Normalizing the spatial attention map to generate a spatial weight mask; The spatial weight mask is multiplied with the CNN branch features and the Transformer branch features respectively to enhance the feature activation of the lesion boundary area.
10. The method according to claim 4, characterized in that The expansion rates of the three parallel dilated convolutional layers are 2, 4, and 6 respectively, and after the output features are spliced through channels, the method for enhancing the sensitivity of the lesion boundary is as follows: Input the hole convolution splicing features into the boundary-sensitive convolution layer to generate the initial boundary probability map; spatially modulating the original splicing features using the initial boundary probability map; Perform 1×1 convolution dimensionality reduction on the modulated features to generate multi-scale fusion features; The multi-scale fusion features are concatenated with the initial boundary probability map along the channel dimension, and the final boundary enhancement features are generated through 1×1 convolution.
Citation Information
Patent Citations
Automatic segmentation of liver and tumor in CT image based on residual UNet and efficient multi-scale attention method
CN116994113A
Polyp cutting method and device
CN119251244A
Cited By
Feature enhancement system and method for thyroid and breast surgery CT image
CN120725905A
A feature enhancement system and method for breast and thyroid surgery CT images
CN120725905B
Automatic lung lesion detection and diagnosis method based on CT image
CN120953287A
Automatic detection and diagnosis method of lung lesions based on CT image
CN120953287B
CNN-ViT-Meta fusion model-based pulmonary tuberculosis intelligent identification method
CN121121406A