Lung infection image segmentation method based on bidirectional guide multi-scale feature decoding
Through the hierarchical Swin Transformer encoder and the progressive bidirectional guided multi-scale feature decoding strategy, the multi-scale characteristics and boundary positioning problems in the segmentation of lung infection image lesions are solved, and high-precision lung infection image segmentation is achieved to meet clinical needs.
Patent Information
- Application Number
- CN202511204585.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing methods for segmenting lung infection image lesions suffer from insufficient multi-scale feature expression, imperfect fusion of encoding and decoding features, lack of feature guidance mechanism, and insufficient boundary positioning accuracy, resulting in inconsistent segmentation results that are difficult to meet clinical needs.
A hierarchical Swin Transformer encoder and a progressive bidirectional guided multi-scale feature decoding strategy are adopted to achieve accurate segmentation of lung infection lesions through adaptive preprocessing, hierarchical feature encoding, progressive bidirectional guided feature fusion and multi-scale decoding.
The segmentation accuracy and boundary positioning accuracy are improved to meet clinical diagnosis needs. The segmentation accuracy is improved by 8-15%, the boundary IoU index is improved by 12%, and the sensitivity of small-scale lesion detection is improved by 20%. The processing time is about 2 seconds, which meets the requirements of clinical real-time processing.
Smart Images

Figure CN120707864A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and in particular relates to a lung infection image segmentation method based on bidirectional guided multi-scale feature decoding, which can accurately segment lesions in various lung infection images. Background Art
[0002] Lung infections are a major threat to human health, including pneumonia caused by the novel coronavirus, bacterial pneumonia, viral pneumonia, fungal pneumonia, and other inflammatory lung diseases. Chest CT imaging is an important tool for diagnosing lung infections, and accurate lesion segmentation is crucial.
[0003] Currently, lesion segmentation in lung infection images relies primarily on subjective judgment and manual annotation by physicians, which presents significant challenges. Manual segmentation is inefficient and unable to meet the needs of large-scale screening. Subjectivity is high, and diagnostic criteria vary among physicians, leading to inconsistent segmentation results. Accuracy is also limited, making it easy to miss or misidentify early lesions and those with blurred boundaries.
[0004] Although the existing deep learning-based medical image segmentation methods have alleviated the above problems to a certain extent, they still have important technical defects. First, the multi-scale feature expression is insufficient. Existing methods usually only use feature information of a single scale, which makes it difficult to effectively deal with the multi-scale characteristics of lung infection lesions and cannot accurately segment tiny ground-glass shadows and large areas of consolidation at the same time. Secondly, the encoding and decoding feature fusion mechanism is imperfect. The traditional encoder-decoder structure lacks effective utilization of the complementarity between encoding features and decoding features, resulting in insufficient fusion of semantic information and detail information. Thirdly, the feature guidance mechanism is missing. Existing methods lack effective feature guidance strategies and cannot fully utilize the mutual promotion between features at different levels. Finally, the boundary positioning accuracy is limited. Pulmonary infection lesions usually have blurred boundaries and low contrast with normal tissues. Existing methods are insufficient to accurately locate the boundaries of lesions. Summary of the Invention
[0005] In response to the above problems, the purpose of the present invention is to provide a lung infection image segmentation method based on bidirectional guided multi-scale feature decoding. This method adopts a hierarchical shift window Transformer (Swin Tranformer) encoder, and through an innovative bidirectional guided feature fusion mechanism and an adaptive multi-scale decoding strategy, fully utilizes the complementarity of global and local features, effectively solves the segmentation difficulties caused by the multi-scale characteristics of lung infection lesions, and achieves accurate segmentation of lesions of different types and sizes.
[0006] To achieve the above objectives, the present invention provides a lung infection intelligent segmentation method based on a bidirectional guided multi-scale feature decoding network, comprising the following steps:
[0007] Step 1: Medical Image Preprocessing and Standardization
[0008] Receive lung CT image data, perform adaptive preprocessing operations, generate standardized image data suitable for deep learning model processing, and provide high-quality input for subsequent feature extraction.
[0009] The preprocessing process is as follows: first, the window width and window position are adaptively adjusted to optimize the contrast between lung tissue and lesion area; then, noise removal is performed, and the edge-preserving bilateral filtering algorithm is used to remove noise interference in CT scan; finally, pixel value normalization and size standardization are performed, and the pixel value normalization process maps the HU value range to Interval, size standardization adjusts all images to uniform pixels, and uses bilinear interpolation to maintain image quality. The output of this step is the standardized image data As input data for subsequent Transformer feature encoding.
[0010] Step 2: Hierarchical Swin Transformer feature encoding
[0011] Standardized image data output from step 1 based on the Swin Transformer architecture Perform hierarchical feature encoding to generate multi-scale feature representation .
[0012] The feature encoding process: the input standardized image data The image is divided into non-overlapping blocks, and each block is mapped to a feature vector through a linear embedding layer and input into the encoder; the encoder consists of four progressive downsampling stages, each of which consists of a series of Swin Transformer blocks, generating feature maps of four different resolutions to obtain a multi-scale feature representation. , the feature dimension increases layer by layer, and the resolution decreases layer by layer.
[0013] Step 3: Progressive bidirectional guided multi-scale feature decoding
[0014] Based on the encoding features generated in step 2 The method uses a progressive bidirectionally guided multi-scale feature decoding strategy. Through an iterative process from deep to shallow layers, it achieves adaptive enhancement of coded features, bidirectional guided fusion, and multi-scale decoding, generating highly accurate lung infection lesion segmentation results. This strategy unifies the traditional independent enhancement, fusion, and decoding steps into a collaboratively optimized progressive processing flow.
[0015] 3.1 Deep Feature Decoding Initialization
[0016] Deepest encoding features As the starting feature of the entire progressive decoding process, cross-layer feature connection is first performed:
[0017]
[0018] in Indicates 2x bilinear upsampling, Represents channel dimension splicing. This cross-layer connection combines deep semantic information with shallow detail information to form a fusion feature containing multi-level information. The feature is then enhanced to obtain the enhanced encoding feature:
[0019]
[0020] in Represents a combination of operations, including cascaded convolution to extract spatial features, batch normalization to stabilize the training process, and linear rectification activation function.
[0021] Enhanced encoding features Directly input the multi-scale decoder to generate initial decoding features ,The decoder uses a four-way parallel processing strategy to process feature information of different resolutions. The original resolution path directly processes the feature , maintaining complete spatial detail information. The 1 / 2 resolution path is processed by adaptive average pooling and channel attention modules:
[0022]
[0023] in is the channel attention module, It is 2×2 average pooling.
[0024] 1 / 4 resolution path through:
[0025]
[0026] Generate semantically rich feature representations that incorporate more semantic information while retaining important details.
[0027] 1 / 8 resolution path through:
[0028]
[0029] Provide high-level semantic guidance information to ensure the semantic consistency of the segmentation results. for Each processing path is equipped with a channel attention module to adaptively adjust the importance of feature channels by learning the dependencies between channels.
[0030] In the multi-scale feature aggregation stage, the four-way features are adaptively fused to generate the initial decoding features. :
[0031]
[0032] in express times bilinear upsampling. The initial decoding features generated in this step It contains rich multi-scale semantic information, providing high-quality starting features for subsequent progressive bidirectional iterative guided processing.
[0033] 3.2 Progressive Bidirectional Guided Iterative Processing
[0034] For hierarchical indexes , performing progressive processing from deep to shallow layers in sequence. Each processing cycle includes three collaboratively optimized sub-processes: guided multi-scale feature enhancement, bidirectional guided feature fusion, and multi-scale feature decoding and segmentation output:
[0035] 3.2.1 Guided Multi-Scale Feature Enhancement
[0036] Based on the decoding features of the current level Generate spatial guidance signals to encode shallow features Perform adaptive enhancement.
[0037] Decoding features go through , using the spatial distribution characteristics of the current layer decoding features to generate guidance weights ; This space guide signal It can highlight the spatial area related to the lesion and suppress the background interference of normal tissue.
[0038] Connect and fuse deep semantic information with shallow detail information:
[0039]
[0040] In particular, when hour, This cross-layer connection combines deep semantic information with shallow detail information to form a fusion feature containing multi-level information.
[0041] Use spatial guidance signals to differentially enhance connection features and generate enhanced coding features :
[0042]
[0043] in Represents element-wise multiplication, spatial guidance signal As spatial attention weights, it achieves differentiated enhancement of features at different spatial locations. Residual connections ensure the integrity of feature information and prevent the gradient vanishing problem in deep networks.
[0044] 3.2.2 Bidirectional guided feature fusion
[0045] The bidirectional guidance mechanism first performs adaptive gating enhancement on the enhanced encoding features. Feature importance weights are generated through a multi-head attention gating mechanism, and then the gating operation is performed:
[0046]
[0047]
[0048] in Ensure the stability of feature distribution, Transform the complex relationships between learning features, represents the activation function, The representation gating mechanism enables adaptive selection and enhancement of encoding features.
[0049] The cross-guidance mechanism is the core technology of bidirectional fusion. This mechanism uses the enhanced encoding features to guide the updating process of decoding features:
[0050]
[0051] in It is a multi-head attention mechanism, and the number of attention heads is set to 8. By taking the decoded features as , encoding features as and , realizing the dynamic query and acquisition of relevant information in the encoding features by the decoding features.
[0052] Feature fusion adopts a learning-based adaptive weighting strategy and then performs weighted fusion:
[0053]
[0054]
[0055] in It is global average pooling. Compared with fixed weight fusion, this strategy dynamically balances the contribution of encoding features and decoding features by learning adaptive weights, achieving the optimal feature combination in different regions.
[0056] Finally, the fusion features are further optimized through the channel-space dual attention mechanism:
[0057]
[0058]
[0059]
[0060] in is the global maximum pooling, Indicates passing through 7×7 convolution.
[0061] 3.2.3 Multi-scale feature decoding and segmentation output
[0062] Generative fusion features , directly input the multi-scale decoder to generate decoding features .
[0063] From the above steps, we can see that the decoding features Bidirectional guided feature fusion process acting on different levels, based on decoding features Generated spatial guidance signal It acts on the multi-scale feature enhancement process at different levels. Start by generating the final decoding features in a layer-by-layer manner . Decoding features pass Activation map generates the final segmentation result ,The segmentation results accurately identify the spatial distribution and boundary contours of lung infection lesions.
[0064] The beneficial effects of the present invention are as follows:
[0065] By leveraging the complementarity of encoding and decoding features through a bidirectional guided feature fusion mechanism, segmentation accuracy is improved by 8-15% compared to traditional one-way feature transfer methods. The hierarchical Transformer encoder effectively establishes long-range dependencies, better handling lesions of varying sizes and shapes, and improving detection sensitivity for small-scale lesions by over 20%. A multi-scale parallel decoding strategy enables precise segmentation of different lesion types, achieving segmentation accuracy exceeding 90% for various lesion types, such as ground-glass opacities and consolidation. The guided enhancement mechanism and dual attention optimization significantly improve boundary localization accuracy, increasing the boundary IoU metric by 12%, meeting clinical diagnostic requirements. The algorithm processes a single image in approximately 2 seconds on a standard GPU, meeting clinical real-time processing requirements and demonstrating good deployment feasibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a schematic diagram of the structure of the present invention;
[0067] Figure 2 To guide the multi-scale feature enhancement module;
[0068] Figure 3To guide the feature fusion module;
[0069] Figure 4 It is a multi-scale feature decoding module. DETAILED DESCRIPTION
[0070] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0071] Pulmonary infection image segmentation method based on bidirectional guided multi-scale feature decoding, such as Figure 1 As shown, the following process is included:
[0072] Example 1: Specific implementation of hierarchical Transformer feature encoding
[0073] Receive lung CT image data, perform adaptive preprocessing operations, generate standardized image data suitable for deep learning model processing, and provide high-quality input for subsequent feature extraction.
[0074] The preprocessing process first performs adaptive adjustment of the window width and window position. According to the HU (Hounsfield unit) value distribution characteristics of the lung tissue, the window width is dynamically set to 1500HU and the window position is set to -600HU to optimize the contrast between the lung tissue and the lesion area. Then, the noise removal operation is performed, and the edge-preserving bilateral filtering algorithm is used to remove the noise interference in the CT scan. The filtering parameters are Maintain the clarity of lesion boundaries while removing noise.
[0075] The pixel value normalization process maps the HU value range to The normalized formula is:
[0076]
[0077] in is the original HU value, is the normalized pixel value. Size normalization resizes all images to 512×512 pixels and uses bilinear interpolation to maintain image quality. The output of this step is the normalized image data As input data for subsequent Transformer feature encoding.
[0078] This example describes the technical implementation of Transformer feature encoding in detail. It uses a pre-trained SwinTransformer-Base network as the encoder backbone, adapting it to the medical imaging domain through transfer learning. The encoder input is a standardized 512×512×3 CT image, and outputs a four-level feature representation.
[0079] The image patch embedding stage segments the input image into 4×4 non-overlapping patches, each of which is mapped to a 96-dimensional feature vector via a linear embedding layer. Position encoding uses a learnable 2D position embedding to assign a unique position identifier to each patch position.
[0080] The first stage contains two Swin Transformer Blocks, processing a 128×128×96 feature map. The second to fourth stages implement downsampling and feature dimension expansion through patch merging operations. Patch merging merges adjacent 2×2 patches into a single patch and expands the feature dimension to twice the original through a linear layer. The number of blocks in each stage is , the feature dimension is .
[0081] Standardized image data output from step 1 based on the Swin Transformer architecture Perform hierarchical feature encoding to generate multi-scale feature representation The details are as follows:
[0082] The feature encoding process first inputs the image The image is divided into 4×4 non-overlapping blocks, and each block is mapped to The n-dimensional feature vector is input to the encoder. The encoder consists of four progressive downsampling stages, each of which consists of consecutive Swin Transformer blocks.
[0083] The first stage generates a feature map F1 with a resolution of 128×128 and a feature dimension of 96. The second stage reduces the resolution to 64×64 and increases the feature dimension to 192 by patch merging. The third stage continues to downsample to 32×32 resolution, with a feature dimension of 384 dimensions, and outputs a feature map The fourth stage generates the deepest feature map with a resolution of 16×16 and 768 dimensions. .
[0084] Each Swin Transformer block uses a shifting window mechanism, with a window size of 7×7 and a number of attention heads of {3, 6, 12, 24}. This shifting window mechanism enables cross-window information exchange through a (3, 3) shift operation, shifting the window three pixels to the lower right. The multilayer perceptron uses the GELU activation function, with the hidden layer dimension being four times the input dimension.
[0085] The multi-scale feature representation output by this step It contains hierarchical information from detailed texture to high-level semantics, with feature dimensions increasing layer by layer and resolution decreasing layer by layer, providing basic data for subsequent guided multi-scale feature enhancement.
[0086] Example 2: Technical Implementation of Guided Multi-Scale Feature Enhancement
[0087] Combine Figure 2 This embodiment describes the detailed implementation of the guidance enhancement module. This module implements iterative enhancement of features through a cyclic guidance mechanism, which is one of the core innovations of the present invention.
[0088] Based on the encoding features generated in step 2 The method uses a progressive bidirectionally guided multi-scale feature decoding strategy. Through an iterative process from deep to shallow layers, it achieves adaptive enhancement of coded features, bidirectional guided fusion, and multi-scale decoding, generating highly accurate lung infection lesion segmentation results. This strategy unifies the traditional independent enhancement, fusion, and decoding steps into a collaboratively optimized progressive processing flow.
[0089] 3.1 Deep Feature Decoding Initialization
[0090] Deepest encoding features As the starting feature of the entire progressive decoding process, cross-layer feature connection is first performed:
[0091]
[0092] in Indicates 2x bilinear upsampling, Represents channel dimension splicing. This cross-layer connection combines deep semantic information with shallow detail information to form a fusion feature containing multi-level information. The feature is then enhanced to obtain the enhanced encoding feature:
[0093]
[0094] in Represents a combination of operations, including convolution to extract spatial features, batch normalization to stabilize the training process, and linear rectification activation function.
[0095] Enhanced encoding features Directly input the multi-scale decoder to generate initial decoding features ,The decoder uses a four-way parallel processing strategy to process feature information of different resolutions. The original resolution path directly processes the feature , maintaining complete spatial detail information. The 1 / 2 resolution path is processed by adaptive average pooling and channel attention modules:
[0096]
[0097] in is the channel attention module, It is 2×2 average pooling.
[0098] 1 / 4 resolution path through:
[0099]
[0100] Generate semantically rich feature representations that incorporate more semantic information while retaining important details.
[0101] 1 / 8 resolution path through:
[0102]
[0103] Provide high-level semantic guidance information to ensure the semantic consistency of the segmentation results. for Each processing path is equipped with a channel attention module to adaptively adjust the importance of feature channels by learning the dependencies between channels.
[0104] In the multi-scale feature aggregation stage, the four-way features are adaptively fused to generate the initial decoding features. :
[0105]
[0106] in express times bilinear upsampling. The initial decoding features generated in this step It contains rich multi-scale semantic information, providing high-quality starting features for subsequent progressive bidirectional iterative guided processing.
[0107] For hierarchical indexes , performing progressive processing from deep to shallow layers in sequence. Each processing cycle includes three collaboratively optimized sub-processes: guided multi-scale feature enhancement, bidirectional guided feature fusion, and multi-scale feature decoding and segmentation output:
[0108] Based on the decoding features of the current level Generate spatial guidance signals to encode shallow features Perform adaptive enhancement.
[0109] Generate guidance weights using the spatial distribution characteristics of the current layer’s decoded features:
[0110]
[0111] The spatial guidance signal It can highlight the spatial area related to the lesion and suppress the background interference of normal tissue.
[0112] Connect and fuse deep semantic information with shallow detail information:
[0113]
[0114] In particular, when hour, This cross-layer connection combines deep semantic information with shallow detail information to form a fusion feature containing multi-level information.
[0115] Use spatial guidance signals to differentially enhance connection features and generate enhanced coding features :
[0116]
[0117] in Represents element-wise multiplication, spatial guidance signal As spatial attention weights, it achieves differentiated enhancement of features at different spatial locations. Residual connections ensure the integrity of feature information and prevent the gradient vanishing problem in deep networks.
[0118] The module receives the encoder characteristics and the guidance signal fed back by the decoder The generation process of the guidance signal adopts an adaptive threshold strategy, which dynamically adjusts the generation parameters of the saliency map according to the statistical characteristics of features at different levels.
[0119] In the process of cross-layer feature connection, in order to ensure the consistency of feature dimensions, a feature alignment strategy with channel attention weighting is adopted. Layer feature connection:
[0120]
[0121] The saliency map generation adopts a multi-scale fusion strategy to improve the guidance accuracy. , first generate three saliency maps of different scales:
[0122] Original scale:
[0123] Downsampling scale:
[0124]
[0125] Upsampling scale:
[0126] The final saliency map is generated by weighted fusion:
[0127]
[0128] The weight Obtained through learning.
[0129] The feature enhancement process adds a residual dense connection mechanism to enhance the feature propagation effect:
[0130]
[0131] This dense connection design can better propagate gradients and improve training stability.
[0132] Example 3: Innovative Implementation of Guided Feature Fusion
[0133] Combine Figure 3 This embodiment details the technical implementation of bidirectional guided fusion. This module uses an attention-guided feature interaction mechanism to achieve deep fusion of encoding features and decoding features.
[0134] Adaptive gating enhancement uses a multi-layer gating strategy to improve feature selection accuracy. The first layer of gating is based on global feature statistics:
[0135]
[0136] The second level of gating is based on local feature distribution:
[0137]
[0138] The final gating weight is:
[0139]
[0140] The cross-guidance mechanism is the core technology of bidirectional guidance. This mechanism uses the enhanced encoding features to guide the updating process of decoding features:
[0141]
[0142] in, This mechanism is called multi-head attention, which divides the feature space into multiple subspaces, and each attention head focuses on different feature aspects. and Respectively 、 , which come from the decoding features and coded features This design enables the decoded features to query relevant information in the encoded features. The number of attention heads is set to 8, with each head's dimension being 1 / 8 of the total dimension, ensuring expressiveness while controlling computational complexity. Through this cross-attention mechanism, the decoded features can capture rich semantic information in the encoded features, while the encoded features can also be adaptively adjusted based on the requirements of the decoding task.
[0143] The learning-based weighted fusion strategy uses an attention-weighted mechanism to replace simple linear combinations. Feature fusion uses a learning-based adaptive weighting strategy and then performs weighted fusion:
[0144]
[0145]
[0146] in It is global average pooling. Compared with fixed weight fusion, this strategy dynamically balances the contribution of encoding features and decoding features by learning adaptive weights, achieving the optimal feature combination in different regions.
[0147] Finally, the fusion features are further optimized through the channel-space dual attention mechanism:
[0148]
[0149]
[0150]
[0151] in is the global maximum pooling, Indicates passing through 7×7 convolution.
[0152] This serialized design can better model the dependencies between features than parallel processing.
[0153] Example 4: Parallel implementation of multi-scale feature decoding
[0154] Combine Figure 4 This embodiment illustrates the specific implementation of multi-scale decoding. This module adopts an adaptive scale selection strategy to dynamically adjust the contribution weight of each scale according to the feature content.
[0155] Generative fusion features , directly input the multi-scale decoder to generate decoding features Similarly, the original resolution path directly processes features , maintaining complete spatial detail information. The processing process of the 1 / 2 resolution path, 1 / 4 resolution path, and 1 / 8 resolution path is as follows:
[0156]
[0157]
[0158]
[0159] In the feature aggregation stage, multi-scale features are adaptively fused to generate decoding features. :
[0160]
[0161] This multi-scale aggregation mechanism can simultaneously utilize the advantages of features at different resolutions, maintaining fine boundary details while incorporating rich semantic information.
[0162] From the above steps, we can see that the decoding features Bidirectional guided feature fusion process acting on different levels, based on decoding features Generated spatial guidance signal It acts on the multi-scale feature enhancement process at different levels. Start by generating the final decoding features in a layer-by-layer manner . Decoding features pass The activation map generates the final segmentation mask ,The segmentation results accurately identify the spatial distribution and boundary contours of lung infection lesions.
[0163] Example 5: Loss Function Design and System Verification
[0164] This embodiment details the loss function design and system performance verification.
[0165] The loss function uses an adaptive weighted combination of four losses:
[0166] Improved Focal loss:
[0167]
[0168] in To balance the weights among categories, they are dynamically adjusted based on the category frequencies. is the predicted target category probability. It is a focusing parameter used to adjust the weight of difficult and easy samples.
[0169] Generalized Dice loss:
[0170]
[0171] in, is the category weight, For the The true label of each pixel, For the The prediction results of pixels, For the intersection, Indicates quantity, is the smoothing parameter.
[0172] Boundary-aware loss:
[0173]
[0174] in represents the gradient operator, is the true segmentation mask, To predict the segmentation mask, represents the square of the L2 norm, is the weight coefficient, is the binary cross entropy loss function, Represents a boundary extraction operation.
[0175] Deep Supervision Loss:
[0176]
[0177] in Indicates the The weight coefficient of the layer, Indicates the Layer-space saliency map.
[0178] The total loss is:
[0179]
[0180] Weight Determined via grid search.
[0181] The system was validated using a multi-center dataset containing 5,000 CT images of lung infections, covering four types of infections: pneumonia caused by the new coronavirus, bacterial pneumonia, viral pneumonia, and fungal pneumonia. The training set consisted of 3,500 cases, the validation set consisted of 750 cases, and the test set consisted of 750 cases. The evaluation indicators used were regional overlap accuracy (Dice, IoU) and boundary positioning accuracy (HD95, ASD). Dice measures the degree of overlap between the predicted area and the true area; IoU is the intersection over union ratio, which represents the ratio of the intersection to the union of the predicted area and the true area; HD95 is the 95% Hausdorff distance, which represents the 95% quantile of the distance between the predicted boundary and the true boundary. ASD is the average surface distance, which represents the average distance from all points on the predicted boundary to the true boundary.
[0182] Table 1 Segmentation performance comparison
[0183]
[0184] Table 2 Segmentation results of different lung infection types by the present invention
[0185]
[0186] Table 3 Analysis of the calculation efficiency of the present invention
[0187]
[0188] The experimental results, shown in Tables 1, 2, and 3, demonstrate that the proposed method significantly outperforms existing methods in key metrics such as the Dice coefficient, IoU, and boundary accuracy. The bidirectional guidance mechanism improves segmentation accuracy by 5.6% and boundary localization accuracy by 25% compared to the basic Swin-UNet. The multi-scale decoding strategy increases the sensitivity of small-scale lesions by 18%, meeting the clinical need for accurate diagnosis. The system's processing speed on a standard GPU meets real-time requirements, demonstrating good feasibility for clinical deployment.
[0189] Through systematic experimental verification, the technical advancement and practicality of the present invention in the intelligent segmentation task of lung infection are fully demonstrated, providing an effective technical solution for medical image segmentation.
Claims
1. A lung infection image segmentation method based on bidirectional guided multi-scale feature decoding, characterized by: The following steps are involved: Step 1: Receive lung CT image data and perform preprocessing to generate standardized image data; Step 2: Based on the Swin Transformer architecture, hierarchical feature encoding is performed on the standardized image data to generate multi-scale feature representation; Step 3: Based on multi-scale feature representation, a progressive bidirectional guided multi-scale feature decoding strategy is adopted to achieve adaptive enhancement of coding features, bidirectional guided fusion and multi-scale decoding to generate the lung infection lesion segmentation result map.
2. The lung infection image segmentation method based on bidirectional guided multi-scale feature decoding according to claim 1 is characterized in that: The preprocessing is specifically as follows: first, adaptively adjust the window width and window position to optimize the contrast between lung tissue and lesion area; then perform noise removal operation and use edge-preserving bilateral filtering algorithm to remove noise interference in CT scan; finally, perform pixel value normalization and size standardization. The pixel value normalization process maps the HU value range to Interval, size normalization adjusts all images to uniform pixels, and bilinear interpolation is used to maintain image quality.
3. The lung infection image segmentation method based on bidirectional guided multi-scale feature decoding according to claim 1, characterized in that: The step 2 is specifically implemented as follows: inputting standardized image data The image is divided into non-overlapping blocks, and each block is mapped to a feature vector through a linear embedding layer and input into the encoder; the encoder consists of four progressive downsampling stages, each of which consists of a series of Swin Transformer blocks, generating feature maps of four different resolutions to obtain a multi-scale feature representation. , the feature dimension increases layer by layer, and the resolution decreases layer by layer.
4. The lung infection image segmentation method based on bidirectional guided multi-scale feature decoding according to claim 1, characterized in that: The specific implementation process of step three is as follows: Step 3.1 Deep feature decoding initialization: feature As the starting feature of the entire progressive decoding process, feature enhancement is performed to obtain enhanced coding features, and the initial decoding features are generated through the multi-scale decoder , Step 3.2 Based on initial decoding features , and executes the progressive bidirectional guided iterative processing from deep to shallow layers in sequence. Each processing cycle includes three collaborative optimization sub-processes: guided multi-scale feature enhancement, bidirectional guided feature fusion, and multi-scale feature decoding and segmentation output to generate the final decoding feature. , decoding features pass Activation map generates the final segmentation result ,The segmentation results identify the spatial distribution and boundary contours of the lung infection lesions.
5. The lung infection image segmentation method based on bidirectional guided multi-scale feature decoding according to claim 4 is characterized in that: Generating initial decoding features The specific process is as follows: The features As the starting feature of the entire progressive decoding process, cross-layer feature connection is first performed to connect the features With 2x bilinear upsampling Splicing is performed to obtain fusion features containing multi-level information ; Then the feature go through Combining operations and features Perform residual connection to complete feature enhancement and obtain enhanced coding features ; described Combination operations, including cascaded convolution to extract spatial features, batch normalization, and linear rectification activation functions; Enhanced encoding features Directly input the multi-scale decoder to generate initial decoding features ,The multi-scale decoder adopts a four-way parallel processing strategy to process feature information at different resolutions, as follows: The original resolution path processes features directly ; The 1 / 2 resolution path will be After adaptive average pooling and channel attention module Processing, obtaining features ; 1 / 4 resolution path through: After 2 times downsampling and Add element by element, then pass through adaptive average pooling and channel attention module Processing, obtaining features ; 1 / 8 resolution path through: After 2 times downsampling and Add element by element, then pass through adaptive average pooling and channel attention module Processing, obtaining features ; The multi-scale feature aggregation stage combines the four-way features 、 、 、 After the corresponding bilinear upsampling, the adaptive fusion is completed by element-by-element addition to generate the initial decoding features. .
6. The lung infection image segmentation method based on bidirectional guided multi-scale feature decoding according to claim 5, characterized in that: The progressive bidirectional guided iterative process is specifically implemented as follows: Step 3.2.1 Guided multi-scale feature enhancement: For hierarchical index , based on the decoding features of the current level Generate spatial guidance signals to encode features Perform adaptive enhancement to generate enhanced coding features ; 3.2.2 Bidirectional guided feature fusion: First, enhance the encoding features Adaptive gate enhancement is performed, and then the enhanced encoding features are used to guide the updating process of the decoding features through the cross-guidance mechanism, and a learning-based adaptive weighting strategy is used for feature fusion. Finally, the fused features are optimized through the channel-space dual attention mechanism to obtain the optimized features. ; 3.2.3 Multi-scale feature decoding and segmentation output: Feature-based , directly input the multi-scale decoder to generate decoding features ; Decoding features Bidirectional guided feature fusion process acting on different levels, based on decoding features Generated spatial guidance signal It acts on the guided multi-scale feature enhancement process at different levels; From the initial decoded features Start by generating the final decoding features in a layer-by-layer manner ; Decoding features pass Activation map generates the final segmentation result ,The segmentation results identify the spatial distribution and boundary contours of the lung infection lesions.
7. The lung infection image segmentation method based on bidirectional guided multi-scale feature decoding according to claim 6, characterized in that: The adaptive enhancement described in step 3.2.1 is specifically implemented as follows: Using the spatial distribution characteristics of the current layer decoding features, go through Combine operations to generate guidance weights ; Through Upsampling and Splicing, connecting and fusing deep semantic information with shallow detail information to obtain features , and when hour, ; Will go through After combining the operation Perform element-wise multiplication and then connect to the input through the residual connection Perform element-by-element addition, use the spatial guidance signal to differentially enhance the connection features, and generate enhanced coding features .
8. The lung infection image segmentation method based on bidirectional guided multi-scale feature decoding according to claim 7, characterized in that: The bidirectional guided feature fusion described in step 3.2.2 is specifically implemented as follows: The bidirectional guidance mechanism first enhances the coding features Perform adaptive gating enhancement: Encoding features will be enhanced Through the normalization layer, linear layer and Implement the adaptive selection and enhancement of encoding features by the gating mechanism, and then and Perform element-by-element multiplication and finally enhance the encoding features with the input through residual connection , adding them together gives ; The enhanced encoding features are used to guide the updating process of decoding features through the cross-guidance mechanism. Calculated by two linear layers respectively, the key Sum :feature Get the query after the linear layer , output by the multi-head attention mechanism ; Adopting learning-based adaptive weighting strategy for feature fusion: and After two global average poolings, they are added together and then passed through The function gets the output score ;at last and Multiply element-wise and add and The result of element-by-element multiplication is the fusion feature ; Finally, the fusion feature is optimized through the channel-space dual attention mechanism: fusion feature After the linear layer and After the activation function, it is sequentially transformed by linear mapping and Function output characteristics ; After global maximum pooling and global average pooling, they are spliced and then convolved. Function output characteristics ;at last 、 、 The three are multiplied element by element and added from the residual connection , and get the optimized features .
Citation Information
Patent Citations
Lung CT image segmentation method based on mixed Swin Transform U-Net
CN117274147A
Insulator fault detection method based on multi-scale expansion and perception codec
CN118657716A