Pneumonia segmentation method based on multi-connection axial graph convolutional network
Through multi-connected axial graph convolution network (MAGNN) to capture the long-range dependence and spatial relationship in pneumonia segmentation, the problem of insufficient segmentation accuracy in the prior art is solved, and higher pneumonia segmentation accuracy and computing efficiency are achieved.
Patent Information
- Application Number
- CN202510320745.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The existing convolutional neural networks are difficult to fully capture the long-range dependence and spatial relationships of lung infection areas in pneumonia segmentation, affecting the segmentation accuracy.
Multi-connected axial graph convolution network (MAGNN) is used to generate the initial region of interest mask through threshold segmentation and morphological processing, and a multi-connected axial graph convolution layer and a multi-region compression extraction layer are used to capture local and global features, and encode and decode processing are performed to improve segmentation accuracy.
Effectively capture spatial and semantic information of images, improves the accuracy of pneumonia segmentation, reduces the amount of calculation, and shows excellent segmentation performance on public data sets.
Smart Images

Figure CN120298422A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and particularly relates to a pneumonia segmentation method based on a multi-connected axial graph convolutional network. Background Art
[0002] In recent years, pneumonia segmentation based on computed tomography (CT) images has played an important role in clinical diagnosis. Traditional convolutional neural networks mainly rely on local feature extraction when dealing with image segmentation problems, but for the complex boundaries and asymmetric structures in pneumonia segmentation, the performance of convolutional neural networks is limited.
[0003] Therefore, existing segmentation methods are difficult to fully capture the long-range dependencies and spatial relationships in the lung infection area, thus affecting the segmentation accuracy. Summary of the Invention
[0004] The purpose of the present invention is to provide a pneumonia segmentation method based on a multi-connected axial graph convolutional network to solve the technical problem that it is difficult to fully capture the long-range dependencies and spatial relationships in the lung infection area in the prior art, thus affecting the segmentation accuracy.
[0005] To solve the above technical problem, the present invention specifically provides the following technical solutions:
[0006] A pneumonia segmentation method based on a multi-connected axial graph convolutional network, comprising the following steps:
[0007] Obtain a chest CT image;
[0008] Adopt threshold segmentation and morphological processing in the chest CT image to generate an initial region of interest mask for the lungs;
[0009] Input the chest CT image and the initial region of interest mask for the lungs into a multi-connected axial graph convolutional network MAGNN, extract the feature map for pneumonia infection area segmentation in the chest CT image and encode it;
[0010] Decode and upsample the encoded feature map to obtain a pneumonia infection area segmentation map in the chest CT image.
[0011] As a preferred solution of the present invention, the method for generating the initial region of interest mask for the lungs includes:
[0012] Perform binarization processing on the chest CT image using the Otsu method, retain the left lung region and the right lung region in the chest CT image through connected component analysis, and apply a morphological operation Morph() to ensure the alignment of the mask boundaries to obtain the initial region of interest mask for the lungs;
[0013] The generation expression of the initial region of interest mask for the lungs is:
[0014] Seg = MorPh(OtsuThreshold(I));
[0015] Wherein, Seg is the initial region of interest mask of the lungs, I is the chest CT image, OtsuThreshold is the identifier of the Otsu method function, and Morph is the identifier of the Morph() function of the morphological operation.
[0016] As a preferred embodiment of the present invention, the structure of the multi-connected axial graph convolutional network MAGNN consists of multiple layers of encoders, and each layer of encoder includes a multi-connected axial graph convolutional layer MAA and a multi-region compression extraction layer MSE;
[0017] Among them, the multi-connected axial graph convolutional layer MAA is used to separate the feature information of the width axis and the height axis, and determine the dependence relationship between the local feature and the global feature by calculating the maximum difference between each node feature and its neighborhood node feature;
[0018] The multi-region compression extraction layer MSE is used to adaptively compress and extract the features of different channels and regions of interest;
[0019] The operation expression of the multi-connected axial graph convolutional network MAGNN is:
[0020] Z (l) = MAA(MSE(Z (l―1) , Seg (l―1) ));
[0021] Wherein, Z (l) is the feature map output by the l-th layer encoder, Z (l―1) is the feature map output by the (l-1)-th layer encoder, Seg (l―1) is the region of interest mask corresponding to the (l-1)-th layer encoder, MSE is the identifier of the multi-region compression extraction layer MSE, and MAA is the identifier of the multi-connected axial graph convolutional layer MAA.
[0022] As a preferred embodiment of the present invention, the decoding and upsampling processing are implemented by multiple layers of decoders, and the decoding and upsampling processing expression is:
[0023]
[0024] Wherein, is the feature map output by the l-th layer decoder, is the feature map output by the (l + 1)-th layer decoder, Z (l) is the feature map output by the l-th layer encoder, Seg (l) is the region of interest mask corresponding to the l-th layer encoder, and Decoder (l)is the identifier of the l-th layer encoder.
[0025] As a preferred embodiment of the present invention, the method for adaptively compressing and extracting features of different channels and regions of interest by the multi-region compression and extraction layer MSE includes:
[0026] The first step, channel compression and activation: By extracting the spatial information of the input feature map of the multi-region compression and extraction layer MSE in each channel and adjusting the activation value of the channel through an activation function to dynamically enhance the features related to a specific channel;
[0027] Among them, the channel compression operation F c-squeeze Extracts the spatial information from the input feature map;
[0028] cs = F c―squeeze (Z);
[0029] In the formula, Z is the input feature map of the multi-region compression and extraction layer MSE, cs is the descriptor of the feature map after channel compression, and F c-squeeze Is the channel compression operation identifier;
[0030] The channel activation operation F c-excite Uses the trainable weight W in the channel activation operation cs To perform activation processing on cs to obtain the channel scaling coefficient;
[0031] c = F c―excite (cs, W cs );
[0032] In the formula, ce is the channel scaling coefficient, F c-excite Is the channel activation operation identifier, and W cs Is the trainable weight in the channel activation operation;
[0033] The second step, region compression and activation: Perform global average pooling on the input feature map of the multi-region compression and extraction layer MSE, extract the global spatial information of each region of interest, and enhance the feature representation of each region through an activation function;
[0034] Among them, the region compression operation F r-squeeze Performs global average pooling on each region of interest region m Of the input feature map to extract the global information at the region level;
[0035]
[0036] In the formula, Z is the input feature map, |region m | is the number of spatial points in the m-th region of interest region m , Denote the feature value of the k-th channel of the input feature map Z at the spatial position (i, j), rs m is the descriptor of the feature map after region compression, F r-squeeze is the region compression operation identifier;
[0037] Region activation operation F r-excite Utilize the trainable weight W in the region activation operation rs for rs m to perform Activation point Processing , and obtain the region scaling coefficient;
[0038]
[0039] In the formula, re is the region scaling coefficient, σ is the Sigmoid activation function, δ is the ReLU activation function,, F r-excite is the region activation operation identifier, W rs is the trainable weight in the region activation operation, including and
[0040] Step 3, recalibration operation: Scale the channels and regions of the input feature map according to the channel scaling coefficient and the region scaling coefficient respectively;
[0041] Among them, based on the channel scaling coefficient ce and the region scaling coefficient re, use the recalibration function F scale to adjust the input feature map;
[0042] F scale (Z, re, ce) = ce ⊙ ∑ m (re m ·Z·1((i, j) ∈ region m ));
[0043] In the formula, F scale (Z, re, ce) is the recalibration function, which uses the obtained re and ce vectors to recalibrate the feature Z, re m is the region scaling coefficient in region m , 1((i, j) ∈ region m ) is the indicator function, and ⊙ represents element-wise multiplication.
[0044] As a preferred embodiment of the present invention, the method for the multi-connected axial graph convolutional layer MAA to separate the feature information of the width axis and the height axis includes:
[0045] First step, separation processing of the width axis and height axis: Axially segment the input feature map of the multi-connected axial graph convolutional layer MAA along the width axis and height axis, and separately extract the feature components in the width direction and the feature components in the height direction;
[0046] Second step, capturing of local features and global features: Apply the attention mechanism to both the feature components of the width axis and the height axis for dependency modeling. Capture local features through the attention mechanism to emphasize the detailed correlation in the short-distance space, and at the same time capture global features to model the context relationship in the long-distance space;
[0047] Third step, fusion of width axis and height axis information: Recombine the features processed on the width axis and the height axis, and form a global feature representation through weighted summation or a specific fusion mechanism to ensure that the information in the spatial dimension can be synergistically optimized and improve the segmentation performance.
[0048] As a preferred solution of the present invention, the method for determining the dependency relationship between local features and global features by the multi-connected axial graph convolutional layer MAA by calculating the maximum difference between each node feature and its neighboring node features includes:
[0049] Calculate the maximum difference between each node feature and its neighboring node features to obtain the dependency relationship between local features:
[0050] The expression for calculating the maximum difference between the feature at each node position (i, j) of the input feature map of the multi-connected axial graph convolutional layer MAA and the feature at the neighboring node position is:
[0051]
[0052] In the formula, Δz p,(i,h) is the maximum difference between the feature at the node position (i, j) and the feature at the neighboring node position, N(i, j) is the neighborhood of the node position (i, j), and z (i,j) is the feature value at the node position (i, j), and z (k,l) is the feature of the neighboring node position (k, l).
[0053] As a preferred solution of the present invention, the loss function of the multi-connected axial graph convolutional network MAGNN is a weighted combination of Dice Loss and Cross-Entropy Loss,
[0054] The evaluation metrics of the multi-connected axial graph convolutional network MAGNN include:
[0055] Dice Coefficient: Used to evaluate the overlap degree between the segmentation result and the true label, and the numerical range is from 0 to 1. The larger the value, the closer the segmentation result is to the true region.
[0056] Intersection over Union (IoU): That is, the ratio of the intersection of the segmented region and the ground truth region to the union, which is used to quantify the accuracy of the segmentation result. The larger the value, the better.
[0057] Hausdorff Distance 95th Percentile (HD95): Measures the maximum distance between the segmentation boundary and the ground truth boundary. The 95th percentile is taken to reduce the influence of outliers. The smaller the value, the better, indicating more precise boundary alignment.
[0058] The present invention has the following beneficial effects compared with the prior art:
[0059] By introducing a multi-connected axial attention mechanism and a multi-connected compression extraction module, the present invention effectively captures the spatial and semantic information of the image in the graph convolutional network, and combines efficient boundary alignment to improve the accuracy of pneumonia segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained based on the provided drawings.
[0061] Figure 1 It is a flowchart of the pneumonia segmentation method based on the multi-connected axial graph convolutional network provided by the embodiment of the present invention;
[0062] Figure 2 It is a mask of the region of interest in the lung provided by the embodiment of the present invention;
[0063] Figure 3 It is a flowchart of the MSE processing of the multi-region compression extraction layer provided by the embodiment of the present invention;
[0064] Figure 4 It is a flowchart of the MAA processing of the multi-connected axial graph convolutional layer provided by the embodiment of the present invention;
[0065] Figure 5 It is a structural diagram of the pneumonia segmentation model based on the multi-connected axial graph convolutional network provided by the embodiment of the present invention;
[0066] Figure 6 It is an evaluation diagram of the model effect on the MosMedData dataset provided by the embodiment of the present invention;
[0067] Figure 7This is the model effect evaluation diagram on the COVID-19 CT Lung and Infection Segmentation dataset provided by the embodiments of the present invention. Detailed implementation manners
[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0069] As Figure 1 and Figure 5 shown, the present invention provides a pneumonia segmentation method based on a multi-connected axial graph convolutional network, including the following steps:
[0070] Obtain chest CT images;
[0071] Adopt threshold segmentation and morphological processing in the chest CT images to generate an initial region of interest mask for the lungs;
[0072] Input the chest CT images and the initial region of interest mask for the lungs into the multi-connected axial graph convolutional network MAGNN to extract the feature maps for pneumonia infection region segmentation in the chest CT images and encode them;
[0073] Decode and upsample the encoded feature maps to obtain the pneumonia infection region segmentation map in the chest CT images.
[0074] By adding a multi-connected axial attention mechanism (i.e., MAA) to the graph convolutional network (i.e., MAGNN), the present invention realizes the long-range dependence modeling of the infection region, combines the design of the multi-region layer (i.e., MSE) to achieve channel compression and spatial compression, captures global and local features, improves the feature expression ability, and effectively reduces the computational complexity.
[0075] Specifically, in the present invention, the multi-region compression extraction layer MSE compresses and extracts the features of different channels and regions of interest to ensure that the model can fuse rich context information at different scales and improve the segmentation results.
[0076] In the present invention, the multi-connected axial graph convolutional layer MAA captures local and global dependence relationships by respectively performing feature modeling on the width axis and height axis of the image. This way of separating axis information processing can effectively reduce the computational complexity while ensuring that the features in each direction can be fully transmitted and captured.
[0077] The method for generating the initial region of interest mask for the lungs includes:
[0078] The Otsu method is used to binarize the chest CT image, and the left and right lung regions in the chest CT image are retained through connected component analysis. In addition, morphological operation Morph() is applied to ensure the boundary alignment of the mask, and the initial region of interest mask of the lungs is obtained;
[0079] The generation expression of the initial region of interest mask of the lungs is:
[0080] Seg = MorPh(OtsuThreshold(I));
[0081] In the formula, Seg is the initial region of interest mask of the lungs, I is the chest CT image, OtsuThreshold is the identifier of the Otsu method function, and Morph is the identifier of the Morph() function of the morphological operation.
[0082] The process of threshold segmentation and morphological processing:
[0083] Using the pixel value range of the CT image, the lung region (left and right lung regions) is extracted based on the set threshold.
[0084] Noise and small artifacts are removed through morphological operations (such as dilation and erosion) to generate the initial mask Seg (as Figure 2 shown).
[0085] The structure of the multi-connected axial graph convolutional network MAGNN consists of multiple layers of encoders, and each layer of encoder includes a multi-connected axial graph convolutional layer MAA and a multi-region compression extraction layer MSE;
[0086] Among them, the multi-connected axial graph convolutional layer MAA is used to separate the feature information of the width axis and the height axis, and determine the dependence relationship between the local feature and the global feature by calculating the maximum difference between each node feature and its neighboring node features;
[0087] The multi-region compression extraction layer MSE is used to adaptively compress and extract the features of different channels and regions of interest;
[0088] The operation expression of the multi-connected axial graph convolutional network MAGNN is:
[0089] Z (l) = MAA(MSE(Z (l―1) ,Seg (l―1) ));
[0090] In the formula, Z (l) is the feature map output by the l-th layer encoder, Z (l―1) is the feature map output by the (l - 1)-th layer encoder, and Seg (l―1)is the region of interest mask corresponding to the (l-1)-th layer encoder, MSE is the identifier of the multi-region compression extraction layer MSE, and MAA is the identifier of the multi-connected axial graph convolution layer MAA.
[0091] Seg (l―1) is the region of interest mask corresponding to the (l-1)-th layer encoder, and is obtained by performing an interpolation operation on Seg (l―2) and reducing the size by half.
[0092] The decoding and upsampling processes are implemented by a multi-layer decoder, and the decoding and upsampling process expressions are:
[0093]
[0094] In the formula, is the feature map output by the l-th layer decoder, is the feature map output by the (l + 1)-th layer decoder, Z (l) is the feature map output by the l-th layer encoder, Seg (l) is the region of interest mask corresponding to the l-th layer encoder, Decoder (l) is the identifier of the l-th layer encoder.
[0095] The Squeeze-and-Excitation (SE) layer adaptively adjusts the importance of each feature map, mainly performs feature recalibration at the channel level, and improves the feature representation ability. However, in medical images, especially the distribution of lung lesions often spans multiple regions of interest (ROIs), and these regions have different features. The SE layer adjusts at the channel level and cannot effectively capture the spatial relationship and discontinuous feature distribution between these regions. The present invention uses a multi-region compression extraction layer MSE to expand the function of the SE layer, not only compresses and recalibrates the channels, but also performs feature extraction at the region level. The MSE layer combines channel compression and region compression operations as follows:
[0096] As Figure 3 shown, the method for the multi-region compression extraction layer MSE to adaptively compress and extract the features of different channels and regions of interest includes:
[0097] The first step, channel compression and activation: by extracting the spatial information of the input feature map of the multi-region compression extraction layer MSE in each channel, and adjusting the activation value of the channel through an activation function to dynamically enhance the features related to a specific channel;
[0098] Among them, the channel compression operation F c-squeeze extracts the spatial information from the input feature map;
[0099] cs = F c―squeeze (Z);
[0100] where \(Z\) is the input feature map of the multi-region compression extraction layer MSE, \(cs\) is the descriptor of the feature map after channel compression, and \(F\) c-squeeze is the channel compression operation identifier;
[0101] The channel activation operation \(F\) c-excite uses the trainable weight \(W\) in the channel activation operation cs to activate \(cs\) to obtain the channel scaling coefficient;
[0102] \(ce = F\) c―excite (\(cs, W\) cs );
[0103] where \(ce\) is the channel scaling coefficient, \(F\) c-excite is the channel activation operation identifier, and \(W\) cs is the trainable weight in the channel activation operation;
[0104] Second, region compression and activation: perform global average pooling on the input feature map of the multi-region compression extraction layer MSE, extract the global spatial information of each region of interest, and enhance the feature representation of each region through an activation function;
[0105] Among them, the region compression operation \(F\) r-squeeze performs global average pooling on each region of interest \(region\) of the input feature map m to extract the global information at the region level;
[0106]
[0107] where \(Z\) is the input feature map, \(|region|\) m is the number of spatial points in the \(m\)-th region of interest \(region\) m , represents the feature value of the \(k\)-th channel of the input feature map \(Z\) at the spatial position \((i, j)\), and \(rs\) m is the descriptor of the feature map after region compression, and \(F\) r-squeeze is the region compression operation identifier;
[0108] The region activation operation \(F\) r-excite uses the trainable weight \(W\) in the region activation operation rs to perform m on \(rs\) Activation point Processing to obtain the region scaling coefficient;
[0109]
[0110] where \(re\) is the region scaling coefficient, \(\sigma\) is the Sigmoid activation function, \(\delta\) is the ReLU activation function, and \(F\)r-excite is the regional activation operation identifier, W rs is the trainable weight in the regional activation operation, including and
[0111] Step 3, recalibration operation: According to the channel scaling factor and the regional scaling factor, scale the channels and regions of the input feature map respectively to ensure that the model can adaptively strengthen the features of key regions and channels. This extraction and optimization adjustment of features helps to significantly improve the segmentation effect, thereby enhancing the model performance;
[0112] Among them, based on the channel scaling factor ce and the regional scaling factor re, use the recalibration function F scale to adjust the input feature map;
[0113] F scale (Z, re, ce) = ce ⊙ ∑ m (re m ·Z·1((i, j) ∈ region m ));
[0114] In the formula, F scale (Z, re, ce) is the recalibration function, which uses the obtained re and ce vectors to recalibrate the feature Z. re m is the regional scaling factor in region m , 1((i, j) ∈ region m ) is the indicator function to ensure that the regional scaling factor re is only applied to the regions in region m . ⊙ represents element-wise multiplication and applies the scaling of the channel layer. m The multi-connected axial graph convolutional layer MAA captures local and global dependencies by respectively performing feature modeling on the width axis and height axis of the image. This way of separating-axis information processing can effectively reduce the computational complexity while ensuring that the features in each direction can be fully transmitted and captured, as follows:
[0115] As
[0116] shown, the method for the multi-connected axial graph convolutional layer MAA to separate the feature information of the width axis and height axis includes: Figure 4
[0117] Step 1, separation processing of the width axis and height axis: Axially segment the input feature map of the multi-connected axial graph convolutional layer MAA along the width axis and height axis, and respectively extract the feature components in the width direction and the feature components in the height direction;
[0118]
[0118] Step 2, Capture of local and global features: The attention mechanism is adopted for dependency modeling of the feature components on both the width axis and the height axis. Through the attention mechanism, local features are captured to emphasize the detailed correlation in the short-distance space, and at the same time, global features are captured to model the context relationship in the long-distance space;
[0119] Step 3, Fusion of width-axis and height-axis information: The features processed on the width axis and the height axis are recombined, and a global feature representation is formed through weighted summation or a specific fusion mechanism to ensure that the information in the spatial dimension can be synergistically optimized to improve the segmentation performance.
[0120] In order to capture the local and global feature dependencies in the present invention, the MAA layer not only calculates the features of a single node, but also calculates the relationship between each node and its neighboring nodes. By calculating the maximum difference between the node features and the features of its neighboring nodes, the model can understand the particularity of the node in the local area and its relationship with other nodes.
[0121] The method for the multi-connected axial graph convolutional layer MAA to determine the dependency relationship between local features and global features by calculating the maximum difference between the features of each node and the features of its neighboring nodes includes:
[0122] Calculate the maximum difference between the features of each node and the features of its neighboring nodes to obtain the dependency relationship between local features:
[0123] The expression for calculating the maximum difference between the feature at each node position (i, j) of the input feature map of the multi-connected axial graph convolutional layer MAA and the feature at the neighboring node position is:
[0124]
[0125] In the formula, Δz p,(i,j) is the maximum difference between the feature at the node position (i, j) and the feature at the neighboring node position, N(i, j) is the neighborhood of the node position (i, j), and z (i,j) is the feature value at the node position (i, j), and z (k,l) is the feature of the neighboring node position (k, l).
[0126] The capture of local and global feature dependencies in the present invention means that the model can distinguish and effectively model the features in different regions of the image. Specifically:
[0127] Local feature dependency: Represents the relationship between a node and its neighboring nodes within its local range. For example, the maximum difference, the relative distance between nodes, or the similarity within the neighborhood, etc. This part of the features is very important for dealing with details and local changes.
[0128] Global feature dependencies: refer to the relationships between nodes and the entire image or other nodes within a larger scope. This part of the features is crucial for capturing global context information and global consistency.
[0129] The local features obtained in the present invention are crucial for detailed regions (such as the boundaries of lung lesions). By calculating the maximum difference between each node and its neighboring nodes, MAA can more accurately capture the boundary information of the lesion area, thereby enhancing the model's segmentation ability in detailed regions (especially complex lung lesions).
[0130] The global features obtained in the present invention help the model understand the relationships between regions in the image. In medical images, the distribution of lesions is usually greatly affected by the global context (for example, the spread of lung inflammation), so global information is very important for improving the overall segmentation effect.
[0131] Therefore, by adding a multi-connected axial attention mechanism to the graph convolutional network, the present invention realizes long-range dependence modeling for the infected area, enhances the model's segmentation ability in detailed regions (especially complex lung lesions), and improves the overall segmentation effect.
[0132] The loss function of the multi-connected axial graph convolutional network MAGNN is a weighted combination of the Dice loss and the cross-entropy loss.
[0133] The evaluation metrics of the multi-connected axial graph convolutional network MAGNN include:
[0134] Dice Coefficient: used to evaluate the overlap degree between the segmentation result and the ground truth label, with a numerical range of 0 to 1. The larger the value, the closer the segmentation result is to the real area.
[0135] Intersection over Union (IoU): that is, the ratio of the intersection of the segmented area and the ground truth area to the union, used to quantify the accuracy of the segmentation result. The larger the value, the better.
[0136] Hausdorff Distance 95th Percentile (HD95): measures the maximum distance between the segmentation boundary and the ground truth boundary. The 95th percentile is taken to reduce the influence of outliers. The smaller the value, the better, indicating more accurate boundary alignment.
[0137] Through experiments on multiple public datasets, the present invention proves that the MAGNN model achieves state-of-the-art performance in the pneumonia CT image segmentation task and significantly reduces the number of parameters, as follows:
[0138] As Figure 6 shown, in terms of the performance on the MosMedData dataset, the results indicate that the MAGNN model exhibits the highest segmentation performance, with a Dice coefficient of 0.8611, a Jaccard index of 0.7591, and the lowest HD95 value of 1.5986. Compared with traditional models (such as Unet and Unet++), MAGNN improves by approximately 5% in terms of the Dice coefficient while only requiring 3.5 million parameters. When compared with high-capacity Transformer-based models (such as TransUNet (109.5 million parameters) and H2Former (33.7 million parameters)), MAGNN achieves higher accuracy with fewer parameters.
[0139] It is worth noting that the performance of Transformer-based models (such as TransUNet and H2Former) on the MMDD dataset is lower than expected, which may be due to the limitation of the small dataset size. Transformer models generally require large datasets to effectively learn complex long-range dependencies and are difficult to fully exploit their advantages in the case of limited data. In contrast, the design of MAGNN is optimized for small datasets and can achieve a good balance between parameter efficiency and segmentation accuracy.
[0140] Furthermore, although the UNext model only has 1.5 million parameters and its lightweight convolutional architecture demonstrates excellent efficiency, its simplified design limits the ability to capture complex long-range dependencies. Therefore, in terms of accurately segmenting complex anatomical structures, it performs less well than MAGNN. The above results indicate that MAGNN achieves an excellent trade-off between model complexity and segmentation performance and is suitable for application in medical image segmentation tasks with limited data.
[0141] As Figure 7 shown, in terms of the performance on the COVID-19 CT Lung and Infection Segmentation dataset, the results indicate that compared with the MosMedData dataset, CLISD provides a larger amount of data and high-quality annotations, which improves the performance of most models. The proposed MAGNN model performs best on the CLISD dataset, with a Dice coefficient of 0.9022, a Jaccard index of 0.8219, and an HD95 value of only 1.0.
[0142] It should be noted that although Attention Unet introduces the attention mechanism, its performance on the CLISD dataset is significantly worse, with a Dice coefficient of only 0.8181 and an HD95 value as high as 2.9984. This performance difference is mainly attributed to the implementation method of its attention mechanism. Attention Unet adopts localized attention. Although this method can enhance the features of specific regions, its limitation in the spatial range reduces the overall performance of the model. In the CLISD dataset, capturing extensive spatial relationships is crucial for achieving accurate segmentation.
[0143] In contrast, MAGNN adopts axial attention operations, which can simultaneously model local and global dependencies in multiple regions of interest (ROIs). This mechanism enables MAGNN to better capture the complex spatial relationships in the pulmonary infection regions, thus achieving excellent segmentation performance on the CLISD dataset. These results indicate that MAGNN is superior to traditional attention mechanism models in terms of both accuracy and generalization ability, especially outstanding in medical image segmentation tasks that require global dependencies.
[0144] The present invention effectively captures the spatial and semantic information of images in the graph convolutional network by introducing a multi-connection axial attention mechanism and a multi-connection compression extraction module, and combines efficient boundary alignment to improve the accuracy of pneumonia segmentation.
[0145] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.
Claims
1. A pneumonia segmentation method based on a multi-connected axial graph convolutional network, characterized in that Including the following steps: Obtain a chest CT image; Adopt threshold segmentation and morphological processing in the chest CT image to generate an initial region of interest mask for the lungs; Input the chest CT image and the initial region of interest mask for the lungs into a multi-connected axial graph convolutional network MAGNN, extract the feature map for pneumonia infection region segmentation in the chest CT image and encode it; Decode and upsample the encoded feature map to obtain a pneumonia infection region segmentation map in the chest CT image.
2. The pneumonia segmentation method based on a multi-connected axial graph convolutional network according to claim 1, wherein: The method for generating the initial region of interest mask for the lungs includes: Perform binarization processing on the chest CT image using the Otsu method, retain the left lung region and the right lung region in the chest CT image through connected component analysis, and apply a morphological operation Morph() to ensure the alignment of the mask boundary to obtain the initial region of interest mask for the lungs; The generation expression of the initial region of interest mask for the lungs is: Seg = MorPh(OtsuThreshold(I)); In the formula, Seg is the initial region of interest mask for the lungs, I is the chest CT image, OtsuThreshold is the identifier of the Otsu method function, and Morph is the identifier of the Morph() function of the morphological operation.
3. The pneumonia segmentation method based on a multi-connected axial graph convolutional network according to claim 2, wherein: The structure of the multi-connected axial graph convolutional network MAGNN consists of multiple layers of encoders, and each layer of encoder includes a multi-connected axial graph convolutional layer MAA and a multi-region compression and extraction layer MSE; Among them, the multi-connected axial graph convolutional layer MAA is used to separate the feature information of the width axis and the height axis, and determine the dependence relationship between the local feature and the global feature by calculating the maximum difference between each node feature and its neighboring node features; The multi-region compression and extraction layer MSE is used to adaptively compress and extract the features of different channels and regions of interest; The operation expression of the multi-connected axial graph convolutional network MAGNN is: Z (l) = MAA(MSE(Z (l―1) , Seg (l―1) )); where Z (l) is the feature map output by the l-th layer encoder, and Z (l―1) is the feature map output by the (l-1)-th layer encoder. Seg (l―1) is the region of interest mask corresponding to the (l-1)-th layer encoder. MSE is the identifier of the multi-region compression extraction layer MSE, and MAA is the identifier of the multi-connected axial graph convolution layer MAA.
4. The pneumonia segmentation method based on a multi-connected axial graph convolutional network according to claim 3, characterized in that: The decoding and upsampling processing is implemented by multiple layers of decoders, and the decoding and upsampling processing expression is: In the formula, is the feature map output by the l-th layer decoder, is the feature map output by the (l + 1)-th layer decoder, Z (l) is the feature map output by the l-th layer encoder, Seg (l) is the region of interest mask corresponding to the l-th layer encoder, Decoder (l) is the identifier of the l-th layer encoder.
5. A pneumonia segmentation method based on a multi-connected axial graph convolutional network according to claim 4, characterized in that: The method for the multi-region compression and extraction layer MSE to adaptively compress and extract the features of different channels and regions of interest includes: The first step, channel compression and activation: Extract the spatial information of the input feature map of the multi-region compression and extraction layer MSE in each channel, and adjust the activation value of the channel through an activation function to dynamically enhance the features related to a specific channel; Among them, the channel compression operation F c-squeeze Extracts spatial information from the input feature map; cs = F c―squeeze (Z); Wherein, Z is the input feature map of the MSE of the multi-region compression extraction layer, cs is the descriptor of the feature map after channel compression, and F c-squeeze is the channel compression operation identifier; Channel activation operation F c-excite Using the trainable weight W in the channel activation operation cs Perform activation processing on cs to obtain the channel scaling factor; ce = F c―excite (cs, W cs ) where ce is the channel scaling factor, F c-excite is the channel activation operation identifier, and W cs is the trainable weight in the channel activation operation; The second step, region compression and activation: Perform global average pooling on the input feature map of the multi-region compression and extraction layer MSE, extract the global spatial information of each region of interest, and enhance the feature representation of each region through an activation function; Among them, the region compression operation F r-squeeze performs global average pooling on each region of interest of the input feature map m to extract global information at the region level; Wherein, Z is the input feature map, and ∣region m ∣ is the number of spatial points in the m-th region of interest region m , represents the feature value of the k-th channel of the input feature map Z at the spatial position (i, j), rs m is the descriptor of the feature map after region compression, F r-squeeze is the region compression operation identifier; Region activation operation F r-excite Using the trainable weight W in the region activation operation rs for rs m perform Activation processing to obtain the region scaling factor; where re is the region scaling factor, σ is the Sigmoid activation function, δ is the ReLU activation function, and F r-excite is the region activation operation identifier, and W rs is the trainable weight in the region activation operation, including and The third step, recalibration operation: Scale the channels and regions of the input feature map respectively according to the channel scaling coefficient and the region scaling coefficient; Among them, based on the channel scaling factor ce and the region scaling factor re, the input feature map is adjusted using the recalibration function F scale Adjust the input feature map; Fscale(Z,re,ce) = ce ⊙ ∑ m (re m ·Z·1((i,j) ∈ region m ); where, F scale (Z, re, ce) is a recalibration function that recalibrates the feature Z using the obtained re and ce vectors, and re m is the region scaling factor in m region, 1((i,j) ∈ region m )) is an indicator function, and ⊙ represents element-wise multiplication.
6. The pneumonia segmentation method based on a multi-connected axial graph convolutional network according to claim 5, characterized in that: The method for the multi-connected axial graph convolutional layer MAA to separate the feature information of the width axis and the height axis includes: The first step, separation processing of the width axis and the height axis: Axially segment the input feature map of the multi-connected axial graph convolutional layer MAA along the width axis and the height axis, and extract the feature components in the width direction and the feature components in the height direction respectively; Step 2, capturing local and global features: The attention mechanism is adopted for dependency modeling of the feature components on both the width axis and the height axis. The local features are captured through the attention mechanism, emphasizing the detailed correlation in the short-distance space, and at the same time, the global features are captured to model the context relationship in the long-distance space; Step 3, fusing the information on the width axis and the height axis: The features processed on the width axis and the height axis are recombined, and a global feature representation is formed through weighted summation or a specific fusion mechanism.
7. A pneumonia segmentation method based on a multi-connected axial graph convolutional network according to claim 6, characterized in that: The method for the multi-connected axial graph convolutional layer MAA to determine the dependency relationship between local and global features by calculating the maximum difference between each node feature and its neighborhood node features includes: Calculating the maximum difference between each node feature and its neighborhood node features to obtain the dependency relationship between local features: The expression for calculating the maximum difference between the feature at each node position (i, j) of the input feature map of the multi-connected axial graph convolutional layer MAA and the feature at the neighborhood node position is: where Δz p,(i,j) is the maximum difference between the feature at the node position (i, j) and the features at the neighboring node positions, N(i, j) is the neighborhood of the node position (i, j), z (i,j) is the feature value at the node position (i, j), and z (k,l) is the feature of the neighboring node position (k, l).
8. A pneumonia segmentation method based on a multi-connected axial graph convolutional network according to claim 1, characterized in that: The loss function of the multi-connected axial graph convolutional network MAGNN is a weighted combination of the Dice loss and the cross-entropy loss. The evaluation metrics of the multi-connected axial graph convolutional network MAGNN include: Dice coefficient: used to evaluate the overlap degree between the segmentation result and the ground truth label; Intersection over Union (IoU): used to quantify the accuracy of the segmentation result; 95% Hausdorff distance: measures the maximum distance between the segmentation boundary and the ground truth boundary.
Citation Information
Patent Citations
Medical image segmentation method based on global and local feature reconstruction network
CN114612479A
New coronal pneumonia CT focus segmentation method based on gated axial attention
CN114998579A
Lung CT image segmentation method based on mixed Swin Transform U-Net
CN117274147A
Lung CT image processing method and system for lung injury prediction
CN117934480A
Image-based edge attention gate medical image segmentation method and device
CN117953208A