Cardiac medical image segmentation system based on global-local fusion attention mechanism
Through the global-local fusion attention mechanism and loss function module, combined with the global-local fusion attention module and the spatial gated enhancement unit, the problem of insufficient modeling of global information and local information, channel information and spatial information in the medical image segmentation model is solved, and efficient pixel-level segmentation accuracy is achieved.
Patent Information
- Application Number
- CN202510732523.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing medical image segmentation models have difficulty balancing global and local information when processing complex structures, lack modeling of channel and spatial information, have high computational complexity, and lack pixel-level segmentation accuracy.
A global-local fusion attention mechanism is adopted. Through the global-local fusion attention module and loss function module, global context perception and local detail focus are combined to dynamically allocate weights and optimize feature expression. The GLFA and SGE modules are introduced in the ResNet encoder to improve feature extraction capabilities and segmentation accuracy.
It significantly improves the feature expression ability and performance of medical image segmentation models, especially in complex anatomical structure segmentation tasks, improves segmentation accuracy and robustness, and reduces computational complexity.
Smart Images

Figure CN120672775A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a cardiac medical image segmentation system based on a global-local fusion attention mechanism. Background Art
[0002] With the rapid development of medical imaging technology, medical images have become an important tool to assist doctors in disease diagnosis. In recent years, the vigorous development of deep learning technology has brought revolutionary breakthroughs in automated medical image segmentation. In particular, the segmentation method based on convolutional neural networks (CNN), with its powerful feature extraction capabilities and end-to-end learning capabilities, has achieved remarkable results in the field of medical image segmentation. At present, convolutional neural networks have become the mainstream technical framework for medical image segmentation tasks. Among them, segmentation models represented by fully convolutional networks (FCN) and UNet are widely used in segmentation tasks of various medical images such as heart, lungs, and brain. These networks can achieve accurate segmentation of regions of interest in medical images through multi-scale feature extraction and gradual upsampling. However, the complexity of medical image segmentation tasks also places higher demands on existing models, which are mainly reflected in the following aspects:
[0003] 1. Insufficient modeling of global information and local information: The anatomical structures in medical images are usually complex and diverse, such as the four-chamber structure of the heart, the distribution of vascular networks, and the morphological changes of tumors. These structures require both global contextual information to understand the overall morphology and local detail information to accurately characterize the boundaries. However, most existing segmentation models tend to focus on one aspect of global information or local information when extracting features, and it is difficult to take into account the characteristics of both at the same time. For example, global modeling methods can capture long-range dependencies, but tend to ignore local details, resulting in blurred boundaries; while local modeling methods, although sensitive to details, lack a grasp of the overall structure and are prone to local misjudgments. Existing segmentation models find it difficult to simultaneously take into account global information and local information modeling, which significantly limits the performance of segmentation models on complex structures.
[0004] 2. Insufficient utilization of channel information and spatial information: In medical image segmentation tasks, the contributions of features in different channels and spatial positions to segmentation are uneven. Some modules in the classic attention mechanism adaptively adjust the importance of each channel by modeling the relationship between channels, thereby improving the segmentation performance. However, these methods mainly focus on modeling channel feature information and ignore the role of spatial information. Spatial information is particularly important for medical image segmentation tasks, especially when dealing with blurred boundaries or tiny structures. Lack of attention to spatial information may lead to segmentation results that are not precise enough. In addition, although some attention mechanisms that combine channel and spatial information can model these two types of information at the same time, their computational cost is high and it is difficult to adapt to the economic requirements of medical image segmentation tasks.
[0005] 3. Computational complexity: Medical image segmentation tasks typically involve high-resolution three-dimensional or two-dimensional image data, which significantly increases the computational overhead of the segmentation model. For example, existing complex attention mechanisms consume a large amount of computing resources when processing high-resolution feature maps, making it difficult to meet the real-time and resource-efficiency requirements of practical applications. Furthermore, while some lightweight attention mechanisms have alleviated this problem to some extent by reducing the number of parameters and computational overhead, these methods generally lack the ability to model local information, resulting in suboptimal performance on tasks such as segmentation boundaries and fine-grained objects.
[0006] 4. High-precision pixel-level segmentation: Medical image segmentation tasks require precise classification of each pixel, placing higher demands on the segmentation model's ability to process details. While global attention mechanisms can capture long-range dependencies, they tend to overfocus on global information and neglect detailed modeling of local features, leading to blurred boundaries or object detection errors. Furthermore, when faced with medical images with complex anatomical structures and blurred lesion boundaries, the accuracy and robustness of existing segmentation models still have significant room for improvement.
[0007] In recent years, researchers have proposed some lightweight attention mechanisms. These methods have alleviated the above problems to a certain extent by reducing the number of parameters and computational overhead. However, these methods generally ignore the ability to model local information, resulting in insufficient segmentation accuracy for fine boundary structures. In addition, medical image segmentation tasks require higher pixel-level accuracy. Traditional global attention mechanisms may focus too much on long-range dependencies and ignore local features, resulting in blurred segmentation boundaries or target detection errors. To this end, the present invention proposes a cardiac medical image segmentation system based on a global-local fusion attention mechanism. Summary of the Invention
[0008] The purpose of the present invention is to provide a cardiac medical image segmentation system based on a global-local fusion attention mechanism, which can solve the problems of insufficient channel information and spatial information modeling, neglect of local features and high computational complexity in the existing attention mechanism in medical image segmentation tasks.
[0009] The technical solutions adopted by the present invention are as follows:
[0010] A cardiac medical image segmentation system based on a global-local fusion attention mechanism is characterized by comprising: a global-local fusion attention mechanism module, a global-local fusion attention segmentation network architecture, and a loss function module;
[0011] The global-local fusion attention mechanism module dynamically assigns weights to different regions by combining global context perception with local detail focus. Its global module extracts overall structure or long-range dependencies, while the local module captures subtle features of neighboring regions. The adaptive fusion strategy is used to complement and integrate these two types of information to achieve a multi-level collaborative representation of key features of complex data.
[0012] Preferably, in the global-local fusion attention mechanism module: input feature map First, local average pooling (LAP) is used to extract local spatial information, and the input is converted into This paper selects S=7; then uses two branches to extract global information and local information respectively; the global branch reduces the dimension of the local feature map through global average pooling (GAP) and converts it into a global feature map To adapt to the one-dimensional convolution operation; then, a one-dimensional convolution kernel (Conv1d) is used to learn cross-channel dependencies in the channel dimension and generate channel attention weights;
[0013] The local branch is directly in the local feature map Modeling spatial information on the local spatial context to capture the local spatial context; in order to reduce the computational overhead, first The number of channels is reduced with a compression rate of r = 4. Then, two consecutive 7×7 large-kernel two-dimensional convolutions are performed to extract local spatial information: the first large-kernel convolution reduces the number of channels from C to C / r, and extracts local spatial features through BatchNormalization (BN) and ReLU activation; the second large-kernel convolution restores the number of channels from C / r to C, and again performs BN operation to generate local spatial attention weights;
[0014] Finally, the attention maps generated by the global branch and the local branch are fused, and the resolution of the fused attention map is adjusted to the same as the input feature map. Consistent, then the fused attention map is combined with the original input feature map Multiply channel by channel to generate an optimized feature map; this operation highlights important areas while suppressing redundant information, achieving feature optimization that simultaneously enhances channel sensitivity and spatial positioning capabilities.
[0015] The global-local fusion attention segmentation network architecture is as follows: the input image is firstly extracted with multi-level local features through the ResNet encoder; a high-dimensional embedding sequence is generated through linear projection, and spatial information is injected through position encoding; the global context dependency is modeled through a multi-layer TransformerLayer (including multi-head self-attention and feedforward network) in the encoding stage, and the feature distribution is optimized through layer normalization (LayerNormal); in the jump connection stage, before the features of each layer of the encoder are fused with the corresponding layer of the decoder, the GLFA attention mechanism is introduced to dynamically calibrate the feature importance weights, suppress redundant information and enhance the response of key areas; the decoder part adopts a multi-level progressive upsampling structure, and each level improves the resolution through transposed convolution, while integrating The SGE module is used as a spatial gated enhancement unit; in the process of step-by-step upsampling, traditional decoders are prone to losing fine-grained spatial details (especially edges and tiny structures) due to the gradual recovery of feature map resolution, and simple channel splicing operations are difficult to effectively filter key information in the multi-scale features transmitted by the encoder; the SGE module divides the feature map into multiple subgroups through the grouped spatial attention mechanism, and calculates the spatial attention weight map of each subgroup separately, dynamically enhancing the feature response of the target area and suppressing irrelevant background noise; through the collaborative optimization of channel attention and spatial mask, the feature expression is enhanced to enhance the detail recovery ability; after the upsampled features are channel-spliced with the encoder features weighted by GLFA, the local structure is refined through convolution, and finally a pixel-level segmentation mask is generated. The improved network significantly improves the robustness of cross-level feature fusion and the accuracy of spatial detail reconstruction through the synergy of GLFA and SGE modules;
[0016] The loss function module combines cross-entropy loss (Cross-EntropyLoss) and Dice loss (DiceLoss) to quantify the difference between the pixel-level segmentation mask segmentation result and the true value, and guides the optimization algorithm pixel-level segmentation mask parameters, and finally outputs a high-resolution pixel-level segmentation mask.
[0017] The loss function module loss function can be expressed as:
[0018]
[0019] Among them, L CE is the cross entropy loss, which is used to measure the difference between the category distribution predicted by the model and the true distribution; L Dice is the Dice loss, which is used to evaluate the degree of overlap between the predicted results and the true labels; N is the total number of all pixels in the image; C is the total number of categories, yi,c is the true label of pixel i belonging to category C, using one-hot encoding; y i is the probability value predicted by the model; is the true label, which is represented by binary value; α is the weight coefficient, and in this invention, α=0.5.
[0020] The technical effects achieved by the present invention are:
[0021] This paper proposes a medical image segmentation system based on the global-local fused attention mechanism (GLFA). By combining global channel information, local spatial information, and collaboratively modeling global and local features, this method significantly improves the segmentation model's feature representation and performance. Specifically, the GLFA mechanism utilizes a joint modeling of local and global branches to fully exploit the potential of channel and spatial information.
[0022] In response to the defects and improvement needs of existing technologies in medical image segmentation tasks, the present invention provides a medical image segmentation method and its network architecture based on the global-local fusion attention mechanism (GLFA). The invention aims to solve the problems of insufficient channel information and spatial information modeling, neglect of local features, and high computational complexity in the existing attention mechanism in medical image segmentation tasks. In addition, the present invention significantly improves the feature extraction capability and segmentation accuracy of the segmentation network by embedding the global-local fusion attention module into the classic segmentation network (TransUNet). Through innovative module design and optimization strategies, the method of the present invention demonstrates excellent performance and robustness in processing complex anatomical structure segmentation tasks.
[0023] This system addresses the problem that existing global attention mechanisms, while capable of capturing long-range dependencies, tend to overfocus on global information and neglect detailed modeling of local features, leading to blurred boundaries or misdetection. Existing attention mechanisms also suffer from the disadvantage of excessive computational effort. By combining global channel information, local spatial information, and collaborative modeling of global and local features, the segmentation model's feature representation and performance can be significantly improved. Furthermore, we use one-dimensional convolution in the global branch to reduce computational effort by compressing the number of channels in the local branch. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a system block diagram of a cardiac medical image segmentation system based on a global-local fusion attention mechanism of the present invention;
[0025] Figure 2 is a schematic diagram of the GLFA module in the present invention;
[0026] Figure 3 It is a schematic diagram of the SGE module in the present invention. DETAILED DESCRIPTION
[0027] In order to make the purpose and advantages of the present invention more clearly understood, the present invention is described in detail below with reference to the following examples. It should be understood that the following text is only used to describe one or more specific embodiments of the present invention and does not strictly limit the scope of protection of the present invention.
[0028] like Figure 1 As shown, a cardiac medical image segmentation system based on a global-local fusion attention mechanism is characterized by comprising: a global-local fusion attention mechanism module, a global-local fusion attention segmentation network architecture, and a loss function module;
[0029] The global-local fusion attention mechanism module dynamically assigns weights to different regions by combining global context perception with local detail focus. Its global module extracts overall structure or long-range dependencies, while the local module captures subtle features of neighboring regions. The adaptive fusion strategy is used to complement and integrate these two types of information to achieve a multi-level collaborative representation of key features of complex data.
[0030] Preferably, in the global-local fusion attention mechanism module: input feature map First, local average pooling (LAP) is used to extract local spatial information, and the input is converted into This paper selects S=7; then uses two branches to extract global information and local information respectively; the global branch reduces the dimension of the local feature map through global average pooling (GAP) and converts it into a global feature map To adapt to the one-dimensional convolution operation; then, a one-dimensional convolution kernel (Conv1d) is used to learn cross-channel dependencies in the channel dimension and generate channel attention weights;
[0031] The local branch is directly in the local feature map Modeling spatial information on the local spatial context to capture the local spatial context; in order to reduce the computational overhead, first The number of channels is reduced with a compression rate of r = 4. Then, two consecutive 7×7 large-kernel two-dimensional convolutions are performed to extract local spatial information: the first large-kernel convolution reduces the number of channels from C to C / r, and extracts local spatial features through BatchNormalization (BN) and ReLU activation; the second large-kernel convolution restores the number of channels from C / r to C, and again performs BN operation to generate local spatial attention weights;
[0032] Finally, the attention maps generated by the global branch and the local branch are fused, and the resolution of the fused attention map is adjusted to the same as the input feature map. Then the fused attention map is combined with the original input feature map Multiply channel by channel to generate an optimized feature map; this operation highlights important areas while suppressing redundant information, achieving feature optimization that simultaneously enhances channel sensitivity and spatial positioning capabilities.
[0033] The global-local fusion attention mechanism module described in the present invention: The specific process is as follows:
[0034] Input feature map (B is the batch size, C is the number of channels, H and W are height and width) After local average pooling processing: Local Average Pooling (LAP): A local average pooling operation with a sliding window size of 7×7 is used to downsample each local area of the input feature map to extract local spatial information; the local feature map is output
[0035] Global channel branch design; the global branch is designed to model the global dependencies across channels; Global Average Pooling (GAP) is used to perform local average pooling on the feature maps. Use the global branch to reduce the dimension of the local feature map through global average pooling (GAP) and convert it into a global feature map
[0036] Dynamic one-dimensional convolution kernel design:
[0037] To avoid information loss caused by a fixed convolution kernel size, a dynamically adjusted one-dimensional convolution kernel is designed. The convolution kernel size k is dynamically calculated based on the number of input channels C:
[0038]
[0039] in,
[0040]
[0041] k is the one-dimensional convolution kernel size, which is always an odd number to ensure the symmetry of the convolution operation. C is the number of channels, and γ and b are hyperparameters that control the compression factor and base bias of the convolution kernel size, respectively.
[0042] Cross-channel attention generation: using one-dimensional convolution pairs Perform feature interaction to generate channel attention weights The formula is:
[0043]
[0044] Among them, σ is the Sigmoid activation function, which is used to normalize the attention weight to [0,1].
[0045] Local spatial feature extraction; the local branch focuses on capturing spatial detail information;
[0046] Will The number of channels is compressed from r=4 to C / 4 to reduce the amount of computation. Local spatial features are then extracted through two consecutive 7x7 large kernel convolutions. The expression is as follows:
[0047]
[0048] Among them, F local BN is the batch normalization for local spatial features, and ReLU is the activation function.
[0049] Spatial attention generation: The feature map is restored to the original input resolution through adaptive average pooling to generate spatial attention weights.
[0050] Global and local attention are fused; the attention weights generated by the global branch and the local branch are fused to generate an optimized attention weight map. The global and local attention weights are balanced by the hyperparameter β (default 0.5). The expression of the fused attention weight is as follows:
[0051] A fusion =β·A global +(1-β)·A local
[0052] Among them, A fusion is the fused attention weight, A global is the global branch attention weight, A local is the local branch attention weight.
[0053] The fused attention weights are then combined with the input feature map through a channel-by-channel multiplication operation to highlight key areas and suppress redundant information, thereby optimizing feature expression. The final output feature map expression is as follows:
[0054]
[0055] Among them, X output is the final output feature map, A fusion is the fused attention weight.
[0056] This paper improves the classic segmentation network TransUNet based on the GLFA mechanism and proposes a new medical image segmentation network architecture, the Global-Local Fusion Attention Segmentation Network. This network effectively improves segmentation accuracy and the ability to recover boundary details through the coordinated optimization of the encoder and decoder.
[0057] The global-local fusion attention segmentation network architecture is as follows: the input image is firstly extracted with multi-level local features through the ResNet encoder; a high-dimensional embedding sequence is generated through linear projection, and spatial information is injected through position encoding; the global context dependency is modeled through a multi-layer TransformerLayer (including multi-head self-attention and feedforward network) in the encoding stage, and the feature distribution is optimized through layer normalization (LayerNormal); in the jump connection stage, before the features of each layer of the encoder are fused with the corresponding layer of the decoder, the GLFA attention mechanism is introduced to dynamically calibrate the feature importance weights, suppress redundant information and enhance the response of key areas; the decoder part adopts a multi-level progressive upsampling structure, and each level improves the resolution through transposed convolution, while integrating The SGE module is used as a spatial gated enhancement unit; in the process of step-by-step upsampling, traditional decoders are prone to losing fine-grained spatial details (especially edges and tiny structures) due to the gradual recovery of feature map resolution, and simple channel splicing operations are difficult to effectively filter key information in the multi-scale features transmitted by the encoder; the SGE module divides the feature map into multiple subgroups through the grouped spatial attention mechanism, and calculates the spatial attention weight map of each subgroup separately, dynamically enhancing the feature response of the target area and suppressing irrelevant background noise; through the collaborative optimization of channel attention and spatial mask, the feature expression is enhanced to enhance the detail recovery ability; after the upsampled features are channel-spliced with the encoder features weighted by GLFA, the local structure is refined through convolution, and finally a pixel-level segmentation mask is generated. The improved network significantly improves the robustness of cross-level feature fusion and the accuracy of spatial detail reconstruction through the synergy of GLFA and SGE modules;
[0058] The loss function module combines cross-entropy loss (Cross-EntropyLoss) and Dice loss (DiceLoss) to quantify the difference between the pixel-level segmentation mask segmentation result and the true value, and guides the optimization algorithm pixel-level segmentation mask parameters, and finally outputs a high-resolution pixel-level segmentation mask.
[0059] The loss function module loss function can be expressed as:
[0060]
[0061] Among them, L CE is the cross entropy loss, which is used to measure the difference between the category distribution predicted by the model and the true distribution; L Dice is the Dice loss, which is used to evaluate the degree of overlap between the predicted results and the true labels; N is the total number of all pixels in the image; C is the total number of categories, y i,c is the true label of pixel i belonging to category C, using one-hot encoding; y i is the probability value predicted by the model; is the true label, which is represented by binary value; α is the weight coefficient, and in this invention, α=0.5.
[0062] The present invention is in the specific experimental process
[0063] The experimental platform used in this study is shown in Table 1. The datasets used include the ACDC heart dataset and the CAMUS dataset. All network structures were evaluated using the same training and test sets.
[0064] The ACDC (Automated Cardiac Diagnosis Challenge) cardiac dataset is a widely used medical image segmentation dataset focused on the automated segmentation of cardiac MRI images. This dataset contains five cardiac pathology types (normal, myocardial hypertrophy, dilated cardiomyopathy, abnormal right ventricle, and post-myocardial infarction heart), each with data from 100 patients. Each data set includes dual-phase images of end-diastole (ED) and end-systole (ES), along with expert manual annotations of the left ventricle (LV), right ventricle (RV), and myocardium (Myocardium). The training set and test set contain 80 and 20 patients, respectively. In this experiment, the training set contains 1000 MRI slice images, and the test set contains 200 MRI slice images. The batch size during training is 2, and the images are resized to 256×256 pixels, with a training batch size of 100. To quantify the experimental results, the Dice coefficient, precision, and recall are used as performance evaluation metrics. The expressions for these three metrics are as follows:
[0065]
[0066] Where |A∩B| represents the number of pixels in the overlap between A and B, |A| represents the number of pixels in the predicted positive class, and |B| represents the number of pixels in the ground-truth positive class. TP is the number of true positive samples correctly classified as positive, while FP is the number of false negative samples that are actually positive but incorrectly classified as negative.
[0067] Table 1 Experimental platform
[0068]
[0069] Comparison with other attention mechanisms: To verify the effectiveness of the dual-branch attention mechanism proposed in this paper, we used the ACDC dataset and embedded the current mainstream attention mechanism into the skip connection part of the segmentation network designed in this paper, and conducted a comparative experiment with the method in this paper. The comparison results are shown in the following table:
[0070] Table 2 Precision stable average value
[0071]
[0072]
[0073] Table 3 Recall stable average value
[0074]
[0075] Tables 2 and 3 show that the GLFA model outperforms all compared models in terms of both precision and recall. For precision, the average values of Rv, Myo, and Lv for the GLFA model are 0.8067, 0.8546, and 0.9303, respectively. These values represent improvements of 0.35%, 1.60%, and 0.19% over the next-best SE model, and 3.55%, 4.74%, and 2.82% over the NONE model. For recall, the average values of Rv, Myo, and Lv for the GLFA model are 0.8345, 0.8853, and 0.9296, respectively. These values represent improvements of 1.12%, 1.12%, and 0.07% over the SE model, and 8.26%, 8.00%, and 4.61% over the NONE model. Combined with data analysis, the GLFA model showed better feature extraction and task generalization capabilities compared with the attention mechanism models CBAM, ECA and BAM, especially in the myocardial (Myo) segmentation task, further verifying its significant performance advantage in complex medical image segmentation tasks.
[0076] To clarify the contribution of different modules to overall performance and reveal the key factors that improve model performance, we conducted ablation experiments using the ACDC dataset. We selected TransUNet as the baseline, and the experimental results are shown in Table 4.
[0077] Table 4 Ablation data (Dice coefficient)
[0078]
[0079]
[0080] Table 4 shows that both the GLFA module and the SGE module play a significant role in improving model performance, and their combination further enhances performance. In the Rv task, the baseline model's Dice coefficient is 87.34. This increases to 90.14 (a 3.20% increase) with the addition of the GLFA module, and to 90.37 (a 3.47% increase) with the addition of the SGE module. When both the GLFA and SGE modules are added, the Dice coefficient reaches 91.17, a 4.39% improvement over the baseline. In the Myo task, the baseline model's Dice coefficient is 86.47. This increases to 89.88 (a 3.94% increase) with the addition of the GLFA module, and to 87.89 (a 1.64% increase) with the addition of the SGE module. The combined Dice coefficient reaches 90.02 (a 4.10% improvement over the baseline). In the Lv task, the addition of the GLFA module alone increased the Dice coefficient from 94.74 to 95.25 (an increase of 0.54%). The addition of the SGE module alone did not significantly improve it (95.02%). However, after combining the two, the Dice coefficient reached 95.90, an increase of 1.22% compared to the baseline.
[0081] Overall, the GLFA module achieved significant performance improvements across all tasks, with particularly strong performance in the Rv and Myo tasks. The SGE module also demonstrated performance gains across multiple tasks. When used in conjunction with the SGE modules, their synergistic effect further enhanced the model's feature representation and segmentation performance, significantly improving segmentation accuracy across various tasks. These experimental results fully validate the effectiveness and complementary nature of the two module designs, providing important support for further optimization of model performance.
[0082] In the present invention, Figure 1 System block diagram of cardiac medical image segmentation system based on global-local fusion attention mechanism; Figure 1 As can be seen, the input image first passes through the ResNet encoder to extract multi-level local features. This is then linearly projected to generate a high-dimensional embedding sequence, and spatial information is injected through positional encoding. During the encoding phase, global contextual dependencies are modeled using a multi-layer Transformer Layer (including multi-head self-attention and feedforward networks), and layer normalization (LayerNormal) is used to optimize feature distribution. During the skip connection phase, before the features of each encoder layer are fused with the corresponding decoder layer, the GLFA attention mechanism is introduced to dynamically calibrate feature importance weights, suppress redundant information, and enhance the response of key areas.
[0083] The decoder part adopts a multi-level progressive upsampling structure, each level improves the resolution by transposed convolution, and integrates the SGE module (spatial gated enhancement unit) such as Figure 3As shown in the figure. During the step-by-step upsampling process, the traditional decoder is prone to losing fine-grained spatial details (especially edges and tiny structures) due to the gradual recovery of the feature map resolution, while the simple channel splicing operation is difficult to effectively filter the key information in the multi-scale features transmitted by the encoder. The SGE module divides the feature map into multiple subgroups through the grouped spatial attention mechanism, calculates the spatial attention weight map of each subgroup separately, dynamically enhances the feature response of the target area and suppresses irrelevant background noise. The feature expression is optimized through the synergistic effect of channel attention and spatial mask, and the detail recovery ability is enhanced; after the upsampled features are channel-spliced with the encoder features weighted by GLFA, the local structure is further refined through convolution, and finally a pixel-level segmentation mask is generated. The improved network significantly improves the robustness of cross-level feature fusion and the accuracy of spatial detail reconstruction through the synergy of GLFA and SGE modules.
[0084] In the present invention, Figure 2 is a schematic diagram of the GLFA module; Figure 2 The following figure shows the principle of GLFA. Input feature map First, local average pooling (LAP) is used to extract local spatial information, and the input is converted into In this paper, S=7 is selected. Then, two branches are used to extract global information and local information respectively. The global branch reduces the dimension of the local feature map through global average pooling (GAP) and converts it into a global feature map. To adapt to the one-dimensional convolution operation. Then, a one-dimensional convolution kernel (Conv1d) is used to learn cross-channel dependencies in the channel dimension and generate channel attention weights. The convolution kernel size k of this one-dimensional convolution is dynamically calculated and is proportional to the input C. The formula is as follows:
[0085]
[0086] in:
[0087]
[0088] C is the number of channels, γ and b are hyperparameters that control the compression factor and base bias of the convolution kernel size, respectively; the default values are γ = 2 and b = 1. k is the one-dimensional convolution kernel size, which is always an odd number to ensure symmetry of the convolution operation. The output of the global branch is a global attention map, which is resized to the same resolution as the original input feature map through adaptive average pooling or expansion.
[0089] The local branch is directly in the local feature map To reduce computational overhead, we first model spatial information on The number of channels is reduced with a compression rate of r = 4, and then the local spatial information is extracted through two consecutive 7×7 large-kernel two-dimensional convolutions: the first large-kernel convolution reduces the number of channels from C to C / r, and extracts local spatial features through BatchNormalization (BN) and ReLU activation; the second large-kernel convolution restores the number of channels from C / r to C, and again performs BN operation to generate local spatial attention weights.
[0090] Finally, the attention maps generated by the global branch and the local branch are fused, and the resolution of the fused attention map is adjusted to the same as the input feature map. Consistent, then the fused attention map is combined with the original input feature map Multiplying the channels one by one generates an optimized feature map. This operation highlights important areas while suppressing redundant information, achieving feature optimization that simultaneously enhances channel sensitivity and spatial localization capabilities.
[0091] In the present invention, Figure 3 This is a schematic diagram of the SGE module; the Spatial Group Enhance (Spatial Group Enhance, SGE) module, as a new lightweight spatial attention mechanism, shows unique advantages in local feature modeling. Its core idea is to divide the input feature map into several subgroups according to the channel dimension, and independently calculate the spatial attention weight map for each group of features. Specifically, each group of features first generates channel statistics through global average pooling, then generates a spatial attention mask through two layers of convolution, and finally normalizes it to a weight value of 0-1 through the Sigmoid function. Through the grouping operation, SGE reduces the computational complexity of spatial attention to 1 / N (N is the number of groups) of the traditional method, while avoiding the information loss caused by channel dimensionality reduction. Experiments show that SGE can effectively improve the ability to capture edge details in semantic segmentation tasks, and is especially suitable for scenes with blurred organ boundaries in medical images.
[0092] The foregoing is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained herein shall, unless otherwise specified or limited, be implemented in accordance with conventional means in the art.
Claims
1. A cardiac medical image segmentation system based on a global-local fusion attention mechanism, characterized by: include: Global-local fusion attention mechanism module, global-local fusion attention segmentation network architecture and loss function module; The global-local fusion attention mechanism module dynamically assigns weights to different regions by combining global context perception with local detail focus. Its global module extracts the overall structure or long-range dependencies, while the local module captures the subtle features of the adjacent area. It uses an adaptive fusion strategy to complement and integrate the two types of information to complete a multi-level collaborative representation of the key features of complex data. The global-local fusion attention segmentation network architecture first extracts multi-level local features from the input image through a ResNet encoder; generates a high-dimensional embedding sequence through linear projection, and injects spatial information through positional encoding; in the encoding stage, global context dependencies are modeled through multiple layers including multi-head self-attention and feedforward networks, and feature distribution is optimized through layer normalization. In the skip connection stage, before the features of each layer of the encoder are fused with the corresponding layer of the decoder, the GLFA attention mechanism is introduced to dynamically calibrate the feature importance weights; The decoder adopts a multi-level progressive upsampling structure, improving the resolution at each level through transposed convolution, and incorporating a spatial gated enhancement unit as an SGE module. The SGE module divides the feature map into multiple subgroups through the grouped spatial attention mechanism, and calculates the spatial attention weight map of each subgroup separately; it optimizes the feature expression through the synergistic combination of channel attention and spatial mask to enhance the detail recovery capability; after channel-wise splicing of the upsampled features and the GLFA-weighted encoder features, it refines the local structure through convolution and finally generates a pixel-level segmentation mask. The loss function module combines cross entropy loss and Dice loss to quantify the difference between the pixel-level segmentation mask segmentation result and the true value, and guides the optimization algorithm pixel-level segmentation mask parameters, and finally outputs the pixel-level segmentation mask.
2. The cardiac medical image segmentation system based on global-local fusion attention mechanism according to claim 1, characterized in that: In the global-local fusion attention mechanism module: input feature map First, local average pooling (LAP) is used to extract local spatial information, and the input is converted into This paper selects S=7; then uses two branches to extract global information and local information respectively; the global branch reduces the dimension of the local feature map through global average pooling (GAP) and converts it into a global feature map To adapt to the one-dimensional convolution operation; then, a one-dimensional convolution kernel (Conv1d) is used to learn cross-channel dependencies in the channel dimension and generate channel attention weights.
3. The cardiac medical image segmentation system based on global-local fusion attention mechanism according to claim 2, characterized in that: In the global-local fusion attention mechanism module: the local branch directly in the local feature map Modeling spatial information on the local space context; first The number of channels is reduced with a compression rate of r = 4, and then two consecutive 7×7 large-kernel two-dimensional convolutions are used to extract local spatial information: the first large-kernel convolution reduces the number of channels from C to C / r, and extracts local spatial features through Batch Normalization (BN) and ReLU activation; the second large-kernel convolution restores the number of channels from C / r to C, and again performs BN operation to generate local spatial attention weights.
4. The cardiac medical image segmentation system based on global-local fusion attention mechanism according to claim 3, characterized in that: In the global-local fusion attention mechanism module: Finally, the attention maps generated by the global branch and the local branch are fused, and the resolution of the fused attention map is adjusted to the same as the input feature map. Consistent, then the fused attention map is combined with the original input feature map Multiply channel by channel to generate the optimized feature map.
5. The cardiac medical image segmentation system based on global-local fusion attention mechanism according to claim 4, characterized in that: The loss function module loss function can be expressed as: Among them, L CE is the cross entropy loss, which is used to measure the difference between the category distribution predicted by the model and the true distribution; L Dice is the Dice loss, which is used to evaluate the degree of overlap between the predicted results and the true labels; N is the total number of all pixels in the image; C is the total number of categories, y i,c is the true label of pixel i belonging to category C, using one-hot encoding; y i is the probability value predicted by the model; is the true label, which is represented by binary value; α is the weight coefficient; α = 0.5.
Citation Information
Patent Citations
3D medical image segmentation method based on cross fusion convolution and deformable attention Transform
CN115830041A
Lung CT image segmentation method based on global and local attention mechanisms
CN117649385A
Liver and tumor segmentation method based on mixed attention
CN119169024A
Cited By
Heterogeneous feature conflict perception fusion method and device for three-dimensional medical image segmentation
CN122336307A
Heterogeneous feature conflict perception fusion method and device for three-dimensional medical image segmentation
CN122336307B