Heart image segmentation method fusing multi-receptive-field convolution and distracting attention

By integrating multi-receptive field convolution and attention shifting into a multimodal cardiac image segmentation model, the problems of high computational complexity and insufficient segmentation accuracy in existing technologies are solved, a balance between efficient global modeling and computational efficiency is achieved, and the accuracy and efficiency of cardiac image segmentation are improved.

CN120783058AActive Publication Date: 2025-10-14NORTHWESTERN POLYTECHNICAL UNIV

Patent Information

Application Number
CN202511242110.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-14
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing technologies for processing three-dimensional medical images, especially cardiac image segmentation, suffer from high computational complexity, insufficient segmentation accuracy, and excessive consumption of computing resources. In particular, in the application of the Transformer architecture, it is difficult to achieve a balance between efficient global modeling and computational efficiency.

Method used

A multimodal cardiac image segmentation model that integrates multi-receptive field convolution and shifted attention is adopted. Through the multi-receptive field deep convolution module and the three-dimensional shifted attention Transformer module, the local feature extraction capability of CNN is combined with the global modeling advantage of Transformer to construct a symmetric encoder-decoder architecture, which reduces computational complexity while improving segmentation accuracy.

Benefits of technology

The accuracy and computational efficiency of cardiac multimodal image segmentation were significantly improved, with an average Dice coefficient of 93.65%, a Jaccard coefficient of 93.1%, and a Hausdorff distance of 42.02 mm. The number of parameters and computational complexity were compressed by 11.94% and 21.49%, respectively, making it suitable for actual clinical scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783058A_ABST
    Figure CN120783058A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical image processing, and provides a heart image segmentation method fusing multi-receptive-field convolution and distracting to solve the technical problem that construction of Transform variants with global modeling capability and controllable calculation complexity is still an urgent breakthrough in the current medical image segmentation field. Comprising the following steps: S1, preprocessing training data to obtain processed training data; s2, constructing a multi-modal heart image segmentation model; s3, training a multi-modal heart image segmentation model through the processed training data to obtain a completely trained segmentation model; and S4, segmenting the to-be-detected heart image by using the completely trained segmentation model to obtain a segmentation result. According to the method, the convolutional neural network and the global self-attention mechanism are deeply fused, so that the precision of heart multi-modal image segmentation is remarkably improved, and the adaptability of the model to a complex anatomical structure is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of medical image processing, and particularly relates to a cardiac image segmentation method fusing multi-receptive field convolution and transfer attention. BACKGROUND

[0002] Precise cardiac structure segmentation is crucial for disease diagnosis, surgical planning, and prognosis evaluation, especially in clinical applications such as cardiac intervention surgery, radiofrequency ablation therapy, and artificial heart implantation. Three-dimensional cardiac CT and MRI images can provide millimeter-level resolution information of cardiac chambers, myocardium, and vascular structures, but face multiple challenges: 1. Anatomical complexity: The heart contains seven main anatomical structures with high dynamic deformation characteristics. For example, the right ventricular wall is only 3-5 mm thick and its shape changes dramatically during systole, leading to significant errors in traditional threshold segmentation methods. In addition, the boundaries between different structures in the heart are ambiguous, such as the mitral valve region between the left ventricle and the left atrium, and the transition region between the myocardium and the cardiac chamber, which greatly complicates accurate segmentation. The thickness of the ventricular wall changes by up to 30-40% during the cardiac cycle, which requires segmentation algorithms to have dynamic adaptation capabilities.

[0003] 2. Modality difference: CT images have low soft tissue contrast (the HU value difference between myocardium and cardiac chamber is about 20-50), while MRI images are based on T1-weighted and T2-weighted imaging characteristics, resulting in significant differences in structure visualization between these two different imaging sequences. For example, in T1-weighted images, the myocardium appears as moderate signal intensity, while in T2-weighted images it appears as high signal intensity. This difference between modalities makes it extremely challenging to develop a universal segmentation algorithm. At the same time, metal implant artifacts in CT images and motion artifacts in MRI images can affect segmentation accuracy.

[0004] 3. Data limitations: Three-dimensional medical image annotation requires professional physicians to outline each layer, which takes about 4-6 hours for a single CT heart annotation. Due to patient privacy protection, the size of public data sets is limited. The public data set such as the Automatic Cardiac Diagnosis Challenge (ACDC) only contains 100 labeled data, which is far from enough to train a deep learning model. In addition, the scanning equipment and parameters used by different medical institutions differ, resulting in inconsistent data distribution, further increasing the difficulty of model generalization.

[0005] 4. Computational complexity: Standard three-dimensional cardiac CT images can have a size of 512x512x300 voxels, and direct application of Transformer global attention will result in computational complexity ), far exceeding the GPU memory capacity. Taking the current mainstream NVIDIA A100 GPU as an example, its 80GB memory can only handle about 256x256x128 voxels of data. This computational complexity not only limits the training efficiency of the model, but also affects its real-time application in clinical practice.

[0006] The rapid development of deep learning technology has driven the paradigm shift of medical image segmentation methods from traditional algorithms to deep learning. Current research focuses on three types of technical architectures: convolutional neural networks (CNNs), Transformers, and their hybrid architectures.

[0007] Since AlexNet achieved a milestone in the 2012 ImageNet competition, CNNs have gradually become the core technology paradigm in image segmentation, thanks to their superior feature learning capabilities. This architecture, through a combination of stacked convolutional and pooling layers, achieves multi-level feature extraction, from local texture to global semantics, and uses fully connected layers to perform classification or regression tasks. In the field of medical image segmentation, CNN technology has been successfully applied to accurately segment key anatomical structures such as the heart, brain tissue, and lung nodules. Its advantage lies primarily in its ability to independently represent complex biomedical features. The fully convolutional network (FCN) pioneered a complete architectural innovation for end-to-end semantic segmentation. By replacing the fully connected layers of traditional CNNs with convolutional layers, it overcomes the limitations of input image size and directly outputs pixel-level predictions. This innovation laid the methodological foundation for the subsequent evolution of medical image segmentation frameworks. As a classic architecture for medical image segmentation, the U-Net significantly improves segmentation accuracy and computational efficiency through its encoder-decoder symmetric structure and cross-layer skip connections. Its encoder extracts abstract features through progressive downsampling, while the decoder restores the spatial resolution of feature maps through deconvolution. Skip connections effectively mitigate information attenuation by fusing shallow details with deep semantic information. Building on the U-Net architecture, R2U-Net innovatively integrates residual modules and recurrent convolutions, achieving improved performance while maintaining the same number of parameters. Residual connections improve gradient propagation efficiency by establishing cross-layer information pathways, while recurrent convolutions enhance the representation of local features by reusing temporal features. LadderNet further enhances the ability to capture multi-scale features by cascading multiple groups of U-Net units to construct a feature pyramid. Targeting the processing needs of 3D medical images, V-Net, using 3D convolution kernels and spatial pooling, has achieved the first end-to-end volumetric data segmentation framework. DeepIGeoS proposes an interactive segmentation paradigm that, by introducing a manual feedback correction mechanism, reduces user interaction while improving the clinical applicability of segmentation results. Notably, despite the significant progress made in medical image segmentation, CNNs' inherent local receptive field still poses a fundamental constraint on modeling global context. This structural characteristic leads to theoretical bottlenecks in the network's long-range dependency modeling and multi-scale feature fusion, which becomes a key factor restricting further improvement of segmentation accuracy.

[0008] With the groundbreaking progress of the Visual Transformer (ViT) in computer vision, its application in medical image segmentation is rapidly gaining momentum. This architecture utilizes a self-attention mechanism to build a global pixel correlation model, overcoming the local receptive field limitations of traditional CNNs and significantly improving the segmentation accuracy of complex anatomical structures. In terms of architectural innovation, Swin-Unet constructs a symmetric Transformer-based encoder-decoder structure by encoding the input image into visual tokens. Its core innovation lies in the dynamic fusion of local detail features with global semantic information through a sliding window multi-head self-attention mechanism. DS-TransUNet builds on this foundation by proposing a dual-branch Swin Transformer architecture, integrating a hierarchical feature learning mechanism into both the encoding and decoding processes for the first time: a progressive feature abstraction strategy is employed in the encoding phase, while a multi-scale feature fusion scheme is implemented in the decoding phase. This significantly improves the segmentation robustness of medical images from different modalities. However, the application of Transformers in medical image segmentation still faces significant challenges. While the self-attention mechanism possesses global modeling capabilities, its computational complexity scales quadratically with the input resolution. This makes video memory usage and computational time particularly problematic when processing high-resolution 3D medical images such as CT and MRI. This computational efficiency bottleneck severely restricts the application of Transformer in clinical scenarios such as real-time surgical navigation. How to achieve a balance between global modeling capabilities and computational efficiency has become a key breakthrough direction in current research.

[0009] Given the complementary advantages of CNNs in local feature extraction and Transformers in long-range dependency modeling, the current field of medical image segmentation is witnessing a wave of innovation in CNN-Transformer hybrid architectures. UNETR innovatively utilizes the Transformer as the core module of the encoder, building a voxel-level correlation model through a global attention mechanism. This is followed by multi-resolution feature reconstruction via a U-shaped decoder, significantly improving the segmentation accuracy of organ boundaries. Taking advantage of the characteristics of three-dimensional medical imaging, TransBTS embeds a Transformer module within a 3D CNN framework, achieving joint modeling of the spatial and channel dimensions through an axial attention mechanism. Its cascaded feature fusion strategy demonstrates excellent performance in brain tumor segmentation tasks. VT-Unet further optimizes the encoder design, constructing a cascaded self-attention module. This synergistic effect of local windowed attention and global cross-attention achieves hierarchical representation of anatomical structures. TransFuse proposes a dual-branch parallel architecture, with a CNN branch responsible for extracting high-resolution spatial details and a Transformer branch for building global semantic associations. Feature complementarity is achieved through a gated fusion module, maintaining network depth while improving computational efficiency. 3DUX-Net introduces a three-dimensional dynamic convolution unit, adapts the morphological features of different anatomical structures through deformable convolution kernels, and combines hierarchical Transformer to achieve multi-scale volume feature extraction.

[0010] SwinUNETR integrates the sliding window mechanism of the Swin Transformer with the voxel reconstruction module of the UNETR, reducing computational complexity while maintaining the ability to model long-range dependencies. This approach is particularly suitable for high-resolution CT image segmentation. MISSFormer proposes a multi-scale self-attention module, generating feature maps with differentiated receptive fields through parallel operation of multiple sets of deformable convolution kernels. Channel-wise attention is then used to achieve adaptive feature weighting. Attention U-Net employs a dual-dimensional attention gating mechanism, both spatial and channel-wise, to dynamically suppress irrelevant background regions through learnable parameters, resulting in outstanding performance in the segmentation of small organs such as the pancreas. nnFormer employs an interleaved convolution-attention block design, capturing wide-scale contextual information through dilated convolutions and then refining edge features through local attention, achieving collaborative modeling of volumetric data. D-Former innovatively introduces a dilated Transformer module, embedding a spatially spaced sampling strategy in the attention computation to effectively expand the receptive field and reduce computational complexity. Notably, CoTr proposes a deformable feature interaction mechanism, dynamically adjusting the geometric distribution of the attention region using a learnable offset, enhancing the accuracy of local feature alignment while maintaining global modeling capabilities. BiTr-Unet builds a bidirectional temporal attention mechanism, achieving context-aware optimization of anatomical structures through a forward-backward feature propagation path. Despite significant progress in these hybrid models, two key challenges remain: First, existing methods underexploit the multi-scale feature extraction potential of CNN modules and lack a systematic strategy for designing combinations of convolutional layers with different receptive fields. Second, the computational complexity of the Transformer module, which is proportional to the square of the input resolution, remains a fundamental issue, particularly when processing 3D medical images, which faces the dual pressures of video memory and computational time.

[0011] Although the Transformer architecture has shown significant advantages in the field of medical image segmentation, its inherent Computational complexity remains a significant constraint on its application in high-resolution medical imaging. To address this, researchers have proposed a variety of lightweight improvements. The SOFT model innovatively proposes a softmax-free attention mechanism, replacing the traditional softmax calculation with a Gaussian kernel function. This reduces computational complexity to a linear level while maintaining global modeling capabilities. EfficientVit achieves balanced local-global feature extraction through the collaborative design of linear attention and depthwise separable convolution. Building on this, the FlattenTransformer introduces focused linear attention and a rank recovery module, effectively alleviating feature degradation through a low-rank approximation strategy. The Swin Transformer constructs a hierarchical windowed attention architecture, alternating between local windowed self-attention (W-MSA) and sliding windowed self-attention (SW-MSA), ensuring cross-window information exchange while achieving linear growth in computational complexity. PVT proposes a progressively sparse attention mechanism, which significantly reduces video memory consumption during image processing by gradually reducing the number of key-value pairs in the deep network stage. However, existing lightweight solutions generally face a trade-off between modeling capabilities and computational efficiency. Although the windowing strategy reduces computational overhead, it weakens the ability to model long-range dependencies; the linear approximation method is prone to lead to feature expression degradation; and the sparse processing may lose key anatomical context information.

[0012] Therefore, how to construct a Transformer variant with global modeling capabilities and controllable computational complexity remains a core issue that urgently needs to be broken through in the field of medical image segmentation. Summary of the Invention

[0013] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a cardiac image segmentation method that integrates multi-receptive field convolution and attention shifting.

[0014] This paper proposes a multimodal cardiac image segmentation (MRFTA) model that integrates multi-receptive field convolution and transfer attention for multimodal three-dimensional cardiac image segmentation. This framework significantly improves the segmentation accuracy and computational efficiency of the entire cardiac structure by integrating multi-receptive field feature extraction and efficient global modeling mechanism. This multimodal cardiac image segmentation model adopts a symmetric encoder-decoder architecture and follows the classic U-shaped topology design. The encoder part consists of multiple multi-receptive field deep convolution downsampling modules (MRFDD), which use multi-receptive field deep convolution blocks (MRFDB) to effectively capture local contextual information without increasing computational costs. This paper innovatively proposes a three-dimensional transfer attention Transformer module (TA3D), which introduces learnable transfer tokens T Avoid query vectors Q With key vector K This mechanism reduces the computational complexity from Reduce to , while not destroying the global modeling capability and ensuring the model segmentation accuracy. The decoder part consists of multiple multi-receptive field depth convolution upsampling modules (MRFDU), combined with the multi-scale encoder features transferred by skip connections, to gradually reconstruct the high-resolution segmentation image.

[0015] To achieve the above object, the technical solution adopted by the present invention is: The cardiac image segmentation method integrating multi-receptive field convolution and attention shifting includes the following steps: S1. Preprocessing the original cardiac image training data to generate a preprocessed training data set; S2. Build a multimodal cardiac image segmentation model integrating multi-receptive field convolution and shifted attention; S3. Using the preprocessed training data set to train the constructed multimodal cardiac image segmentation model to obtain a fully trained segmentation model; S4. Use the trained segmentation model to segment the heart image to be tested and output the segmentation result.

[0016] Preferably, step S1 comprises the following steps: S11, mapping non-continuous anatomical structure identifiers in the original cardiac image training data into 8-dimensional binary vectors, and outputting binary label data; S12, performing anatomical structure space normalization processing on the original cardiac image training data and the binary label data, and uniformly adjusting all data to the RAS coordinate system through affine transformation to obtain image data and label data; S13, processing the image data by bilinear interpolation, and processing the label data by nearest neighbor interpolation, and uniformly resampling the spatial resolution of the processed image data and label data to 1 mm × 1 mm × 1 mm to obtain standardized data; S14, linearly scaling the voxel values ​​of the normalized data to the range of [0, 1], and then processing the data using a contrast-limited adaptive histogram equalization algorithm to obtain preprocessed data; S15, extracting a three-dimensional region of interest from the preprocessed data to obtain extracted data; S16. Perform random intensity shift and random space transformation processing on the extracted data to obtain a preprocessed training data set.

[0017] Preferably, in step S2, the constructed multimodal cardiac image segmentation model integrating multi-receptive field convolution and shifted attention comprises a symmetric encoder-decoder architecture based on TransUnet, a linear projection module, a feature mapping module, multiple multi-receptive field depth convolution downsampling modules, multiple multi-receptive field depth convolution upsampling modules, and multiple three-dimensional shifted attention Transformer modules; a plurality of multi-receptive field deep convolution down-sampling modules, which are sequentially arranged in an encoder path of the TransUnet-based symmetric encoder-decoder architecture, are used for progressively extracting local-long distance features, realizing spatial dimension compression and feature abstraction; a linear projection module, which is arranged in a connection layer between an encoder output end and a bottleneck part of the TransUnet-based symmetric encoder-decoder architecture, is used for converting a convolution feature map sequence into a Transformer-compatible vector sequence; a plurality of three-dimensional transfer attention Transformer modules, which are sequentially arranged in the bottleneck part of the TransUnet-based symmetric encoder-decoder architecture, are used for modeling long-range spatial dependencies through a three-dimensional transfer attention mechanism; a feature mapping module, which is arranged in an interface layer between the bottleneck part of the TransUnet-based symmetric encoder-decoder architecture and a decoder path, is used for remapping a Transformer output sequence into a spatial feature map structure; a plurality of multi-receptive field deep convolution up-sampling modules, which are sequentially arranged in the decoder path of the TransUnet-based symmetric encoder-decoder architecture, are used for gradually restoring spatial resolution through deconvolution and skip connection, fusing multi-scale context information to generate a segmentation map.

[0018] Preferably, each multi-receptive field deep convolution down-sampling module comprises two connected multi-receptive field deep convolution blocks, and each multi-receptive field deep convolution block comprises a batch normalization layer, an activation layer and a multi-receptive field deep convolution module connected in sequence.

[0019] Preferably, the working method of each multi-receptive field deep convolution module is as follows: Step one, parallelly using a depth separable convolution with a convolution kernel size of 3 and 5 to extract features from the preprocessed training data set, to obtain two receptive field feature extraction data; Step two, respectively pre-processing the two receptive field feature extraction data through a batch normalization layer and an activation layer in sequence, to output two receptive field regularization feature data; Step three, after feature splicing of the two receptive field regularization feature data, dynamically recombining the topological connection relationship between feature channels through a channel shuffle operation, to promote the interactive fusion of cross-receptive field feature information, to obtain the output of the multi-receptive field deep convolution module.

[0020] Preferably, each three-dimensional transfer attention Transformer module comprises a transfer attention module and a feedforward network connected in sequence. The transfer attention module is used for extracting global features. The feedforward network is used to perform nonlinear transformation and information processing on the extracted global features to obtain the output of the three-dimensional attention transfer transformer module.

[0021] Preferably, the attention shifting module works as follows: Step 1, global aggregation: define a new transfer token as the query, the transfer token is obtained by pooling the original query token, and apply the Softmax attention calculation operation to the transfer token to aggregate the information from each value to obtain the transfer feature; Step 2: Global broadcast: Use the transferred features as values ​​and the transferred tokens as keys to perform a Softmax attention calculation operation with the query matrix, broadcast the global information from the transferred features to each query token, and achieve an effective mapping from global context to local features to obtain global features.

[0022] Preferably, each multi-receptive field deep convolution upsampling module includes a multi-receptive field deep convolution module, a first convolution layer, a first batch normalization layer, a first activation layer, a second convolution layer, a second batch normalization layer and a second activation layer connected in sequence.

[0023] Preferably, in step S3, during the training process of the constructed multimodal cardiac image segmentation model that integrates multi-receptive field convolution and attention shifting, the training process of the multimodal cardiac image segmentation model that integrates multi-receptive field convolution and attention shifting is constrained by a loss function.

[0024] Preferably, the loss function is given by Dice Loss function and TopK Loss function composition; The expression of the loss function is: ; in, is the loss function, for Dice loss function, for TopK Loss function.

[0025] Compared with the prior art, the present invention has the following beneficial effects: (1) The multimodal cardiac image segmentation model established by the present invention integrates multi-receptive field convolution and attention shifting. By deeply integrating CNN with the global self-attention mechanism, it effectively combines the local feature extraction capability of CNN with the global modeling advantage of the self-attention mechanism. It not only improves the accuracy of cardiac multimodal image segmentation, as shown by simulation experimental data, with an average Dice coefficient of 93.65%, a Jaccard coefficient of 93.1%, and a Hausdorff distance of 42.02 mm; it also enhances the model's adaptability to complex anatomical structures, providing new ideas for the migration application of cross-modal medical image tasks. (2) The present invention uses a multi-receptive field deep convolution module to effectively capture information of different scales in the image using convolution kernels with different receptive fields without increasing the computational complexity. This allows the model to not only retain the inherent inductive bias characteristics of convolution, but also significantly enhance the network's ability to model long-distance spatial dependencies, thereby improving segmentation accuracy while maintaining computational efficiency. Simulation experimental data show that the present invention reduces the number of parameters and computational complexity by 11.94% and 21.49% respectively, demonstrating excellent lightweight characteristics. (3) This paper innovatively introduces transfer tokens through the design of a three-dimensional transfer attention Transformer module. T As a query vector Q and key vector K The medium for transmitting information between Q and K Direct similarity calculation avoids the high-complexity operations in the Transformer. This module achieves effective modeling of global information at a low computational cost, solving the problem of excessive resource consumption in 3D medical image processing and making the model more suitable for actual clinical scenarios. (4) The simulation experimental results show that the multimodal cardiac image segmentation model proposed in this paper, which integrates multi-receptive field convolution and attention shifting, significantly outperforms the existing methods in the cardiac multimodal segmentation task, especially in the segmentation accuracy of edge details and complex structures (such as the boundaries of the atria and ventricles), providing a more reliable automated tool for clinical cardiac imaging analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A framework diagram of a multimodal cardiac image segmentation model integrating multi-receptive field convolution and attention shifting constructed by the present invention; Figure 2 Figure 1 is a framework diagram of the existing linear attention mechanism, the Softmax attention mechanism, and the attention transfer module proposed in this invention; (a) is the linear attention mechanism; (b) is the Softmax attention mechanism; (c) is the attention transfer module; Among them, the blue dotted box represents the global aggregation stage, the black dotted box represents the global broadcast stage, and the red dotted box represents the kernel function within the generalized linear attention; Figure 3 A visual comparison chart of the performance of different models; The size of the circle represents the number of floating-point operations per second; Figure 4 A visual comparison of the segmentation results of the MRFTA model proposed in this invention and other methods on the MM-WHS 2017 dataset; Among them, red, green, dark blue, yellow, cyan, lavender and white represent the left ventricular myocardium (Myo), left atrial blood cavity (LA), left ventricular blood cavity (LV), right atrial blood cavity (RA), right ventricular blood cavity (RV), ascending aorta (AA) and pulmonary artery (PA), respectively. DETAILED DESCRIPTION

[0027] The following is a summary of the embodiments of the present invention. Figures 1 to 4 The technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0028] like Figure 1 As shown, an embodiment of the present invention proposes a cardiac image segmentation method integrating multi-receptive field convolution and attention shifting, comprising the following steps: S1. Preprocessing the original cardiac image training data to generate a preprocessed training dataset, specifically comprising the following steps: S11, mapping non-continuous anatomical structure identifiers in the original cardiac image training data into 8-dimensional binary vectors, and outputting binary label data; S12, performing anatomical structure space normalization processing on the original cardiac image training data and the binary label data, and uniformly adjusting all data to the RAS coordinate system through affine transformation to obtain image data and label data; S13, processing the image data by bilinear interpolation, and processing the label data by nearest neighbor interpolation, and uniformly resampling the spatial resolution of the processed image data and label data to 1 mm × 1 mm × 1 mm to obtain standardized data; S14, linearly scaling the voxel values ​​of the normalized data to the range of [0, 1], and then processing the data using a contrast-limited adaptive histogram equalization algorithm to obtain preprocessed data; S15, extracting a three-dimensional region of interest from the preprocessed data to obtain extracted data; S16. Perform random intensity shift and random space transformation processing on the extracted data to obtain a preprocessed training data set.

[0029] S2. Build a multimodal cardiac image segmentation model integrating multi-receptive field convolution and shifted attention; In an embodiment of the present invention, a multimodal cardiac image segmentation model integrating multi-receptive field convolution and attention shifting is constructed, which includes a symmetric encoder-decoder architecture based on TransUnet, a linear projection module, a feature mapping module, multiple multi-receptive field depth convolution downsampling modules, multiple multi-receptive field depth convolution upsampling modules, and multiple three-dimensional attention shifting Transformer modules; Multiple multi-receptive field deep convolutional downsampling modules are sequentially placed in the encoder path of the symmetric encoder-decoder architecture based on TransUnet to progressively extract local-to-long-range features, achieving spatial dimension compression and feature abstraction; The linear projection module is set in the connection layer between the encoder output and the bottleneck part of the symmetric encoder-decoder architecture based on TransUnet, which is used to convert the convolutional feature map sequence into a Transformer-compatible vector sequence; Multiple 3D transfer attention Transformer modules are sequentially placed in the bottleneck part of the symmetric encoder-decoder architecture based on TransUnet to model long-range spatial dependencies through the 3D transfer attention mechanism; The feature mapping module is set in the interface layer between the bottleneck part and the decoder path of the symmetric encoder-decoder architecture based on TransUnet, and is used to remap the Transformer output sequence into a spatial feature map structure; Multiple multi-receptive field deep convolutional upsampling modules are sequentially set in the decoder path of the symmetric encoder-decoder architecture based on TransUnet, which is used to gradually restore the spatial resolution through deconvolution and skip connections, and fuse multi-scale context information to generate segmentation maps.

[0030] Specifically, each multi-receptive field depth convolution downsampling module consists of two connected multi-receptive field depth convolution blocks, and each multi-receptive field depth convolution block consists of a batch normalization layer, an activation layer (i.e., a ReLU6 activation layer), and a multi-receptive field depth convolution module connected in sequence. The activation function of the activation layer is ReLU6.

[0031] Specifically, the working method of each multi-receptive field depth convolution module is as follows: Step 1: For the input preprocessed training data set, use depth-wise separable convolution with kernel sizes of 3 and 5 in parallel to perform feature extraction, and obtain feature extraction data of two receptive fields; Step 2: Preprocess the feature extraction data of the two receptive fields through the batch normalization layer and the activation layer respectively, and output the regularized feature data of the two receptive fields; Step 3: After concatenating the regularized feature data of the two receptive fields, the topological connection relationship between the feature channels is dynamically reorganized through the channel shuffling operation to promote the interactive fusion of cross-receptive field feature information and obtain the output of the multi-receptive field deep convolution module.

[0032] Specifically, each 3D attention-shifting Transformer module consists of a connected attention-shifting module and a feedforward network. The attention-shifting module is responsible for extracting global features, and the feedforward network is responsible for performing nonlinear transformation and information processing on the extracted global features. The linear layer, activation function, and random dropout operation are used to enhance the model's expressiveness and prevent overfitting. Finally, the output of the 3D attention-shifting Transformer module is obtained. The working method of the attention-shifting module is as follows: Step 1, global aggregation: define a new transfer token as the query, the transfer token is obtained by pooling the original query token, and apply the Softmax attention calculation operation to the transfer token to aggregate the information from each value to obtain the transfer feature; Step 2: Global broadcast: Use the transferred features as values ​​and the transferred tokens as keys to perform a Softmax attention calculation operation with the query matrix, broadcast the global information from the transferred features to each query token, and achieve an effective mapping from global context to local features to obtain global features.

[0033] Specifically, each multi-receptive field depth convolution upsampling module includes a multi-receptive field depth convolution module, a first convolutional layer, a first batch normalization layer, a first activation layer (i.e., a first ReLU activation layer), a second convolutional layer, a second batch normalization layer, and a second activation layer (i.e., a second ReLU activation layer), which are connected in sequence. The activation function of the first activation layer and the second activation layer is ReLU.

[0034] S3. Using the preprocessed training data set to train the constructed multimodal cardiac image segmentation model to obtain a fully trained segmentation model; In an embodiment of the present invention, during the training process of the constructed multimodal cardiac image segmentation model that integrates multi-receptive field convolution and shifted attention, the training process of the multimodal cardiac image segmentation model that integrates multi-receptive field convolution and shifted attention is constrained by a loss function.

[0035] Specifically, the loss function is given by Dice Loss function and TopK Loss function composition; The expression of the loss function is: ; in, is the loss function, for Dice loss function, for TopK Loss function.

[0036] S4. Use the trained segmentation model to segment the heart image to be tested and output the segmentation result.

[0037] The technical solution of the present invention is described in detail below: 1. Data preprocessing: To accurately segment multimodal whole-heart structures, preprocessing of the raw cardiac image training data is required. According to international cardiac image segmentation standards, seven target anatomical structures are defined: left ventricular myocardium (Myo), left atrial cavity (LA), left ventricular cavity (LV), right atrial cavity (RA), right ventricular cavity (RV), ascending aorta (AA), and pulmonary artery (PA). The raw labeled data is encoded using seven non-continuous voxel values ​​(205, 420, 500, 550, 600, 820, and 850), with voxel value 0 representing background regions. First, to adapt to the requirements of deep learning model training, a one-hot encoding strategy is used to map the non-continuous anatomical structure identifiers in the raw cardiac image training data into 8-dimensional binary vectors (0, 1, 2, 3, 4, 5, 6, and 7). Secondly, the original image data and binary label data are normalized in anatomical space. The data are uniformly adjusted to the RAS coordinate system through affine transformation to obtain image data and label data. Since the image data is a continuous variable, smooth interpolation, namely bilinear interpolation, is required to generate smoothly transitioned intensity values ​​to maintain the authenticity of the anatomical structure. Since the label data is a discrete category, nearest neighbor interpolation is required to ensure that the binary nature of the segmentation label is not destroyed and to avoid non-integer results. The spatial resolution of the processed image data and label data is uniformly resampled to 1mm×1mm×1mm to obtain standardized data. Subsequently, intensity normalization is performed on the standardized image data, linearly scaling the voxel values ​​to [0,1] to obtain normalized data. To optimize image quality and maximize the extraction of useful information while reducing interference from equipment differences and noise, an improved contrast-limited adaptive histogram equalization (CLAHE) algorithm was used on the normalized data. An 8×8 checkerboard partition was set along the axial slice direction, and the histogram clipping threshold of each sub-region was 2.0. Bilinear interpolation smoothing was used to eliminate block artifacts to obtain contrast-optimized data. In order to extract the key parts of the data, three-dimensional region of interest (ROI) extraction was implemented, and the contrast-optimized data was clipped to a 96×96×96 effective anatomical region to obtain the extracted data. This process removes irrelevant background areas while preserving the complete cardiac structure, significantly improving the efficiency of subsequent processing. Finally, data enhancement operations were performed, including random intensity offset enhancement (applying a global intensity perturbation of ±10% with a probability of 0.5) and random spatial transformation enhancement (around the Z axis). The radian is randomly rotated and the anisotropic scaling is ±10%), and the preprocessed training dataset is finally obtained.

[0038] 2. Build a multimodal cardiac image segmentation model that integrates multi-receptive field networks and attention shifting: like Figure 1As shown in the figure, the constructed multimodal cardiac image segmentation (MRFTA) model integrating multi-receptive field convolution and transfer attention includes a symmetric encoder-decoder architecture based on TransUnet, a linear projection module, a feature mapping module, multiple multi-receptive field depth convolution downsampling modules, multiple multi-receptive field depth convolution upsampling modules and multiple three-dimensional transfer attention Transformer modules; Specifically, multiple multi-receptive field deep convolutional downsampling modules are sequentially set in the encoder path of the symmetric encoder-decoder architecture based on TransUnet to progressively extract local-to-long-range features, achieving spatial dimension compression and feature abstraction; Specifically, the linear projection module is set in the connection layer between the encoder output and the bottleneck part of the symmetric encoder-decoder architecture based on TransUnet, which converts the convolutional feature map sequence into a Transformer-compatible vector sequence; Specifically, multiple 3D transfer attention Transformer modules are sequentially placed in the bottleneck part of the symmetric encoder-decoder architecture based on TransUnet to model long-range spatial dependencies through the 3D transfer attention mechanism; Specifically, the feature mapping module is set in the interface layer between the bottleneck part and the decoder path of the symmetric encoder-decoder architecture based on TransUnet, and remaps the Transformer output sequence into a spatial feature map structure; Specifically, multiple multi-receptive field deep convolutional upsampling modules are sequentially set in the decoder path of the symmetric encoder-decoder architecture based on TransUnet. They gradually restore the spatial resolution through deconvolution and skip connections, and fuse multi-scale context information to generate segmentation maps.

[0039] 2.1, Encoder: In order to fully utilize the advantages of CNN in extracting local features and Transformer in modeling long-range dependencies, the network uses multi-layer 3D CNN to extract local features at the bottom layer of the encoder, and uses 3D Transformer to deeply model the global relationship between features at the deep layer. Specifically, when the model receives image data ( ) input, multiple multi-receptive field deep convolution downsampling modules (MRFDD) are used to analyze the spatial and depth information of the original data, and a 3×3×3 convolution with a stride of 2 is applied for downsampling to generate a feature map with rich local information. , , where K = 128. To ensure a comprehensive representation of each volume, the generated feature maps Through the linear projection module, that is, using the convolution layer to perform linear projection, the channel dimension is increased from K=128 to k=512. Since the Transformer layer requires a sequence as input, the present invention compresses the spatial dimension and the depth dimension into one dimension to obtain a feature map , ,in, In order to preserve important position information during the segmentation process, this paper introduces a learnable position embedding , and add it to the feature map by direct addition Fusion to get feature embedding , feature map and feature embedding The expression is as follows: (1) (2) in, is the linear projection operation, is the feature map, Embed for position.

[0040] Finally, Input A low-complexity 3D Transfer Attention Transformer module (TA3D) captures long-range spatial dependencies and generates the final feature embedding .

[0041] 2.1.1、Multi-receptive field depth convolution downsampling module (MRFDD): To better extract multi-scale feature information and rich contextual relationships in images while improving model flexibility and efficiency, this paper introduces a multi-receptive field deep convolutional downsampling module (MRFDD) in the encoder. This module extracts multi-scale information by utilizing convolution kernels of varying sizes and reduces resource consumption during multiple convolutions through depthwise convolution. Each MRFDD contains two multi-receptive field deep convolutional blocks (MRFDB). In this paper's MRFDB, batch normalization and ReLU6 activation layers are first used for preprocessing. Next, a multi-receptive field deep convolutional module (MRFD) is used to capture multi-scale and multi-resolution contextual information. Specifically, this paper first uses depthwise separable convolutions with kernel sizes of 3 and 5 in parallel for feature extraction, followed by batch normalization and ReLU6 activation layers, respectively. Finally, because depthwise convolution ignores channel relationships, a channel shuffling operation is performed to reintegrate inter-channel correlations. The expressions for the MRFDB, MRFD, and deep convolutional blocks in this embodiment of the present invention are shown below: (3) (4) (5) in, is a multi-receptive field depth convolution block, It is a multi-receptive field deep convolution module. and are batch normalization and ReLU6 activation, respectively. For feature splicing, for channel shuffling, is a depthwise convolution with a kernel size of s, is the input feature map.

[0042] 2.1.2. 3D Transfer Attention Transformer Module (TA3D): In order to reduce the computational complexity of the model without compromising the long-range modeling capability, this paper introduces a three-dimensional transfer attention Transformer module (TA3D). TA3D consists of a transfer attention module (TA) and a feed-forward network (FFN). The output of a TA3D can be expressed by the following formula: (6) (7) (8) in, represents the attention transfer module, Representation layer normalization, represents a feedforward network, For the In the Transformer layer, the intermediate feature representation after being processed by the attention transfer module and added with residual connection, For the The output of the Transformer layer, For the The output of the Transformer layer, is the current layer index in the Transformer module stack, is the random dropout operation, is the Gaussian error linear unit activation function, and is two linear projections.

[0043] (1) Transfer Attention Module (TA): The root cause of the high complexity of the Transformer is the softmax attention in the Transformer, which has a computational complexity of , as shown in (b) of Figure 2 . Linear attention, which has a computational cost of , is generally considered to have much less model expressiveness than softmax attention, as shown in (a) of Figure 2 . Most people consider these two attention mechanisms as independent methods, and relevant researchers either try to reduce the computational cost of softmax attention or try to improve the performance of linear attention. However, according to most current research results, the existing linear attention with optimized performance still cannot match the model expressiveness of softmax attention, and the existing methods for reducing the computational cost of softmax attention have certain loss in long-range modeling ability. Therefore, the present application introduces a new attention mechanism, namely transfer attention, which skillfully combines the advantages of softmax attention and linear attention, maintains linear computational complexity, and does not sacrifice model expressiveness, as shown in (c) of Figure 2 .

[0044] Specifically, the expressions of softmax attention and linear attention are as follows: (9) (10) wherein, denotes the softmax attention, denotes the linear attention, denotes the query, key and value respectively, denotes the dimension of the key, denotes the kernel function.

[0045] The transfer attention proposed in the present application effectively combines softmax attention and linear attention. Specifically, by introducing a new transfer token T , the similarity calculation between and is avoided, and the expression of the transfer attention is as follows: (11) Specifically, the transfer attention mechanism of the present application is divided into two stages, a global aggregation stage and a global broadcast stage. First, in the global aggregation stage, a new transfer token T is defined as the query, and the softmax attention calculation is applied to it to aggregate information from each value V to obtain the transfer feature . Then, in the global broadcast phase, the features are transferred As a value, transfer token T As key, with the query matrix A second Softmax attention calculation is performed to broadcast the global information from the transfer features to each query token, achieving an effective mapping from global context to local features.

[0046] Diverting attention by introducing transfer tokens T , successfully bypassed and The need for pairwise similarity calculation between two nodes. Set the number of transfer tokens n As a hyperparameter, we achieve The linear complexity of the model is kept, while effectively maintaining the model's long-range modeling capabilities. In terms of acquiring transfer tokens, the present invention adopts a pooling strategy. Furthermore, the transfer attention establishes a generalized linear attention paradigm through two Softmax attention calculation operations, where the two Softmax similarity calculations serve as the kernel function.

[0047] 2.2 Decoder In order to adapt to the input dimension of the 3D CNN decoder, the present invention first applies a feature mapping module to project the sequence data into a standard feature map, and uses a convolution block to convert the channel dimension from k Reduce to K Specifically, the output sequence of Transformer Reshape into feature map After that, the feature maps generated by the encoder are processed using multiple multi-receptive field depth convolution upsampling modules (MRFDU) and 3×3×3 convolution operations. Perform detailed feature extraction and upsampling to generate refined segmentation results In addition, skip connections are applied to fuse encoder and decoder features.

[0048] Multi-receptive field depth convolution upsampling module (MRFDU): To improve the model's performance in generating high-resolution images while maintaining computational efficiency, this paper introduces a multi-receptive field deep convolutional upsampling module (MRFDU). In the MRFDU, multi-scale features are first extracted using MRFD (the MRFD process in this section is the same as that in MRFDD). This is followed by two rounds of convolutional layers, batch normalization layers, and ReLU activation layers. The MRFDU formula is as follows: (12) (13) in, is a multi-receptive field depthwise convolution up-sampling module, is a multi-receptive field depthwise convolution module, is a normal convolution, and are batch normalization and ReLU activation, respectively.

[0049] 3. Training a multi-modal cardiac image segmentation model that integrates multi-receptive field convolution and shift attention: In order to effectively train the model proposed in the present application, the present application combines two complementary loss functions in a ratio of 1:1, which are Dice loss function and TopK loss function. Dice The loss function is used to alleviate the problem of data imbalance, while TopK the loss function focuses on the performance of the model on the most challenging samples.

[0050] Dice The definition of the loss function is: (14) wherein, is the loss function, Dice and are the predicted result and the true result of the i-th sample, respectively, is the total number of samples, is a very small constant to avoid division by zero. N The loss function focuses on the most challenging

[0051] samples, and TopK The definition of the loss function is: TopK (15) wherein, is the loss function, is the predicted result of the i-th sample (corresponding to the j-th most difficult segmentation sample), is the number of most difficult samples considered. TopK Therefore, the final loss function definition of the present application is: (16) K wherein, is the loss function,

[0052] is the loss function, is the loss function. Dice TopK ​​​​​​​

[0053] To validate the effectiveness and feasibility of the proposed method, we trained and tested it on the CT and MRI modalities of the MM-WHS 2017 dataset and compared it with other methods. This dataset, an open-source dataset provided by the 2017 Multimodal Whole Heart Challenge, contains 60 cardiac CT images and 60 cardiac MRI images, covering all cardiac substructures. The goal of this challenge is to accurately segment multimodal whole-heart structures, including the left ventricular myocardium (Myo), left atrial cavity (LA), left ventricular cavity (LV), right atrial cavity (RA), right ventricular cavity (RV), ascending aorta (AA), and pulmonary artery (PA). Since only 20 sets of publicly available annotated data were available, to avoid overfitting, we borrowed the pseudo-label generation method from semi-supervised training. Using the 3D UX-NET network, we generated pseudo-labels for 40 unlabeled images. Thirty-two of these pseudo-labeled images served as the training set, and eight served as the validation set. Finally, the 20 publicly annotated images served as the test set.

[0054] The effects of the present invention can be further illustrated by the following simulation experiments: 1. Simulation conditions: This paper simulates the problem using the Pytorch 1.13.1 framework on an Intel® Xeon® Platinum 8260Y CPU running at 2.40 GHz, 24 GB of GPU memory per GPU, an NVIDIA GeForce RTX 3090 GPU, and the Ubuntu 18.04.1 operating system. The experimental code is written in Python 3.8.

[0055] 2. Simulation details: During the training process, the AdamW algorithm is used with a batch size of 2, a weight decay of 0.0005, an initial learning rate of 0.0001, a maximum of 40,000 iterations, and the best performing model is saved based on the Dice coefficient. In the multi-receptive field deep convolution module, the convolution kernel size is set to [3,5]. In addition, in TA3D, the number of transfer tokens is set to [3,5]. n Set to 64.

[0056] 3. Simulation content: In order to fully illustrate the feasibility and effectiveness of the present invention, the simulation content includes whole-heart structure segmentation experiments, comparative experiments, and ablation experiments on the CT modality and MRI modality of the MM-WHS 2017 dataset.

[0057] 4. Evaluation criteria: The present invention uses three common indicators to evaluate the performance of MRFTA: Dice coefficient, Jaccard coefficient and Hausdorff distance.

[0058] The Dice coefficient is an indicator that measures the similarity between two sets and is often used to evaluate the accuracy of image segmentation. Its calculation formula is: (17) in, A and B Represent the predicted segmentation results and the true segmentation results respectively, Dice The coefficient ranges from 0 to 1. The closer the value is to 1, the higher the overlap between the segmentation result and the true result, and the better the segmentation effect.

[0059] The Jaccard coefficient, also known as the intersection over union (IoU), is another commonly used segmentation evaluation metric. Its calculation formula is: (18) in, A and B They represent the predicted segmentation results and the true segmentation results respectively.

[0060] The Jaccard coefficient also ranges from 0 to 1. The closer the value is to 1, the higher the overlap between the segmentation result and the true result. Compared with the Dice coefficient, the Jaccard coefficient is more sensitive to the boundaries of the segmentation result.

[0061] The Hausdorff distance is an indicator that measures the maximum mismatch between two point sets and is often used to evaluate the boundary accuracy of segmentation results. It is defined as: (19) in, represents the Hausdorff distance, D Indicates a point a and point b The smaller the Hausdorff distance is, the closer the boundary of the segmentation result is to the true boundary.

[0062] 5. Experimental results: (1) Comparative experimental results: The performance comparison results of different models on the MM-WHS 2017 CT dataset are shown in Table 1 below.

[0063] Table 1 Performance comparison of different models on the MM-WHS 2017 CT dataset: ; The results in Table 1 show that the proposed MRFTA model achieved state-of-the-art results across all three performance metrics, with an average Dice coefficient of 93.65%, a Jaccard coefficient of 93.1%, and a Hausdorff distance of 42.02 mm. Furthermore, the MRFTA model achieved the best Dice coefficient for each cardiac substructure.

[0064] The performance comparison results of the MRFTA model and other methods on the MM-WHS 2017 MRI dataset are shown in Table 2 below.

[0065] Table 2 Performance comparison of different models on the MM-WHS 2017 MRI dataset: ; It can be seen from the results in Table 2 that although the MRFTA model did not achieve the optimal Jaccard coefficient, it achieved the optimal average Dice coefficient of 81.21% and the minimum Hausdorff distance of 32.94 mm. In addition, the MRFTA model also obtained the optimal Dice coefficients for the five cardiac substructures of Myo, LA, LV, RV and PA. Moreover, although the SwinUNETR model achieved suboptimal segmentation results on the CT dataset, its performance on the MRI dataset was not satisfactory. In contrast, the MRFTA model of the present invention achieved optimal results on both CT and MRI datasets, demonstrating its better robustness. Compared with the CT dataset, the superiority of the MRFTA model on MRI images with greater noise further demonstrates its stronger generalization ability.

[0066] like Figure 3 As shown in Figure 3, although the VT-Unet model is highly efficient, its segmentation accuracy is far inferior to that of the MRFTA model of the present invention. Compared with other existing methods, the MRFTA model of the present invention achieves the highest Dice coefficient while maintaining a relatively small model size and low computational complexity.

[0067] (2) Experimental results of whole heart structure segmentation: Figure 4 The whole heart segmentation results of the MRFTA model and other segmentation models are intuitively displayed. Figure 4 The generated slice segmentation images and 3D segmentation images show that the segmentation results of the MRFTA model of the present invention are closer to the actual situation. In particular, the MRI segmentation results clearly show that the segmentation results of the MRFTA model are smoother.

[0068] (3) Ablation experiment results: 1. Ablation study of convolution kernel size in MRFD: In order to more effectively optimize the receptive field size in MRFD, the present invention conducted ablation experiments with different convolution kernel sizes, and the results are shown in Table 3 below.

[0069] Table 3 Ablation effects of MRFD with different convolution kernel sizes on the MM-WHS 2017 CT dataset: ; According to the results in Table 3, using a single receptive field for depthwise separable convolution results in poor segmentation performance. Therefore, we choose to use a combination of receptive fields of different sizes to capture image features at multiple scales. After experimenting with various combinations, we determined that the optimal convolution kernel size is [3, 5]. This choice ensures optimal segmentation performance while keeping model complexity manageable.

[0070] 2. Ablation study of MRFD location: The present invention conducted further experiments to explore the position of MRFD in the model, and the results are shown in Table 4 below, where the first group is the model obtained by replacing all MRFD in MRFDD with ordinary convolution with a convolution kernel size of 3, and deleting MRFD in MRFDU, the second group is the model obtained by replacing MRFD in MRFDD with ordinary convolution with a convolution kernel size of 3, the third group is the model obtained by deleting MRFD from MRFDU, the fourth group is the model obtained by adding ordinary convolution with a convolution kernel size of 3 after each MRFD on the basis of MRFDD, and the fifth group is the MRFTA proposed in the present invention.

[0071] Table 4. Location of MRFD in the model structure. Ablation experiment on the MM-WHS 2017 CT dataset. ; The results in Table 4 show that using MRFD alone in either the encoder or decoder does not achieve optimal performance. Directly using MRFD in the encoder instead of traditional convolution for feature extraction can even cause model performance to degrade. However, when supplemented with conventional convolution and MRFD to form MRFDU, optimal performance can be achieved. Furthermore, the results of Group 4 demonstrate the superiority of using MRFDD, which combines MRFD with MRFDU instead of conventional convolution, and confirm the effectiveness of the MRFTA model structure proposed in this paper.

[0072] 3. Ablation study of different modules in MRFTA: To verify the effectiveness of the MRFTA architecture, we conducted an ablation experiment on its core module on the MM-WHS 2017 CT dataset. The results are shown in Table 5 below.

[0073] Table 5 Ablation study results of different modules of MRFTA on the MM-WHS 2017 CT dataset: ; The results in Table 5 show that both TA3D and MRFD improve the model's segmentation performance. Although TA3D slightly increases the model's parameter count, it also reduces the model's FLOPs by 9.36%. MRFD, through structural optimization, successfully reduces the number of parameters and computational complexity by 11.94% and 21.49%, respectively, demonstrating its exceptional lightweight nature. Experimental data demonstrates that each module achieves a synergistic optimization of parameter efficiency and computational resource consumption while improving segmentation accuracy.

[0074] 4. Ablation study of different attention mechanisms: To verify the effectiveness of TA3D, we conducted ablation experiments with different attention mechanisms. The results are shown in Table 6 below.

[0075] Table 6 Performance evaluation results of different attention mechanisms on the MM-WHS 2017 CT dataset: ; As shown in the experimental results in Table 6, the multi-head self-attention mechanism based on Softmax attention achieves higher segmentation accuracy than Softmax-free attention. It is particularly noteworthy that the transfer attention mechanism proposed in this invention achieves optimal performance in all three indicators: Dice coefficient, Jaccard coefficient, and Hausdorff distance, highlighting the significant advantages of global context modeling. In terms of computational efficiency, TA3D achieves the highest segmentation accuracy while maintaining a computational complexity comparable to that of linear attention. These experimental results fully demonstrate that TA3D achieves Pareto optimality between model accuracy and computational efficiency through its innovative attention transfer strategy.

[0076] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A cardiac image segmentation method integrating multi-receptive field convolution and attention shifting, characterized in that: The following steps are involved: S1. Preprocessing the original cardiac image training data to generate a preprocessed training data set; S2. Build a multimodal cardiac image segmentation model integrating multi-receptive field convolution and shifted attention; The constructed multimodal cardiac image segmentation model integrating multi-receptive field convolution and shifted attention includes a symmetric encoder-decoder architecture based on TransUnet, a linear projection module, a feature mapping module, multiple multi-receptive field deep convolution downsampling modules, multiple multi-receptive field deep convolution upsampling modules, and multiple 3D shifted attention Transformer modules; Multiple multi-receptive field deep convolutional downsampling modules are sequentially set in the encoder path of the symmetric encoder-decoder architecture based on TransUnet; A linear projection module is placed in the interface between the encoder output and the bottleneck of the TransUnet-based symmetric encoder-decoder architecture; Multiple 3D attention-shifting Transformer modules are sequentially placed in the bottleneck of a symmetric encoder-decoder architecture based on TransUnet; The feature mapping module is placed at the interface layer between the bottleneck part and the decoder path of the symmetric encoder-decoder architecture based on TransUnet; Multiple multi-receptive field depth convolutional upsampling modules are sequentially set in the decoder path of the symmetric encoder-decoder architecture based on TransUnet; S3. Using the preprocessed training data set to train the constructed multimodal cardiac image segmentation model to obtain a fully trained segmentation model; S4. Use the trained segmentation model to segment the heart image to be tested and output the segmentation result.

2. The cardiac image segmentation method integrating multi-receptive field convolution and attention shifting according to claim 1, characterized in that: Step S1 includes the following steps: S11, mapping non-continuous anatomical structure identifiers in the original cardiac image training data into 8-dimensional binary vectors, and outputting binary label data; S12, performing anatomical structure space normalization processing on the original cardiac image training data and the binary label data, and uniformly adjusting all data to the RAS coordinate system through affine transformation to obtain image data and label data; S13, processing the image data by bilinear interpolation, and processing the label data by nearest neighbor interpolation, and uniformly resampling the spatial resolution of the processed image data and label data to 1 mm × 1 mm × 1 mm to obtain standardized data; S14, linearly scaling the voxel values ​​of the normalized data to the range of [0, 1], and then processing the data using a contrast-limited adaptive histogram equalization algorithm to obtain preprocessed data; S15, extracting a three-dimensional region of interest from the preprocessed data to obtain extracted data; S16. Perform random intensity shift and random space transformation processing on the extracted data to obtain a preprocessed training data set.

3. The cardiac image segmentation method integrating multi-receptive field convolution and attention shifting according to claim 1, characterized in that: Each multi-receptive field depth convolution downsampling module contains two connected multi-receptive field depth convolution blocks, and each multi-receptive field depth convolution block contains a batch normalization layer, an activation layer and a multi-receptive field depth convolution module connected in sequence.

4. The cardiac image segmentation method integrating multi-receptive field convolution and attention shifting according to claim 3, characterized in that: Each multi-receptive field depth convolution module works as follows: Step 1: For the input preprocessed training data set, use depth-wise separable convolution with kernel sizes of 3 and 5 in parallel to perform feature extraction, and obtain feature extraction data of two receptive fields; Step 2: Preprocess the feature extraction data of the two receptive fields through the batch normalization layer and the activation layer respectively, and output the regularized feature data of the two receptive fields; Step 3: After concatenating the regularized feature data of the two receptive fields, the topological connection relationship between the feature channels is dynamically reorganized through the channel shuffling operation to promote the interactive fusion of cross-receptive field feature information and obtain the output of the multi-receptive field deep convolution module.

5. The cardiac image segmentation method integrating multi-receptive field convolution and attention shifting according to claim 1, characterized in that: Each 3D attention-shifting Transformer module contains a connected attention-shifting module and a feed-forward network; The attention transfer module is used to extract global features; The feedforward network is used to perform nonlinear transformation and information processing on the extracted global features to obtain the output of the three-dimensional attention transfer transformer module.

6. The cardiac image segmentation method integrating multi-receptive field convolution and attention shifting according to claim 5, characterized in that: The attention shifting module works as follows: Step 1, global aggregation: define a new transfer token as the query, the transfer token is obtained by pooling the original query token, and apply the Softmax attention calculation operation to the transfer token to aggregate the information from each value to obtain the transfer feature; Step 2: Global broadcast: Use the transferred features as values ​​and the transferred tokens as keys to perform a Softmax attention calculation operation with the query matrix, broadcast the global information from the transferred features to each query token, and achieve an effective mapping from global context to local features to obtain global features.

7. The cardiac image segmentation method integrating multi-receptive field convolution and attention shifting according to claim 1, characterized in that: Each multi-receptive field depth convolution upsampling module includes a multi-receptive field depth convolution module, a first convolution layer, a first batch normalization layer, a first activation layer, a second convolution layer, a second batch normalization layer and a second activation layer connected in sequence.

8. The cardiac image segmentation method integrating multi-receptive field convolution and attention shifting according to claim 1, characterized in that: In step S3, during the training process of the constructed multimodal cardiac image segmentation model integrating multi-receptive field convolution and shifted attention, the training process of the multimodal cardiac image segmentation model integrating multi-receptive field convolution and shifted attention is constrained by a loss function.

9. The cardiac image segmentation method integrating multi-receptive field convolution and attention shifting according to claim 8, characterized in that: The loss function is given by Dice Loss function and TopK Loss function composition; The expression of the loss function is: ; in, is the loss function, for Dice loss function, for TopK Loss function.

Citation Information

Patent Citations

  • Pollen image classification method based on cross attention distillation Transformer

    CN113887610A

  • Medical image automatic segmentation method of U-shaped network based on fusion convolution and attention mechanism

    CN117474866A

  • CNN and attention-based MRI brain tumor segmentation method

    CN117764953A

  • Transform-CNN medical image segmentation method and system based on multi-scale fusion semantic enhancement

    CN120318256A

Cited By

  • Muscle segmentation method based on state space model and boundary perception attention

    CN121213594A

  • Brain tumor image segmentation method based on dynamic token processing and multi-scale pyramid attention and related equipment

    CN121482072A

  • Multi-modal detection model and method for small-scale unmanned aerial vehicle

    CN121861525A

  • Double-branch Laplace gated graph polymerization multi-mode meningioma segmentation method

    CN122116088A