Implementation Method of Cross-Layer Constrained Transformer Network Based on Hierarchical Feature Compensation
By introducing an efficient cross-attention mechanism in medical image segmentation, the problems of loss of detail information and lack of dependency modeling during downsampling are solved, and a more accurate and stable medical image segmentation effect is achieved.
Patent Information
- Application Number
- CN202510221564.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing medical image segmentation methods are prone to loss of detailed information during downsampling, and traditional CNN models lack modeling of dependencies between features when feature fusion, making it difficult to make full use of low-dimensional information to optimize deep feature expression.
A cross-layer constraint Transformer network CCFormer based on hierarchical feature compensation is proposed. By introducing an efficient cross-attention mechanism, the deep features are constrained by using shallow features to ensure that key low-dimensional details are retained during downsampling and upsampling.
CCFormer effectively retains low-level detailed features in medical imaging segmentation tasks, optimizes the expression of deep features, and significantly improves the segmentation ability of complex anatomical structures and detailed features, especially in high-precision edge recognition and fine-grained segmentation.
Smart Images

Figure CN119723296B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and particularly to a method for implementing a cross-layer constrained Transformer network based on hierarchical feature compensation, CCFormer (Cross-Layer Constraint Transformer). CCFormer introduces efficient cross-attention with linear computational complexity, and retains key low-dimensional details during downsampling by imposing constraints from shallow features representing boundaries and textures on deep features representing semantic and structural information. In addition, CCFormer uses efficient cross-attention to optimize the feature fusion between the encoder's low-level features and the decoder's high-level features, ensuring that key detail information is retained during upsampling. This network can be used for various medical image processing tasks such as detecting and analyzing diabetic retinopathy, malignant skin lesions, etc., providing accurate and efficient auxiliary support for clinical diagnosis and treatment. Background Art
[0002] In the field of semantic segmentation, medical image segmentation faces more challenges compared to natural image segmentation due to the complex and variable anatomical structures, rich important details, common low contrast and high noise in medical images, extremely imbalanced ratio between the segmented region and the background region, and scarce labeled data. It is difficult to achieve ideal results using traditional natural image segmentation methods. Convolutional Neural Network (CNN) processes short-term and long-term image features by stacking convolutional operations with different receptive fields, thereby extracting local information of the image and capturing pixel-level fine-grained features. This ability is very effective for detecting and identifying small-sized lesions and subtle anatomical structure changes in medical imaging. The shared weights and local connection characteristics of the convolutional kernel endow it with translational invariance, which is crucial for capturing local texture and edge information in medical images. In addition, the multi-level architecture of CNN can gradually extract high-level semantic features from low-level features such as edges and textures, thereby obtaining features at different scales. This is of great significance for identifying complex anatomical structures and lesions and organs with different sizes and shapes.
[0003] U-Net uses a symmetric encoder-decoder structure and introduces skip connections, showing good segmentation results on various medical images such as MRI, CT, and microscopic images, especially in small-sample data. Many improved versions of U-Net have emerged, such as UNet++, UNet3+, Attention U-Net, and nnUNet, etc. These improved versions further enhance the performance and generality of CNN architectures like U-Net in various medical image segmentation tasks. In medical image segmentation tasks, CNN-based architectures have excellent abilities in recognizing detailed pathological structures and extracting features, while Transformer-based architectures can capture long-range dependencies through multi-head self-attention for global feature modeling. Therefore, hybrid Transformer-CNN methods that combine CNN and Transformer have emerged to utilize the advantages of both. TransUNet uses CNN to extract low-level features and then encodes the features into context sequences through Transformer for global modeling. TransFuse combines Transformer and CNN in a parallel manner, capable of capturing global dependencies and low-level spatial details at a shallower layer. Compared with pure Transformer architectures, hybrid transformer-CNN methods have better detailed feature modeling capabilities. However, current hybrid transformer-CNN methods only limit the fusion of CNN and Transformer to using multi-head self-attention for feature modeling at each fixed stage in the CNN architecture, without considering the connections between the low-level edge, texture, color information represented by the shallow features of the previous stage and the high-level semantic and structural information represented by the deep features of the next stage during the downsampling process, as well as the complementary and enrichment relationship between the shallow features from the encoder and the deep features in the decoder in the skip connections. CNN-based segmentation architectures often supplement the deep features in the decoder with edge, texture, color, etc. information from the shallow features of the encoder through skip connections, but this simple feature fusion lacks modeling of the dependencies between different features and cannot effectively constrain the deep features with the shallow features.
[0004] To address the above challenges, a method is needed to restrict feature selection in the downsampling process and skip connections, preserving features helpful for fine-grained segmentation. To this end, the present invention designs a method using an attention mechanism to achieve feature expression that restricts deep features through shallow features, enabling deep features to take into account low-dimensional details while representing high-dimensional semantic information. Further, the present invention proposes CCFormer with efficient cross-attention having linear computational complexity, using efficient cross-attention to perform feature information constraint on shallow and deep features in the encoder and skip connections, and using shallow information (edges, colors, texture materials) to guide the deep features representing semantics, structures, and local details to express the jointly important parts in the shallow and deep layers, compensating to a certain extent for the information loss caused by convolution in reducing the feature size, and achieving high-performance segmentation with low parameter numbers and computational complexity in both 2D and 3D tasks.
[0005] The CNN-based segmentation network extracts deep high-dimensional spatial features through stacking pooling layers and convolutions for downsampling and supplements edge and texture information for decoder deep features through skip connections. However, the simple information fusion of different features in the skip connection lacks modeling of the dependency relationship between features and is difficult to fully utilize low-dimensional information to optimize deep feature expression. The present invention improves the accuracy of medical image semantic segmentation through efficient cross-attention, using shallow features with richer details such as high-resolution edges and textures in the downsampling stage to guide the low-resolution semantic features in the upsampling process, increasing the attention degree of deep features to important edge and detail information and enhancing the global context modeling ability of the model. Similarly, the present invention uses a cross-layer constraint method for deep features representing category semantics, structural information, and local details guided by shallow features representing image edges, colors, and textures in the downsampling process. Based on this, the present invention proposes CCFormer, which is a linear inter-layer cross-attention Transformer that can make the most of the feature information at each stage in the downsampling and upsampling processes, restricting the expression of deep high-level features through low-level features in the downsampling and supplementing sufficient detail information such as texture and edges for the process of restoring image information in the upsampling. Summary of the Invention
[0006] The object of the present invention is to provide a method for implementing a cross-layer constrained Transformer network based on hierarchical feature compensation in view of the deficiencies of the prior art. The CNN-based segmentation network extracts deep high-dimensional spatial features through downsampling by stacking pooling layers and convolutions and supplements edge and texture information for the decoder deep features through skip connections. However, the simple information fusion of different features in the skip connections lacks the modeling of the dependence relationship between features and it is difficult to make full use of low-dimensional information to optimize the deep feature expression. The present invention improves the accuracy of medical image semantic segmentation through efficient cross-attention, uses shallow features with richer details such as high-resolution edges and textures in the downsampling stage to guide the low-resolution semantic features in the upsampling process, improves the attention degree of deep features to important edge and detail information, and enhances the global context modeling ability of the model. Similarly, the present invention uses shallow features representing image edges, colors, and textures in the downsampling process to guide the cross-layer constraint method of deep features representing category semantics, structural information, and local details. Based on this, the present invention proposes CCFormer, which is a linear layer cross-attention Transformer network. By introducing a Cross-Layer Constraint in Encoder (CCE) module and a Cross-Layer Constraint in Decoder (CCD) module, it can make the most of the feature information at each stage in the downsampling and upsampling processes, restrict the deep high-level feature expression through the downsampled low-level features, and supplement sufficient texture and edge and other detail information for the process of restoring image information in the upsampling. The present invention can achieve more accurate and stable segmentation effects in a variety of 2D and 3D medical image segmentation tasks, showing its broad application prospects in the field of medical image segmentation.
[0007] The technical solutions included in the present invention to solve its technical problems are as follows:
[0008] Step 1: Collect medical image segmentation data sets for multiple different segmentation tasks;
[0009] Step 2: Divide the medical image segmentation data set, and divide the data into a training set, a validation set, and a test set according to a ratio of 8:1:1;
[0010] Step 3: Perform data preprocessing;
[0011] Step 4: Perform data augmentation on the data in Step 3;
[0012] Step 5: Build a neural network for medical image segmentation;
[0013] Step 6: Model training, inference, and saving the model weights with the best performance.
[0014] Step 7: Model performance evaluation and result analysis.
[0015] Advantages of the present invention
[0016] Considering that traditional CNN models are prone to losing detailed information during downsampling, while the Transformer model can capture long-range dependencies but may also ignore spatial details. The CCFormer architecture proposed in the present invention adopts an efficient cross-layer feature constraint mechanism, which can effectively retain low-level detailed features in medical image segmentation tasks while optimizing the expression of deep features. By constraining low-level features such as boundaries and textures, the segmentation ability for complex anatomical structures and detailed features is effectively improved, especially in high-precision edge recognition and fine-grained segmentation. CCFormer not only achieves significant performance improvement but also effectively reduces computational complexity and video memory consumption. By introducing an efficient cross-attention mechanism with linear complexity, it maintains low computational resource requirements without sacrificing segmentation accuracy, making this method more suitable for medical image segmentation tasks with high computational resource requirements in practical applications. The CCFormer proposed in the present invention demonstrates good performance in various medical image segmentation tasks, showing the generalization ability and applicability of medical image segmentation models, greatly improving the accuracy and efficiency of clinical diagnosis and treatment, making important contributions to the development of medical imaging technology, and thus effectively improving the treatment effect of patients. Brief description of the drawings
[0017] Figure 1 It is the overall framework structure diagram of 3D CCFormer.
[0018] Figure 2 It is the structure diagram of the efficient cross-attention module.
[0019] Figure 3 It is the structure diagram of the cross-layer constraint encoder module for 3D data.
[0020] Figure 4 It is the structure diagram of the cross-layer constraint decoder module for 3D data.
[0021] Figure 5 It is the overall framework structure diagram of 2D CCFormer.
[0022] Figure 6 It is the schematic diagram of the module of 2D data in the cross-layer constraint encoder.
[0023] Figure 7 It is the schematic diagram of the module of 2D data in the cross-layer constraint decoder. Detailed implementation manners
[0024] The present invention will be further explained and illustrated below in conjunction with the accompanying drawings and embodiments.
[0025] An implementation method of a cross-layer constrained Transformer network based on hierarchical feature compensation, the structure is as Figure 1-4 shown, aiming to segment medical images, which specifically includes the following steps:
[0026] Step 1: Collect medical image segmentation data sets for multiple different segmentation tasks: ISIC2017, ISIC2018, DSB-18, OIMHS, and SegTHOR, covering different types of medical images including skin cancer, abdominal CT, retinal OCT, and pathological section images.
[0027] Step 2: Divide the data set. Divide the data into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0028] Step 3: Perform data preprocessing. Since there are differences between 2D and 3D data, different processing is required. For 2D data, scale the images to a unified size and normalize their pixel values to the range [0, 1].
[0029] For 3D data, randomly crop a block of [96, 96, 96] and input it into the designed segmentation network. For the cropped data, scale it to between [0, 1] using max-min normalization.
[0030] Step 4: Perform data augmentation on the data preprocessed in Step 3. Use random flipping, random intensity transformation, random flipping, and random scaling to improve data diversity, which helps to improve the generalization of the model.
[0031] Step 5: In order to achieve high-precision medical image segmentation, the present invention proposes a segmentation network (CCFormer), which includes three modules: an Efficient Cross-Attention module, a Cross-Layer Constraint in Encoder (CCE) module, and a Cross-Layer Constraint in Decoder (CCD) module. The structure and functions of CCFormer and its modules are introduced in detail below, and relevant mathematical formulas are given.
[0032] The overall structure of CCFormer is as Figure 1As shown below. First, the input image is transformed through a convolution operation to expand the channel depth, thereby achieving preliminary feature extraction. To reduce overfitting and enhance the generalization ability of the model, a DropBlock layer is applied after the preliminary convolution. Downsampling is achieved by combining max pooling and average pooling, and the channel depth is adjusted through a convolution operation, with a sampling factor of 2. The upsampling path is mirror-executed through a combination of transposed convolution and two layers of convolution, also achieving an upsampling with a sampling factor of 2. Subsequently, Efficient Cross-Attention is applied in the cross-layer constrained encoder module CCE. During this process, shallow features help guide the focus of deep representations. The weighted mapping is restricted to the range (0, 1) through the Sigmoid activation function and then applied to the deep feature representation.
[0033] In the upsampling stage, Efficient Cross-Attention is applied in the cross-layer constrained decoder module CCD; the encoder features and decoder features are adaptively combined through skip connections to form a fused feature, and these fused features are supported by context in terms of structural and texture characteristics. In all convolutional layers in the downsampling and upsampling stages, the LeakyReLU activation function is applied and combined with Instance Normalization to balance feature extraction and maintain spatial details, thereby providing the necessary detail information for reconstruction.
[0034] As Figure 2 shown, to achieve linear computational complexity in CCFormer, a high-performance linearCross-Attention is required. Current Linear Attention uses a simple activation function similar to ReLU or PositiveOrthogonal Random Features (PRFs) to approximate the softmax kernel, reducing the explicit storage and calculation of the attention matrix. However, using an overly simple activation function will make it difficult for the finally obtained attention output to focus on the truly important parts, and using overly complex operations to approximate softmax will generate additional computational burdens. To obtain the performance of approximating the softmax kernel while reducing the operation complexity and computational complexity, the present invention selects a simple activation function LeakyReLU as the basis. By adjusting the directions of each Q and K, the Q and K with a similarity higher than the set threshold are made closer, while the Q and K with a similarity lower than the set threshold are made farther away, thus restoring to a certain extent the concentrated distribution like the original softmax, focusing the attention on the most informative regions, thereby reducing the problem that the weight distribution of the attention is too smooth. Thus, the formula is defined as follows:
[0035]
[0036]
[0037]
[0038] Among them, represents the similarity function, represents the intermediate function, and the function is the focusing function, which represents the degree of attention to the matrix. Inspired by the Flatten Transformer, the norm of the input tensor is retained, and only its direction is adjusted. By default . For the input tensor , a learnable bias initialized to zero is introduced, enabling the model to adapt to and utilize positional features. is a learnable scaling factor that acts on each channel of Q and K to prevent computational overflow and gradient explosion, while represents the LeakyReLU function. The diversity of features also limits the expressive power of linear attention. To further enhance feature representation, the present invention introduces a residual connection to process V, as shown in the formula:
[0039]
[0040] Among them, is the reshaped V. By introducing the convolutional operation and the non-linear activation function LeakyReLU, the rank of the output is enhanced, and by retaining part of the input information, the diversity of features is improved. Using convolutional operation can promote cross-channel information fusion, thereby strengthening the capture of global features, which is crucial for the depiction of complex and subtle structures in medical images. Therefore, the proposed efficient cross-attention mechanism provides performance comparable to that of softmax-scaled dot-product attention while maintaining linear computational complexity.
[0041] For the cross-layer constrained encoder and the cross-layer constrained decoder, when inputting 3D data and 2D data, there are some differences in the size of the feature dimension during the running process. For 3D data:
[0042] As Figure 1 and 3 shown, in the cross-layer constrained encoder, the output of the initial convolutional layer is represented as , and the feature maps generated by each subsequent downsampling stage and the cross-layer constrained encoder (CCE) are represented as and . Among them, each represents each stage (i.e., a specific channel configuration of , and respectively represent the length, width, and depth of the feature map of . For each cross-layer constrained encoder (CCE), represents the shallow feature of the th layer, while represents the subsequent deep feature. Since implementing multi-head cross-layer attention on these layers requires the input sequences Q, K, and V to have a consistent channel depth, a preliminary convolution is needed to adjust to match 's channel depth . Subsequently, adaptive max-pooling and average-pooling operations are used to downsample to the specified resolution , obtaining the downsampled versions and respectively, as shown in the following formula:
[0043]
[0044] where represents the adaptive max-pooling operation, and represents the adaptive average-pooling operation; this configuration significantly reduces the computational burden while maintaining the integrity of performance. For the multi-head attention mechanism, and are reshaped as a group to align with the input format , where B represents the batch size.
[0045] Then, is used as , while is used as . After processing through the efficient cross-attention module, a weighted feature representation is generated, where the number of heads in the first to fourth stages (i.e., - ) are 8, 16, 32, and 64 respectively. To allow the shallow feature ( ) to provide pixel-wise guidance for the optimization of the deep feature ( ), the shape of the weighted feature representation is reshaped to to align with and upsampled to through linear interpolation, denoted as .
[0046] Finally, after being processed by the Sigmoid activation function, the adjusted value is kept within within a range and element - by - element multiply with the deep features to adjust the amplitude and avoid gradient explosion:
[0047]
[0048] wherein, represents the up - sampling operation, represents the shallow features, represents the deep features. This formula ensures the effective alignment and adaptive optimization of the deep features by the shallow features, thereby improving the effectiveness of the encoder in medical image segmentation.
[0049] As Figure 1 and Figure 4 shown, during the step - by - step up - sampling process of the cross - layer constrained decoder, the outputs of its four stages are sequentially named . For example, the skip connection at the concatenates the shallow feature with to generate the concatenated feature . Next, the shallow feature and respectively increase the channel dimension through convolution, then apply group normalization (with every eight channels as a group) and the LeakyReLU activation function to extract finer features, and finally obtain the fine features and as follows:
[0050]
[0051]
[0052] Subsequently, similar to the down - sampling strategy, and both go through adaptive max - pooling and adaptive average - pooling operations to be adjusted to the specified resolution , thereby reducing the computational overhead and generating , and . To be consistent with the multi - head attention format, and are reshaped; thus, and serve as respectively, while serves as . Cross - layer attention is calculated on , to generate the attention maps and , attention map and are then resized by linear interpolation to match and dimensions. Finally, the attention maps and are constrained within the range of by the Sigmoid activation function and multiplied element-wise with and respectively to obtain the weighted features and . The weighted features and are fused by the following formula:
[0053]
[0054] This cross-layer constrained decoder structure enhances the integration of multi-scale features, thus promoting accurate segmentation through a layer-by-layer optimized attention mechanism, especially for fine attention calibration among hierarchically generated features.
[0055] For 2D data, as shown in Figure 5 and 6 , in the cross-layer constrained encoder, the output of the initial convolutional layer is represented as , while the feature maps generated by each subsequent downsampling stage and the cross-layer constrained encoder (CCE) are represented as and respectively. Among them, each represents a specific channel configuration for each stage (i.e., ), , and represent the length, width, and depth of the feature map at the th stage respectively. For each cross-layer constrained encoder (CCE), represents the shallow features of the th layer, while represents the subsequent deep features. Since implementing multi-head cross-layer attention on these layers requires the input sequences Q, K, and V to have a consistent channel depth, it is necessary to adjust through a preliminary convolution to match channel depth . Subsequently, adaptive max-pooling and average-pooling operations are used to downsample to the specified resolution , obtaining the downsampled versions and respectively, as shown in the following formula:
[0056]
[0057] Among them, represents the adaptive max pooling operation, represents the adaptive average pooling operation; this configuration significantly reduces the computational burden while maintaining the integrity of performance. For the multi-head attention mechanism, and are reshaped as a group to align with the input format , where B represents the batch size.
[0058] Then, is used as , while is used as , and after being processed by the efficient cross-attention module , a weighted feature representation is generated, where the number of heads in the first to fourth stages (i.e., - ) are 8, 16, 32, and 64 respectively. To allow the per-pixel guidance of the shallow features ( ) for the optimization of the deep features ( ), the shape of the weighted feature representation is reshaped to , aligned with , and upsampled to by linear interpolation, denoted as .
[0059] Finally, after being processed by the Sigmoid activation function, the adjusted value is kept within the range of , and multiplied element-wise with the deep features to adjust the amplitude and avoid gradient explosion:
[0060]
[0061] Among them, represents the upsampling operation, represents the shallow features, represents the deep features. This formula ensures the effective alignment and adaptive optimization of the deep features by the shallow features, thereby improving the effectiveness of the encoder in medical image segmentation.
[0062] As Figure 5 and Figure 7 shown, during the step-by-step upsampling process of the cross-layer constrained decoder, the outputs of its four stages are sequentially named . For example, the skip connection of the th layer concatenates the shallow feature with to generate the concatenated feature Next, the feature shallows and respectively pass through convolutions to increase the channel dimension, and then apply group normalization (with every eight channels as a group) and the LeakyReLU activation function to extract finer features, finally obtaining the fine features and , as follows:
[0063]
[0064]
[0065] Subsequently, similar to the downsampling strategy, and both undergo adaptive max-pooling and adaptive average-pooling operations, adjusted to the specified resolution , thereby reducing the computational overhead and generating , and . To be consistent with the multi-head attention format, and are reshaped; thus, and serve as respectively, while serves as . Cross-layer attention is calculated on , respectively generating the attention maps and The attention maps and are then resized by bilinear interpolation to match and 's dimensions. Finally, the values of the attention maps and are constrained within the range of by the Sigmoid activation function, and are multiplied element-wise with and respectively, obtaining the weighted features and . The weighted features and are fused through the following formula:
[0066]
[0067] This cross-layer constrained decoder structure enhances the integration of multi-scale features, thereby promoting accurate segmentation through a layer-by-layer optimized attention mechanism, especially for fine attention calibration among hierarchically generated features.
[0068] For 2D and 3D data in the present invention, the network framework remains unchanged, but the dimensions of the data are different.
[0069] Step 6: Use the training set divided in Step 2 for training. This framework is implemented based on CUDA 11.8, Python 3.9, PyTorch 2.0.0, and MONAI 1.3.1, and is trained on dual NVIDIA 4090 GPUs. In terms of hyperparameters, the learning rate is set to 1e-4, the AdamW strategy is used, and validation is performed every 1000 samples trained, and the best weights are saved.
[0070] Step 7: Model performance evaluation and result analysis. Use the weights with the best performance on the validation set obtained in Step 6 to calculate and analyze the results on the test set. The evaluation metrics are Dice Similarity Coefficient, Mean Intersection over Union (mIoU), and 95% Hausdorff Distance (HD95).
[0071] It can be seen from the index results that the present invention has achieved good results in various segmentation tasks, showing its strong competitiveness in medical image segmentation.
[0072] Table 1 Comparison of experimental results on ISIC2017 and ISIC2018 datasets
[0073]
[0074] The test results of the ISIC2017 and ISIC2018 datasets are shown in Table 1. CCFormer is superior to other methods in all indicators. The Dice indicators on the ISIC2017 dataset and the ISIC2018 dataset are 0.39% and 0.89% higher than those of the second-best model respectively. It is worth noting that compared with the second-best model DconnNet, the number of model parameters of CCFormer is smaller, only 36.68% of it.
[0075] Table 2 Comparison of experimental results on DSB-18 dataset
[0076]
[0077] Table 2 shows the test results of the DSB-18 dataset. The mIoU of CCFormer is 88.04% and the Dice is 93.46%, which are superior to several existing methods.
[0078] Table III Comparison of Experimental Results on OIMHS Dataset
[0079]
[0080] The test results of the OIMHS dataset are shown in Table III. CCFormer achieved the best performance with smaller parameter numbers and FLOPs, with Dice reaching 93.71% and mIoU being 88.60%.
[0081] Table IV Comparison of Experimental Results on SegTHOR Dataset
[0082]
[0083] The test results of the SegTHOR dataset are shown in Table IV. The Dice index of CCFormer is 90.96%, mIoU is 83.93%, and HD95 is 2.89, which is better than several existing methods. CFormer shows good performance in both regional and boundary segmentation accuracy.
[0084] Table V Ablation Experiment
[0085]
[0086] To further verify the effectiveness and robustness of CCFormer, ablation experiments were conducted on the DSB-18 dataset. The experimental results in Table V show that integrating cross-layer constraints into the encoder or decoder can improve performance to a certain extent. It is worth noting that introducing cross-attention mechanism in the skip connection can bring more significant performance improvement. In addition, the trade-off between the computational overhead brought by cross-layer constraints and performance improvement is within an acceptable range.
[0087] In summary, CCFormer proposed in the present invention improves the accuracy and robustness of medical image segmentation by integrating cross-layer constraints into the encoder or decoder and using efficient cross-attention to reduce its computing power requirements, and has broad application prospects.
[0088] The above description is only the preferred embodiment of the present invention and does not impose any limitation on the protection scope of the present invention. It should be emphasized that any minor modification or optimization made by those skilled in the art without departing from the basic idea of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A cross-layer constrained Transformer network implementation method based on hierarchical feature compensation, characterized in that: The steps include: Step 1: Collect medical image segmentation datasets for multiple different segmentation tasks; Step 2: Divide the medical image segmentation dataset into a training set, a validation set, and a test set in a ratio of 8:1:1; Step 3: Preprocess the data; Step 4: Perform data enhancement on the data in step 3; Step 5, building a segmentation network for medical image segmentation; the segmentation network integrates an efficient cross attention module, a cross-layer constraint encoder module and a cross-layer constraint decoder module; and the efficient cross attention module is applied to the cross-layer constraint encoder module and the cross-layer constraint decoder module respectively; Step 6: Model training and reasoning and save the best model weights; Step 7: Model performance evaluation and result analysis; The efficient cross-attention module is specifically implemented as follows: We choose a simple activation function LeakyReLU as the basis, and adjust the direction of each Q and K so that Q and K with similarity higher than the set threshold are closer, while Q and K with similarity lower than the set threshold are farther away, thus restoring the original softmax-like centralized distribution and focusing on the most informative area. The formula is defined as follows: sim(Q i ,K j )=φ d (Q i )·φ d (K j ) (1) Among them, sim( ) represents the similarity function, φ d ( ) represents the intermediate function, function F p is the focusing function, k = 3; for the input tensor x∈R B×N×C , positional bias ∈R 1×N×C Introduces a learnable bias initialized to zero, scale∈R 1 ×1×C is a learnable scaling factor that acts on each channel on Q and K to prevent computational overflow and gradient explosion, while θ represents the LeakyReLU function; A residual connection is introduced to process V, as shown in the formula: Att(Q,K,V)=φ d (Q)(φ d (W) T V)+θ(Conv(V')) (4) Among them, V' is the reshaped V. By introducing the convolution operation Conv(V') and the nonlinear activation function LeakyReLU, the output rank is enhanced, and the diversity of features is improved by retaining part of the input information.
2. The method for implementing a cross-layer constrained Transformer network based on hierarchical feature compensation according to claim 1, characterized in that: The proposed cross-layer constrained encoder module imposes constraints from shallow features representing boundaries and textures on deep features representing semantic and structural information, thereby preserving key low-dimensional details during downsampling.
3. The method for implementing a cross-layer constrained Transformer network based on hierarchical feature compensation according to claim 2, characterized in that: The cross-layer constraint encoder module is specifically implemented as follows: The output of the initial convolutional layer is represented as The feature maps generated by each subsequent downsampling stage and cross-layer constrained encoder are represented as and Among them, each C i represents the specific channel configuration at each stage, W i , H i and D i Respectively represent the stage i The length, width and depth of the feature map of the stage; for each cross-layer constrained encoder, represents the shallow features of the i-th layer, and It represents the subsequent deep features; it is adjusted by the preliminary 1×1 convolution To match The channel depth C (i+1) ; Then, adaptive maximum pooling and average pooling operations are used to and Downsample to the specified resolution a i , and get the downsampled versions and Then, Used as and Used as Through efficient cross-attention module processing {Q i ,K i ,V i }, generate weighted feature representation To allow shallow features For deep features Optimized pixel-wise guided, weighted feature representation The shape is reshaped into (C (i+1) ,a i ,a i ,a i ),and Aligned and upsampled to Recorded as Finally, after being processed by the Sigmoid activation function, the adjusted value remains in the range of (0,1) and is consistent with the deep features. Perform element-wise multiplication to regulate magnitude and avoid exploding gradients.
4. The method for implementing a cross-layer constrained Transformer network based on hierarchical feature compensation according to claim 3 is characterized in that: Adaptive max pooling and average pooling operations are used to and Downsample to the specified resolution a i , the specific formula is as follows: Among them, AdaptiveMaxPool() represents the adaptive maximum pooling operation, and AdaptiveAvgPool() represents the adaptive average pooling operation.
5. The method for implementing a cross-layer constrained Transformer network based on hierarchical feature compensation according to claim 3, characterized in that: For the multi-head attention mechanism, and As a set reshaped into the same format as the input (B, a i 3 ,C i+1 ), where B represents the batch size.
6. The method for implementing a cross-layer constrained Transformer network based on hierarchical feature compensation according to claim 3, characterized in that: The number of heads in the first to fourth stages, i.e. stage 1 to stage 4, are 8, 16, 32, and 64 respectively.
7. The method for implementing a cross-layer constrained Transformer network based on hierarchical feature compensation according to claim 3, characterized in that: With deep features Perform element-by-element multiplication and implement the formula as follows: Among them, UpSample() represents the upsampling operation. Represents shallow features, Represents deep features.
8. The method for implementing a cross-layer constrained Transformer network based on hierarchical feature compensation according to claim 3, characterized in that: In the stepwise upsampling process of the cross-layer constrained decoder, the outputs of its four stages are named {F u4 ,F u3 ,F u2 ,F u1 }; stage i The skip connection of the stage and Perform stitching and generate stitching features Next, the feature shallow and The channel dimension is increased by 1×1 convolution, and then group normalization and LeakyReLU activation function are applied to extract finer features, and finally the fine features are obtained respectively. and As shown below: Then, and All are adjusted to the specified resolution b after adaptive maximum pooling and adaptive average pooling operations. i ,generate and To be consistent with the multi-head attention format, and was reshaped; therefore, and As and As Calculate cross-layer attention on {Q1,K,V},{Q2,K,V} and generate attention maps respectively and Attention Map and It is then resized using linear interpolation to match and The dimension of ; finally, the attention map and The value of is constrained to the range (0,1) by the Sigmoid activation function and is respectively and Perform element-by-element multiplication to obtain weighted features and Weighted Features and The feature fusion is performed by the following formula:
Citation Information
Patent Citations
Medical image segmentation method based on residual axial attention
CN119229127A