Multi-attention-guided medical image segmentation method

The multi-attention guided medical image segmentation method addresses the limitations of CNN-Transformer networks by integrating dynamic convolutions and context-aware mechanisms, improving the segmentation of irregular lesions in medical images.

CN120318243APending Publication Date: 2025-07-15SHANDONG INST OF BUSINESS & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510389859.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing medical image segmentation model is not satisfactory when facing irregular lesions, especially the CNN model has limitations in extracting global features, while the Transformer model lacks local texture representation.

Method used

Using a multi-attention-guided medical image segmentation method, by constructing a model including an encoder, a decoder and a jump connection block, a dual attention segmentation block, a decoder segmentation block and a range-variable dynamic convolution are used, combining channel attention and cross attention to form an efficient medical image segmentation network.

Benefits of technology

It improves the segmentation accuracy of irregular lesions, can better restore image details and global context information, and improves the accuracy and efficiency of medical image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318243A_ABST
    Figure CN120318243A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to a multi-attention-guided medical image segmentation method. Preprocessing the input image; a medical image segmentation model is constructed, the preprocessed image passes through a double-attention segmentation block, the spatial dimension and the number of channels of the image are modified, and whether the image is directly transmitted to a decoder or not is judged according to the spatial dimension and the number of channels; if the data are not directly transmitted to the decoder, the data are transmitted to a jump connection block and are transmitted to a decoder segmentation block through jump connection; otherwise, the image is directly transmitted to a decoder, and the spatial dimension and the channel number of the image are recovered; training the constructed medical image segmentation model to obtain a trained medical image segmentation model; and segmenting the medical image by using the trained medical image segmentation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a medical image segmentation method guided by multi-attention. Background Art

[0002] An image is a combination of a graph and an image, containing information about the object being described. As an important information carrier, an image has the characteristics of large amount of information, strong intuitiveness, and simple and easy-to-understand representation. Driven by digital image processing technology, image information is increasingly widely used in many fields such as industry, medicine, and transportation, bringing great convenience to people's production and life.

[0003] In order to effectively extract and utilize the information contained in digital images, it is necessary to segment the images. Image segmentation is a crucial basic link in computer vision. Without correct segmentation, there can be no correct recognition. The combination of medical image segmentation and computer vision has brought a revolutionary change in the field of medical image processing. In recent years, the rapid development of deep learning technology has provided powerful tools and methods for medical segmentation, which can not only reduce the laboriousness and time-consuming of manual work, improve the accuracy and efficiency of segmentation, but also help doctors with rapid clinical diagnosis, surgical planning, and treatment decision-making. However, when faced with irregular lesions, the current medical image segmentation results are often unsatisfactory. Therefore, it is necessary to create fast and accurate segmentation methods in medical image segmentation research.

[0004] Current medical image segmentation models mainly focus on two basic architectures: Convolutional Neural Network (CNN) and Transformer models. The size of the object that the CNN model can detect is closely related to the receptive field dimension of the corresponding network layer. Currently, most methods are reparameterization techniques and pyramid pooling techniques for convolution. Although there have been improvements, the current CNN model has limitations in extracting global features. The Transformer model is a deep learning model based on the self-attention mechanism. In medical image segmentation, the current innovation direction is to modify the attention mechanism from three aspects: channel attention, spatial attention, and their fusion. However, the current model has the problem of insufficient local texture representation. Therefore, some current CNN-Transformer models have become popular techniques in the field of medical image segmentation. Although the processing ability of this encoder-masker structure has been improved in medical image segmentation, most of these networks rely heavily on the CNN backbone, and the segmentation effect is unsatisfactory when faced with irregular lesions. Summary of the Invention

[0005] In order to overcome the problems in the prior art, the present invention proposes a medical image segmentation method guided by multi-attention.

[0006] The technical solution of the present invention to solve the above technical problems is as follows:

[0007] The present invention provides a medical image segmentation method guided by multi-attention, comprising the following steps:

[0008] Preprocess the input image; and construct a medical image segmentation model, the medical image segmentation model includes an encoder, a decoder and a skip connection block, the encoder includes a plurality of dual-attention segmentation blocks, and the decoder includes a plurality of decoder segmentation blocks; the skip connection block is connected between the dual-attention segmentation block and the decoder segmentation block;

[0009] The preprocessed image passes through the dual-attention segmentation block, modifies the spatial dimension and the number of channels of the image, and determines whether to directly input it into the decoder according to the spatial dimension and the number of channels; if it is not directly input into the decoder, it is input into the skip connection block and transmitted to the decoder segmentation block through the skip connection; otherwise, it is directly input into the decoder to restore the spatial dimension and the number of channels of the image;

[0010] Train the constructed medical image segmentation model to obtain a trained medical image segmentation model;

[0011] Use the trained medical image segmentation model to segment medical images.

[0012] Further, the preprocessing includes dividing the input image into a plurality of non-overlapping image blocks, and the size of each image block is equal.

[0013] Further, the dual-attention segmentation block includes transpose attention and efficient attention, and the feature X output after passing through the dual-attention segmentation block eo :

[0014] X eo =EfficientAttention(FFN(Norm(TransposeAttention(FFN(Norm(X))))));

[0015] In the above formula, EfficientAttention is efficient attention, TransposeAttention is transpose attention, FFN is a feed-forward neural network, and Norm is normalization processing.

[0016] Further, the skip connection block includes channel attention and cross attention, and the feature X input into the decoder through the skip connection block skip :

[0017]

[0018] Among them, ChannelAttention is the channel attention operation, Softmax is the activation function, Linear is the linear operation, and X eo is the feature output by the encoder dual attention segmentation block, and x do is the output of the decoder, and d k is the dimension of the key vector.

[0019] Furthermore, the decoder segmentation block includes LayerNorm, segmentation block attention, and a multi-layer perceptron. The decoder segmentation block is expressed as:

[0020] x la = Attention(LN(X skip )) + X skip

[0021] x fo = MLP(LN(x la )) + x la ;

[0022] Among them, MLP is the multi-layer perceptron, LN is LayerNorm, and Attention is the segmentation block attention; x la is the feature output by LayerNorm and the attention residual, and x fo is the feature output by the decoder segmentation block.

[0023] Furthermore, the attention includes: the GeLU activation function, variable dynamic convolution, variable dynamic convolution spatial transformation, and 1×1 convolution; the input feature is passed into a 1×1 convolution, passed through the GeLU activation function into the variable dynamic convolution with a variable range and its spatial transformation, and then passed through multiple 1×1 convolutions to obtain the final result:

[0024] x ge = GELU(Conv1×1(x layer ))

[0025]

[0026] Among them, Conv is the variable dynamic convolution with a variable range, Conv_S is the variable dynamic convolution spatial transformation with a variable range, Conv1×1 is the 1×1 convolution operation, and x ge represents the feature after passing through the GeLU activation function, and x layer represents the feature output by LayerNorm.

[0027] Furthermore, the variable dynamic convolution includes: receiving the feature map output by the GeLU activation function and generating initial offsets using Offsets_Conv; normalizing and activating the offsets through Batch Norm and tanh; using Split to split the processed offsets into components in the x and y directions, which are respectively used as the inputs of Conv_x and Conv_y, and Conv_x and Conv_y calculate the deformed feature map through bilinear interpolation; convolving the deformed feature map with the dynamic convolution kernel; adding the convolution result to the feature map output by the GeLU activation function to obtain the final output feature map.

[0028] Furthermore, calculating the deformed feature map through bilinear interpolation includes: calculating the deformed eigenvalue I' according to the value of the pixel point and its relative position to the deformed coordinates:

[0029] I' = (1 - w1)(1 - w2)I(x1, y1) + w1(1 - w2)I(x2, y1) + (1 - w1)w2I(x1, y2) + w1w2I(x2, y2);

[0030] where w1 is the weight of the interpolation point in the x direction, w2 is the weight of the interpolation point in the y direction; (x1, y1)(x2, y1)(x1, y2)(x2, y2) represent the four pixel points around the deformed coordinates found in the feature map output by the GeLU activation function.

[0031] Furthermore, training the constructed medical image segmentation model to obtain a trained medical image segmentation model includes: training using the gradient descent optimization algorithm; passing the input image into the constructed medical image segmentation model to obtain a prediction result, and calculating the loss value according to the output of the medical image segmentation model and the ground truth label; calculating the gradient and updating the weights of the medical image segmentation model, and continuously optimizing the parameters of the medical image segmentation model within multiple training epochs until the loss function converges or reaches a preset stop condition.

[0032] Furthermore, the loss function adopts the DICE coefficient loss:

[0033]

[0034] where Dice is the DICE coefficient loss, A is the pixel set of the predicted image, and B is the pixel set of the ground truth label.

[0035] Compared with the prior art, the present invention has the following technical effects:

[0036] In order to achieve higher-precision segmentation of irregular lesions, the present invention incorporates a range-variable dynamic convolution into the attention mechanism, enabling the formation of a streamlined attention mechanism that fully understands the context and based on this attention, a decoder segmentation block is composed; secondly, a new skip connection is designed, and channel attention is added before the fusion of query, key, and value to better restore the details of the image; finally, a dual-attention segmentation block is introduced as the encoder and combined with the previous two to form an efficient medical image segmentation network. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0038] Figure 1 It is a flowchart of the present invention;

[0039] Figure 2 It is a structural diagram of the medical image segmentation model of the present invention;

[0040] Figure 3 It is a structural diagram of the dual-attention segmentation block of the present invention;

[0041] Figure 4 It is a structural diagram of the skip connection block of the present invention;

[0042] Figure 5 It is a structural diagram of the range-variable dynamic convolution of the present invention;

[0043] Figure 6 It is a structural diagram of the attention of the present invention;

[0044] Figure 7 It is a structural diagram of the decoder segmentation block of the present invention;

[0045] Figure 8 It is a segmentation result diagram of the present invention on the ISIC2018 dataset;

[0046] Figure 9 It is a segmentation result diagram of the present invention on the Synapse dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following will describe in detail the specific implementation manner, structure, features and effects of the technical solution proposed according to the present invention in combination with the accompanying drawings and preferred embodiments. Specific features, structures or characteristics in one or more embodiments may be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0048] Referring to Figures 1 - 9 , in an embodiment of the present invention, a medical image segmentation method guided by multi-attention is provided, including the following steps:

[0049] Preprocess the input image; and construct a medical image segmentation model, the medical image segmentation model includes an encoder, a decoder and a skip connection block, the encoder includes a plurality of dual-attention segmentation blocks, and the decoder includes a plurality of decoder segmentation blocks; the skip connection block is connected between the dual-attention segmentation block and the decoder segmentation block;

[0050] The preprocessed image passes through the dual-attention segmentation block, modifies the spatial dimension and the number of channels of the image, and determines whether to directly input it into the decoder according to the spatial dimension and the number of channels; if it is not directly input into the decoder, it is input into the skip connection block and transmitted to the decoder segmentation block through the skip connection; otherwise, it is directly input into the decoder to restore the spatial dimension and the number of channels of the image;

[0051] Train the constructed medical image segmentation model to obtain a trained medical image segmentation model;

[0052] Use the trained medical image segmentation model to segment medical images.

[0053] The following will expand on each of the above steps in detail:

[0054] Step 100: Preprocess the input image.

[0055] For the input image Divide it into several non-overlapping image patches, and the size of each image patch is P×P:

[0056]

[0057] Where H represents the height of the image, W represents the width of the image, C is the number of channels, and N is the total number of image patches.

[0058] Map each image patch to a high-dimensional feature space through a linear transformation to obtain the initial feature vector Z0 of the input image, so as to prepare for subsequent model processing:

[0059] Z0 = [z1E; z2E; …; z N E] + E pos ;

[0060] where z N represents the feature vector of the image patch; E is the linear embedding matrix; E pos is the positional encoding to ensure that the network retains the positional spatial information of the image patch.

[0061] Step 200: Construct a medical image segmentation model, the medical image segmentation model includes an encoder, a decoder and a skip connection block, the encoder includes a plurality of dual attention segmentation blocks, and the decoder includes a plurality of decoder segmentation blocks; the skip connection block is connected between the dual attention segmentation block and the decoder segmentation block; wherein, the encoder is used to extract features from the input image, the decoder is used to convert the features generated by the encoder back to the image space, and generate a segmentation result; feature fusion is performed between the encoder and the decoder through the skip connection block.

[0062] Specifically, after the preprocessed image passes through the dual attention segmentation block, the spatial dimension and the number of channels of the image will change accordingly. According to the changes in the spatial dimension and the number of channels of the image, it is judged whether to directly input the image into the decoder. If the features of the image are already suitable enough for the decoder to perform further processing after passing through the dual attention segmentation block, that is, the features of the encoder and the decoder are sufficiently matched in channels and space, and direct splicing can ensure dimension alignment and retain more original information, then it can be directly input into the decoder; otherwise, the low-level features extracted from the encoder stage are first input into the skip connection block, and through the skip connection block, the information of the encoder is transmitted to the decoder segmentation block.

[0063] In this embodiment, the encoder includes a Patch Embedding block and a dual attention segmentation block. The function of the Patch Embedding block is to divide the input image into multiple small patches (Patch) for efficient processing in the low-dimensional space. By dividing the image into smaller regions, the Patch Embedding block not only reduces the computational complexity, but also enables independent processing and analysis of each small patch. Each Patch contains local image information. The Patch Embedding block captures the detailed features of the local region by performing feature extraction on each Patch. Then, it maps these features to the low-dimensional vector space through a linear transformation, converting the feature representation of each Patch into a low-dimensional vector, thereby providing an efficient and representative input for the subsequent deep learning model. This process helps to enhance the model's perception ability of local features and provides a basis for further feature enhancement and refinement.

[0064] Refer to Figure 3, the dual attention segmentation block includes Transpose Attention and Efficient Attention. Among them, the formula for the Transpose Attention is:

[0065] T(Q, K, V) = VC T (K, Q)

[0066] C T (K, Q) = Softmax(K T Q / τ);

[0067] In the above formula, V is the value matrix, C T (K, Q) is the context variable of the Transpose Attention, τ is the temperature parameter, Q is the query matrix, K is the key matrix, and T(Q, K, V) represents the Transpose Attention; τ represents a temperature matrix.

[0068] The formula for the Efficient Attention is:

[0069] E(Q, K, V) = ρ q (Q)(ρ k (K) T V);

[0070] In the above formula, Q is the query matrix, K is the key matrix, V is the value matrix; E(Q, K, V) represents the Efficient Attention; ρ q (Q) represents the normalization operation on the query matrix; ρ k (K) represents the normalization operation on the key matrix.

[0071] Mark the result obtained through the dual attention segmentation block as X eo :

[0072] X eo = EfficientAttention(FFN(Norm(TransposeAttention(FFN(Norm(X))))));

[0073] In the above formula, EfficientAttention is the Efficient Attention, TransposeAttention is the Transpose Attention, FFN is the feed - forward neural network, and Norm is the normalization process.

[0074] Input the input image into the dual attention segmentation block and obtain the corresponding image. Each time passing through the dual attention segmentation block and downsampling, the spatial dimension of the image is halved, and the number of channels becomes twice the original. The specific dimension changes are as follows:

[0075]

[0076] In this embodiment, referring to Figure 4 , skip connection: between each layer of the encoder and the corresponding layer of the decoder, the feature maps of the intermediate layers are retained and passed to the decoder through skip connections. Skip connections help the decoder better recover the details of the image.

[0077] In order to efficiently fuse these two features, the output of the encoder and the output from the decoder, a series of delicate operations are taken.

[0078] First, before being used as the key and value inputs to the attention mechanism, the output of the decoder is linearized. Immediately afterwards, an operation of channel attention is added. The purpose of this step is to further strengthen the relationship of the feature maps in the channel dimension, thereby enhancing the representation ability of the model. The channel attention mechanism usually calculates the weights of each channel and then recalibrates the feature maps based on these weights to emphasize those channels that are more important for the task while suppressing those relatively unimportant channels.

[0079] The tensor after channel attention processing is input into the cross-attention block. The cross-attention mechanism allows the model to focus on and refer to the features at the corresponding position or related positions of the encoder when generating the features at a certain position in the decoder. This mechanism is crucial for capturing global context information because it allows the decoder to make decisions based on the entire input of the encoder when generating each output.

[0080] In the cross-attention block, the output of the decoder after channel attention processing serves as the query, while the output of the encoder serves as the key and value. Through these interaction operations, the model can dynamically adjust the features to generate a feature X for input to the decoder skip .

[0081]

[0082] Among them, ChannelAttention() is the channel attention operation, Softmax() is the activation function, Linear is the linear operation, X eo is the output of a single encoder, x do is the output of a single decoder, d k is the dimension of the key vector.

[0083] In this embodiment, the decoder includes a decoder segmentation block and Patch Expanding. Patch Expanding is used to convert the low-resolution feature map back to a high-resolution image, and this process usually involves an upsampling operation.

[0084] Referring to Figure 5 , for the range-variable dynamic convolution. In the range-variable dynamic convolution, by dynamically adjusting the position and weight of the convolution kernel according to the characteristics of the input feature map to adapt to different input data, the adaptability and performance of the model are improved. In the range-variable dynamic convolution, this idea is further extended so that the convolution operation can not only dynamically adjust the weight according to the input feature map, but also dynamically change the pixel area covered by the convolution kernel.

[0085] The range-variable dynamic convolution includes: receiving the feature map output by the GeLU activation function and generating an initial offset using Offsets_Conv; normalizing and activating the offset through Batch Norm and tanh; using Split to split the processed offset into components in the x and y directions, which are used as the inputs of Conv_x and Conv_y respectively. Conv_x and Conv_y calculate the deformed feature map through bilinear interpolation; convolving the deformed feature map with the dynamic convolution kernel; adding the convolution result to the feature map output by the GeLU activation function to obtain the final output feature map.

[0086] In a specific embodiment, first, for a given input feature map I, the new coordinates of each pixel point in the deformed feature map need to be generated, and this step is the key to dynamically adjusting the position and coverage range of the convolution kernel.

[0087] Specifically, first, the offsets δ x and δy in the x and y directions are generated for each pixel point, where δ ∈ [-1, 1]. These offsets represent the offsets of the new coordinates of each pixel point in the deformed feature map relative to the original position. Then, the deformed coordinates (x', y') are calculated according to the offsets, and the formula is as follows:

[0088] x' = x + δ x

[0089] y' = y + δ y ;

[0090] For the deformed coordinates (x', y'), however, since the deformed coordinates may be non-integer, the corresponding pixel value cannot be directly found in the original feature map. To solve this problem, the method of bilinear interpolation is used to calculate the deformed feature value. Specifically, find the four pixel points (x1, y1), (x2, y1), (x1, y2), (x2, y2) (i.e., the upper left, upper right, lower left, and lower right) around the deformed coordinates in the original feature map, and then use the bilinear interpolation formula to calculate the deformed feature value I' according to the values of these four pixel points and their relative positions to the deformed coordinates:

[0091] I' = (1 - w1)(1 - w2)I(x1, y1) + w1(1 - w2)I(x2, y1) + (1 - w1)w2I(x1, y2) + w1w2I(x2, y2);

[0092] where w1 and w2 are the weights of the interpolation points in the x and y directions, and the determination of the weights depends on the relative position relationship between the four pixel points.

[0093] Finally, take the deformed feature map as the input and perform a convolution operation with the convolution kernel K of the deformable convolution. Since the position and weights of the convolution kernel are dynamically adjusted according to the input feature map, the sampling position and weights of the convolution kernel are not fixed, but are adjusted according to the content of the input feature map through an additional network module or an adaptive mechanism. The position of the convolution kernel can be offset according to the local features and deformation conditions of the input image, so as to better capture the key information in the image. At the same time, the weights of the convolution kernel can also change dynamically according to the input feature map, enabling the convolution operation to respond more flexibly to the changes in the input data. In this way, the convolution operation can more accurately reflect the characteristics of the input data, thereby improving the performance of the model. Therefore, this step can achieve the adaptive processing of the input feature map. By dynamically adjusting the position and weights of the convolution kernel, the final output feature map can be obtained, and this feature map can more accurately reflect the characteristics of the input data, thereby improving the performance of the model.

[0094] Regarding the attention formed based on the above convolution. Refer to Figure 6 , use a range-variable dynamic convolution to replace the traditional convolution layer in the attention framework, and construct an attention module containing multiple components.

[0095] First, pass the input feature into a 1×1 convolution, and then pass it through the GeLU activation function into the range-variable dynamic convolution and its spatial transformation, and finally pass it through multiple 1×1 convolutions again to obtain the final result:

[0096] xge = GELU(Conv1×1(xlayer))

[0097]

[0098] Among them, Conv is a dynamic convolution with a variable range, Conv_S is a spatial transformation of the dynamic convolution with a variable range, Conv1×1 is a 1×1 convolution operation, and x ge represents the feature after passing through the GeLU activation function, and x layer represents the feature output by LayerNorm.

[0099] Regarding the decoder segmentation block formed based on the above attention. Refer to Figure 7 , the structure of the decoder segmentation block includes LayerNorm, segmentation block attention, and a multi-layer perceptron. The inheritance of the residual connection ensures effective feature propagation. The decoder segmentation block can be expressed as:

[0100] x la = Attention(LN(X skip )) + X skip

[0101] x fo = MLP(LN(x la )) + x la ;

[0102] Among them, MLP is a multi-layer perceptron, LN is LayerNorm, and Attention is segmentation block attention; x la is the feature output by LayerNorm and the attention residual, and x fo is the feature output by the decoder segmentation block.

[0103] The obtained features are input into the decoder segmentation block, and the corresponding image is obtained. Each time the decoder segmentation block is passed through, the spatial dimension of the image will double, and the number of channels will be halved. After passing through the decoder segmentation block and upsampling each time, the generated image will be restored to the same dimension and channels as the input image. The specific dimension changes are as follows:

[0104]

[0105] In a specific embodiment, refer to Figure 2, the constructed medical image segmentation model structure includes an encoder, a decoder, and a skip connection module; the encoder includes a first Patch Embedding block, a first dual attention segmentation block, a second Patch Embedding block, a second dual attention segmentation block, a third Patch Embedding block, and a third dual attention segmentation block connected in sequence; the decoder includes a first decoder segmentation block, a first Patch Expanding, a second decoder segmentation block, a second Patch Expanding, a third decoder segmentation block, and a third Patch Expanding;

[0106] The skip connection module includes a first skip connection block and a second skip connection block; the input of the first skip connection block is respectively connected to the output of the first dual attention segmentation block and the output of the second Patch Embedding block, and the output of the first skip connection block is connected to the input of the third decoder segmentation block; the input of the second skip connection block is respectively connected to the output of the second dual attention segmentation block and the output of the first Patch Embedding block, and the output of the second skip connection block is connected to the input of the second decoder segmentation block.

[0107] The image first enters the first Patch Embedding block, is decomposed into multiple small blocks and features are extracted to generate a low-dimensional vector representation; then it enters the first dual attention segmentation block, and through the combined action of spatial attention and channel attention, the feature representation is further refined and enhanced; the feature representation after the first refinement enters the second Patch Embedding block for further feature extraction and vectorization; these new vector representations enter the second dual attention segmentation block again for more in-depth refinement and enhancement; similarly, the feature representation after the second refinement enters the third Patch Embedding block and the dual attention segmentation block for final feature extraction and refinement;

[0108] The first decoder segmentation block processes the output features of the encoder and begins to gradually restore the spatial resolution of the image; the first Patch Expanding: increases the spatial dimension of the feature map through upsampling or similar operations; the second decoder segmentation block further processes the upsampled features and generates a more refined segmentation result; the second Patch Expanding increases the spatial dimension of the feature map again; the third decoder segmentation block processes the features in the final stage of the decoder and generates the final segmentation result; if necessary, the third Patch Expanding can increase the spatial dimension of the feature map again to match the size of the input image.

[0109] The first skip connection block combines the output of the first dual attention segmentation block with the output of the second Patch Embedding block. This combination helps to introduce more context information during the decoding process. The output of the first skip connection block is connected to the input of the third decoder segmentation block. The second skip connection block combines the output of the second dual attention segmentation block with the output of the first Patch Embedding block, which also helps to introduce more context information during the decoding process. The output of the second skip connection block is connected to the input of the second decoder segmentation block.

[0110] By combining the encoder-decoder architecture and the skip connection module, the medical image segmentation task can be effectively processed. The encoder is responsible for extracting features, while the decoder is responsible for converting these features back to the image space. The skip connection module helps to transfer information between the encoder and the decoder, thereby improving the performance of the model.

[0111] Step 300: Train the constructed medical image segmentation model to obtain a trained medical image segmentation model.

[0112] Use the gradient descent optimization algorithm for training; pass the input image into the constructed medical image segmentation model to obtain a prediction result, calculate the loss value according to the output of the medical image segmentation model and the ground truth label; calculate the gradient and update the weights of the medical image segmentation model, and continuously optimize the parameters of the medical image segmentation model within multiple training epochs until the loss function converges or reaches the preset stopping condition.

[0113] For the loss function, mainly use the DICE coefficient loss to measure the overlap degree between the prediction and the ground truth label. The formula is:

[0114]

[0115] where A is the pixel set of the predicted image and B is the pixel set of the ground truth label.

[0116] Step 400: Use the trained medical image segmentation model to segment medical images.

[0117] In the present invention, by integrating range-variable dynamic convolution into the attention mechanism, a streamlined attention mechanism that fully understands the context can be formed and decoder segmentation blocks are composed based on this attention. Secondly, a new skip connection is designed, and channel attention is added before the fusion of query, key, and value to better restore the details of the image. Finally, a dual attention segmentation block is introduced as the encoder and combined with the previous two to form an efficient medical image segmentation network.

[0118] The final results are shown in Tables 1 and 2. The proposed method demonstrates significant improvements over other methods on multiple datasets. Higher results are obtained on multiple metrics such as the left kidney, gallbladder, liver, aorta, pancreas, and average DSC on the Synapse dataset; on the ISIC2017 and ISIC2018 datasets, good results are obtained on multiple metrics such as DSC, specificity, and accuracy. By introducing skip connections and decoder blocks proposed in this method and combining them with the dual attention segmentation block, the accuracy of medical image segmentation can be improved.

[0119] Table 1 compares the proposed method on the Synapse dataset.

[0120]

[0121]

[0122] Table 2 compares the proposed method on the ISIC2017 and ISIC2018 datasets.

[0123]

[0124] The above embodiments are only used to illustrate the technical solutions of the present application, not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A medical image segmentation method guided by multiple attentions, characterized in that Including the following steps: Preprocess the input image; And construct a medical image segmentation model, the medical image segmentation model includes an encoder, a decoder and a skip connection block, the encoder includes a plurality of dual attention segmentation blocks, and the decoder includes a plurality of decoder segmentation blocks; the skip connection block is connected between the dual attention segmentation block and the decoder segmentation block; The preprocessed image passes through the dual attention segmentation block, modifies the spatial dimension and the number of channels of the image, and determines whether to directly input it into the decoder according to the spatial dimension and the number of channels; If it is not directly input into the decoder, it is input into the skip connection block and transmitted to the decoder segmentation block through the skip connection; Otherwise, it is directly input into the decoder to restore the spatial dimension and the number of channels of the image; Train the constructed medical image segmentation model to obtain a trained medical image segmentation model; Use the trained medical image segmentation model to segment medical images.

2. The multi-attention-guided medical image segmentation method according to claim 1, wherein, The preprocessing includes dividing the input image into a plurality of non-overlapping image blocks, and the size of each image block is equal.

3. A multi-attention-guided medical image segmentation method according to claim 1, characterized in that The dual attention segmentation block includes transposed attention and efficient attention, and the feature X output after passing through the dual attention segmentation block eo : X eo = EfficientAttention(FFN(Norm(TransposeAttention(FFN(Norm(X)))))); In the above formula, EfficientAttention is efficient attention, TransposeAttention is transposed attention, FFN is a feed-forward neural network, and Norm is normalization processing.

4. A multi-attention-guided medical image segmentation method according to claim 3, characterized in that The skip connection block includes channel attention and cross attention, and the feature X input to the decoder through the skip connection block skip : Among them, ChannelAttention is the channel attention operation, Softmax is the activation function, Linear is the linear operation, and X eo is the feature output by the encoder double-attention segmentation block, and x do is the output of the decoder, and d k is the dimension of the key vector.

5. A multi-attention-guided medical image segmentation method according to claim 4, characterized in that The decoder segmentation block includes LayerNorm, segmentation block attention and a multi-layer perceptron, and the decoder segmentation block is expressed as: x la = Attention(LN(X skip )) + X skip x fo = MLP(LN(x la )) + x la ; Among them, MLP is a multi-layer perceptron, LN is LayerNorm, and Attention is split-block attention; x la is the feature output by LayerNorm and the attention residual, x fo is the feature output by the decoder split block.

6. A multi-attention-guided medical image segmentation method according to claim 5, wherein The attention includes: GeLU activation function, variable dynamic convolution, variable dynamic convolution spatial transformation and 1×1 convolution; the input feature is input into a 1×1 convolution, passed through the GeLU activation function into the variable dynamic convolution with a variable range and its spatial transformation, and then passed through a plurality of 1×1 convolutions to obtain the final result: x ge = GELU(Conv1×1(x layer )) Among them, Conv is a dynamic convolution with a variable range, Conv_S is a spatial transformation of the dynamic convolution with a variable range, Conv1×1 is a 1×1 convolution operation, and x ge represents the feature after passing through the GeLU activation function, and x layer represents the feature output by LayerNorm.

7. A multi-attention-guided medical image segmentation method according to claim 6, characterized in that, The variable dynamic convolution includes: receiving the feature map I output by the GeLU activation function, and using Offsets_Conv to generate initial offsets; normalizing and activating the offsets through Batch Norm and tanh; using Split to split the processed offsets into components in the x and y directions, respectively as the inputs of Conv_x and Conv_y, and Conv_x and Conv_y calculate the deformed feature map through bilinear interpolation; convolve the deformed feature map with the dynamic convolution kernel; add the convolution result to the feature map output by the GeLU activation function to obtain the final output feature map.

8. A multi-attention-guided medical image segmentation method according to claim 7, characterized in that Calculating the deformed feature map through bilinear interpolation includes: calculating the deformed feature value I' according to the value of the pixel point and the relative position with the deformed coordinates: I' = (1 - w1)(1 - w2)I(x1, y1) + w1(1 - w2)I(x2, y1) + (1 - w1)w2I(x1, y2) + w1w2I(x2, y2); Where, w1 is the weight of the interpolation point in the x direction, w2 is the weight of the interpolation point in the y direction; (x1, y1)(x2, y1)(x1, y2)(x2, y2) represent the four pixel points around the deformed coordinates found in the feature map output by the GeLU activation function.

9. A multi-attention-guided medical image segmentation method according to claim 1, characterized in that, Train the constructed medical image segmentation model to obtain a trained medical image segmentation model, including: using the gradient descent optimization algorithm for training; passing the input image into the constructed medical image segmentation model to obtain a prediction result, and calculating a loss value based on the output of the medical image segmentation model and the ground truth label; calculating the gradient and updating the weights of the medical image segmentation model, and continuously optimizing the parameters of the medical image segmentation model within multiple training epochs until the loss function converges or reaches a preset stopping condition.

10. A multi-attention-guided medical image segmentation method according to claim 9, wherein The loss function uses the DICE coefficient loss: where Dice is the DICE coefficient loss, A is the pixel set of the predicted image, and B is the pixel set of the ground truth label.