Image convolution and attention mixed medical image segmentation method

Through the method of mixing graph convolution and attention, combined with PVTv2 encoder and multi-stitching space attention, the shortcomings of the existing medical image segmentation model in capturing multimodal dependencies are solved, and medical image segmentation with higher accuracy is achieved.

CN120318242APending Publication Date: 2025-07-15SHANDONG INST OF BUSINESS & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510389848.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing medical image segmentation models are difficult to capture both short-term and long-term dependencies of multimodal spatial dimensions simultaneously, resulting in insufficient segmentation accuracy.

Method used

A medical image segmentation model is constructed by combining graph convolution and attention, combining PVTv2 encoder, graph convolution blocks with matrix position offset, upsampling blocks with multi-stitching space attention and innovative activation functions.

Benefits of technology

It improves the accuracy of medical image segmentation, enhances the modeling ability and the perceived feature correlation of complex graph structures, and achieves higher segmentation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318242A_ABST
    Figure CN120318242A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to an image convolution and attention mixed medical image segmentation method. Preprocessing the input image; performing PVTv2 on the preprocessed image, extracting the features of the image, modifying the spatial dimension and the channel number of the image, and judging whether the image is directly transmitted to a decoder or not according to the spatial dimension and the channel number; if the data are not directly transmitted to the decoder, the data are transmitted to the jump connection block and transmitted to the decoder block through jump connection; otherwise, the image is directly transmitted to a decoder block for processing, the processed image is transmitted to an up-sampling block with an activation function, and the image is recovered; training the constructed medical image segmentation model to obtain a trained medical image segmentation model; and segmenting the medical image by using the trained medical image segmentation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a medical image segmentation method combining graph convolution and attention. Background Art

[0002] As a combination of graphics and images, an image carries rich information and precisely conveys various details of the depicted object. As a core information transmission medium, due to its characteristics of information density, intuitiveness, and easy comprehensibility, with the promotion of digital image processing technology, images are widely used in multiple fields such as industry, medicine, and transportation, greatly improving people's production efficiency and quality of life.

[0003] In this context, image segmentation technology, as a crucial step in digital image processing, plays an extremely important role. The accuracy of image segmentation directly affects the effect of subsequent recognition tasks. Therefore, it has a fundamental and key position in the field of computer vision. In the field of medical image processing, the rise of medical image segmentation technology, especially its deep integration with computer vision technology, has promoted a huge transformation in this field. In recent years, the rapid development of deep learning technology has provided strong technical support and efficient solutions for medical image segmentation, not only significantly reducing the burden of manual operations, improving the accuracy and efficiency of segmentation, but also providing strong support for doctors in clinical diagnosis, surgical plan design, and treatment decision-making.

[0004] With the development of deep learning, medical image segmentation technology has made significant progress in lesion recognition and diagnostic accuracy, especially in the application of Convolutional Neural Network (CNN) and Transformer models. Traditional encoder-decoder architectures based on CNN (such as U-Net, ResNet, etc.) have become the mainstream methods for medical image segmentation because they can accurately capture local features. However, due to the locality of the CNN convolution kernel, it is difficult to effectively capture long-range dependencies. For this reason, researchers have improved CNN by introducing attention mechanisms, expanding the convolution receptive field, etc. For example, MA-Unet and APAUNet have improved the segmentation performance through attention mechanisms, while dilated convolution and pyramid pooling methods have further enhanced the model's context capture ability. However, CNN still has limitations in modeling long-range dependencies, which has prompted the introduction of the Transformer model. Transformer was initially applied to natural language processing, but due to its self-attention mechanism's ability to effectively capture long-range dependencies, it has been successfully applied to medical image segmentation. Although both CNN and Transformer have made progress in medical image segmentation, existing models still face limitations, such as difficulty in simultaneously capturing short-term and long-term dependencies in all multi-modal spatial dimensions. Summary of the Invention

[0005] In order to overcome the problems in the prior art, the present invention proposes a medical image segmentation method that combines graph convolution and attention.

[0006] The technical solution of the present invention to solve the above technical problems is as follows:

[0007] The present invention provides a medical image segmentation method that combines graph convolution and attention, including the following steps:

[0008] Preprocess the input image; and construct a medical image segmentation model, the medical image segmentation model includes an encoder, a decoder and a skip connection block, and the skip connection block is connected between the encoder and the decoder; the encoder includes PVTv2; the decoder includes a decoder block and an upsampling block, and the decoder block includes a graph convolution block with bias and multi-splicing spatial attention;

[0009] The preprocessed image passes through PVTv2, extracts the features of the image, modifies the spatial dimension and the number of channels of the image, and determines whether to directly input it into the decoder according to the spatial dimension and the number of channels; if it is not directly input into the decoder, it is input into the skip connection block and transmitted to the decoder block through the skip connection; otherwise, it is directly input into the decoder block for processing, and the processed image is input into the upsampling block with an activation function to restore the image;

[0010] Train the constructed medical image segmentation model to obtain a trained medical image segmentation model;

[0011] Use the trained medical image segmentation model to segment medical images.

[0012] Further, the preprocessing includes dividing the input image into a number of non-overlapping image blocks, and the size of each image block is equal.

[0013] Further, positional encoding is added to the PVTv2, and the positional encoding vector p i is added to the embedding vector z of each image block i and is generated using sine and cosine functions. The final image block representation is:

[0014]

[0015] where is the embedding vector containing position information.

[0016] Furthermore, the PVTv2 uses an improved multi-layer Transformer to extract features of an image; each layer of the Transformer includes a self-attention mechanism and a feed-forward neural network; an embedding vector containing position information After passing through the self-attention mechanism, multi-scale feature fusion, and the feed-forward neural network, the image is aggregated at different scales through a feature pyramid network.

[0017] Furthermore, the graph convolutional block with bias specifically includes:

[0018] The graph convolutional block updates the node features and captures the structural information of the graph; for each feature X i , the result Z G obtained after the operation of the graph convolutional block is:

[0019] Z G = BN(Conv 1×1 (BN(GConv(Conv 1×1 (X i )))))

[0020] where Conv 1×1 is a 1×1 convolution operation, BN is batch normalization, and GConv is a graph convolution operation.

[0021] Furthermore, the multi-splicing spatial attention specifically includes:

[0022] The multi-splicing spatial attention performs weighted sum adjustment on the result obtained after the operation of the graph convolutional block; the result obtained after the operation of the graph convolutional block is spliced along the maximum, minimum, and average values on the channel dimension:

[0023] Attention = Sigmoid(Conv(C(Z G )))*Z G ;

[0024] where C(Z G ) is the result after concatenating the maximum, minimum, and average values on the channel dimension, Conv(·) is a convolution kernel of 7×7, * is the Hadamard product; Sigmoid is an activation function; Attention is an attention operation;

[0025] After the result Z G obtained after the operation of the graph convolutional block enters the multi-splicing spatial attention, the resulting Z A is:

[0026] Z A = Z G ⊙Attention;

[0027] Among them, ⊙ represents an element-wise multiplication operation.

[0028] Furthermore, the upsampling block with an activation function specifically includes:

[0029] The formula for constructing the activation function is as follows:

[0030]

[0031] The result Z obtained by the upsampling block U The formula is:

[0032] Z U = Conv 1×1 (MA(BN(DWC(Up(Z A )))));

[0033] Among them, Up is the upsampling operation, DWC is the depthwise separable convolution operation, BN is the batch normalization, and MA is the activation function.

[0034] Furthermore, training the constructed medical image segmentation model to obtain a trained medical image segmentation model includes: using the gradient descent optimization algorithm for training; passing the input image into the constructed medical image segmentation model to obtain a prediction result, calculating the loss value according to the output of the medical image segmentation model and the ground truth label; calculating the gradient and updating the weights of the medical image segmentation model, and continuously optimizing the parameters of the medical image segmentation model within multiple training epochs until the loss function converges or reaches a preset stopping condition.

[0035] Furthermore, during the training of the constructed medical image segmentation model, the DICE coefficient loss is adopted as the loss function.

[0036] Furthermore, during the training of the constructed medical image segmentation model, the HD95 metric is used to focus on the correct display ratio of pixel gray values in the image, and the IoU metric represents the ratio of the intersection to the union of the prediction result and the ground truth for each class of the medical image segmentation model.

[0037] Compared with the prior art, the present invention has the following technical effects:

[0038] In order to achieve higher-precision segmentation, the present invention adopts graph convolution with matrix position offset. This method adds position offset to the matrix of the graph convolution network, optimizes the processing efficiency of the graph convolution block, enables the graph neural network to more accurately capture the associations between nodes, and enhances the modeling ability of complex graph structures. At the same time, an upsampling block with an innovative activation function is introduced. The addition of the new activation function enhances the non-linear relationship between input features, endows the model with stronger representation ability, and helps the model learn complex patterns and features in the data. In addition, the network also adopts a multi-concatenation spatial attention mechanism. By maximizing the fusion of channel information, it improves the model's attention to the features of the input block and enhances the model's perception ability of the correlation between different features. Finally, the present invention uses PVT v2 as the encoder, combines graph convolution with matrix position offset and the multi-concatenation spatial attention mechanism as the decoder, and combines them to form an efficient medical image segmentation model. Description of the Drawings

[0039] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is the flowchart of the present invention;

[0041] Figure 2 It is the shape diagram of the activation function proposed by the present invention;

[0042] Figure 3 It is the schematic diagram of the graph convolution process with position bias proposed by the present invention;

[0043] Figure 4 It is the model diagram of the present invention, where (a) is the backbone diagram, (b) is the upsampling block, (c) is the decoder block, (d) is the graph convolution block, and (e) is the multi-concatenation spatial attention block;

[0044] Figure 5 It is the schematic diagram of PVT used by the present invention. Detailed Embodiments

[0045] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes the specific implementation manners, structures, features and their effects of the technical solutions proposed according to the present invention in combination with the accompanying drawings and preferred embodiments. Specific features, structures or characteristics in one or more embodiments may be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0046] Referring to Figures 1-5 , in an embodiment of the present invention, a medical image segmentation method combining graph convolution and attention is provided, including the following steps:

[0047] Preprocess the input image; and construct a medical image segmentation model, which includes an encoder, a decoder and a skip connection block, and the skip connection block is connected between the encoder and the decoder; the encoder includes PVTv2; the decoder includes a decoder block and an upsampling block, and the decoder block includes a graph convolution block with bias and multi-splicing spatial attention.

[0048] The preprocessed image passes through PVTv2, extracts the features of the image, modifies the spatial dimension and the number of channels of the image, and determines whether to directly input it into the decoder according to the spatial dimension and the number of channels; if not directly input into the decoder, it is input into the skip connection block and transferred to the decoder block through the skip connection; otherwise, it is directly input into the decoder block for processing, and the processed image is input into the upsampling block with an activation function to restore the image.

[0049] Train the constructed medical image segmentation model to obtain a trained medical image segmentation model.

[0050] Use the trained medical image segmentation model to segment medical images.

[0051] The following details each of the above steps:

[0052] Step 100: Preprocess the input image.

[0053] For the input image Divide it into several non-overlapping image patches, and the size of each image patch is P×P:

[0054]

[0055] where H represents the height of the image, W represents the width of the image, C is the number of channels, and N is the total number of image patches.

[0056] After each block is flattened, it is mapped to an embedding space of a fixed dimension. Suppose the embedding vector of each image block is z i , where i is the index of the block.

[0057] Step 200: Construct a medical image segmentation model. The medical image segmentation model includes an encoder, a decoder, and skip connection blocks. The skip connection blocks are connected between the encoder and the decoder; the encoder includes PVTv2; the decoder includes decoder blocks and upsampling blocks. The decoder blocks include graph convolutional blocks with biases and multi-splicing spatial attention.

[0058] The preprocessed image passes through PVTv2 to extract the features of the image, modify the spatial dimension and the number of channels of the image, and determine whether to directly input it into the decoder according to the spatial dimension and the number of channels; if it is not directly input into the decoder, it is input into the skip connection block and transmitted to the decoder block through the skip connection; otherwise, it is directly input into the decoder block for processing, and the processed image is input into the upsampling block with an activation function to restore the image.

[0059] Specifically, the features of the image after passing through PVTv2 are already suitable enough for the decoder to perform further processing, that is, the features of the encoder and the decoder are sufficiently matched in channels and space, and direct splicing can ensure dimension alignment and retain more original information, then it can be directly input into the decoder; otherwise, the low-level features extracted from the encoder stage are first input into the skip connection block, and through the skip connection block, the information of the encoder is transmitted to the decoder block.

[0060] In this embodiment, referring to Figure 5 , in order to ensure the order of the spatial key information in the image, positional encoding is added to PVTv2. The positional encoding vector p i is added to the embedding vector z of each image block i , which is generated using sine and cosine functions. The final representation of the image block is:

[0061]

[0062] where is the embedding vector containing positional information.

[0063] PVTv2 uses an improved multi-layer Transformer encoder to further extract the features of the image. Each layer of the Transformer includes a self-attention mechanism and a feed-forward neural network. The embedding vector containing positional information, after passing through the self-attention mechanism, multi-scale feature fusion, and feed-forward neural network, aggregates the image from different scales through a feature pyramid network, and this process is achieved through downsampling.

[0064] After passing through the encoder once, the spatial dimension of the image is halved, and the number of channels becomes twice the original. The specific dimension changes are as follows:

[0065]

[0066] Among them, H represents the height of the image, W represents the width of the image, and C is the number of channels.

[0067] In this embodiment, referring to Figure 4 , the decoder includes a decoder block and an upsampling block; the decoder block is used to process the encoded features; the upsampling block is used to map the features output by the decoder block to a higher resolution. Among them, the decoder block includes a graph convolutional block with bias and multi-concatenated spatial attention; the upsampling block includes an activation function.

[0068] The node feature matrix of the graph convolutional block is where N is the number of nodes and C is the feature dimension. The adjacency matrix is where a i,j = 1 indicates that there is an edge between node i and node j, otherwise a i,j = 0, then the form of the graph convolutional block operation is:

[0069] H c = σ(AXW);

[0070] Among them, W is the weight matrix and σ is the activation function.

[0071] The graph convolutional block first calculates the relative position pij between each pair of nodes i and j, and then obtains the corresponding position matrix Set a bias matrix B = bE, combine the matrix P and the matrix B to obtain a position matrix with bias Obtain the relative position p' ij .

[0072] The obtained p' ij is added to the node feature aggregation, so that the aggregation operation depends not only on the node features but also on the relative positions between nodes. For node i, its original graph convolutional aggregation function is:

[0073]

[0074] After adding p' ij , the graph convolutional aggregation function of node i is:

[0075]

[0076] where, h i ' is the output feature of node i, N(i) is the set of neighbor nodes of node i, p'ij is the relative position with bias between node i and node j, W is the weight matrix of graph convolution, and σ is the activation function.

[0077] Then the corresponding graph convolution block operation is:

[0078] H c = σ(A(X + P')W);

[0079] For each feature X i , the result Z G obtained after the graph convolution operation is:

[0080] Z G = BN(Conv 1×1 (BN(GConv(Conv 1×1 (X i )))));

[0081] Among them, Conv 1×1 is a 1×1 convolution operation, BN is batch normalization processing, and GConv is a graph convolution operation.

[0082] After passing through the graph convolution block, it is passed into the corresponding multi - concatenated spatial attention. The multi - concatenated spatial attention concatenates the results obtained after the graph convolution block operation along the maximum, minimum, and average values on the channel dimension:

[0083] Attention = Sigmoid(Conv(C(Z G )))*Z G ;

[0084] Among them, C(Z G ) is the result after concatenating the maximum, minimum, and average values on the channel dimension, Conv(·) is a convolutional layer with a convolution kernel of 7×7 and padding of 3, and * is the Hadamard product.

[0085] In the multi - concatenated spatial attention, first, the most significant eigenvalue is extracted through the max - pooling operation, which helps to emphasize the most representative features in each region, enabling the model to focus more on important local details; at the same time, the least significant features are removed using the min - pooling operation. In this way, those subtle differences and important details that may be ignored in the overall features can be identified, enabling the model to more comprehensively understand the spatial structure of the input tensor; then the average - pooling operation is used to calculate the average value of the input tensor along the channel dimension, thereby obtaining the overall information and providing the model with a global understanding of the input data; finally, the three are concatenated and sent to the convolution for processing, and the final spatial attention weight is obtained through the sigmoid activation function.

[0086] The result Z obtained after the graph convolution block operation G After entering the multi-concatenated spatial attention module, the obtained result Z A is:

[0087] Z A = Z G ⊙ Attention;

[0088] where Attention is the attention operation, and ⊙ represents the element-wise multiplication operation.

[0089] The processed image is fed into the upsampling block, and the sampling block includes an activation function. The activation function plays an important role in improving the model accuracy, accelerating the training speed, and reducing the model calculation amount. Therefore, it is particularly important to construct an activation function suitable for the model.

[0090] In this embodiment, using the LeakyRelu activation function has more advantages in improving the model accuracy than other conventional activation functions. However, since it is discontinuous during backpropagation, it will affect the stability of model convergence. Therefore, a four-segment activation function MA(x) constructed by polynomial curves is designed to maintain the advantages of LeakyRelu and overcome its defects. This activation function is similar in shape to LeakyRelu and satisfies C 1 continuity. The specific construction method is as follows: First, in the intervals [-1,0), (0,1], two cubic polynomial functions M i (x), i = 1, 2, are constructed using the cubic Hermite function and defined as follows:

[0091]

[0092] where,

[0093] α0(x) = (x i+1 - x) 2 (2(x - x i ) + h) / h 3

[0094] α1(x) = (x - x i ) 2 (2(x i+1 - x) + h) / h 3

[0095] β0(x) = (x i+1 - x) 2 (x - x i ) / h 2

[0096] β1(x) = -(x - x i )2 (x i+1 - x) / h 2

[0097] h = x i+1 - x i ;

[0098] Based on the similarity principle with LeakyRelu, M i (x) has function values equal to LeakyRelu at points x1 = -1, x2 = 0, x3 = 1, which are M1(-1) = -0.15, M1(0) = M2(0) = 0, M2(1) = 1 respectively, and the corresponding first-order derivatives are

[0099] From the above function values and derivative values, it can be obtained that:

[0100] M1(x) = ((0.1125x + 0.3375)x + 0.375)x

[0101] M2(x) = ((-1.125x + 1.75)x + 0.375)x

[0102] Then in the intervals (-∞, -1), (1, +∞), two segments of linear polynomial functions are constructed. M A (x) has a function value of 0.15 at the point x = -1 and a first-order derivative of F' = 0.0375.

[0103] At the point x = 1, the function value is 1 and the first-order derivative is F' = 0.5. Based on the above values, it can be obtained that:

[0104] M0(x) = 0.0375x - 0.1125

[0105] M3(x) = 0.5x - 0.5

[0106] Therefore, it can be obtained that:

[0107]

[0108] The result Z obtained after passing through the upsampling block U The formula is:

[0109] Z U = Conv 1×1 (MA(BN(DWC(Up(Z A ))))) ;

[0110] Among them, Up is the upsampling operation, DWC is the depthwise separable convolution operation, BN is the batch normalization, and MA is the activation function proposed by this method.

[0111] After passing through the last set of decoder blocks and upsampling blocks, the generated image will be restored to the same dimensions and channels as the input image. The specific dimensional changes are as follows:

[0112]

[0113] Step 300: Train the constructed medical image segmentation model to obtain a trained medical image segmentation model.

[0114] Use the gradient descent optimization algorithm for training; pass the input image into the constructed medical image segmentation model to obtain a prediction result, calculate the loss value based on the output of the medical image segmentation model and the ground truth label; calculate the gradient and update the weights of the medical image segmentation model, and continuously optimize the parameters of the medical image segmentation model within multiple training epochs until the loss function converges or reaches a preset stopping condition.

[0115] The DICE coefficient is a metric that measures the similarity between the predicted segmentation result and the ground truth segmentation result. The larger the value, the more similar the predicted segmentation result is to the ground truth segmentation result. The formula is:

[0116]

[0117] where A is the set of pixels in the predicted image and B is the set of pixels in the ground truth label.

[0118] The HD95 metric is used to focus on the correct display ratio of pixel gray values in the image. The smaller the value, the better the segmentation effect. The formula is:

[0119]

[0120] where X is the ground truth and Y is the predicted segmentation value.

[0121] The IoU metric is used to represent the ratio of the intersection to the union of the prediction result and the ground truth for each class of the model. The formula is:

[0122]

[0123] where X is the ground truth and Y is the predicted segmentation value.

[0124] Step 400: Use the trained medical image segmentation model to segment medical images.

[0125] In order to achieve higher-precision segmentation, the present invention adopts graph convolution with matrix position offset. This method adds position offset to the matrix of the graph convolution network, optimizes the processing efficiency of the graph convolution block, enables the graph neural network to more accurately capture the correlations between nodes, and enhances the modeling ability of complex graph structures. At the same time, an upsampling block with an innovative activation function is introduced. The addition of the new activation function enhances the non-linear relationship between input features, endows the model with stronger representation ability, and helps the model learn complex patterns and features in the data. In addition, a multi-splicing spatial attention mechanism is also adopted. By maximizing the fusion of channel information, it improves the model's attention to the input block features and enhances the model's perception ability of the correlations between different features. Finally, the present invention uses PVT v2 as the encoder, combines graph convolution with matrix position offset and the multi-splicing spatial attention mechanism as the decoder, and combines them to form an efficient medical image segmentation model.

[0126] Referring to Tables 1 - 3, it can be seen from the experimental results that on the Synapse dataset, the segmentation model proposed by the present method shows higher performance compared with other methods in terms of the aorta, left kidney, right kidney, liver, pancreas, stomach, average DICE, and HD95 metrics; on the ACDC dataset, the segmentation model proposed by the present method has higher segmentation results than other models in terms of average DICE, left ventricle, right ventricle, and myocardium; on the ISIC2018 dataset, the model proposed by the present method shows good results in terms of average DICE and mIoU metrics. Generally speaking, by combining graph convolution with matrix position offset and the multi-splicing spatial attention mechanism to form a decoder block and then combining it with the upsampling block with an innovative activation function and the PVTv2 module to form an overall model, the above excellent experimental results are achieved.

[0127] Table 1: Comparison of the proposed method on the Synapse dataset. DSC shows abdominal organs spleen (Sp), right kidney (RK), left kidney (LK), gallbladder (GB), liver (Liver), stomach (SM), aorta (Aorta), and pancreas (PC)

[0128]

[0129]

[0130] Table 2: Comparison of the proposed method on the ACDC dataset, RV represents the right ventricle, Myo represents the myocardium, and LV represents the left ventricle

[0131] Methods Avg DICE RV Myo LV TransUNet 89.71 86.67 87.27 95.18 MISSFormer 90.86 89.55 88.04 94.99 PVT-CASCADE 91.46 89.97 88.90 95.50 PVT-GCASCADE 91.78 90.18 89.47 95.68 Swin-Unet 88.07 85.77 84.42 94.03 R50+UNet 87.55 87.10 80.63 94.92 MT-UNet 90.43 86.64 89.04 95.62 TransCASCADE 91.63 80.25 89.14 95.50 Ours 91.93 90.25 89.58 95.94

[0132] Table 3: Comparison of the proposed method on the ISIC2018 dataset

[0133]

[0134] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A medical image segmentation method combining graph convolution and attention, characterized in that, It includes the following steps: Preprocess the input image; And construct a medical image segmentation model, which includes an encoder, a decoder, and a skip connection block. The skip connection block is connected between the encoder and the decoder. The encoder includes PVTv2. The decoder includes a decoder block and an upsampling block. The decoder block includes a graph convolutional block with bias and multi-splicing spatial attention; The preprocessed image passes through PVTv2 to extract the features of the image, modify the spatial dimension and the number of channels of the image, and determine whether to directly input it into the decoder based on the spatial dimension and the number of channels. If it is not directly input into the decoder, it is input into the skip connection block and transmitted to the decoder block through the skip connection; Otherwise, it is directly input into the decoder block for processing, and the processed image is input into the upsampling block with an activation function to restore the image; Train the constructed medical image segmentation model to obtain a trained medical image segmentation model; Use the trained medical image segmentation model to segment medical images.

2. The medical image segmentation method combining graph convolution and attention according to claim 1, characterized in that The preprocessing includes dividing the input image into a number of non-overlapping image patches, and the size of each image patch is equal.

3. A medical image segmentation method combining graph convolution and attention according to claim 2, characterized in that Position encoding is added to the PVTv2, and the position encoding vector p i is added to the embedding vector z of each image patch i is generated using sine and cosine functions, and the final image patch representation is: Among them, is an embedding vector containing location information.

4. The medical image segmentation method combining graph convolution and attention according to claim 3, wherein The PVTv2 uses an improved multi-layer Transformer to extract features of an image; each layer of the Transformer includes a self-attention mechanism and a feed-forward neural network; an embedding vector containing position information After passing through the self-attention mechanism, multi-scale feature fusion, and the feed-forward neural network, the image is aggregated at different scales through a feature pyramid network.

5. A medical image segmentation method combining graph convolution and attention, as claimed in claim 1, wherein The graph convolutional block with bias specifically includes: The graph convolutional block updates the node features and captures the structural information of the graph; for each feature X i , the result Z G obtained after the graph convolutional block operation is as follows: Z G = BN(Conv 1×1 (BN(GConv(Conv 1×1 (X i ))))); Among them, Conv 1×1 is a 1×1 convolution operation, BN is batch normalization processing, and GConv is a graph convolution operation.

6. A method for medical image segmentation by mixing graph convolution and attention according to claim 5, characterized in that, The multi-splicing spatial attention specifically includes: The multi-splicing spatial attention performs weighted sum adjustment on the result obtained after the graph convolutional block operation; the result obtained after the graph convolutional block operation is spliced along the maximum value, minimum value, and average value in the channel dimension: Attention=Sigmoid(Conv(C(Z G )))*Z G ; Among them, C(Z G ) is the result after concatenating the maximum value, minimum value, and average value on the channel dimension. Conv(·) has a convolution kernel of 7×7, * is the Hadamard product; Sigmoid is the activation function; Attention is the attention operation; The result Z obtained after the graph convolution block operation G After entering the multi-concatenated spatial attention, the resulting Z A is as follows: Z A = Z G ☉Attention; Among them, ⊙ represents the element-wise multiplication operation.

7. A medical image segmentation method combining graph convolution and attention according to claim 6, characterized in that The upsampling block with an activation function specifically includes: The formula for constructing the activation function is as follows: The result Z obtained by the upsampling block U The formula is: Z U = Conv 1×1 (MA(BN(DWC(Up(Z A ))))); Among them, Up is the upsampling operation, DWC is the depthwise separable convolution operation, BN is the batch normalization, and MA is the activation function.

8. A multi-attention-guided medical image segmentation method according to claim 1, characterized in that Training the constructed medical image segmentation model to obtain a trained medical image segmentation model includes: using the gradient descent optimization algorithm for training; transmitting the input image into the constructed medical image segmentation model to obtain a prediction result, calculating the loss value according to the output of the medical image segmentation model and the ground truth label; calculating the gradient and updating the weights of the medical image segmentation model, and continuously optimizing the parameters of the medical image segmentation model within multiple training cycles until the loss function converges or reaches the preset stop condition.

9. A multi-attention-guided medical image segmentation method according to claim 7, wherein During the training of the constructed medical image segmentation model, the DICE coefficient loss is used as the loss function.

10. A multi-attention-guided medical image segmentation method according to claim 7, wherein During the training of the constructed medical image segmentation model, the HD95 metric is used to focus on the correct display ratio of the pixel gray values in the image, and the IoU metric represents the ratio of the intersection to the union of the prediction result and the ground truth value of each class by the medical image segmentation model.