Image segmentation method based on dynamic and static attention aggregation mechanism

Through the dynamic and static attention aggregation mechanism, an image encoder and prompt encoder are constructed, and combined with dynamic and static attention calculation and efficient decoder structure, the problems of attention sparseness and model overfitting in the existing technology are solved, achieving higher image segmentation accuracy and robustness.

CN120388030APending Publication Date: 2025-07-29GUIZHOU UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510511612.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing image segmentation method has sparse attention weights when processing sequences are long, making it difficult to effectively capture long-distance dependencies and feature context information. The complex decoder structure leads to overfitting of the model and reduces segmentation accuracy.

Method used

The dynamic and static attention aggregation mechanism is adopted to build an image encoder and prompt encoder, combining dynamic and static attention calculation and convolution aggregation, enhance the feature information capture capability, and design an efficient decoder structure, including efficient upsampling blocks and large-core convolution space activation layer, to improve the robustness and segmentation accuracy of the model.

Benefits of technology

It improves the accuracy and robustness of image segmentation, enhances the model's ability to capture feature information, and achieves higher segmentation accuracy and generalization performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388030A_ABST
    Figure CN120388030A_ABST
Patent Text Reader

Abstract

The invention discloses an image segmentation method based on a dynamic and static attention aggregation mechanism. The method comprises the following steps: collecting image data; constructing an image segmentation model, wherein the image segmentation model specifically comprises two types of encoders and a mask decoder; the two types of encoders are a prompt encoder and an image encoder, the prompt encoder encodes a point and a frame, and the image encoder performs feature extraction on an input image so as to encode; the mask decoder carries out unified decoding on the two types of encoders to obtain a mask graph of a corresponding image; and inputting a to-be-segmented image into the trained image segmentation model, and further outputting a mask pattern corresponding to the image to complete image segmentation. The method has the characteristic that the image segmentation precision is improved by enhancing the capturing capability of the model on the feature information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image segmentation method based on a dynamic and static attention aggregation mechanism. Background Art

[0002] In the field of image processing, image segmentation technology plays a crucial role. Its core function is to accurately identify and extract specific structures or regions from images. Most image segmentation methods mainly adopt an encoder-decoder structure. Since the release of the U-Net model, many model structures have incorporated skip connections. With the success of the self-attention mechanism of the classic Transformer, many models have improved and designed many new attention mechanisms according to their own task characteristics. In the prior art, DAUNet introduced a feature fusion guidance module to adaptively integrate features from two branches. MCAFNet designed a cross-layer attention fusion (CAF) module to capture multi-scale features by combining channel and spatial information of different feature layers. Wang Yongqi et al. proposed a segmentation algorithm that improved Ostu double threshold and watershed fitting circles. Fu Jinhao et al. proposed a segmentation method that combined local contrast threshold and halo removal. Yang Yudan et al. proposed a U-shaped network based on three-dimensional cyclic residual convolution to segment spinal CT images. However, when encountering the situation of a long sequence, the attention weights may become sparse in these attention mechanisms, resulting in the model being difficult to effectively capture long-distance dependencies, and at the same time, it cannot effectively obtain the connection between feature contexts.

[0003] There are also methods for improving the decoder. RPDNet proposed a reconstruction regularization parallel decoder network. DAN-PD designed a teacher parallel encoder (TPE) and a domain-aware parallel decoder (DAPD). GFANet uses a context feature gating fusion decoder to fuse multi-level context features. However, an overly complex decoder structure may cause the model to over-rely on a specific data distribution, and at the same time, it will reduce the interpretability of the model and increase the risk of overfitting. In short, the attention mechanisms in the existing image segmentation methods are difficult to obtain context information, and the complex decoder structure will cause the model to overfit, resulting in problems such as low segmentation accuracy. Summary of the Invention

[0004] The object of the present invention is to overcome the above-mentioned drawbacks and propose an image segmentation method based on a dynamic and static attention aggregation mechanism, which improves the accuracy and robustness of the segmentation model by enhancing the model's ability to capture feature information, thereby improving the image segmentation accuracy.

[0005] An image segmentation method based on a dynamic and static attention aggregation mechanism of the present invention, wherein: the specific steps of the method include:

[0006] Step 1: Data collection: Collect the image segmentation dataset;

[0007] Step 2: Model construction: Construct an image segmentation model, including two types of encoders and a mask decoder; the two types of encoders are the prompt encoder and the image encoder. The prompt encoder encodes points and boxes. The image encoder extracts features from the input image and then encodes them. The mask decoder decodes the two types of encoders uniformly to obtain the mask map of the corresponding image;

[0008] The prompt encoder encodes points and bounding boxes. Each point is represented as a vector that contains the position information, foreground embedding, and background embedding of the point. Each bounding box is defined by the position encoding of its upper left and lower right corners, and the upper left and lower right corners are represented as learned embeddings.

[0009] The image encoder: consists of a patch embedding layer and multiple feature extraction and encoding blocks DaSA-EFA; the patch embedding layer: for the input image X, this patch embedding layer performs patch embedding encoding, then normalizes to obtain a series of convolutional projections, and finally enters the feature extraction and encoding block DaSA-EFA for feature extraction;

[0010] The feature extraction and encoding block DaSA-EFA: consists of a dynamic and static attention aggregation layer DaSA Layer, an efficient feature aggregation layer EFA Layer, an adapter layer Adapter Layer, and a multi-layer perceptron layer MLP;

[0011] The dynamic and static attention aggregation layer DaSA Layer includes two parts. One part is for the dynamic and static attention calculation to enhance the model's ability to obtain long-range dependencies, and the other part is for the convolutional aggregation calculation to capture more extensive context information. For a series of input convolutional projections, the feature extraction and encoding block performs the two parts of the calculation in parallel. The formula is:

[0012]

[0013] where, represents the result obtained through the dynamic and static attention calculation, represents the result obtained through the convolutional aggregation calculation. DaS represents the dynamic and static attention calculation, A_Conv represents the convolutional aggregation calculation, and Q, K, and V respectively represent the query, key, and value matrices of the input image X;

[0014] The efficient feature aggregation layer (EFA Layer): Processes image features using dilated convolutions of multiple scales. The result output from the dynamic and static attention aggregation layer is X2, which is passed as input to this efficient feature aggregation layer. First, it is preprocessed through layer normalization, and then through dilated convolution operations of multiple scales to extract features of different scales and types in parallel. A 1×1 convolution is used to further aggregate the results of the convolution, and finally an intermediate feature map is obtained. Subsequently, an addition operation is used to sum the results of different convolutional layers, and then a two-dimensional convolution operation is used again to obtain the final result, which is connected to the original features through a residual connection to obtain the result. The formula is:

[0015]

[0016] The adapter layer (Adapter Layer): Flexibly adjusts the encoder in the channel and spatial dimensions. Global average pooling is used to reduce the resolution of the input feature map, and then a linear layer is used to compress the feature representations of these channels. After that, another linear layer is used to restore its size. Finally, the Sigmoid function is used to generate the weights in the channel dimension, and these weights are multiplied by the input feature map as the input to the next layer.

[0017] The mask decoder: Consists of an efficient upsampling block (EUB) and a large kernel convolutional spatial activation layer (LKSA Layer); The efficient upsampling block (EUB): Consists of multiple groups of transposed convolutions and layer normalization layers; The large kernel convolutional spatial activation layer (LKSA Layer): Consists of grouped dilated convolutions, batch normalization, ReLU activation function, and Sigmoid function.

[0018] Step 3: Train the model: Use the dataset to train the image segmentation model. By means of weighted sum, three types of loss functions, Focal Loss, Dice Loss, and Mask IoU Loss, are combined and used as the final loss function Loss of the model. The specific calculation method is as follows:

[0019] Loss = λ·FL(P, G)+DL(P, G)+IoU Scale·IoU Loss(P, G, pred_iou)

[0020] Among them, FL represents the calculation of the loss function Focal Loss, DL represents the calculation of the loss function Dice Loss, IoU Loss represents the calculation of the loss function Mask IoU Loss, λ represents the weighted value of the Focal Loss of the batch normalization layer, IoU Scale represents the weight that controls the loss function Mask IoU Loss in the total loss, P represents the predicted mask, and G represents the ground truth mask.

[0021] Step 4: Input the image to be segmented into the trained image segmentation model, and then output the mask image corresponding to the image to complete the image segmentation.

[0022] In the above image segmentation method based on the dynamic and static attention aggregation mechanism, in step 2, the specific process of calculating the dynamic and static attention is as follows: First, apply k×k groups of convolutions to all adjacent keys within the grid space k×k to simplify and condense into each key representation. The key K1 obtained through this operation is then concatenated with the query Q, and then through two consecutive 1×1 convolution operations, the attention matrix map M is obtained. The formula is:

[0023] M = [K1, Q]W θ W δ

[0024] where W θ is a convolution operation with a piecewise linear unit activation function, and W δ is a convolution operation without an activation function;

[0025] According to the attention matrix M, calculate the dynamic attention feature map K2 by aggregating all values V. The formula is:

[0026] K2 = M×V

[0027] In the above image segmentation method based on the dynamic and static attention aggregation mechanism, in step 2, the specific process of the convolution aggregation calculation is as follows: First, perform a 1×1 convolution operation, and then obtain features from two perspectives of max pooling and average pooling to obtain the feature information K3 containing rich context:

[0028]

[0029] where represents the intermediate variable of the calculation, Conv represents the convolution operation, Norm represents the normalization operation, GELU represents the activation operation, Avg and Max respectively represent the average pooling and max pooling operations; finally, in an additive combination manner, all the obtained information is fused and aggregated into the feature information.

[0030] In the above image segmentation method based on the dynamic and static attention aggregation mechanism, in step 2, the spatial dimension of the adapter layer Adapter Layer is reduced to one-fourth of the original spatial resolution of the feature map by using a convolutional layer, and then the transposed convolution is used to restore the spatial resolution to ensure the same number of channels as the input.

[0031] The above-mentioned image segmentation method based on the dynamic and static attention aggregation mechanism, wherein: in step two, the transposed convolution of the efficient upsampling block EUB is used to upsample the feature map and increase its spatial resolution;

[0032] The above-mentioned image segmentation method based on the dynamic and static attention aggregation mechanism, wherein: in step two, the layer normalization layer of the efficient upsampling block EUB is used to stabilize the training process and improve the convergence speed of the network. Subsequently, the gated linear unit activation function GELU is used to introduce non-linearity. Finally, a 1×1 convolution is used to reduce the number of channels and fuse the features to obtain the output result EUB(x) of the efficient upsampling block. The specific formula is as follows:

[0033] EUB(x) = C 1×1 (ConvT(GELU(Norm(ConvT(x))))

[0034] where C 1×1 represents a 1×1 convolution, ConvT represents a transposed convolution, GELU represents an activation function, and Norm represents normalization.

[0035] The above-mentioned image segmentation method based on the dynamic and static attention aggregation mechanism, wherein: in step two, the grouped dilated convolution of the large kernel convolutional spatial activation layer LKSA Layer extracts feature information from different directions by two groups of dilated convolutions respectively. Subsequently, batch normalization is used to standardize the data distribution. The two feature maps are integrated together by element-wise addition, and a ReLU activation function is used for non-linear activation. Then, further calculations are performed through 1×1 convolution processing and the batch normalization layer BN to obtain a single-channel feature map, and it is converted into an attention coefficient through the Sigmoid function. Finally, the calculated coefficient is multiplied by the original input to obtain the final output result LKSA(x) of the spatial activation layer. The specific formula is as follows:

[0036] LKSA(x) = x × σ(BN(C 1×1 (RELU(BN(GDC(x)) + BN(GDC(x)))))

[0037] where σ represents the Sigmoid operation, BN represents batch normalization, C 1×1 represents a 1×1 convolution, ReLU represents an activation function, and GDC represents the grouped dilated convolution operation.

[0038] Compared with the prior art, the present invention has obvious beneficial effects. As can be seen from the above solution, the main body of the present invention consists of two types of encoders (an image encoder and a prompt encoder) and a mask decoder. First, for the input two-dimensional image, the image encoder extracts features from it. The image encoder includes a patch embedding layer and N DaSA-EFA feature extraction and encoding blocks. Among them, DaSA is used to obtain dynamic and static feature information, and a convolutional aggregation branch is added to obtain rich context feature information. EFA is used to further process the output of the DaSA part. Multiple groups of dilated convolutions are used to process the input feature information. While expanding the receptive field, it can effectively perform multi-scale feature extraction. Secondly, the adapter technology is introduced to keep it in a learnable state, enabling the model to quickly focus on task-specific features during training and converge quickly. By updating the adapter module, the core structure and parameters of the pre-trained model remain unchanged, thus avoiding catastrophic forgetting and further improving the generalization performance of the model. Then, for the prompt encoder, it mainly encodes points and bounding boxes. Each point is represented as a vector, which contains the position information, foreground embedding, and background embedding of the point. Each bounding box is defined by the position encoding of its upper left and lower right corners, and the upper left and lower right corners are represented as learned embeddings. Finally, in the decoder part, an efficient upsampling layer (EUB Layer) and a large kernel convolutional spatial activation layer (LKSA Layer) are designed and improved to efficiently process the input feature information, making the mask information output by the model have higher accuracy. The main advantages of the present invention:

[0039] 1) Based on the dynamic and static attention aggregation mechanism: A dynamic and static attention aggregation module (DaSA) is designed for the encoder. By constructing a convolutional aggregation branch, rich context information is obtained. To further clarify the effectiveness of feature extraction, an efficient feature aggregation layer (EFA Layer) is constructed, and the output of the DaSA part is subjected to convolutional processing at three scales to achieve accurate position detection and segmentation.

[0040] 2) Construct an efficient activation decoder: An efficient upsampling layer (EUB Layer) and a large kernel convolutional spatial activation layer (LKSA Layer) are designed for the decoder part, extracting attention information from different directions, enabling the model to make more accurate recognition and segmentation of important information.

[0041] In summary, the present invention has the characteristics of improving the image segmentation accuracy by enhancing the model's ability to capture feature information.

[0042] The beneficial effects of the present invention are further described below through specific embodiments. Description of the Drawings

[0043] Figure 1 This is the model structure diagram of the present invention;

[0044] Figure 2 This is the schematic diagram of the image encoder structure in the embodiment of the present invention;

[0045] Figure 3 This is the schematic diagram of the calculation process of the dynamic and static attention aggregation layer in the embodiment of the present invention;

[0046] Figure 4 This is the schematic diagram of the adapter layer structure in the embodiment of the present invention;

[0047] Figure 5 This is the schematic diagram of the mask decoder structure in the embodiment of the present invention. Detailed implementation manners

[0048] The following combines the accompanying drawings and preferred embodiments to detail the specific implementation manners, features and effects of an image segmentation method based on a dynamic and static attention aggregation mechanism proposed according to the present invention.

[0049] See Figure 1 , an image segmentation method based on a dynamic and static attention aggregation mechanism of the present invention, wherein: the specific steps of the method include:

[0050] Step 1: Data collection: Collect an image segmentation data set; the data set is a skin lesion image data set and a thyroid nodule region segmentation data set (TN3K).

[0051] Step 2: Model construction: Design a medical image segmentation model for the constructed large data set, mainly including two types of encoders and a mask decoder. The two encoders are respectively a prompt encoder and an image encoder. Among them, the prompt encoder mainly encodes points and boxes, and the image encoder mainly extracts features from the input image and then encodes. The mask decoder uniformly decodes the two types of encoders and finally obtains the mask map of the corresponding image;

[0052] The prompt encoder encodes points and bounding boxes. Each point is represented as a vector, which contains the position information, foreground embedding and background embedding of the point. Each bounding box is defined by the position encoding of its upper left and lower right corners, and the upper left and lower right corners are represented as learned embeddings.

[0053] As Figure 2 shown, the image encoder: mainly consists of a patch embedding layer and N DaSA-EFA feature extraction and encoding blocks.

[0054] For the patch embedding layer, the input image \(X\in\mathbb{R}^{H\times W\times C}\), the embedding layer will perform patch embedding encoding on it, then perform normalization processing to obtain a series of convolutional projections, and finally enter the DaSA-EFA feature extraction and encoding block for feature extraction.

[0055] The DaSA-EFA feature extraction and encoding block: It consists of a dynamic and static attention aggregation layer (DaSA Layer), an efficient feature aggregation layer (EFA Layer), an adapter layer (Adapter Layer), and a multi-layer perceptron MLP layer.

[0056] The dynamic and static attention aggregation layer (DaSA Layer): It is mainly divided into two parts, that is, the left half branch for dynamic and static attention calculation to enhance the model's ability to obtain long-range dependencies and the right half branch for convolutional aggregation calculation to capture more extensive context information. For a series of input convolutional projections \(X\in\mathbb{R}^{H\times W\times C}\), DaSA performs two parts of calculations in parallel, namely:

[0057]

[0058] Among them, represents the result obtained through dynamic and static attention calculation, represents the result obtained through convolutional aggregation calculation, DaS represents dynamic and static attention calculation, \(A_{Conv}\) represents convolutional aggregation calculation, and \(Q\), \(K\), and \(V\) respectively represent the query, key, and value matrices of the input image \(X\).

[0059] As Figure 3 shown, the specific calculations of the two parts are further described in detail. For the attention calculation of the left half part: Define the key (\(K\)), query (\(Q\)), and value (\(V\)) as \(K = X\), \(Q = X\), \(V = XW\) V . First, apply \(k\times k\) group convolutions to all adjacent keys in the \(k\times k\) grid space to simplify and condense into each key representation, and obtain the key \(K1\in\mathbb{R}^{H\times W\times C}\) through this operation. Subsequently, concatenate \(K1\) and the query \(Q\), and then obtain the attention matrix map \(M\) through two consecutive \(1\times1\) convolution operations (\(W\) θ is a convolution operation with a piecewise linear unit activation function, \(W\) δ is a convolution operation without an activation function):

[0060] \(M=[K1,Q]W\) θ \(W\) δ

[0061] Further referring to the typical self-attention mechanism, according to the attention matrix \(M\), calculate the dynamic attention feature map \(K2\) by aggregating all values \(V\):

[0062] \(K2 = M\times V\)

[0063] For the right - hand side aggregation operation: First, perform a 1×1 convolution operation, and then obtain features from two perspectives of max - pooling and average - pooling to get the feature information K3 with rich context:

[0064]

[0065] Among them, The intermediate variable Conv for calculation represents the convolution operation, Norm represents the normalization operation, GELU represents the activation operation, Avg and Max respectively represent the average - pooling and max - pooling operations. After all the above operations are completed, in an additive combination manner, all the obtained information is fused and aggregated into a kind of feature information.

[0066] The efficient feature aggregation layer (EFA Layer): Processes the image features using dilated convolutions of multiple scales. Let the result output from the DaSA layer be X2∈R^(H×W×C), and it is passed as input to the efficient feature aggregation layer. First, it is pre - processed through layer normalization, and then through dilated convolution operations of multiple scales, that is, D_Conv(Norm(X2)), to parallelly extract features of different scales and types. A 1×1 convolution is used to further aggregate the results of the convolution, and finally an intermediate feature map is obtained. Subsequently, an addition operation is used to sum the results of different convolutional layers, and then a two - dimensional convolution operation is used again to obtain the final result, which is connected with the original feature through a residual connection to obtain the result. That is:.

[0067]

[0068] As Figure 4 shown, the adapter layer: Flexibly adjusts the encoder in the channel and spatial dimensions. It uses global average pooling to reduce the resolution of the input feature map, then applies a linear layer to compress the feature representation of these channels, and then passes through another linear layer to restore its size. Finally, the Sigmoid function is used to generate the weights in the channel dimension, and these weights are multiplied with the input feature map as the input of the next layer. For the spatial dimension, the spatial resolution of the feature map is reduced to one - quarter of the original by using a convolutional layer, and then transposed convolution is used to restore the spatial resolution, ensuring the same number of channels as the input.

[0069] As Figure 5 shown, the mask decoder: Consists of an efficient up - sampling block (EUB) and a large - kernel convolutional spatial activation layer (LKSA Layer).

[0070] The Efficient Upsampling Block (EUB): It consists of multiple groups of transposed convolutions and layer normalization layers. The transposed convolution is used to upsample the feature map and increase its spatial resolution, and the layer normalization is used to stabilize the training process and improve the convergence speed of the network. Subsequently, the Gated Linear Unit (GELU) activation function is used to introduce non-linearity. Finally, a 1×1 convolution is used to reduce the number of channels and fuse the features. The specific formula is as follows:

[0071] EUB(x)0 = C 1×1 (ConvT(GELU(Norm(ConvT(x))))

[0072] where C 1×1 represents the 1×1 convolution, ConvT represents the transposed convolution, GELU represents the activation function, and Norm represents the normalization.

[0073] The Large Kernel Convolution Spatial Activation Layer (LKSA Layer): It consists of grouped dilated convolution, batch normalization, ReLU activation function, and Sigmoid function. Two groups of dilated convolutions extract feature information from different directions respectively. Subsequently, batch normalization is used to standardize the data distribution, and the two feature maps are integrated together by element-wise addition and non-linearly activated using the ReLU activation function. Then, through 1×1 convolution processing and further calculation of the BN layer, a single-channel feature map is obtained and converted into an attention coefficient using the Sigmoid function. Finally, the calculated coefficient is multiplied by the original input to obtain the final result. The specific formula is as follows:

[0074] LKSA(x) = x × σ(BN(C 1×1 (RELU(BN(GDC(x)) + BN(GDC(x)))))

[0075] where σ represents the Sigmoid operation, BN represents batch normalization, C 1×1 represents the 1×1 convolution, ReLU represents the activation function, and GDC represents the grouped dilated convolution operation.

[0076] Step 3: Train the model: Use the dataset constructed previously to train the medical image segmentation network based on the dynamic and static attention aggregation mechanism. Combine three types of loss functions, namely Focal Loss, Dice Loss, and Mask IoU Loss, through weighted sum and use it as the final loss function of the model. The specific calculation method is as follows:

[0077] Loss = λ·FL(P, G) + DL(P, G) + IoU Scale·IoU Loss(P, G, pred_iou)

[0078] Among them, FL represents the calculation of Focal Loss, DL represents the calculation of Dice Loss, IoU Loss represents the calculation of Mask IoU Loss, λ represents the weighting value of Focal Loss, IoU Scale represents the weight controlling Mask IoU Loss in the total loss, P represents the predicted mask, and G represents the ground truth mask.

[0079] Step 4: Input the untrained medical image into the trained medical image segmentation model based on the dynamic and static attention aggregation mechanism, and then output the mask map of the corresponding medical image to complete the medical image segmentation.

[0080] Performance analysis:

[0081] To verify the effectiveness of the present invention, test experiments were carried out on a single NVIDIA 4090 24G GPU, and the experiments showed significant performance improvement.

[0082] Table 1 shows the result analysis of four evaluation metrics of each model on the ISIC-2018 dataset. Among them, the optimal group of data is marked in bold, and the sub-optimal group of data is marked with an underline. From the data in the table, it can be seen that the DoubleU-Net and DS-TransUNet models perform excellently, reaching 89.62% and 91.32% respectively in the mDice metric, and 82.12% and 85.23% in the mIoU metric. However, the model of the present invention achieved better results in the mDice and mIoU metrics, which were 92.85% and 87.08% respectively. Compared with the DS-TransUNet model, the mDice metric increased by 1.53%, and the mIoU metric increased by 1.85%. This shows the excellent segmentation performance of the model of the present invention on the ISIC-2018 dataset and reflects the advantages of the model of the present invention.

[0083] Table 1 Result analysis of the ISIC-2018 dataset on four evaluation metrics: similarity (mDice), mean intersection over union (mIoU), recall, and precision. Among them, the optimal group of data is marked in bold, and the sub-optimal group of data is marked with an underline

[0084]

[0085]

[0086] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. An image segmentation method based on a dynamic and static attention aggregation mechanism, characterized in that: The specific steps of this method include: Step 1: Data collection: Collect an image segmentation dataset; Step 2: Model construction: Construct an image segmentation model, including two types of encoders and a mask decoder; the two types of encoders are a prompt encoder and an image encoder. The prompt encoder encodes points and boxes. The image encoder extracts features from the input image and then encodes them. The mask decoder decodes the two types of encoders uniformly to obtain the mask map of the corresponding image; The prompt encoder encodes points and bounding boxes. Each point is represented as a vector that contains the position information, foreground embedding, and background embedding of the point. Each bounding box is defined by the position encoding of its upper left and lower right corners, and the upper left and lower right corners are represented as learned embeddings; The image encoder: consists of a patch embedding layer and multiple feature extraction and encoding blocks DaSA-EFA; the patch embedding layer: for the input image X, this patch embedding layer performs patch embedding encoding, then normalizes to obtain a series of convolutional projections, and finally enters the feature extraction and encoding block DaSA-EFA for feature extraction; The feature extraction and encoding block DaSA-EFA: consists of a dynamic and static attention aggregation layer DaSA Layer, an efficient feature aggregation layer EFA Layer, an adapter layer Adapter Layer, and a multi-layer perceptron layer MLP; The dynamic and static attention aggregation layer DaSA Layer includes two parts. One part is for the dynamic and static attention calculation to enhance the model's ability to obtain long-range dependencies, and the other part is for the convolutional aggregation calculation to capture more extensive context information. For a series of input convolutional projections, the feature extraction and encoding block performs the two parts of the calculation in parallel. The formula is as follows: Among them, represents the result obtained through dynamic and static attention calculation, represents the result obtained through convolutional aggregation calculation, DaS represents dynamic and static attention calculation, A_Conv represents convolutional aggregation calculation, and Q, K, and V respectively represent the matrices of the query, key, and value of the input image X; The efficient feature aggregation layer (EFA Layer): Processes image features using dilated convolutions of multiple scales. The result output from the dynamic and static attention aggregation layer is X2, which is passed as input to this efficient feature aggregation layer. First, it is preprocessed through layer normalization, and then through dilated convolution operations of multiple scales to parallelly extract features of different scales and types. 1×1 convolution is used to further aggregate the results of the convolution, and finally an intermediate feature map is obtained. Subsequently, an addition operation is used to sum the results of different convolutional layers, and a two-dimensional convolution operation is used once again to obtain the final result, which is connected to the original features through a residual connection to obtain the result. The formula is: The adapter layer Adapter Layer: flexibly adjusts the encoder in the channel and spatial dimensions. It uses global average pooling to reduce the resolution of the input feature map, then applies a linear layer to compress the feature representation of these channels, passes through another linear layer to restore its size, and finally uses the Sigmoid function to generate the weights in the channel dimension, and multiplies these weights with the input feature map as the input of the next layer; The mask decoder: consists of an efficient upsampling block EUB and a large kernel convolutional spatial activation layer LKSA Layer; the efficient upsampling block EUB: consists of multiple groups of transposed convolutions and layer normalization layers; the large kernel convolutional spatial activation layer LKSALayer: consists of grouped dilated convolutions, batch normalization, ReLU activation function, and Sigmoid function; Step 3: Model training: Use the dataset to train the image segmentation model. Combine three types of loss functions, Focal Loss, Dice Loss, and Mask IoU Loss, through weighted sum and use it as the final loss function Loss of the model. The specific calculation method is as follows: Loss = λ·FL(P, G) + DL(P, G) + IoU Scale·IoU Loss(P, G, pred_iou) Among them, FL represents the calculation of the loss function Focal Loss, DL represents the calculation of the loss function Dice Loss, IoU Loss represents the calculation of the loss function Mask IoU Loss, λ represents the weighted value of the Focal Loss of the batch normalization layer, IoU Scale represents the weight that controls the loss function Mask IoU Loss in the total loss, P represents the predicted mask, and G represents the ground truth mask; Step 4: Input the image to be segmented into the trained image segmentation model, and then output the mask image of the corresponding image to complete image segmentation.

2. The image segmentation method based on a dynamic and static attention aggregation mechanism according to claim 1, wherein: In Step 2, the specific process of the dynamic and static attention calculation: First, apply k×k groups of convolutions to all adjacent keys within the grid space k×k to simplify and condense into each key representation. The key K1 obtained through this operation is then concatenated with the query Q, and then through two consecutive 1×1 convolution operations, the attention matrix map M is obtained. The formula is: M = [K1, Q]W θ W δ Among them, W θ is a convolution operation with a piecewise linear unit activation function, and W δ is a convolution operation without an activation function; According to the attention matrix M, calculate the dynamic attention feature map K2 by aggregating all values V. The formula is: K2 = M×V.

3. The image segmentation method based on a dynamic and static attention aggregation mechanism according to claim 1, characterized in that: In Step 2, the specific process of the convolutional aggregation calculation: First, perform a 1×1 convolution operation, and then obtain the feature information K3 with rich context from two perspectives of max pooling and average pooling. The formula is: Among them, represents an intermediate variable for calculation, Conv represents a convolution operation, Norm represents a normalization operation, GELU represents an activation operation, Avg and Max respectively represent average pooling and maximum pooling operations; finally, in an additive combination manner, all the obtained information is fused and converged into feature information.

4. The image segmentation method based on a dynamic and static attention aggregation mechanism according to claim 1, wherein: In Step 2, the spatial dimension of the Adapter Layer is reduced to one-fourth of the original by using a convolutional layer for the spatial resolution of the feature map, and then the transposed convolution is used to restore the spatial resolution to ensure the same number of channels as the input.

5. The image segmentation method based on a dynamic and static attention aggregation mechanism according to claim 1, wherein: In Step 2, the transposed convolution of the Efficient Upsampling Block (EUB) is used to upsample the feature map and increase its spatial resolution.

6. The image segmentation method based on a dynamic and static attention aggregation mechanism according to claim 1, characterized in that: In Step 2, the layer normalization layer of the Efficient Upsampling Block (EUB) is used to stabilize the training process and improve the convergence speed of the network. Subsequently, the gated linear unit activation function GELU is used to introduce non-linearity. Finally, a 1×1 convolution is used to reduce the number of channels and fuse the features to obtain the output result EUB(x) of the efficient upsampling block. The specific formula is as follows: EUB(x) = C 1×1 (ConvT(GELU(Norm(ConvT(x)))) Among them, C 1×1 represents a 1×1 convolution, ConvT represents a transposed convolution, GELU represents an activation function, and Norm represents normalization.

7. A method for image segmentation based on a dynamic and static attention aggregation mechanism according to any one of claims 1 to 6, characterized in that: In Step 2, for the grouped dilated convolution of the Large Kernel Convolution Spatial Activation Layer (LKSA Layer), two groups of dilated convolutions are used to extract feature information from different directions respectively. Subsequently, batch normalization is used to standardize the data distribution. The two feature maps are integrated together by element-wise addition, and the ReLU activation function is used for non-linear activation. Then, further calculation is performed through 1×1 convolution processing and the batch normalization layer BN to obtain a single-channel feature map, and it is converted into an attention coefficient through the Sigmoid function. Finally, the calculated coefficient is multiplied by the original input to obtain the final output result LKSA(x) of the spatial activation layer. The specific formula is as follows: LKSA(x) = x × σ(BN(C 1×1 (RELU(BN(GDC(x)) + BN(GDC(x)))) Among them, σ represents the Sigmoid operation, BN represents batch normalization, C 1×1 represents 1×1 convolution, RELU represents the activation function, and GDC represents the grouped dilated convolution operation.