Image segmentation method based on attention mechanism and multi-dimensional feature fusion
By combining two-dimensional and three-dimensional feature extraction, cross-layer spatial attention, and deep supervision, an image segmentation method based on attention mechanism and multi-dimensional feature fusion is proposed. This method solves the problem that traditional methods cannot meet the high accuracy requirements of specific fields and achieves higher segmentation accuracy and robustness.
Patent Information
- Application Number
- CN202211607936.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-14
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-12-14
AI Technical Summary
Existing image segmentation algorithms, especially traditional methods, struggle to meet the high accuracy requirements of specific domains, and ordinary CNNs are unable to extract global contextual information and improve feature representation capabilities.
An image segmentation method based on attention mechanism and multidimensional feature fusion is adopted. By designing the encoder and decoder, and combining two-dimensional and three-dimensional feature extraction, adaptive fusion, cross-layer spatial attention and deep supervision mechanism, the feature extraction and segmentation results are optimized.
It improves the accuracy of image segmentation, enhances the robustness of the model, solves the limitations of hardware computing power and the vanishing/exploding gradient problem, and improves the ability to extract edge information.
Smart Images

Figure CN116229055B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing and relates to an image segmentation method based on attention mechanism and multidimensional feature fusion. Background Technology
[0002] Current image segmentation algorithms are mainly divided into two categories: one is segmentation methods based on traditional algorithms, and the other is segmentation methods based on deep learning. Traditional segmentation algorithms are further divided into region-based methods and edge-based methods. However, due to the complexity of certain image boundaries and the influence of volumetric effects, traditional methods are difficult to meet the high accuracy requirements of image segmentation in some specific fields.
[0003] Recently, researchers have proposed various deep learning methods based on convolutional neural networks (CNNs) for image segmentation tasks in specific domains, achieving great success. Compared with traditional feature extraction methods, CNNs can train better feature extraction models in a data-driven manner and explore valuable representations suitable for segmentation from the input data. However, ordinary CNNs can only extract local feature information and struggle to learn global context information and improve feature representation capabilities. Therefore, combining multidimensional information and attention mechanisms in the network is crucial. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide an image segmentation method based on attention mechanism and multi-dimensional feature fusion, so as to make full use of the multi-dimensional features of the image and make the segmentation results more accurate.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] An image segmentation method based on attention mechanism and multidimensional feature fusion includes the following steps:
[0007] S1: Acquire and preprocess the image, and randomly select blocks of training samples according to a preset size to form a training set;
[0008] S2: Construct a network model based on multi-dimensional feature information and attention mechanisms, including:
[0009] Encoder: Adopts an adaptive fusion mechanism for 2D and 3D feature extraction, extracts features from multiple axes and dimensions, replaces traditional 3D convolution, and uses 2×2×2 max pooling for pooling layers;
[0010] Decoder: Employs a strategy that combines attention mechanism with multi-scale features and deep supervision, uses trilinear interpolation for resolution restoration, and finally obtains the final segmentation result image through an activation function;
[0011] S3: Parameterize the network, save the model that performs best on the validation data, and validate each data point. After multiple cross-validations, select the optimal model to perform image segmentation.
[0012] Furthermore, the encoder stage specifically includes the following steps:
[0013] Two-dimensional feature maps with maximum activation values on multiple surfaces are obtained by projecting a 3D image through a maximum value mapping.
[0014] Each of them learns its global attention information through a non-local self-attention mechanism;
[0015] The image is restored to 3D using the Unsqueeze operation and then multiplied with the original image.
[0016] MA1 is obtained through an adaptive weighted fusion strategy;
[0017] Feature extraction was performed on the original image using asymmetric convolutions of 1×3×3, 3×1×3, and 3×3×1, respectively.
[0018] MA2 is obtained through an adaptive weighted fusion strategy;
[0019] MA3 is obtained by convolving the original image with a 3×3×3 matrix.
[0020] The final encoder output feature map is obtained by combining MA1, MA2, and MA3 through an adaptive weighted fusion strategy.
[0021] Furthermore, the decoder stage specifically includes the following steps:
[0022] A cross-layer spatial attention module is used to constrain the features of the encoding stage using features from the decoding stage.
[0023] The multi-scale feature aggregation module fully integrates deep semantics, location information, and precise edge information;
[0024] A deep supervision mechanism is used to apply regularization constraints to the output of each encoding stage.
[0025] Furthermore, the loss function for each deeply supervised network branch uses a combination of Focal loss and lovász-softmax loss.
[0026] The beneficial effects of this invention are as follows: This invention extracts features from images from multiple axes and dimensions, improves the robustness of the network based on learning sufficient spatial features, and introduces a deep supervision mechanism to optimize the gradient propagation and convergence speed of the network.
[0027] (1) The present invention performs training on training samples by dividing them into blocks and uses random block selection to solve the problem of hardware computing power limitation.
[0028] (2) In the feature extraction stage, this invention combines the maximum value of multi-axis surfaces to highlight feature information, and combines three-dimensional asymmetric convolution to replace the traditional feature extraction method. Finally, it effectively integrates multi-dimensional information through an adaptive feature fusion mechanism, which effectively filters redundant feature information while enhancing the model's ability to extract edge information.
[0029] (3) In the decoding stage, the present invention adopts a cross-layer spatial attention mechanism to optimize the feature quality of the encoding layer and the decoding layer, and combines a multi-scale feature aggregation mechanism and a deep supervision mechanism to effectively capture edge details and long-distance dependent feature information, while solving the problem of network gradient vanishing / exploding. The present invention can effectively improve the accuracy of network segmentation.
[0030] (4) The upsampling path of the network in this invention restores the resolution through trilinear interpolation, thus avoiding the checkerboard effect of deconvolution.
[0031] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0032] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0033] Figure 1 This is a flowchart of the image segmentation method based on attention mechanism and multi-dimensional feature fusion in this embodiment;
[0034] Figure 2 This is a network model structure diagram of an image segmentation method based on attention mechanism and multi-dimensional feature fusion;
[0035] Figure 3 An adaptive multidimensional feature extraction model;
[0036] Figure 4 For feature attention mechanism model;
[0037] Figure 5 It is a multi-scale feature aggregation model. Detailed Implementation
[0038] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0039] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0040] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0041] Please see Figures 1-5 This embodiment provides an image segmentation method based on attention mechanism and multi-dimensional feature fusion, which specifically includes the following steps:
[0042] Step 1: Take blocks of 32*32*32 pixels from the training samples and use random sampling to input 8 samples into the network for training.
[0043] In this embodiment, the dataset consists of infant brain MR images with a resolution of 256*192*144. Due to the limitations of the server's computing power, the entire image cannot be directly fed into the model for training. Therefore, the image needs to be segmented and sampled. In the training mode, the image is taken as a size of 32*32*32 to match the current hardware computing power.
[0044] Step 2: Build a network model based on attention mechanism and multi-dimensional feature fusion. The overall model diagram is shown below. Figure 2 As shown.
[0045] like Figure 3 As shown, the encoder consists of three parts. The first part replaces traditional 3D convolution with an adaptive 2D and 3D feature extraction mechanism. Through maxima mapping on the three axes combined with an attention mechanism, it effectively learns tissue boundary information. Then, it performs matrix multiplication on the original 3D image and finally performs adaptive feature weighting fusion on the three outputs. The second part uses 3×3×3 convolution to extract features from the original image, serving as supplementary information for residual connections. The third part uses 1×3×3, 3×1×3, and 3×3×1 asymmetric convolutions to capture fine-grained and multidimensional feature information of brain tissue edges. Finally, an adaptive weighted fusion mechanism combines the feature information from the three parts to obtain the encoder output. After obtaining the output, a max pooling mechanism is used for downsampling.
[0046] The decoder consists of two parts. For example... Figure 4 As shown, the first part employs a cross-layer attention mechanism module. This module primarily uses the decoder output features and encoder output features to perform weighted constraints through an attention mechanism, thereby correcting the feature response. For example... Figure 5 As shown, the second part employs a multi-scale feature aggregation module. This module combines the semantic information from the last layer of the encoder, the information from the previous layer of the decoder, and the output information from the cross-layer attention mechanism module, summing the channels to obtain the resulting feature. Upsampling and deep supervision are then applied to this feature to accelerate network convergence and avoid gradient vanishing and exploding. Upsampling is primarily achieved through trilinear interpolation, and the loss function uses a combination of Focal loss and Lovász-softmax loss.
[0047] Step 3: Train the network model on the training set using a network model that combines multidimensional information and attention mechanisms, and tune the network parameters. When validating the model's performance on the validation set, due to the limited sample data, we adopted a 10x10 cross-validation method. Taking the iSeg-2019 Infant Brain Challenge dataset as an example, iSeg-2019 contains 10 datasets. We selected 9 of them as the training set and used the remaining dataset as the validation set. The validation dataset was not used in the network training. Finally, the network model saved the version with the best segmentation result from the validation dataset, and the true performance of the network was measured based on the results of multiple cross-validations.
[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An image segmentation method based on attention mechanism and multidimensional feature fusion, characterized in that: Includes the following steps: S1: Acquire and preprocess the image, and randomly select blocks of training samples according to a preset size to form a training set; S2: Construct a network model based on multi-dimensional feature information and attention mechanisms, including: Encoder: Adopts an adaptive fusion mechanism for 2D and 3D feature extraction, extracts features from multiple axes and dimensions, replaces traditional 3D convolution, and uses 2×2×2 max pooling for pooling layers; The encoder stage specifically includes the following steps: Two-dimensional feature maps with maximum activation values on multiple surfaces are obtained by projecting a 3D image through a maximum value mapping. Each of them learns its global attention information through a non-local self-attention mechanism; The image is restored to 3D using the Unsqueeze operation and then multiplied with the original image. MA1 is obtained through an adaptive weighted fusion strategy; Feature extraction was performed on the original image using asymmetric convolutions of 1×3×3, 3×1×3, and 3×3×1, respectively. MA2 is obtained through an adaptive weighted fusion strategy; MA3 is obtained by convolving the original image with a 3×3×3 matrix. The final encoder output feature map is obtained by combining MA1, MA2, and MA3 through an adaptive weighted fusion strategy. Decoder: Employs a strategy that combines attention mechanism with multi-scale features and deep supervision, uses trilinear interpolation for resolution restoration, and finally obtains the final segmentation result image through an activation function; S3: Parameterize the network, save the model that performs best on the validation data, and validate each data point. After multiple cross-validations, select the optimal model to perform image segmentation.
2. The image segmentation method based on attention mechanism and multi-dimensional feature fusion according to claim 1, characterized in that: The decoder stage specifically includes the following steps: A cross-layer spatial attention module is used to constrain the features of the encoding stage using features from the decoding stage. The multi-scale feature aggregation module fully integrates deep semantics, location information, and precise edge information; A deep supervision mechanism is used to apply regularization constraints to the output of each encoding stage.
3. The image segmentation method based on attention mechanism and multi-dimensional feature fusion according to claim 2, characterized in that: The loss function for each deeply supervised network branch uses a combination of Focal loss and lovász-softmax loss.