Image segmentation method and system based on attention mechanism
The proposed attention-based image segmentation method addresses feature loss and high computational demands in CNNs by enhancing feature extraction and reducing parameter counts, improving segmentation performance and efficiency.
Patent Information
- Application Number
- CN202510394033.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-15
AI Technical Summary
The existing convolutional neural network based on the encoder-decoder structure is prone to lose the features of the region of interest, multi-layer convolution and pooling operations lead to the loss of features, and the backbone feature extraction network parameters are large, which consumes a lot of computing resources and time.
Using an image segmentation method based on attention mechanism, feature extraction is performed through the first, second, third and fourth convolutional layers, feature enhancement, weighting and capture is performed in combination with the first, second and third attention models, and dimensionality reduction is used by the pooling layer and the fifth convolutional layer, and finally the deep feature image is output for segmentation.
It improves the neural network's understanding and processing capabilities of images, focuses on key information adaptively, reduces the number of parameters and the number of operations, enhances significant regional features, reduces computing resource consumption, and improves image segmentation performance and efficiency.
Smart Images

Figure CN120318250A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and particularly to an image segmentation method and system based on an attention mechanism. Background Art
[0002] In recent years, deep learning technology has become an important contributor to computer vision and has been widely applied to the research of image segmentation methods. As an outstanding representative of deep learning technology, the convolutional neural network (CNN) has received extensive attention for its powerful feature representation and global modeling capabilities. Compared with traditional image segmentation methods, the deep image segmentation method based on CNN can automatically extract semantic information in images and perform representation learning on data, effectively overcoming the limitations of traditional methods in manually extracting features. These encoder-decoder based deep CNN structures have achieved satisfactory results on many complex medical image segmentation datasets, verifying their effectiveness in learning the distinction of different regions of interest.
[0003] Currently, networks based on the encoder-decoder structure are generally prone to losing features of regions of interest. Although the stacking of multiple layers of convolution and continuous pooling operations can expand the receptive field and enhance the interaction between local features, they will cause the loss of some features, and incomplete features will also affect the segmentation performance of the network. At the same time, most of the existing backbone feature extraction networks have a relatively large number of parameters, resulting in the consumption of a large amount of computing resources and time. Therefore, it is necessary to design an image segmentation method and system based on an attention mechanism. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and to better and effectively solve the problem that currently, networks based on the encoder-decoder structure are generally prone to losing features of regions of interest. Although the stacking of multiple layers of convolution and continuous pooling operations can expand the receptive field and enhance the interaction between local features, they will cause the loss of some features, and incomplete features will also affect the segmentation performance of the network. At the same time, most of the existing backbone feature extraction networks have a relatively large number of parameters, resulting in the consumption of a large amount of computing resources and time. The present invention provides an image segmentation method and system based on an attention mechanism, which realizes the function of improving the understanding and processing ability of the neural network for the input image by adopting the attention mechanism and enabling the convolutional neural network to adaptively focus on key information in the image. Moreover, through the first attention model, the channel features can be combined with the input image, enhancing the features of significant regions in the image and weakening the features of insignificant regions, not only avoiding unnecessary dimension raising and lowering operations, but also significantly reducing the number of parameters and the number of operations of the model, improving the performance and efficiency of image segmentation.
[0005] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] An image segmentation method based on an attention mechanism, comprising the following steps:
[0007] Step A: Use the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer to perform feature extraction processing on the input image respectively, and obtain the first feature map, the second feature map, the third feature map, and the fourth feature map;
[0008] Step B: Use the first attention model to perform feature enhancement processing on the first feature map and the second feature map respectively, and obtain the first feature enhanced map and the second feature enhanced map;
[0009] Step C: Use the second attention model to perform feature weighting processing on the third feature map and obtain the third feature enhanced map;
[0010] Step D: Use the third attention model to perform feature capture processing on the fourth feature map and obtain the fourth feature enhanced map;
[0011] Step E: Use a pooling layer to perform pooling processing on the input image and obtain a pooled feature map;
[0012] Step F: Stack the first feature enhanced map, the second feature enhanced map, the third feature enhanced map, the fourth feature enhanced map, and the pooled feature map, and use the fifth convolutional layer to perform dimensionality reduction processing to obtain a deep feature image, and then output the deep feature image to a decoder to complete the image segmentation operation.
[0013] For the aforementioned image segmentation method based on an attention mechanism, in Step A, the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer are used to perform feature extraction processing on the input image respectively, and obtain the first feature map, the second feature map, the third feature map, and the fourth feature map, wherein the first convolutional layer uses a 1×1 convolution, the second convolutional layer uses a 3×3 convolution with a dilation rate of 6, the third convolutional layer uses a 3×3 convolution with a dilation rate of 12, and the fourth convolutional layer uses a 3×3 convolution with a dilation rate of 18.
[0014] For the aforementioned image segmentation method based on an attention mechanism, in Step B, the first attention model is used to perform feature enhancement processing on the first feature map and the second feature map respectively, and obtain the first feature enhanced map and the second feature enhanced map, wherein the specific feature enhancement processing steps of the first attention model are as follows:
[0015] Step B1: Perform a first global average pooling on the sample feature map U∈R H×W×C of each channel and output a feature vector of 1×1×C, where the sample feature map is the first feature map or the second feature map, as shown in formula (1).
[0016]
[0017] Among them, z c is the feature vector of the c-th channel of the sample feature map, H and W are the length and width of the sample feature map, C is the channel of the sample feature map, and F sq is the first global average pooling operation, and u c is the eigenvalue at the c-th channel of the sample feature map, c is the channel index, and i and j are the position indices in the height and width directions respectively;
[0018] Step B2, learn the dependencies between channels and generate corresponding weights for each channel. Specifically, use the first fully connected layer, the second fully connected layer, and the non-linear activation function ReLU for learning, and then use the Sigmoid activation function to generate a weight between 0 and 1 for each channel. The first fully connected layer is used to reduce the dimension of the feature vector, and the second fully connected layer is used to restore the dimension to the original number of channels. The specific channel statistic s is shown in formula (2),
[0019] s = F ex (z, W0) = σ(g(z, W0)) = σ(W2δ(W1z)) (2)
[0020] Among them, F ex is the excitation operation, z is the feature vector, W0 is the weight of the fully connected neural network, σ(·) is the Sigmoid function, g(·) is the mapping, δ(·) is the ReLU function, and W2 ∈ R c×(c / r) is the number of channels increased relative to the ratio r, and W1 ∈ R c / r×c is the number of channels decreased relative to the ratio r;
[0021] Step B3, multiply the eigenvalue u c at the c-th channel of the sample feature map by the c-th channel statistic s c to obtain the sample feature enhanced map The sample feature enhanced map is the first feature enhanced map or the second feature enhanced map, as shown in formula (3),
[0022]
[0023] The aforementioned image segmentation method based on the attention mechanism, step C, uses the second attention model to perform feature weighting processing on the third feature map and obtain the third feature enhanced map. The specific steps are as follows,
[0024] Step C1, use the second global average pooling to aggregate the spatial information of each channel of the third feature map and obtain the pooling vector. The second global average pooling is shown in formula (4),
[0025]
[0026] Among them, P(·) is the second global average pooling operation, and U c is the third feature map;
[0027] Step C2, perform one-dimensional convolution on the pooled vector and obtain the one-dimensional convolution output result, where the convolution kernel size k changes dynamically according to the number of channels C, as shown in formula (5),
[0028]
[0029] where ψ(·) is the dynamic change operation, ||.|| odd is the closest odd number, and both γ and b are hyperparameters;
[0030] Step C3, use the Sigmoid function to calculate the activation value of the one-dimensional convolution output result and obtain the feature channel weight value ω c , as shown in formula (6),
[0031] ω c =σ{Conv k [P(U c )]}(6)
[0032] where Conv k is the convolution operation;
[0033] Step C4, multiply the feature channel weight value ω c with the third feature map U c one by one and obtain the third feature enhanced map
[0034] The aforementioned image segmentation method based on the attention mechanism, step D, uses the third attention model to perform feature capture processing on the fourth feature map and obtain the fourth feature enhanced map, where the third attention model specifically adopts the combination of the channel attention mechanism and the spatial attention mechanism to comprehensively capture the feature information in the fourth feature map. The specific steps are as follows.
[0035] Step D1, adopt the channel attention model mechanism to perform feature capture processing on the fourth feature map and output the channel attention feature map M c (F), as shown in formula (7),
[0036] M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)))(7)
[0037] Among them, F is the fourth feature map, AvgPool and MaxPool respectively represent the global average pooling operation and the max pooling operation, and MLP is a multi-layer perceptron;
[0038] Step D2, use the spatial attention mechanism to perform feature capture processing on the attention feature map M c (F) and output the fourth feature enhanced map M s (F), as shown in formula (8):
[0039] M s (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])) (8)
[0040] Among them, f 7×7 is a 7×7 convolution kernel.
[0041] For the aforementioned image segmentation method based on the attention mechanism, in step F, stack the first feature enhanced map, the second feature enhanced map, the third feature enhanced map, the fourth feature enhanced map and the pooled feature map, and use the fifth convolutional layer for dimensionality reduction processing to obtain a deep feature image, and then output the deep feature image to the decoder to complete the image segmentation operation, where the fifth convolutional layer uses a 1×1 convolution.
[0042] An image segmentation system based on the attention mechanism includes a feature extraction module, a feature enhancement module, a feature weighting module, a feature capture module, a pooling module and a stacking and dimensionality reduction module. The feature extraction module is used to perform feature extraction processing on the input image by using the first convolutional layer, the second convolutional layer, the third convolutional layer and the fourth convolutional layer respectively to obtain the first feature map, the second feature map, the third feature map and the fourth feature map; the feature enhancement module is used to perform feature enhancement processing on the first feature map and the second feature map by using the first attention model respectively to obtain the first feature enhanced map and the second feature enhanced map; the feature weighting module is used to perform feature weighting processing on the third feature map by using the second attention model to obtain the third feature enhanced map; the feature capture module is used to perform feature capture processing on the fourth feature map by using the third attention model to obtain the fourth feature enhanced map; the pooling module is used to perform pooling processing on the input image by using a pooling layer to obtain a pooled feature map; the stacking and dimensionality reduction module is used to stack the first feature enhanced map, the second feature enhanced map, the third feature enhanced map, the fourth feature enhanced map and the pooled feature map, and use the fifth convolutional layer for dimensionality reduction processing to obtain a deep feature image, and then output the deep feature image to the decoder to complete the image segmentation operation.
[0043] The beneficial effects of the present invention are as follows: For an image segmentation method and system based on an attention mechanism of the present invention, first, a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer are used to perform feature extraction processing on an input image respectively to obtain a first feature map, a second feature map, a third feature map, and a fourth feature map. Then, a first attention model is used to perform feature enhancement processing on the first feature map and the second feature map respectively to obtain a first feature enhancement map and a second feature enhancement map. Next, a second attention model is used to perform feature weighting processing on the third feature map to obtain a third feature enhancement map. Then, a third attention model is used to perform feature capture processing on the fourth feature map to obtain a fourth feature enhancement map. Subsequently, a pooling layer is used to perform pooling processing on the input image to obtain a pooled feature map. Then, the first feature enhancement map, the second feature enhancement map, the third feature enhancement map, the fourth feature enhancement map, and the pooled feature map are stacked and a fifth convolutional layer is used for dimensionality reduction processing to obtain a deep feature image. Then, the deep feature image is output to a decoder to complete the image segmentation operation, effectively realizing the functions that the image segmentation method and system can improve the understanding and processing ability of a neural network for an input image by using an attention mechanism and enable a convolutional neural network to adaptively focus on key information in the image. Moreover, through the first attention model, channel features can be combined with the input image, enabling the enhancement of significant region features and the weakening of non-significant region features in the image. At the same time, through the second attention model, information between adjacent channels can be exchanged and channel attention calculation can be performed under different channel dimensions, not only avoiding unnecessary dimensionality increase and decrease operations, but also significantly reducing the number of parameters and the number of operations of the model. And through the third attention model, channel attention and spatial attention can be combined to provide a more comprehensive and effective feature extraction ability, which can not only comprehensively capture key information in features, but also improve the feature extraction effect, reduce the number of parameters, and improve the performance and efficiency of image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is the overall flowchart of an image segmentation method based on an attention mechanism of the present invention;
[0045] Figure 2 is the schematic diagram of the operating principle of an image segmentation system based on an attention mechanism of the present invention;
[0046] Figure 3 is the schematic diagram of the operating principle of the first attention model of the present invention;
[0047] Figure 4 is the schematic diagram of the operating principle of the second attention model of the present invention;
[0048] Figure 5 is the schematic diagram of the operating principle of the third attention model of the present invention;
[0049] Figure 6 It is a schematic diagram of the test result of image segmentation using the prior art in the embodiment of the present invention;
[0050] Figure 7 It is a schematic diagram of the test result of image segmentation using the image segmentation method of the present invention in the embodiment of the present invention. Specific embodiments
[0051] Next, the present invention will be further described in conjunction with the accompanying drawings of the specification.
[0052] As Figure 1 and Figure 2 shown, an image segmentation method based on an attention mechanism of the present invention includes the following steps.
[0053] Step A: Use the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer to perform feature extraction processing on the input image respectively and obtain the first feature map, the second feature map, the third feature map, and the fourth feature map. Among them, the first convolutional layer uses 1×1 convolution, the second convolutional layer uses 3×3 convolution with a dilation rate of 6, the third convolutional layer uses 3×3 convolution with a dilation rate of 12, and the fourth convolutional layer uses 3×3 convolution with a dilation rate of 18.
[0054] As Figure 3 shown, in step B: Use the first attention model to perform feature enhancement processing on the first feature map and the second feature map respectively and obtain the first feature enhancement map and the second feature enhancement map. The specific feature enhancement processing steps of the first attention model are as follows.
[0055] Step B1: Perform the first global average pooling on the sample feature map U∈R H×W×C of each channel and output a feature vector of 1×1×C. The sample feature map is the first feature map or the second feature map, as shown in formula (1).
[0056]
[0057] where z c is the feature vector of the c-th channel of the sample feature map, H and W are the length and width of the sample feature map, C is the channel of the sample feature map, F sq is the first global average pooling operation, u c is the feature value at the c-th channel of the sample feature map, c is the channel index, and i and j are the position indices in the height and width directions respectively;
[0058] Step B2, learn the dependencies between channels and generate corresponding weights for each channel. Specifically, use the first fully connected layer, the second fully connected layer, and the non-linear activation function ReLU for learning, and then use the Sigmoid activation function to generate a weight between 0 and 1 for each channel. The first fully connected layer is used to reduce the dimension of the feature vector, and the second fully connected layer is used to restore the dimension to the original number of channels. The specific channel statistic s is shown in formula (2).
[0059] s = F ex (z, W0) = σ(g(z, W0)) = σ(W2δ(W1z)) (2)
[0060] where F ex is the excitaion operation, z is the feature vector, W0 is the weight of the fully connected neural network, σ(·) is the Sigmoid function, g(·) is the mapping, δ(·) is the ReLU function, W2 ∈ R c×(c / r) is the number of channels increased relative to the ratio r, and W1 ∈ R c / r×c is the number of channels decreased relative to the ratio r;
[0061] Step B3, multiply the eigenvalue u c at the c-th channel of the sample feature map c by the c-th channel statistic s to obtain the sample feature enhanced map The sample feature enhanced map
[0062]
[0063] is the first feature enhanced map or the second feature enhanced map, as shown in formula (3). Figure 4 As shown, step C, use the second attention model to perform feature weighting on the third feature map and obtain the third feature enhanced map. The specific steps are as follows
[0064] Among them, through the second attention model, the information between adjacent channels can be exchanged and channel attention calculation can be performed under different channel dimensions, which not only avoids unnecessary dimension raising and lowering operations, but also significantly reduces the number of parameters and the number of operations of the model;
[0065] Step C1, use the second global average pooling to aggregate the spatial information of each channel of the third feature map and obtain the pooling vector. The second global average pooling is shown in formula (4).
[0066]
[0067] where P(·) is the second global average pooling operation, and U c is the third feature map;
[0068] Step C2: Perform one-dimensional convolution on the pooled vector and obtain the one-dimensional convolution output result, where the convolution kernel size k changes dynamically according to the number of channels C, as shown in formula (5).
[0069]
[0070] where ψ(·) is a dynamically changing operation, ||.|| odd is the closest odd number, and both γ and b are hyperparameters;
[0071] Step C3: Use the Sigmoid function to calculate the activation value of the one-dimensional convolution output result and obtain the feature channel weight value ω c , as shown in formula (6).
[0072] ω c = σ{Conv k [P(U c )]}(6)
[0073] where Conv k is the convolution operation;
[0074] Step C4: Multiply the feature channel weight value ω c with the third feature map U c element by element and obtain the third feature enhanced map
[0075] As Figure 5 shown, in step D, use the third attention model to perform feature capture processing on the fourth feature map and obtain the fourth feature enhanced map, where the third attention model specifically uses a combination of channel attention mechanism and spatial attention mechanism to comprehensively capture the feature information in the fourth feature map. The specific steps are as follows
[0076] where the third attention model can combine channel attention and spatial attention and provide a more comprehensive and effective feature extraction ability. It can not only comprehensively capture the key information in the features but also improve the feature extraction effect;
[0077] Step D1: Use the channel attention model mechanism to perform feature capture processing on the fourth feature map and output the channel attention feature map M c (F), as shown in formula (7).
[0078] M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))(7)
[0079] where F is the fourth feature map, AvgPool and MaxPool respectively represent the global average pooling operation and the maximum pooling operation, and MLP is the multi-layer perceptron;
[0080] Step D2, using a spatial attention mechanism to perform feature capture processing on the attention feature map M c (F) and output the fourth feature enhanced map M s (F), as shown in formula (8),
[0081] M s (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])) (8)
[0082] where f 7×7 is a 7×7 convolution kernel.
[0083] Step E, using a pooling layer to perform pooling processing on the input image and obtaining the pooled feature map;
[0084] Step F, stacking the first feature enhanced map, the second feature enhanced map, the third feature enhanced map, the fourth feature enhanced map and the pooled feature map and using a fifth convolutional layer for dimensionality reduction processing to obtain a deep feature image, and then outputting the deep feature image to the decoder to complete the image segmentation operation, where the fifth convolutional layer uses a 1×1 convolution.
[0085] An image segmentation system based on an attention mechanism, including a feature extraction module, a feature enhancement module, a feature weighting module, a feature capture module, a pooling module and a stacking and dimensionality reduction module. The feature extraction module is used to perform feature extraction processing on the input image by using a first convolutional layer, a second convolutional layer, a third convolutional layer and a fourth convolutional layer respectively and obtain a first feature map, a second feature map, a third feature map and a fourth feature map; the feature enhancement module is used to perform feature enhancement processing on the first feature map and the second feature map by using a first attention model respectively and obtain a first feature enhanced map and a second feature enhanced map; the feature weighting module is used to perform feature weighting processing on the third feature map by using a second attention model and obtain a third feature enhanced map; the feature capture module is used to perform feature capture processing on the fourth feature map by using a third attention model and obtain a fourth feature enhanced map; the pooling module is used to perform pooling processing on the input image by using a pooling layer and obtain the pooled feature map; the stacking and dimensionality reduction module is used to stack the first feature enhanced map, the second feature enhanced map, the third feature enhanced map, the fourth feature enhanced map and the pooled feature map and use a fifth convolutional layer for dimensionality reduction processing to obtain a deep feature image, and then output the deep feature image to the decoder to complete the image segmentation operation.
[0086] To better illustrate the usage effect of the present invention, a specific embodiment of using the method of the present invention and its implementation effect are introduced below.
[0087] This embodiment is tested using the target detection dataset Pascal VOC 2007, and the test results are as Figure 6 and Figure 7 shown, where Figure 6 is the image segmentation test result using the prior art, Figure 7 is the image segmentation test result using the image segmentation method of the present invention; the mean intersection over union (MIoU) is used to measure the similarity between the prediction result and the ground truth label in the semantic segmentation task; the value range of MIoU is between 0 and 1, and the closer it is to 1, the higher the similarity between the prediction result and the ground truth label. On the contrary, the closer it is to 0, the lower the similarity; as Figure 7 shown, the MIoU of the method of the present invention is 0.7235, while the MIoU of the prior art is 0.7117.
[0088] In summary, for an image segmentation method and system based on an attention mechanism of the present invention, first, the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer are respectively used to perform feature extraction processing on the input image to obtain the first feature map, the second feature map, the third feature map, and the fourth feature map. Then, the first attention model is used to perform feature enhancement processing on the first feature map and the second feature map respectively to obtain the first feature enhanced map and the second feature enhanced map. Next, the second attention model is used to perform feature weighting processing on the third feature map to obtain the third feature enhanced map. Then, the third attention model is used to perform feature capture processing on the fourth feature map to obtain the fourth feature enhanced map. Subsequently, a pooling layer is used to perform pooling processing on the input image to obtain the pooled feature map. Then, the first feature enhanced map, the second feature enhanced map, the third feature enhanced map, the fourth feature enhanced map, and the pooled feature map are stacked and a fifth convolutional layer is used for dimensionality reduction processing to obtain a deep feature image. Then, the deep feature image is output to the decoder to complete the image segmentation operation, effectively realizing the function that the image segmentation method and system have the ability to improve the understanding and processing ability of the neural network for the input image by using the attention mechanism and enabling the convolutional neural network to adaptively focus on the key information in the image. Moreover, through the first attention model, the channel features can be combined with the input image, enabling the enhancement of the significant region features in the image and the weakening of the insignificant region features. At the same time, through the second attention model, the information between adjacent channels can be exchanged and channel attention calculation can be performed in different channel dimensions, not only avoiding unnecessary dimensionality increase and decrease operations, but also significantly reducing the number of parameters and the number of operations of the model. And through the third attention model, channel attention and spatial attention can be combined to provide a more comprehensive and effective feature extraction ability, not only being able to comprehensively capture the key information in the features, but also improving the feature extraction effect, reducing the number of parameters, and improving the performance and efficiency of image segmentation.
[0089] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. An image segmentation method based on an attention mechanism, characterized in that: including the following steps, Step A: Use the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer to perform feature extraction processing on the input image respectively, and obtain the first feature map, the second feature map, the third feature map, and the fourth feature map; Step B: Use the first attention model to perform feature enhancement processing on the first feature map and the second feature map respectively, and obtain the first feature enhancement map and the second feature enhancement map; Step C: Use the second attention model to perform feature weighting processing on the third feature map and obtain the third feature enhancement map; Step D: Use the third attention model to perform feature capture processing on the fourth feature map and obtain the fourth feature enhancement map; Step E: Use the pooling layer to perform pooling processing on the input image and obtain the pooled feature map; Step F: Stack the first feature enhancement map, the second feature enhancement map, the third feature enhancement map, the fourth feature enhancement map, and the pooled feature map, and use the fifth convolutional layer to perform dimensionality reduction processing to obtain the deep feature image, and then output the deep feature image to the decoder to complete the image segmentation task.
2. The image segmentation method based on the attention mechanism according to claim 1, wherein: Step A: Use the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer to perform feature extraction processing on the input image respectively, and obtain the first feature map, the second feature map, the third feature map, and the fourth feature map, where the first convolutional layer uses 1×1 convolution, the second convolutional layer uses 3×3 convolution with a dilation rate of 6, the third convolutional layer uses 3×3 convolution with a dilation rate of 12, and the fourth convolutional layer uses 3×3 convolution with a dilation rate of 18.
3. The image segmentation method based on the attention mechanism according to claim 2, wherein: Step B: Use the first attention model to perform feature enhancement processing on the first feature map and the second feature map respectively, and obtain the first feature enhancement map and the second feature enhancement map. The specific feature enhancement processing steps of the first attention model are as follows. Step B1: For the sample feature map U ∈ R of each channel H×W×C perform the first global average pooling and output a feature vector of 1×1×C. The sample feature map is the first feature map or the second feature map, as shown in formula (1). Among them, z c is the feature vector of the c-th channel of the sample feature map, H and W are the length and width of the sample feature map, C is the channel of the sample feature map, F sq is the first global average pooling operation, u c is the eigenvalue at the c-th channel of the sample feature map, c is the channel index, and i and j are the position indexes in the height and width directions respectively; Step B2: Learn the dependence between channels and generate corresponding weights for each channel. Specifically, use the first fully connected layer, the second fully connected layer, and the nonlinear activation function ReLU for learning, and then use the Sigmoid activation function to generate a weight between 0 and 1 for each channel. The first fully connected layer is used to reduce the dimension of the feature vector, and the second fully connected layer is used to restore the dimension to the original number of channels. The specific channel statistic s is shown in formula (2). s = F ex (z, W0) = σ(g(z, W0)) = σ(W2δ(W1z)) (2) Among them, F ex is the excitaion operation, z is the feature vector, W0 is the fully connected neural network weight, σ(·) is the Sigmoid function, g(·) is the mapping, δ(·) is the ReLU function, W2 ∈ R c×(c / r) is the number of channels increased relative to the ratio r, W1 ∈ R c / r×c is the number of channels decreased relative to the ratio r; Step B3: Multiply the eigenvalue u at the c-th channel of the sample feature map c by the c-th channel statistic s c to obtain the sample feature enhanced map The sample feature enhanced map is the first feature enhanced map or the second feature enhanced map, as shown in formula (3).
4. The image segmentation method based on the attention mechanism according to claim 3, wherein: Step C: Use the second attention model to perform feature weighting processing on the third feature map and obtain the third feature enhancement map. The specific steps are as follows. Step C1: Use the second global average pooling to aggregate the spatial information of each channel of the third feature map and obtain the pooled vector. The second global average pooling is shown in formula (4). where P(·) is the second global average pooling operation, and U c is the third feature map; Step C2: Perform one-dimensional convolution on the pooled vector and obtain the one-dimensional convolution output result. The size k of the convolution kernel changes dynamically according to the number of channels C, as shown in formula (5). where, ψ(·) is a dynamically changing operation, ||.|| odd is the closest odd number, and both γ and b are hyperparameters; Step C3, calculate the activation value of the one-dimensional convolution output result using the Sigmoid function and obtain the feature channel weight value ω c , as shown in formula (6). ω c = σ{Conv k [P(U c )]}(6) Among them, Conv k is a convolution operation; Step C4, multiply the feature channel weight value ω c with the third feature map U c one by one and obtain the third feature enhanced map 5. A method for image segmentation based on an attention mechanism according to claim 4, characterized in that: Step D: Use the third attention model to perform feature capture processing on the fourth feature map and obtain the fourth feature enhancement map. The third attention model specifically uses a combination of channel attention mechanism and spatial attention mechanism to comprehensively capture the feature information in the fourth feature map. The specific steps are as follows. Step D1, perform feature capture processing on the fourth feature map using a channel attention module and output a channel attention feature map M c (F), as shown in formula (7), M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))(7) Among them, F is the fourth feature map, AvgPool and MaxPool respectively represent global average pooling operation and max pooling operation, and MLP is a multi-layer perceptron; Step D2, using a spatial attention mechanism to perform feature capture processing on the attention feature map M c (F) and output the fourth feature enhancement map M s (F), as shown in formula (8). M s (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)]))(8) Among them, f 7×7 is a 7×7 convolution kernel.
6. The image segmentation method based on the attention mechanism according to claim 5, characterized in that: Step F: Stack the first feature enhanced map, the second feature enhanced map, the third feature enhanced map, the fourth feature enhanced map and the pooled feature map, and perform dimensionality reduction processing using the fifth convolutional layer to obtain a deep feature image, and then output the deep feature image to the decoder to complete the image segmentation task, where the fifth convolutional layer uses 1×1 convolution.
7. An image segmentation system based on an attention mechanism, wherein the specific operation process of the image segmentation system is based on the image segmentation method according to any one of claims 1-6, and is characterized in that: It includes a feature extraction module, a feature enhancement module, a feature weighting module, a feature capture module, a pooling module and a stacking dimensionality reduction module. The feature extraction module is used to perform feature extraction processing on the input image using the first convolutional layer, the second convolutional layer, the third convolutional layer and the fourth convolutional layer respectively, and obtain the first feature map, the second feature map, the third feature map and the fourth feature map; The feature enhancement module is used to perform feature enhancement processing on the first feature map and the second feature map respectively using the first attention model, and obtain the first feature enhanced map and the second feature enhanced map; The feature weighting module is used to perform feature weighting processing on the third feature map using the second attention model, and obtain the third feature enhanced map; The feature capture module is used to perform feature capture processing on the fourth feature map using the third attention model, and obtain the fourth feature enhanced map; The pooling module is used to perform pooling processing on the input image using a pooling layer, and obtain a pooled feature map; The stacking dimensionality reduction module is used to stack the first feature enhanced map, the second feature enhanced map, the third feature enhanced map, the fourth feature enhanced map and the pooled feature map, and perform dimensionality reduction processing using the fifth convolutional layer to obtain a deep feature image, and then output the deep feature image to the decoder to complete the image segmentation task.
Citation Information
Patent Citations
Construction method and application of convolutional neural network based on attention mechanisms
CN108734290A
Mine image super-resolution reconstruction system and method based on overall attention
CN117173024A
Super-resolution image reconstruction method and device
CN118134763A