A refined image semantic segmentation method based on progressive neighbor aggregation

The progressive nearest neighbor aggregation module (PNA) calculates the nearest neighbor pixel correlation and channel attention in the semantic segmentation model, and gradually refines the segmentation results, solving the problem of rough prediction results of the existing model and achieving a higher precision segmentation effect.

CN116246065BActive Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211621971.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-16
Publication Date
2025-07-22
Estimated Expiration
2042-12-16

AI Technical Summary

Technical Problem

The prediction results of the existing semantic segmentation model are relatively rough and cannot effectively retain the detailed information of the target. The initial segmentation results and refinement processes are usually carried out separately, and joint training is lacking.

Method used

The progressive nearest neighbor aggregation module (PNA), including the nearest neighbor aggregation module (NAM) and the self-aggregation module (SAM), is used to gradually refine the segmentation results by calculating the nearest neighbor pixel correlation and channel attention of different resolution features, and integrate the prediction and refinement process into an end-to-end network.

Benefits of technology

Improved segmentation performance, better preserve target details, and improve segmentation accuracy through end-to-end training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246065B_ABST
    Figure CN116246065B_ABST
Patent Text Reader

Abstract

The present invention discloses a refined image semantic segmentation method based on progressive neighborhood aggregation (PNA), which realizes refined image semantic segmentation in a step-by-step refinement manner. PNA includes two main modules, namely, a neighborhood aggregation module (NAM) and a self-aggregation module (SAM). Among them, NAM aggregates semantic features by utilizing the correlation of neighboring pixels in features of different resolutions; SAM enhances important semantic information and supplements missing detailed information through a channel attention mechanism. The present invention gradually refines the segmentation result by using features of different resolutions, improves the segmentation performance, and promotes the development of related applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a refined image semantic segmentation method. Background Art

[0002] The goal of the semantic segmentation task is to predict class labels for each pixel in the input image and divide the image into different semantic regions. Semantic segmentation technology can be widely applied to application scenarios such as autonomous driving, medical image segmentation, and image editing. In the autonomous driving scenario, semantic segmentation technology can accurately analyze target areas such as drivable areas, pedestrians, and vehicles in real time, realize the perception of environmental information around the parking space, and ensure the safety of autonomous driving. In the field of medical image segmentation, applying semantic segmentation technology for lesion segmentation and cancer cell segmentation can greatly reduce the burden on doctors and improve the medical treatment efficiency of patients. With the popularity of online social media such as Weibo and Douyin, the demand for efficient and convenient image editing is also increasing day by day. By using semantic segmentation technology to identify and extract image content, the diverse image editing needs of users can be met.

[0003] Since the emergence of Fully Convolutional Networks (FCN), significant progress has been made in semantic segmentation tasks. Based on FCN, a large number of research works have emerged, which can be mainly divided into two research directions. One is to address the problem that the network prediction results are relatively rough in the boundary regions; the other is to design context information capture modules to solve the multi-scale problem of segmentation targets. For the problem of relatively rough segmentation results, typical solutions include improving the feature resolution through methods such as Dilated Convolution, Deconvolution and Unpooling, and Skip Connection to enhance the segmentation accuracy in the detail regions. Such methods have been applied to many semantic segmentation models, such as SegNet, U-Net, RefineNet, DeepLabV3+, and so on. For the multi-scale problem of the target to be segmented, capturing multi-scale context information is a common solution, such as PSPNet and ASPP. There are also some methods that use attention mechanisms to encode long-range context information. In recent years, Transformers have shown great potential in the field of semantic segmentation. Some works use Transformers as the backbone network for dense pixel-level prediction, such as SETR, Segmenter, and SegFormer. In addition, different from the existing methods that mostly use pixel-level classification for semantic segmentation, there are also some studies that regard semantic segmentation as a prediction problem of a Mask set. The representative works of these studies include MaskFormer and Mask2Former.

[0004] However, the segmentation results predicted by existing semantic segmentation models are still relatively rough and cannot well retain the detailed information of the target. To solve this problem, existing methods usually use features with different resolutions in the backbone network to refine the prediction results because features with relatively higher resolutions contain richer spatial detailed information. By a specific feature fusion method, multiple features with different resolutions at different levels in the network can be fused to supplement the missing detailed information. A common method is to fuse multi-resolution features from the backbone network through skip connections to generate high-quality segmentation prediction results. For example, in FCN and Unet, high-resolution features from the shallow layer are introduced through skip connections, and features with different resolutions are fused using summation operations or concatenation operations. Given the success of Feature Pyramid Networks (FPN) in the field of object detection, the feature fusion strategy based on FPN has been widely applied in the field of semantic segmentation. For example, UperNet designed a semantic segmentation network composed of a pyramid pooling module and a feature pyramid. Considering the problem of semantic misalignment between features with different resolutions, FaPN proposed a feature alignment module to align the upsampled high-level features.

[0005] In addition, some segmentation networks choose to predict the segmentation results step by step from rough to fine or use boundary information to refine the segmentation boundary. For example, CascadedPSP and MagNet achieve the gradual refinement of the segmentation results by repeatedly inputting the prediction results into the refinement modules they designed. However, in these methods, the prediction of the initial segmentation result and the refinement of the segmentation result are carried out separately using two separate models, that is, the two tasks of semantic segmentation and segmentation result refinement cannot be jointly trained. Subsequently, PointRend designed an algorithm to adaptively calculate the pixels with uncertain segmentation results and further refine these confusing pixels. However, this method highly depends on the initial prediction results, and the pixel features used for further refinement also lack sufficient context information, so it is easily misled by incorrect initial segmentation results. Summary of the Invention

[0006] To overcome the deficiencies of the prior art, the present invention provides a refined image semantic segmentation method based on progressive neighborhood aggregation (Progressive Neighbodrhood Aggregation, PNA), which realizes refined image semantic segmentation through a step-by-step refinement method. PNA includes two main modules, namely the Neighborhood Aggregation Module (NAM) and the Self Aggregation Module (SAM). Among them, NAM utilizes the correlation of neighboring pixels in features of different resolutions to aggregate semantic features; SAM enhances important semantic information and supplements lost detailed information through a channel attention mechanism. The present invention gradually refines the segmentation result using features of different resolutions, improves the segmentation performance, and promotes the development of related applications.

[0007] The technical solutions adopted by the present invention to solve its technical problems include the following steps:

[0008] Step 1: Input the input image I ∈ R 3×H×W into the backbone network, and four features of different resolutions are extracted through the backbone network where m ∈ {1, 2, 3, 4};

[0009] The feature F1 passes through a 1x1 Conv, and the number of channels becomes C, corresponding to C semantic categories, that is, the initial segmentation result is obtained

[0010] Step 2: After obtaining F m and P0, the progressive neighborhood aggregation module PNA uses the feature F1 to gradually refine P0. The progressive neighborhood aggregation module PNA includes two modules, namely the neighborhood aggregation module NAM and the self aggregation module SAM;

[0011] In the neighborhood aggregation module NAM, a window of size L*L is set, and the similarity between each position and neighboring pixels within the window is calculated for aggregating semantic features, specifically as follows:

[0012] Step 2-1: First, change the dimensions of F1 and P0 from (C, H1, W1) to (H1, W1, C), and then perform regularization respectively. Determine whether both H1 and W1 are greater than or equal to L. If not, fill 0 on the right or below;

[0013] Step 2-2: Adopt multi-head attention, and set the number of attention heads to h n , and the number of channels corresponding to each attention head is c n = C / h n ; Q and K are obtained by changing the dimensions after F1 passes through the linear mapping shown in Equation (1), and the dimension is (hn , H1, W1, c n ); V is obtained by changing the dimension after the linear mapping shown in Equation (1) from P0, and the dimension is (h n , H1, W1, c n );

[0014] Q = W q F1, K = W k F1, V = W v P0 (1)

[0015] Among them, Q, K, and V respectively correspond to the Query matrix, Key matrix, and Value matrix in the Transformer structure, and F1 and P0 are transformation matrices;

[0016] Step 2 - 3: Calculate the similarity between each position and its neighboring pixels;

[0017] The feature Q at (i, j) in Q i,j has its dimension changed from (h n , 1, 1, c n ) to (h n , 1, c n ); The key matrix of the neighboring pixels of (i, j) within the window is with the dimension of (h n , L×L, c n ); Calculate the similarity S between (i, j) and its neighboring pixels according to Equation (2) i,j , where B is the relative position encoding, coming from the bias matrix which has been introduced previously.

[0018]

[0019] Among them, B is the relative position encoding, defined by the parameterized bias matrix ;

[0020] The obtained S i,j has the dimension size of (h n , 1, L×L), and then calculate the aggregated feature at (i, j) in P0 according to Equation (3)

[0021]

[0022] After calculating the features of all positions, the aggregated feature can be obtained

[0023] Step 3: In the self - aggregation module SAM, adopt the Transformer structure and select multi - head self - attention, where the number of attention heads is set to h s, the dimension c corresponding to each attention head s = C / h s ;

[0024] Step 3-1: First, regularize the feature F1, and then obtain through 1x1Conv and 3x3DWConv which respectively correspond to the Query matrix, Key matrix, and Value matrix in the Transformer structure; as shown in Equation (4):

[0025]

[0026] The feature dimension of is h s , c s , H1×W1;

[0027] Step 3-2: After obtaining , multiply it with the transposed to get the channel attention attn, with the feature dimension of (h s , c s , c s ); after attn is processed by softmax, multiply it with to get the result of self-attention The feature dimension size is (h s , c s , H1×W1), as shown in the following formula:

[0028]

[0029] where α represents a learnable scaling parameter;

[0030] Then change the dimension size of from (h s , c s , H1×W1) to (C, H1, W1), and after passing through a layer of 1x1Conv, obtain the output of the SAM module

[0031] Step 4: After obtaining and , add the two together, and denote the resulting value as After regularization, input it into the feed-forward network FFN, and after adding it to , input it into 1x1Conv for segmentation prediction to obtain the refined segmentation result P1; the entire process is as follows:

[0032]

[0033]

[0034] Step 5: Upsample P1 to the same size as F2; repeat the methods in Steps 2 to 4 to obtain a refined segmentation result P2; and so on, repeat the above process to obtain the final refined result P4.

[0035] Preferably, the backbone network is ResNet101.

[0036] Preferably, L = 7.

[0037] Preferably, the structure and order of the FFN are a linear layer, a GELU activation layer, a Dropout layer, a linear layer, and a Dropout layer.

[0038] The beneficial effects of the present invention are as follows:

[0039] The present invention refines the prediction result of the segmentation network through the proposed progressive neighbor aggregation module, improves the segmentation performance, and promotes the development of related applications. In the present invention, NAM calculates the correlation between pixels and neighboring pixels in features of different resolutions and gradually aggregates these correlations into high-level semantic features. Without destroying the spatial information, SAM calculates the channel-wise self-attention for shallow features to enhance important semantic information, thereby supplementing the detailed information lost in high-level semantic features. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of the method idea of the present invention.

[0041] Figure 2 It is a structural diagram of the method of the present invention, (a) overall structure, (b) NAM module, (c) SAM module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The present invention will be further described below with reference to the drawings and embodiments.

[0043] The present invention aims to solve the problems existing in the related art to a certain extent, and proposes a refined image semantic segmentation method based on progressive neighbor aggregation. The present invention utilizes the near-neighbor similarity characteristics of neighboring pixels in space, and mines and utilizes the structural information of multi-scale features in the backbone network by calculating the near-neighbor similarity matrix, which is different from the existing methods that simply fuse these multi-scale features by adding or concatenating. The core idea is as Figure 1 shown. In addition, the present invention proposes an end-to-end segmentation scheme, which can integrate the process of predicting the initial segmentation result and the process of refining the segmentation result into the same network.

[0044] The Progressive Neighborhood Aggregation (PNA) module designed in the present invention can be applied to existing semantic segmentation networks to form an end-to-end network model. The core idea of PNA is to use the spatial neighborhood correlation in multi-scale features to perform neighborhood aggregation on high-level semantic features. In PNA, the multi-layer features of the backbone network are mainly used for two aspects. One is to use the correlation characteristics between neighboring pixels on features with different resolutions to aggregate high-level semantic features, encoding the neighborhood structure information on the shallow features in the backbone network onto the high-level semantic features. The other is to use the shallow features of the backbone network containing rich detail information to supplement the missing detail information. Therefore, PNA contains two key modules: the Neighborhood Aggregation Module (NAM) and the Self-Aggregation Module (SAM). Among them, NAM is used to calculate the similarity matrix between a pixel and its neighboring pixels, and then used to aggregate and enhance the high-level semantic features. SAM calculates the self-attention on the channels of the shallow features without destroying the spatial information to enhance important semantic information, thereby supplementing the detail information lost in the high-level semantic features.

[0045] The overall process is as shown in the attached Figure 2 (a). In the figure, F m , m ∈ {1, 2, 3, 4} represents multiple resolution features from the backbone network, and P m , m ∈ {1, 2, 3, 4} represents the gradually refined segmentation results, and P0 is the initially predicted rough segmentation result. FFN represents a feed-forward network composed of two linear layers. The calculation process of PNA is as follows: after obtaining the initial segmentation result P0, use the backbone network feature F1 to refine P0 to obtain a more refined segmentation result P1; after obtaining P1, use the backbone network feature F2 to refine P1 to obtain an even more refined segmentation result P2, and so on, gradually using the backbone network feature F m to refine the prediction result, and finally obtain the refined segmentation result P4. In PNA, the prediction result P m-1 and the feature F m obtain the semantically enhanced features after neighborhood aggregation through NAM At the same time, F m obtains the enhanced on the channels through SAM to supplement the missing detail information in . Add the two to obtain the finally fused featurem , the overall process is shown as follows. NAM and SAM will be described separately later.

[0046]

[0047]

[0048] Nearest Neighbor Aggregation Module NAM

[0049] The goal of the NAM module is to aggregate and enhance high-level semantic features using the nearest neighbor similarity matrix calculated on features at different resolutions. Since the predicted segmentation result contains the most semantic feature information, in each NAM module, the refined segmentation result P m-1 from the previous stage is used as the high-level semantic feature. The structure of NAM is shown in Figure 2 (b). Before aggregating P m using F m-1 , first interpolate P m-1 to ensure the same spatial resolution as F m . Then, both are passed through a linear layer to convert them to the same dimension size. For ease of explanation, hereinafter, P ∈ R HW×D and F ∈ R HW×D are used to represent the transformed high-level semantic feature P m-1 and the shallow feature F m .

[0050] For the feature map F ∈ R HW×D , let F i,j ∈ R D represent the feature at position (i, j), and let represent the features of all neighboring nodes of (i, j) within a window of size L×L. Since the present invention is a method for gradually refining the segmentation result from rough to fine, choosing an L×L window can pay more attention to local regions at different scales, capture more fine-grained information, and at the same time reduce the computational complexity. Inspired by the self-attention mechanism in Transformer, the present invention uses the attention mechanism to calculate the similarity matrix between each pixel and its surrounding neighboring pixels for each pixel. Apply linear layers to F and P respectively to obtain Q, K, V. As shown in the following formula,

[0051] Q = W q F, K = W k F, V = W v P

[0052] where W q , W k , W v ∈ R D×D is the weight matrix for linear mapping. Compared with F i,jAnd Similarly, use Q i,j to represent the query vector of (i, j), and use to represent the key matrix of the pixels within the window. As shown in the appendix Figure 2 (b), the similarity matrix between the pixel at (i, j) and its surrounding pixels is calculated by the following formula

[0053]

[0054] where B is the relative position encoding, defined by the parameterized bias matrix . S i,j ∈R L×L is the similarity matrix calculated at (i, j). The aggregated feature at (i, j) is as follows

[0055]

[0056] Based on the above method, all positions in P m can be aggregated through the similarity matrix. By leveraging the multi-resolution shallow features F1, F2, F3, F4, the high-level semantic features are progressively aggregated, and the spatial structure information from the shallow features is also gradually aggregated into the high-level semantic features during this process

[0057] Self-Aggregation Module SAM

[0058] In addition to the high-resolution structural information, the multi-resolution features in the backbone network also contain important detailed information lost in the high-level semantic features. However, directly adding and fusing these features with the semantic features may introduce harmful noise information. Considering this aspect, this patent designs a self-aggregation module to enhance the useful semantic features by calculating the attention between different channels of the features

[0059] Given the advantage of Transformer in modeling global information, it is combined with channel attention to design the self-aggregation module of this patent. For a given shallow feature F m , 1x1 convolution and 3x3 depthwise separable convolution are used to calculate respectively The depthwise separable convolution is used to capture the position information and can reduce the number of parameters and improve the computational efficiency. The calculation formula is as follows

[0060]

[0061] Here After that, the dimension of is changed to R C×HW . Next, the channel-level self-attention is calculated as shown in the following formula

[0062]

[0063] where α is a learnable scaling parameter. After that, perform LN regularization and pass it into a feed-forward network with a residual structure to obtain the self-aggregated features

[0064] Specific implementation method:

[0065] The PNA proposed by the present invention is a network module used to gradually refine the segmentation result, which can be added to the existing segmentation network and trained in an end-to-end manner. For the convenience of description, taking a single image and C semantic classes as an example, ResNet101 is selected as the backbone network, and the specific implementation method is described in combination with PNA.

[0066] Step 1: Input the input image I ∈ R 3×H×W into the backbone network, and four features with different resolutions can be extracted through the backbone network where m ∈ {1, 2, 3, 4}. After the feature F1 passes through 1x1Conv and the number of channels becomes C, corresponding to C semantic classes, the initial segmentation result can be obtained

[0067] Step 2: After obtaining F m , m ∈ {1, 2, 3, 4} and P0, use the PNA proposed in the present invention to gradually refine P0. PNA includes two main modules, NAM and SAM. The specific process is as follows:

[0068] Step 2-1: In the NAM module, use the feature F1 to refine P0. As described above, calculate the similarity between each position and its neighboring pixels within the window and use it to aggregate semantic features. During implementation, the window size L is designed to be 7. When the resolution of the input feature is less than 7×7, padding is required for the feature. The specific steps are as follows:

[0069] Step 2-2: First, change the dimensions of F1 and P0 from (C, H1, W1) to (H1, W1, C), and then perform regularization respectively. Judge whether H1 and W1 satisfy being greater than or equal to L. If not, fill 0 on the right side and the bottom.

[0070] Step 2-3: In NAM, multi-head attention is adopted, and the number of attention heads is set to h n , and the number of channels corresponding to each attention head is c n = C / h n . Q and K are obtained by changing the dimensions after F1 passes through the linear mapping shown in the following formula, and the dimension is (h n , H1, W1, c n ). V is obtained by changing the dimensions after P0 passes through the linear mapping shown in the following formula, and the dimension is (h n,H1,W1,c n )。

[0071] Q = W q F1, K = W k F1, V = W v P0

[0072] Step 2 - 4: Calculate the similarity between each position and its neighboring pixels as described above. The feature Q at (i, j) in Q i,j has a dimension of (h n , 1, 1, c n ) and is changed to (h n , 1, c n ). The key matrix of the neighboring pixels of (i, j) within the window is with a dimension of (h n , L×L, c n ). Calculate the similarity S between (i, j) and its neighboring pixels according to the following formula i,j , where B is the relative position encoding from the bias matrix which has been introduced above.

[0073]

[0074] The obtained S i,j has a dimension of (h n , 1, L×L). Then calculate the aggregated feature at (i, j) in P0 according to the following formula

[0075]

[0076] After calculating the features of all positions, the aggregated feature can be obtained

[0077] Step 3: In the SAM module, use the Transformer structure and select multi - head self - attention, where the number of attention heads is set to h s , and the dimension corresponding to each attention head is c s = C / h s .

[0078] Step 3 - 1: As shown in Figure (c) of the specification, the feature F1 is first regularized and then passed through 1x1Conv and 3x3DWConv to obtain Figure 2 as shown in the following formula:

[0079] The feature dimension of s is (h s , c s , H1×W1).

[0080] Step 3-2: After obtaining , multiply it with the transposed to obtain the attention attn on the channel, with the feature dimension being (h s , c s , c s ). After attn is processed by softmax, multiply it with to obtain the result of self-attention with the feature dimension size being (h s , c s , H1×W1). As shown in the following formula:

[0081]

[0082] Then change the dimension size of from (h s , c s , H1×W1) to (C, H1, W1), and after passing through a 1x1 Conv, obtain the output of the SAM module

[0083] Step 4: After obtaining and , add the two together, and denote the resulting value as After regularization, input it into the feed-forward network FFN, and after adding it with , input it into a 1x1 Conv for segmentation prediction to obtain the refined segmentation result P1. The structure and order of FFN are a linear layer, a GELU activation layer, a Dropout layer, a linear layer, and a Dropout layer. The entire process is as shown in the following formula:

[0084]

[0085]

[0086] Step 5: Upsample P1 to the same size as F2. Then repeat Steps 2 to 4 to obtain the refined segmentation result P2.

[0087] And so on, repeat the above process four times to obtain the final refined result P4.

Claims

1. A refined image semantic segmentation method based on progressive neighbor aggregation, characterized in that Including the following steps: Step 1: Input the input image \(I\in\mathbb{R}\) 3×H×W into the input backbone network, and four features with different resolutions are extracted through the backbone network where \(m\in\{1, 2, 3, 4\}\); Feature F1 passes through a 1x1 Conv, and the number of channels becomes C, corresponding to C semantic categories, thus obtaining the initial segmentation result Step 2: After obtaining F m and P0, the Progressive Neighbor Aggregation module PNA is used to gradually refine P0 using the feature F1. The Progressive Neighbor Aggregation module PNA includes two modules, namely the Neighbor Aggregation module NAM and the Self-Aggregation module SAM; In the Nearest Neighbor Aggregation Module (NAM), set a window of size L*L, and calculate the similarity between each position and its neighboring pixels within the window for aggregating semantic features, specifically as follows: Step 2-1: First, change the dimensions of F1 and P0 from (C, H1, W1) to (H1, W1, C), then perform regularization respectively. Judge whether both H1 and W1 are greater than or equal to L. If not, pad 0s on the right or below; Step 2-2: Use multi-head attention, and set the number of attention heads to h n , and the number of channels corresponding to each attention head is c n = C / h n ; Q and K are obtained by changing the dimension after the linear mapping of F1 as shown in Equation (1), and the dimension is (h n , H1, W1, c n ); V is obtained by changing the dimension after the linear mapping of P0 as shown in Equation (1), and the dimension is (h n , H1, W1, c n ); Q = W q F1, K = W k F1, V = W v P0(1) Among them, Q, K, and V respectively correspond to the Query matrix, Key matrix, and Value matrix in the Transformer structure, and F1 and P0 are transformation matrices; Step 2-3: Calculate the similarity between each position and its neighboring pixels; The feature Q at (i, j) in Q i,j is changed in dimension from (h n , 1, 1, c n ) to (h n , 1, c n ); the key matrix of the neighboring pixels of (i, j) within the window is with dimension (h n , L×L, c n ); calculate the similarity S between (i, j) and the neighboring pixels according to Equation (2) i,j , where B is the relative position encoding from the bias matrix which has been introduced above; Among them, B is the relative position encoding, which is defined by the parameterized bias matrix defined; The obtained S i,j has a dimensionality of (h n , 1, L×L), and then the aggregated feature at (i, j) in P0 is calculated according to Equation (3) After calculating the features at all positions, the aggregated features can be obtained Step 3: In the self-aggregation module SAM, a Transformer structure is adopted, and multi-head self-attention is selected, where the number of attention heads is set to h s , and the dimension c corresponding to each attention head s = C / h s ; Step 3-1: First, regularize the feature F1, and then obtain and through 1x1Conv and 3x3DWConv, which respectively correspond to the Query matrix, Key matrix, and Value matrix in the Transformer structure; as shown in Equation (4): which respectively correspond to the Query matrix, Key matrix, and Value matrix in the Transformer structure; as shown in Equation (4): The characteristic dimension is h s , c s , H1×W1; Step 3-2: After obtaining , multiply it with the transposed to get the channel attention attn, with the feature dimension being (h s , c s , c s ). After being processed by softmax, attn is multiplied with to obtain the result of self-attention with the feature dimension size of (h s , c s , H1×W1), as shown in the following formula: where α represents a learnable scaling parameter; Change again the dimension size of, from (h s , c s , H1×W1) to (C, H1, W1), and obtain the output of the SAM module after passing through a 1x1 Conv Step 4: After obtaining and , add the two together, and denote the resulting value as After regularization, input it into the feed-forward network FFN, and after adding it to , input it into 1x1Conv for segmentation prediction to obtain the refined segmentation result P1; the entire process is as follows: Step 5: Upsample P1 to the same size as F2; repeat the method from Step 2 to Step 4 to obtain the refined segmentation result P2; and so on, repeat the above process to obtain the final refined result P4.

2. The refined image semantic segmentation method based on progressive nearest neighbor aggregation according to claim 1, wherein, The backbone network is ResNet101.

3. A refined image semantic segmentation method based on progressive neighbor aggregation according to claim 1, characterized in that The L = 7.

4. The refined image semantic segmentation method based on progressive neighbor aggregation according to claim 1, wherein The structure and order of the FFN are a linear layer, a GELU activation layer, a Dropout layer, a linear layer, and a Dropout layer.

Citation Information

Patent Citations

  • Real-time traffic scene semantic segmentation method based on deep learning

    CN112802026A

  • Image segmentation method based on residual feature optimization and attention mechanism

    CN115311454A