High-Resolution Remote Sensing Image Semantic Segmentation Method Based on Hierarchical Detail Enhancement
By adopting a semantic segmentation method with hierarchical detail enhancement on high-resolution remote sensing images, the problem of image segmentation effect being affected by structural damage, insufficient local information and blurred boundaries is solved, and precise segmentation and context information recovery are achieved.
Patent Information
- Application Number
- CN202210379126.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-04-12
AI Technical Summary
The semantic segmentation effect of high-resolution remote sensing images is affected by structural damage, insufficient local information and blurred boundaries, making it difficult to achieve refined segmentation while maintaining context information.
Using a hierarchical detail enhancement method, high-resolution remote sensing images are processed through a semantic segmentation network, including boundary and body separation, detail enhancement modules and cross-detail decoder, combining high- and low-level features to restore the context information and details of the image.
It realizes accurate segmentation on high-resolution remote sensing images, especially on data sets with a resolution of 0.3m, improves the segmentation effect of boundary details, and effectively restores the context information of the image.
Smart Images

Figure CN114723948B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular to a semantic segmentation method for high-resolution remote sensing images based on hierarchical detail enhancement. Background Art
[0002] With the progress of satellite technology, the resolution of remote sensing images has been continuously improved, and the application scenarios of remote sensing images have become more extensive. The semantic segmentation of remote sensing images is to distinguish different ground object types in the image according to their different spectral characteristics, and it is an important means for automatic extraction and intelligent recognition of image targets. Although high-resolution remote sensing images provide more ground object features and details, their complexity and diversity also bring difficulties to semantic segmentation on high-resolution remote sensing images.
[0003] The segmentation effect of high-resolution remote sensing images is affected by the following three factors: maintaining high resolution, global context information, and boundary details. First, because high-resolution images are cut into fragments before segmentation operations, the structure is damaged, thus destroying spatial information, detail information, and context information; second, since it is difficult to distinguish categories by only using local information, global context information should be used to help search for category correlations in the image; finally, due to the movement of the satellite, the boundaries of the targets are often blurred, resulting in a decline in the classification effect.
[0004] Therefore, a refined semantic segmentation method for high-resolution remote sensing images while maintaining context information is needed. Summary of the Invention
[0005] The purpose of the present invention is to provide a semantic segmentation method for high-resolution remote sensing images based on hierarchical detail enhancement, which can effectively achieve precise segmentation on high-resolution remote sensing images and achieve better results on a high-resolution remote sensing image dataset with a resolution of 0.3m.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A semantic segmentation method for high-resolution remote sensing images based on hierarchical detail enhancement, the method includes the following steps in sequence:
[0007] (1) Collect high-resolution remote sensing image data, construct a high-resolution remote sensing image dataset, and perform annotation and preprocessing on the high-resolution remote sensing image data;
[0008] (2) Establish a semantic segmentation network, input the high-resolution remote sensing image data in the high-resolution remote sensing image dataset into the semantic segmentation network, and the semantic segmentation network generates a segmentation map.
[0009] The step (2) specifically includes the following steps:
[0010] (2a) Using the hierarchical semantic segmentation network based on the sliding window and attention mechanism as the benchmark model, a total of four feature maps {L 1 , L 2 , L 3 , L 4} with different resolution sizes are generated. After upsampling the feature map L 2 and merging it with the feature map L 1 , the low-level feature map F low is obtained. After upsampling the feature map L 4 and merging it with the feature map L 3 , the high-level feature map F high is obtained;
[0011] (2b) For the low-level feature map F low , the boundary feature map F edge and the main body feature map F body are extracted by separating the boundary and the main body;
[0012] (2c) After obtaining the feature map F′ high from the high-level feature map F high using the detail enhancement module, it is combined with the boundary feature map F edge . The combined feature map passes through the cross convolution and the channel attention module to obtain the feature map F r ;
[0013] (2d) The feature map F r and the main body feature map F body are combined to generate the feature map P 1 for refining the edge;
[0014] (2e) The feature map P 4 is generated by using the cross-detail decoder for the feature map L 2 . The feature map P 1 is generated by using the cross-detail decoder for the feature map P 3 ;
[0015] (2f) The feature map P 2 and the feature map P 3 are combined to generate the prediction feature map, and after linear regression on the prediction feature map, the final segmentation map is generated.
[0016] In step (2a), using the hierarchical semantic segmentation network based on the sliding window and attention mechanism as the benchmark model, a total of four feature maps {L 1 , L 2 , L 3 , L 4} with different resolution sizes are generated, which specifically includes the following steps:
[0017] (2a1) The hierarchical semantic segmentation network based on the sliding window and attention mechanism first divides the output RGB image into non-overlapping patches through the patch segmentation module;
[0018] (2a2) Each non-overlapping patch is regarded as a token, and the token contains the pixel-level features of the image;
[0019] (2a3) The hierarchical semantic segmentation network based on the sliding window and attention mechanism includes 4 stages. Each stage includes multiple semantic segmentation network modules, and each semantic segmentation network module is composed of a multi-head self-attention mechanism and a self-attention mechanism with a moving window;
[0020] (2a4) After each token is first merged into feature maps of different resolution sizes through patch merging, it passes through the semantic segmentation network module to generate feature maps of different resolution sizes {L 1 , L 2 , L 3 , L 4}.
[0021] The specific steps of the said step (2b) include the following steps:
[0022] (2b1) Generate the flow field for the low-level feature map F low . Generate the low-frequency feature map by downsampling the low-level feature map F low , and then upsample the low-frequency feature map to obtain a feature map of the same size as the low-level feature map F low
[0023]
[0023] (2b2) Connect the low-level feature map F low , the low-frequency feature map and the feature map , and then use a 3×3 convolutional layer for compression processing to obtain the low-level prediction feature map M;
[0024] (2b3) Perform feature warping on the low-level feature map F low . Reassign a new coordinate position p i for each pixel point at the position p i in the original standard image as p i + M(p body ). Use the differentiable bilinear sampling mechanism to approximate each pixel point p x in the main feature map F
[0025]
[0026] where N represents p iThe set of 4 nearest neighbor pixels, where n represents a pixel in this set, and m n is the bilinear kernel weight on the warped spatial grid, calculated from the low-level predicted feature map M;
[0027] (2b4) Boundary feature map F edge is obtained by subtracting the low-level feature map F low from the main body feature map F body , that is:
[0028] F edge = F low - F body .
[0029] The specific steps of step (2c) include the following steps:
[0030] (2c1) The detail enhancement module consists of cross convolutions, and the cross convolution is composed of two asymmetric vertical filters k 1×t and k t×1 . Among them, the receptive field size of filter k 1×t is 1×t, and the receptive field size of filter k t×1 is t×1; assuming the input feature of the cross convolution is then the output feature is composed of the convolution of the input feature and the vertical filter plus the bias b, and the calculation formula is as follows:
[0031]
[0032] (2c2) The high-level feature map F high obtains a feature map after passing through the detail enhancement module, and this feature map is upsampled to obtain a feature map F′ low with the same size as the low-level feature map F high ;
[0033] (2c3) The feature map F′ high and the boundary feature map F edge are combined and then pass through the cross convolution and the channel attention module to obtain the feature map F r .
[0034] The specific meaning of step (2d) is:
[0035] Combine the main body feature map F body and the feature map F r to obtain the feature map P 1 for refining the edge, and the calculation formula is as follows:
[0036] P 1 = F r + F body .
[0037] Step (2e) specifically includes the following steps:
[0038] (2e1) The cross-detail decoder obtains the feature maps L 3 and L 4 as well as the category information cls 3 and cls 4 , where where Nc represents the number of categories, represents the k-th class at stage i;
[0039] (2e2) The feature maps L 3 and L 4 correspond to the flags z 3 and z 4 . The flags z 3 and z 4 are linearly transformed into the corresponding k and v through linear transformation. The category information cls 3 and cls 4 are linearly transformed into q through linear transformation. The flags z 3 and z 4 are operated through the cross-attention layer as follows:
[0040] k = z i W k , v = z i W v , q = cls i W q
[0041]
[0042] where W k , W v , W q represent adjustable parameters, D represents the dimension, h represents the number of heads of the cross-attention mechanism set, CA represents the output feature after passing through the cross-attention layer; i is equal to 3 or 4;
[0043] (2e3) The output feature obtained after passing through the cross-attention layer is used as the input feature to enter the multi-layer perceptron layer. The multi-layer perceptron layer adds a cross-detail branch. The calculation formula of the multi-layer perceptron layer is as follows:
[0044] x 1 = FC(x in , θ 1 )
[0045] x out = FC(σ(x 1 + CS(x 1 )), θ2 )
[0046] Among them, x in represents the features input into the improved multi-layer perceptron layer, and x out represents the output features, CS represents the added cross-convolution module branch, and θ 1 and θ 2 represent optional parameters, FC represents the fully connected layer, and σ represents the activation function;
[0047] (2e4) Feature map L 4 After completing the cross-detail decoder, the corresponding feature map P 2 is obtained, and the feature map P 1 completes the cross-detail decoder to obtain the corresponding feature map P 3 .
[0048] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, the present invention can effectively achieve accurate segmentation on high-resolution remote sensing images and has achieved better results on the high-resolution remote sensing image dataset with a resolution of 0.3m; Second, aiming at the disadvantage of blurred boundary details in semantic segmentation of high-resolution remote sensing images, a hierarchical detail enhancement method is adopted, and different methods are used to enhance the details of high-level features and low-level features respectively; Third, the present invention adds a cross-detail decoder to the high-level features and low-level features after detail enhancement, realizes the interaction of context information and global information while realizing detail enhancement, so as to better restore to the high-resolution image mapping; Fourth, different from ordinary single-branch or refinement only for a certain feature layer, the present invention comprehensively utilizes the more context information contained in the high-level features, and the low-level features maintain a higher resolution. After detail enhancement of the high-level features and low-level features, the combination of the two feature maps effectively utilizes the advantages contained in the high-level features and low-level features. Description of the Drawings
[0049] Figure 1 is the flowchart of the method of the present invention;
[0050] Figure 2 is the structural diagram of the semantic segmentation network of the present invention;
[0051] Figure 3 is a color map of a certain data sample on the high-resolution remote sensing image dataset;
[0052] Figure 4 is the annotation map of the data sample;
[0053] Figure 5 is the experimental effect diagram output by the present invention. Detailed Embodiment
[0054] As Figure 1As shown, a semantic segmentation method for high - resolution remote sensing images based on hierarchical detail enhancement, the method includes the following steps in sequence:
[0055] (1) Collect high - resolution remote sensing image data, construct a high - resolution remote sensing image dataset, and perform annotation and pre - processing on the high - resolution remote sensing image data;
[0056] (2) Establish a semantic segmentation network, input the high - resolution remote sensing image data in the high - resolution remote sensing image dataset into the semantic segmentation network, and the semantic segmentation network generates a segmentation map.
[0057] The step (2) specifically includes the following steps:
[0058] (2a) Use a hierarchical semantic segmentation network based on a sliding window and an attention mechanism as a reference model to generate a total of four feature maps {L 1 , L 2 , L 3 , L 4} with different resolution sizes. After upsampling the feature map L 2 and merging it with the feature map L 1 , a low - level feature map F low is obtained. After upsampling the feature map L 4 and merging it with the feature map L 3 , a high - level feature map F high is obtained;
[0059] (2b) Adopt a method of separating the boundary and the main body for the low - level feature map F low to extract a boundary feature map F edge and a main - body feature map F body ;
[0060] (2c) Use a detail enhancement module for the high - level feature map F high to obtain a feature map F′ high , then combine it with the boundary feature map F edge . The combined feature map passes through a cross - convolution and a channel attention module to obtain a feature map F r ;
[0061] (2d) Combine the feature map F r and the main - body feature map F body to generate a feature map P 1 with refined edges;
[0062] (2e) Use a cross - detail decoder for the feature map L 4 to generate a feature map P 2 . Use a cross - detail decoder for the feature map P 1 to generate a feature map P 3 ;
[0063] (2f) Combine the feature map P 2 and the feature map P 3 to generate a predicted feature map, and generate the final segmentation map after performing linear regression on the predicted feature map.
[0064] As Figure 2 shown, in step (2a), using the hierarchical semantic segmentation network based on the sliding window and attention mechanism as the benchmark model, a total of four feature maps {L1, L2, L3, L4} with different resolution sizes are generated, which specifically include the following steps:
[0065] (2a1) The hierarchical semantic segmentation network based on the sliding window and attention mechanism first divides the output RGB image into non-overlapping patches through the patch segmentation module;
[0066] (2a2) Each non-overlapping patch is regarded as a flag, and the flag contains the pixel-level features of the image;
[0067] (2a3) The hierarchical semantic segmentation network based on the sliding window and attention mechanism includes 4 stages, each stage includes multiple semantic segmentation network modules, and each semantic segmentation network module is composed of a multi-head self-attention mechanism and a self-attention mechanism of the moving window;
[0068] (2a4) After each flag is first merged into feature maps of different resolution sizes through patch merging, it passes through the semantic segmentation network module to generate feature maps {L 1 , L 2 , L 3 , L 4} of different resolution sizes.
[0069] The specific steps of step (2b) include the following steps:
[0070] (2b1) Generate a flow field for the low-level feature map F low , generate a low-frequency feature map by downsampling the low-level feature map F low , and then upsample the low-frequency feature map to obtain a feature map of the same size as the low-level feature map F low
[0071] (2b2) Connect the low-level feature map F low , the low-frequency feature map feature map , and then use a 3×3 convolutional layer for compression processing to obtain the low-level predicted feature map M;
[0072] (2b3) For the low-level feature map F low Perform feature distortion, and reassign a new coordinate position p for each pixel at position p on the original standard image i + M(p i ), and approximate each pixel p in the body feature map F i using a differentiable bilinear sampling mechanism. The formula of the differentiable bilinear sampling mechanism is as follows: body where N represents the set of 4 nearest neighbor pixels of p x , n represents a pixel in this set, and m
[0073]
[0074] is the bilinear kernel weight on the distorted spatial grid, calculated from the low-level prediction feature map M; i n edge (2b4) The boundary feature map F
[0075] is obtained by subtracting the low-level feature map F edge from the body feature map F low , that is: body
[0076] F edge = F low - F body .
[0077] Step (2c) specifically includes the following steps:
[0078] (2c1) The detail enhancement module consists of cross convolutions. The cross convolution is composed of two asymmetric vertical filters k 1×t and k t×1 . Among them, the receptive field size of the filter k 1×t is 1×t, and the receptive field size of the filter k t×1 is t×1; assuming the input feature of the cross convolution is then the output feature is composed of the convolution of the input feature and the vertical filter plus the bias b. The calculation formula is as follows:
[0079]
[0080] (2c2) The high-level feature map F high obtains a feature map after passing through the detail enhancement module. Upsample this feature map to obtain a feature map F′ low of the same size as the low-level feature map F high ;
[0081] (2c3) The feature map F′ high and the boundary feature map F edge After combination, the feature map F is obtained through cross convolution and the channel attention module r .
[0082] The specific step (2d) refers to:
[0083] Combine the main feature map F body and the feature map F r to obtain the feature map P for refining the edge 1 , and the calculation formula is as follows:
[0084] P 1 = F r + F body .
[0085] The specific step (2e) includes the following steps:
[0086] (2e1) The cross-detail decoder obtains the feature maps L 3 , L 4 and the category information cls 3 , cls 4 , where where Nc represents the number of categories, represents the k-th class in stage i;
[0087] (2e2) The feature maps L 3 , L 4 correspond to the flags z 3 , z 4 , and the flags z 3 , z 4 are transformed into the corresponding k and v through linear transformation, and the category information cls 3 , cls 4 is transformed into q through linear transformation. The operations of the flags z 3 , z 4 through the cross-attention layer are as follows:
[0088] k = z i W k , v = z i W v , q = cls i W q
[0089]
[0090] where W k , W v , W q represent adjustable parameters, D represents the dimension, h represents the number of heads of the set cross-attention mechanism, CA represents the output feature after passing through the cross-attention layer; i is equal to 3 or 4;
[0091] (2e3) The output features obtained after passing through the cross-attention layer enter the multi-layer perceptron layer as input features. The multi-layer perceptron layer adds a cross-detail branch, and the calculation formula of the multi-layer perceptron layer is as follows:
[0092] x 1 = FC(x in , θ 1 )
[0093] x out = FC(σ(x 1 + CS(x 1 ))), θ 2 )
[0094] Among them, x in represents the features input to the improved multi-layer perceptron layer, x out represents the output features, CS represents the added cross-convolution module branch, θ 1 and θ 2 represent optional parameters, FC represents the fully connected layer, and σ represents the activation function;
[0095] (2e4) The feature map L 4 obtains the corresponding feature map P after completing the cross-detail decoder 2 , and the feature map P 1 completes the cross-detail decoder to obtain the corresponding feature map P 3 .
[0096] As Figure 3 shown, a certain data sample color map on the high-resolution remote sensing image dataset is a remote sensing image with a resolution of 0.3m, and the image contains woods, buildings, farmlands, roads, etc.
[0097] As Figure 4 shown, the data sample annotation map is the image obtained by annotating the Figure 3 high-resolution remote sensing image. The same type of things are annotated with the same color, and different types of things are annotated with different colors. The data sample annotation maps are all manually annotated.
[0098] As Figure 5 shown, the experimental effect diagram output by the method of the present invention. The data sample is input into the high-resolution remote sensing image semantic segmentation network with hierarchical detail enhancement to obtain a segmentation map.
[0099] In summary, the present invention can effectively achieve accurate segmentation on high-resolution remote sensing images and has achieved better results on a high-resolution remote sensing image dataset with a resolution of 0.3 m. Aiming at the disadvantage of blurred boundary details in semantic segmentation of high-resolution remote sensing images, a hierarchical detail enhancement method is adopted to enhance the details of high-level features and low-level features in different ways respectively.
Claims
1. A semantic segmentation method for high-resolution remote sensing images based on hierarchical detail enhancement, characterized in that: This method includes the following steps in sequence: (1) Collect high-resolution remote sensing image data, construct a high-resolution remote sensing image dataset, and perform annotation and preprocessing on the high-resolution remote sensing image data; (2) Establish a semantic segmentation network, input the high-resolution remote sensing image data in the high-resolution remote sensing image dataset into the semantic segmentation network, and the semantic segmentation network generates a segmentation map; The step (2) specifically includes the following steps: (2a) Using the hierarchical semantic segmentation network based on the sliding window and attention mechanism as the benchmark model, a total of four feature maps {L 1 , L 2 , L 3 , L 4} with different resolution sizes are generated. After upsampling the feature map L 2 and merging it with the feature map L 1 , the low-level feature map F low is obtained. After upsampling the feature map L 4 and merging it with the feature map L 3 , the high-level feature map F high is obtained; (2b) For the low-level feature map F low Adopt the method of separating the boundary and the main body to extract the boundary feature map F edge and the main body feature map F body ; (2c) For the high-level feature map F high Use the detail enhancement module to obtain the feature map F' high Then combine it with the boundary feature map F edge After combination, the combined feature map passes through the cross convolution and channel attention module to obtain the feature map F r ; (2d) Combine the feature map F r and the main feature map F body to generate the feature map P for refining the edges 1 ; (2e) For the feature map L 4 Use a cross-detail decoder to generate the feature map P 2 , for the feature map P 1 Use a cross-detail decoder to generate the feature map P 3 ; (2f) Combine the feature map P 2 and the feature map P 3 to generate a predicted feature map, and generate the final segmentation map after performing linear regression on the predicted feature map; The step (2b) specifically includes the following steps: (2b1) Generate the flow field for the low-level feature map F low by performing downsampling on the low-level feature map F low to generate the low-frequency feature map Then, perform upsampling on the low-frequency feature map to obtain a feature map of the same size as the low-level feature map F low (2b2) Connect the low-level feature map F low and the low-frequency feature map feature map After connection, use a 3×3 convolutional layer for compression processing to obtain the low-level predicted feature map M; (2b3) Feature distortion is performed on the low-level feature map F low and a new coordinate position p i is reassigned to each pixel at position p i + M(p i ) on the original standard image. The differentiable bilinear sampling mechanism is used to approximate each pixel p body in the main feature map F x . The formula of the differentiable bilinear sampling mechanism is as follows: where N represents the set of 4 nearest neighbor pixels of p i , n represents a pixel in this set, and m n is the bilinear kernel weight on the warped spatial grid, calculated from the low-level predicted feature map M; (2b4) Boundary Feature Map F edge From the low-level feature map F low Subtracted from the main body feature map F body That is: F edge = F low -F body 。 2. The semantic segmentation method for high-resolution remote sensing images based on hierarchical detail enhancement according to claim 1, characterized in that: In step (2a), using the hierarchical semantic segmentation network based on the sliding window and the attention mechanism as the benchmark model, a total of four feature maps {L 1 , L 2 , L 3 , L 4} with different resolution sizes are generated, which specifically include the following steps: (2a1) The hierarchical semantic segmentation network based on the sliding window and attention mechanism first divides the output RGB image into non-overlapping patches through the patch segmentation module; (2a2) Each non-overlapping patch is regarded as a flag, and the flag contains pixel-level features of the image; (2a3) The hierarchical semantic segmentation network based on the sliding window and attention mechanism includes 4 stages, each stage includes multiple semantic segmentation network modules, and each semantic segmentation network module is composed of a multi-head self-attention mechanism and a self-attention mechanism of the moving window; (2a4) After each logo is first merged into feature maps of different resolution sizes through patching, it passes through a semantic segmentation network module to generate feature maps {L 1 , L 2 , L 3 , L 4} of different resolution sizes.
3. The semantic segmentation method for high-resolution remote sensing images based on hierarchical detail enhancement according to claim 1, characterized in that: The step (2c) specifically includes the following steps: (2c1) The detailed enhancement module consists of cross convolutions, and the cross convolution consists of two asymmetric vertical filters k 1×t and k t×1 . Among them, the receptive field size of filter k 1×t is 1×t, and the receptive field size of filter k t×1 is t×1; assuming that the input feature of the cross convolution is , then the output feature is composed of the convolution of the input feature and the vertical filter plus the bias b, and the calculation formula is as follows: (2c2) High-level feature map F high The feature map obtained after passing through the detail enhancement module is upsampled to obtain a feature map F with the same size as the low-level feature map low F' high ; (2c3) Feature map F' high and the boundary feature map F edge are combined and then passed through a cross convolution and a channel attention module to obtain the feature map F r .
4. The semantic segmentation method for high-resolution remote sensing images based on hierarchical detail enhancement according to claim 1, characterized in that: The step (2d) specifically refers to: Combine the main feature map F body and the feature map F r to obtain the refined edge feature map P 1 , and the calculation formula is as follows: P 1 = F r + F body .
5. The semantic segmentation method for high-resolution remote sensing images based on hierarchical detail enhancement according to claim 1, characterized in that: The step (2e) specifically includes the following steps: (2e1) The cross-detail decoder obtains the feature map L from the backbone network 3 、L 4 and the category information cls 3 、cls 4 , where where Nc represents the number of categories, represents the k-th class in stage i; (2e2) Feature map L 3 、L 4 correspond to the flag z 3 、z 4 , the flag z 3 、z 4 are transformed into the corresponding k, v, and class information cls 3 、cls 4 are linearly transformed into q, and the flag z 3 、z 4 The operations through the cross-attention layer are as follows: k = z i W k , v = z i W v , q = cls i W q Among them, W k and W v and W q represent adjustable parameters, W k and W v and W q all D represents the dimension, h represents the number of heads of the set cross-attention mechanism, CA represents the output features after passing through the cross-attention layer; i is equal to 3 or 4; (2e3) The output features obtained after passing through the cross-attention layer are used as input features to enter the multi-layer perceptron layer. The multi-layer perceptron layer adds a cross-detail branch, and the calculation formula of the multi-layer perceptron layer is as follows: x 1 = FC(x in , θ 1 ) x out = FC(σ(x 1 + CS(x 1 )), θ 2 ) Among them, x in represents the features input to the improved multi-layer perceptron layer, and x out represents the output features. CS represents the added cross-convolution module branch, and θ 1 and θ 2 represent optional parameters, FC represents the fully connected layer, and σ represents the activation function; (2e4) Feature map L 4 After completing the cross-detail decoder, the corresponding feature map P is obtained 2 , feature map P 1 Completing the cross-detail decoder to obtain the corresponding feature map P 3 .
Citation Information
Patent Citations
Chicken image segmentation method and system based on multi-scale attention network
CN113052848A