Cross-scale graph similarity-guided aggregation system, method and application

By introducing a cross-scale map similarity guided aggregation system in remote sensing image semantic segmentation, the cross-scale map interaction module and multi-scale similarity guided aggregation module are used to solve the neglected problem of span scale correlation and boundary multi-scale characteristics in the existing technology, and a more powerful semantic feature representation and edge feature assistance effects are achieved.

CN115880552BActive Publication Date: 2025-05-23OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211223060.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2022-10-08
Publication Date
2025-05-23
Estimated Expiration
2042-10-08

AI Technical Summary

Technical Problem

The prior art ignores the multi-scale characteristics of cross-scale correlation and boundaries in the semantic segmentation of remote sensing images, resulting in insufficient models in semantic feature extraction and boundary refinement.

Method used

By designing a cross-scale map similarity guided aggregation system, using the cross-scale map interactive module and multi-scale similarity guided aggregation module, a cross-scale map structure is built and semantic features and boundary features are aggregated to enhance the representation ability of the model and the auxiliary effect of edge features.

Benefits of technology

This method can not only meet the interaction between cross-scale targets, but also effectively mine multi-scale boundary information, realize robust aggregation, and significantly improve the representation ability and semantic segmentation effect of remote sensing features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115880552B_ABST
    Figure CN115880552B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image processing, and discloses a cross-scale graph similarity guided aggregation system, method and application for semantic segmentation of remote sensing images. The system includes two independent subtask branches, namely a semantic feature extraction branch and a boundary feature extraction branch. In the semantic feature extraction branch, a cross-scale graph interaction module CGI is introduced to construct a graph structure, and graph convolution is used to infer and aggregate the association relationship between cross-scale nodes to enhance the representation ability of remote sensing features; in the boundary feature extraction branch, a multi-scale similarity guided aggregation module MSA is introduced to extract multi-scale boundary features to improve the auxiliary effect of edge features on semantic segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to a cross-scale graph similarity guided aggregation system, method and application. Background Art

[0002] Remote sensing images are widely used in environmental monitoring, land resource management, and disaster assessment. In this application, semantic segmentation is one of the key technologies for remote sensing images, which is to divide each pixel in the input image into a semantic category. However, remote sensing images have diverse geophysical properties and large computational complexity, so it is difficult to achieve effective semantic segmentation.

[0003] In recent years, convolutional neural networks have greatly promoted the development of semantic segmentation of remote sensing images with their powerful feature extraction capabilities. The fully convolutional network (FCN) first modified the fully connected layer into a convolutional layer, making it a fully convolutional network and realizing end-to-end training at the pixel level. Subsequently, in order to better restore image detail information, an "encoder-decoder" structure was proposed. This structure uses jump connections to connect low-level detail information with high-level semantic information, so that the model can obtain more detail information and enhance the model's prediction ability. However, these methods expose a common shortcoming in the extraction of semantic information. The model is limited to a fixed geometric structure and a limited receptive field. To this end, a multi-scale context fusion technology has emerged. It applies specific techniques such as dilated convolution or pyramid pooling modules to aggregate context. Although it can effectively mine multi-scale context information, it is ignored in terms of cross-scale information interaction. There are often important correlations between cross-scales. This is often crucial for semantic segmentation. In addition, attention-based mechanisms and graph convolutional networks (GCNs) adaptively capture a wide range of dependencies from the channel or spatial dimension, thereby effectively expanding the range of the receptive field. The above strategies solve the multi-scale problem to a certain extent and enhance the representation ability of the model. In addition, in order to further perform fine-grained segmentation of remote sensing images, some methods have recently explored the boundary detection module as an independent branch in parallel, using the extracted edge contour features as a supplement, which is crucial to improving the boundary refinement ability and solving the consistency problem in semantic segmentation.

[0004] Although the above methods are of great significance and value, they still have some defects: 1) For multi-scale models, cross-scale correlations are ignored in the context modeling process. 2) In the boundary detection process, the multi-scale characteristics of the boundary are ignored. 3) Boundary information is sparse, and there is a sample imbalance problem, which leads to unreliable guidance of semantic features. In remote sensing images, the number of boundary pixels often only accounts for a small part of the entire image, and its edge features are usually a sparse matrix, which cannot effectively guide and improve the semantic segmentation effect. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a cross-scale graph similarity guided aggregation system, method and application. The present invention constructs a graph structure through a cross-scale graph interaction module, and uses graph convolution to infer and aggregate the associations between cross-scale nodes to enhance the representation capability of remote sensing features; and aggregates semantic features and boundary features through a multi-scale similarity guided aggregation module to improve the auxiliary effect of edge features on semantic segmentation.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0007] Firstly, the present invention provides a cross-scale graph similarity guided aggregation system, which includes a backbone network and two independent sub-task branches, namely a semantic feature extraction branch and a boundary feature extraction branch. In the semantic feature extraction branch, a cross-scale graph interaction module CGI is introduced to extract semantic features, and in the boundary feature extraction branch, a multi-scale similarity guided aggregation module MSA is introduced to extract multi-scale boundary features.

[0008] The backbone network first uses conventional convolution operations to mine the rich semantic features X of the original image i Then, the dilated convolution method is used to change the receptive field of the convolution operation with different dilation rates to further explore the multi-scale semantic features of the original image F k , generate a multi-scale semantic feature map; finally, the feature map mined by the backbone network is used as input to enter the subsequent two branch networks, where the semantic feature X i Input into the edge feature extraction branch, and F k Input into the semantic feature extraction branch;

[0009] The semantic feature extraction branch includes a cross-scale graph interaction module GCI and a graph convolutional network GCN. In the semantic feature extraction branch, the multi-scale semantic features F of the original image mined by the backbone network are used. k As input, it is input into the cross-scale graph interaction module CGI, and the cross-scale graph model is established by constructing the relationship between graph nodes and edges of different scales. Finally, the graph convolutional network GCN is used to infer and aggregate the association relationship between cross-scale semantic features to enhance the representation ability of the model and extract the cross-scale semantic feature G of the semantic feature. i ;

[0010] The boundary feature extraction branch includes a multi-scale similarity guided aggregation module MSA, which includes a multi-scale boundary feature extraction MBFE unit and a similarity guided aggregation SGA unit. The multi-scale similarity guided aggregation module MSA extracts the rich semantic features X iAs input, a preliminary feature fusion is first performed, and supervised training is performed to obtain the boundary feature B containing boundary information; then, the boundary feature B is input into the MBFE unit; the MBFE unit uses dilated convolutions with different expansion rates to detect multi-scale boundary information and extract the multi-scale boundary feature B i , SGA unit calculates boundary feature B i and the cross-scale semantic feature G output by the semantic feature extraction branch i The similarity between them is calculated and a multiplication operation is performed to aggregate the cross-scale semantic features G i and boundary feature B i The two multi-scale features are used to improve the auxiliary effect of edge features on semantic segmentation, and finally a feature map that combines semantic and edge information is output.

[0011] Furthermore, the cross-scale graph interaction module CGI first integrates the multi-scale semantic feature map generated by the backbone network into a cross-scale feature map through feature splicing, and then performs graph reasoning on the cross-scale feature map according to the graph convolution operation, converts the spatial pixels into nodes in the graph model, and uses the similarity matrix of the calculated nodes as the edge of the graph model; finally, cross-scale reasoning is performed through the message passing mechanism in the graph convolution network GCN to aggregate information, and through the role of cross-scale graph reasoning, semantic analysis interacts between multiple scales.

[0012] Secondly, the present invention also provides an application of the cross-scale graph similarity guided aggregation system as described above, which is used for semantic segmentation of remote sensing images.

[0013] Finally, the present invention provides a method for semantic segmentation using the cross-scale graph similarity guided aggregation system as described above, and the specific method is as follows:

[0014] S1. First, the original image is input into the backbone network. The backbone network uses convolution operations to mine the rich semantic features of the original image X. i On the one hand, by setting different void rates to change the receptive field size of the convolution operation, the multi-scale semantic features F of the original image can be mined. k , generate multi-scale semantic feature maps;

[0015] S2. The multi-scale semantic feature graph generated by the backbone network is input into the cross-scale graph interaction module GCI, with spatial pixels as nodes of the cross-scale graph model, and the similarity matrix of the nodes is calculated as the edge of the graph model. Subsequently, the cross-scale reasoning is performed through the message passing mechanism in the graph convolution network GCN to aggregate information. Through the role of cross-scale graph reasoning, the semantic information interacts between multi-scale features, and finally the cross-scale semantic feature G is obtained. i ;

[0016] S3. In the boundary feature extraction branch, the dilated convolution method is used to mine multi-scale boundary features. The semantic features X output by the backbone network i As input, a preliminary feature fusion is first performed, and supervised training is performed to obtain the boundary feature B containing boundary information; then, the boundary feature B is input into the MBFE unit. For the boundary feature B, the MBFE unit uses a hole convolution with different expansion rates to act on the boundary feature B. After feature mining, boundary features B of different scales are obtained. i , realizing multi-scale boundary feature mining; in order to facilitate the subsequent fusion with semantic features, the mined boundary feature B i The number and semantic features of G i Stay consistent;

[0017] S4. The boundary features mined by the boundary feature extraction branch are sparse matrices, which leads to sample imbalance. The similarity-guided aggregation SGA unit is used to calculate the boundary feature B. i and the semantic feature G output by the semantic feature extraction branch i The similarity between them is calculated to obtain the strongest similarity area, which is used as the weight matrix to perform weighted fusion of the original semantic features;

[0018] S5. Finally, the feature map that integrates semantic and edge information is output.

[0019] Furthermore, in step S2, for the multi-scale semantic features F generated by the backbone network k , the spatial pixel points of the feature are regarded as nodes and their sizes are converted to F k ∈R n×d , the cross-scale node set is F = Each node f encodes a different region in the original image. The values ​​of n and d are determined by the spatial and channel sizes of the multi-scale semantic features. The edges of the graph are defined as pairwise similarity calculations between image regions, and the relationship is constructed by the following equation: in and is a regular convolution whose parameters are learned by back-propagation, and They represent the i-th node at the p-th scale and the j-th node at the q-th scale respectively.

[0020] Furthermore, in step S4, the similarity-guided aggregation SGA unit performs similarity calculation in the following specific steps: given a multi-scale semantic graph feature and boundary features Lowercase letters g and b represent the corresponding feature maps respectively. The subscript k represents the number of features currently. k = {1, 2, 3, 4, 5}. The superscript n represents the position of the feature map. First, the similarity between the two is calculated. in are two nonlinear transformations, and They represent the parameters of nonlinear transformation respectively, and the superscript T represents matrix transpose. The function is used to calculate the influence value between the jth position on the boundary and the ith position on the semantic map; then, matrix multiplication is performed between the multi-scale boundary features and the similarity matrix Among them, α is the parameter obtained by back propagation. According to the above calculation, the boundary area of ​​the same category will be activated with a higher weight than other irrelevant areas.

[0021] Compared with the prior art, the present invention has the following advantages:

[0022] The present invention can not only satisfy the interaction between cross-scale targets, but also mine multi-scale boundary information, realize robust aggregation, and greatly improve the representation ability of remote sensing features. Specifically, the present invention designs a cross-scale graph interaction (CGI) module, which establishes a cross-scale graph structure and performs adaptive graph reasoning to capture cross-scale semantic related information. The present invention also designs a multi-scale similarity-guided aggregation (MSA) module to mine multi-scale boundary information and complete reliable boundary feature guidance. The module consists of a multi-scale boundary feature extraction (MBFE) unit and a similarity-guided aggregation (SGA) unit. The MBFE unit uses Atos convolution with different dilation rates to detect multi-scale boundary information, and the SGA unit uses similarity calculation to enhance the robustness and reliability of boundary guidance. Through numerical experiments conducted on two benchmark remote sensing datasets, the semantic segmentation method proposed by the present invention is superior to the most advanced method in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0024] Figure 1 is a system architecture diagram of the present invention;

[0025] Figure 2 Extracting branch structure diagram for semantic features of the present invention;

[0026] Figure 3 Similarity guides for the present invention are shown in the polymerized unit structure diagram. DETAILED DESCRIPTION

[0027] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0028] Example 1

[0029] The present invention proposes a cross-scale graph similarity guided aggregation network (CGSAN) for semantic segmentation of remote sensing images. Based on the existing graph convolutional network for semantic segmentation, two independent subtask branches are designed to obtain the cross-scale correlation between semantic features and the multi-scale characteristics of boundary features, and finally a cross-scale feature map containing semantic features and boundary features is obtained.

[0030] The cross-scale graph similarity guided aggregation system of this embodiment combines Figure 1 As shown in the figure, it includes a backbone network and two independent sub-task branches, namely the semantic feature extraction branch and the boundary feature extraction branch. Each part is introduced below.

[0031] The backbone network consists of two parts: one is to use conventional convolution operations to mine the rich semantic features of the original image X i ; First, we use the atrous convolution technique (ASPP) to change the receptive field of the convolution operation with different dilation rates, thereby mining the multi-scale semantic features of the original image F k , generating a multi-scale semantic feature map. This part is about how to obtain multi-scale features in semantic segmentation in the prior art, and will not be repeated here. The feature map mined by the backbone network is used as input to the subsequent two branch networks, where the semantic feature X i Input into the boundary feature extraction branch, and F k Input into the semantic feature extraction branch.

[0032] Combination Figure 2 As shown in the figure, the semantic feature extraction branch includes the cross-scale graph interaction module GCI and the graph convolutional network GCN. In the semantic feature extraction branch, the multi-scale semantic features F of the original image mined by the backbone network are used. k As input, it is input into the designed cross-scale graph interaction module GCI, and the cross-scale graph structure is established by constructing the relationship between graph nodes and edges of different scales. Finally, the graph convolutional network GCN is used to infer and aggregate the association relationship between cross-scale semantic features to enhance the representation ability of the model and extract the multi-scale feature G of the semantic feature. i .

[0033] The cross-scale graph interaction module CGI, for the multi-scale semantic feature map generated by the backbone network, is first integrated into a larger cross-scale feature map by feature splicing (here, the existing multi-scale features are spliced ​​into a feature map of a larger scale, for example, two feature maps of size 3X3 are spliced ​​to obtain a larger feature map of 6X3). Then, according to the conventional graph convolution operation, the cross-scale feature map is subjected to graph reasoning, that is, the construction of the edges and points of the graph, as well as the aggregation of information. Specifically: the spatial pixel points are converted into nodes in the graph model, and the similarity matrix is ​​calculated in the graph convolution operation, and the similarity matrix of the calculated node is used as the edge of the graph model. When calculating the correlation, the correlation score of the two nodes will be obtained. If the two nodes have a strong semantic relationship, a higher correlation score will be obtained, indicating that the two nodes are strongly correlated. Finally, the cross-scale reasoning is performed through the message passing mechanism in the graph convolution network GCN to aggregate information. Through the role of cross-scale graph reasoning, the semantic analysis interacts between multiple scales, thereby achieving the purpose of mining a wider range of semantic information.

[0034] Combination Figure 3 As shown in FIG. 1 , the boundary feature extraction branch includes a multi-scale similarity-guided aggregation module MSA, which includes a multi-scale boundary feature extraction MBFE unit and a similarity-guided aggregation SGA unit. The multi-scale similarity-guided aggregation module MSA converts the semantically rich feature X i As input, a preliminary feature fusion is first performed and supervised training is performed to obtain the boundary feature B containing boundary information; then, the boundary feature B is input into the MBFE unit. The MBFE unit uses dilated convolutions with different expansion rates to detect multi-scale boundary information and extract the multi-scale boundary feature B. i .

[0035] Similarity-guided aggregation SGA unit calculates boundary feature B i and the semantic feature G output by the semantic feature extraction branch i The similarity between them is calculated and a multiplication operation is performed to aggregate the two multi-scale features of semantic features and boundary features to improve the auxiliary effect of edge features on semantic segmentation.

[0036] Similarity-guided aggregation of SGA units on boundary features B i and the semantic feature G output by the semantic feature extraction branch i Perform similarity calculation, obtain the strongest similarity area through calculation, and use this as the weight matrix to perform weighted fusion on the original semantic features. In this way, the edge features are enhanced. This solves the problem of sample imbalance to a certain extent. The similarity calculation and weighted fusion method are described in detail in the method steps in the subsequent embodiment 3.

[0037] The dilated convolution of the MBFE unit with different dilation rates is an effective method to generate boundary features of different scales. Intuitively speaking, this strategy helps to align the generated boundary features with semantic features, thereby promoting subsequent aggregation. The boundary features mined by the boundary feature extraction branch are sparse matrices, which leads to the problem of sample imbalance. The present invention uses a similarity analysis method to search for edge regions with strong correlation to strengthen and guide the fusion of edge features. Using boundary features to guide semantic feature learning by calculating similarity is more convincing than the traditional semantic feature aggregation by adding or connecting elements. Therefore, a similarity-guided aggregation unit SGA is designed to solve the sparse boundary matrix problem in feature fusion. Inspired by the attention mechanism, the present invention calculates the similarity between semantic features and boundary features at the corresponding scale, thereby highlighting the effective boundary information of semantic features. According to the above calculation, the boundary regions of the same category will be given higher weights than other irrelevant regions. Since the boundary features are enhanced, the problem of intra-class consistency in semantic segmentation is solved to a certain extent, thereby improving the prediction accuracy of the model near the edge.

[0038] Example 2

[0039] This embodiment provides an application of a cross-scale graph similarity guided aggregation system for remote sensing image semantic segmentation. The composition and functions of the system refer to the contents of Embodiment 1, which will not be repeated here.

[0040] Example 3

[0041] This embodiment provides a semantic segmentation method, which is performed using the cross-scale graph similarity guided aggregation system as described in Embodiment 1. The specific method is as follows:

[0042] S1. First, the original image is input into the backbone network. The backbone network uses convolution operations to mine the rich semantic features of the original image X. i On the one hand, by setting different void rates to change the receptive field size of the convolution operation, the multi-scale semantic features F of the original image can be mined. k , generate multi-scale semantic feature maps;

[0043] S2. The multi-scale semantic feature graph generated by the backbone network is input into the cross-scale graph interaction module GCI, with spatial pixels as nodes of the cross-scale graph model, and the similarity matrix of the nodes is calculated as the edge of the graph model. Subsequently, the cross-scale reasoning is performed through the message passing mechanism in the graph convolutional network GCN to aggregate information. Through the role of cross-scale graph reasoning, semantic information interacts between multi-scale features, thereby achieving the purpose of mining a wider range of semantic information, and finally obtaining the cross-scale semantic feature G i .

[0044] More specifically, in step S2, for the multi-scale semantic features F generated by the backbone network k , the spatial pixel points of the feature are regarded as nodes and their sizes are converted to F k ∈R n×d , the cross-scale node set is Each node f encodes a different region in the original image. The values ​​of n and d are determined by the spatial and channel sizes of the multi-scale semantic features. The edges of the graph are defined as pairwise similarity calculations between image regions, and the relationship is constructed by the following equation: in and is a regular convolution whose parameters are learned by back-propagation, and They represent the i-th node at the p-th scale and the j-th node at the q-th scale respectively. The present invention can mine feature graphs at five scales, so the maximum values ​​of p and q are 5. As can be seen from the above equation, if the two calculation areas have a strong semantic relationship, a higher correlation score will be obtained. After constructing the cross-scale graph model of node F and edge R, cross-scale reasoning is performed through the message passing mechanism in GCN to aggregate information. Through the role of cross-scale graph reasoning, the model can contain broader and more diverse cross-scale semantic information.

[0045] S3. In the boundary feature extraction branch, the dilated convolution method is used to mine multi-scale boundary features. The semantic features X output by the backbone network i As input, a preliminary feature fusion is first performed, and supervised training is performed to obtain the boundary feature B containing boundary information; then, the boundary feature B is input into the MBFE unit. For the boundary feature B, the MBFE unit uses a hole convolution with different expansion rates to act on the boundary feature B. After feature mining, boundary features B of different scales are obtained. i , realizing multi-scale boundary feature mining; in order to facilitate the subsequent fusion with semantic features, the mined boundary feature B i The number and semantic features of G i Be consistent.

[0046] S4. The boundary features mined by the boundary feature extraction branch are sparse matrices, which leads to sample imbalance. The similarity analysis method is used to search for edge areas with strong correlation to strengthen and guide the fusion of edge features and semantic features. The present invention uses the similarity-guided aggregation SGA unit to calculate the boundary feature B i and the semantic feature G output by the semantic feature extraction branch i The similarity between them is calculated, and the strongest similarity area is obtained, which is used as the weight matrix to perform weighted fusion of the original semantic features.

[0047] We believe that using boundary features to guide semantic feature learning by calculating similarity is more convincing than the traditional semantic feature aggregation by adding or connecting elements. Therefore, a similarity-guided aggregation SGA unit is designed to solve the sparse boundary matrix problem in feature fusion. Inspired by the attention mechanism, the similarity between semantic features and boundary features is calculated at the corresponding scale to highlight the effective boundary information of semantic features. According to the above calculation, the boundary area of ​​the same category will be given a higher weight than other irrelevant areas. Since the edge features are enhanced, the problem of intra-class consistency in semantic segmentation is solved to a certain extent, thereby improving the prediction accuracy of the model near the edge. The edge features will perform similarity calculations with the semantic features, and the strongest similarity areas will be obtained by calculation, and the original semantic features will be weighted and fused using this as the weight matrix. In this way, the edge features are enhanced. This solves the problem of sample imbalance to a certain extent. Specifically, in step S4, the similarity-guided aggregation SGA unit performs similarity calculations in the following specific steps: Given a multi-scale semantic graph feature and boundary features Lowercase letters g and b represent the corresponding feature maps respectively. The subscript k represents the number of features currently. k = {1, 2, 3, 4, 5}. The superscript n represents the position of the feature map. First, the similarity between the two is calculated. in are two nonlinear transformations, and They represent the parameters of nonlinear transformation respectively, and the superscript T represents matrix transpose. The function is used to calculate the influence value between the jth position on the boundary and the ith position on the semantic map; then, matrix multiplication is performed between the multi-scale boundary features and the similarity matrix Among them, α is the parameter obtained by back propagation. According to the above calculation, the boundary area of ​​the same category will be activated with a much higher weight than other irrelevant areas, which not only ensures the intra-class consistency of the object, but also solves the problem of sparse edge pixels.

[0048] S5. The final output is a feature map that combines semantic and edge information.

[0049] In summary, the present invention firstly takes the original image as input and uses the backbone network to learn the multi-scale feature map. Subsequently, two independent sub-branches are used to obtain cross-scale semantic information and mine multi-scale boundary information. Specifically: the cross-scale graph interaction module (CGI) establishes a cross-scale graph structure by designing nodes and edges, and uses GCN to infer the aggregation of cross-scale features; the multi-scale similarity-guided aggregation module (MSA) is composed of a multi-scale boundary feature extraction unit (MBFE) and a similarity-guided aggregation unit (SGA). The MBFE unit uses a hole convolution with different expansion rates to detect multi-scale boundary information, and the SGA unit calculates the similarity between semantic features and boundary features, and performs multiplication operations to aggregate two multi-scale features. The CGSAN proposed in the present invention can not only meet the interaction between cross-scale targets, but also mine multi-scale boundary information and achieve robust aggregation, which greatly improves the representation ability of remote sensing features and better solves the problem of semantic segmentation of remote sensing images.

[0050] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Any changes, modifications, additions or substitutions made by ordinary technicians in this technical field within the essential scope of the present invention should fall within the protection scope of the present invention.

Claims

1. Cross-scale graph similarity guided aggregation system, It is characterized in that It includes a backbone network and two independent sub-task branches, namely the semantic feature extraction branch and the boundary feature extraction branch. In the semantic feature extraction branch, a cross-scale graph interaction module CGI is introduced to extract semantic features. In the boundary feature extraction branch, a multi-scale similarity-guided aggregation module MSA is introduced to extract multi-scale boundary features. The backbone network first uses conventional convolution operations to mine the rich semantic features X of the original image i ; Then, the dilated convolution method is used to change the receptive field of the convolution operation with different dilation rates to further explore the multi-scale semantic features of the original image F k , generate a multi-scale semantic feature map; finally, the feature map mined by the backbone network is used as input to enter the subsequent two branch networks, where the semantic feature X i Input into the boundary feature extraction branch, and F k Input into the semantic feature extraction branch; The semantic feature extraction branch includes a cross-scale graph interaction module GCI and a graph convolutional network GCN. In the semantic feature extraction branch, the multi-scale semantic features F of the original image mined by the backbone network are used. k As input, it is input into the cross-scale graph interaction module CGI, and the cross-scale graph model is established by constructing the relationship between graph nodes and edges of different scales. Finally, the graph convolutional network GCN is used to infer and aggregate the correlation between cross-scale semantic features to extract the cross-scale semantic features G of the semantic features. i ; The boundary feature extraction branch includes a multi-scale similarity guided aggregation module MSA, which includes a multi-scale boundary feature extraction MBFE unit and a similarity guided aggregation SGA unit. The multi-scale similarity guided aggregation module MSA extracts the rich semantic features X i As input, a preliminary feature fusion is first performed, and supervised training is performed to obtain the boundary feature B containing boundary information; then, the boundary feature B is input into the MBFE unit, and the MBFE unit uses the hole convolution with different expansion rates to detect multi-scale boundary information and extract the multi-scale boundary feature B i ; Similarity guides the aggregation SGA unit to calculate the boundary feature B i and the cross-scale semantic feature G output by the semantic feature extraction branch i The similarity between them is calculated and a multiplication operation is performed to aggregate the cross-scale semantic features G i and boundary feature B i , to improve the auxiliary effect of edge features on semantic segmentation, and finally output a feature map that combines semantic and edge information.

2. The cross-scale graph similarity guided aggregation system according to claim 1, It is characterized in that The cross-scale graph interaction module CGI first integrates the multi-scale semantic feature map generated by the backbone network into a cross-scale feature map through feature splicing, and then performs graph reasoning on the cross-scale feature map according to the graph convolution operation, converts the spatial pixels into nodes in the graph model, and uses the similarity matrix of the calculated nodes as the edge of the graph model; finally, cross-scale reasoning is performed through the message passing mechanism in the graph convolution network GCN to aggregate information, and through the role of cross-scale graph reasoning, semantic analysis interacts between multiple scales.

3. Application of the cross-scale graph similarity guided aggregation system according to claim 1, It is characterized in that Used for semantic segmentation of remote sensing images.

4. A method for semantic segmentation using the cross-scale graph similarity guided aggregation system according to claim 1, It is characterized in that The specific method is as follows: S1. First, the original image is input into the backbone network. The backbone network uses convolution operations to mine the rich semantic features of the original image X. i On the one hand, by setting different void rates to change the receptive field size of the convolution operation, the multi-scale semantic features F of the original image can be mined. k , generate multi-scale semantic feature maps; S2. For the multi-scale semantic feature maps generated by the backbone network, input them into the cross-scale graph interaction module GCI. Use spatial pixel points as the nodes of the cross-scale graph model, and simultaneously calculate the similarity matrix of the nodes as the edges of the graph model. Subsequently, perform cross-scale inference through the message passing mechanism in the graph convolutional network GCN to aggregate information. Through the effect of cross-scale graph inference, semantic information is interacted between multi-scale features, and finally the cross-scale semantic feature G is obtained. i ; S3. In the boundary feature extraction branch, the dilated convolution method is used to mine multi-scale boundary features. The semantic features X output by the backbone network i As input, a preliminary feature fusion is first performed, and supervised training is performed to obtain the boundary feature B containing boundary information; then, the boundary feature B is input into the MBFE unit. For the boundary feature B, the MBFE unit uses a hole convolution with different expansion rates to act on the boundary feature B. After feature mining, boundary features B of different scales are obtained. i , realizing multi-scale boundary feature mining; in order to facilitate the subsequent fusion with semantic features, the mined boundary feature B i The number and semantic features of G i Stay consistent; S4. The boundary features mined by the boundary feature extraction branch are sparse matrices, which leads to sample imbalance. The similarity-guided aggregation SGA unit is used to calculate the boundary feature B. i and the semantic feature G output by the semantic feature extraction branch i The similarity between them is calculated to obtain the strongest similarity area, which is used as the weight matrix to perform weighted fusion of the original semantic features; S5. Finally, the feature map that integrates semantic and edge information is output.

5. The method for semantic segmentation according to claim 4, It is characterized in that In step S2, for the multi-scale semantic features F generated by the backbone network k , the spatial pixel points of the feature are regarded as nodes and their sizes are converted to F k ∈R n×d , the cross-scale node set is Each node f encodes a different region in the original image. The values ​​of n and d are determined by the spatial and channel sizes of the multi-scale semantic features. The edges of the graph are defined as pairwise similarity calculations between image regions, and the relationship is constructed by the following equation: in and is a regular convolution whose parameters are learned by back-propagation, and They represent the i-th node at the p-th scale and the j-th node at the q-th scale respectively.

6. The method for semantic segmentation according to claim 5, It is characterized in that In step S4, the similarity-guided aggregation SGA unit performs similarity calculation in the following specific steps: given a multi-scale semantic graph feature and boundary features First, calculate the similarity between the two in are two nonlinear transformations, and They represent the parameters of nonlinear transformation, and the superscript T represents matrix transposition. The function is used to calculate the influence value between the jth position on the boundary and the ith position on the semantic map; then, matrix multiplication is performed between the multi-scale boundary features and the similarity matrix Among them, α is the parameter obtained by back propagation. According to the above calculation, the boundary area of ​​the same category will be activated with a higher weight than other irrelevant areas.

Citation Information

Patent Citations

  • Cross-view gait recognition method based on spatio-temporal information enhancement and multi-scale saliency feature extraction

    CN113947814A

  • Miao nationality costume image semantic segmentation method

    CN114037833A