A feature fusion method based on heterogeneous graph attention mechanism
Through the feature fusion method based on the heterogeneous map attention mechanism, the problem of simple and rough extraction of multispectral image features is solved, and more comprehensive and accurate feature representation is achieved, which improves the accuracy of multispectral image processing and the generalization ability of the model.
Patent Information
- Application Number
- CN202310575850.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-22
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2043-05-22
AI Technical Summary
The feature extraction methods of existing multispectral images are simple and rough, and cannot effectively express the complex information of the image, resulting in poor generalization capabilities of the model and limited feature dimensions, making it difficult to meet the needs of multiple visual tasks.
A feature fusion method based on the heterogeneous graph attention mechanism is adopted, and multi-spectral data features are extracted through manifold learning and empty spectrum embedding, and the feature maps are connected and aggregated in combination with the graph neural network and attention mechanism to form a heterogeneous graph with fused multi-dimensional dimensions.
It improves the accuracy and stability of feature extraction, reduces redundant information, enhances the complementarity of features, and improves the accuracy of multispectral image processing and the generalization ability of the model.
Smart Images

Figure CN116645579B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a feature fusion method based on a heterogeneous graph attention mechanism. Background Art
[0002] A heterogeneous graph is a graph composed of different types of nodes and edges. In a heterogeneous graph, each node and edge has its own specific type and attributes, and the connections between these nodes and edges may also be diverse. Heterogeneous graphs are commonly used in modeling many complex systems, such as social networks, protein interaction networks, and knowledge graphs. In these applications, node and edge types typically represent different entities or relationships, such as people, organizations, objects, and relationships. In short, a heterogeneous graph is a graph model characterized by diversity and complexity.
[0003] Multispectral images contain information from multiple bands, enabling the extraction of rich spatial and spectral information, overcoming the information limitations of single-modality images. Traditional feature extraction methods for multispectral images often produce relatively simple and crude features, failing to adequately represent complex image information. They are ineffective for nonlinear features, have limitations in expressing multi-scale features, suffer from poor model generalization, and have a limited dimensionality in the extracted features.
[0004] Therefore, how to effectively extract multispectral feature information to facilitate subsequent image classification, target detection and other visual tasks of multispectral images has become a problem that needs to be solved. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a feature fusion method based on the heterogeneous graph attention mechanism. The method can embed multispectral images from three dimensions to obtain a heterogeneous graph that fuses the features of multiple graphs. It can obtain a graph structure that fuses information from more dimensions and improve the effect of feature fusion.
[0006] The technical solution of the present invention to solve the above technical problems is:
[0007] A feature fusion method based on a heterogeneous graph attention mechanism includes the following steps:
[0008] S1. Acquire multispectral images and extract and reduce data features through manifold learning and spatial spectrum embedding methods;
[0009] S2. Extract the features of infrared spectrum waves from the reduced-dimensional multispectral data to obtain physical feature maps, obtain spatial feature maps through spatial superpixel segmentation technology, and construct spectral feature maps based on spectral similarity;
[0010] S3. Analyze the nodes and edges of the three obtained feature graphs, and use the edge and node linking method based on graph neural network to connect the three feature graphs of different types of nodes and edges;
[0011] S4. Use the attention-based graph pooling method to extract and aggregate information from the nodes of the new graph, and finally obtain a heterogeneous graph that integrates multi-source features.
[0012] Preferably, in step S1, the multispectral image is captured by a multispectral camera that can simultaneously capture three or more spectral bands.
[0013] Preferably, in step S2, the feature extraction and dimensionality reduction are a fusion form of the band array coding of the multispectral image represented by the augmented vector and the spatial pixel neighborhood coding of the multispectral image, and the fused data information is embedded in the space spectrum to complete the weight distribution of the similarity of the spectral features of different pixels, and the local neighborhood space and spectral information are similarly classified and the feature dimensionality is reduced by manifold learning.
[0014] Preferably, in step S2, the three feature maps are obtained by: using the spectral data after dimension reduction and combining it with the infrared spectral features to extract the physical feature map of the spectral data; using the linear iterative clustering method to determine the superpixel neighbor node information, constructing the edge connection relationship between the nodes based on the spatial connectivity relationship of the superpixels, and extracting the spatial feature map; combining the spectral feature similarity of the target, designing the adjacency matrix, sampling and recombining from different spectral band dimensions to obtain the target spectral feature distribution, and using the graph neural network to effectively represent the spectral data residing on the smooth manifold.
[0015] Preferably, in step S3, the linked network model is a graph autoencoder, wherein the graph autoencoder can adopt a graph convolutional autoencoder, a variational graph convolutional autoencoder, or an adversarial regularized graph autoencoder.
[0016] Preferably, in step S4, the graph pooling method may adopt DiffPool, SAGPool, or ASAP.
[0017] Through the above technical solutions, it is known that the present invention discloses a feature fusion method based on the heterogeneous graph attention mechanism, which has the following advantages compared with the prior art:
[0018] The method extracts feature maps from multiple dimensions, including spatial features, physical features, and spectral features. These features can effectively capture information of different dimensions and obtain a more comprehensive and accurate feature representation. This method can also reduce redundant information between feature maps and increase the complementarity between feature maps, thereby improving the accuracy and stability of multispectral image processing. By fusing heterogeneous feature maps, the noise and interference in the feature extraction process can be reduced, thereby improving the accuracy and stability of feature extraction. In addition, the use of heterogeneous feature fusion methods can make the model more generalizable and obtain better processing results. Heterogeneous graphs that fuse the features of multiple graphs can obtain a graph structure that fuses information of more dimensions, which can improve the effect of feature fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flowchart of a feature fusion method based on a heterogeneous graph attention mechanism of the present invention.
[0020] Figure 2 Schematic diagram of the process of obtaining a fused heterogeneous graph according to the present invention. DETAILED DESCRIPTION
[0021] The present invention will be described in further detail below with reference to the embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0022] See also Figure 1 , a feature fusion method based on heterogeneous graph attention mechanism of the present invention includes the following steps:
[0023] S1. Acquire multispectral images and extract and reduce data features through manifold learning and spatial spectrum embedding methods;
[0024] S2. Extract the features of infrared spectrum waves from the reduced-dimensional multispectral data to obtain physical feature maps, obtain spatial feature maps through spatial superpixel segmentation technology, and construct spectral feature maps based on spectral similarity;
[0025] S3. Analyze the nodes and edges of the three obtained feature graphs, and use the edge and node linking method based on graph neural network to connect the three feature graphs of different types of nodes and edges;
[0026] S4. Use the attention-based graph pooling method to extract and aggregate information from the nodes of the new graph, and finally obtain a heterogeneous graph that integrates multi-source features.
[0027] See also Figure 1 In step S1, the acquired multispectral data is composed of multispectral images of four bands.
[0028] In addition, the augmented vector of the fused spectral and spatial information can be expressed as follows:
[0029] x=(u,v,b1,b2,...,b B )=(x 1 ,x 2 ,...,x B+2 ) T (1)
[0030] Where h(u,v) is a pixel on the image grid plane, (b1,b2,b3,b B ) is the band array.
[0031] In the specific embodiment, we obtain images of 4 bands, so B=4.
[0032] In addition, the feature extraction and dimensionality reduction of the spatial spectrum information embedding and manifold learning are based on the augmented vector of L labeled pixels. As training data, after normalization, for any x i The same category classification is performed in the supervised mode, and the local neighborhood of pixels is constructed by the k-nearest neighbor algorithm. The manifold learning framework is further combined to encode the manifold local topology of the neighborhood data for feature dimensionality reduction.
[0033]
[0034] Among them, y i is x i The learned m-dimensional features, {W ij} is the input data and Di=∑ j W ij The positive weight of the similarity between the original x i and x j The feature similarity constraint between them can be obtained by y i and y j The Euclidean distance metric between the features is used to maintain
[0035] In a specific embodiment, the number of marked elements in the augmented vector is 6.
[0036] In addition, the weights of the local domain embedding based on the spatial spectrum polynomial can be calculated by Laplace embedding and locality preserving projection, that is:
[0037]
[0038] See also Figure 1 ,In step S2, the physical characteristic graph, including equivalent temperature, equivalent area physical characteristics, is represented as a graph through a random walk graph embedding method.
[0039] In addition, in the spatial feature map, the multispectral image is first segmented into superpixels using the SLIC algorithm. By calculating the spatial distance and spectral distance between pixel points and balancing the weights, the superpixel cluster center and range boundary are iteratively updated. The iteration is stopped when the error between the new cluster center and the old cluster center is less than a certain range, and a multispectral image data composed of superpixels is obtained. The edge connection relationship between the nodes is constructed based on the spatial connectivity relationship of the superpixels.
[0040] Specifically, the expressions of spatial distance and spectral distance are:
[0041]
[0042] Where, d c is the squared difference of the spectrum, d(S x ,S y ) is the spectral angular distance, d s The square of the distance is balanced by adjusting m, which is generally 50.
[0043] In a specific embodiment, the number of iterations is 15, and the error between the cluster center and the old cluster center is less than 0.01.
[0044] In addition, the spectral feature map is constructed by a semi-supervised adjacency matrix method; specifically, the method is constructed based on information provided by a limited number of labeled data and a large amount of unlabeled data, uses a Dirichlet process mixture model based on variational inference to construct pseudo labels, and implements spatial spectral adjacency matrix construction based on an intrinsic clustering algorithm in the data sample.
[0045] See also Figure 1 ,In step S3, the nodes and edges of the three obtained feature graphs are analyzed, and the network structure based on ,graph autoencoder is used to connect the feature graphs of three different types of nodes and edges.
[0046] Specifically, each given graph is analyzed, and the node feature vectors between different graphs are analyzed by cosine similarity, and the nodes with high similarity among the three graphs are retained.
[0047] In a specific example, it is set that the probability greater than 0.7 is retained and the probability less than 0.3 is discarded.
[0048] For the three processed graphs, we use the graph convolutional network to calculate them and get the node representation of each node. For node i, we extract its corresponding node representation z i The formula Z = GCN (X, A) is used to calculate the node representation matrix Z of the graph, where X is the node feature matrix, A is the adjacency matrix, and the i-th row of Z is the representation z of node i. i , that is, z i =Zi .
[0049] Then use the following formula:
[0050]
[0051] in It is the predicted probability between the link nodes (i, j), where σ is the Sigmoid activation function; here, the probability greater than 0.8 is set to be linked, and the probability less than 0.2 is not connected, and a new graph is obtained after linking the three graphs.
[0052] See also Figure 1 ,In step S4, the SAGpool method is used to extract and aggregate information ,from the nodes of the new graph.
[0053] Specifically, the extracted new graph is first passed through a convolution operation of the graph neural network, and GCN learns the feature representation of each node v∈V, that is, the features of the neighboring nodes of each node are aggregated to obtain the feature representation of node v.
[0054] For each node v, a self-attention mechanism is used to calculate an attention score z for each node. X is the characteristic matrix of the graph structure, D represents the degree matrix of the graph, A is the adjacency matrix of the graph, Θ att is the only parameter in the SAGPool layer, and σ is the tanh activation function. The attention score is calculated by considering the characteristics of the node itself and the characteristics of its neighboring nodes. The higher the score, the more important the node is in the current layer.
[0055] Then, we use idx = top-rank(Z, [kN]), where topk is the most important node. The number of retained nodes is determined by the pooling ratio k, and here we set k to 0.5. In this way, we obtain an attention-based mask map, and multiply the mask map with the graph structure of the original input fused heterogeneous information at the corresponding nodes to obtain the final output map, that is, the heterogeneous map that fuses multi-source features.
[0056] The above is a preferred embodiment of the present invention, but the embodiment of the present invention is not limited to the above content. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A feature fusion method based on heterogeneous graph attention mechanism, characterized in that: The following steps are involved: S1. Acquire multispectral images and extract and reduce data features through manifold learning and spatial spectrum embedding methods; S2. Extract the features of infrared spectrum waves from the reduced-dimensional multispectral data to obtain physical feature maps, obtain spatial feature maps through spatial superpixel segmentation technology, and construct spectral feature maps based on spectral similarity; S3. Analyze the nodes and edges of the three obtained feature graphs, and use the edge and node linking method based on graph neural network to connect the three feature graphs of different types of nodes and edges; S4. Use the attention-based graph pooling method to extract and aggregate information from the nodes of the new graph, and finally obtain a heterogeneous graph that integrates multi-source features.
2. A feature fusion method based on heterogeneous graph attention mechanism according to claim 1, characterized in that: In step S1 , the multispectral image is captured by a multispectral camera that can simultaneously capture three or more spectral bands.
3. The feature fusion method based on heterogeneous graph attention mechanism according to claim 1 is characterized in that: In step S2, the feature extraction and dimensionality reduction are a fusion of the band array encoding of the multispectral image represented by the augmented vector and the spatial pixel neighborhood encoding of the multispectral image. The fused data information is embedded in the spatial spectrum to complete the weight distribution of the similarity of the spectral features of different pixels, and the local neighborhood space and spectral information are similarly classified and the feature dimensionality is reduced through manifold learning.
4. The feature fusion method based on heterogeneous graph attention mechanism according to claim 1 is characterized in that: In step S2, the three feature maps are obtained by: using the spectral data after dimension reduction and combining it with the infrared spectral features to extract the physical feature map of the spectral data; using the linear iterative clustering method to determine the superpixel neighbor node information, constructing the edge connection relationship between the nodes based on the spatial connectivity relationship of the superpixels, and extracting the spatial feature map; combining the spectral feature similarity of the target, designing the adjacency matrix, sampling and recombining from different spectral band dimensions to obtain the target spectral feature distribution, and using the graph neural network to effectively represent the spectral data residing on the smooth manifold.
5. The feature fusion method based on heterogeneous graph attention mechanism according to claim 1 is characterized in that: In step S3, the linked network model is a graph autoencoder, wherein the graph autoencoder can adopt a graph convolutional autoencoder, a variational graph convolutional autoencoder, or an adversarial regularized graph autoencoder.
6. A feature fusion method based on heterogeneous graph attention mechanism according to claim 1, characterized in that: In step S4, the graph pooling method may adopt DiffPool, SAGPool, or ASAP.
Citation Information
Patent Citations
Implementation method for hyperspectral medicine component analysis based on graph convolutional neural network
CN113269196A
Infrared spectrum-based point target feature extraction method
CN116049641A