Network alignment method based on individual-environment double-view cross fusion
By adopting the individual-environment dual-view cross-fusion method in the network alignment method, combined with the graph attention network and hypergraph convolutional network, the problem of insufficient network alignment efficiency and robustness in the prior art is solved, and a more efficient and accurate network alignment effect is achieved.
Patent Information
- Application Number
- CN202411973075.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-16
AI Technical Summary
Existing network alignment methods have challenges in processing efficiency and insufficient data volume, and lack of comprehensive modeling of node individual and environmental levels, resulting in insufficient robustness to data noise.
The network alignment method based on individual-environment dual-view cross-fusion is adopted. Through steps such as feature extraction, dual-view embedding modeling, graph embedding cross-fusion and similarity matrix calculation, combined with graph attention network and hypergraph convolution network, the node individual embedding and environment embedding are combined to calculate the similarity matrix of the source network and the target network.
It improves the efficiency and accuracy of network alignment, enhances the robustness of data noise, can more comprehensively model the interaction between individual nodes and environment, and improves the performance of network alignment models.
Smart Images

Figure CN120011825A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network alignment, and in particular to a network alignment method based on cross-fusion of individual-environment dual views. Background Art
[0002] A network, also known as a graph, is a collection of nodes connected by edges. Different networks are interconnected, and nodes belonging to the same entity can appear in multiple networks. For example, on social media, the same user often has accounts in different social networks. These interconnections enable complex relationships and overlaps that can occur between different networks. The task of finding nodes with the same identity in different networks is called network alignment. Nodes belonging to the same individual in different networks are called anchor nodes. Network alignment is fundamental to various downstream tasks such as cross-social network recommendations, protein interaction analysis, and point cloud alignment.
[0003] Network alignment techniques encounter challenges such as efficiency issues and insufficient data. Most existing work frames network alignment as a supervised learning task. Although these methods have achieved good results in terms of accuracy and efficiency, obtaining prior anchor nodes across different networks is a challenging and time-consuming process in the real world. Matrix decomposition-based methods formulate the network alignment problem as a bipartite graph matching problem. These methods often have high time complexity and cannot fully utilize the attribute information of network nodes.
[0004] In recent years, many studies have turned to network alignment methods based on unsupervised graph representation learning. Based on the neighborhood modeling ability of graph neural networks, such methods take the attribute features or structural features of nodes as input and learn the embedding representation of each node. Finally, the network alignment problem is transformed into a vector similarity calculation problem. GAlign and Grad-Align models take node attribute features as input, which also makes them sensitive to attribute noise. For example, if the profile information provided by the same user in two social networks is significantly different, such as her / his birthday and location. Even if the two nodes have similar topological structures, attribute-based methods are likely to classify them as different users. CONE and HackGAN models take node structural features as input and usually do not perform well on real datasets. This is because the two networks may not have topological consistency and the same node may have very different local structures in different networks. For example, a person who is actively seeking a job may have more connections with other users on LinkedIn than Douban. In order to reduce the impact of attribute noise and structural noise on model performance, the SANA model comprehensively considers attribute features and structural features when calculating node similarity, and proposes a graph attention model with shared parameters to alleviate the inconsistency of embeddings of different networks.
[0005] Although the above methods have achieved good modeling of node attributes and structures based on graph neural models, they focus on modeling at the individual node level, lack richer user representation modeling, and are still insufficient in robustness to data noise. Specifically, graph neural models such as graph convolution and graph attention rely on the neighbor nodes of the current node to update the node representation, only using the node neighborhood information and focusing on the information interaction between individual nodes. The noise of neighbor nodes will directly affect the representation of the current node. In addition to the similarity at the individual level, the similarity of the environment in which the nodes are located is also an important factor in evaluating whether two nodes are aligned. The node environment is a more macroscopic view that reflects the direct and indirect connections between nodes. For example, a user's direct and indirect friends on a social network can be considered as the user's social environment. Modeling user representation from an environmental perspective can reduce the model's sensitivity to data noise to a certain extent. However, existing methods have ignored the interaction between users and the environment, making it difficult to discover high-order connections in the network, affecting the performance of the network alignment model. Summary of the invention
[0006] The technical problem to be solved by the present invention is that, in view of the above-mentioned problems in the prior art, the present invention provides a network alignment method based on the cross-fusion of individual-environment dual views, which has simple principle, convenient operation and high processing efficiency.
[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0008] A network alignment method based on cross-fusion of individual-environment dual views, comprising:
[0009] Step S1: Feature extraction: Obtain the attribute features and structural features of the nodes in the graph as the input of the post-graph embedding model;
[0010] Step S2: Dual-view embedding modeling: Using attribute features or structural features as input, model the embedding representation of the graph network from two perspectives, namely, the individual and the environment.
[0011] Step S3: Graph embedding cross fusion: Based on multi-head cross attention, the node individual embedding and the environment embedding are fused;
[0012] Step S4: Similarity matrix calculation and refinement: Based on the individual embedding and environment embedding output by each layer of the graph embedding model, as well as the graph embedding representation after cross-fusion, the similarity matrix of the source network and the target network is jointly calculated.
[0013] As a further improvement of the present invention: in step S1, the node attribute feature It is a binary matrix, which is directly used as input without extraction; the node structure feature is calculated based on the degree distribution of the node's k-hop neighbors, assuming is calculated from the k-hop neighbors of node u, then Express satisfaction The number of k-order neighbors v, the final structural characteristics of node u are calculated as follows:
[0014]
[0015] where β k-1 represents the weight of the k-1 order neighbor, and D is the maximum order specified.
[0016] As a further improvement of the present invention: in step S2, the modeling process includes:
[0017] Given a graph in represents the node set, ε represents the edge set, Represents a set of node attributes. The graph representation learning based on GAT is expressed as the following formula:
[0018] H l+1 =GAT(H l )=σ(C l H l W l )
[0019] in, is the embedding representation matrix, d l represents the embedding representation dimension of the lth layer, σ(.) represents the activation function, and W l is the weight matrix, C l is the attention matrix calculated based on the attention mechanism, Represents the attention weight of node i and its neighbor node j, which is calculated by the following formula:
[0020]
[0021] N i represents all neighbors of node i, where The calculation formula is as follows:
[0022]
[0023] The embedding representation of node i is calculated from its neighbors according to the following formula:
[0024]
[0025] As a further improvement of the present invention: the node embedding based on GAT is highly correlated with the local structure of the network to construct a reconstruction loss function, the goal is to minimize the difference between the node embedding and the adjacency matrix, the specific loss function is as follows:
[0026]
[0027] in represents the adjacency matrix with self-loops added, is a diagonal matrix, where *∈{X,Y} represents the source network or the target network, h i or j It represents the output of the last layer of the multi-layer GAT model.
[0028] As a further improvement of the present invention: the step S4 comprises:
[0029] Step S41: node to hyperedge aggregation;
[0030] Given a hypergraph G hyper , and by aggregating hyperedge e j All connected nodes u on i The embedding representation h_e i , let's learn about hyperedge e j The expression o j , of the form:
[0031]
[0032] Where σ(.) is a ReLU activation function, is a trainable weight matrix, d l is the dimension of the embedding representation at layer l;
[0033] Step S42: Hyperedge to node aggregation;
[0034] After obtaining the representation of the hyperedge, for node u i All participating super edges Perform information aggregation and obtain the updated node representation h_e i , the calculation process is as follows:
[0035]
[0036] in is a trainable weight matrix;
[0037] Step S43: Hyperedge update;
[0038] A hyperedge update module is designed to update the hyperedge representation using the node context-aware embedding representation updated in step S42. The calculation process is as follows:
[0039]
[0040] in is a trainable weight matrix.
[0041] As a further improvement of the present invention: for the hypergraph reconstruction loss, the following loss function is adopted:
[0042]
[0043] in represents the hypergraph matrix constructed by the adjacency matrix of the original network, *∈{X,Y} represents the source network or the target network, h_e i and j ′ It represents the output of the last layer of the multi-layer HGCN model.
[0044] As a further improvement of the present invention: for the noisy hypergraph versions of the source network and the target network and The following contrastive loss function is used to minimize the distance between the context-aware embedding representations of nodes in the original hypergraph network and the noisy hypergraph network:
[0045]
[0046] As a further improvement of the present invention, the hypergraph reconstruction loss function and the contrast loss function are added according to certain weights to obtain the following learning loss of the environment-aware embedding representation for the HGCN model:
[0047]
[0048] Where λ is a hyperparameter used to balance the reconstruction loss and the contrast loss, minimizing and A multi-layer HGCN model with shared parameters is trained for modeling node context-aware embedding representation.
[0049] As a further improvement of the present invention: in the step S3, the individual embedding representation of the node and the environment-aware embedding representation are further fused in the form of cross attention;
[0050] Given the individual representation h and the environment perception representation h_e, the cross-attention based fusion calculation is as follows:
[0051]
[0052] in and W o are all trainable matrices, d ′ = d / H, d is the dimension of the model output vector, H represents the number of heads of multi-head attention, and h i i and h_e i ′They represent the individual embedding representation and environment-aware embedding representation of node i after cross-fusion respectively.
[0053] Compared with the prior art, the advantages of the present invention are:
[0054] 1. The network alignment method based on the cross-fusion of individual-environment dual views of the present invention has a simple principle, convenient operation and high processing efficiency. On the one hand, the present invention designs a multi-layer graph attention network (GATs), which takes node attributes and structural information as input, dynamically assigns different weights to different neighbor nodes, so as to more flexibly model the interaction between individual nodes; on the other hand, the present invention designs a multi-layer hypergraph convolutional network, and the hypergraph allows a hyperedge to connect multiple nodes. Through the information aggregation calculation from node to hyperedge and hyperedge to node, it can simulate the interaction between users and the environment in the real network, thereby effectively modeling the node environment perception representation.
[0055] 2. The network alignment method based on the cross-fusion of individual-environment dual views of the present invention is to further explore and utilize the complementary information of node individual representation and environmental perception representation, and proposes a dual-view representation fusion method based on the cross-attention mechanism. Different from self-attention, cross-attention achieves a high-order fusion of the two types of information by exchanging the query vectors of individual representation and environmental perception representation. Finally, taking into account the limited ability of graph neural networks to represent network structures, the present invention uses the adjacency matrices of the source network and the target network to iteratively refine the node similarity matrix to improve the alignment performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a schematic diagram of the principle of the present invention in a specific embodiment.
[0057] Figure 2 It is a schematic diagram of the data noise robustness test results in a specific embodiment of the present invention.
[0058] Figure 3 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0059] The present invention is further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.
[0060] like Figure 1 and Figure 3 As shown, the network alignment method based on cross-fusion of individual-environment dual views of the present invention includes:
[0061] Step S1: feature extraction;
[0062] The attribute features and structural features of the nodes in the graph are obtained as the input of the subsequent graph embedding model.
[0063] Node attribute characteristics is a binary matrix that is taken directly as input without extraction.
[0064] The structural characteristics of a node are calculated based on the degree distribution of its k-hop neighbors, assuming is calculated from the k-hop neighbors of node u, then Express satisfaction The number of k-order neighbors v, the final structural characteristics of node u are calculated as follows:
[0065]
[0066] where β k-1 represents the weight of the k-1 order neighbor, and D is the maximum order specified.
[0067] Step S2: dual-view embedding modeling;
[0068] Taking attribute features or structural features as input, the embedded representation of the graph network is modeled from two perspectives: individual and environment.
[0069] Specifically, on the one hand, the mutual influence between individual nodes is modeled based on the graph attention network to obtain the individual embedding of the nodes; on the other hand, based on the hypergraph convolutional network, all the first-order neighbors of the node are used as hyperedges to model the interaction between the node environment and obtain the node environment embedding. At the same time, considering the graph reconstruction loss and the noise graph contrast loss, the self-supervised learning of the graph embedding network is realized.
[0070] Step S3: graph embedding cross fusion;
[0071] Based on multi-head cross attention, the node individual embedding and environment embedding are further fused.
[0072] Cross-attention can mine high-order interaction information between individual embeddings and environment embeddings, thereby judging node similarity from a more comprehensive perspective.
[0073] In addition, the noise map contrast loss is also used to achieve self-supervised learning of multi-head cross-attention networks.
[0074] Step S4: similarity matrix calculation and refinement;
[0075] Based on the individual embedding and environment embedding output by each layer of the graph embedding model, as well as the graph embedding representation after cross-fusion, the similarity matrix of the source network and the target network is jointly calculated.
[0076] Since the ability of graph neural networks to represent network structures is limited, they cannot fully utilize the similarity of network structures. Therefore, in this stage, the adjacency matrix of the source network and the target network is used to topologically refine the similarity matrix to improve the accuracy of network alignment.
[0077] In a specific application example, in step S2, the graph embedding modeling process based on the individual-environment dual view is:
[0078] Given a graph in represents the node set, ε represents the edge set, Represents a set of node attributes. The graph representation learning based on GAT is expressed as the following formula:
[0079] H l+1 =GAT(H l )=σ(C l H l W l )
[0080] in, is the embedding representation matrix, d l represents the embedding representation dimension of the lth layer, σ(.) represents the activation function, and W l is the weight matrix, C l is the attention matrix calculated based on the attention mechanism, Represents the attention weight of node i and its neighbor node j, which is calculated by the following formula:
[0081]
[0082] N i represents all neighbors of node i, where The calculation formula is as follows:
[0083]
[0084] Therefore, the embedding representation of node i is calculated from its neighbors as follows:
[0085]
[0086] As a preferred embodiment, the present invention can also construct a multi-layer GAT model to model the embedding of all nodes in the source network and the target network. However, since GAT is designed for a single network, and the semantic information between two real networks is usually inconsistent, if a GAT model is applied to each of the source and target networks, the embedding space between the networks may be different, thereby failing to achieve effective alignment. Therefore, the method of the present invention trains a GAT model with shared parameter weights to maintain the consistency of the topological structure and attributes between the potentially aligned nodes in the two networks.
[0087] If G X and G Y are two isomorphic graphs with the same structure. If any aligned node pair (u, v) is selected, their input vectors are the same, and after being calculated by the multi-layer GAT with shared parameter weights, the output results of each layer of the GAT are also the same. Although the source network and the target network in the real world are not isomorphic graphs, if the aligned nodes can show similarity in attributes and structure, then after being calculated by the multi-layer GAT model with shared weights, their embedded representations are also similar to each other.
[0088] In specific application examples, the multi-layer GAT model can also model node embedding representation without designing a specific training objective function, but the untrained model representation has poor learning ability and it is difficult to ensure a good network alignment effect. Since the node embedding based on GAT is highly related to the local structure of the network, a reconstruction loss function can be constructed, the goal is to minimize the difference between the node embedding and the adjacency matrix. The specific loss function is as follows:
[0089]
[0090] in represents the adjacency matrix with self-loops added, is a diagonal matrix, where *∈{X,Y} represents the source network or the target network. i or j It represents the output of the last layer of the multi-layer GAT model.
[0091] In a specific application example, in step S3, the individual embedding representation of the node and the environment-aware embedding representation are further fused in the form of cross attention.
[0092] Given the individual representation h and the environment perception representation h_e, the cross-attention based fusion calculation is as follows:
[0093]
[0094] in and W O are all trainable matrices, d ′ = d / H, d is the dimension of the model output vector, H represents the number of heads of multi-head attention, and h i ′ and h_E i ′ They represent the individual embedding representation and environment-aware embedding representation of node I after cross-fusion respectively.
[0095] In a specific application example, in step S4, a hypergraph convolution method is adopted, and the process includes:
[0096] Step S41: Node to hyperedge aggregation: Given a hypergraph g hyper , and by aggregating hyperedges E j All connected nodes u on i The embedding representation h_E i , let's learn about hyperedge E j Representation of O j , of the form:
[0097]
[0098] Where σ(.) is a ReLU activation function, is a trainable weight matrix, d l is the dimension of the embedding representation at layer l;
[0099] Step S42: Hyperedge to node aggregation: After obtaining the hyperedge representation, for node u i All participating super edges Perform information aggregation and obtain the updated node representation h_e i , the calculation process is as follows:
[0100]
[0101] in is a trainable weight matrix;
[0102] Step S43: Hyperedge update: In order to further capture the cross-information interaction between environments, a hyperedge update module is designed to update the hyperedge representation using the node environment-aware embedding representation updated in step 2. The calculation process is as follows:
[0103]
[0104] in is a trainable weight matrix.
[0105] For the hypergraph reconstruction loss, since the node representation and the environment representation interact with each other during the hypergraph convolution process, the present invention further adopts the loss function:
[0106]
[0107] in represents the hypergraph matrix constructed by the adjacency matrix of the original network, *∈{X,Y} represents the source network or the target network, h_e i and j ′It represents the output of the last layer of the multi-layer HGCN model.
[0108] Noisy hypergraph versions of the source and target networks and The following contrastive loss function is designed to minimize the distance between the context-aware embedding representations of nodes in the original hypergraph network and the noisy hypergraph network:
[0109]
[0110]
[0111] The hypergraph reconstruction loss function and the contrast loss function are added according to certain weights to obtain the following environment-aware embedding representation learning loss for the HGCN model:
[0112]
[0113] Where λ is a hyperparameter used to balance the reconstruction loss and the contrast loss, minimizing and A multi-layer HGCN model with shared parameters is trained for modeling node context-aware embedding representation.
[0114] Combination Figure 1 As shown, based on the graph attention model and the hypergraph convolution model, the individual embedding representation h of the node and the environmental perception embedding representation h_e can be modeled from the two perspectives of the individual and the environment. However, the training of the graph attention model and the hypergraph convolution model is isolated from each other, and there is no deep fusion between the individual representation and the environmental representation, which limits the consistency of the node embedding representation. In order to enhance the interaction between individual and environmental information and provide a more comprehensive perspective for node similarity judgment, the present invention further adopts the form of cross attention to further fuse the individual embedding representation and the environmental perception embedding representation of the node.
[0115] Given the individual representation h and the environment perception representation h_e, the cross-attention based fusion calculation is as follows:
[0116]
[0117] in and W O are all trainable matrices, d ′ =d / H, d is the dimension of the model output vector, and H represents the number of heads of multi-head attention. i i and h_e i ′They represent the individual embedding representation and environment-aware embedding representation of node i after cross-fusion. Different from self-attention, cross-attention can effectively explore the complementary information and high-order interactions between different vectors by changing the query vector, thereby achieving deep fusion of vectors.
[0118] In order to ensure that the original individual representation and environmental perception information are not affected, residual connections are introduced and normalized:
[0119]
[0120] In order to realize the training of the cross-attention network, it is also necessary to design a loss function for training. Since the node representation after cross-fusion is significantly different from the original representation, the graph reconstruction loss may not be applicable. However, there is still a one-to-one correspondence between the cross-fusion representation of the original network and the cross-fusion representation of the noise network. Therefore, the present invention uses a contrast loss function to perform self-supervised training on the cross-attention network:
[0121]
[0122] where λ is a hyperparameter used to balance the individual and context-aware contrast losses.
[0123] In specific applications, the noise resistance of the proposed method and the baseline method are compared, and experiments are also conducted on Douban and PPI datasets. Structural noise is simulated by randomly removing edges, and attribute noise is simulated by randomly selecting nodes with a certain probability and setting their node attributes to 0. Figure 2 The sensitivity of this method and other three baseline models to data noise on two data sets is shown, where the data noise range is 0.1 to 0.5. Data noise refers to the noise imposed by both structure and attribute. The method of the present invention achieves the best performance overall, and has a stronger noise resistance than other methods, which comes from the modeling of node environment perception features. When the noise level is small, both SLOTAlign and Galign models rely solely on attribute features. As the data noise increases, the performance of the two models decreases significantly.
[0124] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A network alignment method based on cross-fusion of individual-environment dual views, characterized in that: include: Step S1: Feature extraction: Obtain the attribute features and structural features of the nodes in the graph as the input of the post-graph embedding model; Step S2: Dual-view embedding modeling: Using attribute features or structural features as input, model the embedding representation of the graph network from two perspectives, namely, the individual and the environment. Step S3: Graph embedding cross fusion: Based on multi-head cross attention, the node individual embedding and the environment embedding are fused; Step S4: Similarity matrix calculation and refinement: Based on the individual embedding and environment embedding output by each layer of the graph embedding model, as well as the graph embedding representation after cross-fusion, the similarity matrix of the source network and the target network is jointly calculated.
2. The network alignment method based on cross-fusion of individual-environment dual views according to claim 1 is characterized in that: In step S1, the node attribute feature It is a binary matrix, which is directly used as input without extraction; the node structure feature is calculated based on the degree distribution of the node's k-hop neighbors, assuming is calculated from the k-hop neighbors of node u, then Express satisfaction The number of k-order neighbors v, the final structural characteristics of node u are calculated as follows: where β k-1 represents the weight of the k-1 order neighbor, and D is the maximum order specified.
3. The network alignment method based on cross-fusion of individual-environment dual views according to claim 1 is characterized in that: In step S2, the modeling process includes: Given a graph in represents the node set, ε represents the edge set, Represents a set of node attributes. The graph representation learning based on GAT is expressed as the following formula: H l+1 =GAT(H l )=σ(C l H l W l ) in, is the embedding representation matrix, d l represents the embedding representation dimension of the lth layer, σ(.) represents the activation function, and W l is the weight matrix, C l is the attention matrix calculated based on the attention mechanism, Represents the attention weight of node i and its neighbor node j, which is calculated by the following formula: N i represents all neighbors of node i, where The calculation formula is as follows: The embedding representation of node i is calculated from its neighbors according to the following formula:
4. The network alignment method based on cross-fusion of individual-environment dual views according to claim 3 is characterized in that: The node embedding based on GAT is highly correlated with the local structure of the network to construct a reconstruction loss function. The goal is to minimize the difference between the node embedding and the adjacency matrix. The specific loss function is as follows: in represents the adjacency matrix with self-loops added, is a diagonal matrix, where *∈{X,Y} represents the source network or the target network, h i or j It represents the output of the last layer of the multi-layer GAT model.
5. The network alignment method based on cross-fusion of individual-environment dual views according to any one of claims 1 to 4, characterized in that: The step S4 comprises: Step S41: node to hyperedge aggregation; Given a hypergraph G hyper , and by aggregating hyperedge e j All connected nodes u on i The embedding representation h_e i , let's learn about hyperedge e j The expression o j , of the form: Where σ(.) is a ReLU activation function, is a trainable weight matrix, d l is the dimension of the embedding representation at layer l; Step S42: Hyperedge to node aggregation; After obtaining the representation of the hyperedge, for node u i All participating super edges Perform information aggregation and obtain the updated node representation h_e i , the calculation process is as follows: in is a trainable weight matrix; Step S43: Hyperedge update; A hyperedge update module is designed to update the hyperedge representation using the node context-aware embedding representation updated in step S42. The calculation process is as follows: in is a trainable weight matrix.
6. The network alignment method based on cross-fusion of individual-environment dual views according to claim 5 is characterized in that: For the hypergraph reconstruction loss, the following loss function is used: in represents the hypergraph matrix constructed by the adjacency matrix of the original network, *∈{X,Y} represents the source network or the target network, h_e i and j ′ It represents the output of the last layer of the multi-layer HGCN model.
7. The network alignment method based on cross-fusion of individual-environment dual views according to claim 6 is characterized in that: Noisy hypergraph versions of the source and target networks and The following contrastive loss function is used to minimize the distance between the context-aware embedding representations of nodes in the original hypergraph network and the noisy hypergraph network:
8. The network alignment method based on cross-fusion of individual-environment dual views according to claim 7 is characterized in that: The hypergraph reconstruction loss function and the contrast loss function are added according to certain weights to obtain the following environment-aware embedding representation learning loss for the HGCN model: Where λ is a hyperparameter used to balance the reconstruction loss and the contrast loss, minimizing and A multi-layer HGCN model with shared parameters is trained for modeling node context-aware embedding representation.
9. The network alignment method based on cross-fusion of individual-environment dual views according to any one of claims 1 to 4, characterized in that: In step S3, the individual embedding representation of the node and the environment-aware embedding representation are further fused in the form of cross attention; Given the individual representation h and the environment perception representation h_e, the cross-attention based fusion calculation is as follows: Where W j Q , W j K , W j V and W O are all trainable matrices, d ′ = d / H, d is the dimension of the model output vector, H represents the number of heads of multi-head attention, and h i ′ and h_e i ′ They represent the individual embedding representation and environment-aware embedding representation of node i after cross-fusion respectively.