Distributed external node detection method based on energy evaluation and propagation
By introducing energy evaluation and propagation mechanisms into the graph neural network, combining structure perception and joint alignment regularization, the problem of degradation of model effect when new nodes are added in the graph data and misjudgment of off-distribution data detection is solved, and more efficient and accurate off-distribution node detection is achieved.
Patent Information
- Application Number
- CN202510220962.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The existing graph neural network method has a reduced effect when facing the addition of new nodes in graph data, and the out-of-distribution detection method based on softmax confidence scores may lead to overconfidence posterior distribution problem of out-of-distribution data.
A method of detection of external nodes based on energy evaluation and propagation is proposed. By constructing a GNN encoder to generate node embedding, energy fractions are allocated to nodes using energy functions, and structure-aware energy propagation modules and joint alignment regularization are introduced to enhance the generalization ability and detection accuracy of the model.
Effectively identify the external nodes in the graph, enhance the generalization ability of external data, reduce misjudgment, and improve the robustness of the model and detection accuracy.
Smart Images

Figure CN120068929A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning. More precisely, the present invention designs an out-of-distribution node detection method based on energy evaluation and propagation. Background Art
[0002] As an important research direction in deep learning, out-of-distribution detection mainly aims to identify input samples that do not belong to the distribution seen during model training. Since deep learning models usually lack sufficient generalization ability for data outside the training distribution, accurately detecting and distinguishing out-of-distribution samples is crucial for improving the reliability and security of the model.
[0003] Current out-of-distribution detection methods mainly focus on several directions: probability-based detection, discriminative method-based detection, and generative model-based detection. Probability-based detection methods usually use the output probability of the model to judge whether the input is an out-of-distribution sample. When assuming that the training data and test data come from the same distribution, the prediction probability of out-of-distribution samples is often low. Discriminative methods use different discriminative networks to classify input samples, and try to judge whether the input belongs to a known class or an out-of-distribution class by constructing a classifier. In recent years, out-of-distribution detection methods based on generative models have gradually received attention, especially using generative models such as generative adversarial networks or variational autoencoders to capture the latent distribution characteristics of input data. By training the generated samples of the generative model in the target domain, it is possible to judge whether it is an out-of-distribution sample by calculating the difference between the actual sample and the generated sample. Generative models can more flexibly learn the latent structure of data and better adapt to complex out-of-distribution detection tasks.
[0004] Although existing methods have achieved varying degrees of success in different tasks, out-of-distribution detection still faces many challenges. How to design a more efficient and accurate model to reduce the misjudgment of out-of-distribution samples and enhance the robustness of the model remains an urgent problem to be solved. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that in a real-world scenario, graph data usually continues to expand with the continuous acquisition of external knowledge, which means that new nodes of unknown classes may be added to the graph. The difference between the new nodes and the original node distribution may lead to a decline in the effectiveness of existing graph neural network (GNN) methods. At the same time, existing out-of-distribution detection methods based on softmax confidence scores may lead to the problem of overconfident posterior distributions for out-of-distribution data. To solve these problems, the present invention proposes an out-of-distribution node detection method based on energy evaluation and propagation, which endows the GNN with the ability to identify out-of-distribution nodes and improves its generalization ability for out-of-distribution data.
[0006] To achieve the above goals, the technical solution of the out-of-distribution node detection method based on energy evaluation and propagation proposed by the present invention includes the following steps:
[0007] Step 1: Construct a GNN encoder to encode the local neighborhood information of each node into a high-dimensional embedding. This process enables the embedding of each node to not only contain its own features but also reflect the dependence relationship between the node and its surrounding neighbor nodes. In this way, the node embedding can fully capture the complexity of the graph structure while retaining the semantic and structural information of each node in the graph, providing a solid foundation for subsequent tasks.
[0008] (1) First, select a part of the node classes from the graph G as the source node set V S , and these nodes constitute the main data distribution encountered by the model during the training phase; the remaining node classes are used as the out-of-distribution node set V T , representing the test environment different from the training distribution, thus simulating the out-of-distribution generalization scenario. The data partitioning process is as follows:
[0009]
[0010] Among them, V represents the node set of the graph G, V S and V T represent the selected source node set and out-of-distribution node set respectively, C S and C T represent the source node class set and out-of-distribution node class set respectively. Assign the meta-label "source node" to each node in V S , and assign the meta-label "out-of-distribution node" to each node in V T .
[0011] (2) Construct a GNN embedding layer to calculate the embedding representation of the node by aggregating the local neighborhood structure information, and construct the graph representation learning backbone network f θ (·) through multiple GNN embedding layers. The forward propagation process of node v at the k-th layer can be defined as follows:
[0012]
[0013] Among them, represents the embedding of node v at the k-th layer, θ is the learnable parameter, σ is the non-linear activation function (such as Sigmoid), represents the set of neighbor nodes of node v, and AGG is the aggregation function.
[0014] (3) Input the source node set V S and the out-of-distribution node set V T into f θ(·), we get the node embedding matrix. The output of the kth layer can be defined as:
[0015]
[0016] in, The graph represents the kth layer of the learning skeleton, G S and G T Respectively, through V S and V T The sampled subgraph, and Respectively represent G S and G T The node embedding matrix of , their meta-labels are “source node” and “out-of-distribution node” respectively.
[0017] (4) and Splice to get the final output This is the final embedding result of the kth layer.
[0018] Step 2: Use an energy function to generate a corresponding energy score for each node. These scores are consistent with the probability density of the node. The energy value represents the degree of match between the node and its distribution. Nodes with higher energy values are considered to be from outside the distribution. By assigning these energy scores to nodes, the model can effectively distinguish nodes from different distributions in the graph, and then identify out-of-distribution node instances.
[0019] (1) Using the energy function from f θ Extract the energy value of the node and map the node embedding to a scalar to indicate whether the node is an out-of-distribution node. The function is defined as follows:
[0020]
[0021] Among them, v i is the input node, For v i The r-ego subgraph centered at f θ is the learning skeleton for graph representation, C is the category label of the source node set, and T is the temperature coefficient.
[0022] (2) When the energy value of a node exceeds a certain threshold, the node is judged as an out-of-distribution node because it exhibits higher uncertainty and uniformity of category prediction distribution.
[0023] Step 3: Introduce a structure-aware energy propagation module so that the energy values of nodes can be dynamically adjusted according to changes in the graph structure. As the training progresses, the model will automatically optimize the energy value propagation process of nodes, making the energy values more sensitive to the position and structural changes of nodes in the graph. In this way, the model can not only better adapt to changes in the graph structure but also further improve the accuracy and robustness in detecting out-of-distribution nodes.
[0024] (1) To enhance the energy difference between nodes with different distributions, aggregate the energy values of neighbor nodes to the target node. In this process, a reasonable method is to assign higher weights to neighbor nodes with the same distribution as the target node; on the contrary, neighbor nodes with other distributions have less impact on energy aggregation and should therefore be assigned lower weights. In this way, the energy values can better reflect the distribution differences between nodes.
[0025] (2) Use edge energy scores to quantify the weights of neighbor nodes, which are defined as follows:
[0026]
[0027] where, represents the edge energy score between adjacent nodes v i and v j , and are the energy scores of nodes v i and v j at the l-th layer, respectively.
[0028] (3) Aggregate the node energy scores and edge energy scores to update the energy score of the target node. The update process is defined as follows:
[0029]
[0030] where, is the energy score of node v i i, is the set of neighbor nodes of node v i .
[0031] Step 4: Introduce joint alignment regularization to prompt the model to capture more generalizable features under different data distributions. This regularization method helps the model learn across distributions and ensures that it can better identify and adapt to out-of-distribution nodes, thereby improving the overall out-of-distribution detection performance. In this way, the model can not only perform well on the current training data but also maintain strong generalization ability in new and unknown distributions.
[0032] (1) Calculate the entropy of the output distribution to reveal the deviation between the data distribution and the training distribution. A higher entropy value indicates that the model encounters unknown data patterns (i.e., out-of-distribution samples), while a lower entropy value indicates that the model is dealing with familiar patterns in the training set. The entropy value of a node is defined as follows:
[0033]
[0034] where v i is the target node, is the r-ego subgraph centered on node v i , f θ is the graph representation learning framework, and C is the label set.
[0035] (2) Normalize the node energy score and entropy value, and the calculation process is as follows:
[0036]
[0037] where μ s and μ e are the means of the energy score and entropy value respectively, δ e and δ e are the standard deviations of the energy score and entropy value respectively, and are the normalized energy score and entropy value respectively.
[0038] (3) Use joint alignment regularization to minimize the difference between the output entropy and the energy score, ensure prediction consistency, and prevent energy evaluation bias. The joint alignment regularization is defined as follows:
[0039]
[0040] where is the vector of normalized node energy scores, is the vector of normalized entropy.
[0041] (4) Calculate the joint alignment loss, determine the difference between the output entropy of the last layer and the energy score of each layer, and average the differences for all k layers:
[0042]
[0043] where is the vector of normalized node energy scores of the i-th layer, is the vector of entropy of the last layer.
[0044] (5) Combine the joint alignment regularization with the cross-entropy loss function to obtain the overall optimization objective:
[0045]
[0046] Among them, CE is the cross-entropy loss, α controls the decay of the regularization weight, t is the number of iteration steps, and β is a hyperparameter that controls the strength of the regularization term.
[0047] Through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0048] Most existing graph neural network methods are designed based on the independent and identically distributed (i.i.d.) assumption, which means that the training and test data come from the same distribution. Under this assumption, GNNs assume that all data follows the same distribution, ignoring the different characteristics and complex patterns of data with different distributions. We should develop graph neural network models that can identify and adapt to out-of-distribution data. On the other hand, although some methods attempt to identify out-of-distribution data in the dataset, these methods often rely on softmax confidence scores for detection. However, the softmax posterior distribution assumes that each input belongs to a certain observed category. Even if the input is out-of-distribution data, softmax may still assign it to a known category and give a high confidence. This leads to overconfident predictions for out-of-distribution samples. This dependence may lead to the problem of overconfident posterior distributions for out-of-distribution data. In contrast, energy-based models can map the embeddings of each node to a common scalar space, where observed node instances have lower energy values, while unobserved nodes have higher energy values. For this reason, we provide a graph neural network framework EPGNN based on energy evaluation and propagation, which enhances the out-of-distribution generalization ability by using an energy model to identify out-of-distribution nodes in the graph. Specifically, to capture the complex dependencies between nodes in the original graph, we first construct a GNN encoder to generate node embeddings and encapsulate structured neighborhood information. This enables the node embeddings to carry rich semantic and structural information of each local environment in the graph. Then, to enable the model to distinguish nodes following different distributions, we design a plug-and-play out-of-distribution evaluator based on energy. This evaluator carefully designs an energy function to generate energy scores that are theoretically aligned with the probability density of the input nodes. We assign specific energy values to different nodes to help the model distinguish the distribution characteristics of the nodes, so that the model can identify node instances from previously unseen or unknown distributions. In addition, to further enhance the model's ability to detect out-of-distribution data, we introduce a plug-and-play structure-aware energy propagation module, which can improve the adaptability of node energy values as the training progresses. This makes the model more sensitive to changes in the graph structure and better adjusts its energy-based node representations. Finally, the present invention designs joint alignment regularization to enable the model to capture generalizable knowledge between different distributions, thereby improving the out-of-distribution generalization ability of the model.
[0049] In summary, the present invention proposes an out-of-distribution node detection method based on energy evaluation and propagation, named EPGNN, which can effectively detect out-of-distribution nodes in the input graph. An energy-based out-of-distribution evaluator is constructed to distinguish in-distribution nodes from out-of-distribution nodes. In addition, the present invention introduces a structure-aware energy propagation module to achieve effective energy aggregation. At the same time, the present invention also adopts a joint alignment optimization module to better guide the learning process of EPGNN. Description of the Drawings
[0050] Figure 1 is a flowchart of the out-of-distribution node detection method based on energy evaluation and propagation provided by an embodiment of the present invention.
[0051] Figure 2 is a detailed illustration of the out-of-distribution node detection method based on energy evaluation and propagation provided by an embodiment of the present invention. Detailed Embodiments
[0052] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0053] As Figure 1 shown, an embodiment of the present invention is an out-of-distribution node detection method based on energy evaluation and propagation, and the method specifically includes:
[0054] Step1: Input the data graph data G=(V, E, X, A), where V represents the node set of graph G, E represents the edge set, X represents the feature matrix, and A represents the adjacency matrix.
[0055] Step2: Define the backbone network f θ (·), initialize the parameters of the graph neural network model, and the forward propagation process of node v at the k-th layer is defined as follows:
[0056]
[0057] where represents the embedding of node v at the k-th layer, θ is the learnable parameter, σ is the non-linear activation function (such as Sigmoid), represents the set of neighbor nodes of node v, and AGG is the aggregation function.
[0058] Step 3: First, select a part of the node categories from the graph as the source node set, and the remaining node categories as the out-of-distribution node set. The data partitioning process is as follows:
[0059]
[0060] Among them, \(V\) represents the node set of graph \(G\), \(V\) S and \(V\) T represent the selected source node set and out-of-distribution node set respectively, \(C\) S and \(C\) T represent the source node category set and out-of-distribution node category set respectively. Assign the meta-label "source node" to each node in \(V\) S , and assign the meta-label "out-of-distribution node" to each node in \(V\) T .
[0061] Step 4: Construct an \(r\)-ego subgraph centered on each node \(v\) including node \(v\) and all nodes and edges within its \(r\)-hop neighborhood.
[0062] Step 5: Aggregate neighbor information through a multi-layer GNN encoder, and input the source node set \(V\) S and the out-of-distribution node set \(V\) T into \(f\) θ (·) to generate node embeddings \(Z\). The output of the \(k\)-th layer can be defined as:
[0063]
[0064] Among them, represents the \(k\)-th layer of the graph representation learning framework, \(G\) S and \(G\) T represent the subgraphs sampled through \(V\) S and \(V\) T respectively. and represent the node embedding matrices of \(G\) S and \(G\) T respectively, and their meta-labels are "source node" and "out-of-distribution node". Concatenate and to obtain the final output This is the final embedding result of the \(k\)-th layer.
[0065] Step 6: Based on the embedding representation of each node, calculate the energy value \(E(v\) i ). The lower the energy value, the more likely it is to be an in-distribution node. The energy calculation process can be defined as:
[0066]
[0067] Among them, v i is the input node, is the r-ego subgraph centered on v i , f θ is the graph representation learning framework, C is the class label of the source node set, and T is the temperature coefficient.
[0068] Step7: Use the edge energy score to quantify the weights of neighbor nodes, and its definition is as follows:
[0069]
[0070] Among them represents the edge energy score between adjacent nodes v i and v j , and are the energy scores of nodes v i and v j at the l-th layer, respectively.
[0071] Step8: Use the edge energy score as the weight to perform weighted aggregation to update the node energy score, and the update process is defined as follows:
[0072]
[0073] Among them, is the energy score of node v i i, is the set of neighbor nodes of node v i .
[0074] Step9: Calculate the entropy e i of each node in the model output to measure the uncertainty of the prediction distribution, and the entropy value calculation process is defined as follows:
[0075]
[0076] Among them, v i is the target node, is the r-ego subgraph centered on node v i , f θ is the graph representation learning framework, and C is the label set.
[0077] Step10: Normalize the energy scores of nodes at each layer and the entropy value of the last layer to eliminate the dimension difference, and the normalization process is defined as follows:
[0078]
[0079] Among them, μ s and μ eare the mean values of the energy fraction and the entropy value, respectively, and δ s and δ e are the standard deviations of the energy fraction and the entropy value, respectively, and are the normalized energy fraction and entropy value, respectively.
[0080] Step11: Calculate the difference between the output entropy of the last layer and the energy fraction of each layer, and average the differences of all k layers. Using joint alignment regularization, minimize the difference between the output entropy and the energy fraction to ensure prediction consistency. The joint alignment regularization is defined as follows:
[0081]
[0082] where, is the normalized node energy fraction vector of the i-th layer, is the normalized entropy vector of the last layer.
[0083] Step12: Identify the nodes with energy fraction lower than the threshold as in-distribution nodes Calculate the cross-entropy classification loss CE of the in-distribution nodes. The calculation process is defined as follows:
[0084]
[0085] Step13: The total loss is the weighted sum of the classification loss CE and the joint alignment regularization term :
[0086]
[0087] where, α controls the decay of the regularization weight, t is the number of iteration steps, and β is the hyperparameter that controls the strength of the regularization term.
[0088] Step14: Update the GNN parameter θ by backpropagation.
[0089] Step15: Determine whether the maximum number of iterations is satisfied. If satisfied, enter Step16; if not, enter Step5 to continue execution.
[0090] Step16: Output the finally obtained graph representation learning network f θ (·) and the set of nodes predicted as out-of-distribution nodes
[0091] Figure 2The detailed illustration of the present invention is shown. The present invention proposes a new out-of-distribution node detection method based on energy evaluation and propagation (EPGNN). Specifically, EPGNN mainly consists of three parts: an energy-based out-of-distribution evaluation module, a structure-aware energy propagation module, and a joint alignment optimization module. The energy-based out-of-distribution evaluation module first constructs a GNN encoder to obtain node embeddings containing neighborhood structure information, and then uses an energy-based out-of-distribution evaluator to assign corresponding energy values to different nodes. The structure-aware energy propagation module iteratively propagates energy scores between nodes to achieve effective energy aggregation. The joint alignment optimization module guides the learning process of EPGNN by coordinating the relationship between energy scores and output entropy.
[0092] The above-disclosed are only several specific embodiments of the present invention. Those skilled in the art can make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the present invention. However, the embodiments of the present invention are not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. A distributed out-node detection method based on energy evaluation and propagation, characterized in that: The following steps are involved: Step 1: Build a GNN encoder to encode the local neighborhood information of each node into a high-dimensional embedding; This process enables the embedding of each node to not only contain its own characteristics, but also reflect the dependency relationship between the node and its surrounding neighboring nodes. In this way, node embedding can fully capture the complexity of the graph structure while retaining the semantic and structural information of each node in the graph, providing a solid foundation for subsequent tasks. Step 2: Use an energy function to generate a corresponding energy score for each node. These scores are consistent with the probability density of the node. The energy value represents the degree of match between the node and its distribution. Nodes with higher energy values are considered to be from outside the distribution. By assigning these energy scores to nodes, the model can effectively distinguish nodes from different distributions in the graph, and then identify out-of-distribution node instances. Step 3: Introduce a structure-aware energy propagation module so that the energy value of the node can be dynamically adjusted according to the changes in the graph structure. As the training progresses, the model automatically optimizes the propagation process of the node energy value, making the energy value more sensitive to the position and structural changes of the node in the graph. In this way, the model can not only better adapt to the changes in the graph structure, but also further improve the accuracy and robustness when detecting out-of-distribution nodes. Step 4: Introduce joint alignment regularization to enable the model to capture more generalizable features under different data distributions; This regularization method helps the model learn across distributions and ensures that it can better identify and adapt when faced with out-of-distribution nodes, thereby improving the overall out-of-distribution detection performance; in this way, the model can not only perform well on the current training data, but also maintain strong generalization capabilities in new, unknown distributions.
2. The distributed out-node detection method based on energy evaluation and propagation according to claim 1 is characterized in that: The step one comprises: (1) First, select a part of node categories from the graph G as the source node set V S , these nodes constitute the main data distribution encountered by the model during the training phase; the remaining node categories are regarded as the out-of-distribution node set V T , represents a test environment that is different from the training distribution, thereby simulating the out-of-distribution generalization scenario; the data partitioning process is as follows: Among them, V represents the node set of graph G, V S and V T represent the selected source node set and the distributed node set respectively, C S and C T Respectively represent the source node category set and the distribution out-of-node category set, V S Each node in V is assigned a meta-label "source node" T Each node in is assigned the meta-label "distribution out-node"; (2) Construct a GNN embedding layer to calculate the embedded representation of the node by aggregating the local neighborhood structure information, and construct a graph representation learning skeleton network f through multiple layers of GNN embedding layers θ (·); The forward propagation process of node v in the kth layer can be defined as follows: in, represents the embedding of node v in the kth layer, θ is a learnable parameter, σ is a nonlinear activation function (such as Sigmoid), represents the set of neighbor nodes of node v, AGG is the aggregation function; (3) The source node set V S and the distributed external node set V T Input to f θ (·), we get the node embedding matrix; the output of the kth layer can be defined as: in, The graph represents the kth layer of the learning skeleton, G S and G T Respectively, through V S and V T The sampled subgraph, and Respectively represent G S and G T The node embedding matrix of , whose meta-labels are “source node” and “out-of-distribution node” respectively; (4) and Splice to get the final output This is the final embedding result of the kth layer.
3. The distributed out-node detection method based on energy evaluation and propagation according to claim 1 is characterized in that: The step 2 comprises: (1) Using the energy function from f θ Extract the energy value of the node and map the node embedding to a scalar to indicate whether the node belongs to an out-of-distribution node; the function is defined as follows: Among them, v i is the input node, For v i The r-ego subgraph centered at f θ is the learning skeleton for graph representation, C is the category label of the source node set, and T is the temperature coefficient; (2) When the energy value of a node exceeds a certain threshold, the node is judged as an out-of-distribution node because it exhibits higher uncertainty and uniformity of category prediction distribution.
4. The distributed out-node detection method based on energy evaluation and propagation according to claim 1 is characterized in that: The step three comprises: (1) In order to enhance the energy difference between nodes with different distributions, the energy values of neighbor nodes are aggregated to the target node. In this process, a reasonable approach is to assign higher weights to neighbor nodes with the same distribution as the target node. On the contrary, neighbor nodes with other distributions have less impact on energy aggregation and should be assigned lower weights. In this way, the energy value can better reflect the distribution difference between nodes. (2) The edge energy score is used to quantify the weight of neighbor nodes, which is defined as follows: in, Represents the adjacent node v i and v j The edge energy fraction between and The nodes v i and v j Energy fraction at level l; (3) Aggregate node energy scores and edge energy scores to update the energy score of the target node. The update process is defined as follows: in, For the festival i The energy fraction of i, For node v i The set of neighbor nodes.
5. The distributed out-node detection method based on energy evaluation and propagation according to claim 1 is characterized in that: The step 4 comprises: (1) Calculate the entropy of the output distribution to reveal the deviation between the data distribution and the training distribution; a higher entropy value indicates that the model encounters unknown data patterns (i.e., out-of-distribution samples), while a lower entropy value indicates that the model is processing familiar patterns in the training set; the entropy value of a node is defined as follows: Among them, v i is the target node, For node v i The r-ego subgraph centered at f θ is the graph representation learning skeleton, C is the label set; (2) Normalize the node energy score and entropy value. The calculation process is as follows: Among them, μ s and μ e are the means of energy fraction and entropy value, δ s and δ e are the standard deviations of energy fraction and entropy value, respectively, and are the normalized energy score and entropy value, respectively; (3) Use joint alignment regularization to minimize the difference between the output entropy and the energy score, ensure prediction consistency, and prevent energy evaluation bias; joint alignment regularization is defined as follows: in, is the normalized node energy score vector, is the normalized entropy vector; (4) Calculate the joint alignment loss by determining the difference between the last layer output entropy and the energy scores of each layer, and average the differences over all k layers: in, is the normalized node energy score vector of the i-th layer, is the entropy vector of the last layer; (5) Combine the joint alignment regularization with the cross entropy loss function to obtain the overall optimization goal: Among them, CE is the cross entropy loss, α controls the attenuation of the regularization weight, t is the number of iterations, and β is a hyperparameter that controls the strength of the regularization term.
Citation Information
Cited By
Federal learning distribution external generalization detection method based on local attention enhancement and singular vector global modeling
CN121413803A
A federated learning out-of-distribution generalization detection method based on local attention enhancement and singular vector global modeling
CN121413803B