Graph neural network self-adaptive discarding method and system based on information entropy

Through the adaptive discarding method based on information entropy, the message discarding rate in the graph neural network is dynamically adjusted, which solves the overfitting and noise sensitivity problems of the graph neural network in non-Euclidean spatial data processing, and improves the robustness and generalization ability of the model.

CN120337986AActive Publication Date: 2025-07-18NANJING UNIV OF INFORMATION SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510819465.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Graph neural networks face overfitting, oversmoothing and high noise sensitivity when processing non-Euclidean spatial data. Especially on training sets with simple graph structure or single information, the model generalization ability is weak, and it is highly dependent on graph topology, and is susceptible to noise interference.

Method used

Through an adaptive discarding method based on information entropy, the message discarding rate of nodes is dynamically adjusted, and edge index mapping is used by PyTorch Geometric framework, Bernoulli mask is generated, message matrix is scaled, and graph neural network model is input to realize adaptive discarding.

Benefits of technology

It improves the robustness and generalization ability of the model, reduces sample variance, suppresses overfitting and oversmoothing, and enhances performance on complex graph data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337986A_ABST
    Figure CN120337986A_ABST
Patent Text Reader

Abstract

The invention discloses a graph neural network self-adaptive discarding method and system based on information entropy, and relates to the technical field of graph neural network random discarding strategies, and the method comprises the steps: receiving a feature vector of a node, carrying out Softmax normalization processing on the feature vector of the node, calculating an entropy value corresponding to the node based on the information entropy, and carrying out the Softmax normalization processing on the entropy value; normalizing the entropy value corresponding to the node, and multiplying the normalized entropy value by a preset global maximum discarding rate to obtain a personalized discarding probability of the node; mapping the personalized discard probability of the node through an edge index mechanism in a PyTorch Geometric framework to obtain a mapping edge, and sampling based on the mapping edge to obtain a Bernoulli mask; and performing scaling based on the Bernoulli mask and the personalized discarding probability of the node to obtain a disturbed message matrix, and inputting the disturbed message matrix into a pre-established graph neural network model, thereby realizing adaptive discarding of the graph neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of random dropout strategies for graph neural networks, and specifically to an adaptive dropout method and system for graph neural networks based on information entropy. Background Technique

[0002] Graph Neural Networks (GNNs) as a deep learning model specialized for processing data in non-Euclidean spaces have demonstrated powerful modeling capabilities in graph representation learning tasks.

[0003] Although significant progress has been made in the theoretical research and practical applications of graph neural networks, during the actual deployment process, due to factors such as excessive model complexity, convergence of features caused by deep information aggregation, and high sensitivity to graph topologies, key technical challenges such as overfitting, oversmoothing, and high noise sensitivity still exist. Among them, overfitting is mainly due to the high dependence of graph neural networks on node features and graph structures. On training sets with simple graph structures or single information, the model may overly rely on local features of specific nodes and have weak generalization ability in test data. In addition, the imbalance in label distribution also makes graph neural networks overly focus on categories with a larger number of samples and ignore minority categories, further exacerbating overfitting. Deep graph neural networks aggregate information through multiple layers of neighborhoods, resulting in node representations gradually becoming similar, inhibiting the model's ability to identify local patterns, and further exacerbating overfitting. The oversmoothing phenomenon stems from multi-layer information aggregation. As the number of layers increases, node features tend to be similar, and the feature difference weakens, which is particularly obvious in graphs with simple structures or high node similarity. The lack of an effective regularization mechanism exacerbates this problem. High noise sensitivity is due to the high dependence of graph neural networks on graph topologies. Any inaccurate or missing edge information may interfere with the propagation of node features, leading to biased features learned by the model, and thus affecting the performance and generalization ability of the model under poor data quality or noisy data. Summary of the Invention

[0004] To solve the deficiencies mentioned in the above background technique, the purpose of the present invention is to provide an adaptive dropout method and system for graph neural networks based on information entropy.

[0005] In a first aspect, the purpose of the present invention can be achieved through the following technical solutions: An adaptive dropout method for graph neural networks based on information entropy, the method comprising the following steps: Receiving the feature vector of a node, performing Softmax normalization processing on the feature vector of the node, calculating the entropy value corresponding to the node based on information entropy, normalizing the entropy value corresponding to the node, and multiplying it by a preset global maximum dropout rate to obtain the personalized dropout probability of the node; Map the personalized dropout probability of nodes through the edge index mechanism in the PyTorch Geometric framework to obtain mapped edges, and sample based on the mapped edges to obtain a Bernoulli mask; Scale based on the Bernoulli mask and the personalized dropout probability of nodes to obtain a perturbed message matrix, and input the perturbed message matrix into a pre-established graph neural network model, thereby realizing the adaptive dropout of the graph neural network.

[0006] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the process of performing Softmax normalization on the feature vector of the node: Let the node feature matrix be , where the th row represents the feature vector of node , and use the Softmax function to normalize the feature vector into a probability distribution: In the formula, represents the probability that node belongs to a certain specific category, represents the original input logits, The exponential operation of is part of the Softmax function, is the step of normalizing over all categories, represents the index of different categories or features; represents the th node's th feature, represents the th node's th feature.

[0007] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the calculation formula for calculating the entropy value corresponding to the node based on the information entropy is as follows: In the formula, is the information entropy of node , is the probability component of node in the th dimension after being normalized by the Softmax function, is an index variable representing each dimension in the node feature vector; is the total number of dimensions of the node features, is a constant to prevent .

[0008] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: normalizing the entropy value corresponding to the node, including: In the formula, is the normalized information entropy, representing the value of the information entropy of node after being standardized, is the maximum value among the information entropies of all nodes.

[0009] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the personalized discard probability of the node is defined by the following formula for the adaptive discard rate of a single node as: where is the global maximum discard rate hyperparameter, representing the highest discard probability that the node with the maximum entropy can reach.

[0010] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the process of scaling based on the Bernoulli mask and the personalized discard probability of the node to obtain the perturbed message matrix: For each element in the message matrix, according to the adaptive discard rate of the source node , perform Bernoulli sampling as follows, where represents the message from node to node : In the formula, is a random variable, is the Bernoulli distribution, is the success probability of the Bernoulli distribution, is a constant related to node .

[0011] Subsequently, generate the perturbed message matrix , and the calculation formula for its elements is as follows: In the formula, this is the perturbed message, is the original message from node to node . In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the When , retain the message and perform scaling; when When this occurs, the message is completely discarded; the scaling factor ensures that the perturbed message is consistent with the original value in expectation.

[0012] Combined with the first aspect, in some implementations of the first aspect, the method further includes: The pre-established graph neural network model is as follows: In each layer of the graph neural network, each node receives feature information from its neighbor nodes, integrates the neighbor information through an aggregation operation, and performs a non-linear transformation combined with its own features, enabling the graph neural network to achieve node classification, graph classification, and link prediction; Let the undirected graph be denoted as where represents the set of nodes represents the set of edges; the node feature matrix is where represents the feature vector of node , is the feature dimension of the node. The edges describe the relationship between nodes and can be represented by the adjacency matrix where represents the th row of the adjacency matrix, represents the relationship between node and If , it means that there is an edge connection between node and . The node degree vector is denoted as where represents the number of edges connected to node , that is , Calculate the sum of the weights of all edges connected to node . The total degree of the graph is calculated by . When applying the message-passing graph neural network, the message matrix can be represented as where is the message passed between nodes, is the total number of messages passed in the graph, is the dimension of the message; The specific steps of the message-passing graph neural network are as follows: (1) Generation of messages from neighbor nodes: The features of each node act together with the features of its neighbor nodes and the features of the edges through the message generation function . The message generation formula is: In the formula, is at the th layer from node To the node The message passed; And Respectively represent the nodes And node At the Layer feature vectors, Indicates at the Layer from node To node Edge, Is the Layer message generation function; (2), Introduction and calculation of node information entropy: Quantitatively evaluate the information complexity of each node feature in the graph through information entropy; Normalize the feature vector of the node to convert it into a probability distribution; calculate the initial entropy value of each node based on the information entropy formula, and normalize the entropy values of all nodes to obtain the normalized information entropy on a unified scale; (3), Design of adaptive dropout rate: Determine the message dropout rate of each node according to the normalized information entropy. The allocation of the dropout rate is proportional to the entropy value, and the global maximum dropout rate is used as a hyperparameter to control the upper limit of the dropout intensity; (4), Dynamic regulation of message passing graph neural network: Training phase: Randomly sample each message, and decide whether to discard it according to the adaptive dropout rate of the source node. If not discarded, scale it proportionally to keep the overall expected value stable; if discarded, set it directly to zero; Testing phase: Directly use the complete original message matrix without performing the dropout operation.

[0013] In a second aspect, to achieve the above object, the present invention discloses an adaptive dropout system for a graph neural network based on information entropy, including: A probability acquisition module, configured to receive the feature vector of the node, perform Softmax normalization processing on the feature vector of the node, calculate the entropy value corresponding to the node based on information entropy, normalize the entropy value corresponding to the node, and multiply it by a preset global maximum dropout rate to obtain the personalized dropout probability of the node; A sampling module, configured to map the personalized dropout probability of the node through the edge index mechanism in the PyTorch Geometric framework to obtain a mapped edge, and perform sampling based on the mapped edge to obtain a Bernoulli mask; An adaptive dropout module, configured to scale based on the Bernoulli mask and the personalized dropout probability of the node to obtain a perturbed message matrix, and input the perturbed message matrix into a pre-established graph neural network model, thereby realizing the adaptive dropout of the graph neural network.

[0014] In another aspect of the present invention, in order to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, it adopts a method for adaptive dropout of graph neural network based on information entropy as described above.

[0015] Advantages of the present invention: By analyzing the influence of perturbations on node representations, the present invention can gain a deeper understanding of the working mechanisms and performance differences of various random dropout mechanisms during model training. The adaptive dropout strategy based on information entropy can effectively reduce sample variance by focusing on discarding information in the message matrix, thereby improving the stability of the training process and accelerating the convergence of the model. The dynamic dropout rate strategy based on dynamic evaluation of node information entropy is significantly superior to the fixed dropout rate strategy under the premise of ensuring the same average dropout rate. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings; Figure 1 It is a schematic flowchart of the method of the present invention; Figure 2 It is a schematic diagram of the principles of three dropout methods of the present invention; Figure 3 It is a schematic diagram of the principle of the adaptive dropout strategy of the graph neural network based on dynamic evaluation of node information entropy of the present invention; Figure 4 It is a schematic diagram of the variance calculation of random dropout of the present invention; Figure 5 It is a comparison chart of perturbations with different dropout rates in the dataset of the present invention; Figure 6 It is a schematic diagram of the over-smoothing degree of MADGap during training of three methods of the present invention; Figure 7 It is a schematic diagram of the principle of message passing of the graph neural network of the present invention; Figure 8 It is a schematic diagram of the system structure of the present invention. Detailed Embodiments

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0018] Embodiment 1: As Figure 1 shown, an information entropy-based graph neural network adaptive dropout method includes the following steps: S101: Receive the feature vector of the node, perform Softmax normalization on the feature vector of the node, calculate the entropy value corresponding to the node based on information entropy, normalize the entropy value corresponding to the node, and multiply it by the preset global maximum dropout rate to obtain the personalized dropout probability of the node; The process of performing Softmax normalization on the feature vector of the node: Let the node feature matrix be , where the th row represents the feature vector of node . Use the Softmax function to normalize the feature vector into a probability distribution In the formula, represents the probability that node belongs to a certain specific category, represents the original input logits, The exponential operation of is a part of the Softmax function, is the step of normalizing all categories, represents the index of different categories or features; represents the th node and the th feature, represents the th node and the th feature.

[0019] The calculation formula for calculating the entropy value corresponding to the node based on information entropy is as follows: In the formula, is the information entropy of node , is the probability component of node in the th dimension after normalization by the Softmax function, is an index variable representing each dimension in the node feature vector. is the total number of dimensions of the node features, is to prevent constant.

[0020] The normalization of the entropy value corresponding to the node includes: In the formula, is the normalized information entropy, representing the value of the information entropy of the node after standardization. The normalized information entropy is used to unify the scale, making the entropy values of all nodes comparable and avoiding unstable model training due to excessive differences in node features. is the maximum value among the information entropies of all nodes, that is, the maximum information entropy of all node features in the graph. By dividing the information entropy of each node by the maximum value, the information entropy can be normalized to a smaller range, making its value more unified and comparable.

[0021] Suppose there are 3 nodes A, B, and C in the training of the network, and their feature vectors are: Feature vector of node A: 0.5, 0.3, 0.2; Feature vector of node B: 0.1, 0.7, 0.2; Feature vector of node C: 0.6, 0.2, 0.2 Node feature matrix is: Next, apply Softmax normalization to the feature vector of each node. For the feature vector [0.5, 0.3, 0.2] of node A, the calculation formula of Softmax is: First, calculate to values: Then, calculate the Softmax normalization constant: So the normalized probability of node A is: The same process is applied to the feature vectors of node B and node C.

[0022] The formula for information entropy is: For node A, its normalized probability is substituted into the calculation: The same method is applied to Node B and Node C.

[0023] Next, normalize the information entropy. The information entropy of Node A , and the information entropies of Node B and Node C are respectively and .

[0024] Normalize the information entropy of each node. The formula is: The normalized information entropy of Node A: The normalized information entropy of Node B: The normalized information entropy of Node C: S102: Map the personalized dropout probability of the nodes through the mechanism of edge index in the PyTorch Geometric framework to obtain the mapped edges, and sample based on the mapped edges to obtain the Bernoulli mask; PyTorch Geometric (PyG) is a graph neural network (Graph Neural Network) library built on PyTorch, which is specifically used to process graph data. This library provides a flexible way to represent graph data and a series of efficient graph convolution operations. PyG mainly relies on the Message Passing mechanism, which makes the message passing between nodes in the graph the core function of the graph neural network. In PyG, graph data is stored through the Data class, and the edge and node features of the graph are stored in attributes such as edge_index and x. edge_index is a 2×E tensor, where E is the number of edges in the graph, and it stores the connection information of the edges.

[0025] The personalized dropout probability of the said node is defined by the following formula for the adaptive dropout rate of a single node as: In the formula, is the global maximum dropout rate hyperparameter, which represents the highest dropout probability that the node with the maximum entropy can reach.

[0026] The process of scaling based on the Bernoulli mask and the personalized dropout probability of the nodes to obtain the perturbed message matrix: For each element in the message matrix, according to the adaptive dropout rate of the source node , perform the Bernoulli sampling as shown in the following formula, where represents sampling from the node Message to node : Subsequently, a perturbed message matrix is generated , and the calculation formula for the elements therein is as follows: In the formula, This is the perturbed message, indicating that under the dropout strategy, the adjusted value of the message from node to node . This value is a weighted version of the original message , and the weight depends on the dropout strategy. is the original message from node to node . In the absence of a dropout strategy, the message represents the amount of information or features transmitted from node to node .

[0027] When , the message is retained and scaled; when , the message is completely discarded; the scaling factor ensures that the perturbed message is consistent with the original value in expectation.

[0028] Calculate the adaptive dropout rate for each node according to the normalized information entropy . Assuming the maximum dropout rate max_dropout = 0.5, then: Dropout rate of node A: Dropout rate of node B: Dropout rate of node C: Assume that the present invention calculates the message from node A to node B, and based on the dropout rate of node A, perform Bernoulli sampling: According to the perturbation formula: S103: Scale based on the Bernoulli mask and the personalized dropout probability of the nodes to obtain a perturbed message matrix, and input the perturbed message matrix into a pre-established graph neural network model, thereby realizing the adaptive dropout of the graph neural network. Calculating the perturbed message matrix from node A to node B completes the random dropout step based on information entropy. Through this regularization, the robustness and generalization ability of the model can be improved; The pre-established graph neural network model is as follows: Graph neural networks are a class of deep learning models designed specifically for processing graph-structured data, and the core mechanism is message passing. In each layer of the graph neural network, by passing information between nodes, the feature representation of the nodes is gradually updated, as Figure 7 shown. As can be seen from the figure, in each layer of the graph neural network, each node receives feature information from its neighbor nodes, integrates the neighbor information through an aggregation operation, and combines its own features for non-linear transformation, thereby continuously optimizing the representation of the nodes. This mechanism enables the graph neural network to perform tasks such as node classification, graph classification, and link prediction.

[0029] Let an undirected graph be denoted as , where represents the set of nodes represents the set of edges. The node feature matrix is , where represents the feature vector of node , is the feature dimension of the node. The relationship between nodes described by edges can be represented by the adjacency matrix , where represents the th row of the adjacency matrix, represents the relationship between node and . If , it means that there is an edge connection between node and . The node degree vector is denoted as , where represents the number of edges connected to node , that is . Calculate the sum of the weights of all edges connected to node . At the same time, the total degree of the graph is calculated by . When applying the message passing graph neural network on graph G, the message matrix can be expressed as , where is the message passed between nodes, is the total number of messages passed in the graph, is the dimension of the message.

[0030] FromFigure 7 It can be seen that nodes are connected by edges to form an undirected graph. As Figure 2 shown in the input graph in node is connected to its neighbor nodes , etc. by edges. Each node has a feature vector , where represents node , represents the current layer (or time step). Edges represent the relationships between nodes, usually represented by an adjacency matrix , and the element indicates the existence of an edge between node and . The specific steps of the message-passing graph neural network are as follows: Message generation of neighbor nodes. The features of each node act together with the features of its neighbor nodes and the features of the edges through the message generation function to generate a message. The message generation formula is: In the formula, is the message passed from node to node in the -th layer; and represent the feature vectors of node and node in the -th layer respectively, represents the edge from node to node in the -th layer, is the message generation function in the -th layer, which generates a message based on the features of nodes and edges. Figure 2 The dashed box in represents the message generation function

[0031] Aggregation: After receiving the messages, the central node needs to aggregate the information from its neighbor nodes. There are various ways of aggregation, and common ways include operations such as summation, averaging, or maximum value. Through the aggregation operation, the node can combine all the messages from neighbor nodes to generate a new representation, and the formula is as follows: Node update: The result after aggregation will be used as the input for the next node update. The central node It will update its own features according to its current state and the aggregated messages. The update process is completed through a transformation function to generate a new node representation , and the formula is as follows: In the formula, represents the feature representation of node at the th layer; is a transformation function that generates an updated node feature by combining the feature of the current node and the aggregated messages of neighbor nodes. In Figure 2 , the arrow points from the transformation function to the updated node feature , representing the update process of the node feature.

[0032] By continuously performing message generation, message aggregation, and node update operations, the graph neural network can optimize the feature representation of nodes layer by layer. The information transmission of each layer enables nodes to not only utilize their own feature information but also gradually incorporate the global structure information of the graph through interactions with neighbor nodes. After multi-layer information transmission, the features of nodes become more accurate and representative, thus being better used for graph analysis tasks.

[0033] This model is based on the graph neural network (Graph Neural Network) and an information entropy-driven adaptive dropout strategy, aiming to optimize the information flow in the graph neural network by dynamically adjusting the dropout rate of node information. This method can effectively avoid overfitting, over-smoothing, and information redundancy problems, and improve the performance of the model on complex graph data. The present invention innovates on the basis of the traditional graph neural network by using information entropy to dynamically control the information transmission of each edge. Model structure overview: The graph neural network (Graph Neural Network) updates the representation of nodes by aggregating the features of neighbor nodes. Usually, nodes pass messages to their neighbors through an adjacency matrix to update their features. The model of the present invention introduces an information entropy-driven adaptive dropout strategy on the basis of the traditional graph neural network to dynamically adjust the dropout probability of messages. Specifically, the message dropout probability of each node is determined according to the complexity of the features of the node (i.e., information entropy). The higher the information entropy of a node, the greater its message dropout probability.

[0034] To quantitatively evaluate the information complexity of each node feature, the present invention introduces node information entropy as a metric. Information entropy can effectively reflect the uncertainty and diversity of the node feature distribution, thereby characterizing the information richness of nodes in the graph structure. By introducing information entropy, it not only helps to identify key nodes but also provides a theoretical basis and quantitative support for subsequent feature selection and representation learning. The definition is as follows: In the formula, is the node 's eigenvector The -th dimensional component after softmax normalization; is a small constant to prevent ; The information entropy value of a node reflects the distribution balance degree of its features in each dimension, and is an important indicator to measure the information complexity and uncertainty of the node. The higher the entropy value, the more uniform the distribution of the node in each feature dimension, the more complex and rich the information carried, and at the same time, there is higher information redundancy. Therefore, such nodes have a higher tolerance for feature perturbation or random dropout during the message passing process. On the contrary, nodes with lower entropy values often have more concentrated and stronger-structured key information, and their feature distributions show a high degree of bias, representing a more definite and compact semantic expression. In this case, it is necessary to retain its message as completely as possible to avoid the loss of key information. Through feature analysis based on information entropy, a more refined message regulation strategy can be realized, so as to improve the robustness and generalization ability of the model while retaining effective information.

[0035] Specifically, the solution of the present invention will be further elaborated below through embodiments: Based on the implicit regularization effect introduced by the random dropout strategy in the graph neural network, this paper focuses on how to affect the representation learning and generalization ability of the model by regulating the perturbation intensity during the message propagation process. For the convenience of theoretical derivation and analysis, the present invention discusses under a simplified setting, assuming that the used graph neural network model is a single-layer structure and the downstream task is a binary classification problem. In this setting, the representation of each node is obtained by aggregating and transforming the features of its neighbor nodes, so as to form an embedded expression of the local graph structure. By analyzing the influence of perturbation on the node representation, a deeper understanding of the action mechanisms and performance differences of various random dropout mechanisms during the model training process can be achieved. Specifically, it can be expressed as: In the formula, is the updated feature matrix of the node, is the adjacency relationship matrix of the node (after normalization), is the message matrix is the transformation matrix. Then, a sigmoid activation function is used for nonlinear transformation, which is expressed as: During model training, the cross-entropy loss function is usually used to measure the difference between the prediction and the actual label. The expression of the cross-entropy loss function is: In the formula, and respectively represent the output features of nodes and , and are the labels of the nodes.

[0036] During the training process, when applying the dropout method, the original message matrix will no longer be used, but instead the perturbed message matrix will be used, that is, some elements of the message matrix are masked through the dropout operation. Under the framework of the information entropy-driven dropout method, the perturbed message matrix is generated by applying a Bernoulli mask to each element. Specifically, it is expressed as: In the formula, is the global dropout rate, is a mask generated by the Bernoulli distribution, used to control whether each message is dropped. By applying the mask to each element in , the random dropout of the message matrix can be achieved. in the formula is a very small constant (such as ), used to avoid numerical instability and can be ignored in the derivation formula. The perturbed message matrix is substituted into the objective function, and the new expected loss function is denoted as . The expected value of the objective function can be expressed as: Simplified to: Where: For and the two functions are respectively expanded by Taylor series: Substitute the second-order Taylor expansion of and into to get: In the text, is approximately used to replace , because the true label The variance characteristic presents a behavior pattern similar to the predicted probability - the variance reaches its maximum when the probability value is close to 0.5, and gradually decreases as it approaches 0 or 1. This correspondence not only facilitates theoretical understanding but also provides a simplified basis for formula derivation.

[0037] Due to , The formula is finally simplified to: After introducing the random dropout perturbation and performing a Taylor expansion on the loss function, the final expected loss function is obtained: As can be seen from the above formula, this loss function contains a regularization term that can suppress overfitting and oversmoothing phenomena, thereby enhancing the generalization ability and robustness of the model. The regularization term reduces the variance of node features, prompting the model to pay more attention to important features, reducing the influence of redundant information, and making the node representation more stable and reliable. In this way, the random dropout method can improve the performance of the graph neural network model when dealing with complex graph data, especially in high-dimensional tasks, and can improve the robustness and generalization ability of the model.

[0038] The present invention theoretically analyzes the adaptive message dropout strategy based on information entropy and proves its advantages over other dropout strategies from two aspects: (1) By comparing the differences in sample variance between this method and other dropout rate methods, it is proved that the adaptive dropout strategy based on information entropy can effectively reduce the sample variance by focusing on discarding information in the message matrix, thereby improving the stability of the training process and accelerating the convergence of the model. (2) By deriving and comparing the variances of the dynamic dropout rate strategy and the fixed dropout rate strategy, it is proved that the dynamic dropout rate strategy based on the dynamic evaluation of node information entropy is significantly better than the fixed dropout rate strategy under the premise of ensuring the same average dropout rate.

[0039] To prove the effectiveness of the message dropout strategy based on information entropy in reducing sample variance, the specific steps include: establishing a mathematical model of this strategy and deriving the theoretical boundary of its variance reduction. Further, through comparative analysis with existing random dropout methods, the effectiveness of the proposed strategy in improving the training stability of the model and accelerating the convergence process is theoretically clarified.

[0040] Random dropout methods essentially transform into corresponding dropout methods by performing a masking operation on the message matrix The message matrix is usually used to represent the communication process between nodes in a graph neural network. Therefore, the present invention evaluates the influence of these methods by comparing the message matrices in different training cycles, and calculates the sample variance through the norm of the message matrix.

[0041] Such asFigure 4 As shown, assume that the original message matrix is a matrix of all 1s with a size of , that is, the value of each element is 1. Based on this assumption, the sample variance can be calculated by the 1-norm of the message matrix. For simplicity of analysis, it is assumed that each directed edge corresponds to a row vector in the message matrix, the graph is assumed to be an undirected graph, and the degree of each node is , that is, each node is connected to four surrounding nodes by undirected edges. Therefore, there are two edges between every two connected nodes. There are two edges from node 1 to node 11 and from node 11 to node 1 in the graph shown. At this time, it can be seen from the graph that the total number of edges, which is also the total number of rows of the message matrix, is: .

[0042] The random dropout method can be regarded as multiple independent Bernoulli samplings, and the whole process conforms to the binomial distribution. Therefore, the variance of the message matrix can be calculated. The variances of the message matrix under different random dropout methods and this method will be calculated below.

[0043] Dropout: This method randomly drops the same number of features as the number of adjacent nodes after sampling the features on each node. Figure 4 There are nodes in , and the feature dimension of each node is and has adjacent nodes. Therefore, after times of Bernoulli sampling, whenever an element is dropped, it will mask

[0044] elements in the message matrix. Therefore, the variance of the message matrix Substituting the numerical values gives: DropEdge: This method randomly drops all the feature elements on an edge after sampling all the edges. Figure 4 There are edges in (assuming an undirected graph, so there are two edges between every two nodes), and the feature dimension of each edge is . Therefore, after times of Bernoulli sampling. Whenever 1 element is dropped, it will mask

[0045] elements in the message matrix. Therefore, the variance of the message matrix Substitute the numerical values to obtain: (3)DropNode: This method randomly discards all the feature elements on one node after sampling all the nodes. Figure 4 There are nodes, and the feature dimension of each node is and it has adjacent nodes. Therefore, after times of Bernoulli sampling, whenever one element is discarded, it will mask

[0046] elements in the message matrix. Therefore, the variance of the message matrix Substitute the numerical values to obtain: (4)Adaptive message discarding strategy based on information entropy drive: This method discards 1 element after sampling all the feature elements on all the nodes. Figure 4 There are nodes, and the feature dimension of each node is and it has adjacent nodes. Therefore, after

[0047] times of Bernoulli sampling, whenever one element is discarded, it will mask 1 element in the message matrix. Therefore, the variance of the message matrix Substitute the numerical values to obtain: Among them, is the discarding rate, is the discarding probability dynamically calculated according to the node feature entropy, is the number of nodes, is the feature dimension, is the degree of each node Based on the above variance analysis results, the variance characteristics of different random discarding methods can be quantitatively compared, and the following relationships exist among the variances of each sample: In the actual application scenario, the number of nodes in the graph, the average node degree and the feature dimension Far greater than the simplified conditions set in the theoretical analysis of the present invention, usually several times or even dozens of times higher. Under such high-dimensional and large-scale graph structures, the variance introduced by existing random discarding methods during the training process will be significantly amplified, further exacerbating the instability and performance fluctuations of the model. In contrast, the method proposed by the present invention has better performance in controlling the perturbation intensity and suppressing the growth of variance. Therefore, in actual large-scale graph data, the gap in variance control ability between it and other methods will also be significantly enlarged, thus further highlighting the advantages of this method in terms of stability and generalization performance.

[0048] Systematic experiments were conducted on four benchmark graph datasets, Cora, CiteSeer, PubMed, and Flicker, to verify the effectiveness and universality of the adaptive message discarding strategy based on node information entropy. The experiments used GCN and GAT as the basic frameworks, compared with mainstream discarding strategies such as Dropout, DropEdge, and DropNode, and focused on evaluating the performance in two typical graph learning tasks: node classification and link prediction.

[0049] Dataset Introduction and Experimental Parameter Settings The present invention selected five widely used benchmark graph datasets: Cora, CiteSeer, PubMed, and Flicker for experimental evaluation. These datasets cover diverse graph structure types (from academic citation networks to social networks), different scales (from thousands to tens of thousands of nodes), and various graph learning tasks (such as node classification and link prediction), and can comprehensively and systematically evaluate the performance of different message discarding strategies in various actual scenarios. The specific data information of each dataset is shown in Table 1.

[0050] Table 1 Dataset Information Among them, Cora, CiteSeer, and PubMed, as three typical academic paper citation network datasets, where the nodes represent academic papers and the edges represent citation relationships between papers. These datasets are often used for node classification tasks, and through classification experiments based on paper content features (such as keywords, abstracts, etc.), the performance of graph neural networks in academic network analysis can be effectively evaluated.

[0051] The Flicker dataset is a graph structure dataset constructed from the social media platform Flickr. The nodes represent images uploaded by users, and the edges reflect the connection relationships between images based on visual similarity or social associations. This dataset is widely used in image classification tasks, and by classifying image nodes into predefined semantic categories, the representation learning ability and classification performance of graph neural networks in a social network environment can be effectively evaluated.

[0052] This experiment will use the AdamW optimizer for training on RTX5070. The experimental environment is shown in Table 2: Table 2 Experimental environment The training epoch is set to 500, and the model consists of a two-layer graph convolution module. The hidden layer dimension of the two-layer graph convolution is set to 16. The learning rate is 0.005. The regularization rate is , The value is set to .

[0053] On four benchmark datasets, Cora, CiteSeer, PubMed and Flickr, the present invention and traditional random drop strategies (including DropNode, DropEdge and Dropout) were systematically compared in node classification tasks, and the results are shown in Tables 3 and 4. As can be seen from the table, compared with the benchmark graph neural network model without a drop mechanism, all methods that introduce drop strategies have achieved performance improvements in classification accuracy; however, there are significant differences in the performance of different strategies on different datasets, such as DropEdge has the best effect on the Cora dataset, while there is a significant performance degradation on the Flickr dataset.

[0054] Table 3 Comparison of the accuracy of this method and other discarding methods in the GCN backbone network Table 4 Comparison of the accuracy of this method and other discarding methods in the GAT backbone network In contrast, the node information entropy-based adaptive discarding strategy proposed in this paper achieved optimal performance in all test scenarios, demonstrating superior stability and adaptability, and verifying its wide applicability and robustness in graph neural networks. Specifically: 1) In terms of classification accuracy, the present invention is about 5% higher than the traditional discarding strategy on average; 2) In terms of stability, by dynamically adjusting the discarding rate, it effectively balances information retention and noise suppression; 3) In terms of adaptability, it can maintain leading performance for different graph structures and task scenarios.

[0055] In summary, the present invention shows significant advantages in accuracy, stability and generalization ability, verifying the effectiveness and advancement of the adaptive message discarding strategy based on information entropy in graph neural networks.

[0056] Robustness Analysis The present invention evaluates the classification performance of different dropout strategies on the perturbed graph to analyze the robustness of each dropout strategy by means of the odds ratio. To ensure the cleanliness of the initial data, four benchmark datasets, namely Cora, CiteSeer, PubMed, and Flickr, are selected as the experimental objects. The graph structure perturbation is simulated by randomly adding different proportions of edges (the perturbation rate increases from 0% to 30%) to the original graph, and the node classification task is carried out on this basis.

[0057] As can be seen from Tables 5 to 8, as the perturbation rate increases, the adaptive dropout strategy of the present invention has the smallest accuracy drop, indicating its strong anti-interference ability. In the extreme case where the perturbation rate is 30%, the classification accuracies of the Cora, CiteSeer, PubMed, and Flickr datasets decrease by 5.71%, 2.88%, 3.13%, and 1.73% respectively. In addition, from Figure 5 the visualization results, it can also be intuitively observed that the performance degradation amplitude of the present invention at each perturbation level is significantly lower than that of other comparison methods, further verifying the significant advantages of the present method in terms of robustness and stability.

[0058] Table 5 Robustness test results and comparison in Cora dataset Table 6 Robustness test results and comparison in CiteSeer dataset Table 7 Robustness test results and comparison in PubMed dataset Table 8 Robustness test results and comparison in Flickr dataset Analysis of over-smoothing To quantitatively evaluate the over-smoothing phenomenon, the present invention uses MADGap as a measurement index. Among them, MADGap (Mean Absolute Difference Gap) is a quantitative index for measuring the difference in node representations between layers of a graph neural network, and is defined as the mean of the norm difference of the node feature matrices of adjacent layers: where is the total number of layers of the neural network, is the total number of nodes in the graph, is the th feature vector of the th node, is the

[0059] This metric reflects the model's resistance to over-smoothing by calculating the absolute differences of node representations layer by layer and taking the global average. When the MADGap value is low, it indicates that the node features in different layers tend to be homogenized, presenting a risk of over-smoothing; conversely, it means that the model can effectively maintain the distinctiveness of inter-layer features. Therefore, MADGap can be used as an important basis for measuring the model's anti-over-smoothing ability, reflecting the degree of difference in node representations between different layers: a smaller value means that the node representations are highly consistent, indicating an over-smoothing trend in the model; a larger value means that the inter-layer representations have good distinctiveness, indicating that the model can effectively maintain the discriminative ability of nodes during training.

[0060] Figure 6 The results show that as the number of training epochs increases, the MADGap value of the present invention shows a steady upward trend and finally stabilizes, indicating that the model can continuously maintain a high degree of distinctiveness in node representations during training, significantly alleviating the over-smoothing problem. In contrast, the MADGap values of the GCN-DropNode and GCN-DropEdge methods fluctuate greatly during training, and the differences in inter-layer representations are unstable, showing a more obvious over-smoothing phenomenon.

[0061] As Figures 2 - 7 shown, Figure 2 The message matrix in represents the feature information between nodes. In the figure, the connections between 1 and 2, 3, 4, 5 indicate that nodes 2 to 5 are adjacent to node 1. Each matrix element represents the feature information between two nodes. m12 represents the feature information between node 1 and node 2; m13 represents the feature information between node 1 and node 3; m14 represents the feature information between node 1 and node 4; m15 represents the feature information between node 1 and node 5. The feature information of these adjacent nodes constitutes the message matrix in the figure. The darkness of the color in the matrix represents the level of information entropy. The darker the color, the higher the information entropy and the more complex the information flow. The stripes in the matrix represent the discarded information. The subgraphs (a), (b), (c) below the figure show three different random discarding strategies; (a) represents the Dropout method. (b) represents the DroNode method. (c) represents the DroEdget method. The color bar below represents the information entropy values, with a color gradient (from low to high) (light colors represent low information entropy and dark colors represent high information entropy); the striped squares represent the partially discarded information.

[0062] Figure 3The message matrix in [description] represents the feature information between nodes. In the figure, node 1 is connected to nodes 2, 3, 4, and 5, indicating the adjacent nodes 2 to 5 of node 1. Each matrix element represents the feature information between two nodes. m12 represents the feature information between node 1 and node 2; m13 represents the feature information between node 1 and node 3; m14 represents the feature information between node 1 and node 4; m15 represents the feature information between node 1 and node 5. The depth of the color in the matrix represents the level of information entropy, with darker colors indicating higher information entropy. The stripes in the matrix represent the discarded information. The right side of the matrix shows the method of information discarding in the message matrix of this method. The color bar below represents the color gradient of the information entropy values (from low to high); the striped squares represent the discarded partial information.

[0063] Figure 4 In [description], (b) represents a message matrix of a matrix of all 1s with a size of , that is, the value of each element is 1. n represents the number of nodes, that is, Figure 4 In (c) of [description], nodes 1, 2, 3, and node. 1, 2, 3, 4 represent the feature information on nodes 1, 2, 3, and 4 respectively. c represents the feature dimension, that is, the number of information dimensions between each pair of nodes. d represents the degree of the node, that is, the number of adjacent nodes of each node. Node 1 is adjacent to four nodes: node 11, node 12, node 13, and node 14. Node 2 is adjacent to four nodes: node 21, node 22, node 23, and node 24. Node 3 is adjacent to four nodes: node 31, node 32, node 33, and node 34. Node 4 is adjacent to four nodes: node 41, node 42, node 43, and node 44. Figure 4In (a), k represents the total number of feature information on the edges, which is also the total number of edges. 1,11 represents the edge from node 1 to node 11, where the direction is from node 1 to node 11; 11,1 represents the edge from node 1 to node 11, where the direction is from node 11 to node 1; 1,12 represents the edge from node 1 to node 12, where the direction is from node 1 to node 12; 12,1 represents the edge from node 1 to node 12, where the direction is from node 12 to node 1; 1,13 represents the edge from node 1 to node 13, where the direction is from node 1 to node 13; 13,1 represents the edge from node 1 to node 13, where the direction is from node 13 to node 1; 1,14 represents the edge from node 1 to node 14, where the direction is from node 1 to node 14; 14,1 represents the edge from node 1 to node 14, where the direction is from node 14 to node 1; 2,21 represents the edge from node 2 to node 21, where the direction is from node 2 to node 21; 21,2 represents the edge from node 2 to node 21, where the direction is from node 21 to node 2; 2,22 represents the edge from node 2 to node 22, where the direction is from node 2 to node 22; 22,2 represents the edge from node 2 to node 22, where the direction is from node 22 to node 2; 2,23 represents the edge from node 2 to node 23, where the direction is from node 2 to node 23; 23,2 represents the edge from node 2 to node 23, where the direction is from node 23 to node 2; 2,24 represents the edge from node 2 to node 24, where the direction is from node 2 to node 24; 24,2 represents the edge from node 2 to node 24, where the direction is from node 24 to node 2; 3,31 represents the edge from node 3 to node 31, where the direction is from node 3 to node 31; 31,3 represents the edge from node 3 to node 31, where the direction is from node 31 to node 3; 3,32 represents the edge from node 3 to node 32, where the direction is from node 3 to node 32; 32,3 represents the edge from node 3 to node 32, where the direction is from node 32 to node 3; 3,33 represents the edge from node 3 to node 33, where the direction is from node 3 to node 33; 33,3 represents the edge from node 3 to node 33, where the direction is from node 33 to node 3; 3,34 represents the edge from node 3 to node 34, where the direction is from node 3 to node 34; 34,3 represents the edge from node 3 to node 34, where the direction is from node 34 to node 3; 4,41 represents the edge from node 4 to node 41, where the direction is from node 4 to node 41; 41,4 represents the edge from node 4 to node 41, where the direction is from node 41 to node 4; 4,42 represents the edge from node 4 to node 42, where the direction is from node 4 to node 42; 42,4 represents the edge from node 4 to node 42, where the direction is from node 42 to node 4; 4,43 represents the edge from node 4 to node 43, where the direction is from node 4 to node 43; 43,4 represents the edge from node 4 to node 43, where the direction is from node 43 to node 4; 4,44 represents the edge from node 4 to node 44, where the direction is from node 4 to node 44; 44,4 represents the edge from node 4 to node 44, where the direction is from node 44 to node 4.

[0064] Figure 5Shows four sub - graphs, representing the experimental results on four different datasets: Cora, CiteSeer, PubMed, and Flickr. GCN - DropNode represents the experiment of integrating DropNode into the GCN network; GCN - DropEdge represents the experiment of integrating DropEdge into the GCN network; GCN - the present invention represents the experiment of integrating the present method into the GCN network; GCN - Dropout represents the experiment of integrating Dropout into the GCN network. Each sub - graph shows the trend of the performance of different dropout strategies changing with the dropout rate. The abscissa on each sub - graph represents the number of perturbations added to the dataset. 10% means adding 10% of the perturbations; 20% means adding 20% of the perturbations; 30% means adding 30% of the perturbations. The ordinate represents the accuracy under the corresponding perturbations. The experimental results of the Cora dataset are marked on the ordinate with 0.675, 0.700, 0.725, 0.750, 0.775, 0.800, 0.825. The accuracy intervals from 0.675 to 0.825 are 0.025; the experimental results of the CiteSeer dataset are marked on the ordinate with 0.580, 0.600, 0.620, 0.640, 0.660, 0.680, 0.700, 0.720. The accuracy intervals from 0.580 to 0.720 are 0.020; the experimental results of the PubMed dataset are marked on the ordinate with 0.730, 0.740, 0.750, 0.760, 0.770, 0.780, 0.790. The accuracy intervals from 0.730 to 0.790 are 0.010; the experimental results of the Flickr dataset are marked on the ordinate with 0.420, 0.440, 0.460, 0.480, 0.500, 0.520. The accuracy intervals from 0.420 to 0.520 are 0.020. GCN represents the experiment of integrating various dropout methods on this network. The yellow line is the experimental result of the DropNode method under the GCN network; the blue line is the experimental result of the DropEdge method under the GCN network; the red line is the experimental result of the Dropout method under the GCN network; the black line is the experimental result of the present method under the GCN network.

[0065] Figure 6Show the over-smoothing degree of the three methods of DropNode, DropEdge and the present invention in different training groups. GCN-DropNode means integrating DropNode into the GCN network for experiments; GCN-DropEdge means integrating DropEdge into the GCN network for experiments; GCN-the present invention means integrating the present method into the GCN network for experiments. The abscissa is the number of training groups, ranging from 0 to 500 times, with an interval of 100 times. The ordinate is the MADGap over-smoothing value, ranging from 0 to 0.6 with an interval of 0.1. GCN means integrating various dropout methods on this network for experiments. The blue line is the changing trend of the MADGap over-smoothing degree value of the DropEdge method with the increase of the number of training groups under the GCN network; the orange line is the changing trend of the MADGap over-smoothing value of the DropEdge method with the increase of the number of training groups under the GCN network; the green line is the changing trend of the MADGap over-smoothing value of the present method with the increase of the number of training groups under the DropEdge method of the GCN network.

[0066] Figure 7 Shows the processing flow of the graph neural network. Part (a) shows the input graph, including nodes , , , , and the connection relationships between them: and , , are adjacent and are connected by edges to form an undirected graph. Part (b) shows the processing flow of the graph neural network. Each node has a feature vector , where represents node , represents the current layer (or time step). represents the feature vector of node at the th layer; represents the feature vector of node at the th layer; represents the feature vector of node at the th layer; represents the feature vector of node at the th layer. represents the relationship connecting these two nodes. represents the edge from node to node ; represents the edge from node to node The edge of Indicates an edge from node to node of is a message generation function, and messages can be generated through the function. is the message passed from node to node of , and are generated through the message generation function ; , and are generated through the message generation function ; , and are generated through the message generation function . Represents message aggregation; Represents a transformation function. , and After message aggregation, the feature vector of node at the th layer is used to update the node through the transformation function to generate a new node feature representation .

[0067] It can be seen that the method proposed by the present invention is superior in suppressing over-smoothing, enhancing the discriminability of node representations, and maintaining the stability of the model, further verifying the effectiveness and advantages of the present invention.

[0068] Specifically, to verify the effect of the present invention, an experiment on abnormal detection of social networks integrating graph convolutional neural networks is conducted in combination with this application. Data collection 1. Node data: The behavioral data of users (such as posting frequency, number of likes, number of comments, social interaction behaviors, etc.) are the core node features. Usually, the behavioral data of each user are obtained from the social platform through the API interface, including: the posting, like, and comment frequencies of each user; the interaction records of the user (such as with which users the user interacts and the interaction frequency, etc.).

[0069] 2. Graph structure data: The structure information of the social network (such as the friendship and following relationships between users) is used to construct the adjacency matrix of the graph. Specifically, the relationships between users determine the edges in the graph, and common social relationships include: Whether user A and user B follow each other; Whether user A has interacted with user B (such as commenting, liking, etc.).

[0070] Data preprocessing After the data is collected, it needs to go through preprocessing steps to ensure its suitability for subsequent model training and inference.

[0071] 1. Data cleaning: Remove invalid or missing data, such as missing user behavior records or incomplete user information.

[0072] 2. Feature normalization and standardization: To ensure that the model can handle features of different scales, the behavior features of each user need to be normalized or standardized. For example, normalize the posting frequency, number of likes, number of comments, etc. so that they are within the range of [0,1].

[0073] 3. Convert node features to probability distribution: Apply Softmax normalization to the behavior features of each node to convert them into the form of probability distribution: 4. Information entropy calculation: Calculate the entropy value of each node. The higher the entropy value of a node, the more information it represents.

[0074] Feature selection and dropout rate calculation 1. Feature selection: Through entropy value calculation, select features useful for model training. Nodes with higher entropy values usually have richer behavior information, and a lower dropout rate can be used to retain more information.

[0075] 2. Dropout rate calculation: Calculate the adaptive dropout rate based on the entropy value of each node. The dropout rate of a node is positively correlated with the entropy value. The higher the entropy, the higher the dropout rate. The formula for calculating the dropout rate is: Problems to note in specific applications Identification of abnormal users Features of abnormal users: Abnormal users usually have behavior patterns different from most users. For example, there may be an abnormally high posting frequency, a disproportionate number of likes and comments, or highly single behavior. Based on the entropy-based dropout strategy, these abnormal users can be accurately identified by dynamically adjusting the dropout rate.

[0076] Relationship between entropy value and dropout rate: The lower the entropy value of a node, the more single its behavior, while the higher the entropy value of a node, the more complex its behavior. Therefore, nodes with high entropy should retain more information to avoid the loss of important behavior features.

[0077] Data volume and computational overhead Computational Overhead of Large-Scale Social Networks: As the scale of the social network expands, the number of nodes and the dimensionality of features will increase significantly. At this time, the overhead of information entropy calculation and discard rate adjustment may be relatively large. To solve this problem, a distributed computing framework can be used to process graph data in parallel and reduce the computational bottleneck.

[0078] Adjustment of Discard Rate: According to the characteristics of the data and the application scenario, adjust the maximum discard rate and entropy threshold. If the discard rate is set too high, key information may be discarded, resulting in a decline in model performance; if the discard rate is too low, redundant information may not be effectively reduced.

[0079] Example: A node (abnormal user) may be characterized by frequent posting and less interaction with other users. The complexity of its behavior is relatively low, so its entropy value is low, the discard rate is low, and the model can focus on its main behavioral characteristics.

[0080] Reducing the Propagation of Redundant Information Balance between Information Retention and Discard: Through the calculation of entropy values and the dynamic adjustment of discard rates, the model can retain more information on nodes with rich information and discard unnecessary features on nodes with redundant information, avoiding information overload and over-smoothing problems.

[0081] Example: For a node (active user), whose behavioral characteristics are rich and the information entropy is high, the model retains more information through a low discard rate, which helps to improve the learning of the behaviors of normal users.

[0082] Improving the Stability and Robustness of the Model Better Generalization Ability: The adaptive discard mechanism helps the model avoid overfitting during training, reduce the impact of noise and redundant information, and improve the robustness of the model when facing complex and dynamically changing data.

[0083] Accurate Identification of Abnormal User Behaviors Combination of Entropy Value and Discard Strategy: In the detection of abnormal users, users with low information entropy (such as those who post frequently and have no interaction) will have a higher probability of discarding unnecessary information, ensuring that the model focuses on important nodes. Through this strategy, the model can more accurately identify users with abnormal behaviors and improve the detection accuracy.

[0084] Use the social network data shown in Table 9: Node Feature Matrix: Table 9 Social Network Data Adjacency Matrix (representing the interaction relationships between users): Calculate the entropy value of each node Softmax normalizes the features of each node: Apply the Softmax function to the behavioral features of each node (such as posting frequency, number of likes, number of comments) to transform them into a probability distribution.

[0085] For example, for node , the Softmax normalization formula is: Assume that the features of node are , and the present invention calculates the normalized probability of each feature: Therefore, the normalized probability of node is: [0.3935, 0.2187, 0.3878] Calculate the information entropy: Use the information entropy formula: For node , calculate its entropy: Node has an entropy value of 1.0383.

[0086] Perform the same calculation for other nodes. Assume the entropy values are as follows: Node has an entropy value of 1.0383 Node has an entropy value of 0.9060 Node has an entropy value of 0.8887 Node has an entropy value of 1.0532 Node has an entropy value of 0.9300 Node has an entropy value of 0.4020 (abnormal user) Normalize the entropy values Calculate the normalized entropy value of each node: The normalized entropy of each node is as follows: Node : 1.0383 / 1.0532 = 0.9859 Node : 0.9060 / 1.0532 = 0.8603 Node : 0.8887 / 1.0532 = 0.8430 Node : 1.0532 / 1.0532 = 1.0000 Node : 0.9300 / 1.0532 = 0.8833 Node : 0.4020 / 1.0532 = 0.3813 (abnormal user) Calculate the discard rate Calculate the discard rate of each node according to the normalized entropy value. Assume that the global maximum discard rate is 0.5.

[0087] For example, for the node : Calculation result: Node The discard rate of = 0.4929 Node The discard rate of = 0.4302 Node The discard rate of = 0.4215 Node The discard rate of = 0.5000 Node The discard rate of = 0.4417 Node The discard rate of = 0.1906 (abnormal user) Bernoulli sampling and message discarding Bernoulli sampling: According to the discard rate of each node, the present invention uses the Bernoulli distribution for sampling to decide whether to retain each message: Where: when , retain the message and scale it; when , the message is completely discarded.

[0088] Message scaling: Scale the retained messages, and the formula is as follows: Where: when When, the message is retained and scaled; when When, the message is completely discarded; the scaling factor ensures that the perturbed message is consistent with the original value in expectation.

[0089] Message passing and feature update The messages after discarding and scaling will participate in the next node feature aggregation operation to update the features of the nodes. The perturbation (discarding messages) in the information passing process helps to enhance the robustness of the model and reduce the propagation of redundant information.

[0090] Final effect: Node , due to its low discard rate, hardly discards messages, retains more information, and ensures the full utilization of this information by the model.

[0091] Node , due to its high discard rate, messages will be discarded, reducing the interference of redundant information and helping the model to identify abnormal users.

[0092] Embodiment 2: As Figure 8 shown, to achieve the above object, the present invention discloses a graph neural network adaptive discarding system based on information entropy, including: A probability acquisition module 11, configured to receive the feature vector of a node, perform Softmax normalization processing on the feature vector of the node, calculate the entropy value corresponding to the node based on information entropy, normalize the entropy value corresponding to the node, and multiply it by a preset global maximum discard rate to obtain the personalized discard probability of the node; A sampling module 12, configured to map the personalized discard probability of the node through the mechanism of edge index in the PyTorch Geometric framework to obtain a mapped edge, and sample based on the mapped edge to obtain a Bernoulli mask; An adaptive discarding module 13, configured to scale based on the Bernoulli mask and the personalized discard probability of the node to obtain a perturbed message matrix, and input the perturbed message matrix into a pre-established graph neural network model, thereby realizing the adaptive discarding of the graph neural network.

[0093] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.

[0094] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.

[0095] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0096] The above has shown and described the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.

Claims

1. An adaptive dropout method for graph neural networks based on information entropy, characterized in that The method includes the following steps: Receiving the feature vector of a node, performing Softmax normalization on the feature vector of the node, calculating the entropy value corresponding to the node based on information entropy, normalizing the entropy value corresponding to the node, and multiplying it by a preset global maximum dropout rate to obtain the personalized dropout probability of the node; Mapping the personalized dropout probability of the node through the edge index mechanism in the PyTorch Geometric framework to obtain mapped edges, and sampling based on the mapped edges to obtain a Bernoulli mask; Scaling based on the Bernoulli mask and the personalized dropout probability of the node to obtain a perturbed message matrix, and inputting the perturbed message matrix into a pre-established graph neural network model, thereby realizing adaptive dropout of the graph neural network.

2. The adaptive dropout method of the graph neural network based on information entropy according to claim 1, wherein The process of performing Softmax normalization on the feature vector of the node: Let the node feature matrix be , where the th row represents the feature vector of node . Use the Softmax function to normalize the feature vector into a probability distribution: In the formula, represents the probability that a node belongs to a certain specific category, represents the original input logits, and the exponential operation of which is part of the Softmax function, which is a step of normalizing over all categories, represents the index of different categories or features; denotes the th feature of the th node, and denotes the th feature of the th node.

3. The adaptive dropout method for graph neural network based on information entropy according to claim 2, wherein The calculation formula for calculating the entropy value corresponding to the node based on information entropy is as follows: Wherein, is the information entropy of the node , is the probability component of the -th dimension after normalization by the Softmax function for the node , is an index variable, representing each dimension in the node feature vector; is the total number of dimensions of the node features, is a constant to prevent .

4. An adaptive dropout method for graph neural networks based on information entropy according to claim 3, characterized in that, Normalizing the entropy value corresponding to the node includes: In the formula, is the normalized information entropy, representing the value of the information entropy of node after standardization, and is the maximum value among the information entropies of all nodes.

5. A graph neural network adaptive dropout method based on information entropy according to claim 4, characterized in that The personalized discard probability of the node is defined by the following formula for the adaptive discard rate of a single node as follows: Among them is the global maximum dropout rate hyperparameter, representing the highest dropout probability that the node with the maximum entropy can reach.

6. The adaptive dropout method of the graph neural network based on information entropy according to claim 5, characterized in that, The process of scaling based on the Bernoulli mask and the personalized dropout probability of the node to obtain a perturbed message matrix: For each element in the message matrix , according to the adaptive discard rate of the source node , perform Bernoulli sampling, where represents the message from node to node : wherein, is a random variable, is a Bernoulli distribution, is the success probability of the Bernoulli distribution, is related to node is a related constant; Subsequently, a perturbed message matrix is generated , and the calculation formula for its elements is as follows: In the formula, This is the perturbed message, is the node to node of the original message.

7. A graph neural network adaptive dropout method based on information entropy according to claim 6, characterized in that, The When happens, the message is retained and scaled; when happens, the message is completely discarded; the scaling factor ensures that the perturbed message is consistent with the original value in expectation.

8. A method for adaptively discarding a graph neural network based on information entropy according to claim 1, characterized in that, The pre-established graph neural network model is as follows: In each layer of the graph neural network, each node receives feature information from neighbor nodes, integrates the neighbor information through an aggregation operation, and performs a non-linear transformation in combination with its own features, enabling the graph neural network to achieve node classification, graph classification, and link prediction; Let the undirected graph be denoted as , where represents the set of nodes represents the set of edges; The node feature matrix is , where represents the feature vector of node , is the feature dimension of the node. The edge describes the relationship between nodes and can be represented by the adjacency matrix , where represents the th row of the adjacency matrix, represents the relationship between node and . If , it means there is an edge connection between node and . The node degree vector is denoted as , where represents the number of edges connected to node , that is . calculates the sum of the weights of all edges connected to node . The total degree of the graph is calculated by . When applying the message-passing graph neural network, the message matrix can be represented as , where is the message passed between nodes, is the total number of messages passed in the graph, is the dimension of the message. The specific steps of the message passing graph neural network are as follows: (1) Generation of messages from neighbor nodes: The features of each node are generated by a message generation function which, together with the features of its neighbor nodes and the features of the edges, act together. The formula for generating the message is: In the formula, is the message transmitted from node to node in the layer; and respectively represent the feature vectors of node and node in the layer, represents the edge from node to node in the layer, is the message generation function of the layer; (2) Introduction and calculation of node information entropy: Quantitatively evaluating the information complexity of the features of each node in the graph through information entropy; Normalizing the feature vector of the node to convert it into a probability distribution; calculating the initial entropy value of each node based on the information entropy formula, and normalizing the entropy values of all nodes to obtain a normalized information entropy with a unified scale; (3) Design of the adaptive dropout rate: Determining the message dropout rate of each node according to the normalized information entropy, and the distribution of the dropout rate is proportional to the size of the entropy value. The global maximum dropout rate is used as a hyperparameter to control the upper limit of the dropout intensity; (4) Dynamic regulation of the message passing graph neural network: Training stage: Randomly sampling each message, deciding whether to discard it according to the adaptive dropout rate of the source node. If not discarded, it is scaled proportionally to keep the overall expected value stable; if discarded, it is directly set to zero; Testing stage: Directly using the complete original message matrix without performing the dropout operation.

9. An adaptive dropout system for graph neural networks based on information entropy, which adopts an adaptive dropout method for graph neural networks based on information entropy as described in any one of claims 1 to 8, characterized in that, Including: A probability acquisition module for receiving the feature vector of a node, performing Softmax normalization on the feature vector of the node, calculating the entropy value corresponding to the node based on information entropy, normalizing the entropy value corresponding to the node, and multiplying it by a preset global maximum dropout rate to obtain the personalized dropout probability of the node; A sampling module for mapping the personalized dropout probability of the node through the edge index mechanism in the PyTorch Geometric framework to obtain mapped edges, and sampling based on the mapped edges to obtain a Bernoulli mask; An adaptive dropout module is used to scale based on a Bernoulli mask and the personalized dropout probability of nodes to obtain a perturbed message matrix, and the perturbed message matrix is input into a pre-established graph neural network model, thereby realizing the adaptive dropout of the graph neural network.

10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, A computer program capable of running on a processor is stored in the memory. When the processor loads and executes the computer program, a graph neural network adaptive dropout method according to any one of claims 1 to 8 is adopted.

Citation Information

Patent Citations

  • Underwater target feature extraction method based on convolutional neural network (CNN)

    CN107194404A

  • Neural network regularization method based on feature space correlation

    CN111950699A

  • Object intention prediction method and device, computer equipment and storage medium

    CN116662814A

  • Robustness classification method based on graph neural network in edge environment

    CN118916742A

  • Random discarding method and system of graph neural network based on dynamic discarding

    CN119646481A