A method and device for constructing a high-depth graph neural network based on energy constraints
By introducing residual layers and Dirichlet energy constraints into the graph neural network, limiting the Dirichlet energy range, the oversmooth problem of graph neural network is solved, the performance of the deep graph neural network is improved, and efficient node embedding and classification tasks are achieved.
Patent Information
- Application Number
- CN202510461285.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing graph neural network has too smoothing problems, which leads to the performance of deep networks in node classification tasks, making it difficult to build a high-performance deep graph neural network.
By constructing a high-deep graph neural network based on energy constraints, using residual layer and Dirichlet energy constraints, limiting Dirichlet energy within a specified range, determining the value of the combined coefficients, and constructing a target graph neural network.
It solves the problem of oversmoothing of graph neural networks, improves the performance of deep graph neural networks, and can generate large batches of high-deep graph neural networks, which are suitable for tasks such as generating character portraits, network attack detection, and academic knowledge graph discipline classification.
Smart Images

Figure CN119990184B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular, to a method and apparatus for constructing a high-depth graph neural network based on energy constraint. Background Art
[0002] There is an over-smoothing problem in graph neural networks, that is, as the number of network layers increases, the node embeddings of graph data tend to be the same, and different nodes cannot be distinguished, resulting in a decline in its performance in node classification tasks.
[0003] Shallow graph neural networks capture the local structural information of graph data. The deeper the graph neural network, the more globalized the captured features. In the feature extraction tasks of complex graph data such as chemical molecular structures, knowledge graphs, and social network graphs, it is often necessary to obtain the global structural information of graph data to achieve better performance in node classification tasks.
[0004] Although theoretically, the global structural information of graph data can be obtained by stacking network layers, due to the restriction of the over-smoothing problem, the construction of deep graph neural networks poses challenges. Therefore, how to construct a high-performance deep graph neural network is an urgent problem to be solved. Summary of the Invention
[0005] This specification provides a method, apparatus, storage medium, and electronic device for constructing a high-depth graph neural network based on energy constraint to at least partially solve the above problems existing in the prior art.
[0006] This specification adopts the following technical solutions:
[0007] This specification provides a method for constructing a high-depth graph neural network based on energy constraint, including:
[0008] Obtain a to-be-determined graph neural network, where the to-be-determined graph neural network includes multiple residual layers, the residual layer includes a graph convolutional sublayer and a first combination sublayer, the residual layer performs a linear combination of the output of the graph convolutional sublayer and the output of the first combination sublayer to obtain the output of the residual layer, and the first combination sublayer integrates the outputs of the previous residual layers;
[0009] Set the learnable parameter matrix included in the residual layer as the identity matrix;
[0010] Determine the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit;
[0011] Determine the values of the combination coefficients within the range to obtain the target graph neural network.
[0012] Optionally, the first combinatorial sub-layer includes a projection sub-layer and a second combinatorial sub-layer. After projecting the output of the previous residual layer, the projection sub-layer performs a linear combination, and the second combinatorial sub-layer performs a linear combination on the output of the previous residual layer.
[0013] Optionally, determining the range of the combination coefficients included in the linear combination specifically includes:
[0014] Restricting one of the combination coefficients in the combination coefficients of the projection sub-layer and the combination coefficients of the second combinatorial sub-layer to be 0, and determining the range of the combination coefficients included in the linear combination.
[0015] Optionally, the graph convolutional sub-layer includes a learnable parameter matrix.
[0016] Optionally, the graph convolutional sub-layer and the projection sub-layer include a learnable parameter matrix.
[0017] Optionally, the parameter matrices of the graph convolutional sub-layer and the projection sub-layer share parameters.
[0018] Optionally, the method further includes:
[0019] Training the target graph neural network specifically includes:
[0020] Obtaining sample graph data and label data of the sample graph data;
[0021] Inputting the sample graph data into the target graph neural network to obtain a prediction result output by the target graph neural network;
[0022] Determining a cross-entropy loss according to the difference between the prediction result and the label data; determining a unit constraint loss according to the difference between the parameter matrix of each residual layer of the graph neural network and the identity matrix;
[0023] Taking the minimization of the sum of the cross-entropy loss and the unit constraint loss as the objective, and adjusting the parameter matrices of each residual layer of the target graph neural network.
[0024] This specification provides a high-depth graph neural network construction device based on energy constraint. The device includes:
[0025] An acquisition module that acquires a to-be-determined graph neural network, where the to-be-determined graph neural network includes a plurality of residual layers, a graph convolutional sub-layer, and a combined residual layer. The residual layer is used to perform a linear combination on the output of the graph convolutional sub-layer and the output of the combined residual layer to obtain the residual output of the residual layer;
[0026] An assumption module that sets the learnable parameter matrix included in the residual layer to an identity matrix;
[0027] The inference and prediction module determines the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit;
[0028] The determination module determines the values of the combination coefficients within the range to obtain a target graph neural network.
[0029] This specification provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above method for constructing a high-depth graph neural network based on energy constraint.
[0030] This specification provides an electronic device including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above method for constructing a high-depth graph neural network based on energy constraint.
[0031] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0032] In the method for constructing a high-depth graph neural network based on energy constraint provided in this specification, a to-be-determined graph neural network is obtained. The to-be-determined graph neural network includes multiple residual layers. The residual layer performs a linear combination of the output of the graph convolutional sublayer and the output of the first combination sublayer to obtain the output of the residual layer. The first combination sublayer integrates the output of the previous residual layer. The learnable parameter matrix included in the residual layer is set as the identity matrix. By constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit, the range of each combination coefficient included in the linear combination is determined. Within this range, the values of the combination coefficients are determined to obtain a target graph neural network.
[0033] In this method, by restricting the value range of the Dirichlet energy, the node smoothness of the graph data embedding features obtained by the graph neural network is constrained. Also, by restricting the parameter matrix to be the identity matrix, the unknown combination coefficients can be calculated according to the constraint of the Dirichlet energy to obtain combination coefficients that meet the specified Dirichlet energy range. There are two advantages of the present invention. One is that the performance of the graph neural network increases with the increase in depth, that is, the over-smoothing problem of the graph neural network is solved. The other is that a large number of high-depth graph neural networks can be generated. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The drawings described herein are used to provide a further understanding of this specification and form a part of this specification. The schematic embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0035] Figure 1 It is a schematic flowchart of a method for constructing a high-depth graph neural network based on energy constraint in this specification;
[0036] Figure 2 It is a schematic structural diagram of a to-be-determined graph neural network residual layer provided in an embodiment of this specification;
[0037] Figure 3 It is a schematic network structure diagram of a first combinatorial sub-layer provided in an embodiment of this specification;
[0038] Figure 4 It is a schematic diagram of a device for constructing a high-depth graph neural network based on energy constraint provided in this specification;
[0039] Figure 5 It is corresponding to Figure 1 in this specification and is a schematic diagram of an electronic device. Specific embodiments
[0040] To make the purpose, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of this application.
[0041] Regarding the over-smoothing problem of graph neural networks, there is currently a method of adding a residual connection structure to the graph neural network, setting the network layer of the graph neural network as a residual layer, and making the shallow features jump-connect to the deep layer through the residual layer, so that the node embedding of the graph data output by each network layer does not completely depend on the neighbor aggregation result, which can alleviate the over-smoothing phenomenon to a certain extent.
[0042] Generally, the effective depth of a traditional graph neural network (Graph Convolutional Networks, GCN) is 2 - 3 layers. For a graph neural network with a residual connection structure, its effective depth can reach dozens of layers. Although, compared with the traditional graph neural network, the residual connection structure is helpful for increasing the network depth to a certain extent, it still has limitations. How to construct a deeper graph neural network to obtain node embeddings that more fully integrate the global structure of graph data is still a challenging problem at present.
[0043] To quantify and analyze the over-smoothing problem of graphs, Dirichlet energy is proposed as an effective tool. Based on this, researchers have proposed an energy-constrained learning method and constructed a deep graph neural network whose performance improves with the increase of model depth (see Zhou, Kaixiong, Xiao Huang, Daochen Zha, Rui Chen, Li Li, Soo-Hyun Choi, and Xia Hu. Dirichlet energy constrained learning for deep graph neural networks. Advances in Neural Information Processing Systems, 34:21834–21846, 2021). However, the limitation of this method is that it is unable to design high-depth graph neural networks on a large scale, which restricts its universality in practical applications.
[0044] In graph neural networks, Dirichlet energy is a metric for measuring the smoothness of node embeddings in graph data. The smaller the value of Dirichlet energy, the smaller the difference in node embeddings of graph data and the higher the smoothness. As the number of layers of the graph neural network approaches infinity, the Dirichlet energy of graph data will converge to zero, and the node embeddings of different nodes will be relatively close, making it difficult to distinguish different nodes through node embeddings and resulting in a performance decline in classification tasks.
[0045] Given the node embedding matrix of a graph data G with n nodes , the Dirichlet energy of this graph data can be calculated according to the following formula:
[0046]
[0047] where represents the -th element in the adjacency matrix A of the graph data G, and is the degree of the i-th node in the graph data G.
[0048] Based on the above view, if the Dirichlet energy of the node embeddings of graph data can be constrained within a certain range during the training of the neural network, the over-smoothing problem of the graph neural network can be prevented.
[0049] In GCN, the node embeddings of graph data are essentially the aggregation of the features of the neighbor nodes of the graph data. In each network layer, the node aggregation process of the graph data G can be recursively represented by propagation as follows:
[0050]
[0051] where represents the propagation matrix, , represents the degree matrix of the graph data G, represents the normalized adjacency matrix with self-loops, , A represents the adjacency matrix of the graph data G, I represents the identity matrix, is the learnable parameter matrix of the k-th network layer, represents the activation function, represents the node embedding matrix output by the k-th network layer.
[0052] Because in the network structure design stage of the graph neural network, the original input of the graph data is unknown, it is impossible to determine the specific values of the node embedding matrices of each network layer, and apply the above calculation formula of the Dirichlet energy to determine the Dirichlet energy of each network layer. At the same time, due to the existence of the parameter matrix, the calculation difficulty is further increased.
[0053] The following combines the attached drawings to detail the technical solutions provided by each embodiment of this specification.
[0054] Figure 1 is a schematic flowchart of a method for constructing a graph neural network based on energy constraint in this specification, specifically including the following steps:
[0055] S100: Obtain a to-be-determined graph neural network, the to-be-determined graph neural network includes multiple residual layers, the residual layer includes a graph convolutional sub-layer and a first combination sub-layer, the residual layer makes a linear combination of the output of the graph convolutional sub-layer and the output of the first combination sub-layer to obtain the output of the residual layer, and the first combination sub-layer integrates the output of the previous residual layer.
[0056] In this specification, the device for constructing a graph neural network based on energy constraint can be a server or an electronic device such as a desktop computer or a laptop computer. For the convenience of description, only the server is used as the execution subject below to illustrate the method for constructing a deep graph neural network based on energy constraint provided by this specification.
[0057] Generally, a residual layer with a residual connection structure can jump-connect the input of the previous layer to the current residual layer. Then a residual layer will have multiple input data, so that the output of the graph neural network with a residual connection structure depends not only on the operations of the current layer but also on the operations of the previous historical layers.
[0058] In the existing residual layer settings, generally, the internal output obtained from the internal operations within the current residual layer in the forward propagation is directly added to the output of the previous layer to obtain the output of the current residual layer. The output of the residual layer is the data that needs to be input to the next residual layer in the forward propagation.
[0059] In this method, by setting the combination coefficients, the various inputs of the residual layer are linearly combined, so that the magnitude of the combination coefficients determines the influence degree of each input data on the final output of the residual layer.
[0060] Determining the appropriate values of the combination coefficients is the problem to be solved in the subsequent steps of this specification. In this step, the combination coefficients required for the linear combination are unknown data.
[0061] Figure 2 It is a schematic structural diagram of a to-be-determined graph neural network residual layer provided in the embodiments of this specification. In Figure 2 In represents the output of the nth residual layer, represents the linear combination operation. As Figure 2 shown, the residual layer of the to-be-determined graph neural network includes a graph convolutional layer and a first combination sublayer.
[0062] Combined with Figure 2 , for the nth residual layer, the graph convolutional layer receives the output of the previous residual layer, performs a convolution operation on the output of the previous residual layer, and outputs the convolution result as the output of the graph convolutional sublayer. The first combination sublayer receives the outputs of the previous residual layers: , that is, the outputs from the 0th residual layer to the (n - 1)th residual layer. Then, the first combination sublayer integrates the outputs of the previous residual layers to obtain the output of the first combination sublayer. Then, this residual layer performs a linear combination on the output of this graph convolutional sublayer and the output of this first combination sublayer to obtain the output
[0063] of this residual layer.
[0064] In the method for constructing a high-depth graph neural network provided in this specification, the first combination sublayer integrates the outputs of the previous residual layers. Here, the "previous residual layers" can be all the residual layers before the current residual layer, or can be the specified residual layers before the current residual layer. This specification does not limit which specific residual layer is specified. Among them, the integration operation can include a linear combination operation, or can include a projection operation and a linear combination operation.
[0064] When the previous residual layers are all the residual layers before the current residual layer, the output of the nth residual layer can be expressed as . Among them, represents the combination coefficient corresponding to the output of the ith residual layer, represents the convolution operation performed by the graph convolutional sublayer of the nth residual layer, is the learnable parameter matrix in the convolution operation of the nth residual layer.
[0065] The current residual layer linearly combines the intra-layer operations of the current residual layer and the output of the previous residual layer through combination coefficients to obtain the output of the current residual layer.
[0066] S102: Set the learnable parameter matrix included in the residual layer as the identity matrix.
[0067] In this specification, in order to simplify the calculation formula of the Dirichlet energy, the learnable parameter matrix included in the residual layer is set as the identity matrix to eliminate the computational complexity increased by multiplying the parameter matrix.
[0068] In this step, during the inference calculation of the Dirichlet energy, the learnable parameter matrix is assumed to be the identity matrix and substituted into the calculation. After determining the target graph neural network, the parameter matrix can be adjusted to be close to the identity matrix during the training process to achieve this assumption.
[0069] In different network structures, the parameter matrices included in the residual layer are different.
[0070] In one embodiment, the graph convolutional layer of the residual layer includes a learnable parameter matrix. Then, the learnable parameter matrix of the graph convolutional layer can be set as the identity matrix.
[0071] S104: Determine the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit.
[0072] The smaller the value of the Dirichlet energy, the smoother the node embedding. If the value of the Dirichlet energy is too large, it means that the node embeddings of different nodes are too separated, which will also have a negative impact on the classification task. Therefore, it is necessary to limit the Dirichlet energy within a suitable range.
[0073] The restricted range of the Dirichlet energy can be specified as needed, that is, the required specified energy upper limit and specified energy lower limit can be freely selected. Then, the specified energy upper limit and specified energy lower limit are substituted into the calculation formula of the Dirichlet energy to deduce the range of each combination coefficient.
[0074] The following explains the specific derivation of the combination coefficients in this specification.
[0075] When the network structure of a graph neural network is determined, that is, the operations to be performed by each residual layer and the order of the operations are fixed, the data recurrence relationship between the input and output of this residual layer is also determined accordingly, which means that the propagation recurrence formula of this residual layer is determined.
[0076] The message propagation mechanism in GCN can be expressed as , that is Based on the calculation formula of the Dirichlet energy, the energy recurrence relationship between residual layers can be deduced as follows:
[0077]
[0078]
[0079] Where, represents the Dirichlet energy of the residual output of the k-th residual layer, that is, the Dirichlet energy of the output feature map of the k-th residual layer obtained through the operation of the to-be-determined graph neural network. is the augmented normalized Laplacian operator the eigenvalue closest to 1 in is the eigenvalue closest to 0 in is the augmented normalized Laplacian operator, . is the square of the smallest singular value of is the square of the largest singular value of
[0080] By adjusting the parameter matrix during the training process, can be restricted within a range close to the identity matrix, that is, is achieved, where the eigenvalues of are all less than , is a sufficiently small constant.
[0081] The current residual layer linearly combines the intra-layer operation of the current residual layer and the output of the previous residual layer through the combination coefficients to obtain the output of the current residual layer.
[0082] Then the process of obtaining the output of the n-th residual layer can be expressed as: . Where, represents the combination coefficients used.
[0083] By setting the parameter matrix to the identity matrix, that is, considering that the matrix in is small enough, the influence of the matrix on the residual output of the residual layer can be ignored. At this time, .
[0084] Based on this, the expressions of the upper and lower bounds of the Dirichlet energy of the graph neural network can be further deduced as follows:
[0085]
[0086]
[0087] Among them, is the k-th combination coefficient, is the propagation matrix and its minimum eigenvalue ranges between and
[0088] After obtaining the expressions for the upper and lower bounds of the Dirichlet energy, by specifying the upper bound of the energy and the lower bound of the energy, substituting them into the above upper and lower bounds of the energy, according to the constraint range of the value range of
[0089] For example, the specified upper bound of the energy can be set to 100 and the specified lower bound of the energy can be set to 0.01 to calculate the pending graph neural network that meets the constraint conditions.
[0090] In this method, by constraining the lower bound of the Dirichlet energy, it is ensured that it is not overly smooth, and at the same time, by setting the upper bound of the Dirichlet energy, it is ensured that the node embeddings do not separate excessively.
[0091] S106: In the said range, determine the values of the respective combination coefficients to obtain the target graph neural network.
[0092] In this step S104, only the value ranges of the respective combination coefficients are calculated. In the value range of each coefficient, the specific values of the combination coefficients can be arbitrarily selected.
[0093] For example, if there are two combination coefficients, and the range of combination coefficient A is calculated as , and the range of combination coefficient B is , then combination coefficient A can take a value of in , and combination coefficient B can take any value of in . The combination of and
[0094] can be represented as a value set, and this value set contains countless possibilities.
[0095] After determining the specific values of each combination coefficient, only the learnable parameter matrix in the propagation recurrence formula of the graph neural network is unknown, and the propagation recurrence formula of the graph neural network is completely determined, indicating that the construction of the graph neural network is completed. Based on this propagation recurrence formula, training and inference tasks can be carried out.
[0096] The application of the constructed target graph neural network in this specification is not limited and can be used for user profiling, intrusion detection, knowledge graph classification, etc.
[0097] Based on the above Figure 1 The graph neural network construction method based on energy constraint shown above is used to obtain a to-be-determined graph neural network. The to-be-determined graph neural network includes multiple residual layers. The residual layer includes a graph convolution sublayer and a first combination sublayer. The residual layer performs a linear combination of the output of the graph convolution sublayer and the output of the first combination sublayer to obtain the output of the residual layer. The first combination sublayer integrates the outputs of the previous residual layers, sets the learnable parameter matrix included in the residual layer as the identity matrix, determines the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit, and determines the values of each combination coefficient within this range to obtain the target graph neural network.
[0098] In this method, by restricting the value range of the Dirichlet energy, the node smoothness of the graph data embedding features obtained by the graph neural network is constrained. And by restricting the parameter matrix to be the identity matrix, the unknown combination coefficients can be calculated according to the constraint of the Dirichlet energy to obtain the combination coefficients that meet the specified Dirichlet energy range. There are two advantages of the present invention. One is that the performance of the graph neural network increases with the increase of depth, that is, the over-smoothing problem of the graph neural network is solved. The other is that a large number of high-depth graph neural networks can be generated.
[0099] According to the deep neural network construction method provided in this specification, a 128-layer deep graph neural network can be constructed. After being trained, the deep graph neural network can be used for tasks such as generating portraits, network attack detection, and academic knowledge graph subject classification.
[0100] In an embodiment of this specification, the first combination sublayer of the to-be-determined graph neural network may further include a projection sublayer and a second combination sublayer. After the projection sublayer projects the output of the previous residual layer, it performs a linear combination of the projection results of the outputs of each previous residual layer, and the second combination sublayer performs a linear combination of the outputs of the previous residual layers.
[0101] Figure 3 This is a schematic diagram of the network structure of a first combination sublayer provided in the embodiment of this specification, where represents the output of the (n - 1)-th residual layer, represents the linear combination operation. AsFigure 3 As shown in Figure 3 , the first combinatorial layer includes a projection layer and a second combinatorial layer. The projection layer receives the output of the previous residual layer: , performs a projection operation on each received output, and performs a linear combination on each projection result. The second combinatorial layer receives the output of the previous residual layer: , and directly performs a linear combination on the output of the previous residual layer.
[0102] Then, in this embodiment, in the linear combination operations of the projection layer and the second combinatorial layer, combination coefficients corresponding to the outputs of each previous residual layer are set.
[0103] In the projection operation in this embodiment, for the output of a certain previous residual layer, it is to multiply the output of the previous residual layer by a projection matrix. This projection matrix can be set as a learnable parameter matrix. Then, in the network structure of the to-be-determined graph neural network in this embodiment, the graph convolutional layer and the projection layer contain learnable parameter matrices.
[0104] Let represent the learnable parameter matrix in the projection layer of the nth residual layer. Then, in the projection layer of the nth residual layer, the projection operation performed on the output of the previous residual layer can be expressed as: .
[0105] Then, in this embodiment, two learnable parameter matrices corresponding to the graph convolutional layer and the projection layer are included. During the training process, these two learnable parameter matrices can be adjusted separately. Or, in order to improve the training efficiency, the parameter matrices of the graph convolutional layer and the projection layer can be set to share parameters, and these two parameter matrices can be adjusted synchronously.
[0106] In Figure 3 the network structure shown in Figure 3 , the functions of the projection layer and the second combinatorial layer are both linear combinations of the outputs of the previous residual layer. In an embodiment of this specification, by setting the combination coefficient corresponding to the projection layer to 0, or the coefficient corresponding to the second combinatorial layer to 0, it can be controlled that only one linear combination operation is retained in the first combinatorial layer.
[0107] That is, in the above step S104, when calculating the range of each combination coefficient, it can be restricted that one of the combination coefficients of the projection layer and the combination coefficient of the second combinatorial layer is 0 to calculate the range of the combination coefficients included in the linear combination.
[0108] When the combination coefficient of the second combinatorial layer is restricted to 0, it is equivalent that the operation of the second combinatorial layer has no influence on the output of the entire residual layer. The output of the nth residual layer can be expressed as: .
[0109] When the combination coefficient of the projection sublayer is 0, it is equivalent to the operation of the projection sublayer having no effect on the output of the entire residual layer. The output of the nth residual layer can be expressed as: . It can be found that in this case, the network structure of the first combination sublayer including the projection sublayer and the second combination sublayer in this embodiment is equivalent to the network structure of the separate first combination sublayer described in the above step S100.
[0110] In the assumption of S102 above, the parameter matrix is set to the identity matrix. Then, correspondingly, during the training process of the graph neural network in this specification, it is necessary to ensure that the parameter matrix is close to the identity matrix.
[0111] The training method is as follows:
[0112] First, the server obtains the sample graph data and the label data of the sample graph data. The sample graph data is input into the target graph neural network determined in step S104, and the prediction result output by the target graph neural network is obtained.
[0113] Then, the server determines the cross-entropy loss according to the difference between the prediction result and the label data, and determines the unit constraint loss according to the difference between the parameter matrix of each residual layer of the graph neural network and the identity matrix. With the goal of minimizing the sum of the cross-entropy loss and the unit constraint loss, the target graph neural network is trained to adjust the parameter matrix of each residual layer of the graph neural network.
[0114] Among them, the setting of the unit constraint loss can ensure that in the graph neural network trained by this method, the parameter matrix is close to the identity matrix.
[0115] In the above step S100, when the previous residual layer is the specified residual layer before the current residual layer, several optional network structures of the to-be-determined graph neural network are provided below.
[0116] In one embodiment, the specified residual layer can be the previous residual layer of the current residual layer. The residual layer of the to-be-determined graph neural network includes a graph convolutional sublayer, a projection sublayer, and a second combination sublayer. The propagation recurrence formula of the to-be-determined graph neural network can be expressed as follows:
[0117]
[0118]
[0119]
[0120] Among them, is the original input data of the first residual layer, represents the intermediate output of the nth residual layer, is the learnable parameter matrix in the convolutional sublayer, is the learnable parameter matrix in the projection sub-layer, represents the propagation matrix, represents the residual output of the nth residual layer, represents the activation function. 、 、 are the combination coefficients.
[0121] Based on this embodiment, in the above step S104, determining the value range of each combination coefficient can specifically be to determine the value range of 、 、 .
[0122] For example, in this embodiment, it can be calculated that , .
[0123] In this embodiment, the convolutional sub-layer and the projection sub-layer can share the parameter matrix to reduce the computational amount during training and improve the training efficiency. When the convolutional sub-layer and the projection sub-layer share the parameter matrix, the above propagation recurrence formula can be abbreviated as:
[0124]
[0125]
[0126] where is the parameter matrix shared by the convolutional sub-layer and the projection sub-layer in the nth residual layer.
[0127] An example of the selection of a combination coefficient is to take , , .
[0128] At this time, the propagation recurrence formula of the graph neural network obtained can be specifically:
[0129]
[0130]
[0131] In another embodiment of this specification, the specified residual layer can be all even residual layers before the current residual layer. An optional residual layer of the graph neural network includes a graph convolutional sub-layer and a first combination sub-layer. In the convolutional sub-layer, a learnable parameter matrix is included. The propagation recurrence formula of this graph neural network is as follows:
[0132]
[0133]
[0134]
[0135] Among them, is the original input data of the first residual layer, represents the intermediate output of the nth residual layer, is a learnable parameter matrix, represents the propagation matrix, represents the output of the nth residual layer, represents the activation function, represents the output of the even residual layer, represents the linear combination of all even residual layers before the current residual layer. is the combination coefficient to be solved in this embodiment.
[0136] Based on this embodiment, in the above step S104, to determine the value range of the combination coefficient, it can be specifically to determine the value.
[0137] For example, in this embodiment, it can be calculated that the range is . Then in step S106, can be selected.
[0138] At this time, the propagation recurrence formula of the graph neural network can be specifically:
[0139]
[0140]
[0141]
[0142] In another embodiment of this specification, the specified residual layer can be the input of the first residual layer (the output of the 0th residual layer) and the previous residual layer of the current residual layer. An optional residual layer of the graph neural network includes a convolutional sublayer and a first combination sublayer. In the convolutional sublayer, a learnable parameter matrix is included. The propagation recurrence formula of this graph neural network is as follows:
[0143]
[0144]
[0145] Among them, is the original input data of the first residual layer, is a learnable parameter matrix, represents the propagation matrix, represents the output of the nth residual layer, represents the propagation matrix, represents an activation function. , , are combination coefficients.
[0146] Based on this embodiment, in the above step S104, determining the value ranges of the respective combination coefficients can specifically be to determine the value ranges of , , .
[0147] For example, in this embodiment, it can be calculated that , , .
[0148] Then an example of the selection of a combination coefficient is to take , , . At this time, the propagation recurrence formula of the graph neural network can be specifically obtained as:
[0149]
[0150]
[0151] In one embodiment, the target graph neural network constructed in this specification can be applied to predicting user portraits. A user portrait is the result of analyzing information such as a user's behavior patterns, interest preferences, and social relationships, extracting the user's key attributes (such as interest tags, occupation categories), and classifying them.
[0152] In this embodiment, the user information that can be extracted from a social platform website includes static attributes such as the user's gender, age, geographical location, and dynamic behavior attributes such as browsing history, interaction frequency, and activity participation. And a social network graph is constructed with users as nodes and the user's static attributes and dynamic behavior attributes as edges.
[0153] Input the social network graph of the user whose user portrait is to be determined into the trained target graph neural network, obtain the probabilities corresponding to the classification of the user into multiple user groups output by the target graph neural network, and the user group classification corresponding to the highest probability can be used as the user portrait of this user.
[0154] In one embodiment, the target graph neural network constructed in this specification can be applied to network attack detection. The goal of network attack detection is to identify potential abnormal behaviors and malicious activities from complex network traffic data, such as Distributed Denial of Service (DDoS) attacks, data leakage, and network intrusion.
[0155] The network monitoring system continuously collects real-time traffic data, including features such as communication frequency, access time, and packet size, constructs a graph, and inputs it into the target graph neural network. The traffic data is constructed into an undirected graph. Each node in the undirected graph represents an IP address or a host, and each edge represents the communication behavior between two nodes. The node features include static information of the host such as IP type, location, and dynamic behavior features such as the number of traffic packets, connection frequency, and access time distribution.
[0156] Input the undirected graph corresponding to the constructed traffic data into the trained target graph neural network to obtain the anomaly probability of each node. The nodes with probabilities higher than the threshold are regarded as risk nodes, and these risk nodes may be the nodes where attacks occur.
[0157] In one embodiment, the target graph neural network constructed in this specification can be applied to subject classification based on an academic knowledge graph. An academic knowledge graph is a graph-structured data composed of academic papers, authors, institutions, and the citation relationships between them. The nodes represent papers, and the edges represent citation relationships or the association relationships between authors and papers. The task objective is to predict the subject field to which each paper belongs, which may include computer science, physics, biology, mathematics, chemistry, etc.
[0158] The academic knowledge graph can be obtained from existing public academic datasets (such as the ArXiv dataset, Microsoft Academic Graph, etc.). In the academic knowledge graph, the nodes represent academic papers, which contain attributes such as title, keywords, abstract, publication year, etc. The edges represent the citation relationships between papers or the association relationships between papers and authors. The node features are to generate vector features by embedding text attributes (such as title, abstract); numerical encoding is performed on non-text attributes (such as publication year).
[0159] Input the academic knowledge graph to be classified into the trained target graph neural network to obtain the probabilities of the academic knowledge graph belonging to each subject field, and select the subject field corresponding to the maximum probability as the subject field of the academic knowledge graph to be classified.
[0160] The above is the method for constructing a graph neural network based on energy constraint provided in this specification. Based on the same idea, this specification also provides a corresponding device for constructing a graph neural network based on energy constraint, as Figure 5 shown.
[0161] Figure 4 It is a schematic diagram of a device for constructing a high-depth graph neural network based on energy constraint provided in this specification, specifically including:
[0162] An acquisition module 200, configured to acquire a to-be-determined graph neural network, where the to-be-determined graph neural network includes a plurality of residual layers, the residual layer includes a graph convolution sub-layer and a first combination sub-layer, the residual layer performs a linear combination on the output of the graph convolution sub-layer and the output of the first combination sub-layer to obtain the output of the residual layer, and the first combination sub-layer integrates the output of the previous residual layer;
[0163] A hypothesis module 202, configured to set the learnable parameter matrix included in the residual layer as an identity matrix;
[0164] An inference and prediction module 204, configured to determine the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit;
[0165] A determination module 206, configured to determine the values of the combination coefficients in the range to obtain a target graph neural network.
[0166] Optionally, the first combination sub-layer includes a projection sub-layer and a second combination sub-layer. After the projection sub-layer projects the output of the previous residual layer, a linear combination is performed, and the second combination sub-layer performs a linear combination on the output of the previous residual layer.
[0167] Optionally, the inference and prediction module 204 is specifically configured to limit one of the combination coefficients of the projection sub-layer and the combination coefficients of the second combination sub-layer to be 0, and determine the range of the combination coefficients included in the linear combination.
[0168] Optionally, the graph convolution sub-layer includes a learnable parameter matrix.
[0169] Optionally, the graph convolution sub-layer and the projection sub-layer include learnable parameter matrices.
[0170] Optionally, the parameter matrices of the graph convolution sub-layer and the projection sub-layer share parameters.
[0171] Optionally, the apparatus further includes a training module 208;
[0172] The training module 208 is configured to acquire sample graph data and label data of the sample graph data, input the sample graph data into the target graph neural network to obtain a prediction result output by the target graph neural network, determine a cross-entropy loss according to the difference between the prediction result and the label data; determine a unit constraint loss according to the difference between the parameter matrix of each residual layer of the graph neural network and the identity matrix, and adjust the parameter matrices of each residual layer of the target graph neural network with the goal of minimizing the sum of the cross-entropy loss and the unit constraint loss.
[0173] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-mentioned Figure 1 method for constructing a high-depth graph neural network based on energy constraint provided.
[0174] This specification also provides Figure 5 a schematic structural diagram of the electronic device shown. As Figure 5 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned Figure 1 method for constructing a graph neural network based on energy constraint. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logical devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logical device.
[0175] Improvements to a technology can be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. The designer can program by themselves to "integrate" a digital system on a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0176] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0177] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0178] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0179] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0180] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0181] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0182] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0183] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0184] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0185] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0186] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0187] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, system, or computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0188] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0189] Each embodiment in this specification is described in a progressive manner. For the identical or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.
[0190] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this application.
Claims
1. A method for constructing a high-depth graph neural network based on energy constraints, characterized in that Including: Obtain a to-be-determined graph neural network for network attack detection. The to-be-determined graph neural network includes multiple residual layers. Each residual layer contains a graph convolutional sublayer and a first combination sublayer. The residual layer performs a linear combination of the output of the graph convolutional sublayer and the output of the first combination sublayer to obtain the output of the residual layer. The first combination sublayer integrates the outputs of previous residual layers. The first combination sublayer includes a projection sublayer and a second combination sublayer. The projection sublayer projects the output of the previous residual layer and then performs a linear combination. The second combination sublayer performs a linear combination of the output of the previous residual layer. Set the learnable parameter matrix included in the residual layer as an identity matrix. Determine the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit; wherein, the specified energy upper limit is the square of the sum of all combination coefficients ; the specified energy lower limit is the square of the sum of the products of all combination coefficients and the minimum eigenvalue of the propagation matrix to the k power, and the range of is between (-1, 1], k = 0, 1, …, n, where n is the number of residual layers; In the range, determine the values of the combination coefficients to obtain a target graph neural network. Wherein, the trained target neural network is used to identify abnormal behaviors and malicious activities from network traffic data.
2. The method according to claim 1, characterized in that, Determine the range of the combination coefficients included in the linear combination, specifically including: Restrict one of the combination coefficients of the projection sublayer and the combination coefficient of the second combination sublayer to be 0, and determine the range of the combination coefficients included in the linear combination.
3. The method according to claim 1, wherein The graph convolutional sublayer includes a learnable parameter matrix.
4. The method according to claim 1, characterized in that, The graph convolutional sublayer and the projection sublayer include learnable parameter matrices.
5. The method according to claim 4, characterized in that, The parameter matrices of the graph convolutional sublayer and the projection sublayer share parameters.
6. The method according to claim 1, wherein The method further includes: Train the target graph neural network, specifically including: Obtain sample graph data and label data of the sample graph data. Input the sample graph data into the target graph neural network to obtain a prediction result output by the target graph neural network. Determine the cross-entropy loss according to the difference between the prediction result and the label data. Determine the unit constraint loss according to the difference between the parameter matrix of each residual layer of the graph neural network and the identity matrix. With the goal of minimizing the sum of the cross-entropy loss and the unit constraint loss, adjust the parameter matrices of each residual layer of the target graph neural network.
7. An apparatus for constructing a high-depth graph neural network based on energy constraint, characterized in that Including: An acquisition module that acquires a to-be-determined graph neural network for network attack detection. The to-be-determined graph neural network includes multiple residual layers. Each residual layer contains a graph convolutional sublayer and a first combination sublayer. The residual layer performs a linear combination of the output of the graph convolutional sublayer and the output of the first combination sublayer to obtain the output of the residual layer. The first combination sublayer integrates the outputs of previous residual layers. The first combination sublayer includes a projection sublayer and a second combination sublayer. The projection sublayer projects the output of the previous residual layer and then performs a linear combination. The second combination sublayer performs a linear combination of the output of the previous residual layer. A hypothesis module that sets the learnable parameter matrix included in the residual layer as an identity matrix. The inference and prediction module determines the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit; wherein, the specified energy upper limit is the square of the sum of all combination coefficients ; the specified energy lower limit is the square of the sum of the products of all combination coefficients and the minimum eigenvalue of the propagation matrix to the k power, and the range of is between (-1, 1], k = 0, 1, …, n, where n is the number of residual layers; A determination module that determines the values of the combination coefficients in the range to obtain a target graph neural network. Wherein, the trained target neural network is used to identify abnormal behaviors and malicious activities from network traffic data.
8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 6 above is implemented.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1 to 6 above is implemented.