High-depth map neural network construction method and device based on energy constraint

By introducing energy constraint-based methods into graph neural networks, the oversmoothing problem in deep graph neural networks is solved, network performance is improved, and the design of large-scale high-deep graph neural networks is realized.

CN119990184AActive Publication Date: 2025-05-13ZHEJIANG LAB
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510461285.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing technology faces too smoothing problems when building deep graph neural networks, resulting in the node embedding tending to be consistent and the different nodes cannot be distinguished, which in turn affects the performance of node classification tasks.

Method used

By introducing an energy constraint-based method into the graph neural network, the specific steps include obtaining the pending graph neural network, setting the learnable parameter matrix of the residual layer as a unit matrix, and determining the range of the combined coefficients through the Dirichlet energy constraint, and finally obtaining the target graph neural network.

Benefits of technology

It effectively solves the problem of oversmoothing of graph neural networks, improves the performance of deep networks, and enables graph neural networks to more fully integrate the global structural information of graph data, and can design high-deep graph neural networks on a large scale.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990184A_ABST
    Figure CN119990184A_ABST
Patent Text Reader

Abstract

The invention discloses a high-depth map neural network construction method and device based on energy constraint, and the method comprises the steps: obtaining a to-be-determined map neural network which comprises a plurality of residual layers, each residual layer comprises a map convolution sub-layer and a first combination sub-layer, and enabling the residual layers to carry out the linear combination of the output of the map convolution sub-layer and the first combination sub-layer, the first combination sub-layer linearly combines projections of one and the other of two sets of linear combinations of the outputs of the respective residual layers before. According to the method, a learnable parameter matrix of a residual layer is set as a unit matrix, Dirichlet energy of the residual layer is constrained in a specified range, each combination coefficient is determined, and a target map neural network is obtained. By constructing the model, the method can be used for generation of figure portraits, network attack detection, academic knowledge graph subject classification and the like. The method has the advantages that the over-smoothing problem of the graph neural network is solved, and a large batch of high-depth graph neural networks can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for constructing a high-depth graph neural network based on energy constraints. Background Art

[0002] Graph neural networks have an over-smoothing problem, that is, as the number of network layers increases, the node embeddings of graph data tend to be consistent, and different nodes cannot be distinguished, resulting in a decrease in its performance in node classification tasks.

[0003] Shallow graph neural networks capture local structural information of graph data. The deeper the graph neural network, the more global the features captured. In feature extraction tasks of complex graph data such as chemical molecular structures, knowledge graphs, and social network graphs, it is often necessary to obtain global structural information of graph data in order to achieve better performance in node classification tasks.

[0004] Although in theory the global structural information of graph data can be obtained by stacking network layers, the construction of deep graph neural networks is challenging due to the over-smoothing problem. Therefore, how to build a high-performance deep graph neural network is an urgent problem to be solved. Summary of the invention

[0005] This specification provides a method, device, storage medium and electronic device for constructing a high-depth graph neural network based on energy constraints to at least partially solve the above-mentioned problems existing in the prior art.

[0006] This manual adopts the following technical solutions: This specification provides a method for constructing a high-depth graph neural network based on energy constraints, including: Obtain a pending graph neural network, wherein the pending graph neural network includes a plurality of residual layers, wherein the residual layer includes a graph convolution sublayer and a first combination sublayer, wherein the residual layer linearly combines the output of the graph convolution sublayer with the output of the first combination sublayer to obtain the output of the residual layer, and the first combination sublayer integrates the output of the previous residual layer; Setting the learnable parameter matrix contained in the residual layer to the unit matrix; Determining the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit; Within the range, the values ​​of the combination coefficients are determined to obtain the target graph neural network.

[0007] Optionally, the first combination sublayer includes a projection sublayer and a second combination sublayer, the projection sublayer projects the output of the previous residual layer and performs linear combination, and the second combination sublayer performs linear combination on the output of the previous residual layer.

[0008] Optionally, determining a range of combination coefficients included in the linear combination specifically includes: The combination coefficients of the projection sublayer and the second combination sublayer are restricted, one of the combination coefficients being 0, and the range of the combination coefficients included in the linear combination is determined.

[0009] Optionally, the graph convolution sublayer contains a learnable parameter matrix.

[0010] Optionally, the graph convolution sublayer and the projection sublayer contain learnable parameter matrices.

[0011] Optionally, parameter matrices of the graph convolution sublayer and the projection sublayer share parameters.

[0012] Optionally, the method further comprises: Training the target graph neural network specifically includes: Obtaining sample graph data and label data of the sample graph data; Inputting the sample graph data into the target graph neural network to obtain a prediction result output by the target graph neural network; Determine the cross entropy loss based on the difference between the prediction result and the label data; determine the unit constraint loss based on the difference between the parameter matrix of each residual layer of the graph neural network and the unit matrix; With the goal of minimizing the sum of the cross entropy loss and the unit constraint loss, the parameter matrices of the residual layers of the target graph neural network are adjusted.

[0013] This specification provides a device for constructing a high-depth graph neural network based on energy constraints, the device comprising: An acquisition module is used to acquire a pending graph neural network, wherein the pending graph neural network includes multiple residual layers, a graph convolution sublayer and a combined residual layer, wherein the residual layer is used to linearly combine the output of the graph convolution sublayer with the output of the combined residual layer to obtain a residual output of the residual layer; Assume that the module sets the learnable parameter matrix contained in the residual layer to the unit matrix; An inference prediction module determines the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit; A determination module determines the values ​​of each combination coefficient within the range to obtain a target graph neural network.

[0014] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for constructing a high-depth graph neural network based on energy constraints.

[0015] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for constructing a high-depth graph neural network based on energy constraints is implemented.

[0016] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects: In the energy-constrained high-depth graph neural network construction method provided in this specification, a to-be-determined graph neural network is obtained, which includes multiple residual layers, the residual layer includes a graph convolution sublayer and a first combination sublayer, the residual layer linearly combines the output of the graph convolution sublayer with the output of the first combination sublayer to obtain the output of the residual layer, the first combination sublayer integrates the output of the previous residual layer, and sets the learnable parameter matrix contained in the residual layer to the unit matrix. By constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit, the range of each combination coefficient contained in the linear combination is determined, and within this range, the value of each combination coefficient is determined to obtain the target graph neural network.

[0017] In this method, by limiting the value range of Dirichlet energy, the node smoothness of the graph data embedded features obtained by the graph neural network is constrained, and by limiting the parameter matrix to the unit matrix, the unknown combination coefficients can be calculated according to the constraints of Dirichlet energy to obtain the combination coefficients that meet the specified Dirichlet energy range. The advantages of the present invention are twofold. One is that the performance of the graph neural network increases with the increase of depth, that is, the over-smoothing problem of the graph neural network is solved. The other is that a large number of high-depth graph neural networks can be generated. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings: Figure 1 A schematic diagram of a process for constructing a high-depth graph neural network based on energy constraints in this specification; Figure 2 A schematic diagram of the structure of a residual layer of a graph neural network to be determined provided in an embodiment of this specification; Figure 3 A schematic diagram of a network structure of a first combination sublayer provided in an embodiment of this specification; Figure 4A schematic diagram of a device for constructing a high-depth graph neural network based on energy constraints provided in this specification; Figure 5 The corresponding Figure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0020] To address the over-smoothing problem of graph neural networks, there is currently a method that adds a residual connection structure to the graph neural network. The network layer of the graph neural network is set to a residual layer, and the shallow features are jump-connected to the deep layer through the residual layer. This makes the node embedding of the graph data output by each network layer not completely dependent on the neighbor aggregation result, which can alleviate the over-smoothing phenomenon to a certain extent.

[0021] Generally, the effective depth of traditional graph convolutional networks (GCN) is 2-3 layers. With the addition of residual connection structure, the effective depth of graph neural networks can reach dozens of layers. Although the residual connection structure helps to increase the network depth compared to traditional graph neural networks, it still has limitations. How to achieve deeper graph neural network construction to obtain node embedding that more fully integrates the global structure of graph data is still a challenging problem.

[0022] In order to quantify and analyze the problem of graph oversmoothing, Dirichlet energy was proposed as an effective tool. Based on this, researchers proposed an energy-constrained learning method and constructed a deep graph neural network whose performance improved with the depth of the model (see Zhou, Kaixiong, Xiao Huang, Daochen Zha, Rui Chen, Li Li, Soo-Hyun Choi, and Xia Hu. Dirichlet energy constrained learning for deep graph neural networks. Advances in Neural Information Processing Systems, 34:21834–21846, 2021). However, the limitation of this method is that it is impossible to design high-depth graph neural networks on a large scale, which limits its universality in practical applications.

[0023] In graph neural networks, Dirichlet energy is an indicator to measure the smoothness of graph data node embedding. The smaller the value of Dirichlet energy, the smaller the difference in embedding of each node of the graph data, and the higher the smoothness. As the number of layers of the graph neural network approaches infinity, the Dirichlet energy of the graph data will converge to zero, the node embeddings of different nodes will be close, and it will be difficult to distinguish different nodes through node embeddings, resulting in reduced performance in classification tasks.

[0024] Given a graph data G with n nodes, the node embedding matrix , the Dirichlet energy of the graph data It can be calculated as follows:

[0025] in, The adjacency matrix A of the graph data G represents elements, is the degree of the i-th node in the graph data G.

[0026] Based on the above viewpoints, if the Dirichlet energy embedded in the graph data nodes can be constrained within a certain range during the training of the neural network, the over-smoothing problem of the graph neural network can be prevented.

[0027] In GCN, the node embedding of graph data is essentially the aggregation of neighbor node features of graph data. At each network layer, the node aggregation process of graph data G can be expressed by propagation recursion as follows:

[0028] in, represents the propagation matrix, , Represents the degree matrix of the graph data G, represents the normalized adjacency matrix with self-loops, , A represents the adjacency matrix of the graph data G, I represents the identity matrix, is the learnable parameter matrix of the kth network layer, represents the activation function, Represents the node embedding matrix output by the kth network layer.

[0029] Because the original input of the graph data is unknown during the network structure design phase of the graph neural network, it is impossible to determine the specific values ​​of the node embedding matrix of each network layer, and to apply the above Dirichlet energy calculation formula to determine the Dirichlet energy of each network layer. At the same time, the existence of the parameter matrix further increases the difficulty of calculation.

[0030] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.

[0031] Figure 1 This is a flow chart of a method for constructing a graph neural network based on energy constraints in this specification, which specifically includes the following steps: S100: Obtain a pending graph neural network, wherein the pending graph neural network includes multiple residual layers, wherein the residual layer includes a graph convolution sublayer and a first combination sublayer, wherein the residual layer linearly combines the output of the graph convolution sublayer with the output of the first combination sublayer to obtain the output of the residual layer, and the first combination sublayer integrates the output of the previous residual layer.

[0032] In this specification, the device used to construct a graph neural network based on energy constraints can be a server or an electronic device such as a desktop computer, a laptop computer, etc. For the sake of ease of description, the following only takes the server as the execution subject to illustrate the advanced graph neural network construction method based on energy constraints provided in this specification.

[0033] Usually, a residual layer with a residual connection structure can jump-connect the input of the previous layer to the current residual layer. Then a residual layer will have multiple input data, so that the output of the graph neural network with a residual connection structure depends not only on the operation of the current layer, but also on the operation of the previous historical layer.

[0034] In the existing residual layer setting, the internal output obtained by the internal operation of the current residual layer in the forward propagation is generally directly added to the output of the previous layer to obtain the output of the current residual layer. The output of the residual layer is the data that needs to be input to the next residual layer in the forward propagation.

[0035] In this method, the inputs of the residual layer are linearly combined by setting the combination coefficient, so that the influence of each input data on the final output of the residual layer is determined by the size of the combination coefficient.

[0036] Determining the appropriate value of the combination coefficient is a problem to be solved in the subsequent steps of this specification. In this step, the combination coefficient required for the linear combination is unknown data.

[0037] Figure 2 This is a schematic diagram of the structure of a residual layer of a graph neural network to be determined in an embodiment of this specification. Figure 2 middle, represents the output of the nth residual layer, represents a linear combination operation. Figure 2 As shown, the residual layer of the proposed graph neural network includes a graph convolution layer and a first combination sublayer.

[0038] Combination Figure 2 , for the nth residual layer, the graph convolution layer receives the output of the previous residual layer , the output of the previous residual layer Perform convolution operation and output the convolution result as the output of the graph convolution sublayer. The first combination sublayer receives the output of the previous residual layers: , that is, starting from the 0th residual layer to the output of the n-1th residual layer, and then the first combination sublayer integrates the outputs of the previous residual layers to obtain the output of the first combination sublayer. Then, the residual layer linearly combines the output of the graph convolution sublayer and the output of the first combination sublayer to obtain the output of the residual layer .

[0039] In the method for constructing a high-depth graph neural network provided in this specification, the first combination sublayer integrates the output of the previous residual layer. The "previous residual layer" here can be all residual layers before the current residual layer, or it can be a specified residual layer before the current residual layer. This specification does not limit which residual layer is specifically specified. Among them, the integration operation can include a linear combination operation, and can also include a projection operation and a linear combination operation.

[0040] When the previous residual layer is all the residual layers before the current residual layer, the output of the nth residual layer can be expressed as, .in, represents the combination coefficient corresponding to the output of the i-th residual layer, represents the convolution operation performed by the graph convolution sublayer of the nth residual layer, is the learnable parameter matrix in the convolution operation of the nth residual layer.

[0041] The current residual layer linearly combines the intra-layer operation of the current residual layer with the output of the previous residual layer through the combination coefficient to obtain the output of the current residual layer.

[0042] S102: Setting the learnable parameter matrix included in the residual layer to a unit matrix.

[0043] In this specification, in order to simplify the calculation formula of the Dirichlet energy, the learnable parameter matrix contained in the residual layer is set to the unit matrix to eliminate the computational complexity increased by multiplying the parameter matrix.

[0044] This step is to assume that the learnable parameter matrix is ​​the unit matrix when calculating the Dirichlet energy. After the target graph neural network is determined, the parameter matrix can be adjusted to be close to the unit matrix during the training process to achieve this assumption.

[0045] In different network structures, the parameter matrix contained in the residual layer is different.

[0046] In one embodiment, the graph convolution layer of the residual layer includes a learnable parameter matrix. The learnable parameter matrix of the graph convolution layer can be set to a unit matrix.

[0047] S104: Determine the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit.

[0048] The smaller the value of Dirichlet energy, the smoother the node embedding. If the value of Dirichlet energy is too large, it means that the node embeddings of different nodes are too separated, which will also have a negative impact on the classification task. Therefore, it is necessary to limit the Dirichlet energy to a suitable range.

[0049] The limit range of Dirichlet energy can be specified as needed, that is, the required upper and lower limits of specified energy can be freely selected. The upper and lower limits of specified energy are brought into the calculation formula of Dirichlet energy to deduce the range of each combination coefficient.

[0050] The specific derivation of the combination coefficient in this specification is explained below.

[0051] When the network structure of a graph neural network is determined, that is, the operations that each residual layer needs to perform and the order of operations are fixed, the data recursive relationship between the input and output of the residual layer is also determined, which means that the propagation recursion of the residual layer is determined.

[0052] The message propagation mechanism in GCN can be expressed as ,Right now Based on the calculation formula of Dirichlet energy, the energy recursive relationship between each residual layer can be derived as follows:

[0053]

[0054] in, Represents the Dirichlet energy of the residual output of the k-th residual layer, that is, the Dirichlet energy of the output feature map of the k-th residual layer obtained by the operation of the undetermined graph neural network. is the augmented normalized Laplacian The eigenvalue closest to 1 in yes The eigenvalue closest to 0 is is the augmented normalized Laplace operator, . yes The square of the smallest singular value of yes The square of the largest singular value of .

[0055] By adjusting the parameter matrix during training , you can Restricted to a range close to the identity matrix, that is, to achieve ,in, The eigenvalues ​​of , is a sufficiently small constant.

[0056] The current residual layer linearly combines the intra-layer operation of the current residual layer with the output of the previous residual layer through the combination coefficient to obtain the output of the current residual layer.

[0057] Then the output of the nth residual layer is The process of obtaining can be expressed as: .in, Indicates the coefficient used for each combination.

[0058] By setting the parameter matrix to the identity matrix, we assume that The matrix in Small enough, the matrix The impact on the residual output of the residual layer can be ignored. .

[0059] Based on this, we can further derive the expressions of the upper and lower limits of the Dirichlet energy of the graph neural network, as follows:

[0060]

[0061] in, is the kth combination coefficient, is the propagation matrix The minimum eigenvalue of between.

[0062] After obtaining the expressions of Dirichlet energy upper and lower limits, we can substitute the energy upper and lower limits mentioned above by specifying the energy upper and lower limits. The constraint range is calculated The value range of .

[0063] For example, the upper limit of the specified energy can be set to 100 and the lower limit of the specified energy can be set to 0.01 to calculate the undetermined graph neural network that meets the constraints.

[0064] In this method, the lower limit of the Dirichlet energy is constrained to ensure that it is not too smooth, and the upper limit of the Dirichlet energy is set to ensure that the node embedding is not too separated.

[0065] S106: Determine the values ​​of the combination coefficients within the range to obtain the target graph neural network.

[0066] In this step S104, only the value range of each combination coefficient is calculated. Within the value range of each coefficient, a specific value of the combination coefficient can be arbitrarily selected.

[0067] For example, if there are two combination coefficients, and the range of the combination coefficient A is calculated to be , the range of the combination coefficient B is , then the combination coefficient A can be The value is , the combination coefficient B can be Any value in . and The combination of can be expressed as a value set, which contains countless possibilities.

[0068] Therefore, by applying the graph neural network determination method provided in this specification, a group of deep graph neural networks can be determined. When different combination coefficients are selected, the target graph neural network is also different. Therefore, by selecting different combination coefficients, the method of this specification can determine a large number of high-depth graph neural networks.

[0069] After determining the specific values ​​of each combination coefficient, the only unknown quantity in the propagation recursion formula of the graph neural network is the learnable parameter matrix. The propagation recursion formula of the graph neural network is also completely determined, indicating that the construction of the graph neural network is complete and training and reasoning tasks can be performed based on the propagation recursion formula.

[0070] This specification does not limit the application of the constructed target graph neural network, which can be used for user profiling, intrusion detection, knowledge graph classification, etc.

[0071] Based on the above Figure 1 The energy-constrained graph neural network construction method shown in the figure obtains a pending graph neural network, wherein the pending graph neural network includes multiple residual layers, and the residual layer includes a graph convolution sublayer and a first combination sublayer. The residual layer linearly combines the output of the graph convolution sublayer with the output of the first combination sublayer to obtain the output of the residual layer. The first combination sublayer integrates the output of the previous residual layer, and sets the learnable parameter matrix contained in the residual layer to the unit matrix. By constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit, the range of each combination coefficient contained in the linear combination is determined. Within this range, the value of each combination coefficient is determined to obtain the target graph neural network.

[0072] In this method, by limiting the value range of Dirichlet energy, the node smoothness of the graph data embedded features obtained by the graph neural network is constrained, and by limiting the parameter matrix to the unit matrix, the unknown combination coefficients can be calculated according to the constraints of Dirichlet energy to obtain the combination coefficients that meet the specified Dirichlet energy range. The advantages of the present invention are twofold. One is that the performance of the graph neural network increases with the increase of depth, that is, the over-smoothing problem of the graph neural network is solved. The other is that a large number of high-depth graph neural networks can be generated.

[0073] According to the deep neural network construction method provided in this specification, a 128-layer deep graph neural network can be constructed. After training, the deep graph neural network can be used to generate character portraits, network attack detection, and academic knowledge graph subject classification.

[0074] In one embodiment of the present specification, the first combination sublayer of the to-be-determined graph neural network may further include a projection sublayer and a second combination sublayer. After the projection sublayer projects the output of the previous residual layer, the projection results of the outputs of each previous residual layer are linearly combined, and the second combination sublayer linearly combines the outputs of the previous residual layers.

[0075] Figure 3 This is a schematic diagram of a network structure of a first combination sublayer provided in an embodiment of this specification, wherein: represents the output of the n-1th residual layer, represents a linear combination operation. Figure 3 As shown, the first combination sublayer includes a projection sublayer and a second combination sublayer, and the projection sublayer receives the output of the previous residual layer: , perform projection operations on each received output, and perform linear combinations on each projection result. The second combination sublayer receives the output of the previous residual layer: , directly make a linear combination of the output of the previous residual layer.

[0076] In this embodiment, in the linear combination operation of the projection sub-layer and the second combination sub-layer, combination coefficients corresponding to the outputs of the previous residual layers are set.

[0077] The projection operation in this embodiment, for the output of a previous residual layer, is to multiply the output of the previous residual layer by a projection matrix. The projection matrix can be set as a learnable parameter matrix. In the network structure of the graph neural network to be determined in this embodiment, the graph convolution sublayer and the projection sublayer contain learnable parameter matrices.

[0078] by represents the learnable parameter matrix in the projection sublayer of the nth residual layer. Then, in the projection sublayer of the nth residual layer, the projection operation on the output of the previous residual layer can be expressed as: .

[0079] In this embodiment, two learnable parameter matrices corresponding to the graph convolution sublayer and the projection sublayer are included. During the training process, the two learnable parameter matrices can be adjusted separately. Alternatively, in order to improve the training efficiency, the parameter matrices of the graph convolution sublayer and the projection sublayer can be set to share parameters, and the two parameter matrices can be adjusted synchronously.

[0080] exist Figure 3 In the network structure shown, the projection sublayer and the second combination sublayer both function as a linear combination of the output of the previous residual layer. In one embodiment of the present specification, by setting the combination coefficient corresponding to the projection sublayer to 0, or the coefficient corresponding to the second combination sublayer to 0, only one linear combination operation can be retained in the first combination sublayer.

[0081] That is, in the above step S104, when calculating the range of each combination coefficient, the combination coefficient of the projection sublayer and the combination coefficient of the second combination sublayer may be limited, one of the combination coefficients being 0, to calculate the range of the combination coefficients included in the linear combination.

[0082] When the combination coefficient of the second combination sublayer is limited to 0, it is equivalent to the operation of the second combination sublayer having no effect on the output of the entire residual layer. The output of the nth residual layer can be expressed as: .

[0083] When the combination coefficient of the projection sub-layer is restricted to 0, it is equivalent to the operation of the projection sub-layer having no effect on the output of the entire residual layer. The output of the nth residual layer can be expressed as: It can be found that in this case, the network structure of the first combination sublayer including the projection sublayer and the second combination sublayer in this embodiment is equivalent to the network structure of the single first combination sublayer described in the above step S100.

[0084] In the above assumption of S102, the parameter matrix is ​​set to the unit matrix. Accordingly, in the training process of the graph neural network in this specification, it is necessary to ensure that the parameter matrix is ​​close to the unit matrix.

[0085] The training method is as follows: First, the server obtains sample graph data and label data of the sample graph data, and inputs the sample graph data into the target graph neural network determined in step S104 to obtain the prediction result output by the target graph neural network.

[0086] Then, the server determines the cross entropy loss based on the difference between the prediction result and the label data, and determines the unit constraint loss based on the difference between the parameter matrix of each residual layer of the graph neural network and the unit matrix. With the goal of minimizing the sum of the cross entropy loss and the unit constraint loss, the target graph neural network is trained and the parameter matrices of each residual layer of the graph neural network are adjusted.

[0087] Among them, the setting of the unit constraint loss can ensure that the parameter matrix of the graph neural network trained in this method is close to the unit matrix.

[0088] In the above step S100, when the previous residual layer is a specified residual layer before the current residual layer, several optional network structures of the graph neural network to be determined are provided below.

[0089] In one embodiment, the specified residual layer may be a residual layer before the current residual layer. The residual layer of the undetermined graph neural network includes a graph convolution sublayer, a projection sublayer, and a second combination sublayer. The propagation recursion formula of the undetermined graph neural network can be expressed as follows:

[0090]

[0091]

[0092] in, is the original input data of the first residual layer, represents the intermediate output of the nth residual layer, is the learnable parameter matrix in the convolutional sublayer, is the learnable parameter matrix in the projection sublayer, represents the propagation matrix, represents the residual output of the nth residual layer, Represents the activation function. , , is the combination coefficient.

[0093] Based on this embodiment, in the above step S104, the value range of each combination coefficient is determined, which can be specifically determined as follows: , , The value range of .

[0094] For example, in this embodiment, it can be calculated , .

[0095] In this embodiment, the convolution sublayer and the projection sublayer can share a parameter matrix to reduce the amount of calculation during training and improve training efficiency. When the convolution sublayer and the projection sublayer share a parameter matrix, the above propagation recursion can be abbreviated as:

[0096]

[0097]

[0098] in, is the parameter matrix shared by the convolution sublayer and the projection sublayer in the nth residual layer.

[0099] An example of selecting a combination coefficient is to take , , .

[0100] At this point, the propagation recursion formula of the graph neural network can be obtained as follows:

[0101]

[0102]

[0103] In another embodiment of the present specification, the specified residual layer may be all even-numbered residual layers before the current residual layer. An optional residual layer of a graph neural network includes a graph convolution sublayer and a first combination sublayer. The convolution sublayer includes a learnable parameter matrix. The propagation recursion formula of the graph neural network is as follows:

[0104]

[0105]

[0106] in, is the original input data of the first residual layer, represents the intermediate output of the nth residual layer, is the learnable parameter matrix, represents the propagation matrix, represents the output of the nth residual layer, represents the activation function, represents the output of the even residual layer, represents the linear combination of all even-numbered residual layers before the current residual layer. It is the combination coefficient that needs to be solved in this embodiment.

[0107] Based on this embodiment, in the above step S104, the value range of the combination coefficient is determined, which can be specifically determined as follows: The value of .

[0108] For example, in this embodiment, it can be calculated The range is Then in step S106, you can select .

[0109] At this point, the propagation recursion formula of the graph neural network can be obtained as follows:

[0110]

[0111]

[0112] In another embodiment of the present specification, the specified residual layer can be the input of the first residual layer (the output of the 0th residual layer) and the previous residual layer of the current residual layer. An optional residual layer of a graph neural network includes a convolution sublayer and a first combination sublayer. The convolution sublayer includes a learnable parameter matrix. The propagation recursion formula of the graph neural network is as follows:

[0113]

[0114] in, is the original input data of the first residual layer, is the learnable parameter matrix, represents the propagation matrix, represents the output of the nth residual layer, represents the propagation matrix, Represents the activation function. , , is the combination coefficient.

[0115] Based on this embodiment, in the above step S104, the value range of each combination coefficient is determined, which can be specifically determined as follows: , , The value range of .

[0116] For example, in this embodiment, it can be calculated , , .

[0117] Then an example of selecting a combination coefficient is: , , At this point, the propagation recursion formula of the graph neural network can be obtained as follows:

[0118]

[0119] In one embodiment, the target graph neural network constructed in this specification can be applied to predict user portraits. User portraits are the result of extracting and classifying user key attributes (such as interest tags and occupation categories) by analyzing user behavior patterns, interest preferences, social relationships and other information.

[0120] In this embodiment, the user information that can be extracted from the social platform website includes static attributes such as gender, age, and geographic location of the user, as well as dynamic behavior attributes such as browsing history, interaction frequency, and activity participation. The social network graph is constructed with the user as the node and the user's static attributes and dynamic behavior attributes as the edge.

[0121] The social network graph of the user whose user profile is to be determined is input into the trained target graph neural network to obtain the probability of the user corresponding to multiple user group classifications output by the target graph neural network. The user group classification corresponding to the highest probability can be used as the user profile of the user.

[0122] In one embodiment, the target graph neural network constructed in this specification can be applied to network attack detection. The goal of network attack detection is to identify potential abnormal behaviors and malicious activities from complex network traffic data, such as distributed denial of service (DDoS) attacks, data leaks, and network intrusions.

[0123] The network monitoring system continuously collects real-time traffic data, including communication frequency, access time, packet size and other features, and constructs a graph to input into the target graph neural network. The traffic data is constructed as an undirected graph. Each node of the undirected graph represents an IP address or host, and each edge represents the communication behavior between two nodes. The node features include static information of the host (IP type, location) and dynamic behavior features (number of traffic packets, connection frequency, and access time distribution).

[0124] The undirected graph corresponding to the constructed traffic data is input into the trained target graph neural network to obtain the abnormal probability of each node. The nodes with a probability higher than the threshold are regarded as risk nodes, which may be nodes where attacks occur.

[0125] In one embodiment, the target graph neural network constructed in this specification can be applied to subject classification based on academic knowledge graph. Academic knowledge graph is a graph structure data consisting of academic papers, authors, institutions and the citation relationship between them. Nodes represent papers, and edges represent citation relationships or the relationship between authors and papers. The task goal is to predict the subject area to which each paper belongs, which may include computer science, physics, biology, mathematics, chemistry, etc.

[0126] Academic knowledge graphs can be obtained from existing public academic datasets (such as ArXiv datasets, Microsoft Academic Graph, etc.). In academic knowledge graphs, nodes represent academic papers, including attributes such as title, keywords, abstract, and publication year. Edges represent citation relationships between papers, or the relationship between papers and authors. Node features are generated by embedding text attributes (such as title and abstract) into vector features; non-text attributes (such as publication year) are numerically encoded.

[0127] The academic knowledge graph to be classified is input into the trained target graph neural network to obtain the probability that the academic knowledge graph belongs to each subject field, and the subject field corresponding to the maximum probability is selected as the subject field of the academic knowledge graph to be classified.

[0128] The above is a graph neural network construction method based on energy constraints provided in this specification. Based on the same idea, this specification also provides a corresponding graph neural network construction device based on energy constraints, such as Figure 5 shown.

[0129] Figure 4 A schematic diagram of a high-depth graph neural network construction device based on energy constraints provided in this specification specifically includes: An acquisition module 200 is used to acquire a pending graph neural network, wherein the pending graph neural network includes a plurality of residual layers, wherein the residual layer includes a graph convolution sublayer and a first combination sublayer, wherein the residual layer linearly combines the output of the graph convolution sublayer with the output of the first combination sublayer to obtain the output of the residual layer, and the first combination sublayer integrates the output of the previous residual layer; Assume that module 202 is used to set the learnable parameter matrix contained in the residual layer to a unit matrix; An inference prediction module 204, configured to determine a range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit; The determination module 206 is used to determine the values ​​of each combination coefficient within the range to obtain the target graph neural network.

[0130] Optionally, the first combination sublayer includes a projection sublayer and a second combination sublayer, the projection sublayer projects the output of the previous residual layer and performs linear combination, and the second combination sublayer performs linear combination on the output of the previous residual layer.

[0131] Optionally, the inference prediction module 204 is specifically used to limit the combination coefficients of the projection sublayer and the combination coefficients of the second combination sublayer, one of which is 0, to determine the range of the combination coefficients included in the linear combination.

[0132] Optionally, the graph convolution sublayer contains a learnable parameter matrix.

[0133] Optionally, the graph convolution sublayer and the projection sublayer contain learnable parameter matrices.

[0134] Optionally, parameter matrices of the graph convolution sublayer and the projection sublayer share parameters.

[0135] Optionally, the device further comprises a training module 208; The training module 208 is used to obtain sample graph data and label data of the sample graph data, input the sample graph data into the target graph neural network, obtain a prediction result output by the target graph neural network, and determine the cross entropy loss based on the difference between the prediction result and the label data; determine the unit constraint loss based on the difference between the parameter matrix of each residual layer of the graph neural network and the unit matrix, and adjust the parameter matrix of each residual layer of the target graph neural network with the goal of minimizing the sum of the cross entropy loss and the unit constraint loss.

[0136] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A method for constructing a high-depth graph neural network based on energy constraints is provided.

[0137] This manual also provides Figure 5 The schematic structure diagram of the electronic device shown in FIG. Figure 5As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The energy-constrained graph neural network construction method. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0138] For the improvement of a technology, it can be clearly distinguished whether it is a hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or a software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0139] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.

[0140] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0141] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0142] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0143] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0144] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0146] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0147] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0148] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0149] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0150] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0151] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0152] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0153] The above description is only an embodiment of this specification and is not intended to limit this specification. For those skilled in the art, this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included in the scope of the claims of this application.

Claims

1. A method for constructing a high-depth graph neural network based on energy constraints, characterized in that: include: Obtain a pending graph neural network, wherein the pending graph neural network includes a plurality of residual layers, wherein the residual layer includes a graph convolution sublayer and a first combination sublayer, wherein the residual layer linearly combines the output of the graph convolution sublayer with the output of the first combination sublayer to obtain the output of the residual layer, and the first combination sublayer integrates the output of the previous residual layer; Setting the learnable parameter matrix contained in the residual layer to the unit matrix; Determining the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit; Within the range, the values ​​of the combination coefficients are determined to obtain the target graph neural network.

2. The method according to claim 1, characterized in that The first combination sublayer includes a projection sublayer and a second combination sublayer. The projection sublayer projects the output of the previous residual layer and performs linear combination. The second combination sublayer performs linear combination on the output of the previous residual layer.

3. The method according to claim 2, characterized in that Determining the range of the combination coefficients included in the linear combination specifically includes: The combination coefficients of the projection sublayer and the second combination sublayer are restricted, one of the combination coefficients being 0, and the range of the combination coefficients included in the linear combination is determined.

4. The method according to claim 1, characterized in that The graph convolution sublayer contains a learnable parameter matrix.

5. The method according to claim 2, characterized in that The graph convolution sublayer and the projection sublayer contain learnable parameter matrices.

6. The method according to claim 5, characterized in that The parameter matrices of the graph convolution sublayer and the projection sublayer share parameters.

7. The method according to claim 1, characterized in that The method further comprises: Training the target graph neural network specifically includes: Obtaining sample graph data and label data of the sample graph data; Inputting the sample graph data into the target graph neural network to obtain a prediction result output by the target graph neural network; Determine the cross entropy loss based on the difference between the prediction result and the label data; determine the unit constraint loss based on the difference between the parameter matrix of each residual layer of the graph neural network and the unit matrix; With the goal of minimizing the sum of the cross entropy loss and the unit constraint loss, the parameter matrices of the residual layers of the target graph neural network are adjusted.

8. A device for constructing a high-depth graph neural network based on energy constraints, characterized in that: include: An acquisition module is provided to acquire a pending graph neural network, wherein the pending graph neural network includes a plurality of residual layers, wherein the residual layer includes a graph convolution sublayer and a first combination sublayer, wherein the residual layer linearly combines the output of the graph convolution sublayer with the output of the first combination sublayer to obtain the output of the residual layer, and the first combination sublayer integrates the previous residual layer; Assume that the module sets the learnable parameter matrix contained in the residual layer to the unit matrix; An inference prediction module determines the range of each combination coefficient included in the linear combination by constraining the Dirichlet energy of the residual layer between a specified energy upper limit and a specified energy lower limit; A determination module determines the values ​​of each combination coefficient within the range to obtain a target graph neural network.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Data classification method and device based on unified optimization target framework graph neural network

    CN112733933A

  • Traffic flow prediction method of attention-based deep residual space-time diagram convolutional network

    CN115691129A

  • Recommendation method capable of fusing various graph structure information and having interpretability

    CN116610862A

  • Broadband data oscillation positioning method and system based on depth map neural network

    CN119399604A

  • Virtual power plant collaborative optimization scheduling method and system based on deep reinforcement learning

    CN119494521A