Graph neural network structure learning method and system based on self-supervised learning

By using a self-supervised learning method to generate a multi-view adjacency matrix and introduce a self-constraint mechanism, the performance degradation problem of graph neural networks under incomplete graph structures and noise is solved, and more efficient feature propagation and node classification accuracy are achieved.

CN120911537AActive Publication Date: 2025-11-07XIAN AERONAUTICAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511011342.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-11-07
Estimated Expiration
2045-07-22

Smart Images

  • Figure CN120911537A_ABST
    Figure CN120911537A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of graph neural networks, in particular to a graph neural network structure learning method and system based on self-supervised learning, and the method comprises the steps: obtaining a node feature matrix and an adjacent matrix; performing symmetric normalization processing on the adjacent matrix through an adjacent converter, and performing multi-order information fusion to obtain a fused adjacent matrix; obtaining a binary mask matrix through the node feature matrix, disturbing the node feature matrix through the binary mask matrix to obtain a new feature matrix, and obtaining a new node feature matrix through the new feature matrix and the fusion adjacent matrix; determining a classifier and a total loss function by fusing the adjacent matrix and the new node feature matrix; and learning of the graph neural network structure of self-supervised learning is carried out through the classifier and the total loss function. According to the invention, the accuracy of graph neural network structure learning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of graph neural networks, in particular to a graph neural network structure learning method and system based on self-supervised learning. BACKGROUND

[0002] When learning in the existing graph neural network structure, firstly, a traditional GNN (Graph Neural Network) model is seriously dependent on an initial graph structure, and performance is significantly reduced when the input graph structure is incomplete or contains noise; secondly, node features tend to be homogenized (over-smoothed) after multi-round feature propagation, so that different categories of nodes are difficult to distinguish. Therefore, a self-supervised graph stochastic neural network (SCGSNN, Self-Constructing Graph Stochastic Neural Network) is designed, an adjacency matrix is dynamically learned through multi-view generation structure, a self-constraint mechanism is introduced to constrain the consistency of generated features and original features, and a hybrid multi-level information aggregation is adopted to fuse local structure information. SUMMARY

[0003] The application provides a graph neural network structure learning method and system based on self-supervised learning, and is used for solving the problem that an existing graph neural network is difficult to distinguish due to incomplete graph structure, noise and over-smoothed node features.

[0004] The purpose of the application can be achieved by the following technical solutions. The first aspect of the application provides a graph neural network structure learning method based on self-supervised learning, comprising: obtaining a node feature matrix, generating an adjacency matrix through multi-view generation structure according to the node feature matrix; symmetrically normalizing the adjacency matrix through an adjacency converter to obtain a symmetrically normalized adjacency matrix; performing multi-order information fusion on the symmetrically normalized adjacency matrix to obtain a fused adjacency matrix; obtaining a binary mask matrix through the node feature matrix, perturbing the node feature matrix through the binary mask matrix to obtain a new feature matrix, and obtaining a new node feature matrix through the new feature matrix and the fused adjacency matrix; determining a classifier through the fused adjacency matrix and the new node feature matrix, determining a total loss function according to the difference between the feature vectors of the nodes in the new node feature matrix and the node feature matrix, and the cross entropy between the class probability distribution predicted by the model and the One-Hot encoding of the real label, and learning the graph neural network structure through the classifier and the total loss function.

[0005] Further, the step of obtaining the node feature matrix and generating an adjacency matrix based on the node feature matrix through a multi-view generation structure includes: Obtain the feature vector of each node in the graph structure, and concatenate the feature vectors of all nodes to obtain the node feature matrix; A node matrix is ​​generated using a multilayer perceptron based on the node feature matrix. A sparse matrix and a similarity matrix are obtained using the K-nearest neighbor algorithm based on the node matrix. An adjacency matrix is ​​obtained using the sparse matrix and the similarity matrix.

[0006] Further, the step of generating a node matrix using a multilayer perceptron based on the node feature matrix, obtaining a sparse matrix and a similarity matrix using the K-nearest neighbor algorithm based on the node matrix, and obtaining an adjacency matrix using the sparse matrix and the similarity matrix includes: The specific process for obtaining the similarity matrix is ​​as follows:

[0007] In the formula, Represents the node in the node matrix The eigenvectors of the row, Represents the node in the node matrix The eigenvectors of the row, Represents the node in the node matrix The magnitude of the eigenvector of a row. Represents the node in the node matrix The magnitude of the eigenvector of a row. In a similarity matrix, the coordinates are... The corresponding element; The similarity matrix is ​​determined by the elements corresponding to all coordinates in the similarity matrix; The specific process for obtaining the sparse matrix is ​​as follows: Each row of the similarity matrix is ​​labeled according to its size. The labeling process is as follows: the first K largest elements in each row of the similarity matrix are labeled as 1, and the remaining elements in each row are labeled as 0. All row elements of the similarity matrix are labeled according to the labeling process. The labeled matrix is ​​called the sparse matrix. The specific process for obtaining the adjacency matrix is ​​as follows:

[0008] In the formula, Represents a sparse matrix. Let ⊙ denote a similar matrix, and ⊙ denote the Hadamard product. This represents the adjacency matrix.

[0009] Further, the step of performing symmetric normalization on the adjacency matrix using an adjacency converter to obtain a symmetric normalized adjacency matrix includes:

[0010] In the formula, represents the initial adjacency matrix, represents the activation function, represents the processed symmetric normalized adjacency matrix; represents the degree matrix, represents the transposed matrix of the initial adjacency matrix processed by the activation function, represents the inverse matrix of the square root of the degree matrix; wherein the degree matrix is a diagonal matrix; wherein the degree matrix each element on the diagonal line is:

[0011] In the formula, represents the corresponding element in the initial adjacency matrix with coordinates , represents the number of all columns, represents the corresponding element in the degree matrix with coordinates .

[0012] Further, the multi-order information fusion on the symmetric normalized adjacency matrix to obtain a fused adjacency matrix comprises:

[0013] In the formula, represents the order symmetric normalized adjacency matrix, represents the weight coefficient of the order symmetric normalized adjacency matrix, represents the highest order number, represents the fused adjacency matrix.

[0014] Further, the binary mask matrix is obtained through the node feature matrix, the new feature matrix is obtained by perturbing the node feature matrix through the binary mask matrix, and the new node feature matrix is obtained through the new feature matrix and the fused adjacency matrix, comprising: The binary mask matrix is obtained according to the Bernoulli distribution of the node feature matrix; The acquisition process of the new feature matrix is:

[0015] In the formula, represents the node feature matrix, represents the binary mask matrix, represents the new feature matrix, and represents the Hadamard product. The new feature matrix and the fusion adjacency matrix are input into the GNN to obtain a new node feature matrix.

[0016] Further, the classifier is determined by fusing the adjacency matrix and the new node feature matrix, and the classifier includes: The formula corresponding to the first layer in the classifier is specifically represented as:

[0017] In the formula, denotes the fusion adjacency matrix, denotes the new node feature matrix, denotes the first layer weight matrix, is an activation function, denotes a hidden layer node feature matrix; The formula corresponding to the second layer in the classifier is specifically represented as:

[0018] In the formula, denotes the second layer weight matrix, denotes a normalized exponential function, denotes a probability distribution matrix of node classification; The first layer weight matrix and the second layer weight matrix are both obtained by random initialization and optimized by a back propagation algorithm in a training process.

[0019] Further, the total loss function is determined according to a difference between feature vectors of nodes in the new node feature matrix and the node feature matrix, and a cross entropy between a class probability distribution predicted by a model and a One-Hot encoding of a real label, and the total loss function includes: The total loss function is composed of two parts; one part is a classification loss, and the other part is a self-constraint loss; The classification loss function is specifically represented as:

[0020] In the formula, denotes a node belongs to a class , and the one-hot encoding value, denotes a node belongs to a class , and the predicted probability; denotes a number of all classes, denotes a natural constant-based logarithmic function, denotes a number of all nodes in a set of labeled nodes, denotes a classification loss value; The self-constraint loss function is specifically represented as:

[0021] In the formula, represents the feature vector of the i-th node in the new node feature matrix, represents the number of all nodes, represents the feature vector of the i-th node in the node feature matrix, represents the feature vector of the i-th node in the node feature matrix, represents the square of the Euclidean distance between the two feature vectors; represents the self-constraint loss value; represents the total loss value. The total loss function is specifically represented as:

[0022] In the formula, represents the total loss value.

[0023] The second aspect of the present application provides a graph neural network structure learning system based on self-supervised learning, comprising: A multi-view generation module is used to obtain a node feature matrix, generate an adjacency matrix through a multi-view generation structure according to the node feature matrix; A conversion module is used to perform symmetric normalization processing on the adjacency matrix through an adjacency converter to obtain a symmetric normalized adjacency matrix; A fusion module is used to perform multi-order information fusion on the symmetric normalized adjacency matrix to obtain a fused adjacency matrix; A self-constraint module is used to obtain a binary mask matrix through the node feature matrix, perturb the node feature matrix through the binary mask matrix to obtain a new feature matrix, and obtain a new node feature matrix through the new feature matrix and the fused adjacency matrix; A classification module is used to determine a classifier through the fused adjacency matrix and the new node feature matrix, determine a total loss function through the difference between the feature vectors of the nodes in the new node feature matrix and the node feature matrix, the cross entropy between the class probability distribution predicted by the model and the One-Hot encoding of the real label, and perform self-supervised learning of the graph neural network structure through the classifier and the total loss function.

[0024] The third aspect of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the graph neural network structure learning method based on self-supervised learning when executing the computer program.

[0025] Compared with the prior art, the beneficial effects of the present application are: obtaining a node feature matrix, generating an adjacency matrix according to the node feature matrix through a multi-view generation structure; improving the robustness of the graph structure, avoiding dependence on a single input graph structure, alleviating performance degradation caused by incomplete or noisy original graph structures, enhancing expression ability, and capturing more rich node relationships; performing symmetric normalization processing on the adjacency matrix through an adjacency converter to obtain a symmetric normalized adjacency matrix; avoiding unstable feature values caused by too large node degree difference through normalization, and making feature propagation more efficient through symmetric processing to speed up model training; performing multi-order information fusion on the symmetric normalized adjacency matrix to obtain a fused adjacency matrix; alleviating oversmoothing and enhancing local perception through fusion; obtaining a binary mask matrix through the node feature matrix, perturbing the node feature matrix through the binary mask matrix to obtain a new feature matrix, and obtaining a new node feature matrix through the new feature matrix and the fused adjacency matrix; improving model robustness and preventing overfitting; determining a classifier through the fused adjacency matrix and the new node feature matrix; determining a total loss function according to the difference between the feature vectors of the nodes in the new node feature matrix and the node feature matrix, and the cross entropy between the class probability distribution predicted by the model and the One-Hot encoding of the real label; learning the self-supervised learning graph neural network structure through the classifier and the total loss function, balancing supervised and self-supervised signals, jointly updating the adjacency matrix generator and the classifier parameters, realizing joint learning of the graph structure and the node representation, and improving the accuracy of the graph neural network structure learning. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0027] Figure 1 A step flowchart of a graph neural network structure learning method based on self-supervised learning is provided for the present application. Figure 2 A module flowchart of a graph neural network structure learning system based on self-supervised learning is provided for the present application. DETAILED DESCRIPTION

[0028] In order to make the person skilled in the art better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0029] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0030] In view of the problems in the background art, a graph neural network structure learning method and system based on self-supervised learning are designed, which has important practical significance.

[0031] As shown in Figure 1 The first aspect of the present application is to provide a graph neural network structure learning method based on self-supervised learning, comprising the following steps: Step S001: Obtain a node feature matrix, and generate an adjacency matrix according to the node feature matrix through multi-view generation structure.

[0032] It should be noted that in order to solve the problem of trend homogenization after iterative updating of the existing graph result, a plurality of adjacency matrices of different views are generated from the original graph to avoid the deviation of a single view.

[0033] Specifically, the feature vector corresponding to each node in the graph structure is obtained, and the feature vectors of all nodes are spliced to obtain a node feature matrix; According to the node feature matrix, an adjacency matrix is generated through multi-view generation structure; wherein each matrix represents a view of the graph structure; In this embodiment, the multi-view generation structure is a multi-layer perceptron K nearest neighbor (MLP-KNN, Multi-layer Perceptron K-Nearest Neighbor) processor. The adjacency matrix generated by the multi-view generation processor is an asymmetric and non-normalized adjacency matrix.

[0034] It should be noted that a learnable parameterization (LP) processor can be used instead of the MLP-KNN processor.

[0035] The node feature matrix is processed by the MLP-KNN processor to obtain an adjacency matrix. The specific process is as follows: First, a node matrix is generated from the node feature matrix by a multi-layer perceptron (MLP), and a sparse matrix and a similarity matrix are obtained from the node matrix by a K-nearest neighbor (KNN) algorithm. The adjacency matrix is obtained from the sparse matrix and the similarity matrix. It should be noted that the adjacency matrix is a first-order adjacency matrix.

[0036] The MLP and KNN are known technologies and will not be described in detail here.

[0037] The specific process of obtaining the sparse matrix and the similarity matrix from the node matrix by KNN is as follows: The feature vector of each row in the node matrix is obtained. The similarity between any two rows is calculated by cosine similarity, and the similarity between any two rows is taken as an element of the similarity matrix. The specific formula is as follows:

[0038] In the formula, represents the feature vector of the i-th row in the node matrix, represents the feature vector of the j-th row in the node matrix, represents the norm of the feature vector of the i-th row in the node matrix, represents the norm of the feature vector of the j-th row in the node matrix, represents the element corresponding to the coordinates (i, j) in the similarity matrix. The similarity matrix is determined by the elements corresponding to all coordinates in the similarity matrix. At this point, the similarity matrix is obtained.

[0039]

[0040] At this point, the similarity matrix is obtained.

[0041] ​​​​According to the size of each row element of the similarity matrix, each row element is marked, and the marking process of each row element is specifically that the first K largest elements in each row of the similarity matrix are marked as 1, and the remaining elements in each row are marked as 0; all row elements of the similarity matrix are marked according to the marking process of each row element, and the marked matrix of the similarity matrix is a sparse matrix. In this embodiment, K is a preset number threshold, and in this embodiment, K = 20. In this embodiment, the preset number threshold K is not specifically limited, and the implementer can be determined according to the specific situation.

[0042] At this point, the sparse matrix is obtained.

[0043] By the sparse matrix and the similarity matrix, the adjacency matrix is obtained, and the formula is specifically:

[0044] In the formula, the sparse matrix is represented by S, the similarity matrix is represented by S, and the adjacency matrix is represented by A.

[0045] At this point, the adjacency matrix is obtained by the above method.

[0046] Step S002: symmetric normalization processing is performed on the adjacency matrix by the adjacency converter to obtain a symmetric normalized adjacency matrix.

[0047] It should be noted that, in order to reduce the influence of the difference between the data, the generated asymmetric and non-normalized adjacency matrix is converted into a symmetric and normalized form, which is convenient for subsequent processing.

[0048] Specifically, the symmetric processing is performed on each generated adjacency matrix, and then the matrix is normalized.

[0049] The symmetric normalization processing is performed on the adjacency matrix by the adjacency converter to obtain a symmetric normalized adjacency matrix; The symmetric normalization processing of the adjacency matrix by the adjacency converter is specifically represented as:

[0050] In the formula, the initial adjacency matrix is represented by A, the activation function is represented by f (where the corresponding activation function is different when different processors are used), the processed symmetric normalized adjacency matrix is represented by A, the degree matrix (which is a diagonal matrix) is represented by D, and the transposed matrix of the initial adjacency matrix processed by the activation function is represented by A. An inverse matrix of a square root of a degree matrix.

[0051] wherein each element on the diagonal of the degree matrix is:

[0052] wherein, denotes an element in the initial adjacency matrix with coordinates corresponding element, denotes the number of all columns, denotes an element in the degree matrix with coordinates corresponding element.

[0053] Thus far, the symmetric normalized adjacency matrix is obtained by the above method.

[0054] Step S003: obtaining a fused adjacency matrix by multi-order information fusion on the symmetric normalized adjacency matrix.

[0055] It should be noted that, in order to alleviate the problem of over-smoothing, the local information of different levels can be fused to solve the problem, so that more local features can be retained.

[0056] Specifically, the fused adjacency matrix is obtained by multi-order information fusion on the symmetric normalized adjacency matrix, and is specifically expressed by a formula as follows:

[0057] wherein, denotes the order symmetric normalized adjacency matrix, denotes the weight coefficient of the order symmetric normalized adjacency matrix, denotes the highest order number, denotes the fused adjacency matrix.

[0058] Thus far, the fused adjacency matrix is obtained by the above method.

[0059] Step S004: obtaining a binary mask matrix by the node feature matrix, perturbing the node feature matrix by the binary mask matrix, obtaining a new feature matrix, and obtaining a new node feature matrix by the new feature matrix and the fused adjacency matrix.

[0060] It should be noted that, in order to enhance the robustness of the model, the binary mask matrix is generated to perturb the node feature matrix, so as to enhance the robustness of the model.

[0061] Specifically, a binary mask matrix is obtained through the node feature matrix; elements in the binary mask matrix are 0 or 1; the binary mask matrix is obtained through Bernoulli distribution according to the node feature matrix; in the Bernoulli distribution in this embodiment, the retention probability is 0.8. The Bernoulli distribution is a known technology, and will not be described in detail here.

[0062] The node feature matrix is disturbed through the binary mask matrix to obtain a new feature matrix; the obtaining process of the new feature matrix is specifically represented by a formula as follows:

[0063] In the formula, X represents the node feature matrix, B represents the binary mask matrix, and X' represents the new feature matrix. In the formula, X represents the node feature matrix, In the formula, B represents the binary mask matrix, In the formula, X' represents the new feature matrix, and represents Hadamard product.

[0064] The new feature matrix and the fusion adjacency matrix are input into a GNN (Graph Neural Network) to obtain a new node feature matrix.

[0065] Thus, the new node feature matrix is obtained through the above method.

[0066] Step S005: a classifier is determined through the fusion adjacency matrix and the new node feature matrix; a total loss function is determined according to the difference between the feature vectors of the nodes in the new node feature matrix and the node feature matrix, and the cross entropy between the class probability distribution predicted by the model and the One-Hot encoding of the real label; and learning of the self-supervised learning graph neural network structure is performed through the classifier and the total loss function.

[0067] It should be noted that the classifier and the objective function are the last step of the SCGSNN (Self-Constructing Graph Stochastic Neural Network), which is responsible for mapping the optimized graph structure and node features to the prediction result, and guiding the model training through the joint loss function.

[0068] Specifically, a two-layer GCN (Graph Convolutional Network) is used as the classifier in this embodiment. In the formula, A represents the fusion adjacency matrix, and X' represents the new node feature matrix.

[0069] In the formula, A represents the fusion adjacency matrix, In the formula, A represents the fusion adjacency matrix, In the formula, X' represents the new node feature matrix, denotes the first layer weight matrix, is an activation function, denotes the hidden layer node feature matrix.

[0070] wherein the formula corresponding to the second layer in the classifier is specifically represented as:

[0071] wherein, denotes the fusion adjacency matrix, denotes the hidden layer node feature matrix, denotes the second layer weight matrix, denotes the normalized exponential function, denotes the probability distribution matrix (or prediction score matrix) of node classification.

[0072] wherein the first layer weight matrix and the second layer weight matrix are obtained by random initialization and optimized through the back propagation algorithm in the training process.

[0073] The objective function jointly optimizes the classification task and the self-constraint target, and the loss function thereof is composed of two parts; one part is the classification loss, and the other part is the self-constraint loss.

[0074] wherein the classification loss function is specifically represented as:

[0075] wherein, denotes the node belongs to the one-hot encoding value of the category , denotes the node belongs to the predicted probability of the category ; denotes the number of all categories, denotes the natural constant based logarithmic function, denotes the number of all nodes in the set of labeled nodes, denotes the classification loss value. Wherein, the value of is 1 or 0, that is, when the real category of the node is , then is 1, otherwise is 0.

[0076] wherein the self-constraint loss function is specifically represented as:

[0077] wherein, denotes the first a feature vector of the i-th node, denotes the number of all nodes, denotes a feature vector of the i-th node in the node feature matrix, a feature vector of the i-th node, denotes the square of the Euclidean distance between two feature vectors; denotes a self-constrained loss value.

[0078] According to the classification loss function and the self-constrained loss function, a total loss function is obtained; specifically represented as:

[0079] In the formula, denotes a classification loss value, denotes a self-constrained loss value, denotes a total loss value.

[0080] Finally, the self-supervised learning of the graph neural network structure learning method is verified through the classifier and the total loss function; Among them, the self-supervised learning of the graph neural network structure learning method is experimentally analyzed, and the specific analysis results are as follows: First of all, according to experience, it is determined that the order number in the multi-level information fusion process of the several order symmetric normalized adjacency matrices is 3, that is, the classification effect is the best when the order number is 3; wherein the weight coefficient of the first order symmetric normalized adjacency matrix is 1; the weight coefficient of the second order symmetric normalized adjacency matrix is in the range of 1 / 5 to 1; the weight coefficient of the third order symmetric normalized adjacency matrix is in the range of 1 / 9 to 1.

[0081] The experimental results show that the SCGSNN model is significantly better than the existing methods on multiple benchmark data sets, and the classification accuracy on the Cora, Citeseer and Pubmed data sets reaches 77.5%, 74.8% and 75.1% respectively, which is improved by 5.1%, 18.8% and 28.8% respectively compared with the traditional GCN model, and even on the large-scale ogbn-arxiv data set, an accuracy of 57.8% is achieved, which is 8.7% better than the comparison method. Through the ablation experiment verification, the SCGSNN (Self-Constructing Graph Stochastic Neural Network) introduced by the self-constraint mechanism is 5.1% (76.4% vs 71.3%) higher in accuracy than the SCGSNN (Self-Constrained Graph Stochastic Network-No self-constraint) without self-constraint on the Cora data set, and shows stronger robustness when processing perturbed data. In addition, the SCGSNN effectively alleviates the over-smoothing problem and still maintains a feature discrimination degree (MADGap index, wherein, MADGap is Mean Absolute Deviation Gap) of 48% when the network layer is increased to 16 layers, which is much higher than the 12% of GCN. The multi-view generation structure and the mixed multi-level aggregation strategy make the model integrate more rich local information, and the self-constraint mechanism ensures the rationality of the generated graph structure by minimizing the feature reconstruction error (reduced by more than 30%). These improvements make SCGSNN a general and robust graph structure learning framework, which realizes performance breakthrough in node classification tasks and shows good scalability.

[0082] As shown in Figure 2 The second aspect of the present application is to provide a graph neural network structure learning system based on self-supervised learning, comprising: The multi-view generation module 101 is used for obtaining a node feature matrix, and generating an adjacency matrix through a multi-view generation structure according to the node feature matrix; The conversion module 102 is used for performing symmetric normalization processing on the adjacency matrix through an adjacency converter to obtain a symmetric normalized adjacency matrix; The fusion module 103 is used for performing multi-order information fusion on the symmetric normalized adjacency matrix to obtain a fused adjacency matrix; The self-constraint module 104 is used for obtaining a binary mask matrix through the node feature matrix, performing perturbation on the node feature matrix through the binary mask matrix to obtain a new feature matrix, and obtaining a new node feature matrix through the new feature matrix and the fused adjacency matrix; The classification module 105 is configured to determine a classifier by fusing the adjacency matrix and the new node feature matrix, determine a total loss function according to a difference between feature vectors of nodes in the new node feature matrix and the node feature matrix, and a cross entropy between a class probability distribution predicted by the model and a One-Hot encoding of a real label, and learn the structure of the graph neural network by self-supervised learning based on the classifier and the total loss function.

[0083] The third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements a method for learning a structure of a graph neural network based on self-supervised learning when executing the computer program.

[0084] Those skilled in the art will understand that embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code.

[0085] The present application is described with reference to flowcharts and / or block diagrams of methods, systems, and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce an apparatus that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for performing the function specified by one or more flows and / or blocks

[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction means, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for performing the function specified by one or more flows and / or blocks

[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable data processing apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a process for implementing the functions specified in the flowcharts and / or block diagrams.Figure 1 steps of a function specified in the flow or flows and / or block. Figure 1 steps of a function specified in the flow or flows and / or block.

[0088] It should be noted that the above-mentioned embodiments are only used to illustrate but not to limit the technical solutions of the present application. Although the present application has been described in detail with reference to the above-mentioned embodiments, those of ordinary skill in the art should understand that the specific embodiments of the present application can be modified or equivalent replacements without departing from the spirit and scope of the present application, and any modifications or equivalent replacements should be covered in the protection scope of the present application.

Claims

1. A method for learning a graph neural network structure based on self-supervised learning, characterized in that, The method comprises the following steps: obtaining a node feature matrix, and generating an adjacency matrix according to the node feature matrix through a multi-view generation structure; symmetrically normalizing the adjacency matrix through an adjacency converter to obtain a symmetrically normalized adjacency matrix; performing multi-order information fusion on the symmetrically normalized adjacency matrix to obtain a fused adjacency matrix; obtaining a binary mask matrix from the node feature matrix, perturbing the node feature matrix through the binary mask matrix to obtain a new feature matrix, and obtaining a new node feature matrix from the new feature matrix and the fused adjacency matrix; determining a classifier from the fused adjacency matrix and the new node feature matrix, determining a total loss function according to the difference between the feature vectors of the nodes in the node feature matrix and the new node feature matrix, and the cross entropy between the class probability distribution predicted by the model and the One-Hot encoding of the real label, and learning the self-supervised learning graph neural network structure through the classifier and the total loss function.

2. The method of claim 1, wherein, The method comprises the following steps: obtaining a node feature matrix, and generating an adjacency matrix according to the node feature matrix through a multi-view generation structure; obtaining a node feature matrix by concatenating the feature vectors of all nodes, and generating a node matrix through a multi-layer perceptron according to the node feature matrix, and obtaining a sparse matrix and a similarity matrix through a K nearest neighbor algorithm according to the node matrix; and obtaining the adjacency matrix through the sparse matrix and the similarity matrix.

3. The method of claim 2, wherein, The method comprises the following steps: obtaining a node feature matrix, and generating an adjacency matrix according to the node feature matrix through a multi-view generation structure; obtaining a node feature matrix by concatenating the feature vectors of all nodes, and generating a node matrix through a multi-layer perceptron according to the node feature matrix, and obtaining a sparse matrix and a similarity matrix through a K nearest neighbor algorithm according to the node matrix; and obtaining the adjacency matrix through the sparse matrix and the similarity matrix. wherein denotes the eigenvector of the node matrix in the row, denotes the eigenvector of the node matrix in the row, denotes the length of the eigenvector of the node matrix in the row, denotes the length of the eigenvector of the node matrix in the row, denotes the corresponding element of the similarity matrix in the coordinates and The method comprises the following steps: The specific process for obtaining the similarity matrix is as follows: determining the similarity matrix through the elements corresponding to all coordinates in the similarity matrix; The specific process for obtaining the sparse matrix is as follows: wherein denotes a sparse matrix, denotes a similarity matrix, and denotes a Hadamard product, denotes an adjacency matrix.

4. The method of claim 1, wherein, labeling each row element according to the size of each row element in the similarity matrix, wherein the process of labeling each row element is specifically as follows: labeling the first K largest elements in each row of the similarity matrix as 1, and labeling the remaining elements in each row as 0; labeling all row elements of the similarity matrix according to the process of labeling each row element, and the labeled matrix of the similarity matrix is denoted as a sparse matrix; In the formula, denotes an initial adjacency matrix, denotes an activation function, denotes a processed symmetric normalized adjacency matrix; denotes a degree matrix, denotes a transpose matrix of the matrix after processing the initial adjacency matrix by the activation function, denotes an inverse matrix of the square root of the degree matrix; The specific process for obtaining the adjacency matrix is as follows: where each element on the diagonal of the degree matrix D is: D = (dij) wherein denotes the element of the initial adjacency matrix with coordinates corresponding to the element denotes the number of all columns, denotes the element of the degree matrix with coordinates corresponding to the element 5. The method of claim 1, wherein, The specific process for obtaining the symmetrically normalized adjacency matrix is as follows: wherein, denotes the order symmetric normalized adjacency matrix, denotes the weight coefficient of the denotes the highest order number, denotes the fused adjacency matrix.

6. The method of claim 1, wherein, wherein the degree matrix is a diagonal matrix; The method comprises the following steps: The specific process for obtaining the fused adjacency matrix is as follows: wherein denotes a node feature matrix, denotes a binary mask matrix, denotes a new feature matrix, and denotes a Hadamard product; The method comprises the following steps:

7. The method of claim 1, wherein, obtaining a binary mask matrix from the node feature matrix, perturbing the node feature matrix through the binary mask matrix to obtain a new feature matrix, and obtaining a new node feature matrix from the new feature matrix and the fused adjacency matrix; obtaining a binary mask matrix from the node feature matrix through a Bernoulli distribution; The method comprises the following steps: inputting the new feature matrix and the fused adjacency matrix into a GNN to obtain a new node feature matrix. The method comprises the following steps: The formula corresponding to the first layer in the classifier is specifically represented as follows: wherein, denotes a fused adjacency matrix, denotes a new node feature matrix, denotes a first layer weight matrix, is an activation function, denotes a hidden layer node feature matrix; The formula corresponding to the second layer in the classifier is specifically represented as: wherein denotes a second layer weight matrix, denotes a normalized exponential function, denotes a probability distribution matrix of node classification; The first layer weight matrix and the second layer weight matrix are obtained by random initialization and optimized by the back propagation algorithm in the training process.

8. The method of claim 1, wherein, The total loss function is determined according to the difference between the feature vectors of the nodes in the new node feature matrix and the node feature matrix, the cross entropy between the class probability distribution predicted by the model and the One-Hot encoding of the real label, and the total loss function comprises: The total loss function is composed of two parts; one part is the classification loss, and the other part is the self-constraint loss; The classification loss function is specifically represented as: wherein, represents a node belongs to a class of one-hot encoding values, represents a node belongs to a class of a predicted probability; represents the number of all classes, represents a natural constant-based logarithmic function, represents the number of all nodes in a set of labeled nodes, represents a classification loss value; The self-constraint loss function is specifically represented as: wherein, denotes the eigenvector of the i-th node in the new node feature matrix, denotes the eigenvector of the i-th node in the new node feature matrix, denotes the number of all nodes, denotes the eigenvector of the i-th node in the new node feature matrix, denotes the eigenvector of the i-th node in the new node feature matrix, denotes the square of the Euclidean distance between two eigenvectors; denotes the self-constrained loss value; The total loss function is specifically represented as: In the formula, represents the total loss value.

9. A self-supervised learning based graph neural network structure learning system, characterized in that, Comprises: Multi-view generation module: for obtaining the node feature matrix, generating the adjacency matrix according to the node feature matrix through the multi-view generation structure; Conversion module: for symmetric normalization processing of the adjacency matrix through the adjacency converter, obtaining the symmetric normalized adjacency matrix; Fusion module: for multi-order information fusion of the symmetric normalized adjacency matrix, obtaining the fused adjacency matrix; Self-constraint module: for obtaining the binary mask matrix through the node feature matrix, perturbing the node feature matrix through the binary mask matrix, obtaining the new feature matrix, and obtaining the new node feature matrix through the new feature matrix and the fused adjacency matrix; Classification module: for determining the classifier through the fused adjacency matrix and the new node feature matrix; determining the total loss function according to the difference between the feature vectors of the nodes in the new node feature matrix and the node feature matrix, the cross entropy between the class probability distribution predicted by the model and the One-Hot encoding of the real label; and learning the self-supervised learning graph neural network structure through the classifier and the total loss function.

10. An electronic device, comprising: The computer program stored in the memory and executable on the processor, when the processor executes the computer program, realizes the self-supervised learning graph neural network structure learning method according to any one of claims 1-8. The computer program stored in the memory and executable on the processor, when the processor executes the computer program, realizes the self-supervised learning graph neural network structure learning method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Graph clustering method and device based on multi-order neighbor information transfer fusion clustering network

    CN114792113A

  • Network embedding method based on attention mechanism

    CN118094252A

  • Node classification method for large-scale self-supervised graph representation

    CN119832329A

  • System and method for structure learning for graph neural networks

    US20220101103A1