Graph self-supervised learning method based on combination of mask strategy and contrast learning
By combining masking strategies and comparison learning methods, a graph self-supervised learning method was designed, which solved the problem that existing methods were difficult to fully tap the potential of graph data, and achieved stronger graph representation learning effect and generalization ability.
Patent Information
- Application Number
- CN202510015581.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-13
AI Technical Summary
Existing self-supervised learning methods are difficult to fully tap the potential of graph data, especially in designing effective self-supervised tasks and processing heterogeneity and dynamics of graph data.
A graph self-supervised learning method based on a combination of masking strategy and contrast learning is adopted to realize self-supervised learning of graph data through data augmentation, masking strategy and joint optimization of comparison loss and alignment loss.
This method can more effectively capture the different granularity features of graph data, improve the generalization ability and robustness of graph representation learning, and significantly improve the performance of graph neural networks in practical applications.
Smart Images

Figure CN119990239A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graph data processing, and in particular to a graph self-supervised learning method based on combining a mask strategy with contrastive learning. Background Art
[0002] With the advent of the big data era, graph data, as a complex and highly structured data form, has shown broad application prospects in multiple fields. Unlike traditional Euclidean data (such as images and text), graph data consists of nodes and edges, which can effectively represent entities and their relationships or interactions. For example, in e-commerce, graph data is used to describe the interactive relationship between users and products to achieve accurate recommendations; in the field of chemistry, molecular structures are modeled as graphs to assist in drug discovery; in the academic field, citation networks are connected through citation relationships between papers to facilitate classification and retrieval.
[0003] The notable feature of graph data is its irregularity. Each node may be connected to a different number of other nodes, which makes it difficult to directly apply traditional deep learning methods. Therefore, graph neural networks (GNNs) came into being, specifically designed for graph structures. As an important research direction of GNNs, graph representation learning is committed to extracting effective low-dimensional feature representations from graph data. These representations not only reflect the attributes and relationships of nodes, but also capture the structural characteristics of the entire graph. In order to further improve the effect of graph representation learning, the application of self-supervised learning in GNNs has received widespread attention. Self-supervised learning generates supervision signals from raw data without relying on manually labeled data, which significantly reduces the cost and workload of data annotation.
[0004] In graph data analysis, self-supervised learning is mainly divided into two categories: generative and contrastive. Generative methods learn feature representations by reconstructing input data or its transformed form, such as autoencoders and mask reconstruction networks, such as GraphMAE, etc. Contrastive methods learn discriminative features by comparing the relationship between different data samples, such as GraphCL, GCA, etc. These methods have achieved remarkable results in improving the performance of GNNs, but still face many challenges, including how to design effective self-supervised tasks to fully capture the complex structure of graph data, how to ensure the generalization ability of feature representation in a variety of downstream tasks, and how to deal with the heterogeneity and dynamics of graph data.
[0005] In summary, graph neural networks have shown strong capabilities in processing complex graph data, and the introduction of self-supervised learning has further improved the learning effect of GNN without the need for a large amount of labeled data. However, current self-supervised learning methods are difficult to fully tap the potential of graph data. Summary of the invention
[0006] The purpose of this invention is to provide a graph self-supervised learning method based on the combination of mask strategy and contrastive learning, to deeply explore the self-supervised learning method in graph neural network, to improve the effect of graph representation learning, to solve the shortcomings of existing methods, and to promote the widespread application and development of GNN in practical applications.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] A graph self-supervised learning method based on a combination of mask strategy and contrastive learning includes the following steps:
[0009] Get the original image data;
[0010] Perform data enhancement processing on the original graph data to obtain an enhanced graph;
[0011] Combine the masking strategy to mask different parts of the original graph, obtain the masked graph, and obtain the adjacency matrix of the masked graph;
[0012] The original image, enhanced image and mask image are input into the feature encoder to obtain enhanced image features, original image features and mask features respectively;
[0013] Input the adjacency matrix and mask features into the regressor to obtain the reconstructed features;
[0014] Compute the contrast loss based on the original image features and the enhanced image features, and compute the alignment loss based on the original image features and the reconstructed features;
[0015] Feature encoders and regressors are trained based on contrast loss and alignment loss to achieve graph self-supervised learning.
[0016] The goal of the contrast loss is to minimize the similarity between positive sample pairs while maximizing the similarity between positive and negative samples, expressed as:
[0017]
[0018] Among them, z i represents the i-th original image feature, represents the i-th enhanced graph feature, SIM represents the similarity function, τ is the similarity scaling parameter, and N is the number of data in a training batch.
[0019] The goal of the alignment loss is to enhance the model's ability to model local features through reconstruction tasks, and to predict and restore the masked part through the feature information of the existing unmasked part, which can be expressed as:
[0020]
[0021] Among them, z i and The i-th original image feature and the i-th reconstructed feature respectively.
[0022] A total weighted loss function is constructed based on the contrast loss and the alignment loss, and the contrast loss and the alignment loss are jointly optimized. The total weighted loss function is expressed as:
[0023]
[0024] The weighting coefficient λ is adjusted according to the experiment to balance the learning of global and local features.
[0025] The training of the feature encoder and the regressor based on the contrast loss and the alignment loss is specifically: training is performed using stochastic gradient descent based on the total weighted loss function.
[0026] The data enhancement process includes:
[0027] Node deletion: Randomly delete a preset proportion of nodes to enhance the model's adaptability to changes in graph structure;
[0028] Edge deletion: Randomly delete a preset proportion of edges to simulate local connection changes in the graph;
[0029] Feature perturbation: Randomly perturb the features of nodes or edges to enhance the robustness of the model to feature changes.
[0030] The method of combining the mask strategy to mask different parts of the original graph specifically includes: masking some nodes or edges of the graph to generate a masked graph.
[0031] The type of mask can be flexibly configured according to task requirements to obtain different masking strategies.
[0032] The feature encoder adopts a graph neural network.
[0033] The original image, enhanced image and mask image are input into the feature encoder to obtain the enhanced image feature, the original image feature and the mask feature respectively, specifically:
[0034] The node and edge information of the input graph data is encoded through multiple layers of graph convolutional layers or other variants of graph neural networks to generate embedded representations of the nodes and graphs, and obtain low-dimensional feature vectors of the nodes and the representation of the graph as a whole.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1) The present invention designs a new self-supervised learning framework, which combines the masking strategy and the contrastive learning mechanism, and can model the different granularity features of the data to obtain a more generalized feature representation of the graph data. Specifically, the present invention designs contrast and alignment losses, so that the model can align the graph representation at different granularities, thereby learning the features of the graph data at different granularities. The contrast loss minimizes the difference between positive sample pairs and maximizes the difference between positive and negative samples to make the model insensitive to the global perturbation of the data, thereby learning the global features; the alignment loss requires the model to use the unmasked information to restore the masked part, thereby aligning the features of the masked part and the plaintext part, and learning the local features. A regressor is used to decouple the task of reconstructing the graph data. By minimizing the contrast and alignment losses, the model can comprehensively learn the feature representations of the graph data at different levels, thereby obtaining a more robust and effective self-supervised learning framework.
[0037] 2) The present invention uses a decoupling method to convert the reconstruction loss of the mask strategy into an alignment loss in the feature space, which reduces the complexity of the system while improving the ability to model local features. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic diagram of the data processing flow of the graph self-supervised learning method of the present invention. DETAILED DESCRIPTION
[0039] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0040] Among the current mainstream graph representation learning methods, contrastive learning methods excel in capturing the global features of graph data, but have certain deficiencies in local feature modeling. To solve this problem, this embodiment proposes a graph self-supervised learning method based on a combination of mask strategy and contrastive learning, introduces the mask strategy into the contrastive learning framework, and uses the decoupling idea to decouple the reconstruction of the mask into the alignment in the feature space, thereby obtaining a more comprehensive and robust self-supervised learning framework.
[0041] Specifically, this embodiment provides a graph self-supervised learning method based on a combination of mask strategy and contrastive learning, such as Figure 1 As shown, the following steps are included:
[0042] S1, obtain the original image data.
[0043] S2, perform data enhancement processing on the original graph data to obtain an enhanced graph.
[0044] Data augmentation techniques are applied to the input graph data. The augmented graph data includes operations such as node and edge deletion and feature perturbation. These augmentation operations can simulate various perturbations of graph data and help train the model for robustness. The specific steps are as follows:
[0045] S21, node deletion: randomly delete a preset proportion of nodes to enhance the model's adaptability to graph structure changes;
[0046] S22, edge deletion: randomly delete a preset proportion of edges to simulate local connectivity changes in the graph;
[0047] S23, feature perturbation: Randomly perturb the features of nodes or edges to enhance the robustness of the model to feature changes.
[0048] S3, mask different parts of the original graph in combination with the masking strategy to obtain a masked graph, and obtain the adjacency matrix of the masked graph.
[0049] Based on the above enhanced graph, some nodes or edges of the graph are further masked to generate a "masked graph", namely a masked graph. The masked parts are usually some node features or edge information, which will be masked and restored through the learning of subsequent models.
[0050] In this embodiment, the type of mask can be flexibly configured according to task requirements to obtain different mask strategies.
[0051] S4, the original image, enhanced image and mask image are input into the feature encoder to obtain enhanced image features, original image features and mask features respectively.
[0052] The feature encoder uses a graph neural network (GNN). The graph neural network can extract low-dimensional feature representations of nodes and the entire graph from the structure of the graph by aggregating neighborhood information.
[0053] The node and edge information of the input graph data is feature encoded through a multi-layer graph convolution layer (GCN) or other variants of graph neural networks to generate an embedding representation of the node and the graph, and obtain a low-dimensional feature vector of the node and a representation of the entire graph.
[0054] S5, input the adjacency matrix and mask features into the regressor to obtain the reconstructed features.
[0055] S6, calculates the contrast loss based on the original image features and the enhanced image features, and calculates the alignment loss based on the original image features and the reconstructed features.
[0056] The goal of the contrast loss is to minimize the similarity between positive sample pairs while maximizing the similarity between positive and negative samples. It uses the cross entropy loss based on normalized temperature scaling (NT-Xent) expressed as:
[0057]
[0058] Among them, z i represents the i-th original image feature, Represents the i-th enhanced graph feature, SIM represents the similarity function, τ is the similarity scaling parameter, which is used to control the distribution range of the similarity. The larger its value is, the less sensitive it is to the change of similarity, which will make the distribution curve of the feature smoother. The smaller its value is, the more fitting the feature distribution curve will be. N is the number of data in a training batch.
[0059] The goal of the alignment loss is to enhance the model's ability to model local features through reconstruction tasks. By using the feature information of the existing unmasked part, the masked part is predicted and restored. By minimizing the alignment loss, the model can learn the local structural information of the graph. The alignment loss is expressed as:
[0060]
[0061] Among them, z i and The i-th original image feature and the i-th reconstructed feature respectively.
[0062] S7, trains feature encoders and regressors based on contrast loss and alignment loss to achieve graph self-supervised learning.
[0063] In order to effectively learn the global and local features of graph data, this embodiment is based on contrast loss and alignment loss. According to task requirements, the two are combined into a total weighted loss function in a weighted manner, and the contrast loss and alignment loss are jointly optimized. The total weighted loss function is expressed as:
[0064]
[0065] The weighting coefficient λ is adjusted according to the experiment to balance the learning of global and local features.
[0066] Stochastic gradient descent (SGD) is used for training based on the total weighted loss function.
[0067] This embodiment uses social network data sets (such as IMDB, COLLAB) and bioinformatics data sets (such as PROTEINS, NCI1) to verify the effectiveness of the method of the present invention. Among them, bioinformatics data sets include MUTAG, NCI1 and PROTEINS, which mainly contain data related to molecules and proteins, and are widely used to study the relationship between chemical molecular structure and biological function, especially in drug design and molecular property prediction. In addition, social network data sets that reflect social network interactions and structures are also selected, including IMDB-BINARY, IMDB-MULTI and COLLAB. These data sets focus on the cooperative relationship between movie actors and the cooperation model between scientific researchers, respectively, and can effectively evaluate the model's ability to represent complex social structures.
[0068] This embodiment selects two classic graph kernel methods, namely Weisfeiler-Lehman subtree kernel (WL) and deep graph kernel (DGK) as the objects to be compared. Subsequently, the baseline method in the graph self-supervised learning method is also compared. Here, two methods from the generative type and the comparative type are selected respectively. Specifically, this embodiment selects the GraphMAE method based on mask modeling and the most classic graph comparative learning method GraphCL. Table 1 shows the comparison of the results of each method.
[0069] Table 1
[0070]
[0071]
[0072] Experimental results show that the method of the present invention has obvious performance advantages on most datasets, including NCI1, PROTEINS, IMDB-BINARY, IMDB-MULTI and COLLAB datasets. On these datasets, the method of the present invention has achieved SOTA results on various tasks, showing stronger generalization ability and adaptability to diverse graph data. This includes all social datasets, and given that these datasets capture the inherent complex structure and relationships in social networks, the excellent performance achieved by DecoupledLocalGCL highlights its ability to understand these relationships and the complexity of social network data.
[0073] As for bioinformatics datasets, the present invention achieved SOTA results on NCI1 and PROTEINS datasets, and on the MUTAG dataset, although the method of the present invention was slightly inferior to GraphMAE, the performance gap was very small. The excellent performance of the GraphMAE method on this dataset can be attributed to the dependence of the molecular functions of highly specific chemical compounds in the dataset on the local properties of the molecules. However, the method of the present invention still achieved competitive results on small-scale, highly specific chemical datasets.
[0074] In general, the present invention performs well in a variety of tasks, especially in unsupervised graph classification tasks, significantly improving the performance and generalization ability of the model.
[0075] In summary, the present invention first performs multiple enhancement processing on the graph data, and masks different parts of the graph data in combination with the masking strategy, so as to obtain different versions of the graph as input. These inputs are encoded by the feature encoder to obtain a low-dimensional representation of the graph in the feature space. On this basis, the present invention designs contrast and alignment losses so that the model can align the graph representation at different granularities, thereby learning the features of the graph data at different granularities. The contrast loss minimizes the difference between positive sample pairs and maximizes the difference between positive and negative samples to make the model insensitive to the global perturbation of the data, thereby learning the global features; the alignment loss requires the model to use the unmasked information to restore the masked part, thereby aligning the features of the masked part and the plaintext part, and learning the local features. Traditional masking strategies usually require the feature space to be converted back to the data space through reconstruction and the reconstruction loss is used on it, but the present invention uses a regressor to decouple the task of reconstructing the graph data, so that the model can focus more on the alignment task based on local features, reduce the system complexity, and also improve the model performance. By minimizing the contrast and alignment losses, the model can comprehensively learn the feature representations of different levels of graph data. In downstream tasks in the fields of bioinformatics and social networks, the present invention can achieve excellent performance, verifying the effectiveness of the present invention.
[0076] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A graph self-supervised learning method based on a combination of mask strategy and contrastive learning, characterized in that: The following steps are involved: Get the original image data; Perform data enhancement processing on the original graph data to obtain an enhanced graph; Combine the masking strategy to mask different parts of the original graph, obtain the masked graph, and obtain the adjacency matrix of the masked graph; The original image, enhanced image and mask image are input into the feature encoder to obtain enhanced image features, original image features and mask features respectively; Input the adjacency matrix and mask features into the regressor to obtain the reconstructed features; Compute the contrast loss based on the original image features and the enhanced image features, and compute the alignment loss based on the original image features and the reconstructed features; Feature encoders and regressors are trained based on contrast loss and alignment loss to achieve graph self-supervised learning.
2. According to claim 1, a graph self-supervised learning method based on a combination of mask strategy and contrastive learning is characterized in that: The goal of the contrast loss is to minimize the similarity between positive sample pairs while maximizing the similarity between positive and negative samples, expressed as: Among them, z i represents the i-th original image feature, represents the i-th enhanced graph feature, SIM represents the similarity function, τ is the similarity scaling parameter, and N is the number of data in a training batch.
3. According to claim 2, a graph self-supervised learning method based on a combination of mask strategy and contrastive learning is characterized in that: The goal of the alignment loss is to enhance the model's ability to model local features through reconstruction tasks, and to predict and restore the masked part through the feature information of the existing unmasked part, which can be expressed as: Among them, z i and The i-th original image feature and the i-th reconstructed feature respectively.
4. According to claim 3, a graph self-supervised learning method based on a combination of mask strategy and contrastive learning is characterized in that: A total weighted loss function is constructed based on the contrast loss and the alignment loss, and the contrast loss and the alignment loss are jointly optimized. The total weighted loss function is expressed as: The weighting coefficient λ is adjusted according to the experiment to balance the learning of global and local features.
5. According to claim 4, a graph self-supervised learning method based on a combination of mask strategy and contrastive learning is characterized in that: The training of the feature encoder and the regressor based on the contrast loss and the alignment loss is specifically: training is performed using stochastic gradient descent based on the total weighted loss function.
6. According to claim 1, a graph self-supervised learning method based on a combination of mask strategy and contrastive learning is characterized in that: The data enhancement process includes: Node deletion: Randomly delete a preset proportion of nodes to enhance the model's adaptability to changes in graph structure; Edge deletion: Randomly delete a preset proportion of edges to simulate local connection changes in the graph; Feature perturbation: Randomly perturb the features of nodes or edges to enhance the robustness of the model to feature changes.
7. According to claim 1, a graph self-supervised learning method based on a combination of mask strategy and contrastive learning is characterized in that: The method of combining the mask strategy to mask different parts of the original graph specifically includes: masking some nodes or edges of the graph to generate a masked graph.
8. According to claim 1, a graph self-supervised learning method based on a combination of mask strategy and contrastive learning is characterized in that: The type of mask can be flexibly configured according to task requirements to obtain different masking strategies.
9. According to claim 1, a graph self-supervised learning method based on a combination of mask strategy and contrastive learning is characterized in that: The feature encoder adopts a graph neural network.
10. The graph self-supervised learning method based on the combination of mask strategy and contrastive learning according to claim 1, characterized in that: The original image, enhanced image and mask image are input into the feature encoder to obtain the enhanced image feature, the original image feature and the mask feature respectively, specifically: The node and edge information of the input graph data is encoded through multiple layers of graph convolutional layers or other variants of graph neural networks to generate embedded representations of the nodes and graphs, and obtain low-dimensional feature vectors of the nodes and the representation of the graph as a whole.
Citation Information
Cited By
Water quality data prediction method and system based on topological structure evolution dynamic graph
CN120688183A
Engineering cost intelligent evaluation and auditing system based on self-supervised learning
CN121684837A