A Self-Supervised Learning Method, System, Device and Medium for Text Attribute Graph
By introducing accessibility embedding and topology enhancement comparison learning in the text attribute graph, and integrating structural and semantic information, the problem of insufficient information integration in the existing technology is solved, and the accuracy of node classification and the self-supervised learning effect of the model is improved.
Patent Information
- Application Number
- CN202510451214.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The prior art is difficult to effectively integrate the structural and semantic information in text attribute graphs, resulting in insufficient performance in self-supervised learning, especially in low-resource scenarios.
By introducing accessibility embedding and topology enhancement comparison learning methods, combining graph neural networks and language models, global topology information and text semantic information are integrated, modal fusion node representations are generated, and a small sample node classification is performed.
The performance of downstream tasks is significantly improved, especially the accuracy of node classification under low resource conditions, and the model representation ability is improved through the supplementary and modal fusion of global topological information.
Smart Images

Figure CN119962612B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of graph neural networks and natural language processing, and particularly to a self-supervised learning method, system, device and medium for text attributed graphs. Background Art
[0002] Text Attributed Graphs (TAGs) are a powerful data modeling paradigm that combines structured graph information with unstructured text attributes and are widely used in academic citation networks and e-commerce scenarios. Nodes in text attributed graphs usually contain rich text information, such as the title, abstract, and keywords of a paper, while edges represent the relationships between nodes, such as the citation relationship of papers, making text attributed graphs not only capable of capturing the structural dependencies between entities but also encoding rich semantic information and becoming an ideal modeling tool for tasks such as node classification. In the modeling process of text attributed graphs, how to effectively integrate the topological information of the graph modality and the semantic information of the text modality is the core challenge in generating high-quality node representations. Existing research mainly focuses on three methods: prompt-based learning, supervised learning, and self-supervised learning. Prompt-based methods make full use of the potential of pre-trained language models by aligning downstream tasks with the pre-training objectives of language models. Representative methods include aligning graph and text modality embeddings through contrastive learning and using hard prompts and soft prompts to achieve few-shot node classification. However, the design of such prompts has a great impact on performance and it is difficult to process complex graph structure information. Supervised learning methods rely on large-scale labeled data to optimize the combination of graph and text information, such as realizing the interactive training of graph neural networks and language models through the variational expectation maximization framework. However, this method has limited performance in low-resource scenarios. Self-supervised learning methods mine the internal relationships between the graph and text modalities through specific pre-training tasks, such as multi-label classification and contrastive learning, but most lack a unified framework to fully integrate the global and local information of text attributed graphs. Summary of the Invention
[0003] The present invention aims to at least solve one of the technical problems existing in the related art. For this reason, the present invention provides a self-supervised learning method, system, device and medium for text attributed graphs, which can make up for the lack of global topological information by introducing reachability embeddings, can more comprehensively integrate the structural and semantic information in text attributed graphs, and can significantly improve the performance of downstream tasks through few-shot node classification based on an interactive language model.
[0004] The present invention provides a self-supervised learning method for text attributed graphs, including:
[0005] S1: Obtain a text attributed graph data set, preprocess the text attributed graph data set, and extract the node text attributes and node graph topological structure of the text attributed graph;
[0006] S2: Vectorize the node text attributes through the embedding layer to obtain the text attribute vector;
[0007] S3: Process the node graph topology through random walk to generate the reachability embedding of the nodes;
[0008] S4: Take the reachability embedding of the nodes and the node graph topology as the input of the graph neural network to obtain the node vector and the neighbor vector;
[0009] S5: Perform compression processing on the neighbor vector, and align the node vector and the compressed neighbor vector through the alignment projector to obtain the node embedding;
[0010] S6: Align the node embedding and the text attribute vector, concatenate the node vector, the compressed neighbor vector, and the text attribute vector and input them into the encoder layer to generate the modality-fused node representation;
[0011] S7: Perform few-shot node classification through the modality-fused node representation.
[0012] Furthermore, step S1 includes:
[0013] S11: Extract the node graph topology of the text attribute graph through the graph neural network;
[0014] S12: Extract the node text attributes of the text attribute graph through the pre-trained language module.
[0015] Furthermore, step S3 includes:
[0016] S31: Calculate the cumulative probability of the nodes through random walk, and the calculation expression of the cumulative probability is:
[0017]
[0018] where is the cumulative probability of reaching node from the starting node after steps of random walk, is the transition probability of reaching node from the starting node after
[0019] S32: Sort in descending order, and select the top nodes with the highest cumulative probability as the anchor set;
[0020] S33: Calculate the reachability embedding of each node to the anchor, and the calculation expression is:
[0021]
[0022] Among them, is the reachability embedding of the node , is the probability that the random walk transfers from the node to the node .
[0023] Furthermore, in step S5, select the neighbors with the largest cosine similarity to the target node, and compress the information of the neighbors into learnable query vectors. The neighbor compression process is expressed as:
[0024]
[0025] Among them, is the compressed neighbor vector, is the neighbor compression module, is the query vector, , is the first learnable query vector, is the th learnable query vector, is the neighbor vector of the node , , is the node 's first neighbor, is the node 's th neighbor, is the trainable parameter.
[0026] Furthermore, the alignment projector is a multi-layer perceptron with trainable parameters.
[0027] Furthermore, align the node embedding and the text attribute vector through the topological enhancement contrast learning module.
[0028] Furthermore, the calculation expression of the modality fusion node representation is:
[0029]
[0030] Among them, is the modality fusion node representation, is the encoder layer, is the node vector of the node , is the compressed neighbor vector, is the node The text attribute vector, is a trainable parameter of the model.
[0031] The present invention also provides a text attribute graph self-supervised learning system for performing any one of the above-described text attribute graph self-supervised learning methods, including:
[0032] An acquisition and preprocessing module, which is used to acquire a text attribute graph data set, preprocess the text attribute graph data set, and extract the node text attributes and the node graph topology structure of the text attribute graph;
[0033] A text attribute vector acquisition module, which is used to vectorize the node text attributes through an embedding layer to obtain a text attribute vector;
[0034] A random walk module, which processes the node graph topology structure through random walk to generate reachability embeddings of the nodes;
[0035] An input module, which takes the reachability embeddings of the nodes and the node graph topology structure as inputs of a graph neural network to obtain node vectors and neighbor vectors;
[0036] An alignment module, which is used to perform compression processing on the neighbor vectors, and align the node vectors and the compressed neighbor vectors through an alignment projector to obtain node embeddings;
[0037] A generation module, which is used to align the node embeddings and the text attribute vectors, splice the node vectors, the compressed neighbor vectors, and the text attribute vectors and input them into an encoder layer to generate a modality-fused node representation;
[0038] A classification module, which performs few-shot node classification through the modality-fused node representation.
[0039] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of any one of the above-described text attribute graph self-supervised learning methods are implemented.
[0040] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above-described text attribute graph self-supervised learning methods are implemented.
[0041] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects:
[0042] A self-supervised learning method, system, device and medium for text attribute graphs provided by the present invention can make up for the lack of global topological information by introducing reachability embedding, can more comprehensively integrate the structural and semantic information in the text attribute graph, and can significantly improve the performance of downstream tasks through few-shot node classification based on an interactive language model.
[0043] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0044] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the implementation examples or the prior art descriptions. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0045] Figure 1 It is a schematic flowchart of a self-supervised learning method for text attribute graphs provided by the present invention.
[0046] Figure 2 It is a schematic diagram of the model framework of a self-supervised learning method for text attribute graphs provided by the present invention.
[0047] Figure 3 It is a schematic diagram of the structure of a self-supervised learning system for text attribute graphs provided by the present invention.
[0048] Figure 4 It is the comparison result of the parameter sensitivity of a self-supervised learning method for text attribute graphs provided by the present invention.
[0049] Figure 5 It is a schematic diagram of the structure of an electronic device provided by the present invention.
[0050] Reference Signs:
[0051] 101, Acquisition and Preprocessing Module; 102, Text Attribute Vector Acquisition Module; 103, Random Walk Module; 104, Input Module; 105, Alignment Module; 106, Generation Module; 107, Classification Module; 201, Processor; 202, Communication Bus; 203, Communication Interface; 204, Memory. Detailed Embodiments
[0052] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but shall not be used to limit the scope of the present invention.
[0053] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0054] The following Figures 1 to 5 describes a self-supervised learning method, system, device, and medium for text attribute graphs of the present invention.
[0055] As Figure 1 shown, a self-supervised learning method for text attribute graphs includes:
[0056] S1: Obtain a text attribute graph data set, preprocess the text attribute graph data set, and extract the node text attributes and node graph topologies of the text attribute graph;
[0057] In some specific embodiments of the present invention, five benchmark data sets for text attribute graph node classification tasks are selected for evaluation, including the OGBN-Arxiv data set, Children data set, History data set, Computers data set, and Photo data set. The OGBN-Arxiv data set is a directed citation graph, where the nodes represent ArXiv papers in the field of computer science, and the edges represent the directed citation relationships between papers; the Children data set, History data set, Computers data set, and Photo data set are all Amazon data sets, which are used after being processed, where the nodes represent products, and the edges are formed according to the frequency of co-browsing or co-purchasing of two products.
[0058] S11: Extract the node graph topologies of the text attribute graph through a graph neural network;
[0059] S12: Extract the node text attributes of the text attribute graph through the pre-trained language module.
[0060] The pre-trained language module includes multiple stacked Transformer encoders. Through the pre-trained language module, extract the text embedding of the target node containing text attributes. The calculation expression of the text embedding of node is:
[0061]
[0062] where, is the text embedding of node , is the pre-trained language module, is the text attribute vector of node , are the trainable parameters of the model.
[0063] S2: Vectorize the node text attributes through the embedding layer to obtain the text attribute vector;
[0064] Decompose the node text attributes through the tokenizer and vectorize them through the embedding layer to obtain the text attribute vector.
[0065] S3: Process the node graph topology through random walk to generate the reachability embedding of the nodes;
[0066] Graph neural networks (GNNs) can effectively capture local neighborhood information but have difficulties in representing global topologies, resulting in poor performance in tasks that require long-range dependencies or global graph contexts. To address this issue, reachability embeddings are introduced into GNNs to enhance the model's expressive power through global topology information.
[0067] Construct graph , where is the set of nodes, is the set of edges. Each node is associated with a text information, such as a paper in an academic citation graph. Edge represents the relationship between nodes. In each step, the probability that the random walk transfers from the current node to its neighbor node is:
[0068]
[0069] where, is the probability that the random walk transfers from the current node to its neighbor node , is the node Degree;
[0070] For the path from node to node the intermediate node visited by steps of random walk is , The transition probability that steps of random walk reaches node from the starting node to node is: is:
[0071]
[0072] Wherein, is the product.
[0073] S31: Calculate the cumulative probability of nodes through random walk, and the calculation expression of the cumulative probability is:
[0074]
[0075] Wherein, is steps of random walk reaches node from the starting node the cumulative probability of, is steps of random walk reaches node from the starting node the transition probability of, is the node set; reflects the importance of node in the graph;
[0076] S32: Sort in descending order, and select the nodes with the highest cumulative probability as the anchor set; as shown in the following formula:
[0077]
[0078]
[0079] Wherein, is the anchor set, is the cumulative probability of node is the cumulative probability of node
[0080] S33: Calculate the reachability embedding of each node to the anchor, and the calculation expression is:
[0081]
[0082] Among them, is the reachability embedding of the node , is the probability that the random walk transfers from the node to the node .
[0083] Taking the reachability embedding as the input of the graph neural network effectively captures the global topological information and enhances the ability of the graph neural network to jointly model the local and global topological information of the graph through the access probability of the key anchor points.
[0084] S4: By taking the reachability embedding of the node and the node graph topology as the input of the graph neural network, obtain the node vector and the neighbor vector;
[0085] Based on the message passing mechanism, the graph neural network learns the representation of the target node by iteratively aggregating the local neighbor information. The representation of the target node is calculated by the following expression:
[0086]
[0087] Among them, is the node feature, is the edge set, is the trainable parameter of the graph neural network, is the graph neural network.
[0088] S5: Compress the neighbor vector, align the node vector and the compressed neighbor vector through the alignment projector to obtain the node embedding;
[0089] Select the neighbors with the largest cosine similarity to the target node, and compress the information of the neighbors into the learnable query vectors. The neighbor compression process is expressed as:
[0090]
[0091] Among them, is the compressed neighbor vector, is the neighbor compression module, is the query vector, , is the first learnable query vector, is the th learnable query vector, is the neighbor vector of the node , is the first neighbor of node , is the -th neighbor of node , being the trainable parameter.
[0092] The alignment projector is a multi-layer perceptron with trainable parameters,
[0093]
[0094]
[0095] wherein, is the node vector of node after alignment, is the alignment projector, is the trainable parameter of the alignment projector, is the neighbor vector of node after alignment;
[0096] The node embedding is composed of the node vector after alignment and the neighbor vector after alignment.
[0097] S6: Align the node embedding and the text attribute vector, concatenate the node vector, the compressed neighbor vector, and the text attribute vector and input them into the encoder layer to generate a modality-fused node representation;
[0098] Align the node embedding and the text embedding through topology-enhanced contrastive learning. The present invention inputs two special tokens and into the encoder layer, wherein, is the start of the graph sequence, is the end of the graph sequence, and are used to distinguish the graph sequence and the text sequence.
[0099] Concatenate the node vector, the compressed neighbor vector, the text attribute vector, the start token of the graph sequence, and the end token of the graph sequence and input them into the encoder layer to generate a modality-fused node representation. The calculation expression of the modality-fused node representation is:
[0100]
[0101] wherein, is the modality-fused node representation, is the text attribute vector of node , being the trainable parameter of the model.
[0102] S7: Perform few-shot node classification using the node representation obtained through modal fusion. Use a classifier to complete the few-shot node classification task based on the node representation obtained through modal fusion.
[0103] The process of obtaining the node representation through modal fusion is the construction process of an interaction-based language model. The framework of the interaction-based language model is as Figure 2 shown.
[0104] The present invention deeply integrates local / global graph topology information and text semantic information, providing a powerful representation for downstream tasks.
[0105] As Figure 3 shown, a text attribute graph self-supervised learning system for performing the above-mentioned text attribute graph self-supervised learning method includes:
[0106] The acquisition and preprocessing module 101 is used to acquire a text attribute graph data set, preprocess the text attribute graph data set, and extract the node text attributes and node graph topology structure of the text attribute graph;
[0107] The text attribute vector acquisition module 102 is used to vectorize the node text attributes through an embedding layer to obtain text attribute vectors;
[0108] The random walk module 103 processes the node graph topology structure through random walk to generate reachability embeddings of nodes;
[0109] The input module 104 takes the reachability embeddings of nodes and the node graph topology structure as the input of a graph neural network to obtain node vectors and neighbor vectors;
[0110] The alignment module 105 is used to perform compression processing on neighbor vectors, and align the node vectors and the compressed neighbor vectors through an alignment projector to obtain node embeddings;
[0111] The generation module 106 is used to align the node embeddings and text attribute vectors, splice the node vectors, the compressed neighbor vectors, and the text attribute vectors and input them into an encoder layer to generate node representations through modal fusion;
[0112] The classification module 107 performs few-shot node classification using the node representation obtained through modal fusion.
[0113] Through the collaborative work of the above modules, introducing reachability embeddings to make up for the lack of global topology information can more comprehensively integrate the structural and semantic information in the text attribute graph. Performing few-shot node classification through an interaction-based language model can significantly improve the performance of downstream tasks.
[0114] Design a topology-enhanced pre-training strategy and adopt a two-stage training method to train the framework.
[0115] The present invention jointly optimizes a model by combining four self-supervised learning tasks, including Topology-Enhanced Contrastive Learning (TECL), Topology-Enhanced Masked Language Modeling (TEMLM), Topology-Enhanced Node-Text Matching (TENTM), and Topology-Enhanced Knowledge Distillation (TEKD).
[0116] TECL aims to align the node embeddings obtained from a GNN with the text embeddings generated by a pre-trained language model (PLM). Node embeddings and text embeddings are usually in different spaces, resulting in limited interaction between the two modalities. TECL narrows this gap by aligning graph and text representations, enabling the model to encode topological and semantic information in a unified way. The present invention uses the Information Noise-Contrastive Estimation loss (InfoNCE loss for short) as the loss function, and the objective of TECL is:
[0117]
[0118] where, is the first TECL loss, is the second TECL loss, is the node 's text embedding, is the total TECL loss, is the node 's node vector, is the exponential function, is the batch instance, is the trainable parameter for scaling the similarity in contrastive learning, is the cosine similarity.
[0119] Masked Language Modeling (MLM) is a basic pre-training task for language models. It randomly masks certain tokens in a text sequence and then predicts these masked tokens based on the context, thereby promoting the learning of context-aware word representations and capturing semantic relationships. However, traditional MLM ignores the topological structure of the graph, which is crucial for TAGs. To overcome this limitation, the present invention designs TEMLM, which integrates the topological information of the graph into the masked token prediction. Specifically, tokens are masked at a specified masking rate, and the prediction is based on the subgraph sequence and the unmasked tokens. Let be the masked text sequence,
[0120] The loss function of TEMLM can be expressed as:
[0121]
[0122] where, is the TEMLM loss, is the set of masked tokens, is the true label of the th masked token, and
[0123] is the probability predicted by the model. Since topological information is incorporated in the TEMLM task, it should be able to more complexly facilitate better semantic representation learning. Inspired by multi-modal pre-training, the present invention proposes Topology-Enhanced Node-Text Matching (TENTM) to learn the fine-grained alignment between the graph and text modalities in TAGs. TAGs essentially represent multi-modal information: the topological structure of the graph represents structural dependencies, and the text attributes represent rich semantic information. Existing methods usually have difficulty bridging the gap between these two modalities. TENTM aligns subgraph and text representations, enabling the model to learn more cohesive and complementary embeddings, thus better performing multi-modal integration in downstream tasks. The alignment task framework of TENTM is a binary classification task aiming to judge the correct match between a 1-hop subgraph (local topology) and its associated text, in the form as follows:
[0124]
[0125]
[0126] where, is the probability predicted for node and node , is the sigmoid function, is the binary classification task head, is the Topology-Enhanced Node-Text Matching module, is the text attribute vector of node , is the TENTM loss, is the true label of node and node
[0127] Although the model learns rich interaction embeddings that integrate topology and semantics, during the interaction process, it may "forget" the key information of each modality. Therefore, the embeddings may not fully represent the topology of the graph and the semantics of the text. To address this issue, the present invention proposes Topology-Enhanced Knowledge Distillation (TEKD), which performs distillation during the training process by preserving topology, semantics, and the interaction between them. The core of the distillation task is to ensure that the output embeddings retain the similarity between the original node embeddings and text embeddings This is achieved by minimizing the cosine similarity between them. The loss of TEKD can be expressed as:
[0128]
[0129] where is the TEKD loss, is the cosine similarity. TEKD enables the model to maintain and optimize the topological, semantic, and interaction relationships between them.
[0130] In the pre-training stage, the present invention jointly optimizes the model by combining four self-supervised learning tasks. The overall training objective is to minimize the sum of the loss functions of each task, and the calculation expression is:
[0131]
[0132] where is the total loss.
[0133] In the downstream tasks, the present invention applies the pre-trained model to the node classification task. For this purpose, a new classifier is introduced, and only the query token and the classifier parameters are updated during the fine-tuning stage, and the remaining model parameters are kept frozen. In this way, the model can efficiently utilize the topological and text information learned during pre-training in downstream tasks.
[0134] In the few-shot learning scenario in downstream tasks, the performance of the node representations generated by the model is evaluated.
[0135] Pre-training part
[0136] After initializing the model parameters, under the joint optimization objective of self-supervised tasks (TECL, TEMLM, TENTM, TEKD), large-scale unlabeled text attribute graph data is used for pre-training. Among them: the maximum step length of random walk is set to 9, the number of anchor points is set to 768, the number of query tokens is set to 4, the maximum length of the 1-hop subgraph sequence is set to 32, and the maximum length of the text sequence is truncated to 128. The AdamW optimizer is adopted, the learning rate is 5e-5, and it is trained for 5 epochs.
[0137] Lightweight fine-tuning part
[0138] On the basis of completing topological enhanced contrastive learning, only the parameters of the query tokens and the newly added classifier layer are adjusted, and other parameters remain frozen. During the fine-tuning process, a small amount of labeled data (few-shot scenario) is used to train on the node classification task. The AdamW optimizer is adopted, with a learning rate of 2e-2, and at most 20 epochs are trained, and the model with the highest validation score is selected.
[0139] The experimental dataset of the present invention is shown in Table 1,
[0140] Table 1 Experimental dataset of the present invention
[0141]
[0142] To evaluate the performance of the model in low-resource scenarios, the present invention conducts experiments on few-shot node classification tasks, and the experimental results of the few-shot node classification tasks are shown in Table 2.
[0143] Table 2 Experimental results of few-shot node classification tasks
[0144]
[0145] In the few-shot setting, randomly select samples for each class, where is selected from {3, 5, 10}. For the OGBN-Arxiv dataset, Computers dataset, and Photo dataset, they are divided into training set, validation set, and test set according to time. For the Children and History datasets, they are randomly divided into training set, validation set, and test set at a ratio of 60 / 20 / 20. To ensure the reliability of the experimental results, 20 experiments are repeated on each dataset, and the average results and standard deviations are reported. GOODER is the model proposed by the present invention, and w / o RE means replacing the reachability embedding with text embeddings from bert-base-uncased. The following conclusions can be drawn:
[0146] For GNN-based methods such as GraphSAGE (Graph Sample and Aggregate, a graph neural network model for graph node embedding learning) and GCN (Graph Convolutional Network), self-supervised learning often degrades the performance of node classification. A possible reason is the gap between pre-training and downstream task objectives, which prevents pre-training from learning a suitable initialization for downstream tasks. Compared with GCN, methods like DGI (Deep Graph Infomax, a model for unsupervised graph embedding learning) and G2P2 (Graph-Grounded Pre-training and Prompting) show an average accuracy drop of 1.2% to 1.8% in different few-shot settings. For PLM-based methods, BERT (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture) and RoBERTa (Robustly Optimized BERT Pretraining Approach, an improved BERT pre-training method) generally have lower performance than GNN-based methods. RoBERTa's average accuracy is 8.9% to 9.5% lower than that of GCN. On the one hand, compared with GNN, PLM has more parameters. In few-shot settings, the existing data is not sufficient to effectively train PLM. On the other hand, in low-resource scenarios, GNN can make more effective use of limited labeled data because they not only learn the representations of individual nodes but also capture the relationships between nodes. This enables GNN to extract more information from fewer data points.
[0147] For TAG-based methods, in most cases, their performance is better than other baseline methods. This is because these methods learn information from both graph and text modalities during pre-training, obtaining more comprehensive knowledge than using GNN or PLM alone. Therefore, they provide better initialization for downstream tasks. GOODER always outperforms all baseline methods. Different from traditional TAG-based methods, GOODER integrates global topological information, effectively bridges the graph modality and the text modality, and adopts a tailored self-supervised learning algorithm to efficiently extract prior knowledge. Removing reachability embeddings (GOODER w / o RE) leads to a performance drop, especially on the Computers and Photo datasets, with an average accuracy drop of 11.5% and 5.3% respectively compared to the full GOODER model.
[0148] Tables 3 and 4 demonstrate the powerful performance of GOODER after topological-enhanced pre-training. Table 3 shows the ablation experiment results on the History dataset, where each component of GOODER is systematically ablated by removing TECL, TEMLM, TENTM, and TEKD to verify the powerful performance of GOODER; Table 4 shows the ablation experiment results on the Children dataset, where each component of GOODER is systematically ablated by removing TECL, TEMLM, TENTM, and TEKD to verify the powerful performance of GOODER;
[0149] Table 3 Ablation experiment results on the History dataset
[0150]
[0151] Table 4 Ablation experiment results on the Children dataset
[0152]
[0153] Removing any self-supervised learning method from GOODER and training with the topological-enhanced pre-training strategy will lead to a performance drop. Removing TECL causes the most significant drop because it disrupts the modality alignment and hinders other methods from effectively capturing cross-modal information. Failing to extract prior knowledge during pre-training results in poor few-shot node classification performance. Notably, removing TEKD leads to a significant performance drop, highlighting its importance in preserving modality-specific information, considering the risk of information loss during the modality interaction in GOODER.
[0154] As Figure 4 shown, the sensitivity of several key hyperparameters of GOODER was evaluated on the History dataset.
[0155] Figure 4 Figure (a) in shows the impact of the number of query vectors on the accuracy of GOODER. The number of query vectors is the number of learnable vectors used to compress neighborhood information. As the number of query vectors increases, the model performance initially improves but saturates after exceeding a certain threshold. This indicates that the optimal number of query vectors can effectively compress neighborhood information, while too many query vectors introduce redundancy and computational inefficiency. Therefore, choosing an appropriate number of query vectors is crucial for balancing performance and resource utilization.
[0156] Figure 4Figure (b) in [description] shows the impact of the number of first-order neighbors on the GOODER accuracy. The number of first-order neighbors is the number of neighbor nodes selected from the first-order subgraph constructed around the target node. Increasing the number of first-order neighbors can enhance the model's perception of the local graph topology, but excessive neighbor information will introduce noise and reduce performance. This indicates that local graph information is beneficial within a limited range, but over-reliance may hinder the model's ability to capture the global representation.
[0157] Figure 4 Figure (c) in [description] shows the impact of the number of random walk steps on the GOODER accuracy. The number of random walk steps is the number of steps taken in the random walk for constructing the reachability embedding. Adjusting the number of random walk steps directly affects the model's ability to capture the global graph topology. When the step size is too small, the model cannot fully utilize the global topology information; when the step size is too large, the model may introduce irrelevant information, resulting in performance degradation. Therefore, choosing an appropriate step size is crucial for effectively capturing the global context.
[0158] Figure 4 Figure (d) in [description] shows the impact of the number of anchor points on the GOODER accuracy. The number of anchor points is the dimension of the reachability embedding. The number of anchor points determines the model's representation of the global graph structure. An appropriate number of anchor points can help the model capture the global topology information more comprehensively. However, too many anchor points will increase the computational overhead and introduce redundancy. Therefore, in practical applications, the number of anchor points should be adjusted according to the balance between efficiency and performance.
[0159] Figure 5 An example of a block diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor 201, a communication interface 203, a memory 204, and a communication bus 202. Among them, the processor 201, the communication interface 203, and the memory 204 communicate with each other through the communication bus 202. The processor 201 can call the logical instructions in the memory 204 to execute a text attribute graph self-supervised learning method.
[0160] In addition, when the logical instructions in the above-mentioned memory 204 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0161] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute a self-supervised learning method for a text attribute graph provided by the above-mentioned various methods.
[0162] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a self-supervised learning method for a text attribute graph provided by the above-mentioned various methods.
[0163] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0164] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical discs, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.
[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A self-supervised learning method for text attribute graphs, characterized in that Including: S1: Obtain a text attribute graph dataset, preprocess the text attribute graph dataset, and extract the node text attributes and node graph topological structure of the text attribute graph; S2: Vectorize the node text attributes through an embedding layer to obtain text attribute vectors; S3: Process the node graph topological structure through random walk to generate reachability embeddings of the nodes; S31: Calculate the cumulative probability of the nodes through random walk, and the calculation expression of the cumulative probability is: Among them, is the cumulative probability that a step random walk reaches node ; is the transition probability that a step random walk reaches node ; is the set of nodes; S32: Sort in descending order and select the with the highest cumulative probability as the anchor point set; nodes as the anchor point set; S33: Calculate the reachability embedding of each node to the anchor point, and the calculation expression is: Among them, is the reachability embedding of the node ; is the probability that the random walk transfers from the node to the node ; S4: Use the reachability embeddings of the nodes and the node graph topological structure as the input of a graph neural network to obtain node vectors and neighbor vectors; S5: Compress the neighbor vectors, and align the node vectors and the compressed neighbor vectors through an alignment projector to obtain node embeddings; S6: Align the node embeddings and the text attribute vectors, concatenate the node vectors, the compressed neighbor vectors, and the text attribute vectors and input them into an encoder layer to generate a modality-fused node representation; S7: Perform few-shot node classification through the modality-fused node representation.
2. The self-supervised learning method for text attribute graph according to claim 1, wherein Step S1 includes: S11: Extract the node graph topological structure of the text attribute graph through a graph neural network; S12: Extract the node text attributes of the text attribute graph through a pre-trained language module.
3. A self-supervised learning method for text attribute graphs according to claim 1, characterized in that In step S5, select the neighbors with the largest cosine similarity to the target node, and compress the information of the neighbors into learnable query vectors. The neighbor compression process is expressed as: Among them, is the compressed neighbor vector, is the neighbor compression module, is the query vector, , is the first learnable query vector, is the m th learnable query vector, is the neighbor vector of node , , is the first neighbor of node , is the th neighbor of node , are trainable parameters.
4. A self-supervised learning method for text attribute graphs according to claim 1, characterized in that, The alignment projector is a multi-layer perceptron with trainable parameters.
5. A self-supervised learning method for text attribute graphs according to claim 1, characterized in that, Align the node embeddings and the text attribute vectors through a topology-enhanced contrast learning module; The topology-enhanced contrast learning module narrows the gap between the node embeddings and the text attribute vectors by aligning the graph and text representations, enabling the model to encode topological and semantic information in a unified manner.
6. A self-supervised learning method for text attribute graph according to claim 1, characterized in that, The calculation expression of the modality-fused node representation is: Among them, is the node representation of modal fusion, is the encoder layer, is the node 's node vector, is the compressed neighbor vector, is the node 's text attribute vector, are the trainable parameters of the model.
7. A self-supervised learning system for text attribute graphs, characterized in that For implementing a self-supervised learning method for a text attribute graph as described in any one of claims 1 to 6, including: An acquisition and preprocessing module, which is used to obtain a text attribute graph dataset, preprocess the text attribute graph dataset, and extract the node text attributes and node graph topological structure of the text attribute graph; A text attribute vector acquisition module, which is used to vectorize the node text attributes through an embedding layer to obtain text attribute vectors; A random walk module, which processes the node graph topological structure through random walk to generate reachability embeddings of the nodes; S31: Calculate the cumulative probability of the nodes through random walk, and the calculation expression of the cumulative probability is: Among them, is the cumulative probability that a step random walk reaches node , is the transition probability that a step random walk reaches node , is the set of nodes; S32: Sort in descending order, and select the nodes with the highest cumulative probability as the anchor set; S33: Calculate the reachability embedding of each node to the anchor point, and the calculation expression is: Among them, is the reachability embedding of the node ; is the probability that the random walk transfers from the node to the node ; An input module, which uses the reachability embeddings of the nodes and the node graph topological structure as the input of a graph neural network to obtain node vectors and neighbor vectors; An alignment module, which is used to compress the neighbor vectors, and align the node vectors and the compressed neighbor vectors through an alignment projector to obtain node embeddings; A generation module, which is used to align the node embeddings and the text attribute vectors, concatenate the node vectors, the compressed neighbor vectors, and the text attribute vectors and input them into an encoder layer to generate a modality-fused node representation; A classification module, which performs few-shot node classification through node representations of modality fusion.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of a text attribute graph self-supervised learning method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of a text attribute graph self-supervised learning method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Graph embedding link prediction method fusing topological structure and node attributes
CN111709474A
Text graph node classification method based on pre-training language model and depth prompt
CN119046730A