Text attribute graph self-supervised learning method, system and device and medium

By introducing accessibility embedding and interaction-based language models into the text attribute graph, the problem of integrating topological and semantic information in the text attribute graph is solved, and the performance of node classification is significantly improved, especially in low-resource scenarios.

CN119962612AActive Publication Date: 2025-05-09NANKAI UNIV

Patent Information

Application Number
CN202510451214.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-09
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate the topological information of the graph modal and the semantic information of the text modal in the text attribute diagram, resulting in the low quality of the node representation, especially in low-resource scenarios.

Method used

By introducing accessibility embedding, the lack of global topological information is made up for, and combined with an interaction-based language model, the classification of few sample nodes is performed, which significantly improves the performance of downstream tasks.

Benefits of technology

It realizes a more comprehensive integration of the structural and semantic information in the text attribute graph, which significantly improves the performance of node classification, especially in low-resource scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962612A_ABST
    Figure CN119962612A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of graph neural network and natural language processing, and provides a text attribute graph self-supervised learning method, system and device and a medium, and the method comprises the steps: obtaining a text attribute graph data set, and carrying out the preprocessing; processing the graph topology structure of the node through random walk to generate reachability embedding of the node; the reachability embedding of the nodes and the graph topology structure of the nodes are used as the input of the graph neural network; aligning the node vector with the compressed neighbor vector through an alignment projector to obtain node embedding; after node embedding and text attribute vector alignment, splicing and inputting to an encoder layer to generate modal fusion node representation; and carrying out few-sample node classification through modal fusion node representation. According to the method, reachability embedding is introduced to make up for the deficiency of global topological information, few-sample node classification is carried out through an interaction-based language model, and the performance of downstream tasks is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of graph neural networks and natural language processing technology, and in particular to a text attribute graph self-supervised learning method, system, device and medium. Background Art

[0002] Text Attributed Graphs (TAGs) are a powerful data modeling paradigm that combines structured graph information with unstructured text attributes. They are widely used in academic citation networks and e-commerce scenarios. The nodes in a text attribute graph usually contain rich text information, such as the title, abstract, and keywords of a paper, while the edges represent the relationships between nodes, such as the citation relationship of a paper. This allows the text attribute graph to not only capture the structural dependencies between entities, but also encode rich semantic information, making it an ideal modeling tool for tasks such as node classification. In the modeling process of text attribute graphs, how to effectively integrate the topological information of the graph modality and the semantic information of the text modality is the core challenge of generating high-quality node representations. Existing research mainly focuses on three methods: prompt-based learning, supervised learning, and self-supervised learning. Prompt-based methods fully utilize the potential of pre-trained language models by aligning downstream tasks with the pre-training objectives of language models. Representative methods include aligning graph and text modality embeddings through contrastive learning, and using hard and soft prompts to achieve few-shot node classification. However, the design of prompts in this method has a great impact on performance and it is difficult to process complex graph structure information. Supervised learning methods rely on large-scale annotated data to optimize the combination of graph and text information, such as using the variational expectation maximization framework to achieve interactive training of graph neural networks and language models, but this method has limited performance in low-resource scenarios. Self-supervised learning methods explore the intrinsic relationship between graph and text modalities through specific pre-training tasks, such as multi-label classification and contrastive learning, but most of them lack a unified framework to fully integrate the global and local information of text attribute graphs. Summary of the invention

[0003] The present invention aims to solve at least one of the technical problems existing in the related art. To this end, the present invention provides a text attribute graph self-supervised learning method, system, device and medium, which can more comprehensively integrate the structural and semantic information in the text attribute graph by introducing reachability embedding to make up for the lack of global topological information, and perform few-sample node classification through an interactive language model, significantly improving the performance of downstream tasks.

[0004] The present invention provides a text attribute graph self-supervised learning method, comprising: S1: Obtain a text attribute graph dataset, preprocess the text attribute graph dataset, and extract the node text attributes and node graph topology structure of the text attribute graph; S2: vectorize the node text attributes through the embedding layer to obtain the text attribute vector; S3: Process the node graph topology through random walks to generate node reachability embeddings; S4: Obtain node vectors and neighbor vectors by taking the node reachability embedding and node graph topology as the input of the graph neural network; S5: compress the neighbor vector, align the node vector and the compressed neighbor vector through the alignment projector to obtain the node embedding; S6: Align the node embedding and text attribute vector, concatenate the node vector, compressed neighbor vector, and text attribute vector and input them into the encoder layer to generate the modality-fused node representation; S7: Few-shot node classification via modality-fused node representation.

[0005] Furthermore, step S1 includes: S11: Extracting the node graph topology of text attribute graph through graph neural network; S12: Extract node text attributes of the text attribute graph through the pre-trained language module.

[0006] Furthermore, step S3 includes: S31: The cumulative probability of the node is calculated by random walk. The calculation expression of the cumulative probability is: in, for Step random walk from the starting node Arrival Node The cumulative probability of for Step random walk from the starting node Arrival Node The transition probability, is a node set; S32: Sort in descending order and select the one with the highest cumulative probability Nodes are used as anchor points. S33: Calculate the reachability embedding of each node to the anchor point. The calculation expression is: in, For Node The reachability embedding, Random walk slave node Transfer to Node probability.

[0007] Furthermore, in step S5, the node with the largest cosine similarity to the target node is selected. neighbors, through the neighbor compression module The information of neighbors is compressed into In the learnable query vector, the neighbor compression process is expressed as: in, is the compressed neighbor vector, Compression module for neighbors, is the query vector, , is the first learned query vector, For the The learned query vectors, For Node Neighbor vector, , For Node The first neighbor of For Node No. Neighbors, is a trainable parameter.

[0008] Furthermore, the alignment projector is a multi-layer perceptron with trainable parameters.

[0009] Furthermore, the node embeddings and text attribute vectors are aligned via a topology-enhanced contrastive learning module.

[0010] Furthermore, the computational expression of the node representation of modal fusion is: in, is the node representation of modal fusion, is the encoder layer, For Node The node vector of is the compressed neighbor vector, For Node The text attribute vector of are the trainable parameters of the model.

[0011] The present invention also provides a text attribute graph self-supervised learning system for executing any of the above-mentioned text attribute graph self-supervised learning methods, comprising: An acquisition and preprocessing module, wherein the acquisition and preprocessing module is used to acquire a text attribute graph data set, preprocess the text attribute graph data set, and extract node text attributes and a node graph topological structure of the text attribute graph; A text attribute vector acquisition module, which is used to vectorize node text attributes through an embedding layer to obtain a text attribute vector; A random walk module, wherein the random walk module processes the topological structure of the node graph through random walks to generate reachability embeddings of the nodes; An input module, wherein the input module obtains a node vector and a neighbor vector by taking the node reachability embedding and the node graph topology structure as inputs of the graph neural network; An alignment module, wherein the alignment module is used to compress the neighbor vector, align the node vector with the compressed neighbor vector through an alignment projector, and obtain a node embedding; A generation module, which is used to align the node embedding and the text attribute vector, concatenate the node vector, the compressed neighbor vector, and the text attribute vector and input them into the encoder layer to generate a modality-fused node representation; A classification module performs few-sample node classification through modality-fused node representation.

[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the above-mentioned text attribute graph self-supervised learning methods are implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described text attribute graph self-supervised learning methods.

[0014] The above one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects: The present invention provides a method, system, device and medium for self-supervised learning of text attribute graphs. By introducing reachability embedding to compensate for the lack of global topological information, the present invention can more comprehensively integrate the structural and semantic information in the text attribute graph, and perform few-sample node classification through an interactive language model, thereby significantly improving the performance of downstream tasks.

[0015] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1It is a flow chart of a text attribute graph self-supervised learning method provided by the present invention.

[0018] Figure 2 It is a schematic diagram of a model framework of a text attribute graph self-supervised learning method provided by the present invention.

[0019] Figure 3 It is a structural schematic diagram of a text attribute graph self-supervised learning system provided by the present invention.

[0020] Figure 4 It is a comparison result of parameter sensitivity of a text attribute graph self-supervised learning method provided by the present invention.

[0021] Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention.

[0022] Reference numerals: 101. Acquisition and preprocessing module; 102. Text attribute vector acquisition module; 103. Random walk module; 104. Input module; 105. Alignment module; 106. Generation module; 107. Classification module; 201. Processor; 202. Communication bus; 203. Communication interface; 204. Memory. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical scheme and advantages of the present invention clearer, the technical scheme in the present invention will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0024] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0025] Combine the following Figures 1 to 5 A text attribute graph self-supervised learning method, system, device and medium of the present invention are described.

[0026] like Figure 1 As shown, a text attribute graph self-supervised learning method includes: S1: Obtain a text attribute graph dataset, preprocess the text attribute graph dataset, and extract the node text attributes and node graph topology structure of the text attribute graph; In some specific embodiments of the present invention, five benchmark datasets for text attribute graph node classification tasks are selected for evaluation, including the OGBN-Arxiv dataset, the Children dataset, the History dataset, the Computers dataset, and the Photo dataset. The OGBN-Arxiv dataset is a directed citation graph, in which nodes represent ArXiv papers in the field of computer science, and edges represent directed citation relationships between papers; the Children dataset, the History dataset, the Computers dataset, and the Photo dataset are all Amazon datasets, which are used after processing, in which nodes represent products, and edges are formed according to the frequency with which two products are browsed or purchased together.

[0027] S11: Extracting the node graph topology of text attribute graph through graph neural network; S12: Extract node text attributes of the text attribute graph through the pre-trained language module.

[0028] The pre-trained language module includes a multi-layer stacked Transformer encoder, which extracts The text embedding of the target node with text attributes, node The text embedding calculation expression is: in, For Node Text embedding, For pre-trained language modules, For Node The text attribute vector of are the trainable parameters of the model.

[0029] S2: vectorize the node text attributes through the embedding layer to obtain the text attribute vector; The node text attributes are decomposed by the word segmenter and vectorized by the embedding layer to obtain the text attribute vector.

[0030] S3: Process the node graph topology through random walks to generate node reachability embeddings; Graph Neural Networks (GNNs) effectively capture local neighborhood information, but have difficulty representing global topology, resulting in poor performance in tasks that require long-range dependencies or global graph context. To address this problem, reachability embedding is introduced into GNNs to enhance the model's expressiveness through global topological information.

[0031] Build the graph ,in is a node set, is the edge set, each node , associate a text information, such as a paper in an academic citation graph, edge , represents the relationship between nodes. In each step, the random walk starts from the current node Transfer to its neighbor node The probability is: in, For random walk from the current node Transfer to its neighbor node The probability of For Node degree; For slave nodes To Node Path , ,pass The intermediate nodes visited by the random walk are , Step random walk from the starting node Arrival Node The transition probability for: in, is the product.

[0032] S31: The cumulative probability of the node is calculated by random walk. The calculation expression of the cumulative probability is: in, for Step random walk from the starting node Arrival Node The cumulative probability of for Step random walk from the starting node Arrival Node The transition probability, is a node set; Reflects the node Importance in the diagram; S32: Sort in descending order and select the one with the highest cumulative probability nodes as a set of anchor points; as shown in the following formula: in, is the anchor point set, For Node The cumulative probability of For Node The cumulative probability of .

[0033] S33: Calculate the reachability embedding of each node to the anchor point. The calculation expression is: in, For Node The reachability embedding, Random walk slave node Transfer to Node probability.

[0034] Taking reachability embedding as the input of graph neural network effectively captures the global topological information, and enhances the graph neural network's ability to jointly model local and global topological information of the graph through the access probability of key anchor points.

[0035] S4: Obtain node vectors and neighbor vectors by taking the node reachability embedding and node graph topology as the input of the graph neural network; Graph neural networks are based on a message passing mechanism and learn the representation of target nodes by iteratively aggregating local neighbor information. Representation The calculation expression is: in, is the node feature, is the edge set, are the trainable parameters of the graph neural network, It is a graph neural network.

[0036] S5: compress the neighbor vector, align the node vector and the compressed neighbor vector through the alignment projector to obtain the node embedding; Select the node with the largest cosine similarity to the target node neighbors, through the neighbor compression module The information of neighbors is compressed into In the learnable query vector, the neighbor compression process is expressed as: in, is the compressed neighbor vector, Compression module for neighbors, is the query vector, , is the first learned query vector, For the The learned query vectors, For Node Neighbor vector, , For Node The first neighbor of For Node No. Neighbors, is a trainable parameter.

[0037] The alignment projector is a multi-layer perceptron with trainable parameters, in, For Node The aligned node vectors, To align the projector, To align the projector trainable parameters, For Node Aligned neighbor vectors; The node embeddings are composed of the aligned node vectors and the aligned neighbor vectors.

[0038] S6: Align the node embedding and text attribute vector, concatenate the node vector, compressed neighbor vector, and text attribute vector and input them into the encoder layer to generate the modality-fused node representation; By topologically enhanced contrastive learning, the node embedding and text embedding are aligned. and Input to the encoder layer, where is the beginning of the graph sequence, is the end of the graph sequence, as well as Used to distinguish between image sequences and text sequences.

[0039] The node vector, compressed neighbor vector, text attribute vector, start marker of the graph sequence, and end marker of the graph sequence are concatenated and input into the encoder layer to generate the node representation of modal fusion. The calculation expression of the node representation of modal fusion is: in, is the node representation of modal fusion, For Node The text attribute vector of are the trainable parameters of the model.

[0040] S7: Few-shot node classification based on modal fusion node representation. Use the classifier to complete the few-shot node classification task based on modal fusion node representation.

[0041] The process of obtaining the node representation of modal fusion is the process of constructing an interactive language model. The interactive language model framework is as follows: Figure 2 shown.

[0042] The present invention deeply integrates local / global graph topology information and text semantic information, providing powerful representation for downstream tasks.

[0043] like Figure 3 As shown, a text attribute graph self-supervised learning system is used to execute the above-mentioned text attribute graph self-supervised learning method, including: The acquisition and preprocessing module 101 is used to acquire a text attribute graph data set, preprocess the text attribute graph data set, and extract node text attributes and a node graph topology structure of the text attribute graph; The text attribute vector acquisition module 102 is used to vectorize the node text attributes through the embedding layer to obtain the text attribute vector; The random walk module 103 processes the node graph topology structure through random walk to generate reachability embedding of the node; The input module 104 obtains a node vector and a neighbor vector by taking the node reachability embedding and the node graph topology structure as the input of the graph neural network; The alignment module 105 is used to compress the neighbor vector, align the node vector with the compressed neighbor vector through the alignment projector, and obtain the node embedding; The generation module 106 is used to align the node embedding and the text attribute vector, concatenate the node vector, the compressed neighbor vector, and the text attribute vector and input them into the encoder layer to generate a modality-fused node representation; The classification module 107 performs few-sample node classification through modality fusion node representation.

[0044] Through the collaborative work of the above modules, the introduction of reachability embedding makes up for the lack of global topological information, which can more comprehensively integrate the structural and semantic information in the text attribute graph, and perform few-sample node classification through an interactive language model, significantly improving the performance of downstream tasks.

[0045] A topology-enhanced pre-training strategy is designed, and the framework is trained using a two-stage training method.

[0046] The present invention jointly optimizes the model by combining four self-supervised learning tasks, including topology enhanced contrastive learning (TECL), topology enhanced masked language modeling (TEMLM), topology enhanced node-text matching (TENTM), and topology enhanced knowledge distillation (TEKD).

[0047] TECL aims to align node embeddings obtained from GNNs with text embeddings generated by pre-trained language models (PLMs). Node embeddings and text embeddings are usually located in different spaces, resulting in limited interaction between the two modalities. TECL narrows this gap by aligning graph and text representations, allowing the model to encode topological and semantic information in a unified way. This paper uses Information Noise-Contrastive Estimation loss (InfoNCE loss for short) as the loss function. The goal of TECL is: in, For the first TECL loss, For the second TECL loss, For Node Text embedding, is the total TECL loss, For Node The node vector of is an exponential function, For batch processing instances, is a trainable parameter for scaling similarity in contrastive learning, is the cosine similarity.

[0048] Masked Language Modeling (MLM) is a basic pre-training task for language models, which randomly masks certain tokens in a text sequence and then predicts these masked tokens based on the context, thereby facilitating contextualized word representation learning and capturing semantic relationships. However, traditional MLM ignores the topological structure of the graph, which is crucial for TAGs. To overcome this limitation, we design TEMLM, which integrates the topological information of the graph into masked token prediction. Specifically, the tokens are masked at a specified masking rate, and the prediction is based on the subgraph sequence and the unmasked tokens. Let is the masked text sequence, The loss function of TEMLM can be expressed as: in, is the TEMLM loss, To mask the set of markers, For the The true labels of the masked labels, is the probability predicted by the model. Since the topological information is incorporated in the TEMLM task, it should be able to promote better semantic representation learning with greater complexity.

[0049] Inspired by multimodal pre-training, this paper proposes topology-enhanced node-text matching to learn fine-grained alignment between graph and text modalities in TAGs. TAGs essentially represent multimodal information: the topological structure of the graph represents structural dependencies, and the text attributes represent rich semantic information. Existing methods often have difficulty bridging the gap between these two modalities. TENTM aligns subgraph and text representations, enabling the model to learn more cohesive and complementary embeddings, thereby better performing multimodal integration in downstream tasks. The alignment task framework of TENTM is a binary classification task that aims to determine the correct match between a 1-hop subgraph (local topology) and its related text, and its form is as follows: in, For Node and nodes The predicted probability, is the sigmoid function, is the binary classification task head, Enhanced node-text matching module for topology, For Node The text attribute vector of For TENTM loss, For Node and nodes The true label, positive samples are 1 and negative samples are 0. During training, positive samples consist of correctly matched sub-image-text pairs, while negative samples are generated by using the cosine similarity between embeddings within a batch. Following the TECL guidance, more difficult negative samples, i.e. pairs with higher similarity, are selected.

[0050] Although the model learns rich interactive embeddings that integrate topology and semantics, it may "forget" key information of each modality during the interaction process. Therefore, the embedding may fail to fully represent the topological structure of the graph and the semantics of the text. To solve this problem, this paper proposes Topology Enhanced Knowledge Distillation (TEKD), which performs distillation during the training process by maintaining topology, semantics, and the interactions between them. The core of the distillation task is to ensure that the output embedding retains the original node embedding and text embedding This is achieved by minimizing the cosine similarity between them. The loss of TEKD can be expressed as: in, For TEKD loss, For cosine similarity, TEKD enables the model to maintain and optimize topology, semantics, and the interactions between them.

[0051] In the pre-training stage, the present invention jointly optimizes the model by combining four self-supervised learning tasks. The overall training goal is to minimize the sum of the loss functions of each task, and the calculation expression is: in, is the total loss.

[0052] In the downstream task, the present invention applies the pre-trained model to the node classification task. To this end, a new classifier is introduced , and only update the query token during the fine-tuning phase and classifier The remaining model parameters remain frozen. In this way, the model can efficiently utilize the topological and textual information learned in pre-training in downstream tasks.

[0053] We evaluate the performance of the node representations generated by the model using few-shot learning scenarios in downstream tasks.

[0054] Pre-training part After initializing the model parameters, pre-training is performed using large-scale unlabeled text attribute graph data under the joint optimization objective of self-supervised tasks (TECL, TEMLM, TENTM, TEKD). The maximum step size of random walks is set to 9, the number of anchor points is set to 768, the number of query tokens is set to 4, the maximum length of a 1-hop subgraph sequence is set to 32, and the maximum length of a text sequence is truncated to 128. The AdamW optimizer is used, the learning rate is 5e-5, and the training is performed for 5 cycles.

[0055] Lightweight fine-tuning part After completing topology enhanced contrastive learning, only the parameters of the query token and the newly added classifier layer are adjusted. The parameters of the model are set, and the other parameters remain frozen. During fine-tuning, a small amount of labeled data (few-sample scenario) is used for training on the node classification task. The AdamW optimizer is used with a learning rate of 2e-2, and a maximum of 20 cycles are trained, and the model with the highest validation score is selected.

[0056] The experimental data set of the present invention is shown in Table 1. Table 1 Experimental data set of the present invention In order to evaluate the performance of the model in low-resource scenarios, the present invention conducted experiments on the few-sample node classification task. The experimental results of the few-sample node classification task are shown in Table 2.

[0057] Table 2 Experimental results of few-sample node classification task In the few-shot setting, we randomly select samples, of which Select from {3, 5, 10}. For the OGBN-Arxiv dataset, Computers dataset and Photo dataset, they are divided into training set, validation set and test set according to time. For the Children and History datasets, they are randomly divided into training set, validation set and test set in a ratio of 60 / 20 / 20. To ensure the reliability of the experimental results, the experiment was repeated 20 times on each dataset, and the average results and standard deviations were reported. GOODER is the model proposed in this invention, w / o RE means replacing the reachability embedding with the text embedding from bert-base-uncased. The following conclusions can be drawn: For GNN-based methods, GraphSAGE (Graph Sample and Aggregate) and GCN (Graph Convolutional Network), self-supervised learning often degrades the performance of node classification. One possible reason is the gap between pre-training and downstream task objectives, which prevents pre-training from learning a suitable initialization for downstream tasks. Compared with GCN, DGI (Deep Graph Infomax) and G2P2 (Graph-Grounded Pre-training and Prompting) methods show an average accuracy drop of 1.2% to 1.8% under different few-shot settings. For PLM-based methods, BERT (Bidirectional Encoder Representations from Transformers) and RoBERTa (Robustly Optimized BERTPretraining Approach) usually perform worse than GNN-based methods. RoBERTa's average accuracy is 8.9% to 9.5% lower than GCN. On the one hand, PLMs have more parameters compared to GNNs. In the few-shot setting, the existing data is insufficient to effectively train PLMs. On the other hand, GNNs are able to more effectively utilize limited labeled data in low-resource scenarios because they not only learn representations of individual nodes but also capture the relationships between nodes. This enables GNNs to extract more information from fewer data points.

[0058] For TAG-based methods, they outperform other baseline methods in most cases. This is because these methods learn information from both graph and text modalities during pre-training, acquiring more comprehensive knowledge than GNN or PLM alone. Therefore, they provide better initialization for downstream tasks. GOODER consistently outperforms all baseline methods. Unlike traditional TAG-based methods, GOODER incorporates global topological information, effectively bridges graph and text modalities, and adopts a tailored self-supervised learning algorithm to efficiently extract prior knowledge. Removing reachability embedding (GOODER w / o RE) leads to a drop in performance, especially on Computers and Photo datasets, with an average accuracy drop of 11.5% and 5.3% compared to the full GOODER model, respectively.

[0059] Tables 3 and 4 show the powerful performance of GOODER after topology enhancement pre-training. Table 3 shows the results of ablation experiments on the History dataset, which systematically ablate GOODER components by removing TECL, TEMLM, TENTM, and TEKD to verify the powerful performance of GOODER; Table 4 shows the results of ablation experiments on the Children dataset, which systematically ablate GOODER components by removing TECL, TEMLM, TENTM, and TEKD to verify the powerful performance of GOODER; Table 3 Ablation experiment results on the History dataset Table 4 Ablation experiment results on the Children dataset Removing any self-supervised learning method from GOODER, trained with a topology-enhanced pre-training strategy, results in a drop in performance. Removing TECL results in the most significant drop as it destroys modal alignment, preventing other methods from effectively capturing cross-modal information. Failure to extract prior knowledge during pre-training leads to poor performance in few-shot node classification. Notably, removing TEKD results in a large drop in performance, highlighting its importance in preserving modality-specific information, given the risk of information loss during modal interactions in GOODER.

[0060] like Figure 4 As shown in Figure 2, the sensitivity of several key hyperparameters of GOODER is evaluated on the History dataset.

[0061] Figure 4 Figure (a) shows the effect of the number of query vectors on the accuracy of GOODER, where the number of query vectors is the number of learnable vectors used to compress neighborhood information. As the number of query vectors increases, the model performance initially improves, but tends to saturate after exceeding a certain threshold. This shows that the optimal number of query vectors can effectively compress neighborhood information, while too many query vectors introduce redundancy and computational inefficiency. Therefore, choosing the right number of query vectors is crucial to balancing performance and resource utilization.

[0062] Figure 4 Figure (b) shows the effect of the number of first-order neighbors on the accuracy of GOODER, where the number of first-order neighbors is the number of neighbor nodes selected from the first-order subgraph built around the target node. Increasing the number of first-order neighbors can improve the model's perception of the local graph topology, but too much neighbor information will introduce noise and reduce performance. This shows that local graph information is beneficial within a limited range, but over-reliance may hinder the model's ability to capture global representations.

[0063] Figure 4 Figure (c) shows the effect of random walk steps on GOODER accuracy. The random walk steps are the number of steps taken in the random walk, which are used to build the reachability embedding. Adjusting the random walk steps directly affects the model's ability to capture the global graph topology. When the step size is too small, the model cannot fully utilize the global topology information; when the step size is too large, the model may introduce irrelevant information, resulting in performance degradation. Therefore, choosing an appropriate step size is crucial to effectively capture the global context.

[0064] Figure 4 Figure (d) shows the effect of the number of anchors on the accuracy of GOODER, where the number of anchors is the dimension of reachability embedding. The number of anchors determines the model's representation of the global graph structure. An appropriate number of anchors can help the model capture global topological information more comprehensively. However, too many anchors will increase computational overhead and introduce redundancy. Therefore, in practical applications, the number of anchors should be adjusted based on the balance between efficiency and performance.

[0065] Figure 5 A block diagram of an electronic device is shown as an example. Figure 5 As shown, the electronic device may include: a processor 201 (processor), a communication interface 203 (Communications Interface), a memory 204 (memory) and a communication bus 202, wherein the processor 201, the communication interface 203, and the memory 204 communicate with each other through the communication bus 202. The processor 201 may call the logic instructions in the memory 204 to execute a text attribute graph self-supervised learning method.

[0066] In addition, the logic instructions in the above-mentioned memory 204 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0067] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute a text attribute graph self-supervised learning method provided by the above-mentioned methods.

[0068] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a text attribute graph self-supervised learning method provided by the above methods.

[0069] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0070] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A text attribute graph self-supervised learning method, characterized in that: include: S1: Obtain a text attribute graph dataset, preprocess the text attribute graph dataset, and extract the node text attributes and node graph topology structure of the text attribute graph; S2: vectorize the node text attributes through the embedding layer to obtain the text attribute vector; S3: Process the node graph topology through random walks to generate node reachability embeddings; S4: Obtain node vectors and neighbor vectors by taking the node reachability embedding and node graph topology as the input of the graph neural network; S5: compress the neighbor vector, align the node vector and the compressed neighbor vector through the alignment projector to obtain the node embedding; S6: Align the node embedding and text attribute vector, concatenate the node vector, compressed neighbor vector, and text attribute vector and input them into the encoder layer to generate the modality-fused node representation; S7: Few-shot node classification via modality-fused node representation.

2. A text attribute graph self-supervised learning method according to claim 1, characterized in that: The S1 step includes: S11: Extracting the node graph topology of text attribute graph through graph neural network; S12: Extract node text attributes of the text attribute graph through the pre-trained language module.

3. A text attribute graph self-supervised learning method according to claim 1, characterized in that: The S3 steps include: S31: The cumulative probability of the node is calculated by random walk. The calculation expression of the cumulative probability is: in, for Step random walk from the starting node Arrival Node The cumulative probability of for Step random walk from the starting node Arrival Node The transition probability, is a node set; S32: Sort in descending order and select the one with the highest cumulative probability Nodes are used as anchor points. S33: Calculate the reachability embedding of each node to the anchor point. The calculation expression is: in, For Node The reachability embedding, Random walk slave node Transfer to Node probability.

4. A text attribute graph self-supervised learning method according to claim 1, characterized in that: In step S5, select the node with the largest cosine similarity to the target node neighbors, through the neighbor compression module The information of neighbors is compressed into In the learnable query vector, the neighbor compression process is expressed as: in, is the compressed neighbor vector, Compression module for neighbors, is the query vector, , is the first learned query vector, For the m The learned query vectors, For Node Neighbor vector, , For Node The first neighbor of For Node No. Neighbors, is a trainable parameter.

5. A text attribute graph self-supervised learning method according to claim 1, characterized in that: The alignment projector is a multi-layer perceptron with trainable parameters.

6. A text attribute graph self-supervised learning method according to claim 1, characterized in that: Node embeddings and text attribute vectors are aligned via a topology-enhanced contrastive learning module.

7. A text attribute graph self-supervised learning method according to claim 1, characterized in that: The calculation expression of the node representation of modal fusion is: in, is the node representation of modal fusion, is the encoder layer, For Node The node vector of is the compressed neighbor vector, For Node The text attribute vector of are the trainable parameters of the model.

8. A text attribute graph self-supervised learning system, characterized in that: Used to perform a text attribute graph self-supervised learning method as claimed in any one of claims 1 to 7, comprising: An acquisition and preprocessing module, wherein the acquisition and preprocessing module is used to acquire a text attribute graph data set, preprocess the text attribute graph data set, and extract node text attributes and a node graph topological structure of the text attribute graph; A text attribute vector acquisition module, which is used to vectorize node text attributes through an embedding layer to obtain a text attribute vector; A random walk module, wherein the random walk module processes the topological structure of the node graph through random walks to generate reachability embeddings of the nodes; An input module, wherein the input module obtains a node vector and a neighbor vector by taking the node reachability embedding and the node graph topology structure as inputs of the graph neural network; An alignment module, wherein the alignment module is used to compress the neighbor vector, align the node vector with the compressed neighbor vector through an alignment projector, and obtain a node embedding; A generation module, which is used to align the node embedding and the text attribute vector, concatenate the node vector, the compressed neighbor vector, and the text attribute vector and input them into the encoder layer to generate a modality-fused node representation; A classification module performs few-sample node classification through modality-fused node representation.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the text attribute graph self-supervised learning method as described in any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a text attribute graph self-supervised learning method as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Network link prediction method based on multiple semantic influences of multiple neighbor nodes

    CN110851491A

  • Graph embedding link prediction method fusing topological structure and node attributes

    CN111709474A

  • Text graph node classification method based on pre-training language model and depth prompt

    CN119046730A

  • Carbon energy flow tracing optimization method and system

    CN119295109A

Cited By

  • Different matching text attribute graph node classification method and system based on optimal transmission

    CN120973943A