Multi-scale contrast learning method based on large model enhanced graph structure
By generating semantic edges in text attribute graphs and combining multi-scale comparison learning frameworks, the limitations of graph construction and insufficient integration of text semantics and graph topology are solved, and more powerful text attribute graph modeling and task performance improvement are achieved.
Patent Information
- Application Number
- CN202510205874.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has limitations in graph construction and insufficient fusion of text semantics and graph topology when processing text attribute graphs, resulting in insufficient information loss and generalization capabilities.
By combining the semantic understanding ability of large language models and the topological learning ability of graph neural networks, semantic edges are generated to enhance the graph structure, and trained through a multi-scale comparison learning framework to achieve efficient fusion of graph and text modalities.
It significantly improves the modeling ability and task performance of text attribute graphs, enhances the ability to mine the relationship between text features and graph structure, and improves the integrity and discriminantity of node feature representation.
Smart Images

Figure CN120124705A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of graph representation learning, data augmentation technology, and large language model technology, and particularly relates to a multi-scale contrast learning method based on a large model to enhance graph structure. Background Art
[0002] Text-Attributed Graphs (TAGs) are a type of graph structure with rich text information on nodes. In recent years, due to their unique advantage of combining the text node feature space with graph structure information, text-attributed graphs have received extensive attention. Such graphs have broad application prospects in multiple fields such as social networks, e-commerce platforms, and citation networks. Traditional Graph Neural Networks (GNNs) models, such as GCN (Graph Convolutional Networks), GAT (Graph Attention Networks), and GraphSAGE (Graph Sample and Aggregate), mainly rely on the message passing mechanism to learn features from the graph structure. These models have achieved certain success in various graph-related tasks. However, when dealing with TAG tasks (i.e., various graph learning tasks on text-attributed graphs), their capabilities are still limited, mainly due to the insufficient mining of graph structure information and the over-simplification of text feature processing.
[0003] With the rapid development of Large Language Models (LLMs) in the field of natural language processing in recent years, some studies have attempted to combine the powerful text processing capabilities of LLMs with GNNs to address the challenges in TAG tasks. These methods can be roughly divided into the following three categories: LLMs as enhancers, using LLMs to encode the original text data to generate high-quality feature representations or convert text information into node features; LLMs as predictors, converting graph data into a format that can be directly processed by language models; LLMs as aligners, mapping the graph and text modalities into a shared embedding space through LLMs to achieve their seamless interaction.
[0004] Although the existing methods have made some progress in TAG tasks, they still face the following two major challenges: 1. Limitations in graph construction: Traditional rule-based graph construction methods are difficult to comprehensively capture the information in real-world text attribute data. This is mainly manifested as problems such as topological noise caused by the imperfection in the process of converting raw data into a graph, the sparsity of the graph structure (usually caused by the power-law degree distribution characteristics of most graphs), and information loss in the overly selective attribute construction process. Even for a well-constructed graph, it may be difficult to fully support specific downstream tasks due to the weak or irrelevant correlation between edge information and task objectives.
[0005] 2. Insufficient integration of text semantics and graph topology: Existing GNN-based methods mainly focus on enhancing the internal structure of the graph while ignoring the rich information contained in text attributes. Although some recent studies have attempted to combine text and graph and achieved certain results in text processing and graph learning in low-resource scenarios, these methods usually adopt coarse-grained text processing strategies, and the graph and text encoders are designed separately, failing to fully explore the internal relationship between text features and graph structure. This limitation leads to information loss, unstable training dynamics, and insufficient generalization ability in cross-task applications.
[0006] Therefore, those skilled in the art are committed to developing a multi-scale contrast learning method based on large models to enhance the graph structure. Summary of the Invention
[0007] In view of the above defects of the prior art, the technical problem to be solved by the present invention is the limitations in graph construction and the insufficient integration of text semantics and graph topology in the prior art.
[0008] The applicant analyzes that the information in the text attribute graph consists of two parts: the graph structure and text attributes. By combining the semantic understanding ability of the large language model with the topological learning ability of the graph neural network, and using the powerful semantic understanding and text processing ability of the large language model, the semantic relationship between nodes is extracted from text attributes, and semantic edges reflecting the relative distance of nodes in the semantic space are generated as a new topological structure, which can make up for the information loss and feature deficiency caused by the fixed sampling method and specific edge generation strategy of the original graph structure, and enhance the representation learning ability of the text attribute graph; combining the semantic edges with the original sub-graph set, that is, using semantic edges to enhance the original sub-graph set (referred to as graph enhancement) to generate an enhanced sub-graph set, which provides the necessary data support for multi-scale learning. Compared with the original sub-graph set, the enhanced sub-graph set contains richer information features and a wider range of perception capabilities; multi-scale information can help the self-supervised contrast learning framework for training and optimization, align the information of these two heterogeneous modalities of graph structure and text attributes, and achieve excellent learning and task performance. Specifically as follows: 1. LLM-based graph mining and semantic edge construction: Utilize the semantic understanding ability of the large language model (LLM) to calculate the semantic relationship distance of nodes in the text feature space and generate the corresponding similarity matrix. Based on the similarity matrix, "semantic edges" are generated through graph construction methods, effectively capturing the subtle semantic relationships hidden in node texts, which are often ignored by the graph construction methods of the prior art, and are particularly suitable for processing large-scale network text data; 2. Enhance subgraph-level construction and efficient processing of the graph: After successfully constructing the semantic edges, further fuse the semantic edges with the original graph to construct an enhanced subgraph set. Different from the methods of the prior art that operate on the entire graph, the applicant conducts fusion processing at the subgraph level and uses the multi-dimensional random walk (MRW) algorithm for sampling to obtain the enhanced subgraph set, which greatly improves the computational efficiency and scalability while ensuring high accuracy; 3. Multi-scale feature learning and contrast learning framework: To efficiently fuse the graph and text modalities and fully explore their potential in downstream tasks, the applicant uses an improved self-supervised contrast learning framework to achieve precise alignment of the graph and text through multi-scale feature learning. Specifically, the applicant adopts an enhanced subgraph set graph encoder and an original subgraph set graph encoder structure based on the graph neural network (GNN). The original subgraph set graph encoder processes the original subgraph set; the enhanced subgraph set graph encoder processes the enhanced subgraph set containing semantic edges to capture richer topological information. Concatenate the features generated by the enhanced subgraph set graph encoder and the original subgraph set graph encoder structures, and perform projection mapping and adaptive fusion through a projector. In addition, for text attributes, the applicant adopts a Transformer-based architecture to generate text representation vectors and combines them with the graph representation vectors for joint optimization in the self-supervised contrast learning framework.
[0009] In one embodiment of the present invention, a multi-scale contrast learning method based on a large model-enhanced graph structure is provided, including the following steps: S100. Graph structure sampling: For the original text attribute graph, extract the graph structure and use a random process algorithm to sample the graph structure to obtain the original subgraph set; S200. Enhancement of the original subgraph set: Encode and analyze the text attributes of the original subgraph set to obtain a similarity matrix, then use a graph construction method to generate semantic edges, and use the semantic edges to enhance the original subgraph set to obtain the enhanced subgraph set; S300. Obtaining graph representation vectors: Select an enhanced subgraph set graph encoder and an original subgraph set graph encoder to encode the enhanced subgraph set and the original subgraph set respectively to obtain the enhanced subgraph set graph representation vector and the original subgraph set graph representation vector; S400. Graph representation vector splicing: splice the graph representation vector of the enhancer graph set and the graph representation vector of the original graph set, and then use a projector for mapping to obtain the graph representation vector with multi-scale features; S500. Text attribute encoding: for the text attributes of the original text attribute graph, use a text encoder to encode the text attributes to obtain the text representation vector; S600. Contrastive learning training: perform L2 regularization on the graph representation vector with multi-scale features and the text representation vector, then calculate the similarity to form a similarity matrix, calculate the cross-entropy loss function of rows and columns, add them to obtain the contrastive learning loss function, use the Adam optimizer to adjust the learning rate, and update the parameters of the enhancer graph set graph encoder, the original graph set graph encoder, the text encoder, and the projector until the loss value converges to complete the contrastive learning process; S700. Classification result testing: use the original text attribute graph for testing as the input, and obtain the final classification result through calculations by the enhancer graph set graph encoder, the original graph set graph encoder, the text encoder, and the projector.
[0010] Optionally, in the multi-scale contrastive learning method based on the large model enhanced graph structure in the above embodiment, the step S100 includes: S110. Graph structure extraction: the original text attribute graph includes the node, the edge association matrix, and the text of the node. Use the edge association matrix of the node to construct the edge relationship between nodes and extract the graph structure; S120. Graph structure sampling: use a random process algorithm to sample the graph structure to obtain the original graph set.
[0011] Further, in the multi-scale contrastive learning method based on the large model enhanced graph structure in the above embodiment, the step S120 includes: S121. Select the initial root node: randomly select multiple root nodes from the graph structure; S122. Select a specific root node: the root node is used as the starting point for subgraph sampling. According to the degree of each root node, in accordance with the probability distribution proportional to the degree, the greater the degree of the node, the higher the selection probability. Use the above probability distribution for random sampling to select a root node u ; S123. Sample neighbor nodes: randomly select an adjacent node from the neighbor nodes of the root node u , replace the root node with the adjacent node u′ , and add the root node u′ to the node set of the subgraph; u , and add the root node u to the node set of the subgraph; S124. Generate a subgraph, and loop through steps S122 - S123 until a specified number of nodes are collected, and use the node set of the subgraph to induce a subgraph; S125. Generate a set of subgraphs, and loop through the above steps S121 - S124 to generate a set number of subgraphs to form the original set of subgraphs.
[0012] Furthermore, in the multi - scale contrastive learning method based on enhancing the graph structure with a large model in the above - mentioned embodiment, adjust the number of root nodes according to the size of the original text attribute graph. The larger the original text attribute graph, the more root nodes are selected.
[0013] Furthermore, in the multi - scale contrastive learning method based on enhancing the graph structure with a large model in the above - mentioned embodiment, the range of the number of selected root nodes is 8, 16, 32, 64,..., 2 k , k which is an integer greater than 6.
[0014] Preferably, in the multi - scale contrastive learning method based on enhancing the graph structure with a large model in the above - mentioned embodiment, the number of selected root nodes is 64.
[0015] Furthermore, in the multi - scale contrastive learning method based on enhancing the graph structure with a large model in the above - mentioned embodiment, the specified number is 5% - 25% of the number of nodes in the original text attribute graph.
[0016] Preferably, in the multi - scale contrastive learning method based on enhancing the graph structure with a large model in the above - mentioned embodiment, the specified number is 10% of the number of nodes in the original text attribute graph.
[0017] Furthermore, in the multi - scale contrastive learning method based on enhancing the graph structure with a large model in the above - mentioned embodiment, the set number is selected as 2 - 5 times the number of nodes in the original text attribute graph divided by the initial number of root nodes.
[0018] Furthermore, in the multi - scale contrastive learning method based on enhancing the graph structure with a large model in the above - mentioned embodiment, when the original text attribute graph has 20000 nodes and the initial number of root nodes is 64, 1000 subgraphs are selected, and the total number of root nodes in the set of subgraphs is 64000, which is 2 - 5 times the number of nodes in the original text attribute graph.
[0019] Preferably, in the multi - scale contrastive learning method based on enhancing the graph structure with a large model in the above - mentioned embodiment, the set number is selected as 3 times the number of nodes in the original text attribute graph divided by the initial number of root nodes.
[0020] Optionally, in the multi - scale contrastive learning method based on enhancing the graph structure with a large model in the above - mentioned embodiment, the random process algorithm includes the multi - dimensional random walk (MRW) algorithm.
[0021] Optionally, in the multi-scale contrastive learning method based on a large model enhanced graph structure in any of the above embodiments, the step S200 includes: S210. Calculate the similarity matrix. Use a large language model to encode the text attributes in the original text attribute graph to obtain text encodings, and then use a similarity calculation algorithm to calculate the similarity of the text encodings to obtain a similarity matrix; S220. Generate semantic edges. Based on the similarity matrix, use a graph construction method to generate semantic edges; S230. Obtain an enhanced subgraph set. Integrate the semantic edges into the original subgraph set to obtain an enhanced subgraph set.
[0022] Optionally, in the multi-scale contrastive learning method based on a large model enhanced graph structure in any of the above embodiments, the large language model includes the Llama3 series, the Qwen2.5 series, and the GLM4 series.
[0023] Further, in the multi-scale contrastive learning method based on a large model enhanced graph structure in the above embodiment, the Llama3 series includes Llama3-8B and Llama3-70B.
[0024] Further, in the multi-scale contrastive learning method based on a large model enhanced graph structure in the above embodiment, the Qwen2.5 series includes Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-3B, Qwen2.5-7B, Qwen2.5-14B, and Qwen2.5-72B.
[0025] Further, in the multi-scale contrastive learning method based on a large model enhanced graph structure in the above embodiment, the GLM4 series includes GLM4-9B, GLM4-9B-chat, and GLM4-9B-chat-hf.
[0026] Optionally, in the multi-scale contrastive learning method based on a large model enhanced graph structure in any of the above embodiments, the similarity calculation algorithm includes the cosine similarity algorithm.
[0027] Optionally, in the multi-scale contrastive learning method based on a large model enhanced graph structure in any of the above embodiments, the graph construction method includes the K-nearest neighbor algorithm and the maximum threshold algorithm.
[0028] Further, in the multi-scale contrastive learning method based on a large model enhanced graph structure in the above embodiment, the K-nearest neighbor algorithm generates K semantic edges by selecting the K nodes with the largest similarity between a certain node and other nodes; the maximum threshold algorithm generates semantic edges by selecting the nodes with a similarity greater than the set threshold between a certain node and other nodes.
[0029] Optionally, in the multi-scale contrastive learning method based on a large model to enhance the graph structure in any of the above embodiments, step S300 includes: S310. Select a graph encoder, and select an enhanced sub-graph set graph encoder and an original sub-graph set graph encoder; S320. Encode the enhanced sub-graph set, and encode the enhanced sub-graph set through the enhanced sub-graph set graph encoder to generate an enhanced sub-graph set graph representation vector; S330. Encode the original sub-graph set, and encode the original sub-graph set through the original sub-graph set graph encoder to generate an original sub-graph set graph representation vector.
[0030] Optionally, in the multi-scale contrastive learning method based on a large model to enhance the graph structure in any of the above embodiments, the enhanced sub-graph set graph encoder and the original sub-graph set graph encoder are based on a graph neural network, GNN (Graph Neural Network).
[0031] Optionally, in the multi-scale contrastive learning method based on a large model to enhance the graph structure in any of the above embodiments, the graph neural network includes but is not limited to GCN (Graph Convolutional Network), GraphSAGE (Graph Sample and Aggregate), GAT (Graph Attention Network).
[0032] Optionally, in the multi-scale contrastive learning method based on a large model to enhance the graph structure in any of the above embodiments, step S400 includes: S410. Concatenate the graph representation vectors, concatenate the enhanced sub-graph set graph representation vector and the original sub-graph set graph representation vector in the vertical dimension to obtain a concatenated graph representation vector; S420. Linearly map the graph representation vector, and linearly map the concatenated graph representation vector using a projector to obtain a graph representation vector that fuses the multi-scale features of the enhanced sub-graph set and the original sub-graph set.
[0033] Furthermore, in the multi-scale contrastive learning method based on a large model to enhance the graph structure in the above embodiment, the projector is based on a linear neural network.
[0034] Optionally, in the multi-scale contrastive learning method based on a large model to enhance the graph structure in any of the above embodiments, the text encoder uses a Transformer structure.
[0035] Optionally, in the multi-scale contrastive learning method based on a large model to enhance the graph structure in any of the above embodiments, step S500 includes: S510. Text attribute preprocessing: Preprocess the text attributes in the original text attribute graph, convert the text attributes into embedding representations (token embeddings) according to the pre-trained vocabulary, and generate position encodings (position embeddings) at the same time. S520. Construction of text encoder input: Add the embedding representation and the position encoding to form the text encoder input, that is, the text representation. S530. Processing by multi-head attention mechanism: Input the text representation into the text encoder, and extract features from the text representation through the multi-head attention mechanism (Multi-Head Attention) to obtain the feature representation. S540. Feature representation update: Calculate the feature representation through a feed-forward neural network (Feed-Forward Neural Network, FFN) to complete non-linear transformation and feature update, and obtain the text representation vector.
[0036] Optionally, in the multi-scale contrast learning method based on the large model enhanced graph structure in any of the above embodiments, the step S600 includes: S610. Regularization processing: Perform L2 regularization on the graph representation vectors and text representation vectors of the multi-scale features, normalize the norms of the graph representation vectors and text representation vectors of each multi-scale feature to 1 to ensure consistent ranges and eliminate the influence of features at different scales. S620. Similarity calculation: Calculate the dot product similarity between the normalized graph representation vectors and text representation vectors of the multi-scale features to form a similarity matrix, where the rows represent the graph representation vectors of each multi-scale feature and the columns represent each text representation vector. S630. Cross-entropy loss calculation: Calculate the cross-entropy loss from the perspectives of the graph and the text for the similarity matrix. When calculating from the graph perspective, use the similarity scores of each row as the prediction distribution to calculate the cross-entropy loss. When calculating from the text perspective, use the similarity scores of each column as the prediction distribution to calculate the cross-entropy loss, and then add the cross-entropy losses of the rows and columns to obtain the contrast learning loss function. S640. Contrast learning training: Use the Adam optimizer to optimize the contrast learning loss function, update the parameters of the augmented sub-graph set graph encoder, the original sub-graph set graph encoder, the text encoder, and the projector, and loop through this step until the loss value of the contrast learning loss function converges to complete the contrast learning process.
[0037] Furthermore, in the multi-scale contrast learning method based on the large model enhanced graph structure in the above embodiment, the step S640 includes: S641. Parameter initialization: Initialize the parameters of the enhancer graph set graph encoder, the original graph set graph encoder, the text encoder, and the projector. S642. Loss function optimization: Set hyperparameters, use the Adam optimizer to optimize the contrastive learning loss function, and update the parameters of the enhancer graph set graph encoder, the original graph set graph encoder, the text encoder, and the projector. S643. Training completion: Loop through step S642 until the loss value of the contrastive learning loss function converges, completing the contrastive learning training process.
[0038] Further, in the multi-scale contrastive learning method based on the large model enhanced graph structure in the above embodiment, the hyperparameters include the learning rate.
[0039] Further, in the multi-scale contrastive learning method based on the large model enhanced graph structure in the above embodiment, the step S700 includes: S710. Test input: Use the original text attribute graph for testing as the input, and sequentially pass it through the enhancer graph set graph encoder, the original graph set graph encoder, the text encoder, and the projector to obtain the graph representation vector and the text representation vector of the multi-scale features. S720. Node classification probability calculation: Calculate the similarity between the graph representation vector and the text representation vector of the multi-scale features to obtain the node classification probability. S730. Classification result output: Select the node category with the highest probability as the final classification result.
[0040] The applicant effectively makes up for the deficiencies of the prior art in graph construction and the fusion of text semantics and graph topology by combining the semantic understanding ability of the large language model with the topological learning ability of the graph neural network, and significantly improves the modeling ability and task performance of the text attribute graph.
[0041] The following will further illustrate the concept, specific structure, and technical effects of the present invention with reference to the accompanying drawings to fully understand the purpose, features, and effects of the present invention. Description of the Drawings
[0042] Figure 1 is a flowchart showing the multi-scale contrastive learning method based on the large model enhanced graph structure according to an exemplary embodiment; Figure 2 is a flowchart showing the graph structure sampling according to an exemplary embodiment; Figure 3 is a flowchart showing the semantic edge enhanced original graph set according to an exemplary embodiment; Figure 4 is a flowchart showing the contrastive learning training according to an exemplary embodiment. Detailed Embodiments
[0043] The following describes multiple preferred embodiments of the present invention with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.
[0044] In the drawings, components with the same structure are denoted by the same numerical labels, and components with similar structures or functions everywhere are denoted by similar numerical labels. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. To make the illustration clearer, the thickness of some components is schematically exaggerated appropriately in the drawings.
[0045] The inventor has designed a multi-scale contrastive learning method based on a large model-enhanced graph structure, as Figure 1 shown, which includes the following steps: S100. Graph structure sampling. For the original text attribute graph, extract the graph structure, and use a random process algorithm to sample the graph structure to obtain an original subset of graphs; specifically including: S110. Graph structure extraction. The original text attribute graph contains the node, the incidence matrix of the edges, and the text of the nodes. Use the incidence matrix of the nodes and edges to construct the edge relationship between the nodes, and extract the graph structure; S120. Graph structure sampling. Use a random process algorithm to sample the graph structure to obtain an original subset of graphs. The random process algorithm includes a multi-dimensional random walk (MRW) algorithm, as Figure 2 shown, specifically including: S121. Select the initial root nodes. Randomly select multiple root nodes from the graph structure, and adjust the number of root nodes according to the size of the original text attribute graph. The larger the original text attribute graph, the more root nodes are selected. The number of selected root nodes is 64; S122. Select a specific root node. The root node is used as the starting point for subgraph sampling. According to the degree of each root node, according to the probability distribution proportional to the degree, the higher the degree of the node, the higher the selection probability. Use the above probability distribution for random sampling to select a root node u ; S123. Sample neighbor nodes. Randomly select an adjacent node u from the neighbor nodes of the root node u′ , replace the root node u′ with the adjacent node u , and add the root node u to the node set of the subgraph; S124. Generate a subgraph. Loop steps S122 - S123 until a specified number of nodes are collected. The specified number is 10% of the number of nodes in the original text attribute graph, and use the node set of the subgraph to induce a subgraph; S125. Generate a sub - atlas set. Repeat the above steps S121 - S124 in a loop to generate a set number of sub - graphs, which form the original sub - atlas set.
[0046] S200. Enhance the original sub - atlas set. Encode and analyze the text attributes of the original sub - atlas set to obtain a similarity matrix, then use a graph construction method to generate semantic edges, and use the semantic edges to enhance the original sub - atlas set to obtain an enhanced sub - atlas set. As Figure 3 shown, it specifically includes: S210. Calculate the similarity matrix. Use a large - language model to encode the text attributes in the original text attribute graph to obtain text encodings. The large - language model selects Llama3 - 8B in the Llama3 series, and then use a similarity calculation algorithm to calculate the similarity of the text encodings. The similarity calculation algorithm selects the cosine similarity algorithm to obtain the similarity matrix. S220. Generate semantic edges. Based on the similarity matrix, use a graph construction method to generate semantic edges. The graph construction method selects the K - Nearest Neighbor algorithm. By selecting the K nodes with the largest similarity between a certain node and other nodes, generate K semantic edges. S230. Obtain the enhanced sub - atlas set. Integrate the semantic edges into the original sub - atlas set to obtain the enhanced sub - atlas set.
[0047] S300. Obtain graph representation vectors. Select an enhanced sub - atlas set graph encoder and an original sub - atlas set graph encoder to encode the enhanced sub - atlas set and the original sub - atlas set respectively, and obtain the enhanced sub - atlas set graph representation vectors and the original sub - atlas set graph representation vectors respectively. Specifically include: S310. Select graph encoders. Select an enhanced sub - atlas set graph encoder and an original sub - atlas set graph encoder. The enhanced sub - atlas set graph encoder and the original sub - atlas set graph encoder are based on the Graph Neural Network (GNN), and the graph neural network selects the Graph Convolutional Network (GCN). S320. Encode the enhanced sub - atlas set. Encode the enhanced sub - atlas set through the enhanced sub - atlas set graph encoder to generate enhanced sub - atlas set graph representation vectors. S330. Encode the original sub - atlas set. Encode the original sub - atlas set through the original sub - atlas set graph encoder to generate original sub - atlas set graph representation vectors.
[0048] S400. Concatenate graph representation vectors. Concatenate the enhanced sub - atlas set graph representation vectors and the original sub - atlas set graph representation vectors, and then use a projector for mapping to obtain graph representation vectors with multi - scale features. Specifically include: S410. Graph representation vector concatenation: Concatenate the graph representation vector of the enhancer graph set and the graph representation vector of the original graph set in the vertical dimension to obtain the concatenated graph representation vector. S420. Linear mapping of graph representation vector: Use a projector to perform a linear mapping on the concatenated graph representation vector. The projector is based on a linear neural network to obtain a graph representation vector that fuses the multi-scale features of the enhancer graph set and the original graph set.
[0049] S500. Text attribute encoding: For the text attributes of the original text attribute graph, use a text encoder to encode the text attributes. The text encoder uses a Transformer structure to obtain a text representation vector. Specifically, it includes: S510. Text attribute preprocessing: Preprocess the text attributes in the original text attribute graph, convert the text attributes into embedding representations (token embeddings) according to the pre-trained vocabulary, and generate position encodings (position embeddings) at the same time. S520. Construction of text encoder input: Add the embedding representation and the position encoding to form the text encoder input, that is, the text representation. S530. Multi-head attention mechanism processing: Input the text representation into the text encoder, and perform feature extraction on the text representation through the multi-head attention mechanism (Multi-Head Attention) to obtain a feature representation. S540. Feature representation update: Calculate the feature representation through a feed-forward neural network (Feed-Forward Neural Network, FFN) to obtain a text representation vector.
[0050] S600. Contrastive learning training: Perform L2 regularization on the graph representation vector of the multi-scale features and the text representation vector, then calculate the similarity to form a similarity matrix, calculate the cross-entropy loss function of the rows and columns, add them up to obtain the contrastive learning loss function, and use the Adam optimizer to adjust the learning rate to update the parameters of the enhancer graph set graph encoder, the original graph set graph encoder, the text encoder, and the projector until the loss value converges to complete the contrastive learning process. Specifically, it includes: S610. Regularization processing: Perform L2 regularization on the graph representation vector of the multi-scale features and the text representation vector, normalize the norms of each graph representation vector of the multi-scale features and the text representation vector to 1 to ensure consistent ranges and eliminate the influence of features of different scales. S620. Similarity calculation: Calculate the dot product similarity between the normalized graph representation vector of the multi-scale features and the text representation vector to form a similarity matrix, where the rows represent each graph representation vector of the multi-scale features and the columns represent each text representation vector. S630. Cross-entropy loss calculation: Calculate the cross-entropy loss for the similarity matrix from both the graph and text perspectives. When calculating from the graph perspective, use the similarity scores of each row as the predicted distribution to calculate the cross-entropy loss. When calculating from the text perspective, use the similarity scores of each column as the predicted distribution to calculate the cross-entropy loss. Then add the cross-entropy losses of the rows and columns to obtain the contrastive learning loss function. S640. Contrastive learning training: Use the Adam optimizer to optimize the contrastive learning loss function and update the parameters of the enhancer graph set graph encoder, original graph set graph encoder, text encoder, and projector. Loop through this step until the loss value of the contrastive learning loss function converges, completing the contrastive learning process. As Figure 4 shown, it specifically includes: S641. Parameter initialization: Initialize the parameters of the enhancer graph set graph encoder, original graph set graph encoder, text encoder, and projector. S642. Loss function optimization: Set hyperparameters, including the learning rate, and use the Adam optimizer to optimize the contrastive learning loss function and update the parameters of the enhancer graph set graph encoder, original graph set graph encoder, text encoder, and projector. S643. Training completion: Loop through step S642 until the loss value of the contrastive learning loss function converges, completing the contrastive learning training process.
[0051] S700. Classification result testing: Use the original text attribute graph for testing as the input, and through calculations by the enhancer graph set graph encoder, original graph set graph encoder, text encoder, and projector, obtain the final classification result. Specifically, it includes: S710. Test input: Use the original text attribute graph for testing as the input, and sequentially pass it through the enhancer graph set graph encoder, original graph set graph encoder, text encoder, and projector to obtain the graph representation vector and text representation vector with multi-scale features. S720. Node classification probability calculation: Calculate the similarity between the graph representation vector and text representation vector with multi-scale features to obtain the node classification probability. S730. Classification result output: Select the node category with the highest probability as the final classification result.
[0052] To verify the technical effects of this patent, the applicant systematically experimentally compared and verified the effectiveness of the method in the graph node classification task. The specific experimental scheme and conclusions are as follows: First, select the graph neural network methods (GCN, GAT, GraphSAGE) of the prior art and the methods based on large language models of the prior art (Llama3 series using Llama3-8B, Llama3-70B, Qwen series using QWen2.5-7B, QWen2.5-72B) as the benchmark comparison objects in this embodiment. The same evaluation method is used in the experiment, and the performance difference is quantified through the node classification accuracy index.
[0053] The experimental results are shown in the following table. The task type selected for the experiment is the classification problem of graph nodes. Four graph datasets are sampled in the experiment, and four columns are shown in the table, namely Cora, Photo, History, and ogbn-ArXiv. Two classification accuracy indexes are selected in the experiment, namely Accuracy and Macro-F1, and the larger the two indexes, the better. Each row in the table is the benchmark method for comparison, and the last row is this embodiment.
[0054] The experimental results show that this embodiment achieves an average 40% improvement in node classification accuracy compared with the graph neural network methods of the prior art, and an average 15% improvement in accuracy compared with the methods based on large language models.
[0055] To further verify the necessity of the technical solution of this embodiment, the applicant designed an ablation experiment. The experimental results are shown in the following table. Each column represents four datasets respectively, and the rows represent the benchmark methods. Removing graph augmentation and removing multi-scale contrast learning are selected as the benchmark comparison objects of this embodiment, and the last row is this embodiment.
[0056] The experimental results show that when graph augmentation is removed, the node classification accuracy drops by an average of 7.5% under the four datasets (Cora, Photo, History, and ogbn-ArXiv); when the multi-scale contrast learning module is removed, the accuracy drops by 12%. This result proves that graph augmentation and multi-scale contrast learning have a significant synergistic effect on the contribution of technical effects.
[0057] In summary, the present invention dynamically embeds text attribute information into the graph topology space through the semantic edge enhancement technology, effectively making up for the information loss problem of the original graph structure. Combining with the multi-scale contrast learning mechanism, the efficient fusion of graph structure information and text semantic information is realized. Compared with the methods that solely rely on graph neural networks or large language models, the present invention has made substantial improvements in the integrity and discriminability of node feature representations, providing a reliable technical implementation path for text-enhanced graph learning tasks.
[0058] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art shall fall within the protection scope determined by the claims.
Claims
1. A multi-scale contrastive learning method based on large model enhanced graph structure, characterized in that: The following steps are involved: S100, graph structure sampling, extracting the graph structure from the original text attribute graph, and sampling the graph structure using a random process algorithm to obtain an original sub-graph set; S200, enhancing the original sub-atlas, encoding and analyzing the text attributes of the original sub-atlas to obtain a similarity matrix, then using a graph construction method to generate semantic edges, and using the semantic edges to enhance the original sub-atlas to obtain an enhanced sub-atlas; S300, obtaining a graph representation vector, selecting an enhanced sub-atlas graph encoder and an original sub-atlas graph encoder to respectively encode the enhanced sub-atlas and the original sub-atlas, to obtain an enhanced sub-atlas graph representation vector and an original sub-atlas graph representation vector respectively; S400, concatenating graph representation vectors, concatenating the enhanced sub-atlas graph representation vectors and the original sub-atlas graph representation vectors, and then mapping them using a projector to obtain a graph representation vector of multi-scale features; S500, text attribute encoding, encoding the text attributes of the original text attribute graph using a text encoder to obtain a text representation vector; S600, contrastive learning training, performing L2 regularization on the graph representation vector of the multi-scale feature and the text representation vector, and then performing similarity calculation to form a similarity matrix, calculating the cross entropy loss function of the rows and columns, adding them to obtain the contrastive learning loss function, using the Adam optimizer to adjust the learning rate, updating the parameters of the enhanced sub-atlas graph encoder, the original sub-atlas graph encoder, the text encoder, and the projector until the loss value converges, and completing the contrastive learning process; S700, classification result test, taking the original text attribute graph for testing as input, and calculating through the enhanced sub-atlas graph encoder, the original sub-atlas graph encoder, the text encoder and the projector to obtain the final classification result.
2. The multi-scale contrastive learning method based on large model enhanced graph structure as claimed in claim 1, characterized in that: The step S100 includes: S110, graph structure extraction, the original text attribute graph includes a node, an edge association matrix and a node text, and the node, edge relationships between nodes are constructed using the node, edge association matrix, and the graph structure is extracted; S120 , sampling the graph structure by using a random process algorithm to sample the graph structure to obtain the original sub-graph set.
3. The multi-scale contrastive learning method based on large model enhanced graph structure as claimed in claim 2, characterized in that: The step S120 includes: S121, selecting an initial root node, and randomly selecting multiple root nodes from the graph structure; S122, select a specific root node, the root node is used as the starting point of the subgraph sampling, and according to the degree of each root node, a probability distribution proportional to the degree is used. The node with a larger degree has a higher selection probability. Random sampling is performed using the probability distribution to select a root node. u ; S123, sampling neighbor nodes, from the root node u Randomly select an adjacent node from the neighboring nodes u′ , using the adjacent nodes u′ Replace the root node u , and the root node u The set of nodes added to the subgraph; S124, generating a subgraph, looping steps S122-S123 until a specified number of nodes are collected, and inducing a subgraph using the node set of the subgraph; S125 , generating a sub-atlas set, looping through steps S121 - S124 , generating a set number of sub-atlases to form the original sub-atlas set.
4. The multi-scale contrastive learning method based on large model enhanced graph structure according to claim 1, characterized in that: The step S200 includes: S210, similarity matrix calculation, using a large language model to encode the text attributes in the original text attribute graph to obtain text encoding, and then using a similarity calculation algorithm to perform similarity calculation on the text encoding to obtain a similarity matrix; S220, generating semantic edges, generating semantic edges by using the graph construction method according to the similarity matrix; S230, obtaining an enhanced sub-atlas, fusing the semantic edge into the original sub-atlas, to obtain an enhanced sub-atlas.
5. The multi-scale contrastive learning method based on large model enhanced graph structure according to claim 1, characterized in that: The step S300 includes: S310, selecting a graph encoder, selecting an enhanced sub-atlas graph encoder and an original sub-atlas graph encoder; S320, enhanced subatlas encoding, encoding the enhanced subatlas through the enhanced subatlas graph encoder to generate the enhanced subatlas graph representation vector; S330 , encoding the original sub-atlas, encoding the original sub-atlas by the original sub-atlas graph encoder to generate the original sub-atlas graph representation vector.
6. The multi-scale contrastive learning method based on large model enhanced graph structure according to claim 1, characterized in that: The step S400 includes: S410, splicing graph representation vectors, splicing the enhanced sub-graph set graph representation vectors with the original sub-graph set graph representation vectors in the vertical dimension to obtain a spliced graph representation vector; S420, linear mapping of the graph representation vector: using a projector to linearly map the spliced graph representation vector to obtain a graph representation vector that fuses the multi-scale features of the enhanced sub-atlas and the original sub-atlas.
7. The multi-scale contrastive learning method based on large model enhanced graph structure according to claim 1, characterized in that: The step S500 includes: S510, preprocessing text attributes, preprocessing the text attributes in the original text attribute graph, and converting the text attributes into embedded representations according to a pre-trained vocabulary, and generating position codes at the same time; S520, constructing a text encoder input, adding the embedding representation and the positional encoding to form a text encoder input, i.e., a text representation; S530, multi-head attention mechanism processing, inputting the text representation into the text encoder, performing feature extraction on the text representation through the multi-head attention mechanism, and obtaining a feature representation; S540, feature representation update: the feature representation is calculated through a feedforward neural network to complete nonlinear transformation and feature update to obtain a text representation vector.
8. The multi-scale contrastive learning method based on large model enhanced graph structure according to claim 1, characterized in that: The step S600 includes: S610, regularization processing, performing L2 regularization on the graph representation vector and the text representation vector of the multi-scale feature, normalizing the modulus of each graph representation vector and text representation vector of the multi-scale feature to 1, ensuring a consistent range, and eliminating the influence of features of different scales; S620, similarity calculation, calculating the dot product similarity between the normalized graph representation vector and the text representation vector of the multi-scale feature to form a similarity matrix, in which a row represents the graph representation vector of each multi-scale feature and a column represents each text representation vector; S630, calculating the cross entropy loss, calculating the cross entropy loss for the similarity matrix from the perspective of the graph and the text respectively, when calculating from the perspective of the graph, taking the similarity score of each row as the predicted distribution, calculating the cross entropy loss, when calculating from the perspective of the text, taking the similarity score of each column as the predicted distribution, calculating the cross entropy loss, and then adding the cross entropy losses of the rows and columns to obtain a contrastive learning loss function; S640, contrastive learning training, using the Adam optimizer to optimize the contrastive learning loss function, update the parameters of the enhanced sub-atlas image encoder, the original sub-atlas image encoder, the text encoder and the projector, and execute this step repeatedly until the loss value of the contrastive learning loss function converges, completing the contrastive learning process.
9. The multi-scale contrastive learning method based on large model enhanced graph structure as claimed in claim 8, characterized in that: The step S640 includes: S641, parameter initialization, initializing parameters of the enhanced sub-atlas image encoder, the original sub-atlas image encoder, the text encoder and the projector; S642, loss function optimization, setting hyperparameters, using Adam optimizer to optimize the contrastive learning loss function, and updating parameters of the enhanced sub-atlas encoder, the original sub-atlas encoder, the text encoder, and the projector; S643, the training is completed, and the step S642 is executed in a loop until the loss value of the contrastive learning loss function converges, thereby completing the contrastive learning training process.
10. The multi-scale contrastive learning method based on large model enhanced graph structure according to claim 1, characterized in that: The step S700 includes: S710, test input, taking the original text attribute graph for test as input, sequentially passing through the enhanced sub-atlas graph encoder, the original sub-atlas graph encoder, the text encoder and the projector, to obtain the graph representation vector and the text representation vector of the multi-scale feature; S720, calculating node classification probability, calculating the similarity between the graph representation vector of the multi-scale feature and the text representation vector, and obtaining the node classification probability; S730: Output the classification results, and select the node category with the highest probability as the final classification result.