Graph data analysis method based on fine tuning during testing
By adopting a test-based fine-tuning method in graph data analysis, pre-training and fine-tuning the basic graph model during testing, and introducing jump coding and task gate mechanisms, the problems of large computing resource consumption and large distance information capture errors in the existing technology are solved, and efficient and accurate graph data analysis is achieved.
Patent Information
- Application Number
- CN202510050123.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art has problems such as large computing resource consumption and large distance information capture errors when processing graph data, and is inefficient when processing large-scale graph data.
The graph data analysis method based on fine-tuning during testing is adopted. By pre-training the basic graph model and fine-tuning during testing, jump encoding and adjustable jump encoding settings are introduced, and the distance information between different nodes is captured using task invariant gates and task-related gates.
It effectively reduces the computing resource consumption in the inference stage, improves the inference efficiency, reduces the error in distance information capture, is suitable for processing large-scale graph data, and has better versatility and flexibility.
Smart Images

Figure CN119962628A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer science and technology, and in particular to a graph data analysis method based on fine-tuning during testing. Background Art
[0002] In the field of computer science and technology, graph data analysis is an important technology used to process and analyze complex relational and structural data. Graph neural network (GNN) is a deep learning model specifically designed for processing graph data. GNN can automatically extract useful information from the graph by learning the features of nodes and edges in the graph. However, traditional GNNs have some limitations when processing graph data, such as they cannot effectively capture the distance information between different nodes in the graph and require a lot of computing resources in the reasoning stage. Existing solutions mainly capture the distance information in the graph by introducing hop encoding. Hop encoding is a fixed-length vector that is set as a learnable parameter according to the distance from the target node. This technique can fuse the subgraph structure encoding into the node embedding and help highlight the differences between nodes at different distances. In addition, the node embedding can be expanded by hop encoding by introducing a gating module and a projector, and converted into a language-tagged embedding space. In this way, the importance of neighbors from different hops can be adjusted according to specific downstream tasks. Although existing solutions have solved the problem of graph data analysis to a certain extent, some problems still exist. First, existing solutions require a lot of computing resources in the reasoning stage, which may be unrealistic for some application scenarios. Secondly, existing solutions may have some errors or deviations when capturing the distance information between different nodes in the graph. In addition, existing solutions may be inefficient when processing large-scale graph data. Therefore, how to design an efficient, accurate and low-resource graph data analysis method is the main problem currently facing this field. Summary of the invention
[0003] In view of the above-mentioned deficiencies in the prior art, the present invention provides a graph data analysis method based on fine-tuning during testing.
[0004] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0005] A graph data analysis method based on test-time fine-tuning includes the following steps:
[0006] S1, collect graph pre-training dataset;
[0007] S2. Build a graph basic model, and use the graph pre-training dataset to pre-train the graph basic model to obtain the pre-trained graph basic model;
[0008] S3, setting a test-time fine-tuning phase, and improving the pre-trained graph base model in the test-time fine-tuning phase to obtain an improved graph base model;
[0009] S4. Obtain a graph data set to be analyzed, and use the improved graph basic model to analyze the graph data set to be analyzed to obtain a graph data analysis result.
[0010] Furthermore, in step S1, the graph pre-training dataset includes nodes and edges; nodes represent entities in the graph, and edges represent relationships between entities.
[0011] Further, in step S3, the pre-trained graph base model is improved in the test-time fine-tuning phase to obtain an improved graph base model, including the following steps:
[0012] A1. Obtain a graph dataset with test labels, use the graph dataset with test labels, update the parameters of the pre-trained graph base model through supervised loss, and obtain a preliminary improved graph base model;
[0013] A2. Introducing skip coding and adjustable skip coding settings, and improving the preliminary improved graphics base model according to the set backbone components to obtain an improved graphics base model.
[0014] Furthermore, in step A2, jump coding and adjustable jump coding settings are introduced, and the preliminary improved graph base model is improved according to the set backbone components to obtain an improved graph base model. The specific process is: jump coding and adjustable jump coding settings are introduced, the target node and its neighborhood are encoded, and the distance information between different nodes is captured according to the set task invariant gate and task-related gate to improve the preliminary improved graph base model and obtain an improved graph base model.
[0015] The present invention has the following beneficial effects:
[0016] (1) Efficiency: The method proposed in the present invention introduces an additional parameter adjustment stage in the inference stage, namely, the fine-tuning stage during testing. This stage can effectively reduce the computing resource consumption in the inference stage and improve the inference efficiency. Compared with the existing technology, the present invention greatly speeds up the inference speed while ensuring the prediction accuracy, and is suitable for processing large-scale graph data.
[0017] (2) Accuracy: The backbone architecture of the present invention adopts an innovative design. By encoding the target node and its neighborhood and using task-invariant gates and task-dependent gates, the distance information between different nodes in the graph can be effectively captured, reducing the generation of errors and deviations. Compared with the prior art, the node embedding of the present invention is more accurate and can better reflect the structural relationship in the graph.
[0018] (3) Flexibility: The present invention uses a designed test fine-tuning phase to allow the model to be effectively fine-tuned during test time to quickly adapt to different test tasks. Compared with the prior art, the present invention has better versatility and flexibility and can be applied to graph data analysis in different fields and scenarios.
[0019] (4) Practicality: The design concept of the present invention is simple and clear, easy to implement and deploy. Compared with the existing technology, the present invention has higher practical value and feasibility in practical application;
[0020] (5) Scalability: The graph-based model of the present invention is designed with scalability in mind and can be applied to other graph data fields such as molecular graphs and transportation networks. Compared with the prior art, the present invention has a wider range of applications and better expansion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 The figure is a flowchart of a graph data analysis method based on test-time fine-tuning. DETAILED DESCRIPTION
[0022] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0023] like Figure 1 As shown, a graph data analysis method based on test-time fine-tuning includes the following steps:
[0024] S1. Collect graph pre-training dataset.
[0025] In an optional embodiment of the present invention, the graph pre-training dataset includes nodes and edges; nodes represent entities in the graph, and edges represent relationships between entities. The number of nodes and edges can be adjusted according to actual application scenarios. Specifically, the present invention uses an existing PPI (protein-protein interaction) dataset as a graph pre-training dataset.
[0026] S2. Build a graph basic model, and use the graph pre-training dataset to pre-train the graph basic model to obtain the pre-trained graph basic model.
[0027] In an optional embodiment of the present invention, the present invention constructs a graph foundation model (GFMS). GFMS is a model based on a graph neural network that can process large-scale graph data. The construction process of GFMS includes steps such as defining a model architecture and initializing model parameters. The model architecture can refer to the existing GFMS architecture, specifically GraphSAGE. After constructing GFMS, the present invention uses a graph pre-training dataset to pre-train the graph foundation model. The purpose of pre-training is to make the model parameters adapt to the structure and characteristics of the graph data to a certain extent.
[0028] S3. Setting a test-time fine-tuning phase, and improving the pre-trained graph base model in the test-time fine-tuning phase to obtain an improved graph base model.
[0029] In an optional embodiment of the present invention, the present invention improves the pre-trained graph base model in the fine-tuning stage during testing to obtain an improved graph base model, including the following steps:
[0030] A1. Obtain a graph dataset with test labels, use the graph dataset with test labels, update the parameters of the pre-trained graph base model through supervised loss, and obtain a preliminary improved graph base model.
[0031] Specifically, the present invention introduces a test-time fine-tuning phase before inference. The goal of the test-time fine-tuning phase is to update the pre-trained model parameters through supervised loss by using a small number of examples of the graph dataset with test labels. The present invention uses the existing Adam optimizer as the optimization algorithm in the adjustment process of the test-time fine-tuning phase.
[0032] A2. Introducing skip coding and adjustable skip coding settings, and improving the preliminary improved graphics base model according to the set backbone components to obtain an improved graphics base model.
[0033] The present invention introduces jump coding and adjustable jump coding settings, and improves the preliminary improved graph basic model according to the set backbone components to obtain the improved graph basic model. The specific process is: introducing jump coding and adjustable jump coding settings, encoding the target node and its neighborhood, and capturing the distance information between different nodes according to the set task invariant gate and task related gate to improve the preliminary improved graph basic model to obtain the improved graph basic model.
[0034] The present invention can capture the distance information between different nodes in the graph by introducing skip coding and adjustable skip coding settings, thereby improving the prediction accuracy. The adjustable skip coding settings can be adjusted according to the actual application scenario.
[0035] In the test-time fine-tuning stage, the present invention improves the initially improved graph base model according to the set backbone components so as to effectively process large-scale graph data.
[0036] After adjusting the model parameters in the fine-tuning phase during the test, the improved graph basic model is tested and evaluated to verify the performance of the improved graph basic model. The present invention selects the same data set as the graph pre-training data set as the test data set. The evaluation indicators include accuracy, recall rate and F1 value.
[0037] S4. Obtain a graph data set to be analyzed, and use the improved graph basic model to analyze the graph data set to be analyzed to obtain a graph data analysis result.
[0038] In an optional embodiment of the present invention, the present invention obtains the graph data set to be analyzed and inputs it into the improved graph basic model to analyze the graph data set to be analyzed, thereby obtaining the graph data analysis result.
[0039] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0040] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0041] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0042] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
[0043] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A graph data analysis method based on fine-tuning during testing, characterized in that: The following steps are involved: S1, collect graph pre-training dataset; S2. Build a graph basic model, and use the graph pre-training dataset to pre-train the graph basic model to obtain the pre-trained graph basic model; S3, setting a test-time fine-tuning phase, and improving the pre-trained graph base model in the test-time fine-tuning phase to obtain an improved graph base model; S4. Obtain a graph data set to be analyzed, and use the improved graph basic model to analyze the graph data set to be analyzed to obtain a graph data analysis result.
2. The graph data analysis method based on test-time fine-tuning according to claim 1, characterized in that: In step S1, the graph pre-training dataset contains nodes and edges; nodes represent entities in the graph, and edges represent relationships between entities.
3. The graph data analysis method based on test-time fine-tuning according to claim 1, characterized in that: In step S3, the pre-trained graph base model is improved in the test-time fine-tuning phase to obtain an improved graph base model, including the following steps: A1. Obtain a graph dataset with test labels, use the graph dataset with test labels, update the parameters of the pre-trained graph base model through supervised loss, and obtain a preliminary improved graph base model; A2. Introducing skip coding and adjustable skip coding settings, and improving the preliminary improved graphics base model according to the set backbone components to obtain an improved graphics base model.
4. The graph data analysis method based on test-time fine-tuning according to claim 3 is characterized in that: In step A2, jump coding and adjustable jump coding settings are introduced, and the preliminary improved graph base model is improved according to the set backbone components to obtain the improved graph base model. The specific process is: jump coding and adjustable jump coding settings are introduced, the target node and its neighborhood are encoded, and the distance information between different nodes is captured according to the set task invariant gate and task-related gate to improve the preliminary improved graph base model to obtain the improved graph base model.