Vertical domain LLM training method and device, medium and equipment

By converting the data in the knowledge graph into a hierarchical tree structure and injecting it into the training of vertical LLM, the problems of insufficient vertical knowledge and fusion are solved, and the generalization ability of the model and the reduction of hallucinations are achieved.

CN120106175APending Publication Date: 2025-06-06RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510252992.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the process of training vertical LLM, insufficient vertical knowledge leads to severe hallucinations of the model, and it is a technical challenge to face a large amount of vertical knowledge how to deeply integrate it into the model.

Method used

By accessing the knowledge graph, the graph data is converted into a data sequence of hierarchical tree structure, and projected into the mark embedding space of the base LLM, injected into the vertical dataset, and updated the base LLM to obtain the vertical LLM.

Benefits of technology

Effectively alleviate model illusions, improve model generalization capabilities, realize deep integration of knowledge graphs and LLM, improve modeling efficiency and reduce overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106175A_ABST
    Figure CN120106175A_ABST
Patent Text Reader

Abstract

The invention provides a vertical domain LLM training method and device, a medium and equipment, and the method comprises the steps: accessing at least one knowledge graph, carrying out the graph data conversion of the data in the knowledge graph, converting the graph data into a data sequence of a hierarchical tree structure according to the nodes, edges and structures in the graph, and traversing the hierarchical tree structure, projecting the data sequence to a mark embedding space of the base LLM; reading a mark embedding space of the base LLM, obtaining a data sequence obtained based on graph data conversion, and injecting the data sequence into the vertical domain data set to obtain the vertical domain data set with the injected graph data; and updating the base LLM according to the vertical domain data set of the injection graph data to obtain a vertical domain LLM.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a vertical LLM training method, device, medium and equipment. Background Art

[0002] Large Language Model (LLM) is an important research direction and technological breakthrough in the field of natural language processing (NLP) in recent years. The core of large language models lies in their scale and complexity. They usually have billions or even trillions of parameters, which are much larger than previous models. This enables them to process and understand extremely complex language structures and semantics.

[0003] With the development of LLM, the model has shown powerful capabilities, but the computing power required for model training is also increasing. In order to save the training cost of the model, vertical domain LLM has received more and more attention. Vertical domain LLM (vertical domain LLM) refers to a model that uses industry data to train or fine-tune general LLM in a specific field (such as e-commerce, finance, education, medicine, code, mathematics, etc.) to adapt to the professional knowledge and skill requirements of a specific field. The main difference between vertical domain LLM and general LLM (General LLM) lies in their application scenarios and professional knowledge. General LLM uses a large amount of general text data for pre-training, and has cross-task versatility and cross-domain versatility. Vertical domain LLM focuses on specific fields and trains or fine-tunes them using industry data to provide more professional and practical services. The application of vertical domain LLM in specific fields can significantly improve production efficiency and reduce production costs.

[0004] However, in the process of training vertical domain LLM, on the one hand, if the vertical domain knowledge is insufficient, it will cause serious model hallucinations; on the other hand, faced with a large amount of vertical domain knowledge, how to deeply integrate it into the model is a technical challenge. Summary of the invention

[0005] In view of this, the present application provides a vertical domain LLM training method, device, medium and equipment, the main purpose of which is to deeply integrate vertical domain knowledge with LLM, reduce model hallucinations and improve model generalization capabilities.

[0006] According to one aspect of the present application, a vertical domain LLM training method is provided, comprising:

[0007] Access at least one knowledge graph, perform graph data conversion on the data in the knowledge graph, so as to convert the graph data into a data sequence of a hierarchical tree structure according to the nodes, edges, and structures in the graph, and traverse the hierarchical tree structure to project the data sequence into the tag embedding space of the base LLM;

[0008] Read the tag embedding space of the base LLM, obtain the data sequence obtained based on the graph data conversion, inject the data sequence into the vertical domain data set, and obtain the vertical domain data set injected with the graph data;

[0009] According to the vertical domain data set of the injection map data, the base LLM is updated to obtain the vertical domain LLM.

[0010] In one implementation, the step of converting the data in the knowledge graph into a data sequence of a hierarchical tree structure according to the nodes, edges, and structures in the graph includes:

[0011] Step 1: Based on the data conversion template, determine the central node of the graph data and the environmental information formed by the edges between the nodes;

[0012] Step 2: Starting from the central node, a sampling computation tree of a fixed shape is constructed according to the environmental information;

[0013] Step 3: For each hop number of neighbor nodes of the central node, define the neighbor sample size n i , where n i represents the sample size of the i-th jump;

[0014] Step 4: In the computation tree, take the central node as the root node and start from the 1-hop neighbor set To begin, randomly select n 1 nodes form a new neighbor set if The size is less than n 1 , then use placeholder nodes to supplement n 1 indivual;

[0015] Step 5: For each selected node, recursively use n 2 neighbors as child nodes of the selected node, and the insufficient node set is filled with placeholder nodes;

[0016] Step 6: Repeat steps 4 to 5 above until the central node and the neighboring nodes are converted into node sequences of fixed length.

[0017] In one implementation, a text graph learning model is used to implement the graph data conversion.

[0018] In one implementation, the method further includes:

[0019] The fine-tuning of the language model LM and the training of the graph neural network GNN are decoupled into two stages, and the text graph learning model is trained. In the first stage, for a pre-trained LM, efficient parameter fine-tuning is performed using node classification labels; in the second stage, the fine-tuned LM is used to generate node embeddings, and the node embeddings are obtained by removing the LM head layer, which can be used by any GNN for training the same task, so as to train the GNN on the generated node embeddings.

[0020] In one implementation, the method further includes:

[0021] Based on the knowledge graph, graph information extraction is performed, wherein graph information related to downstream tasks and / or graph information related to specific samples in specific tasks is extracted;

[0022] Based on the extracted graph information, the vertical domain LLM is fine-tuned.

[0023] In one implementation, graph information extraction is performed based on the knowledge graph, including:

[0024] For node-level tasks, the extracted graph information includes: the entity itself, entity attributes, and N adjacent nodes around the entity, where N is less than or equal to 2;

[0025] For edge-level tasks, the extracted graph information includes: the edge itself, the edge type information, the two nodes of the graph triple: the two vertices V1 and V2 of the edge, and the edge information of the grid determined by continuing to verify from V1 and V2;

[0026] For graph-level tasks, the extracted graph information includes the graph itself and graph type information.

[0027] In one implementation, fine-tuning the vertical domain LLM based on the extracted graph information includes:

[0028] The graph information is used as a sample to fine-tune the vertical domain LLM in the form of context learning, wherein the Prompt template is updated according to the graph information, and the updated prompt is used as the prompt text of the vertical domain LLM.

[0029] In one implementation, fine-tuning the vertical domain LLM based on the extracted graph information includes:

[0030] The graph information is used as the main sample and context, and as the training sample for the downstream multi-task learning of the vertical domain LLM, and the downstream multi-task learning model is trained to perform multi-task fine-tuning on the vertical domain LLM according to the multi-task learning model.

[0031] In one implementation, the multi-task fine-tuning of the vertical domain LLM according to the multi-task learning model includes:

[0032] Determine that the downstream multi-tasks of the multi-task learning model include node classification task, link prediction task and node description task;

[0033] Taking the task descriptions of node classification task, link prediction task and node description task as examples, the vertical domain LLM is fine-tuned in the form of contextual learning, wherein the task descriptions of node classification task, link prediction task and node description task are added to the Prompt template, and the updated prompt is used as the prompt text of the vertical domain LLM.

[0034] The present application also provides a vertical LLM training device, comprising:

[0035] A graph data conversion unit, used to access at least one knowledge graph, and perform graph data conversion on the data in the knowledge graph, so as to convert the graph data into a data sequence of a hierarchical tree structure according to the nodes, edges, and structures in the graph;

[0036] A graph data mapping unit, used to traverse the hierarchical tree structure and project the data sequence into the label embedding space of the base LLM;

[0037] A graph data injection unit is used to read the tag embedding space of the base LLM, obtain a data sequence obtained based on the graph data conversion, and inject the data sequence into the vertical domain data set to obtain the vertical domain data set injected with the graph data;

[0038] The model updating unit is used to update the base LLM according to the vertical domain data set of the injection map data to obtain the vertical domain LLM.

[0039] In one implementation, the graph data conversion unit specifically performs the following steps:

[0040] Step 1: Based on the data conversion template, determine the central node of the graph data and the environmental information formed by the edges between the nodes;

[0041] Step 2: Starting from the central node, a sampling computation tree of a fixed shape is constructed according to the environmental information;

[0042] Step 3: For each hop number of neighbor nodes of the central node, define the neighbor sample size n i , where n i represents the sample size of the i-th jump;

[0043] Step 4: In the computation tree, take the central node as the root node and start from the 1-hop neighbor set To begin, randomly select n 1nodes form a new neighbor set if The size is less than n 1 , then use placeholder nodes to supplement n 1 indivual;

[0044] Step 5: For each selected node, recursively use n 2 neighbors as child nodes of the selected node, and the insufficient node set is filled with placeholder nodes;

[0045] Step 6: Repeat steps 4 to 5 above until the central node and the neighboring nodes are converted into node sequences of fixed length.

[0046] In one implementation, the graph data conversion unit uses a text graph learning model to implement the graph data conversion.

[0047] In one implementation, it also includes:

[0048] The text graph learning model training unit is used to decouple the fine-tuning of the language model LM and the training of the graph neural network GNN into two stages, and train the text graph learning model, wherein, in the first stage, for a pre-trained LM, efficient parameter fine-tuning is performed using node classification labels; in the second stage, the fine-tuned LM is used to generate node embeddings, and the node embeddings are obtained by removing the LM head layer, which can be used by any GNN for training the same task, so as to train the GNN on the generated node embeddings.

[0049] In one implementation, the method further includes:

[0050] A graph information extraction unit, configured to extract graph information based on the knowledge graph, wherein graph information related to downstream tasks and / or graph information related to specific samples in specific tasks is extracted;

[0051] The model fine-tuning unit is used to fine-tune the vertical domain LLM based on the extracted graph information.

[0052] In one implementation, the graph information extraction unit is specifically used for, for node-level tasks, the extracted graph information includes: the entity itself, entity attributes, and N adjacent nodes around the entity, where N is less than or equal to 2; for edge-level tasks, the extracted graph information includes: the edge itself, the type information of the edge, two nodes of the graph triplet: the two vertices V1 and V2 of the edge, and an edge information of the grid determined by continuing to verify from V1 and V2; for graph-level tasks, the extracted graph information includes the graph itself and graph type information.

[0053] In one implementation, the model fine-tuning unit is specifically used to use the graph information as a sample to fine-tune the vertical domain LLM in the form of context learning, wherein the Prompt template is updated according to the graph information, and the updated prompt is used as the prompt text of the vertical domain LLM.

[0054] In one implementation, the model fine-tuning unit is specifically used to use the graph information as the main sample and context, as a training sample for the downstream multi-task learning of the vertical domain LLM, to train the downstream multi-task learning model, so as to perform multi-task fine-tuning on the vertical domain LLM according to the multi-task learning model.

[0055] In one implementation, the model fine-tuning unit is specifically used to determine that the downstream multiple tasks of the multi-task learning model include a node classification task, a link prediction task, and a node description task; taking the task descriptions of the node classification task, the link prediction task, and the node description task as examples, and fine-tuning the vertical domain LLM in the form of context learning, wherein the task descriptions of the node classification task, the link prediction task, and the node description task are added to the Prompt template, and the updated prompt is used as the prompt text of the vertical domain LLM.

[0056] According to one aspect of the present application, a storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above method when running.

[0057] According to one aspect of the present application, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the above method.

[0058] Through the above technical solution, the present application provides a vertical domain LLM training method, device, medium and equipment, which directly models knowledge from a "graph" perspective for vertical domain scenarios, injects and internalizes into the training scheme of the large language model LLM, realizes automatic injection of graphs in the vertical domain, and quickly acquires new vertical domain knowledge. Due to the authenticity and objectivity of knowledge graphs, by injecting knowledge of knowledge graphs into the model, it can effectively reduce model hallucinations and improve the generalization ability of the model. For "knowledge graph + LLM", most of the graphs are currently used as external knowledge bases of LLM, such as through RAG retrieval, or in a certain link of the application process (such as query rewriting, entity linking), etc. This application reduces the two-stage implementation method, improves modeling efficiency and reduces overhead.

[0059] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0061] Figure 1 A flowchart of a vertical domain LLM training method provided in an embodiment of the present application is shown;

[0062] Figure 2 An example schematic diagram of a vertical domain LLM training method provided in an embodiment of the present application is shown;

[0063] Figure 3 A schematic diagram of image data conversion in a vertical domain LLM training method provided in an embodiment of the present application is shown;

[0064] Figure 4 A schematic diagram of the structure of a vertical LLM training device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0065] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only embodiments of a part of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present application. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0066] First, the terms involved are explained.

[0067] 1. Large Language Model

[0068] Large Language Model (LLM) is a deep learning model with a large number of parameters, usually billions to trillions. These models extract language structures and patterns from massive text data through self-supervised learning, and can understand and generate natural language text. The core advantage of LLM is that they can handle many types of text tasks without specialized training for each task, and are widely used in natural language processing tasks such as text generation, translation, question answering, and dialogue systems.

[0069] 2. Knowledge Graph

[0070] A knowledge graph is a knowledge base organized in a graph structure, where nodes represent entities (such as people, places, events, etc.) and edges represent the relationships between entities. Knowledge graphs can store and express rich semantic information and support complex query and reasoning tasks. In the business and scientific research fields, knowledge graphs are widely used in search engine optimization, recommendation systems, question-answering systems, and intelligent assistants to help machines understand and handle complex real-world problems.

[0071] 3. Knowledge Links

[0072] Knowledge linking refers to the process of connecting entities in text with corresponding entries in external knowledge bases (such as knowledge graphs). This technology helps to enhance the understanding of text. By mapping entities in text to specific information in the knowledge base, it can provide additional semantic information for natural language processing tasks and improve the accuracy and effectiveness of tasks.

[0073] 4. Level Traversal BFS

[0074] BFS (Breadth-First Search on Hierarchical Graphs) is an algorithm for traversing a graph with a hierarchical structure. Similar to the traditional breadth-first search, it visits nodes layer by layer in hierarchical order starting from the root node, but is particularly useful in hierarchical graphs because it can effectively explore tree-like or hierarchical data. This traversal method is particularly effective when dealing with application scenarios such as knowledge graphs, family trees, or organizational charts.

[0075] 5. ICL contextual learning

[0076] ICL (In-Context Learning) Contextual learning refers to the ability of a model to guide its output by providing examples or context without explicit fine-tuning. In ICL, the model learns tasks by observing patterns or examples in the input, rather than through the traditional supervised training process. This approach is particularly effective in large language models and can be used to solve a variety of tasks without additional training data or parameter updates.

[0077] 6. Prompt for learning

[0078] Prompt learning is a natural language processing technique in which the model generates text based on a provided prompt. The prompt can be a question, the beginning of a sentence, or a specific task description, and the model will generate relevant text responses based on these prompts. This method is widely used in dialogue systems, text generation, and question-answering systems, and can flexibly guide the model to generate output in a specific format or style.

[0079] 7. Multi-task learning

[0080] Multi-task learning (MTL) is a machine learning paradigm in which the model learns multiple related tasks simultaneously. This approach improves learning efficiency and generalization ability by sharing underlying representations or certain parameters, especially when resources are limited. Multi-task learning can promote knowledge transfer between different tasks, thereby improving overall performance and model flexibility.

[0081] For the current large vertical domain models, on the one hand, the LLM illusion is very serious, so knowledge graphs can be used to help correct it; on the other hand, how a large amount of knowledge can be applied to the training of LLM is a challenge. For example, life apps are in the vertical field of people's livelihood, involving categories of goods or services such as takeout, retail, and medicine. Therefore, they involve multiple vertical domain knowledge graphs, which are large in volume, heterogeneous, and cross different industries and categories. Therefore, the current way to obtain vertical domain LLM is to use the knowledge graph as an external knowledge base of LLM, such as through RAG retrieval, or in a certain link of the application process (such as query rewriting, entity linking). So far, the knowledge graph itself has not been deeply integrated or deeply bound with LLM.

[0082] Therefore, this application mainly considers and improves from the following aspects:

[0083] 1. In the vertical domain, knowledge graphs and big models have not been well linked, and the effect of "1+1>2" has not been achieved;

[0084] 2. From a graph perspective, directly internalize the knowledge graph and inject it into the training of LLM to build a vertical KG-LLM.

[0085] See also Figure 1 , shows a flow chart of a vertical domain LLM training method provided by an embodiment of the present application. The vertical domain LLM training method includes the following steps S101-S103.

[0086] S101: Access at least one knowledge graph, perform graph data conversion on the data in the knowledge graph, so as to convert the graph data into a data sequence of a hierarchical tree structure according to the nodes, edges, and structures in the graph, traverse the hierarchical tree structure, and project the data sequence into the tag embedding space of the base LLM.

[0087] The construction of knowledge graph mainly includes the following key links. 1. Data acquisition and preprocessing. Selecting a suitable data source is the first step in building a knowledge graph. The data source can be a public data set such as Wikipedia, Freebase, DBpedia, etc., or it can be private data such as an internal enterprise database. After acquiring the data, data cleaning is required, including removing erroneous, duplicated or incomplete information to ensure the reliability and accuracy of the data. 2. Entity recognition and relationship extraction. Entity recognition is to extract entities in the knowledge graph from text, usually relying on natural language processing technology such as named entity recognition (NER). Relationship extraction is to extract semantic relationships between entities from text. 3. Knowledge fusion. Knowledge fusion includes steps such as entity alignment, entity disambiguation and attribute alignment, aiming to ensure the consistency of data from different sources. This process is achieved through technologies such as entity alignment and reference resolution to improve data quality. 4. Knowledge storage and calculation, select the appropriate storage method according to needs, such as graph database or distributed storage. Perform knowledge reasoning and graph mining in the repository to discover implicit relationships, such as the shortest path, subgraph query, etc. 5. Knowledge application. Present knowledge to users in the most appropriate way, such as for semantic search, knowledge question answering, or recommendation systems.

[0088] Through the above steps, knowledge graphs of various sub-fields can be constructed. Taking the aforementioned life APP as an example, the vertical domain knowledge graphs involved in its LLM may include retail graphs and medical graphs, which are not limited in this embodiment of the present application.

[0089] After accessing the knowledge graph, the graph data structure needs to be converted to obtain a serialized form that LLM can understand. The knowledge graph is a knowledge base organized in a graph structure, where nodes represent entities (such as people, places, events, etc.) and edges represent the relationship between entities. Therefore, the embodiment of the present application proposes to convert the graph data into a data sequence of a hierarchical tree structure based on the nodes, edges, and structures in the graph. The specific conversion steps may include:

[0090] Step 1: Based on the data conversion template, determine the central node of the graph data and the environmental information formed by the edges between the nodes;

[0091] Step 2: Starting from the central node, a sampling computational tree of a fixed shape is constructed based on the environmental information;

[0092] Step 3: For each hop of the central node's neighbor nodes, define the neighbor sample size n i , where n i represents the sample size of the i-th jump;

[0093] Step 4: Take the central node as the root node in the computation tree and start from the 1-hop neighbor set To begin, randomly select n 1 nodes form a new neighbor set if The size is less than n 1 , then use placeholder nodes to supplement n 1 indivual;

[0094] Step 5: For each selected node, recursively use n 2 neighbors as child nodes of the selected node, and the insufficient node set is filled with placeholder nodes;

[0095] Step 6: Repeat steps 4 to 5 above until the central node and the neighboring nodes are converted into node sequences of fixed length.

[0096] Through the above steps, the graph data is converted into node sequences, and then these sequences can be mapped to the labeled embedding space of the base LLM through a versatile projector, thereby achieving in-depth understanding and processing of graph structured data.

[0097] S102: Read the tag embedding space of the base LLM, obtain the data sequence obtained based on the graph data conversion, inject the data sequence into the vertical domain dataset, and obtain the vertical domain dataset injected with the graph data.

[0098] S103: Update the base LLM according to the vertical domain data set of the injection map data to obtain the vertical domain LLM.

[0099] Among them, the base LLM is a general LLM or an initial vertical domain LLM. For example, the initial vertical domain LLM under the life APP platform, as the base LLM, injects the graph data into the data set of the base LLM through the improved method provided in the embodiment of the present application, so as to update the base LLM with the updated data set and obtain the target vertical domain LLM.

[0100] It can be seen that in the embodiments of the present application, for vertical domain scenarios, knowledge is directly modeled from a "graph" perspective, and the training scheme is injected and internalized into the large language model LLM to realize automatic injection of the graph in the vertical domain, and quickly acquire new vertical domain knowledge. Due to the true objectivity of knowledge graph knowledge, by injecting the knowledge of the knowledge graph into the model, the model illusion can be effectively reduced and the generalization ability of the model can be improved. For "knowledge graph + LLM", most of the graphs are currently used as external knowledge bases of LLM, such as through RAG retrieval, or in a certain link of the application process (such as query rewriting, entity linking), etc. This application reduces the two-stage implementation method, improves modeling efficiency, and reduces overhead.

[0101] The following is an example to illustrate the embodiments of the present application.

[0102] See also Figure 2 , shows an example schematic diagram of a vertical domain LLM training method provided in an embodiment of the present application.

[0103] Figure 2 This paper describes a process of converting graph structure data and injecting it into a large language model (LLM), especially for vertical domain scenarios. First, the whole process is explained in the following four points:

[0104] (1) Graph Perspective Modeling for Vertical Scenarios

[0105] This process is the first method to model knowledge from a graph perspective for specific vertical scenarios. This means that it is not a general graph data processing method, but a customized solution for a specific industry or field.

[0106] (2)Graph->LLM Sequence format conversion process:

[0107] A key part of the process is to convert the graph structured data into a sequence format that LLM can process. This process includes converting the nodes and edges in the graph into a hierarchical tree structure, then traversing this structure and passing the graph information to LLM in a clear and complete manner.

[0108] (3) Multi-task learning (MoE)

[0109] In order to cope with diverse downstream tasks, the process adopts multi-task learning (MoE, Mixture of Experts). This means that the model can handle different types of tasks at the same time, such as:

[0110] Node-level tasks: involve classification or other operations on individual nodes in the graph.

[0111] Edge-level tasks: involve the analysis of edges in the graph, such as link prediction.

[0112] Gaph-level tasks: involve the analysis of the entire graph, such as graph classification or description.

[0113] (4) Graph Attention and Subgraph Selection:

[0114] The concept of "subgraph selection" is defined in the process, which means that according to different training tasks, the system will select and extract relevant subgraph information. These subgraph information are provided to LLM as examples (Demonstration) and also serve as sample context for multi-task learning (MoE). This selective information extraction ensures that the model can focus on the most important part of the graph for the current task.

[0115] The key processing steps are described in detail below.

[0116] 1. Data Conversion

[0117] Data conversion can be completed using data conversion templates. The data conversion model refers to the data conversion method, not a specific text template. The purpose of data conversion is to convert graph structure data into a serialized form that can be understood by large language models (LLMs).

[0118] See also Figure 3 , shows an example of converting a graph structure into a node sequence of a hierarchical tree. Data conversion can be performed through the following steps.

[0119] Step 1: Based on the data conversion template, determine the central node of the graph data and the environmental information formed by the edges between the nodes;

[0120] Step 2: Starting from the central node, a sampling computation tree of a fixed shape is constructed according to the environmental information, for example, node A is used as the central node;

[0121] Step 3: For each hop of the central node’s neighbor nodes, define a neighbor sample size n 1 、n 2 , ...where n i represents the sample size of the i-th jump;

[0122] Step 4: The computational tree constructed takes the central node as the root node and starts from the 1-hop neighbor set To begin, randomly select n 1 nodes form a new neighbor set if The size is less than n 1 , then use placeholder nodes to supplement n 1 indivual;

[0123] Step 5: For each selected node, recursively use n 2 neighbors as child nodes of the selected node, and the insufficient node set is filled with placeholder nodes;

[0124] Step 6: Through this layer-order traversal, the detailed information of the central node and the neighboring nodes is converted into a node sequence of fixed length.

[0125] As a result, graph data can be converted into node sequences. These sequences are then mapped to the token embedding space of a large language model through a versatile projector, thereby achieving in-depth understanding and processing of graph structured data.

[0126] In practice, a text graph learning model can be used to implement the graph data conversion.

[0127] In this implementation, it is necessary to obtain or train a text graph learning model, and then use the text graph learning model to convert the knowledge graph data. The text graph learning module is obtained by decoupling the fine-tuning of the language model LM and the training of the graph neural network GNN into two stages, and training the text graph learning model. In the first stage, for a pre-trained LM, the node classification labels are used to perform efficient parameter fine-tuning; in the second stage, the fine-tuned LM is used to generate node embeddings. The node embeddings are obtained by removing the LM head layer and can be used by any GNN for the same task training, so that the GNN is trained on the generated node embeddings.

[0128] For example, the SimTeG model is used for text graph learning to perform data conversion. Because LLM requires sequence-like features, but graph data is not a sequence feature, the model first learns the graph as embedding (vector representation) and then converts it into sequence features.

[0129] SimTeG (Simple approach for Textual Graph learning) is a new approach for representation learning of textual graphs (TGs). A textual graph is a special type of graph whose nodes correspond to text (sentences or documents), and this type of graph is very common in real-world applications. The core contribution of SimTeG is to demonstrate a simple but effective method to improve representation learning of textual graphs by leveraging pre-trained language models, thereby significantly improving the performance of GNNs on tasks such as node classification and link prediction.

[0130] SimTeG has several key processing flows or features as follows:

[0131] 1. Parameter-Efficient Fine-Tuning (PEFT): SimTeG first performs parameter-efficient fine-tuning on a pre-trained language model (LM) using labels from downstream tasks (e.g., node classification).

[0132] 2. Node Embeddings Generation: Node embeddings are generated using the fine-tuned LM. These embeddings are obtained by removing the head layer and can be further used by any graph neural network (GNN) for the same task training.

[0133] 3. Two-Stage Training Approach: This approach decouples the fine-tuning of the language model and the training of the GNN into two stages. The first stage is to fine-tune the LM using the downstream task loss, and the second stage is to train the GNN on the generated node embeddings.

[0134] 4. Simple yet Effective Approach: Unlike previous work, SimTeG does not innovate in the framework, model, or task, but takes a very simple approach by fine-tuning on the pre-trained LM and then using these features to improve the performance of GNN.

[0135] 5. Significant Performance Improvement: Through extensive experiments, SimTeG significantly improves the performance of various GNNs on multiple graph benchmarks. In particular, when including additional supporting text provided by large language models (LLMs), SimTeG achieves accuracy comparable to the state-of-the-art (SOTA) performance on the OGBN-Arxiv dataset.

[0136] Therefore, this application adopts the SimTeG model for text graph learning and realizes graph data conversion, which is a more efficient and accurate data conversion method.

[0137] 2. Sub-image selection

[0138] Based on the Graph Attention mechanism, relevant information is selectively extracted from the knowledge graph. Graph information extraction can be performed from two levels: 1. What types of information should be extracted that are related to downstream specific tasks; 2. What information should be extracted that is related to specific samples within a specific task.

[0139] For example,

[0140] 1. For node-level tasks, such as entity classification, its associated information can be defined as:

[0141] the entity itself;

[0142] Attributes of the entity;

[0143] N adjacent nodes around the entity, N <= 2.

[0144] 2. For edge-level tasks, such as relation extraction RE, we define its associated information as:

[0145] The edge itself, the type of edge, etc.

[0146] The two nodes of the graph triple, i.e. the two vertices v1 and v2 of the “edge”;

[0147] ·Continue to extend from v1 and v2, each of them has 1 "edge".

[0148] 3. For graph-level tasks, such as text, similar to sentence sentiment analysis, we define its associated information as:

[0149] Graph itself, which is the type of the entire graph;

[0150] The task does not depend on the properties of a certain node or edge, but needs to consider the information of the entire graph.

[0151] After the graph information is extracted, the vertical domain LLM can be fine-tuned based on the extracted graph information. For example, the graph information is specifically applied to:

[0152] 1. As a Demonstration example, in the form of ICL, it helps promote LLM training. At this time, it is combined with the Prompt template to become a longer prompt and input to LLM.

[0153] 2. As the main sample and the "context" of the sample, it is used to enhance the sample itself and directly used as a training sample for the downstream MoE task.

[0154] 3. Multi-task learning

[0155] Multi-Task Learning (MTL) is a machine learning method that aims to solve multiple related tasks at the same time. It improves the overall learning effect through shared representation and information. Specifically, multi-task learning uses the correlation or common features between different tasks to improve the performance of each task. Multi-task learning can be defined as a machine learning method that learns multiple related tasks together based on shared representation. This method can not only improve the generalization ability of the model on each task, but also reduce the number of model parameters, thereby improving training efficiency. In addition, multi-task learning is also a method of inferential transfer learning, which is achieved by using the training signals of related tasks in the main task.

[0156] In the embodiment of the present application, LLM-independent multi-task fine-tuning is implemented.

[0157] LLM-independent multi-task fine-tuning can be understood as how LLM understands multiple tasks. For example, three task modes (entity classification, link prediction, and Graph description) are added to the LLM prompt, and fine-tuning of the model such as few shots is performed based on the ICL of the large model.

[0158] Specifically, you can freeze the LLM first and fine-tune the projection layer. For example, based on some open source frameworks (such as lora open source technology), it can be understood that the main parameters of LLM will not be back-propagated and only some parameters will be learned.

[0159] This fine-tuning process is performed according to different downstream tasks. The purpose of this process is to ensure that the embeddings of the graph nodes are well aligned with the token embedding space of the large language models (LLMs). This alignment enables LLMs to effectively handle graph-related tasks such as node classification, link prediction, and node description without any modification to the parameters of the LLMs.

[0160] The main technical points involved include:

[0161] 1. Multi-task learning: LLaGA uses three key graph tasks to adjust the projector: node classification, link prediction, and node description. These three tasks help the projector understand graph data from different perspectives. For example, by describing the node classification in the prompt of LLN, LLM can understand the task from three directions.

[0162] An example is given below.

[0163] Node classification: [Nongfu Spring] is [mineral water];

[0164] Node description: [Nongfu Spring] is [a mineral water category, its price is XX, it is sweet and delicious, and it is a similar brand to Baishuishan];

[0165] Link prediction: The same brand as [Nongfu Spring] in the mineral water category is [Baishuishan].

[0166] 2. Node description task: Different from traditional graph analysis, the node description task aims to align node embeddings with specific descriptive text. This innovative task can provide rich semantic explanations and provide a deeper understanding for graph-based predictions.

[0167] 3. Training format: During training, questions and answers are organized in a chat format. For example, for the node description task, the question may be "Please describe the central node:", and the answer is a description of the central node, such as "The central node represents a paper on [topic], which is about [description]".

[0168] Therefore, in one implementation, the vertical domain LLM is fine-tuned based on the extracted graph information, including: using the graph information as the main sample and context, as the training sample for the downstream multi-task learning of the vertical domain LLM, training the downstream multi-task learning model, and fine-tuning the vertical domain LLM according to the multi-task learning model. Specifically, it is determined that the downstream multi-tasks of the multi-task learning model include node classification tasks, link prediction tasks, and node description tasks; the task descriptions of the node classification tasks, link prediction tasks, and node description tasks are used as examples, and the vertical domain LLM is fine-tuned in the form of context learning, wherein the task descriptions of the node classification tasks, link prediction tasks, and node description tasks are added to the Prompt template, and the updated prompt is used as the prompt text of the vertical domain LLM.

[0169] In summary, the present application embodiment provides a vertical domain LLM training method, and its main technical advantages are reflected in the following aspects:

[0170] 1. In vertical scenarios, knowledge is directly modeled from a “graph” perspective and injected into the training method of the large language model (LLM);

[0171] 2. Defined the Graph->LLM Sequence conversion, including vertex-edge information representation->hierarchical tree traversal, clearly and completely bringing the Graph information into LLM;

[0172] 3. In the face of multiple downstream tasks, MoE multi-task learning is adopted, including node, edge, and graph multi-level learning tasks;

[0173] 4. The concept of "subgraph selection" is defined. Different graph information is extracted for different training tasks, given to LLM as Demonstration, and given to MoE as sample context to achieve fine-tuning of LLM.

[0174] Corresponding to the above method, the present application embodiment also provides a vertical domain LLM training device, see Figure 4 ,include:

[0175] A graph data conversion unit, used to access at least one knowledge graph, and perform graph data conversion on the data in the knowledge graph, so as to convert the graph data into a data sequence of a hierarchical tree structure according to the nodes, edges, and structures in the graph;

[0176] A graph data mapping unit, used to traverse the hierarchical tree structure and project the data sequence into the label embedding space of the base LLM;

[0177] A graph data injection unit is used to read the tag embedding space of the base LLM, obtain a data sequence obtained based on the graph data conversion, and inject the data sequence into the vertical domain data set to obtain the vertical domain data set injected with the graph data;

[0178] The model updating unit is used to update the base LLM according to the vertical domain data set of the injection map data to obtain the vertical domain LLM.

[0179] In one implementation, the graph data conversion unit specifically performs the following steps:

[0180] Step 1: Based on the data conversion template, determine the central node of the graph data and the environmental information formed by the edges between the nodes;

[0181] Step 2: Starting from the central node, a sampling computation tree of a fixed shape is constructed according to the environmental information;

[0182] Step 3: For each hop number of neighbor nodes of the central node, define the neighbor sample size n i , where n irepresents the sample size of the i-th jump;

[0183] Step 4: In the computation tree, take the central node as the root node and start from the 1-hop neighbor set To begin, randomly select n 1 nodes form a new neighbor set if The size is less than n 1 , then use placeholder nodes to supplement n 1 indivual;

[0184] Step 5: For each selected node, recursively use n 2 neighbors as child nodes of the selected node, and the insufficient node set is filled with placeholder nodes;

[0185] Step 6: Repeat steps 4 to 5 above until the central node and the neighboring nodes are converted into node sequences of fixed length.

[0186] In one implementation, the graph data conversion unit uses a text graph learning model to implement the graph data conversion.

[0187] In one implementation, it also includes:

[0188] The text graph learning model training unit is used to decouple the fine-tuning of the language model LM and the training of the graph neural network GNN into two stages, and train the text graph learning model, wherein, in the first stage, for a pre-trained LM, efficient parameter fine-tuning is performed using node classification labels; in the second stage, the fine-tuned LM is used to generate node embeddings, and the node embeddings are obtained by removing the LM head layer, which can be used by any GNN for training the same task, so as to train the GNN on the generated node embeddings.

[0189] In one implementation, the method further includes:

[0190] A graph information extraction unit, configured to extract graph information based on the knowledge graph, wherein graph information related to downstream tasks and / or graph information related to specific samples in specific tasks is extracted;

[0191] The model fine-tuning unit is used to fine-tune the vertical domain LLM based on the extracted graph information.

[0192] In one implementation,

[0193] The graph information extraction unit is specifically used for, for node-level tasks, the extracted graph information includes: the entity itself, entity attributes, and N adjacent nodes around the entity, where N is less than or equal to 2; for edge-level tasks, the extracted graph information includes: the edge itself, the type information of the edge, two nodes of the graph triplet: the two vertices V1 and V2 of the edge, and the edge information of the grid determined by continuing to verify from V1 and V2; for graph-level tasks, the extracted graph information includes the graph itself and graph type information.

[0194] In one implementation,

[0195] The model fine-tuning unit is specifically used to use the graph information as a sample to fine-tune the vertical domain LLM in the form of context learning, wherein the Prompt template is updated according to the graph information, and the updated prompt is used as the prompt text of the vertical domain LLM.

[0196] In one implementation,

[0197] The model fine-tuning unit is specifically used to use the graph information as the main sample and context, as the training sample for the downstream multi-task learning of the vertical domain LLM, to train the downstream multi-task learning model, so as to perform multi-task fine-tuning on the vertical domain LLM according to the multi-task learning model.

[0198] In one implementation,

[0199] The model fine-tuning unit is specifically used to determine that the downstream multiple tasks of the multi-task learning model include node classification tasks, link prediction tasks and node description tasks; taking the task descriptions of the node classification task, link prediction task and node description task as examples, fine-tuning the vertical domain LLM in the form of context learning, wherein the task descriptions of the node classification task, link prediction task and node description task are added to the Prompt template, and the updated prompt is used as the prompt text of the vertical domain LLM.

[0200] An embodiment of the present application further provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.

[0201] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:

[0202] Access at least one knowledge graph, perform graph data conversion on the data in the knowledge graph, so as to convert the graph data into a data sequence of a hierarchical tree structure according to the nodes, edges, and structures in the graph, and traverse the hierarchical tree structure to project the data sequence into the tag embedding space of the base LLM;

[0203] Read the tag embedding space of the base LLM, obtain the data sequence obtained based on the graph data conversion, inject the data sequence into the vertical domain data set, and obtain the vertical domain data set injected with the graph data;

[0204] According to the vertical domain data set of the injection map data, the base LLM is updated to obtain the vertical domain LLM.

[0205] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.

[0206] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0207] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0208] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0209] Access at least one knowledge graph, perform graph data conversion on the data in the knowledge graph, so as to convert the graph data into a data sequence of a hierarchical tree structure according to the nodes, edges, and structures in the graph, and traverse the hierarchical tree structure to project the data sequence into the tag embedding space of the base LLM;

[0210] Read the tag embedding space of the base LLM, obtain the data sequence obtained based on the graph data conversion, inject the data sequence into the vertical domain data set, and obtain the vertical domain data set injected with the graph data;

[0211] According to the vertical domain data set of the injection map data, the base LLM is updated to obtain the vertical domain LLM.

[0212] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0213] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0214] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0215] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0216] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0217] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0218] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0219] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A vertical LLM training method, characterized in that: include: Access at least one knowledge graph, perform graph data conversion on the data in the knowledge graph, so as to convert the graph data into a data sequence of a hierarchical tree structure according to the nodes, edges, and structures in the graph, and traverse the hierarchical tree structure to project the data sequence into the tag embedding space of the base LLM; Read the tag embedding space of the base LLM, obtain the data sequence obtained based on the graph data conversion, inject the data sequence into the vertical domain data set, and obtain the vertical domain data set injected with the graph data; According to the vertical domain data set of the injection map data, the base LLM is updated to obtain the vertical domain LLM.

2. The method according to claim 1, characterized in that The step of converting the data in the knowledge graph into graph data to convert the graph data into a data sequence of a hierarchical tree structure according to the nodes, edges, and structures in the graph includes: Step 1: Based on the data conversion template, determine the central node of the graph data and the environmental information formed by the edges between the nodes; Step 2: Starting from the central node, a sampling computation tree of a fixed shape is constructed according to the environmental information; Step 3: For each hop number of neighbor nodes of the central node, define the neighbor sample size n i , where n i represents the sample size of the i-th jump; Step 4: In the computation tree, take the central node as the root node and start from the 1-hop neighbor set Initially, n1 nodes are randomly selected to form a new neighbor set. If If the size is less than n1, then placeholder nodes are used to supplement it to n1; Step 5: For each selected node, recursively use n2 neighbors as child nodes of the selected node, and fill the insufficient node set with placeholder nodes; Step 6: Repeat steps 4 to 5 above until the central node and the neighboring nodes are converted into node sequences of fixed length.

3. The method according to claim 1, characterized in that A text graph learning model is used to implement the graph data conversion.

4. The method according to claim 3, characterized in that The method further comprises: The fine-tuning of the language model LM and the training of the graph neural network GNN are decoupled into two stages, and the text graph learning model is trained. In the first stage, for a pre-trained LM, efficient parameter fine-tuning is performed using node classification labels; in the second stage, the fine-tuned LM is used to generate node embeddings, and the node embeddings are obtained by removing the LM head layer, which can be used by any GNN for training the same task, so as to train the GNN on the generated node embeddings.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: Based on the knowledge graph, graph information extraction is performed, wherein graph information related to downstream tasks and / or graph information related to specific samples in specific tasks is extracted; Based on the extracted graph information, the vertical domain LLM is fine-tuned.

6. The method according to claim 5, characterized in that Based on the knowledge graph, graph information extraction is performed, including: For node-level tasks, the extracted graph information includes: the entity itself, entity attributes, and N adjacent nodes around the entity, where N is less than or equal to 2; For edge-level tasks, the extracted graph information includes: the edge itself, the edge type information, the two nodes of the graph triple: the two vertices V1 and V2 of the edge, and the edge information of the grid determined by continuing to verify from V1 and V2; For graph-level tasks, the extracted graph information includes the graph itself and graph type information.

7. The method according to claim 5, characterized in that The vertical domain LLM is fine-tuned based on the extracted graph information, including: The graph information is used as a sample to fine-tune the vertical domain LLM in the form of context learning, wherein the Prompt template is updated according to the graph information, and the updated prompt is used as the prompt text of the vertical domain LLM.

8. A vertical LLM training device, characterized in that: include: A graph data conversion unit, used to access at least one knowledge graph, and perform graph data conversion on the data in the knowledge graph, so as to convert the graph data into a data sequence of a hierarchical tree structure according to the nodes, edges, and structures in the graph; A graph data mapping unit, used to traverse the hierarchical tree structure and project the data sequence into the label embedding space of the base LLM; A graph data injection unit is used to read the tag embedding space of the base LLM, obtain a data sequence obtained based on the graph data conversion, and inject the data sequence into the vertical domain data set to obtain the vertical domain data set injected with the graph data; The model updating unit is used to update the base LLM according to the vertical domain data set of the injection map data to obtain the vertical domain LLM.

9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.