Large language model training method based on knowledge graph

By building professional knowledge graphs and entity alignment technologies, adjusting the word embedding layer of the large language model, the problem of inability to trace the source and insufficient adaptability of the knowledge generated by large language model is solved, and high-quality and accurate natural language text generation is achieved.

CN120069087APending Publication Date: 2025-05-30CHONGQING NORMAL UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510232675.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing large language models cannot trace the source when generating knowledge, lack interpretation, and in the enterprise intranet scenario, the external network data does not meet the needs of the enterprise, resulting in insufficient adaptability of the model, and learning semantic ambiguity samples, affecting the effectiveness of the training set.

Method used

By building a professional knowledge graph in a specific field, using the depth-first search algorithm and graph convolutional neural network, entity knowledge correlation analysis and entity alignment are performed, the expanded vocabulary list is generated, the word embedding layer of the pre-trained large language model is adjusted, and the internal knowledge graph is generated in combination with intranet data to perform knowledge fusion and fine-tuning.

Benefits of technology

It realizes that when large language models process and generate natural language text, they dynamically acquire and utilize related knowledge, improve the quality and accuracy of output text, reduce the cost and time of knowledge graph construction, and optimize the training process of large language models, so that the content generated is more accurate and reliable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069087A_ABST
    Figure CN120069087A_ABST
Patent Text Reader

Abstract

The invention discloses a big language model training method based on a knowledge graph, which comprises the following steps: collecting professional text data in a specific field, extracting entities and relationships from the professional text data, constructing a field knowledge graph, and constructing the knowledge graph based on data of different data sources on the basis of an external network; and introducing the domain knowledge graph into an intranet, storing the domain knowledge graph into a knowledge base in the intranet, extracting terminologies from the domain knowledge graph, analyzing word frequencies of the terminologies in the professional text data, and selecting the terminologies with the word frequencies higher than a set threshold value as newly added high-frequency terminologies. According to the method, the professional knowledge graph of the target field is constructed, the big language model is finely adjusted based on the professional knowledge graph, and the professional knowledge big language model is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model training methods, and more specifically, it relates to a large language model training method based on a knowledge graph. Background Art

[0002] With the continuous development of computer hardware systems, significant progress has been made in the field of artificial intelligence, especially in natural language processing (NLP) technology. Large-scale pre-trained language models (such as GPT-4, etc.) have powerful language understanding and generation capabilities through training on a vast amount of general corpora. They can generate coherent and logical texts and even show near-human-level dialogue and communication capabilities in certain situations, which undoubtedly opens up a new path for the development of artificial intelligence. Both large language models and knowledge graphs are highly influential and promising technologies in the contemporary field of artificial intelligence. Large language models have brought a revolutionary change to the field of natural language processing with their excellent text generation and understanding capabilities. However, the biggest problem is that the generated knowledge cannot trace its source, belonging to a "knowledge black box" and lacking interpretability. Traditional knowledge graph construction mainly relies on manual rules and a small amount of labeled data. However, these methods have limitations in dealing with large-scale, diverse, and dynamically changing data. In the prior art, the deployment and training of large language models usually rely on external network data and computing resources. In the enterprise intranet scenario, this will bring some problems. For example, external network data may not meet the enterprise's demand scenarios, resulting in insufficient adaptability of the model in business scenarios. For example, external network data may contain information in fields, topics, styles, etc. that are irrelevant to the enterprise, while lacking information specific to or important in the enterprise. When multiple data are fused, the entity expression forms in different data may be different, resulting in different triples constructed for the same entity in the knowledge graphs built from different source data, which will cause samples with semantic ambiguity to be learned during subsequent training of the large language model. The purpose of knowledge graph fusion is to match the entities and relationships in the knowledge graphs from different constructors in various fields to obtain a more complete and richer knowledge graph. However, due to the subjectivity of knowledge graph constructors and the non-uniqueness of knowledge, there are often entities that have different representations but the same meaning in different knowledge graphs, affecting the effectiveness of the training set of the large language model. Summary of the Invention

[0003] (1) Technical Problems to be Solved

[0004] In view of the problems existing in the prior art, the present invention provides a large language model training method based on a knowledge graph to solve the technical problems mentioned in the background art.

[0005] (2) Technical Solutions

[0006] To achieve the above object, the present invention provides the following technical solutions: A large language model training method based on a knowledge graph, comprising the following steps:

[0007] Step 1: Collect professional text data in a specific field and extract entities and relationships therefrom to construct a domain knowledge graph. The knowledge graph is constructed based on data from different data sources, and its basis is the external network.

[0008] Step 2: Bring the domain knowledge graph into the internal network and store it in the knowledge base in the internal network. Extract professional terms from the domain knowledge graph, analyze the word frequency of the professional terms in the professional text data, and select the professional terms with a word frequency higher than the set threshold as new high-frequency professional terms.

[0009] Step 3: Use the depth-first search algorithm to determine the search path vector of each entity node corresponding to each entity in each knowledge graph based on the search path corresponding to each entity in each knowledge graph; determine the entity knowledge relevance based on the attribute information between any two entities on each knowledge graph and the search path vectors of the entity nodes corresponding to the two entities.

[0010] Step 4: Generate new tokens for each new high-frequency professional term, integrate the new tokens into the vocabulary of the pre-trained large language model to obtain an expanded vocabulary; initialize the word embedding vectors of the tokens using the entity embedding vectors in the knowledge graph, and adjust the word embedding layer of the pre-trained large language model based on the word embedding vectors of the new tokens in the expanded vocabulary and perform retraining.

[0011] Step 5: Generate an internal knowledge graph based on the internal network data; use knowledge fusion technology to merge and fuse the internal knowledge graph and the domain knowledge graph from the external network to form a complete and consistent aggregated application knowledge graph; when deploying the application, use the aggregated application knowledge graph as the external knowledge base of the large language model, so that the large language model can dynamically obtain and utilize relevant knowledge from the application knowledge graph when processing and generating natural language text, so as to improve the quality and accuracy of the output text.

[0012] Step 6: Use the clustering algorithm to obtain the clustering results of the entity nodes corresponding to the entities in each knowledge graph based on the weighted entity association graph of each knowledge graph; determine the entity embedding distance between different entities in the two knowledge graphs based on the structural differences of the entity nodes corresponding to different entities in the clustering results of the two knowledge graphs; use the graph convolutional neural network to obtain the alignment results of the entities in the two knowledge graphs based on the entity embedding distance, attribute information and context information between the entities in any two knowledge graphs; complete the training of the large language model for knowledge-based question answering based on the alignment results of the entities in all knowledge graphs.

[0013] Step 7: Use the question-answer pairs constructed based on the domain knowledge graph as the instruction fine-tuning dataset after expert review and improvement, and use the instruction fine-tuning dataset to fine-tune the retrained large language model; perform performance evaluation and continuous optimization on the fine-tuned large language model.

[0014] The present invention is further configured such that the method for constructing a knowledge graph based on data from different data sources is as follows: obtaining text data from different sources using different data collection methods; taking the text data of each source as a type of original data, and processing each type of original data using entity named entity recognition technology and relation extraction technology to obtain a preset number of triples, and constructing a knowledge graph of each type of original data based on the preset number of triples.

[0015] The present invention is further configured such that extracting professional terms from the domain knowledge graph includes: extracting entities from the domain knowledge graph to obtain an entity set, further processing the entity set, removing common vocabulary, and taking the remaining entity set as professional terms.

[0016] The present invention is further configured such that the method further includes: in the intranet, selecting a pre-trained large language model, and then using the domain knowledge graph combined with intranet data to generate training data to fine-tune the large language model.

[0017] The present invention is further configured such that using the domain knowledge graph combined with intranet data to generate training data includes the following steps: retrieving entities, attributes, and relations related to intranet data from the domain knowledge graph using a query method; generating training samples according to the retrieved entities, attributes, and relations using a template method or a generation method; and converting the training samples into the input-output format required by the large language model using a formatting method.

[0018] The present invention is further configured such that initializing the word embedding vector of a token using the entity embedding vector in the knowledge graph, and adjusting the word embedding layer of the pre-trained large language model and performing retraining based on the word embedding vectors of the new tokens in the expanded vocabulary includes: for the pre-trained large language model pre-trained on general data, converting the entities in the knowledge graph into entity embedding vectors, taking the entity embedding vectors as the initial values of the word embedding vectors of the tokens in the pre-trained large language model, and using the word embedding vectors of the new tokens in the expanded vocabulary to adjust the word embedding layer of the pre-trained large language model, and then retraining the adjusted large language model based on general data to adapt to the expanded vocabulary.

[0019] The present invention is further configured such that the method for determining the entity knowledge relevance based on the attribute information between any two entities on each knowledge graph and the search path vectors of the entity nodes corresponding to the two entities is as follows: determining the attribute similarity between the two entities based on the difference in the attribute information between the two entities in each knowledge graph; taking the sum of the negative value of the attribute similarity between the two entities and the metric distance between the search path vectors of the entity nodes corresponding to the two entities as the first calculation factor; and taking the data mapping result of the first calculation factor as the entity knowledge relevance between the two entities.

[0020] The present invention is further configured such that using the question-answer pairs constructed based on the domain knowledge graph as the instruction fine-tuning dataset after expert review and improvement includes: generating questions for specific domain knowledge points using a question generation template based on the entities and relationships in the domain knowledge graph, and generating answers corresponding to the questions, and having experts review and improve the questions and answers, and constructing the reviewed and improved questions and answers into an instruction fine-tuning dataset.

[0021] The present invention is further configured such that when using the fine-tuned large language model for various applications, it includes the following sub-steps: obtaining the input text from the intranet data or user input according to the application requirements; converting the input text into a vector representation using an encoder; generating the output text using a decoder based on the vector representation; and converting the output text into a readable and displayable format according to the application requirements and returning it to the user or storing it in the intranet.

[0022] (III) Advantageous Effects

[0023] Compared with the prior art, the present invention provides a method for training a large language model based on a knowledge graph, having the following advantageous effects:

[0024] The present invention constructs a professional knowledge graph for the target domain. Based on the professional knowledge graph, the large language model is fine-tuned to generate a professional knowledge large language model. Based on the professional knowledge large language model and guiding prompt words, a structured answer to the target question is generated, and the structured answer is converted into a triple answer. Based on the triple answer, an answer knowledge graph for the target question is constructed, and the answer knowledge graph is rendered and displayed. The present invention can realize the construction of a knowledge graph based on a language model and the fine-tuning of a large language model based on a knowledge graph, enabling the knowledge graph and the large language model to work together. The large language model can parse and generate information that conforms to the structure of the knowledge graph, thereby constructing a more complete and rich knowledge graph. While ensuring the quality and accuracy of the knowledge graph, the construction cost and time are reduced. The knowledge graph can optimize and guide the training process of the large language model, making the generated content more accurate and reliable. Through the mutual empowerment of knowledge graph technology and large language model technology, efficient management of knowledge, intelligent and precise answering of professional knowledge questions are realized, and the quality and efficiency of solution decision-making are improved. In addition, the present invention obtains the search path and search path vector of each entity through the relevance between each entity and its surrounding entities in each knowledge graph. Secondly, based on the search path vector and the attribute information between similar entities, the entity knowledge relevance between entities is determined, and a weighted entity association graph is constructed based on the entity knowledge relevance, improving the effectiveness of constructing the adjacency matrix during subsequent entity alignment; secondly, according to the clustering results of different weighted entity association graphs, the entity embedding distance of entities in different knowledge graphs is determined. The beneficial effect is that by measuring the structural similarity and semantic information between the subtrees where each entity node is located, the probability that the entity is replaced by the embedding vector in the subsequent neural network model as the tail entity in the triple can be accurately reflected, improving the entity alignment effect; secondly, based on the entity alignment effect, the knowledge graph is fused and complemented, avoiding semantic ambiguity and noise interference in the original data, improving the effectiveness of the subsequent large language model training set, and enhancing the reply performance of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is the overall flowchart of the device in the unused state in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0027] It should be pointed out that unless otherwise specified, all technical and scientific terms used in this application have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0028] Embodiment 1:

[0029] Please refer to Figure 1 , a large language model training method based on a knowledge graph, comprising the following steps:

[0030] Step 1: Collect professional text data in a specific domain and extract entities and relationships therefrom to construct a domain knowledge graph. The knowledge graph is constructed based on data from different data sources, and its basis is the external network;

[0031] Step 2: Bring the domain knowledge graph into the internal network and store it in the knowledge base in the internal network. Extract professional terms from the domain knowledge graph, analyze the word frequencies of the professional terms in the professional text data, and select the professional terms with word frequencies higher than the set threshold as new high-frequency professional terms;

[0032] Step 3: Use the depth-first search algorithm to determine the search path vector of each entity node corresponding to each entity in each knowledge graph based on the search path corresponding to each entity in each knowledge graph; Determine the entity knowledge relevance based on the attribute information between any two entities on each knowledge graph and the search path vectors of the entity nodes corresponding to the two entities;

[0033] Step 4: Generate new word tokens for each new high-frequency professional term, integrate the new word tokens into the vocabulary of the pre-trained large language model to obtain an expanded vocabulary; Initialize the word embedding vectors of the word tokens using the entity embedding vectors in the knowledge graph, and adjust the word embedding layer of the pre-trained large language model based on the word embedding vectors of the new word tokens in the expanded vocabulary and perform re-training;

[0034] Step 5: Generate an internal knowledge graph based on the internal network data; Use knowledge fusion technology to merge and fuse the internal knowledge graph and the domain knowledge graph from the external network to form a complete and consistent aggregated application knowledge graph; When deploying the application, use the aggregated application knowledge graph as the external knowledge base of the large language model, so that the large language model can dynamically obtain and utilize relevant knowledge from the application knowledge graph when processing and generating natural language text to improve the quality and accuracy of the output text;

[0035] Step 6: Use the clustering algorithm to obtain the clustering results of the entity nodes corresponding to the entities in each knowledge graph based on the weighted entity association graph of each knowledge graph; Determine the entity embedding distance between different entities in the two knowledge graphs based on the structural differences between the clustering results where the entity nodes corresponding to different entities in the two knowledge graphs are located; Use a graph convolutional neural network to obtain the alignment results of the entities in the two knowledge graphs based on the entity embedding distance, the attribute information of the entities, and the context information between the entities in any two knowledge graphs; Complete the training of the large language model for knowledge-based question answering based on the alignment results of the entities in all knowledge graphs;

[0036] Step 7: Use the question-answer pairs constructed based on the domain knowledge graph as an instruction fine-tuning dataset after expert review and improvement, and use the instruction fine-tuning dataset to fine-tune the retrained large language model; perform performance evaluation and continuous optimization on the fine-tuned large language model.

[0037] In a further embodiment of the present invention, the method for constructing a knowledge graph based on data from different data sources is as follows: obtaining text data from different sources using different data collection methods; taking the text data of each source as a type of original data, and processing each type of original data using entity named entity recognition technology and relationship extraction technology to obtain a preset number of triples, and constructing a knowledge graph of each type of original data based on the preset number of triples.

[0038] In a further embodiment of the present invention, extracting professional terms from the domain knowledge graph includes: extracting entities from the domain knowledge graph to obtain an entity set, further processing the entity set, removing general vocabulary, and using the remaining entity set as professional terms.

[0039] In a further embodiment of the present invention, the method further includes: in the intranet, selecting a pre-trained large language model, and then using the domain knowledge graph combined with intranet data to generate training data to fine-tune the large language model.

[0040] Embodiment 2:

[0041] A method for training a large language model based on a knowledge graph, comprising the following steps:

[0042] Step 1: Collect professional text data in a specific domain, extract entities and relationships therefrom, construct a domain knowledge graph, and construct a knowledge graph based on data from different data sources, the basis of which is based on the extranet;

[0043] Step 2: Bring the domain knowledge graph into the intranet and store it in the knowledge base in the intranet, extract professional terms from the domain knowledge graph, analyze the word frequency of the professional terms in the professional text data, and select the professional terms with a word frequency higher than the set threshold as new high-frequency professional terms;

[0044] Step 3: Use the depth-first search algorithm to determine the search path vector of each entity node corresponding to each entity in each knowledge graph based on the search path corresponding to each entity in each knowledge graph; determine the entity knowledge relevance based on the attribute information between any two entities on each knowledge graph and the search path vectors of the entity nodes corresponding to the two entities;

[0045] Step 4: Generate new lemmas for each newly added high-frequency professional term, integrate the new lemmas into the vocabulary of the pre-trained large language model to obtain an extended vocabulary; initialize the word embedding vectors of the lemmas using the entity embedding vectors in the knowledge graph, and adjust the word embedding layer of the pre-trained large language model based on the word embedding vectors of the newly added lemmas in the extended vocabulary and perform retraining;

[0046] Step 5: Generate an internal knowledge graph based on the intranet data; use knowledge fusion technology to merge and integrate the internal knowledge graph and the domain knowledge graph from the extranet to form a complete and consistent aggregated application knowledge graph; when deploying the application, use the aggregated application knowledge graph as the external knowledge base of the large language model, so that the large language model can dynamically obtain and utilize relevant knowledge from the application knowledge graph when processing and generating natural language text to improve the quality and accuracy of the output text; adopt a clustering algorithm to obtain the clustering results of the entity corresponding entity nodes in each knowledge graph based on the weighted entity association graph of each knowledge graph; determine the entity embedding distance between different entities in the two knowledge graphs based on the structural differences of the entity corresponding entity nodes in the clustering results of different entities in the two knowledge graphs; use a graph convolutional neural network to obtain the alignment results of the entities in the two knowledge graphs based on the entity embedding distance, entity attribute information, and context information between any two entities in the knowledge graphs; complete the training of the large language model for knowledge-based question answering based on the alignment results of the entities in all knowledge graphs; use the question-answer pairs constructed based on the domain knowledge graph as the instruction fine-tuning dataset after expert review and improvement, and use the instruction fine-tuning dataset to fine-tune the retrained large language model; perform performance evaluation and continuous optimization on the fine-tuned large language model.

[0047] Generating training data by combining the domain knowledge graph with intranet data includes the following steps: Retrieving entities, attributes, and relationships related to intranet data from the domain knowledge graph using a query method; Generating training samples according to the retrieved entities, attributes, and relationships using a template method or a generation method; Converting the training samples into the input-output format required by the large language model using a formatting method; Initializing the word embedding vectors of tokens with the entity embedding vectors in the knowledge graph, and adjusting the word embedding layer of the pre-trained large language model based on the word embedding vectors of the newly added tokens in the expanded vocabulary and performing retraining, including: For the pre-trained large language model obtained by pre-training on general data, converting the entities in the knowledge graph into entity embedding vectors, using the entity embedding vectors as the initial values of the word embedding vectors of the tokens in the pre-trained large language model, and adjusting the word embedding layer of the pre-trained large language model using the word embedding vectors of the newly added tokens in the expanded vocabulary, and then retraining the adjusted large language model based on general data to adapt to the expanded vocabulary; The method for determining the entity knowledge relevance based on the attribute information between any two entities on each knowledge graph and the search path vectors of the entity nodes corresponding to the two entities is: Determining the attribute similarity between two entities based on the difference in the attribute information between the two entities in each knowledge graph; Taking the sum of the negative value of the attribute similarity between the two entities and the metric distance between the search path vectors of the entity nodes corresponding to the two entities as the first calculation factor; Taking the data mapping result of the first calculation factor as the entity knowledge relevance between the two entities; Using the question-answer pairs constructed based on the domain knowledge graph as the instruction fine-tuning dataset after expert review and improvement, including: Based on the entities and relationships in the domain knowledge graph, using a question generation template to generate questions for specific domain knowledge points and generating answers corresponding to the questions, and having experts review and improve the questions and answers, and constructing the reviewed and improved questions and answers into an instruction fine-tuning dataset; When using the fine-tuned large language model for various applications, it includes the following sub-steps: Obtaining the input text from the intranet data or user input according to the application requirements; Converting the input text into a vector representation using an encoder; Generating the output text using a decoder according to the vector representation; According to the application requirements, converting the output text into a readable and displayable format and returning it to the user or storing it in the intranet.

[0048] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A large language model training method based on knowledge graph, comprising the following steps: Step 1: Collect professional text data in a specific field and extract entities and relationships from it to build a domain knowledge graph. The knowledge graph is built based on data from different data sources, and its foundation is based on the external network; Step 2: Bring the domain knowledge graph into the intranet and store it in the knowledge base in the intranet, extract professional terms from the domain knowledge graph, analyze the frequency of professional terms in professional text data, and select professional terms with a frequency higher than a set threshold as new high-frequency professional terms; Step 3: Determine the search path vector of the entity node corresponding to each entity based on the search path corresponding to each entity in each knowledge graph using a depth-first search algorithm; determine the entity knowledge association based on the attribute information between any two entities on each knowledge graph and the search path vector of the entity nodes corresponding to the two entities; Step 4: Generate new word-grams for each new high-frequency professional term, integrate the new word-grams into the vocabulary of the pre-trained large language model to obtain an expanded vocabulary; use the entity embedding vector in the knowledge graph to initialize the word embedding vector of the word-gram, and adjust the word embedding layer of the pre-trained large language model based on the word embedding vector of the new word-gram in the expanded vocabulary and retrain; Step 5: Generate an internal knowledge graph based on intranet data; The internal knowledge graph and the domain knowledge graph from the external network are merged and fused using knowledge fusion technology to form a complete and consistent aggregated application knowledge graph; when deploying the application, the aggregated application knowledge graph is used as the external knowledge base of the large language model, so that the large language model can dynamically obtain and use related knowledge from the application knowledge graph when processing and generating natural language text, so as to improve the quality and accuracy of the output text; Step 6: Use a clustering algorithm to obtain the clustering results of entity nodes corresponding to entities in each knowledge graph based on the weighted entity association graph of each knowledge graph; determine the entity embedding distance between different entities in the two knowledge graphs based on the structural differences in the clustering results of entity nodes corresponding to different entities in the two knowledge graphs; use a graph convolutional neural network to obtain the alignment results of entities in the two knowledge graphs based on the entity embedding distance between entities in any two knowledge graphs, the attribute information of the entities, and the context information; complete the training of the large language model for knowledge question answering based on the alignment results of entities in all knowledge graphs; Step 7: Use the question-answer pairs constructed based on the domain knowledge graph as the instruction fine-tuning dataset after expert review and improvement. Use the instruction fine-tuning dataset to fine-tune the retrained large language model; perform performance evaluation and continuous optimization on the fine-tuned large language model.

2. According to the method for training a large language model based on a knowledge graph according to claim 1, it is characterized by: The method for constructing a knowledge graph based on data from different data sources is as follows: using different data collection methods to obtain text data from different sources; treating each type of text data from a source as a type of raw data, using entity naming recognition technology and relationship extraction technology to process each type of raw data to obtain a preset number of triples, and constructing a knowledge graph for each type of raw data based on the preset number of triples.

3. According to the method for training a large language model based on a knowledge graph according to claim 1, it is characterized by: The extracting of professional terms from the domain knowledge graph includes: extracting entities from the domain knowledge graph to obtain entity sets, further processing the entity sets, eliminating common words, and using the remaining entity sets as professional terms.

4. According to the method for training a large language model based on a knowledge graph according to claim 1, it is characterized by: The method also includes: selecting a pre-trained large language model in the intranet, and then using the domain knowledge graph in combination with intranet data to generate training data to fine-tune the large language model.

5. The method for training a large language model based on a knowledge graph according to claim 1, wherein: The training data is generated by combining the domain knowledge graph with intranet data, including the following steps: using a query method to retrieve entities, attributes and relationships related to the intranet data from the domain knowledge graph; using a template method or a generation method to generate training samples based on the retrieved entities, attributes and relationships; using a formatting method to convert the training samples into the input and output formats required by the large language model.

6. The method for training a large language model based on a knowledge graph according to claim 5, characterized in that: The method uses the entity embedding vector in the knowledge graph to initialize the word embedding vector of the word unit, and adjusts the word embedding layer of the pre-trained large language model based on the word embedding vector of the newly added word unit in the expanded vocabulary and retrains it, including: for the pre-trained large language model obtained by pre-training on general data, converting the entities in the knowledge graph into entity embedding vectors, using the entity embedding vectors as the initial values ​​of the word embedding vectors of the word units in the pre-trained large language model, and using the word embedding vectors of the newly added word units in the expanded vocabulary to adjust the word embedding layer of the pre-trained large language model, and retraining the adjusted large language model based on the general data to adapt to the expanded vocabulary.

7. The method for training a large language model based on a knowledge graph according to claim 6, characterized in that: The method for determining the entity knowledge association based on the attribute information between any two entities on each knowledge graph and the search path vectors of the entity nodes corresponding to the two entities is as follows: determining the attribute similarity between the two entities based on the difference in attribute information between the two entities in each knowledge graph; taking the inverse of the attribute similarity between the two entities and the sum of the metric distance between the search path vectors of the entity nodes corresponding to the two entities as the first calculation factor; The data mapping result of the first calculation factor is used as the entity knowledge association between the two entities.

8. The method for training a large language model based on a knowledge graph according to claim 1, characterized in that: The question-answer pairs constructed based on the domain knowledge graph are reviewed and improved by experts as an instruction fine-tuning dataset, including: based on the entities and relationships in the domain knowledge graph, using question generation templates to generate questions for specific domain knowledge points, and generating answers corresponding to the questions, and having experts review and improve the questions and answers, and constructing the reviewed and improved questions and answers as an instruction fine-tuning dataset.

9. The method for training a large language model based on a knowledge graph according to claim 1, characterized in that: When using a fine-tuned large language model for various applications, the following sub-steps are included: according to application requirements, obtain input text from intranet data or user input; use an encoder to convert the input text into a vector representation; use a decoder to generate output text based on the vector representation; According to application requirements, the output text is converted into a readable and displayable format and returned to the user or stored in the intranet.

Citation Information

Cited By

  • Life significance sense detection method and device

    CN120596640A

  • Knowledge dynamic construction method and system based on question and answer large model

    CN121809634A

  • Large language model multi-stage guided training system and method for steel production field

    CN122615414A