Entity prediction method and device combining embedded model and instruction fine-tuning large model
By combining embedded models and instruction fine-tuning large models, the problem of poor performance of large models in entity prediction tasks is solved, and more efficient and reliable entity prediction results are achieved.
Patent Information
- Application Number
- CN202510197803.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-24
AI Technical Summary
The existing entity prediction methods have poor performance when using large models, mainly because the large models do not adapt to structured tasks, the data required for fine-tuning the large models is difficult and costly, and there are hallucinations in the big models, and the generated predictive entities may not exist in the knowledge graph.
Combining the entity prediction method of embedded models and instructions fine-tune large models, fine-tune instructions are generated by embedding models, fine-tune large models to adapt to entity prediction tasks, QLoRA method is used to quantify large model parameters to reduce video memory requirements, and hyperparameters are adjusted to optimize model performance.
It effectively improves the performance of large models in entity prediction tasks, reduces the memory requirements and training time of fine-tuning large models, and improves the reliability and practicality of prediction results.
Smart Images

Figure CN120197677A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge graph completion, and particularly to an entity prediction method and device that combines an embedding model and a large model with instruction fine-tuning. Background Art
[0002] A knowledge graph stores facts in the real world in a multi-relational structure, where nodes represent entities and edges are labeled with relationships, describing facts in the form of triples, such as (head entity, relationship, tail entity). Knowledge graphs often face the problem of incompleteness. Knowledge graph completion aims to solve the problem of missing links in the knowledge graph, making the knowledge graph more complete and providing more support for the application of related downstream tasks. This task aims to infer missing facts based on existing facts to complete the knowledge graph. For the entity prediction task in the knowledge graph completion task, an incomplete fact is given, where the head entity or the tail entity is missing, and the missing entity needs to be predicted to complete the knowledge graph completion. For example, in the tail entity prediction task, it is expressed as given the head entity and the relationship, the tail entity needs to be predicted.
[0003] The methods used in existing entity prediction tasks can be mainly divided into two types. One is the method based on knowledge embedding, and the other is the method based on text generation.
[0004] The method based on knowledge embedding usually trains an embedding model (such as TransE) to learn the vector representations of entities and relationships in the knowledge graph, thereby learning the structure of the knowledge graph, and then calculating the probability of predicting entities. With the development of pre-trained language models (such as BERT), new methods use pre-trained language models to learn the encoded representations of entities and relationships in the knowledge graph.
[0005] The method based on text generation is mainly inspired by generative language models (such as T5, BART), and transforms the entity prediction task into a text generation task. These methods first convert the entity prediction task into natural language, and then directly give the answer using a generative language model. With the rapid development of large models (such as LLaMA, GPT), these methods turn to use large models for text generation.
[0006] However, since large models are mainly used for general text generation tasks and perform poorly in structured tasks such as entity prediction, it is necessary to fine-tune large models to adapt to a certain type of specific structured task through instruction fine-tuning technology. The fine-tuning effect of large models depends heavily on the quality of the fine-tuning instruction set, and generating a large number of high-quality supervised fine-tuning instruction sets requires a large amount of human resources and time costs, and fine-tuning cannot solve the hallucination phenomenon of large models for specific tasks. The above problems together lead to the poor performance of large models in entity prediction tasks.
[0007] As described above, the existing entity prediction solutions mainly have the following disadvantages:
[0008] (1) Large models are not suitable for entity prediction tasks, and directly using large models results in poor performance. Large models perform excellently in natural language generation tasks. However, since large models are mainly targeted at general text generation objectives during the pre-training process, rather than specific tasks such as knowledge graph completion or entity relationship prediction, their performance for structured tasks like entity prediction is often inferior to that of dedicated models. Directly using large models often fails to generate high-quality entity prediction results.
[0009] (2) It is difficult to generate the data required for the large model fine-tuning method, and the generation cost is high. Existing fine-tuning methods have extremely high requirements for high-quality supervised data, and there are multiple challenges in generating such data. Constructing precisely labeled supervised fine-tuning data requires a large amount of time and human resources investment.
[0010] (3) There is an hallucination phenomenon in large models, and the predicted entities generated may not exist in the knowledge graph. There is an hallucination phenomenon in large models during the generation task, that is, the content they output may not conform to the facts. For entity prediction tasks, this may lead to the generated entities or relationships not existing in the existing knowledge graph, thus reducing the reliability and practicality of the prediction. Summary of the Invention
[0011] This application aims to solve at least one of the technical problems in the related art to some extent.
[0012] To this end, the first object of this application is to propose an entity prediction method that combines an embedding model and instruction fine-tuning of a large model. By combining the advantages of the embedding model and the large language model, generating fine-tuning instructions through knowledge embedding, and fine-tuning the large model through the instruction fine-tuning method, the performance of the large model in entity prediction tasks is effectively improved, meeting the dynamic and efficient knowledge acquisition requirements.
[0013] The second object of this application is to propose an entity prediction device that combines an embedding model and instruction fine-tuning of a large model.
[0014] The third object of this application is to propose a computer device.
[0015] The fourth object of this application is to propose a non-transitory computer-readable storage medium.
[0016] To achieve the above object, the first aspect embodiment of this application proposes an entity prediction method that combines an embedding model and instruction fine-tuning of a large model, including:
[0017] Obtain a dataset for entity prediction, and divide the dataset into embedding model training data, fine-tuning instruction generation data, and hyperparameter adjustment data;
[0018] Train an embedding model with the embedding model training data;
[0019] Generate data and an embedding model through fine-tuning instructions, fill in the fine-tuning instruction template to obtain a fine-tuning instruction dataset, where the fine-tuning instruction template includes an input question, relevant information, a candidate entity list, and a correct answer;
[0020] Use the fine-tuning instruction dataset to fine-tune the large model. When fine-tuning, adjust the hyperparameters of the large model using hyperparameter adjustment data, and select the optimal hyperparameter combination and the optimal fine-tuned large model;
[0021] Given an entity prediction task, generate a candidate entity list through the embedding model, and use the fine-tuned large model to select a predicted entity from the candidate entity list.
[0022] Optionally, in an embodiment of the present application, for the entity prediction task, the input question is a triple missing a head entity or a tail entity, the relevant information is a set of first-order adjacent entity relationship triples of entities related to the input question in the knowledge graph, the correct answer is the missing entity in the triple, and the generation process of the candidate entity list includes:
[0023] Input the input question into the embedding model, output a candidate entity list sorted by probability from high to low, and retain the top K entities in a truncated manner.
[0024] Optionally, in an embodiment of the present application, fine-tuning the large model includes:
[0025] Use the PEFT fine-tuning framework to perform instruction fine-tuning on the large model, and use the QLoRA quantization method to quantize the large model parameters.
[0026] Optionally, in an embodiment of the present application, completing the entity prediction task includes:
[0027] Input the triple missing a head entity or a tail entity into the embedding model, output a candidate entity list sorted by probability from high to low, and retain the top K entities in a truncated manner;
[0028] Input the triple missing a head entity or a tail entity, the candidate entity list, and the relevant information into the fine-tuned large model, sort the candidate entity list through the fine-tuned large model, and select the entity with the highest probability in the sorted list as the predicted entity.
[0029] To achieve the above object, an embodiment of the second aspect of the present invention proposes an entity prediction device combining an embedding model and an instruction fine-tuned large model, including:
[0030] A dataset processing module, configured to obtain a dataset for entity prediction, and split the dataset into embedding model training data, fine-tuning instruction generation data, and hyperparameter adjustment data;
[0031] An embedding model training module for training an embedding model with embedding model training data;
[0032] A fine-tuning instruction data generation module for populating a fine-tuning instruction template with fine-tuning instruction generation data and the embedding model to obtain a fine-tuning instruction data set, where the fine-tuning instruction template includes an input question, relevant information, a candidate entity list, and a correct answer;
[0033] A large model fine-tuning module for fine-tuning a large model using the fine-tuning instruction data set. During fine-tuning, hyperparameter adjustment data is used to adjust the hyperparameters of the large model to select the optimal hyperparameter combination and the optimal fine-tuned large model;
[0034] An entity prediction module for a given entity prediction task, generating a candidate entity list through the embedding model, and using the fine-tuned large model to select a predicted entity from the candidate entity list.
[0035] Optionally, in an embodiment of the present application, for the entity prediction task, the input question is a triple missing a head entity or a tail entity, the relevant information is a set of first-order adjacent entity relationship triples of entities related to the input question in the knowledge graph, the correct answer is the missing entity in the triple, and the generation process of the candidate entity list includes:
[0036] Input the input question into the embedding model, output a candidate entity list sorted by probability from high to low, and retain the top K entities in a truncated manner.
[0037] Optionally, in an embodiment of the present application, the large model fine-tuning module is specifically used for:
[0038] Use the PEFT fine-tuning framework to perform instruction fine-tuning on the large model and use the QLoRA quantization method to quantize the large model parameters.
[0039] Optionally, in an embodiment of the present application, the entity prediction module is specifically used for:
[0040] Input the triple missing a head entity or a tail entity into the embedding model, output a candidate entity list sorted by probability from high to low, and retain the top K entities in a truncated manner;
[0041] Input the triple missing a head entity or a tail entity, the candidate entity list, and the relevant information into the fine-tuned large model, sort the candidate entity list through the fine-tuned large model, and select the entity with the highest probability in the sorted list as the predicted entity.
[0042] To achieve the above object, an embodiment of the third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned entity prediction method combining the embedding model and the instruction fine-tuning large model is implemented.
[0043] To achieve the above object, an embodiment of the fourth aspect of the present invention provides a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by a processor, the above-mentioned entity prediction method combining the embedding model and the instruction fine-tuning large model can be executed.
[0044] In the entity prediction method of combining the embedding model and the instruction fine-tuning large model according to the embodiments of the present application, aiming at the problem that the large model is not suitable for the entity prediction task and the direct use of the large model has poor effects, a technology for adapting the instruction fine-tuning large model to entity prediction is proposed. The large model is fine-tuned by the instruction fine-tuning method, and the QLoRA method is combined to quantize the large model parameters, reducing the video memory resources required for fine-tuning the large model, improving the fine-tuning speed, and adjusting the hyperparameters through the validation set to obtain the model with the best effect, so that the fine-tuned large model is adapted to the knowledge graph completion task; aiming at the problem that the data generation required by the method of fine-tuning the large model is difficult and the generation cost is high, a technology for generating fine-tuning instructions by combining the embedding model is proposed. The embedding model is trained to learn the representations of entities and relationships in the knowledge graph, the probability ranking of entities is predicted using the embedding model, and the entities with high occurrence probabilities are intercepted as candidate entities and filled into the instruction template to generate discriminative instructions; aiming at the problem that the large model has a hallucination phenomenon and the predicted entities generated may not exist in the knowledge graph, a large model entity prediction technology combining the embedding model is proposed. For an entity prediction task, a candidate entity list is generated by combining the trained embedding model, and the fine-tuned large model is used to find the most likely answer from the candidate entity list to complete the entity prediction task.
[0045] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. Description of the Drawings
[0046] The above-mentioned and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0047] Figure 1 is a schematic flowchart of an entity prediction method combining an embedding model and an instruction fine-tuning large model provided by Embodiment 1 of the present application;
[0048] Figure 2 is an example diagram of the content format of the fine-tuning instruction provided by Embodiment 2 of the present application;
[0049] Figure 3Schematic structural diagram of an entity prediction device combining an embedding model and an instruction fine-tuned large model provided by an embodiment of the present application. Detailed implementation manners
[0050] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, but should not be construed as limiting the present application.
[0051] The entity prediction method and device combining an embedding model and an instruction fine-tuned large model according to an embodiment of the present application will be described below with reference to the accompanying drawings.
[0052] Figure 1 Schematic flowchart of an entity prediction method combining an embedding model and an instruction fine-tuned large model provided by Embodiment 1 of the present application.
[0053] As Figure 1 shown, the entity prediction method combining an embedding model and an instruction fine-tuned large model includes the following steps:
[0054] Step 101, obtain a dataset for entity prediction, and divide the dataset into embedding model training data, fine-tuning instruction generation data, and hyperparameter adjustment data;
[0055] In this embodiment, after preprocessing the dataset for entity prediction, the dataset is divided into four parts: embedding model training data, fine-tuning instruction generation data, hyperparameter adjustment data, and model effect verification data.
[0056] The embedding model training data accounts for 50% of the dataset, and this part of the data is used to train the embedding model to learn the embedding representations of entities and relationships in the dataset.
[0057] The fine-tuning instruction generation data accounts for 10% of the dataset. Based on the trained embedding model, this part of the data generates an entity prediction list to generate fine-tuning instructions for fine-tuning the large model.
[0058] The hyperparameter adjustment data accounts for 20% of the dataset, and this part of the data is used to find the optimal hyperparameter combination for fine-tuning the large model, and the hyperparameters can be adjusted to select the model with the best effect.
[0059] The model effect verification data accounts for 20% of the dataset, and this part of the data is used to test the effect of the model designed by the method.
[0060] Step 102, train the embedding model with the embedding model training data;
[0061] In this embodiment, the training embedding model learns the representations of entities and relationships in the knowledge graph, and the model maps entities and relationships into vectors of a unified dimension. Select common knowledge graph embedding models, such as TransE, ConvE, ComplEx, etc., and use the embedding model training data processed by the dataset preprocessing subsystem to train the embedding model, and use the hyperparameter adjustment data to adjust the hyperparameters to select the model with the best effect in multiple rounds of training.
[0062] Based on the supervised data, the embedding model is trained. The model will learn the embedding representations of entities and relationships in the knowledge graph. In the entity prediction task, for example, in the tail entity prediction task, given the head entity and the relationship, the embedding model can generate a list of candidate entities sorted in descending order of the probability of the tail entity predicted by the model.
[0063] Step 103, generate data and an embedding model through fine-tuning instructions, and fill in the fine-tuning instruction template to obtain a fine-tuning instruction dataset, where the fine-tuning instruction template includes an input question, relevant information, a list of candidate entities, and the correct answer;
[0064] In this embodiment, the format of a fine-tuning instruction template mainly includes four parts: input question, list of candidate entities, relevant information, and correct answer. An example of the content of the fine-tuning instruction format is shown Figure 2 as follows.
[0065] The input question is the question given in the dataset. For the entity prediction task, this question is a triple missing the head entity or the tail entity, and the entity prediction task requires predicting the missing entity.
[0066] The relevant information is the set of first-order adjacent entity relationship triples of the entities related to the input question in the knowledge graph, but does not include the correct answer to the question itself. Using this knowledge can improve the fine-tuning effect of the large model.
[0067] The list of candidate entities is the list of predicted entities of the trained embedding model for the input question. This list is sorted in descending order of the probability of the predicted entities appearing in the model, and in order to prevent too many predicted entities from exceeding the context limit of the large model, a truncation method is used to only select the top K entities with high rankings.
[0068] The correct answer is the answer given in the dataset for this question. For the entity prediction task, this answer is the entity missing from the incomplete triple in the input question.
[0069] By combining the data and the embedding model generated by the fine-tuning instructions, filling in the fine-tuning instruction template, a fine-tuning instruction dataset is generated for fine-tuning the large model.
[0070] Step 104: Fine-tune the large model using the fine-tuning instruction dataset. When fine-tuning, use hyperparameter tuning data to adjust the hyperparameters of the large model, and select the optimal hyperparameter combination and the optimal fine-tuned large model.
[0071] In this embodiment, the large model is fine-tuned to adapt to the entity prediction task. Using the generated fine-tuning instruction set, the PEFT fine-tuning framework is used to perform instruction fine-tuning on the large model, which can effectively improve the performance of the large model in the entity prediction task.
[0072] When fine-tuning, in order to save the GPU memory resources used for fine-tuning, the QLoRA quantization method is adopted to quantize the large model parameters from 8-bit to 4-bit, which can greatly reduce the GPU memory requirement threshold and improve the training speed, and will not have too much impact on the model performance. After fine-tuning, the model can effectively re-rank the entity prediction candidate list generated by the embedding model, thereby further optimizing the entity prediction accuracy.
[0073] During the fine-tuning process, use hyperparameter tuning data to adjust the hyperparameters, and select the optimal hyperparameter combination and the optimal fine-tuned model.
[0074] Step 105: Given an entity prediction task, generate a candidate entity list through the embedding model, and use the fine-tuned large model to select the predicted entity from the candidate entity list.
[0075] In this embodiment, for a given entity prediction task, each question in the task is a triple missing the head entity or the tail entity. First, the embedding model trained by the embedding model training subsystem predicts the missing entity list, which represents the ranking of the embedding model's prediction probabilities for the missing entities from high to low, and truncate this list to take the top K and hand them over to the fine-tuned large model obtained by the large model fine-tuning subsystem. Based on the given Prompt, use this large model to complete the re-ranking of the candidate entity list, and select the entity with the highest probability in the sorted list as the method prediction result.
[0076] In this embodiment, model performance verification data is used to verify the performance of this application.
[0077] The entity prediction method combining an embedding model and an instruction-tuned large model according to the embodiments of the present application proposes a technique for generating fine-tuning instructions by combining an embedding model, trains the embedding model to learn the representations of entities and relationships in a knowledge graph, uses the embedding model to predict the probability ranking of entities, intercepts the top K entities with occurrence probabilities as candidate entities, and fills them into an instruction template to generate discriminative instructions; proposes a technique for adapting an instruction-tuned large model to entity prediction, fine-tunes the large model through an instruction fine-tuning method, combines the QLoRA method to quantize the parameters of the large model, reduces the video memory resources required for fine-tuning the large model, improves the fine-tuning speed, and adjusts the hyperparameters through a validation set to obtain the best model, so that the fine-tuned large model adapts to the knowledge graph completion task; proposes a large model entity prediction technique combining an embedding model. For an entity prediction task, combines the trained embedding model to generate a list of candidate entities, and uses the fine-tuned large model to find the most likely answer from the list of candidate entities to complete the entity prediction task.
[0078] To implement the above embodiments, the present application also proposes an entity prediction device combining an embedding model and an instruction-tuned large model.
[0079] Figure 3 FIG. is a schematic structural diagram of an entity prediction device combining an embedding model and an instruction-tuned large model provided by an embodiment of the present application.
[0080] As Figure 3 shown, the entity prediction device combining an embedding model and an instruction-tuned large model includes:
[0081] A dataset processing module, configured to obtain a dataset for entity prediction, and divide the dataset into embedding model training data, fine-tuning instruction generation data, and hyperparameter adjustment data;
[0082] An embedding model training module, configured to train an embedding model through the embedding model training data;
[0083] A fine-tuning instruction data generation module, configured to fill a fine-tuning instruction template through the fine-tuning instruction generation data and the embedding model to obtain a fine-tuning instruction dataset, where the fine-tuning instruction template includes an input question, relevant information, a list of candidate entities, and a correct answer;
[0084] A large model fine-tuning module, configured to fine-tune the large model using the fine-tuning instruction dataset, and when fine-tuning, adjust the hyperparameters of the large model using the hyperparameter adjustment data to select the optimal hyperparameter combination and the optimal fine-tuned large model;
[0085] An entity prediction module, configured to, given an entity prediction task, generate a list of candidate entities through the embedding model, and select a predicted entity from the list of candidate entities using the fine-tuned large model.
[0086] Optionally, in an embodiment of the present application, for the entity prediction task, the input question is a triple missing the head entity or the tail entity, the relevant information is the set of first-order adjacent entity relationship triples of the entities related to the input question in the knowledge graph, the correct answer is the entity missing in the triple, and the process of generating the candidate entity list includes:
[0087] Input the input question into the embedding model, output a candidate entity list sorted by probability from high to low, and retain the top K entities in a truncated manner.
[0088] Optionally, in an embodiment of the present application, the large model fine-tuning module is specifically used for:
[0089] Use the PEFT fine-tuning framework to perform instruction fine-tuning on the large model, and use the QLoRA quantization method to quantize the parameters of the large model.
[0090] Optionally, in an embodiment of the present application, the entity prediction module is specifically used for:
[0091] Input the triple missing the head entity or the tail entity into the embedding model, output a candidate entity list sorted by probability from high to low, and retain the top K entities in a truncated manner;
[0092] Input the triple missing the head entity or the tail entity, the candidate entity list, and the relevant information into the fine-tuned large model, sort the candidate entity list through the fine-tuned large model, and select the entity with the highest probability in the sorted list as the predicted entity.
[0093] It should be noted that the foregoing explanation of the embodiment of the entity prediction method combining the embedding model and the instruction fine-tuned large model also applies to the entity prediction device combining the embedding model and the instruction fine-tuned large model in this embodiment, and will not be repeated here.
[0094] To implement the above embodiments, the present invention also proposes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in the above embodiments is implemented.
[0095] To implement the above embodiments, the present invention also proposes a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the above embodiments is implemented.
[0096] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0097] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of these features. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0098] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application belong.
[0099] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise appropriately processing if necessary, and then storing it in a computer memory.
[0100] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0101] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0102] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0103] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. An entity prediction method combining an embedding model and an instruction fine-tuning large model, characterized in that: include: Obtaining a dataset of entity prediction, and dividing the dataset into embedding model training data, fine-tuning instruction generation data, and hyperparameter adjustment data; Train the embedding model using the embedding model training data; Generating data and embedding models through the fine-tuning instruction, filling a fine-tuning instruction template, and obtaining a fine-tuning instruction dataset, wherein the fine-tuning instruction template includes an input question, relevant information, a candidate entity list, and a correct answer; Fine-tune the large model using the fine-tuning instruction data set, and during fine-tuning, use the hyperparameter adjustment data to adjust the hyperparameters of the large model, and select the optimal hyperparameter combination and the optimal fine-tuned large model; Given an entity prediction task, a candidate entity list is generated through the embedding model, and a fine-tuned large model is used to select the predicted entity from the candidate entity list.
2. The method according to claim 1, characterized in that For the entity prediction task, the input question is a triplet with a missing head entity or tail entity, the relevant information is a set of first-order adjacent entity relationship triples in the knowledge graph that are related to the input question, the correct answer is the missing entity in the triplet, and the generation process of the candidate entity list includes: The input question is fed into the embedding model, which outputs a list of candidate entities sorted from high to low probability, and retains the top K entities by truncation.
3. The method according to claim 1, characterized in that Fine-tune large models, including: The PEFT fine-tuning framework is used to fine-tune the instructions of the large model, and the QLoRA quantization method is used to quantize the parameters of the large model.
4. The method according to claim 2, characterized in that Complete entity prediction tasks, including: Input the triples with missing head or tail entities into the embedding model, output a list of candidate entities sorted by probability from high to low, and retain the top K entities by truncation; The triples of missing head entities or tail entities, the candidate entity list, and related information are input into the fine-tuned large model, the candidate entity list is sorted by the fine-tuned large model, and the entity with the highest probability in the sorted list is selected as the predicted entity.
5. An entity prediction device combining an embedded model and an instruction fine-tuning large model, characterized in that: include: A data set processing module, used for obtaining a data set for entity prediction, and dividing the data set into embedding model training data, fine-tuning instruction generation data, and hyperparameter adjustment data; An embedding model training module, used to train the embedding model using embedding model training data; A fine-tuning instruction data generation module, used to generate data and embedding models through the fine-tuning instructions, fill in the fine-tuning instruction template, and obtain a fine-tuning instruction data set, wherein the fine-tuning instruction template includes an input question, related information, a candidate entity list, and a correct answer; A large model fine-tuning module, used to use the fine-tuning instruction data set to fine-tune the large model. During fine-tuning, the hyperparameter adjustment data is used to adjust the hyperparameters of the large model, and the optimal hyperparameter combination and the optimal fine-tuned large model are selected; The entity prediction module is used to generate a candidate entity list through the embedding model for a given entity prediction task, and select the predicted entity from the candidate entity list using the fine-tuned large model.
6. The device according to claim 5, characterized in that For the entity prediction task, the input question is a triplet with a missing head entity or tail entity, the relevant information is a set of first-order adjacent entity relationship triples in the knowledge graph that are related to the input question, the correct answer is the missing entity in the triplet, and the generation process of the candidate entity list includes: The input question is fed into the embedding model, which outputs a list of candidate entities sorted from high to low probability, and retains the top K entities by truncation.
7. The device according to claim 4, characterized in that The large model fine-tuning module is specifically used for: The PEFT fine-tuning framework is used to fine-tune the instructions of the large model, and the QLoRA quantization method is used to quantize the parameters of the large model.
8. The device according to claim 4, characterized in that The entity prediction module is specifically used for: Input the triples with missing head or tail entities into the embedding model, output a list of candidate entities sorted by probability from high to low, and retain the top K entities by truncation; The triples of missing head entities or tail entities, the candidate entity list, and related information are input into the fine-tuned large model, the candidate entity list is sorted by the fine-tuned large model, and the entity with the highest probability in the sorted list is selected as the predicted entity.
9. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1 to 4 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.