A knowledge question and answer model construction method, a knowledge question and answer system, and a reasoning method

Through two-stage training and pattern-item linking strategies, the cross-knowledge base adaptability and query program transparency of the knowledge question answering model are improved, the problem of poor generalization of language models in existing technologies is solved, and more efficient reasoning of answers to complex questions is achieved.

CN119829722BActive Publication Date: 2025-10-17INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510013772.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-10-17
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

When dealing with complex questions, existing knowledge base question answering technologies have poor language model generalization, making it difficult to perform answer reasoning tasks for complex questions on unseen knowledge bases, and the generated query programs lack transparency and explainability.

Method used

A two-stage training strategy is adopted, including function understanding pre-training and instruction fine-tuning training, to construct function definition code and function call instances, combined with the pattern item linking strategy to improve the model's semantic understanding of complex problems and cross-knowledge base query capabilities.

Benefits of technology

It improves the generalization ability of the knowledge question answering model on different knowledge bases and the transparency of the query program, can generate answers more accurately, and improves the efficiency and explainability of answer reasoning for complex questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119829722B_ABST
    Figure CN119829722B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing a knowledge question-answering model, which includes using a pre-trained language model as a base model and performing two-stage training in the following manner to obtain a knowledge question-answering model: function understanding pre-training: constructing multiple function definition codes and multiple function call instances to train the base model so that the model learns the function execution principle; instruction fine-tuning training: obtaining task instructions, query instances, and query programs corresponding to the query instances to train the initial model so that the model learns the function combination programming principle. The technical solution of the present invention trains the knowledge question-answering model by proposing a two-stage training strategy, so that the model has the ability to understand functions used for knowledge programming, and the ability to convert complex problems into query programs based on the function understanding ability, thereby solving the problem in the prior art that the model tends to memorize the program itself and has difficulty learning semantic parsing, thereby improving the generalization of the model for out-of-domain queries with unseen semantics.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of knowledge reasoning, in particular to the field of knowledge reasoning, and more particularly to a knowledge question answering model construction method, a knowledge question answering system and a reasoning method. BACKGROUND

[0002] Knowledge base question answering technology is a key technology in the field of knowledge reasoning, which focuses on understanding and analyzing natural language questions raised by users, and then accurately matching these questions with relevant entities or concepts in the knowledge base to achieve fast query or intelligent generation of answers. This technology has two advantages, on the one hand, it can refer to the rich resources of external knowledge base to provide more professional and reliable responses for users, which effectively solves the limitations of traditional search engines in the face of knowledge-intensive problems, enabling users to more accurately obtain the information they need; on the other hand, knowledge base question answering technology can deeply understand the complex intentions of users, improving the accuracy and efficiency of information retrieval, which significantly reduces the time required by users to search for answers in a large amount of data, enabling faster response to user needs and providing better service. Based on these characteristics, knowledge base question answering technology has broad application potential in building professional knowledge question answering systems, especially in the fields of law, medicine and other professional knowledge-intensive and highly specialized fields.

[0003] Early knowledge base question answering technology mainly focuses on answering simple questions, which usually only involves a single fact. However, with the continuous progress of deep learning technology and the growing demand for application, knowledge base question answering technology has begun to deal with more complex problems, which often involve multi-hop reasoning, constraint relationships, numerical operations or combinations of these elements. For these complex problems, methods based on semantic parsing or information retrieval can be used to realize the reasoning of answers.

[0004] The method based on information retrieval realizes the reasoning of answers to complex problems by sequentially performing subgraph retrieval and answer sorting. In the subgraph retrieval stage, the entities mentioned in the question text are first identified, and then these entities are linked to the knowledge base, and then the subgraphs associated with these entities are extracted to form a set of candidate answers. In the answer sorting stage, a sorting model is used to model the problem query and the candidate answers to predict and select the highest score answer. Although such a processing method can realize the reasoning of answers to complex problems, the effectiveness of this method depends largely on the recall ability of candidate answers in the subgraph retrieval stage, when the answer is not in the extracted subgraph or further processing of the query result is needed to get the correct answer, such methods often fail to give the correct response. In addition, the reasoning process of this method lacks transparency, showing a black box feature, resulting in poor explainability of the answers obtained by reasoning.

[0005] The method based on semantic parsing parses the question into a corresponding query program through a trained language model, and executes the generated query program on the knowledge base to obtain the answer. For complex questions containing numerical operations and logical operations, this method can express the query semantics through flexible knowledge programming. However, this method relies on a large number of question-program annotations to endow the language model with the ability of semantic parsing and code generation, and has poor generalization for out-of-domain questions containing unseen query semantics.

[0006] Compared with the information retrieval-based method, the semantic parsing-based method can provide a query program with better interpretability. Therefore, most existing knowledge base question answering technologies use the semantic parsing-based method to infer the answer of a complex question. Although this method can parse the question through a trained language model and generate a corresponding query program to obtain the answer of the question, the language model is usually trained based on the independent and identically distributed assumption of data, which makes the language model tend to memorize the program itself and is difficult to learn the semantic parsing process. Moreover, the language model is usually trained independently on different knowledge bases, which can make the model learn the program skeleton and the knowledge base pattern item corresponding to the parameter in the program at the same time during the training process. However, the language model trained in this way only has good pattern item corresponding ability in the learned knowledge base, and is difficult to align the parameter in the program with the pattern item on the knowledge base in other unseen knowledge bases, which makes the query program generated by the language model difficult to find the correct answer when executed on the unseen knowledge base.

[0007] In summary, the existing technology has two deficiencies in performing the answer inference task of a complex question. On the one hand, the language model is usually trained based on the independent and identically distributed assumption of program data, which makes the language model tend to memorize the program itself and is difficult to learn the semantic parsing process, making the language model have poor generalization for out-of-domain questions containing unseen query semantics. On the other hand, the language model is usually trained independently on different knowledge bases, which makes the language model difficult to align the parameter in the program with the pattern item on the knowledge base in other unseen knowledge bases, resulting in the difficulty of the query program generated by the language model in performing the answer inference task of a complex question on the unseen knowledge base. SUMMARY

[0008] Therefore, the purpose of the present application is to overcome the above-mentioned deficiencies of the prior art, and to provide a knowledge question answering model construction method, a knowledge question answering system and a knowledge question answering inference method.

[0009] The purpose of the present application is achieved by the following technical solutions.

[0010] According to a first aspect of the present application, a method for constructing a knowledge question answering model is provided, the method comprising taking a pre-trained language model as a base model and performing two-stage training in the following manner to obtain a knowledge question answering model: function understanding pre-training: constructing a plurality of function definition codes and a plurality of function call instances to train the base model so that the model learns function execution principles; wherein the function understanding pre-training comprises: constructing a plurality of function definition codes and a plurality of function call instances; wherein each function definition code is a code text for creating a function; and each function call instance is a code text for executing a multi-step function call to answer a specific query; performing word segmentation processing on each function definition code and each function call instance respectively to convert each function definition code and each function call instance into sequence data, one function definition code or one function call instance corresponding to one sequence data; wherein each sequence data comprises a plurality of elements, each element representing a code snippet in the corresponding code text; taking the sequence data as input, generating a predicted output in a recursive prediction manner, and training the base model according to a preset target function until the model converges to obtain an initial model; instruction fine-tuning training: obtaining task instructions, query instances, and query programs corresponding to the query instances to train the initial model so that the model learns function combination programming principles; wherein the instruction fine-tuning training comprises: obtaining task instructions, a plurality of query instances, and query programs corresponding to each query instance from an existing data set; taking the task instructions and the query instances as input, the predicted query programs as output, and training the initial model according to a preset fine-tuning manner and a preset target function until convergence to obtain the knowledge question answering model.

[0011] In some embodiments of the present application, the function call instances are constructed in the following manner: a dependency graph corresponding to the preset function library is constructed, a plurality of solving paths are obtained by traversing the dependency graph, and a problem template is set for each solving path; wherein the dependency graph comprises a plurality of nodes and a plurality of directed edges connecting two nodes, the nodes represent functions in the function library, and the directed edges represent the direction of data flow between functions; each solving path represents a function call order that needs to be followed to answer a specific query category, and each solving path comprises a plurality of functions arranged in sequence; the problem template is a structured sentence for a specific query category, and the problem template comprises a plurality of placeholders; based on the function types included in each solving path and the problem template corresponding to the solving path, text parameters corresponding to each solving path are obtained by sampling parameters from an existing knowledge base; the sampled text parameters are filled into the placeholders of the problem template of the corresponding solving path to obtain a query for each solving path; and a query program is generated based on the solving path and the corresponding query.

[0012] In some embodiments of the present application, the pre-trained language model is a large language model.

[0013] In some embodiments of the present application, the preset target function is a cross-entropy loss function.

[0014] In some embodiments of the present application, the preset fine-tuning manner is lora fine-tuning.

[0015] According to a second aspect of the present application, a knowledge question answering system is provided for receiving a query and generating an answer according to a knowledge base, the system comprising: a knowledge base comprising a plurality of entities, relationships between the entities and entity concepts; a knowledge question answering model constructed by the method of the first aspect of the present application, for generating an initial query procedure corresponding to the query; a linking module for parameter linking the initial query procedure with the knowledge base according to a preset processing manner to obtain a target query procedure; and an execution module for executing the target query procedure to retrieve an answer from the knowledge base.

[0016] In some embodiments of the present application, the preset processing manner comprises: extracting a linking entity and a linking pattern item from the initial query procedure, wherein the linking pattern item is a relationship or an entity concept; calculating semantic similarity between the linking entity and each entity in the knowledge base to obtain a set of entity similarities, and calculating semantic similarity between the linking pattern item and each relationship and each entity concept in the knowledge base to obtain a set of pattern item similarities; selecting one or more entities ranked in the front from the set of entity similarities in descending order of similarity as the linking entities, and selecting one or more pattern items ranked in the front from the set of pattern item similarities in descending order of similarity as the linking pattern items; and replacing the linking entity and the linking pattern item in the initial query procedure with the linking entities and the linking pattern items respectively to obtain the target query procedure.

[0017] According to a third aspect of the present application, a knowledge question answering reasoning method is provided, the method comprising: step S1, obtaining a query to be processed; and step S2, obtaining an answer to the query to be processed by using the knowledge question answering system according to the second aspect of the present application.

[0018] Compared with the prior art, the present application has the following advantages: (1) a two-stage training strategy is proposed, and function definition code and function call instances are constructed to train the knowledge question answering model, so that the model has the ability to understand functions for knowledge programming, and the ability to convert complex problems into query procedures based on the function understanding ability; (2) a pattern item linking strategy is proposed, and by aligning the parameters in the program with the pattern items on the knowledge base, the query procedure generated by the knowledge question answering model can execute the answer reasoning task of complex problems on different knowledge bases. BRIEF DESCRIPTION OF DRAWINGS

[0019] The embodiments of the present application will be further described below with reference to the accompanying drawings, in which:

[0020] Figure 1A flowchart of a knowledge question and answer model construction method according to an embodiment of the present application is shown in FIG. 1.

[0021] Figure 2 A schematic diagram of a knowledge question and answer system according to an embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION

[0022] To make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0023] As mentioned in the background section, the prior art has two deficiencies in performing the answer inference task of complex problems. On the one hand, language models are usually trained based on independent and identically distributed program data, which makes the language model tend to memorize the program itself and makes it difficult to learn the semantic analysis process, so that the language model has poor generalization for problems outside the domain containing unseen query semantics. On the other hand, the language model tends to be trained independently on different knowledge bases, which makes it difficult for the language model to align the parameters in the program with the pattern items on the knowledge base on other unseen knowledge bases, resulting in the query program generated by the language model being difficult to perform the answer inference task of complex problems on unseen knowledge bases.

[0024] To solve the above problems, the inventors propose a two-stage training strategy and a pattern item linking strategy. The two-stage training strategy includes function understanding pre-training and instruction fine-tuning training. The function understanding pre-training trains the model by constructing multiple function definition codes and multiple function call instances to enable the model to learn the function execution principle. The instruction fine-tuning training fine-tunes the model after the function understanding pre-training to enable the model to learn the function combination programming principle. Through the two-stage training, the model can have the understanding ability of the function of knowledge programming, and the ability to convert complex problems into query programs based on the function understanding ability. Moreover, the model learns the process of problem semantic analysis during the two-stage training, rather than simply memorizing the program patterns in the training data, which helps to improve the generalization of the model for problems outside the domain containing unseen query semantics. The pattern item linking strategy refers to parameter linking processing of the query program generated by the knowledge question and answer model constructed by the two-stage training, to align the parameters in the program with the pattern items on the knowledge base, so that the query program can perform the answer inference task of complex problems on different knowledge bases.

[0025] In summary, as Figure 1As shown, the present application provides a knowledge question and answer model construction method, which comprises taking a pre-trained language model as a base model and performing two-stage training in the following manner to obtain a knowledge question and answer model: function understanding pre-training: constructing a plurality of function definition codes and a plurality of function call instances to train the base model to enable the model to learn function execution principles; wherein the function understanding pre-training comprises: constructing a plurality of function definition codes and a plurality of function call instances; wherein each function definition code is a code text for creating a function; each function call instance is a code text for executing a multi-step function call to answer a specific query; each function definition code and each function call instance is subjected to word segmentation processing to convert each function definition code and each function call instance into sequence data, one function definition code or one function call instance corresponding to one sequence data; wherein each sequence data comprises a plurality of elements, each element representing a code segment in the corresponding code text; taking the sequence data as input, generating a predicted output in a recursive prediction manner, and training the base model according to a preset target function until the model converges to obtain an initial model; instruction fine-tuning training: obtaining task instructions, query instances and query programs corresponding to the query instances to train the initial model to enable the model to learn function combination programming principles; wherein the instruction fine-tuning training comprises: obtaining task instructions, a plurality of query instances and query programs corresponding to each query instance from an existing data set; taking the task instructions and the query instances as input, the predicted query programs as output, training the initial model according to a preset fine-tuning manner according to a preset target function until the knowledge question and answer model is obtained.

[0026] In order to better understand the present application, the function understanding pre-training and the instruction fine-tuning training will be described in detail below in conjunction with specific embodiments.

[0027] I. Function understanding pre-training

[0028] The function understanding pre-training aims to enable the model to learn and understand the usage of basic functions. Through pre-training, it can ensure that the model performs true semantic understanding and program generation, rather than memorizing the patterns of programs in the training set, which is of great significance for the model to generalize to unseen queries in the real world. In order to achieve this, the training data used in the function understanding pre-training stage consists of two types of code style data, namely function definition codes and function call instance codes.

[0029] 1.1 Function definition code construction

[0030] The function definition code gives the input and output parameter types of the function, function comments, and specific description of usage, which can help the model understand the basic usage of the function from the perspective of basic information. The function definition code is written in python style, and the main body includes function name, reasonable parameters of function and function comment. The function name uses camel case, which can express clear and intuitive semantics; in the function parameter list, the python type library is used to limit the reasonable parameter types expected by the function; in the function comment, the function description, input and output semantics, function parameter semantics and a simple use example of the function are given in a standard format. These function definition codes can help the model understand the basic information of the function comprehensively. It is worth noting that the specific code of the function body is not given in the function definition code, because in the function understanding process, the model is not expected to learn the implementation of the function on the specific knowledge base, but to learn the understanding of the reasonable parameters of the function, the output result expected by taking a certain parameter as input, and the scene in which the function should be called.

[0031] In order to better understand the function definition code, a simple explanation is given with the function definition code example shown in table 1. As can be seen from table 1, the function definition code is a definition of a Python function, and the function name is Relate. The design purpose of this function is to find the tail entity connected to the given entity set with the given relationship according to the specified direction (direction is optional) in the case of giving an entity set and relationship name. Among them, Relate is the function name; entities is a parameter, the type is annotated as list, which means it should be a list containing entities; Relation_name is a parameter, the type is annotated as str, which means it is a string representing the relationship name; direction is an optional parameter, the type is annotated as Optional[str], which means it can be a string or None; Description indicates whether the input entity is the head entity or the tail entity of the relationship, and if the parameter is not set, it is assumed that the input entity is the head entity; Usage is used to describe how to use the function Relate, which may include some usage examples or scenarios; Arguments lists the parameters of the function and explains whether they are required or optional; pass is a null operation statement used in Python to represent an empty code block, which occupies the place to indicate that the specific implementation of the function will be added later.

[0032] Table 1

[0033]

[0034] 1.2 Function call instance construction

[0035] The function call instance gives a simple and complex use scenario of each function, and can help the model understand the specific use of the function from the perspective of application. For the function call instance code, synthesis is selected on the WikiData knowledge base. The call instance is a simple use sample of a function, including the user query question itself, multiple function call steps, and thinking notes before each function call. According to an embodiment of the application, the function call instance is constructed in the following manner: a corresponding dependency graph is constructed based on a preset function library, a plurality of solving paths are obtained by traversing the dependency graph, and a problem template is set for each solving path; wherein the dependency graph includes a plurality of nodes and a plurality of directed edges connecting two nodes, the nodes represent the functions in the function library, and the directed edges represent the data flow direction between the functions; each solving path represents the function call order that needs to be followed to answer a specific query category, and each solving path includes a plurality of functions arranged in order; the problem template is a structured sentence for a specific query category, and the problem template includes a plurality of placeholders; based on the function types included in each solving path and the problem template corresponding to the solving path, text parameters corresponding to each solving path are obtained by sampling parameters from an existing knowledge base; the text parameters obtained by sampling are filled into the placeholders of the problem template of the corresponding solving path to obtain the query of each solving path; and a query program is generated based on the solving path and the corresponding query.

[0036] In order to better understand the construction process of the function call instance, the following will be described in detail in combination with the foregoing embodiments.

[0037] As pre-training data, in order to enable the model to fully understand the use of each function in the function library, it is necessary to ensure that the function call instance has a relatively rich data pattern. In order to ensure this, the construction process of the function call instance is formalized into four steps of solving path construction, parameter sampling, query generation, and program generation.

[0038] Among them, the goal of solving path construction is to construct a plurality of solving paths to generate a plurality of problem instances. Although problems in the real world are infinite, the sequence of a series of functions required to solve the problem is limited. These paths represent the core ideas required to solve the problem. By enumerating the possible solving paths and sampling different values for the function parameters involved in each path, a plurality of problem instances can be derived from a single path; parameter sampling for different solving paths can generate diversified problem instances reflecting different core solution ideas, ensuring that the function call instances used for training are as diverse as possible.

[0039] The target of parameter sampling is to generate a function sequence with complete parameters, so that the function call can be correctly executed and produce the expected results, thereby generating different problems based on the solution idea (solution path). In the parameter sampling process, introducing a certain randomness not only enables the generation of function sequences corresponding to different problems, but also ensures the robustness of the function call instance, so that the model can adapt to different types of function call instances.

[0040] The target of query generation is to generate the specific problem solved by the function sequence obtained by parameter sampling. Each solution path is provided with a problem template, and the entities and relationships obtained after parameter sampling are filled into the problem template. Rewriting the problem using a large language model can generate a query corresponding to the function sequence.

[0041] The target of program generation is to generate a query program with program comments. In the function sequence obtained by parameter sampling, add comments before each function call to form the final function call instance. The purpose of this processing is to enable the function call instance to contain the explanation information required for the model to understand. Adding comments before each function call explains the purpose and expected effect of the call, so that the function call instance formed can not only help the model understand the specific purpose of each function, but also enable the model to learn how to combine multiple functions to solve complex problems.

[0042] Specifically, the implementation process of solution path construction, parameter sampling, query generation, and program generation is as follows.

[0043] Solution path construction: In order to ensure that the program skeletons constructed are different, it is necessary to first construct different solution paths. To achieve this, first construct a dependency graph representing the dependency relationship between functions based on the preset function library and perform a reasonableness test. The dependency graph treats functions as nodes and the dependency relationship between functions as edges. For example, if the output type of function A is the input parameter type of function B, there is a dependency edge from function A to function B between the two functions. Traverse the dependency graph to obtain multiple parameter transmission reasonable solution paths. These solution paths are referred to as solutions. After manually confirming the reasonableness of the solutions, these solutions are collectively formed into a library, referred to as the solution library. Among them, a solution is defined as a sequence of function executions traversing the dependency graph, using to represent each function in the solution path, and subscript to represent the step in the sequence, to represent the preset function library, all functions are enumerated from which, a solution (solution path) can be represented as: wherein, represents the solution, , , , represents each function in the solving path, and the subscript represents the execution order, , , represents the edge between functions.

[0044] It should be noted that the solution library represents a set of valid and reasonable solutions, and all solving paths (solutions) need to meet validity and reasonableness. Among them, validity means that in all the solutions traversed in the dependency graph, only the function as the first function call is valid. For example, the Find function for linking and finding entities with a specific name, because only such functions can accept the user's ambiguous natural language description as a parameter, thereby starting the subsequent precise query. Reasonableness refers to the fact that for some function call paths, their parameter passing paths may be valid, but do not have reasonable query semantics. For example, the FindAll function is generally called without a specific entity name, and queries all entities in the knowledge base, so it will not appear in the same query branch as Find; for the call path “[Find, FilterLiteral, FilterLiteral, FilterLiteral]”, an entity query is performed, and the entity set is filtered three times by literal constraints. In the real world, it is obviously difficult to access such complex queries from users, so these semantically unreasonable solving paths need to be filtered out.

[0045] Parameter sampling: for each solving path, due to the diversity of inputs and outputs of each function, even after the solution is established, it is still not possible to determine under what conditions to find which information (inference results / outputs). In order to determine the specific query semantics of each solving path, the text parameters of the functions in the solving path need to be determined, such as: the problem template corresponding to the solving path “[FindAll, FilterLiteral, Count]” is in the form of “How many <cn>whose <k> <op> <v>After filling the corresponding placeholders with the sampled text parameters, the complete solution and question can be obtained.

[0046] It should be noted that in the process of parameter sampling, not only the reasonable parameter of the function itself needs to be met, but also the dependency between functions needs to be considered. For example, in the solution path: ['QueryEntity', 'FilterLiteral', 'QueryLiteral'], the entity needs to be determined through the entity name ('QueryEntity') and a certain specific literal ('FilterLiteral'), and then the literal value of the entity ('QueryLiteral') is queried. At this time, the literal used to determine the entity should not be used again in the query, which will lead to unreasonable query semantics. Therefore, in the process of parameter sampling, the dependency between functions also needs to be considered.

[0047] Query generation: for the solution path that completes parameter sampling, it needs to be converted into the corresponding natural language question. Since the length of the function sequence needs to be controlled (i.e. the number of jumps of the query, the query in the real world generally will not exceed 3 jumps), a limited number of solution paths (33 solution paths are reserved in the present application) can be obtained. In order to ensure the quality of the natural language question, for each solution path, 1-2 question templates are written for each solution path, and after parameter sampling, the sampled parameters are filled into the placeholders corresponding to the question template to obtain the complete question. However, the question obtained after parameter filling may have grammatical errors or awkward questioning methods, such as: the complete question How many country whose population > 3000,000,000? (from the question template How many <cn>whose <k> <op> <v>To address this, a prompt can be constructed to guide the large language model to rewrite the question. A simple description of these functions is provided in the prompt, along with the meaning of each placeholder, giving instructions to guide the large model to modify the question syntax and rewrite the question as a more natural real-world query as much as possible, such as the above question can be rewritten as: "How many countries have a population greater than 3,000,000?".

[0048] Program generation: After the construction of the solving path, parameter sampling and query generation, a complete program is obtained, which is composed of multiple function calls, and the function called at each step is different. In order to facilitate the model to understand the query program, it is also necessary to add annotations to each function in the query program, so as to obtain a query program with function annotations as a function call instance.

[0049] Based on the foregoing four steps of solving path construction, parameter sampling, query generation and program generation, the function call instance shown in Table 2 can be obtained. Among them, the function call instance in Table 2 is a program example for understanding the function library of knowledge query, which uses a three-step process to answer the question: "Is the number of children of David Newman equal to 3?". In this program example, step 1: query entity, use the QueryEntity function to query the entity named "David Newman" from the knowledge base; step 2: use the QueryLiteral function to retrieve the attribute value according to the entity found in step 1 and the specified literal key "number of children"; step 3: use the Verify function to verify whether the literal value retrieved in step 2 is equal to 3.

[0050] Table 2

[0051]

[0052] 1.3 Base model training

[0053] Through the foregoing embodiments and specific implementation steps, a plurality of function definition codes and a plurality of function call instances can be obtained, and each function definition code and each function call instance is subjected to word segmentation processing to convert each function definition code and each function call instance into sequence data; taking the sequence data as input, generating a prediction output according to a recursive prediction manner, and training the base model according to a preset target function until the model converges to obtain an initial model. Among them, a sequence data can be represented as In the training process, the base model passes through the element to predict the element , the element and the element to predict the element , the element to predict , and the prediction loss of each sequence data is calculated according to a cross-entropy loss function to update the parameters of the base model until the model converges.

[0054] According to an embodiment of the present application, the pre-trained language model is a large language model. The large language model can adopt an open source large language model structure such as llama2-7B, qwen7B, etc., and the present application does not make specific limitations on the large language model. It should be noted that the large language model is adopted because the large language model (LLM) has achieved remarkable results in various natural language processing tasks, and the large language model has superior semantic understanding ability and few-shot learning ability, which can more accurately understand the user's input intent and quickly adapt and optimize the model performance under limited training samples.

[0055] II. Instruction fine-tuning training

[0056] The purpose of the instruction fine-tuning training is to make the initial model obtained after the function understanding pre-training further learn the process of converting complex problems into complete query programs, ensuring that the model performs true semantic understanding and program generation, rather than memorizing the patterns of programs in the training set.

[0057] Specifically, when performing instruction fine-tuning training, first, the task instruction, the plurality of query instances, and the query program corresponding to each query instance are obtained from the existing data set; then, the task instruction and the query instance are taken as input, the predicted query program is taken as output, and the initial model is trained according to the preset fine-tuning mode and the preset target function until the knowledge question and answer model is obtained by convergence. It should be noted that the task instruction, the plurality of query instances, and the query program corresponding to each query instance can be obtained from the existing KBQA data set, and the forms of the obtained task instruction, query instance, and query program corresponding to each query instance are consistent with the examples shown in Table 3. In Table 3, the task instruction indicates that the task of the model is to generate a query program with steps; the query instance is "Who is the winner of the Oscar Best Adapted Screenplay Award for the movie 'The Life of Mrs. Miniver'?"; and the query program indicates how to call the function to answer the question of the query instance.

[0058] Table 3

[0059]

[0060] According to one embodiment of the present invention, the preset fine-tuning method is Lora fine-tuning. Among them, Lora fine-tuning is an existing parameter-efficient fine-tuning method, which allows parameter optimization while keeping most of the weights of the initial model unchanged. It is achieved by only adjusting the low-rank matrix of a specific layer in the initial model, thereby greatly reducing the number of parameters that need to be optimized. It should be noted that in the fine-tuning process, the task instructions and query instances are used as input, the predicted query program is used as output, and the cross entropy loss is used to calculate the prediction loss. The specific calculation process is expressed as follows:

[0061]

[0062] in, represents the cross-validation loss, which quantifies the difference between the predicted query program and the true query program; Indicates the number of query instances; represents the natural logarithm function, Represents a given initial model parameter In the case of the initial model for the input Correctly predict output The probability of Represents input The corresponding prediction query program The label of the element, Expresses the sequence length of the predicted query program. It should be noted that, unlike the recursive prediction method used in function comprehension pre-training, during instruction fine-tuning training, the initial model predicts the entire query program based on the task instructions and query instances. The predicted query program can be viewed as sequence data consisting of multiple elements, each of which is a code snippet within the query program.

[0063] The knowledge question answering model constructed based on the above embodiment, such as Figure 2 As shown, the present application also provides a knowledge question-answering system for receiving a query and generating an answer according to a knowledge base, the system comprising: a knowledge base comprising a plurality of entities, relationships between the entities, and entity concepts; a knowledge question-answering model constructed by the method of the preceding embodiment, for generating an initial query procedure corresponding to the query; a linking module for parameter linking the initial query procedure with the knowledge base according to a preset processing mode, to obtain a target query procedure; and an execution module for executing the target query procedure to retrieve an answer from the knowledge base. The purpose of the linking module is to align the parameters in the initial query procedure with the content stored in the knowledge base. The reason for this is that the entities and relationship names in the query may not be consistent with the content stored in the knowledge base, and the parameters in the initial query procedure can only be extracted from the query text, which may result in the initial query procedure being unable to be normally executed on the knowledge base. Therefore, the entities and relationship parameters in the initial query procedure need to be extracted, and the extracted entities and relationship parameters need to be aligned with the content stored in the knowledge base, i.e., the parameters extracted from the query are replaced by the real content stored in the knowledge base, to ensure the normal execution of the procedure.

[0064] According to one embodiment of the present application, the preset processing mode is: extracting a to-be-linked entity and a to-be-linked pattern item from the initial query procedure, wherein the to-be-linked pattern item is a relationship or an entity concept; calculating the semantic similarity between the to-be-linked entity and each entity in the knowledge base to obtain a set of entity similarities, and calculating the semantic similarity between the to-be-linked pattern item and each relationship and each entity concept in the knowledge base to obtain a set of pattern item similarities; selecting one or more entities ranked in the front from the set of entity similarities in order of similarity from high to low as linking entities, and selecting one or more pattern items ranked in the front from the set of pattern item similarities in order of similarity from high to low as linking pattern items; and replacing the to-be-linked entity and the to-be-linked pattern item in the initial query procedure with the linking entity and the linking pattern item, respectively, to obtain the target query procedure.

[0065] In order to better understand the working principle of the knowledge question-answering system, a specific query example is used below to illustrate the working principle.

[0066] Suppose that the knowledge base in the knowledge question-answering system stores information of movies, including movie names, directors, actors, etc. Based on the content stored in the knowledge base, a query is constructed: "Who is the director of The Shawshank Redemption?". Taking this query as an example, the knowledge question-answering system is used to reason the answer to the query, and the specific reasoning process is as follows.

[0067] Firstly, the knowledge question answering model generates an initial query procedure corresponding to the query; wherein the initial query procedure is composed of two steps, the first step is to query the entity corresponding to the movie "The Shawshank Redemption" in the knowledge base: e0=QueryEntity("The Shawshank Redemption"); the second step is to find the tail entity of the fact composed of the entity and the "director" relationship based on the entity: e1=Relate(e0, "director"), and finally e1 is taken as the final answer: end(e1).

[0068] Then, the linking module links the initial query procedure and the knowledge base according to the preset processing mode to obtain a target query procedure. Assuming that the knowledge base has the following movie entities: "The Shawshank Redemption", "The Godfather" and "Forrest Gump", the semantic similarity between these entities and "The Shawshank Redemption" is calculated; assuming that the similarity scores are: "The Shawshank Redemption" - 0.95, "The Godfather" - 0.20, "Forrest Gump" - 0.15, the entities are sorted according to the similarity, and the top k entities with the highest similarity (assuming k=1) are selected, and the selected entity is "The Shawshank Redemption"; further, the pattern item parameter "director" is extracted from the generated initial query procedure, and it is assumed that the relationships related to the movie "The Shawshank Redemption" in the knowledge base are director, actor and type, and the semantic similarity between these relationships and "director" is calculated, assuming that the similarity scores are director - 0.90, actor - 0.10, and type - 0.05, the entities are sorted according to the similarity, and the top k entities with the highest similarity (assuming k=1) are selected, and the selected pattern item is director; the target query procedure is obtained by replacing the initial query procedure with "The Shawshank Redemption" and the director respectively.

[0069] Finally, the execution module executes the target query procedure to retrieve the answer from the knowledge base, and obtains that the director of "The Shawshank Redemption" is Frank Darabont.

[0070] Based on the knowledge question answering system described in the foregoing embodiments, the present application further proposes a knowledge question answering reasoning method, the method comprising: step S1, obtaining a query to be processed; step S2, using the knowledge question answering system as described in the foregoing embodiments to obtain an answer to the query to be processed.

[0071] The beneficial effects of the present application are: (1) a two-stage training strategy is proposed, and function definition code and function call instances are constructed to train the knowledge question and answer model, so that the model has the understanding ability of the function used for knowledge programming, and the ability to convert complex problems into query programs based on the function understanding ability; (2) a pattern item linking strategy is proposed, and by aligning the parameters in the program with the pattern items on the knowledge base, the query program generated by the knowledge question and answer model can execute the answer reasoning task of the complex problem on different knowledge bases.

[0072] It should be noted that although the above describes each step in a specific order, it does not mean that each step must be performed in the above specific order, in fact, some of the steps can be executed concurrently, or even changed in order, as long as the required function can be realized.

[0073] The present application can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions stored therein to implement various aspects of the present application.

[0074] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, and any suitable combination of the foregoing. A non-transitory, computer-readable storage medium, as used herein, is expressly intended to encompass a computer readable medium that is tangible and excludes transitory signals.

[0075] The embodiments of the present application have been described above, the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles, practical applications or technical improvements in the art of the embodiments, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.< / v> < / op> < / k> < / cn> < / v> < / op> < / k> < / cn>

Claims

1. A method for constructing a knowledge question answering model, characterized in that: The method includes using a pre-trained language model as a base model and performing two-stage training to obtain a knowledge question answering model as follows: Function understanding pre-training: Build multiple function definition codes and multiple function call instances to train the base model so that the model can learn the function execution principles. Function understanding pre-training includes: Constructing multiple function definition codes and multiple function call instances; wherein each function definition code is code text for creating a function; and each function call instance is code text for executing a multi-step function call to answer a specific query; Performing word segmentation processing on each function definition code and each function call instance respectively to convert each function definition code and each function call instance into sequence data, where one function definition code or one function call instance corresponds to one sequence data; wherein each sequence data includes multiple elements, each element representing a code fragment in the corresponding code text; Taking sequence data as input, generating prediction output according to a recursive prediction method, and training the base model according to a preset objective function until the model converges to obtain an initial model, wherein the sequence data includes multiple elements arranged in order, and generating prediction output according to the recursive prediction method means predicting the second to the last elements in the sequence data in sequence, and predicting each element based on all previous elements; Instruction fine-tuning training: Obtain task instructions, query instances, and the query programs corresponding to the query instances to train the initial model so that the model can learn the principles of functional composition programming. Instruction fine-tuning training includes: Obtaining a task instruction, a plurality of query instances, and a query program corresponding to each query instance from an existing data set; Taking task instructions and query instances as input and predicted query programs as output, the initial model is trained according to the preset objective function using a preset fine-tuning method until convergence to obtain a knowledge question-answering model, where the query program indicates how to call a function to answer the question of the query instance.

2. The method according to claim 1, characterized in that Construct a function call instance as follows: Based on a preset function library, a corresponding dependency graph is constructed, the dependency graph is traversed to obtain multiple solution paths, and a problem template is set for each solution path. The dependency graph includes multiple nodes and multiple directed edges connecting two nodes. The nodes represent functions in the function library, and the directed edges represent the direction of data flow between functions. Each solution path represents the function call sequence that needs to be followed to answer a specific query category, and each solution path includes multiple functions arranged in sequence. The problem template is a structured statement for a specific query category, and the problem template includes multiple placeholders. Based on the function type contained in each solution path and the problem template corresponding to the solution path, parameter sampling is performed from the existing knowledge base to obtain the text parameters corresponding to each solution path; Fill the sampled text parameters into the placeholders of the problem template corresponding to the solution path to obtain the query of each solution path; A query program is generated based on the solution paths and their corresponding queries.

3. The method according to claim 2, characterized in that The pre-trained language model is a large language model.

4. The method according to claim 3, characterized in that The preset objective function is the cross entropy loss function.

5. The method according to claim 4, characterized in that The preset fine-tuning method is lora fine-tuning.

6. A knowledge question answering system for receiving queries and generating answers based on a knowledge base, characterized in that: The system comprises: A knowledge base, which includes multiple entities, relationships between entities, and entity concepts; The knowledge question answering model constructed by the method according to any one of claims 1 to 5, used to generate an initial query program corresponding to the query; A linking module is used to link the initial query program with the knowledge base according to a preset processing method to obtain a target query program; The execution module executes the target query program to retrieve the answer from the knowledge base.

7. The system according to claim 6, characterized in that The preset processing method is: Extracting entities to be linked and schema items to be linked from an initial query program, wherein the schema items to be linked are relationships or entity concepts; Calculate the semantic similarity between the entity to be linked and each entity in the knowledge base to obtain an entity similarity set, and the semantic similarity between the pattern item to be linked and each relationship and each entity concept in the knowledge base to obtain a pattern item similarity set; Selecting one or more entities ranked first in the entity similarity set in descending order of similarity as link entities, and selecting one or more pattern items ranked first in the pattern item similarity set in descending order of similarity as link pattern items; The target query program is obtained by using the link entity and the link pattern item to replace the to-be-linked entity and the to-be-linked pattern item in the initial query program respectively.

8. A knowledge question answering reasoning method, characterized in that: The method comprises: Step S1: Obtain pending queries; Step S2: Use the knowledge question answering system as described in any one of claims 6-7 to obtain the answer to the query to be processed.

9. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of any one of the methods of claims 1-5 and 8.

10. An electronic device, characterized in that: include: one or more processors, and a memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method of any one of claims 1-5 and 8 by executing the executable instructions.

Citation Information

Patent Citations

  • Session text matching method and device, storage medium and equipment

    CN117574877A

  • Vertical type government affair large model service method and system based on interactive learning

    CN118820448A