Method and device for generating questions and answers based on knowledge graph

By classifying problem, identifying entity and linking entities in the RAG method, combining knowledge graphs and ASP rules, the problem of poor searching information in the knowledge graph data source is solved, achieving more efficient information retrieval and more accurate answer generation.

CN119938846APending Publication Date: 2025-05-06SHANGHAI HANSHUO ZHIRONG INFORMATION TECHNOLOGY TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510024189.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing RAG methods are difficult to effectively retrieve information in the knowledge base with knowledge graphs as the data source, resulting in poor retrieval information ability and low inference performance of the model.

Method used

By text classification, entity recognition and entity linking received questions, candidate entities are determined and similarity calculations are performed with entities in the knowledge graph to obtain the target entity. Then, the search strategy and ASP rules are determined according to the classification of the questions, and the entity set is processed to generate the fact set, and finally the answer is generated through the large language model.

Benefits of technology

Perform efficient searches to improve the inference performance of the model and generate more accurate answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938846A_ABST
    Figure CN119938846A_ABST
Patent Text Reader

Abstract

The invention relates to a method and device for generating questions and answers based on a knowledge graph. According to the method, received questions are analyzed through a text classification model, and the classification of the questions is determined; performing entity recognition through an entity recognition model according to the question, and determining candidate entities in the question; performing similarity calculation on the candidate entities and entities in the knowledge graph to obtain target entities corresponding to the candidate entities in the knowledge graph; determining a retrieval strategy according to the classification of the questions, and retrieving the target entities through a knowledge base according to the retrieval strategy to generate an entity set; determining a corresponding ASP rule according to the classification of the problem, and processing the entity set according to the ASP rule to obtain a fact set; and arranging the fact set through a large language model to generate answers to the corresponding questions. The method has the advantages that retrieval can be effectively carried out, the reasoning performance of the model is improved, and more accurate answers are generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technology applications, and in particular to a method and device for generating questions and answers based on a knowledge graph. Background Art

[0002] In the field of question answering (e.g., artificial intelligence products, intelligent customer service system products), the current "Retrieval-Augmented Generation" (RAG) method faces the following two major challenges:

[0003] Challenge 1: Existing document-based RAGs are mainly trained and evaluated on open-domain tasks. However, knowledge is usually stored in a structured form, such as knowledge graphs and structured knowledge bases, which makes it difficult for document-based retrievers to effectively retrieve information. This makes it difficult for RAGs to obtain effective information in knowledge bases that use knowledge graphs as data sources.

[0004] Challenge 2: When faced with common multi-constraint problems, existing RAG methods often introduce a large amount of irrelevant knowledge in the retrieval stage, resulting in excessively long context length, which not only increases economic costs but also affects the reasoning performance of the model.

[0005] At present, no effective solution has been proposed for the problem in the prior art that the RAG method cannot effectively retrieve information and affects the reasoning performance of the model, resulting in poor information retrieval ability and low reasoning performance of the model. Summary of the invention

[0006] The purpose of the present invention is to provide a method and device for generating questions and answers based on a knowledge graph in order to address the deficiencies in the prior art, so as to solve the technical problems in the prior art that the RAG method cannot effectively retrieve information and affects the reasoning performance of the model, resulting in poor information retrieval capability and low reasoning performance of the model.

[0007] To achieve the above object, the technical solution adopted by the present invention is:

[0008] The present invention provides a method for generating questions and answers based on a knowledge graph, comprising: analyzing a received question through a text classification model to determine the classification of the question; performing entity recognition through an entity recognition model based on the question to determine candidate entities in the question; calculating similarities between the candidate entity and an entity in the knowledge graph to obtain a target entity corresponding to the candidate entity in the knowledge graph; determining a retrieval strategy based on the classification of the question, and searching the target entity through a knowledge base based on the retrieval strategy to generate an entity set; determining a corresponding ASP rule based on the classification of the question, and processing the entity set based on the ASP rule to obtain a fact set; and organizing the fact set through a large language model to generate an answer to the corresponding question.

[0009] Optionally, the received question is analyzed through a text classification model to determine the classification of the question, including: when the text classification model is a K-BERT model trained with a knowledge graph, segmenting the question to obtain a text sequence after segmentation; and classifying the text sequence after segmentation through the K-BERT model to obtain the classification of the question.

[0010] Further, optionally, entity recognition is performed through an entity recognition model based on the question to determine candidate entities in the question, including: recognizing semantic information in the question through a K-BERT model; processing the semantic information through a bidirectional long short-term memory network to obtain a feature representation sequence; processing the feature representation sequence according to a conditional random field layer to obtain candidate entities.

[0011] Optionally, calculating the similarity between the candidate entity and the entity in the knowledge graph to obtain the target entity corresponding to the candidate entity in the knowledge graph includes: calculating the vector similarity between the candidate entity and attributes and the entities and attributes in the knowledge graph to obtain the vector similarity of the entities and attributes corresponding to the candidate entity and attributes in the knowledge graph; determining the entities and attributes corresponding to the candidate entity and attributes in the knowledge graph based on the vector similarity to obtain the target entity and the attributes of the target entity.

[0012] Further, optionally, a retrieval strategy is determined based on the classification of the question, and the target entity is retrieved through the knowledge base based on the retrieval strategy, and generating an entity set includes: determining the retrieval strategy based on the classification of the question, and generating a corresponding knowledge set for each entity tuple in the target entity based on the retrieval strategy; calculating the vector cosine similarity between the attributes of each identified entity in the knowledge set and the target entity, and determining the score of the attribute; selecting at least two attributes from the knowledge set based on the attribute score to generate an entity set.

[0013] Optionally, the corresponding ASP rules are determined according to the classification of the problem, and the entity set is processed according to the ASP rules to obtain the fact set, including: when the problem is classified as a quantitative statistical problem, the first ASP rule is used to locate all entities that meet the specified constraints and list the names, and all entities that meet the specified constraints and list the names are determined as the fact set; when the problem is classified as a maximum value problem, the second ASP rule is used to calculate the maximum values ​​of all attributes, and the maximum values ​​are represented by ASP triples, and the maximum values ​​represented by ASP triples are determined as the fact set; when the problem is classified as an enumeration problem, the first ASP rule is used to obtain the entity type that meets the specific constraint conditions, and the entity type that meets the specific constraint conditions is determined as the fact set; when the problem is classified as an intersection problem with multiple constraints, On the basis of the first ASP rule, the same constraints are added to determine the entities that meet multiple constraints as the fact set; when the problem is classified as a difference set problem, the third ASP rule is used to exclude entities that meet the first constraint but do not meet the second constraint by using keywords, and the entities that meet the first constraint but do not meet the second constraint are determined as the fact set; when the problem is classified as a single entity problem and a multi-entity problem under finite entity restrictions, ASP refinement is performed based on the corresponding knowledge of single entities and multi-entities under finite entity restrictions to obtain a fact set; when the problem is classified as a single event problem, the fillable field is set as a hard constraint through the fourth ASP rule, and the soft constraint is set through the fifth ASP rule, and the entities screened according to the hard constraints and soft constraints are determined as the fact set.

[0014] The present invention provides a device for generating questions and answers based on a knowledge graph, comprising: a question classification module, used for analyzing received questions through a text classification model to determine the classification of the questions; an entity recognition module, used for performing entity recognition through an entity recognition model based on the questions to determine candidate entities in the questions; an entity linking module, used for calculating similarities between candidate entities and entities in the knowledge graph to obtain target entities corresponding to the candidate entities in the knowledge graph; a knowledge retrieval module, used for determining a retrieval strategy based on the classification of the questions, and searching the target entities through a knowledge base based on the retrieval strategy to generate an entity set; a knowledge refinement module, used for determining corresponding ASP rules based on the classification of the questions, and processing the entity set based on the ASP rules to obtain a fact set; and a problem solving module, used for collating the fact set through a large language model to generate an answer to the corresponding question.

[0015] Optionally, the question classification module includes: a word segmentation unit, which is used to segment the question when the text classification model is a K-BERT model trained with a knowledge graph, to obtain a text sequence after word segmentation; and a question classification unit, which is used to classify the text sequence after word segmentation through the K-BERT model to obtain a classification of the question.

[0016] Furthermore, optionally, the entity recognition module includes: a semantic recognition unit, used to recognize semantic information in the question through a K-BERT model; an information processing unit, used to process the semantic information through a bidirectional long short-term memory network to obtain a feature representation sequence; and an entity recognition unit, used to process the feature representation sequence according to a conditional random field layer to obtain a candidate entity.

[0017] Optionally, the entity linking module includes: a similarity calculation unit, which is used to perform vector similarity calculation on the candidate entities and attributes with the entities and attributes in the knowledge graph, and obtain the vector similarity of the entities and attributes corresponding to the candidate entities and attributes in the knowledge graph; an entity linking unit, which is used to determine the entities and attributes corresponding to the candidate entities and attributes in the knowledge graph based on the vector similarity, and obtain the target entity and the attributes of the target entity.

[0018] The present invention adopts the above technical scheme, and determines the classification of the question by analyzing the received question through a text classification model; performs entity recognition through an entity recognition model based on the question to determine the candidate entity in the question; calculates the similarity between the candidate entity and the entity in the knowledge graph to obtain the target entity corresponding to the candidate entity in the knowledge graph; determines the retrieval strategy based on the classification of the question, and searches the target entity through the knowledge base based on the retrieval strategy to generate an entity set; determines the corresponding ASP rule based on the classification of the question, and processes the entity set according to the ASP rule to obtain a fact set; organizes the fact set through a large language model to generate the answer to the corresponding question. Compared with the prior art, it has the following technical effects: it can effectively perform retrieval to improve the reasoning performance of the model and generate more accurate answers. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flow chart of a method for generating questions and answers based on a knowledge graph according to an embodiment of the present invention;

[0020] Figure 2 is a schematic diagram of data flow in a method for generating questions and answers based on a knowledge graph according to an embodiment of the present invention;

[0021] Figure 3 is a structural schematic diagram of question and answer generation in a method for generating questions and answers based on a knowledge graph according to an embodiment of the present invention;

[0022] Figure 4 is a flowchart of data processing in a method for generating questions and answers based on a knowledge graph according to an embodiment of the present invention;

[0023] Figure 5 It is a schematic diagram of a device for generating questions and answers based on a knowledge graph according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0025] Obviously, the drawings described below are only some examples or embodiments of the present application. For ordinary technicians in this field, the present application can also be applied to other similar scenarios based on these drawings without creative work. In addition, it can also be understood that although the efforts made in this development process may be complicated and lengthy, for ordinary technicians in this field related to the content disclosed in this application, some changes in design, manufacturing or production based on the technical content disclosed in this application are just conventional technical means, and should not be understood as insufficient content disclosed in this application.

[0026] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0027] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantitative limitation, and may represent the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or units (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" / "several" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships, for example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0028] Technical terms involved in this application:

[0029] Answer Set Programming: Answer Set Programming, referred to as ASP;

[0030] Conditional Random Fields: Conditional Random Fields, referred to as CRF;

[0031] Bidirectional Long Short-Term Memory Network: Bidirectional Long Short-Term Memory, referred to as BiLSTM;

[0032] Knowledge-enhanced pre-trained language representation model based on the BERT model: Knowledge-enabled Bidirectional Encoder Representation from Transformers, referred to as K-BERT.

[0033] Example 1

[0034] An exemplary embodiment of the present invention is as follows Figure 1 As shown, Figure 1: is a flow chart of a method for generating questions and answers based on a knowledge graph according to an embodiment of the present invention; a method for generating questions and answers based on a knowledge graph provided in an embodiment of the present application includes:

[0035] Step S100, analyzing the received question through a text classification model to determine the classification of the question;

[0036] Optionally, in step S100, the received question is analyzed through a text classification model to determine the classification of the question, including: when the text classification model is a K-BERT model trained with a knowledge graph, segmenting the question to obtain a text sequence after segmentation; and classifying the text sequence after segmentation through the K-BERT model to obtain the classification of the question.

[0037] Specifically, Figure 2 is a schematic diagram of data flow in a method for generating questions and answers based on a knowledge graph according to an embodiment of the present invention; Figure 2 As shown, corresponding to the question classification part, the questions are classified by the text classification model trained on the question dataset to determine the predefined categories.

[0038] The question classification task requires determining the type of question. The method for generating questions and answers based on a knowledge graph provided in the embodiment of the present application uses a text classification model trained on a data set to identify questions and determine the preset type corresponding to the question. Subsequent knowledge retrieval and problem-solving methods will set corresponding strategies based on different question types.

[0039] In terms of specific implementation, in order to enhance the model's ability in knowledge understanding and improve its overall performance, the method for generating questions and answers based on knowledge graphs provided in the embodiments of the present application uses K-BERT as a pre-training model to integrate structured knowledge graph knowledge into the model. K-BERT enhances the ability to capture entities and semantics by combining knowledge graphs and semantic contexts. Using these word vectors to initialize word nodes helps improve the classification effect of the model. The method for generating questions and answers based on knowledge graphs provided in the embodiments of the present application uses the K-BERT model to embed the segmented text sequence q into a high-dimensional vector space hq, and calculates the dependency relationship between each word and other words through multiple self-attention layers, thereby capturing global 5 context information, and finally uses the classification layer to convert the encoded representation hq into a specific classification result cr. The specific process is:

[0040] hq=K-BRRT(q); formula (1).

[0041] Step S102, performing entity recognition through an entity recognition model according to the question to determine candidate entities in the question;

[0042] Optionally, in step S102, entity recognition is performed through an entity recognition model based on the question to determine candidate entities in the question, including: recognizing semantic information in the question through a K-BERT model; processing the semantic information through a bidirectional long short-term memory network to obtain a feature representation sequence; processing the feature representation sequence according to a conditional random field layer to obtain candidate entities.

[0043] Among them, entity linking is based on the vectorization of the entity sequence in the entity recognition step, and the sequence with the maximum score calculated by the maximum likelihood formula is the corresponding candidate entity, that is, the linked entity.

[0044] Specifically, Figure 2 As shown, in the corresponding entity recognition part, the model trained in the entity recognition corpus is used to process the problem and extract relevant candidate entities and attributes. The subsequent entity linking step relies on the basis provided by the step S102 process.

[0045] The goal of the entity recognition task is to identify candidate entities in the question. The recognition results (i.e., candidate entities in the embodiment of the present application) will be used as input to the entity linking task for entity disambiguation and the knowledge refinement task for knowledge filtering.

[0046] 1) Use a pre-trained deep learning model to obtain semantic information in the question. Since pre-trained models such as BERT mainly rely on large-scale general corpora, they lack specific fields and limit the understanding of knowledge in specific fields. The method for generating questions and answers based on knowledge graphs provided in the embodiments of the present application adopts K-BERT (i.e., the K-BERT model in the embodiments of the present application). The K-BERT model can inject domain knowledge graph information into the text and deal with professional vocabulary and specific concepts more effectively.

[0047] The specific process is as follows: for a given input sequence s={c0, c1, c2, c3, ..., c n-1}, the hidden layer encoding sequence S of K-BERT of S is obtained through the process shown in the formula v , S v =K-BERT(S); Formula (2);

[0048] 2) Add a bidirectional long short-term memory network (BiLSTM, i.e., the bidirectional long short-term memory network in the embodiment of the present application) after the K-BERT model to improve the ability to process long texts. The introduction of BiLSTM enables the model to better capture the long-distance dependencies between words, optimize the text feature expression, and enhance the accuracy of named entity recognition, thereby improving the accuracy of recognition. The specific process is: for the hidden layer encoding S obtained by K-BERT, v, after being processed by the BiLSTM network, a feature representation sequence X (i.e., a feature representation sequence provided in an embodiment of the present application) is obtained;

[0049] 3) Conditional Random Fields (CRF) use the label transfer patterns and constraints learned from the data to ensure that the model accurately constructs a compliant entity recognition sequence and improves the accuracy of label prediction;

[0050] After using BiLSTM to represent text features, the CRF layer is integrated to accurately identify different entity types. In the specific process, assume that the feature representation sequence input to the CRF layer is X, the predicted label sequence is Y, and the score matrix obtained by the Encoder layer is P, where P ij represents the i-th input sequence X i Predicted as the jth label Y j The score of the overall predicted label sequence Y can be obtained as follows:

[0051] On this basis, the probability of Y can be calculated:

[0052] where Y yi,yi+1 Indicates label Y i Transfer to Y i+1 The score of , which constitutes the entire label transfer score matrix M, is the real label sequence, Y x CRF optimizes the model by finding the maximum likelihood function for the predicted label sequence Y score function, and finally obtains the sequence Y with the maximum score. * .

[0053]

[0054] Step S104, calculating the similarity between the candidate entity and the entity in the knowledge graph to obtain the target entity corresponding to the candidate entity in the knowledge graph;

[0055] Optionally, in step S104, similarity is calculated between the candidate entity and the entity in the knowledge graph to obtain the target entity corresponding to the candidate entity in the knowledge graph, including: vector similarity is calculated between the candidate entity and attributes and the entities and attributes in the knowledge graph to obtain the vector similarity of the entities and attributes corresponding to the candidate entity and attributes in the knowledge graph; the entities and attributes corresponding to the candidate entity and attributes in the knowledge graph are determined based on the vector similarity to obtain the target entity and the attributes of the target entity.

[0056] Specifically, Figure 2As shown in the figure, in the entity linking part, based on the pre-trained model, vector similarity and other technologies are used to align the candidate entities and attributes with the target entities and attributes stored in the knowledge graph to identify the entities related to the problem and their constraints;

[0057] In entity linking, vector similarity is calculated to match candidate entities with corresponding entities in the knowledge graph. The generated linked entities are used in the knowledge retrieval module to obtain relevant background knowledge from the knowledge graph.

[0058] Step S106, determining a search strategy according to the classification of the question, and searching the target entity through the knowledge base according to the search strategy to generate an entity set;

[0059] Among them, in the embodiment of the present application, facts are modeled through ASP, and the modeling process includes: converting the problem into a identifiable entity set, and determining the retrieval strategy based on the classification of the problem, generating a corresponding knowledge set for each entity tuple in the target entity based on the retrieval strategy; calculating the vector cosine similarity between the attributes of each identified entity in the knowledge set and the target entity, and determining the score of the attribute; selecting at least two attributes from the knowledge set based on the attribute score to generate an entity set.

[0060] For example: When modeling with ASP, the relationship between entity sets and knowledge sets is usually expressed as facts and rules:

[0061] 1. Solid Modeling:

[0062] entity(e1,"person").

[0063] entity(e2,"book").

[0064] entity(e3,"city").

[0065] The above facts indicate that entity e1 is a person, e2 is a book, and e3 is a city.

[0066] 2. Knowledge Modeling:

[0067] knowledge(e1,"name","Alice").

[0068] knowledge(e1,"age",30).

[0069] knowledge(e2,"title","AI for Everyone").

[0070] knowledge(e2,"author","Alice").

[0071] knowledge(e3,"population",500000).

[0072] Here, the fact indicates that e1's name is Alice, the title of e2 is "AI for Everyone", and the author is Alice.

[0073] Optionally, in step S106, a retrieval strategy is determined based on the classification of the question, and the target entity is retrieved through the knowledge base based on the retrieval strategy. Generating an entity set includes: determining a retrieval strategy based on the classification of the question, and generating a corresponding knowledge set for each entity tuple in the target entity based on the retrieval strategy; calculating the vector cosine similarity between the attributes of each identified entity in the knowledge set and the target entity, and determining the score of the attribute; selecting at least two attributes from the knowledge set based on the attribute score to generate an entity set.

[0074] Specifically, Figure 2 As shown, corresponding to the knowledge retrieval part, based on entity information and problem categories, the knowledge related to problem solving is extracted, and the facts are modeled through ASP to provide support for subsequent steps.

[0075] Among them, the method for generating questions and answers based on knowledge graphs provided in the embodiment of the present application will retrieve relevant entities from the knowledge base to supplement the entity set E. The method for generating questions and answers based on knowledge graphs provided in the embodiment of the present application formulates a specific retrieval strategy according to the type of question, and uses the retrieval strategy to generate a corresponding knowledge set K for each entity tuple, where K i The knowledge set representing the i-th entity is used as part of the retrieval results to simplify the system structure and improve maintainability and reusability.

[0076] In question S, the set of relevant knowledge associated with the i-th identified named entity is denoted as K i ={p i1 , p i2 , p i3 , ..., p in}, k i ∈k,P ij represents the jth attribute, which belongs to the i-th entity. By calculating S and P ij The vector cosine similarity between them is used to obtain the attribute score (ie, the attribute score in the embodiment of the present application), which is used to evaluate the relevance of the attributes, followed by sorting, and finally selecting the top K attributes for subsequent processes.

[0077] Among them, the entity set E is the entity and type of the problem, and the entity tuple is one of the entity set E. For each entity tuple, the corresponding knowledge set is generated by the retrieval strategy. The knowledge set is the attribute set of the entity generated by the retrieval strategy.

[0078] Step S108, determining the corresponding ASP rules according to the classification of the question, and processing the entity set according to the ASP rules to obtain a fact set;

[0079] Optionally, in step S108, the corresponding ASP rule is determined according to the classification of the problem, and the entity set is processed according to the ASP rule to obtain a fact set, including: when the problem is classified as a quantitative statistical problem, the first ASP rule is used to locate all entities that meet the specified constraints and list the names, and all entities that meet the specified constraints and list the names are determined as the fact set; when the problem is classified as a maximum value problem, the second ASP rule is used to calculate the maximum values ​​of all attributes, and the maximum values ​​are represented by ASP triples, and the maximum values ​​represented by ASP triples are determined as the fact set; when the problem is classified as an enumeration problem, the first ASP rule is used to obtain the entity type that meets the specific constraint conditions, and the entity type that meets the specific constraint conditions is determined as the fact set; when the problem is classified as an intersection problem with multiple constraints In this case, the same constraints are added on the basis of the first ASP rule, and the entities that meet the multiple constraints are determined as the fact set; when the problem is classified as a difference set problem, the third ASP rule is used to exclude entities that meet the first constraint but do not meet the second constraint by using keywords, and the entities that meet the first constraint but do not meet the second constraint are determined as the fact set; when the problem is classified as a single entity problem and a multi-entity problem under finite entity restrictions, ASP refinement is performed based on the corresponding knowledge of single entities and multi-entities under finite entity restrictions to obtain the fact set; when the problem is classified as a single event problem, the fillable field is set as a hard constraint through the fourth ASP rule, and the soft constraint is set through the fifth ASP rule, and the entities screened according to the hard constraints and soft constraints are determined as the fact set.

[0080] Specifically, Figure 2 As shown, in the corresponding knowledge refinement part, in the embodiment of the present application, ASP can conveniently represent constraints, and its solver can find all solution sets that satisfy the constraints through automatic reasoning. During the knowledge refinement process, the framework uses the solver embedded in ASP to calculate and reason about the instantiated ASP facts and rules, which can effectively compress the amount of knowledge, thereby reducing the reasoning burden of the large language model when solving problems. In addition, ASP can also directly give answers, reducing the reliance on large language models in some scenarios.

[0081] Among them, ASP has powerful description and solution capabilities. The embedded search and reasoning functions of ASP can search for all possible solution sets under given constraints, which is very suitable for dealing with multiple constraint problems. At the same time, ASP can find the optimal solution that meets the soft constraint requirements under the premise of satisfying the hard constraints, and can reduce the error accumulation caused by text understanding in the process of solving multiple constraint problems.

[0082] When dealing with quantitative statistical problems, that is, calculating the number of entities that meet specified constraints, in order to improve the explainability of the solution process, the method for generating questions and answers based on the knowledge graph provided in the embodiment of the present application adopts ASP rule 7 (that is, the first ASP rule in the embodiment of the present application) to locate all entities that meet the specified constraints and list the names, rather than directly calculating the number.

[0083] Constraint(Entity):-entity(Entity,_,”{entity_constraint}”) formula (7)

[0084] For maximum value type questions (collectively referred to as maximum value type questions together with minimum value type questions), in view of their diverse expressions, the method for generating questions and answers based on knowledge graphs provided in the embodiment of the present application uses ASP rule 8 (i.e., the second ASP rule in the embodiment of the present application) for processing. It does not rush to determine which specific attribute it is, but calculates the maximum value of all relevant attributes and expresses it in the form of ASP triples. Finally, the triples and questions are input into the large language model together to generate answers. This design realizes delayed attribute linking, and the steps for finding the minimum value are similar to those described in formula (8). The maximum value type question is described as follows by formula (8):

[0085] Max_constraint(Entity,”{property_constraint}”,Maxvalue)

[0086] :-entity(Entity,{property_constraint},value),Value

[0087] =#max{{W:entity(_,”{property_constraint}”,W)}},

[0088] MaxValue = Value; Formula (8)

[0089] The ASP triple includes: a subject, a predicate, and an object. In the embodiment of the present application, it is a structured representation method, which is usually used to describe the relationship between the subject, the predicate, and the object. Specifically, the general form of the triple is:

[0090] triple(Subject,Predicate,Object).

[0091] Subject: describes an entity or subject, such as an attribute, object name, or concept.

[0092] Predicate: describes the relationship or attribute between entities, such as is, has, max, min, etc.

[0093] Object: describes the specific attribute value of an entity or a target entity.

[0094] Application of ASP triples:

[0095] Typical scenarios for using triples in ASP include:

[0096] 1. Knowledge representation and reasoning: describing and reasoning about entities and their relationships.

[0097] 2. Attribute calculation and optimization: express the constraints and results such as maximum and minimum values ​​of entity attributes.

[0098] 3. Data modeling: Convert facts and rules into a unified structured representation to facilitate the processing of complex logic.

[0099] Definition and example of ASP triples:

[0100] 1. Triple representation of facts

[0101] Suppose we want to describe some entities and attributes, we can use the following triples to represent them:

[0102] %Entity and attribute relationship

[0103] triple("a","has_price",10).

[0104] triple("b","has_price",15).

[0105] triple("c","has_price",7).

[0106] triple("d","has_price",20).

[0107] 2. Triple Representation of Maximum Problem

[0108] Calculate the maximum value of an attribute through ASP rules and output it in the form of triples:

[0109] % Define the maximum value rule

[0110] max_value(V):-triple(_,"has_price",V),not less_value(V).

[0111] less_value(V):-triple(_,"has_price",V),triple(_,"has_price",V2),V <V2.

[0112] % Use triples to represent the maximum value

[0113] triple(Entity,"max_price",MaxValue):-triple(Entity,"has_price",MaxValue),max_value(MaxValue).

[0114] Input data:

[0115] triple("a","has_price",10).

[0116] triple("b","has_price",15).

[0117] triple("c","has_price",7).

[0118] triple("d","has_price",20).

[0119] Output:

[0120] triple("d","max_price",20).

[0121] For enumeration questions, that is, entity types that meet specific constraints, the method for generating questions and answers based on knowledge graphs provided in the embodiment of the present application quickly obtains answers by directly applying formula (7).

[0122] When dealing with intersection problems with multiple constraints, that is, entities that need to satisfy two constraints at the same time, the method for generating questions and answers based on knowledge graphs provided in the embodiment of the present application adds an identical constraint condition on the basis of formula (7) to satisfy the dual condition requirements, and the result is directly calculated by ASP;

[0123] For the difference set problem, the entity needs to meet the first constraint and exclude the second constraint. The ASP rule is shown in formula (9) (i.e., the third ASP rule in the embodiment of the present application), which uses the not keyword to exclude entities that do not meet the second constraint;

[0124] Constraint(Entity):-entity(Entity,_,”{entity_constraints[0]}”),

[0125] Not entity(Entity,_,”{entity_constraints[1]}”). Formula (9)

[0126] For single-entity problems and multi-entity problems under finite entity constraints, including selection, judgment, comparison, and basic arithmetic problems, these problems usually involve the retrieval and analysis of finite attributes / entities. Through ASP refinement of relevant knowledge, the large language model can achieve better answering effects on these problems; this article uses the prompt word template shown in Table 1 to generate answers, where {context_str} and {query_str} are dynamic fill-in items, which are used to fill in background knowledge and specific problems respectively. In particular, when dealing with calculation problems, it is found that the large language model may tend to generate prediction results based on text patterns, resulting in direct calculations that are often inaccurate. For this reason, in the embodiment of the present application, by adding a prompt of "Please write a program to calculate" before the question, the large language model is guided to generate Python code for calculation to improve the accuracy of the calculation results.

[0127] Table 1 Simple question query template

[0128]

[0129] For example, you need to answer the following questions based on the contextual information I provide rather than prior knowledge, and clearly state the source of the knowledge for your answer in the explanation.

[0130] If relevant knowledge is not included in the context, indicate it in your answer rather than making it up.

[0131] The context information currently provided is as follows. Each tuple represents a fact: {context_str}, and the question is: {query_str}. Format: Please submit your answer in JSON format with answer and explanation as the key. Please strictly follow the format I specified.

[0132] For single event type questions, the method for generating questions and answers based on knowledge graphs provided in the embodiment of the present application uses multiple ASP rules to filter and refine background knowledge. When the alignment process fails, the constraints are relaxed to obtain potential background knowledge, "to ensure that the refined data maintains a high accuracy. The method for generating questions and answers based on knowledge graphs provided in the embodiment of the present application defines an ASP rule as shown in formula (10) (i.e., the fourth ASP rule in the embodiment of the present application) for representing hard constraints. Three fillable fields are defined in the rule, which can be entity, entity type, date, starting point, end point and task type. By defining soft constraints presented in the form of formula (11) (i.e., the fifth ASP rule in the embodiment of the present application), it is applicable to date, starting point and end point, and the penalty values ​​of the above soft constraints are all set to 1. The final output result is determined by screening the event with the lowest total penalty value.

[0133] Constraint(EntityType,Entity,TaskType,Date,Departure,Destination,ID)

[0134] EntityType={entity_type_constraint},

[0135] TaskType={task_constraint}.Formula (10)

[0136] Date_soft_constraint(ID,1):-Formula(11)

[0137] constraint(_,_,_,Date,_,_,ID),

[0138] Date!={date_constraint}.

[0139] Step S110, organizing the fact set through the large language model to generate the answer to the corresponding question.

[0140] Specifically, Figure 2 As shown, for the problem solving part, the large language model is used to generate the corresponding query answer by combining the question content, prompts and knowledge refined using ASP rules.

[0141] In summary, combined Figure 2 As shown, Figure 3 is a structural schematic diagram of question and answer generation in a method for generating question and answer based on a knowledge graph according to an embodiment of the present invention; Figure 3As shown, the process of generating questions and answers based on the knowledge graph includes: question classification, entity recognition, entity linking, knowledge retrieval, knowledge simplification and problem solving, corresponding to steps S100 to S110.

[0142] Combination Figure 2 , Figure 4 is a flowchart of data processing in a method for generating questions and answers based on a knowledge graph according to an embodiment of the present invention; Figure 4 As shown, corresponding to Figure 1 Step S100 to step S110.

[0143] The present invention adopts the above technical scheme, and determines the classification of the question by analyzing the received question through a text classification model; performs entity recognition through an entity recognition model based on the question to determine the candidate entity in the question; calculates the similarity between the candidate entity and the entity in the knowledge graph to obtain the target entity corresponding to the candidate entity in the knowledge graph; determines the retrieval strategy based on the classification of the question, and searches the target entity through the knowledge base based on the retrieval strategy to generate an entity set; determines the corresponding ASP rule based on the classification of the question, and processes the entity set according to the ASP rule to obtain a fact set; organizes the fact set through a large language model to generate the answer to the corresponding question. Compared with the prior art, it has the following technical effects: it can effectively perform retrieval to improve the reasoning performance of the model and generate more accurate answers.

[0144] Example 2

[0145] An exemplary embodiment of the present invention is as follows Figure 5 As shown, Figure 5 is a schematic diagram of a device for generating questions and answers based on a knowledge graph according to an embodiment of the present invention. The device for generating questions and answers based on a knowledge graph provided by an embodiment of the present application includes:

[0146] The question classification module 51 is used to analyze the received question through the text classification model to determine the classification of the question; the entity recognition module 52 is used to perform entity recognition through the entity recognition model based on the question to determine the candidate entity in the question; the entity linking module 53 is used to calculate the similarity between the candidate entity and the entity in the knowledge graph to obtain the target entity corresponding to the candidate entity in the knowledge graph; the knowledge retrieval module 54 is used to determine the retrieval strategy based on the classification of the question, and retrieve the target entity through the knowledge base based on the retrieval strategy to generate an entity set; the knowledge refinement module 55 is used to determine the corresponding ASP rule based on the classification of the question, and process the entity set according to the ASP rule to obtain a fact set; the problem solving module 56 is used to organize the fact set through a large language model to generate an answer to the corresponding question.

[0147] Optionally, the question classification module 51 includes: a word segmentation unit, which is used to segment the question when the text classification model is a K-BERT model trained with a knowledge graph, and obtain a text sequence after word segmentation; a question classification unit, which is used to classify the text sequence after word segmentation through the K-BERT model to obtain the classification of the question.

[0148] Further, optionally, the entity recognition module 52 includes: a semantic recognition unit, used to recognize semantic information in the question through the K-BERT model; an information processing unit, used to process the semantic information through a bidirectional long short-term memory network to obtain a feature representation sequence; an entity recognition unit, used to process the feature representation sequence according to a conditional random field layer to obtain a candidate entity.

[0149] Optionally, the entity linking module 53 includes: a similarity calculation unit, which is used to perform vector similarity calculation on the candidate entities and attributes with the entities and attributes in the knowledge graph to obtain the vector similarity of the entities and attributes corresponding to the candidate entities and attributes in the knowledge graph; an entity linking unit, which is used to determine the entities and attributes corresponding to the candidate entities and attributes in the knowledge graph based on the vector similarity to obtain the target entity and the attributes of the target entity.

[0150] The present invention adopts the above technical scheme, and determines the classification of the question by analyzing the received question through a text classification model; performs entity recognition through an entity recognition model based on the question to determine the candidate entity in the question; calculates the similarity between the candidate entity and the entity in the knowledge graph to obtain the target entity corresponding to the candidate entity in the knowledge graph; determines the retrieval strategy based on the classification of the question, and searches the target entity through the knowledge base based on the retrieval strategy to generate an entity set; determines the corresponding ASP rule based on the classification of the question, and processes the entity set according to the ASP rule to obtain a fact set; organizes the fact set through a large language model to generate the answer to the corresponding question. Compared with the prior art, it has the following technical effects: it can effectively perform retrieval to improve the reasoning performance of the model and generate more accurate answers.

[0151] The above description is only a preferred embodiment of the present invention, and does not limit the implementation mode and protection scope of the present invention. For those skilled in the art, it should be aware that all solutions obtained by equivalent substitutions and obvious changes made using the description and illustrations of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for generating questions and answers based on knowledge graph, characterized in that: include: Analyze the received questions through a text classification model to determine the classification of the questions; Performing entity recognition based on the question using an entity recognition model to determine candidate entities in the question; Calculate the similarity between the candidate entity and the entity in the knowledge graph to obtain the target entity corresponding to the candidate entity in the knowledge graph; Determine a search strategy according to the classification of the problem, and search the target entity through a knowledge base according to the search strategy to generate an entity set; Determine a corresponding ASP rule according to the classification of the problem, and process the entity set according to the ASP rule to obtain a fact set; The fact set is collated through a large language model to generate an answer corresponding to the question.

2. The method for generating questions and answers based on knowledge graph according to claim 1, characterized in that: Analyzing the received question through a text classification model to determine the classification of the question includes: When the text classification model is a K-BERT model trained by the knowledge graph, segmenting the question to obtain a text sequence after segmentation; The text sequence after word segmentation is classified by the K-BERT model to obtain the classification of the problem.

3. The method for generating questions and answers based on knowledge graph according to claim 2, characterized in that: The performing entity recognition by using an entity recognition model according to the question to determine the candidate entities in the question includes: Identify semantic information in the question through the K-BERT model; Processing the semantic information through a bidirectional long short-term memory network to obtain a feature representation sequence; The feature representation sequence is processed according to a conditional random field layer to obtain the candidate entity.

4. The method for generating questions and answers based on knowledge graph according to claim 1, characterized in that: The similarity calculation between the candidate entity and the entity in the knowledge graph to obtain the target entity corresponding to the candidate entity in the knowledge graph includes: Calculate the vector similarity between the candidate entity and attribute and the entity and attribute in the knowledge graph to obtain the vector similarity between the candidate entity and attribute and the entity and attribute corresponding to the candidate entity and attribute in the knowledge graph; The entities and attributes corresponding to the candidate entities and attributes in the knowledge graph are determined based on the vector similarity to obtain the target entity and the attributes of the target entity.

5. The method for generating questions and answers based on knowledge graph according to claim 4, characterized in that: Determining a search strategy based on the classification of the problem, and searching the target entity through a knowledge base based on the search strategy to generate an entity set includes: Determine a retrieval strategy according to the classification of the problem, and generate a corresponding knowledge set for each entity tuple in the target entity according to the retrieval strategy; Calculating the vector cosine similarity between the attribute of each identified entity in the knowledge set and the target entity to determine the score of the attribute; At least two attributes are selected from the knowledge set according to the scores of the attributes to generate the entity set.

6. The method for generating questions and answers based on knowledge graph according to claim 1, characterized in that: The corresponding ASP rules are determined according to the classification of the problem, and the entity set is processed according to the ASP rules to obtain a fact set including: In the case where the problem is classified as a quantitative statistics problem, the first ASP rule is used to locate all entities that meet the specified constraints and list their names, and all entities that meet the specified constraints and list their names are determined as the fact set; In the case where the problem is classified as an extreme value problem, the extreme values ​​of all attributes are calculated using the second ASP rule, and the extreme values ​​are represented by ASP triples, and the extreme values ​​represented by the ASP triples are determined as the fact set; In the case where the problem is classified as an enumeration problem, the first ASP rule is used to obtain entity types that meet specific constraints, and the entity types that meet the specific constraints are determined as the fact set; In the case where the problem is classified as an intersection problem containing multiple constraints, the same constraints are added on the basis of the first ASP rule, and entities satisfying the multiple constraints are determined as the fact set; In the case where the problem is classified as a difference set problem, a third ASP rule is used to exclude entities that meet the first constraint condition but do not meet the second constraint condition by using keywords, and entities that meet the first constraint condition but do not meet the second constraint condition are determined as the fact set; In the case where the problem is classified as a single entity problem and a multi-entity problem under finite entity constraints, ASP refinement is performed based on the single entity and multi-entity corresponding knowledge under finite entity constraints to obtain the fact set; When the problem is classified as a single event problem, the fillable field is set as a hard constraint through the fourth ASP rule, and the soft constraint is set through the fifth ASP rule. The entities obtained by screening according to the hard constraint and the soft constraint are determined as the fact set.

7. A device for generating questions and answers based on a knowledge graph, characterized in that: include: A question classification module is used to analyze the received questions through a text classification model to determine the classification of the questions; An entity recognition module, used to perform entity recognition based on the question through an entity recognition model to determine candidate entities in the question; An entity linking module is used to calculate the similarity between the candidate entity and the entity in the knowledge graph to obtain the target entity corresponding to the candidate entity in the knowledge graph; A knowledge retrieval module, used to determine a retrieval strategy according to the classification of the problem, and to search the target entity through a knowledge base according to the retrieval strategy to generate an entity set; A knowledge refining module, used to determine the corresponding ASP rules according to the classification of the problem, and process the entity set according to the ASP rules to obtain a fact set; The problem-solving module is used to organize the fact set through a large language model to generate an answer corresponding to the question.

8. The device for generating questions and answers based on knowledge graph according to claim 7, characterized in that: The problem classification module includes: A word segmentation unit, used for, when the text classification model is a K-BERT model trained by the knowledge graph, segmenting the question to obtain a text sequence after word segmentation; A question classification unit is used to classify the text sequence after word segmentation through the K-BERT model to obtain the classification of the question.

9. The device for generating questions and answers based on knowledge graph according to claim 8, characterized in that: The entity recognition module includes: A semantic recognition unit, used to recognize semantic information in the question through the K-BERT model; An information processing unit, used for processing the semantic information through a bidirectional long short-term memory network to obtain a feature representation sequence; The entity recognition unit is used to process the feature representation sequence according to the conditional random field layer to obtain the candidate entity.

10. The device for generating questions and answers based on knowledge graph according to claim 7, characterized in that: The entity linking module includes: A similarity calculation unit, used to perform vector similarity calculation on the candidate entities and attributes and the entities and attributes in the knowledge graph to obtain vector similarities of the entities and attributes corresponding to the candidate entities and attributes in the knowledge graph; An entity linking unit is used to determine the entities and attributes corresponding to the candidate entities and attributes in the knowledge graph based on the vector similarity, and obtain the target entity and the attributes of the target entity.

Citation Information

Cited By

  • Knowledge graph retrieval reasoning method and system based on big language model enhancement

    CN120123487A