Intelligent question and answer method, device and equipment

By using a question triplet in an intelligent question-answer system to find the original description from the graph database to generate answers, the answers summarized and fabricated questions in traditional systems are solved, and the reliability of the system and the accuracy of the answers are improved.

CN120011495APending Publication Date: 2025-05-16BEIJING AEROSPACE TITAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411900220.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Due to technical complexity, traditional intelligent question-and-answer systems have information that is not related to the question or fabricated when searching for answers, which reduces the reliability of the system, especially in business areas that require high accuracy and reliability of the results.

Method used

By obtaining the problem description, the problem triplet is extracted and the corresponding text information is found from the pre-constructed graph database based on the problem triplet, and the question answer is generated based on the text information. This method locates the answers to the question through the question triple to the original description of the knowledge document, avoiding the summary and fabrication of the answers.

Benefits of technology

It effectively improves the reliability of the intelligent question-and-answer system, ensures the accuracy and relevance of answers, and is suitable for business areas with high requirements for the reliability of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011495A_ABST
    Figure CN120011495A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent question and answer method, device and equipment, and the method comprises the steps: carrying out the information extraction of question description, and obtaining a question triple set; searching a corresponding text answer from a pre-constructed graph database based on the question triple set; summarizing the text answers to generate question answers. According to the method, the original description (namely, the text answer) of the question answer in the knowledge document is positioned through the question triple during intelligent question answering, and the original description is summarized to form the question answer, so that the problem that the retrieved answer is summarized in each step of the middle level due to the complex overall technology of GraphRAG is solved; therefore, the reliability of the intelligent question answering is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and in particular to an intelligent question-answering method, apparatus and device. Background Art

[0002] With the rapid development of artificial intelligence technology, intelligent question-answering systems have become an important application in the fields of information retrieval and natural language processing. Traditional intelligent question-answering systems are mainly based on Microsoft's GraphRAG (search augmented generation) technology, which uses the relationship information between entities to retrieve and query related graph elements through a graph database, thereby achieving more accurate retrieval and more comprehensive contextual understanding. However, due to the complexity of its overall technology, GraphRAG will summarize the retrieved answers at each step in the intermediate level. In actual applications, there will be information that is irrelevant to the question or fabricated in the answers to the questions, which leads to poor reliability in intelligent question-answering in business fields with high reliability requirements. Therefore, how to improve the reliability of intelligent question-answering in business fields with high requirements for result accuracy and reliability has become a problem that needs to be solved urgently by technicians in this field. Summary of the invention

[0003] In view of this, the present application proposes an intelligent question-answering method, apparatus and device, which can effectively improve the reliability of intelligent question-answering.

[0004] According to a first aspect of the present application, a method for intelligent question answering is provided, comprising:

[0005] Get a description of the problem;

[0006] Extract information from the problem description to obtain a problem triple;

[0007] Based on the question triples, corresponding text information is searched in a pre-built graph database;

[0008] Generate a corresponding answer to the question according to the text information.

[0009] In a possible implementation, the graph database contains more than two knowledge graphs, and different knowledge graphs correspond to different technical fields.

[0010] In a possible implementation, the construction of the graph database includes:

[0011] Acquire knowledge documents;

[0012] Segmenting the knowledge document to obtain segmented text information;

[0013] Extracting information from each text information to obtain each text triple;

[0014] The text information and the text triples are stored in a database to obtain the graph database.

[0015] In a possible implementation, when the texts and the text triples are stored in a database to obtain the graph database, the operation of expanding the text triples to obtain expanded text triples is also included, and then the operation of storing the texts and the text triples in a database to obtain the graph database is performed.

[0016] In a possible implementation, when searching for corresponding text information in a pre-constructed graph database, the corresponding text information is searched by matching the question triple set with the text triples.

[0017] In a possible implementation, when matching the question triple set with the text triple to search for the corresponding text information, the question triple set includes:

[0018] Obtaining entities in the question triplet;

[0019] Find a set of entities similar to the entity;

[0020] Searching for a corresponding relationship set according to the similar entity set;

[0021] The entity set and the relationship set are combined to obtain the question triple set.

[0022] In a possible implementation, when generating the corresponding answer to the question according to the text information, it also includes determining the relevance of the text information to the question, and then executing the operation of generating the corresponding answer to the question according to the text information.

[0023] In a possible implementation, when the searched text information is subjected to the question relevance judgment, it includes:

[0024] When it is determined that the text information is related to the problem, the text information matches the problem description, and the current text information is retained;

[0025] When it is determined that the text information is not relevant to the problem, the text information does not match the problem description, and the current text information is discarded.

[0026] According to a second aspect of the present application, there is provided an intelligent question-answering device, comprising:

[0027] A problem description acquisition module is used to acquire a problem description;

[0028] An information extraction module, used to extract information from the problem description to obtain a problem triple;

[0029] A text information search module, used to search for corresponding text information in a pre-built graph database based on the question triple;

[0030] The question answer generation module is used to generate corresponding question answers based on the text information.

[0031] According to the third aspect of the present application, an intelligent question-and-answer device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method described in the first aspect of the present application.

[0032] In this application, a smart question-answering method includes: obtaining a description of a question for smart question-answering; extracting information from the question description to obtain a set of question triples; searching for corresponding text answers in a pre-built graph database based on the set of question triples; summarizing the text answers to generate answers to the questions. When performing smart question-answering, this application locates the original description of the question answer in the knowledge document (i.e., the text answer) through the question triples, and summarizes the original description to form an answer to the question. This solves the problem that GraphRAG, due to its overall technical complexity, summarizes the retrieved answers at each step of the intermediate level, resulting in technical problems in the answers that are irrelevant to the question or fabricated information, and effectively improves the reliability of smart question-answering.

[0033] Other features and aspects of the present application will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the present application and, together with the description, serve to explain the principles of the present application.

[0035] Figure 1 A flow chart showing an intelligent question-answering method according to an embodiment of the present application is shown;

[0036] Figure 2 A schematic block diagram of an intelligent question-answering device according to an embodiment of the present application is shown;

[0037] Figure 3 A schematic block diagram of an intelligent question-answering device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0038] Various exemplary embodiments, features and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0039] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0040] In addition, in order to better illustrate the present application, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present application can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present application.

[0041] <Method Example>

[0042] Figure 1 An example flow chart of an intelligent question-answering method according to an embodiment of the present application is shown. Figure 1 As shown, the method includes steps S1100-S1400: S1100, obtaining a question description; S1200, extracting information from the question description to obtain a question triple; S1300, searching for corresponding text information in a pre-built graph database based on the question triple; S1400, generating a corresponding answer to the question based on the text information. The present application extracts question triples, searches for the original description corresponding to the question based on the question triples, and then generates the final answer to the question based on the original description. The intelligent question-answering method adopted in the present application locates the answer to the question in the original description through the question triples, and then generates the final answer to the question based on the original description, which solves the technical problem that GraphRAG, due to its overall technical complexity, summarizes the retrieved answers at each step of the intermediate level, resulting in the appearance of irrelevant information or fabricated information in the answers, and effectively improves the reliability of intelligent question-answering.

[0043] Among them, when the problem description is extracted to obtain the problem triple, the problem triple corresponding to the problem description can be extracted by the large language model. Specifically, when the large language model extracts information from the problem description, the problem description is first embedded in the preset prompt word template to obtain the complete prompt word, and then the large language model extracts information according to the complete prompt word to obtain the problem triple. It should be noted that the large language model can be selected according to the actual situation, preferably the Qwen2_5-32B-Instruct (the 32 billion parameter instruction fine-tuning model of Tongyi Qianwen 2.5) model. Among them, it is common knowledge in the art that the large language model performs corresponding operations according to the prompt word, and no further elaboration is made here. It should be pointed out here that the problem triple is a set for representing entities and relationships, including subject entities, predicate relationships and object entities.

[0044] It should be further explained that the preset prompt word template includes a triple extraction requirement, and the specific prompt words can be set according to actual conditions.

[0045] For example, when extracting triples of question descriptions, the prompt words of the large language model can be set as:

[0046] #Please play the role of an information extraction expert, conduct an in-depth analysis of the given text content, and extract all the information that can form a 'subject-predicate-object' structure. Among them, the subject and object can be people, things, objects, place names, and time, and the predicate is a prepositional phrase that includes a verb or has an action meaning.

[0047] """

[0048] {Question}

[0049] """

[0050] Based on the obtained question triples, the corresponding text information is searched in the pre-built database. Before searching for the corresponding text information from the question triples, it is necessary to first build a graph database, and then search for the corresponding text information based on the pre-built graph database. It should be noted that the graph database contains more than two knowledge graphs, and different knowledge graphs correspond to different technical fields. The corresponding graph database is obtained based on each knowledge graph.

[0051] The following uses the process of building a knowledge graph in a database to illustrate the construction process of each knowledge graph. In one possible implementation, when building a knowledge graph in a graph database, first obtain the knowledge document, then segment the knowledge document to obtain each text, then extract information from each text to obtain the corresponding text triple, build a knowledge graph based on each text triple, and finally store the knowledge graph and the corresponding text into the database to obtain the final knowledge graph. It should be pointed out here that the text triple is a set of entities and relationships corresponding to each text, including subject entities, predicate relationships, and object entities.

[0052] When knowledge documents are segmented to obtain texts, each document is first standardized, that is, each document format is converted into a unified preset format. Then, knowledge documents based on each standard format are segmented. When knowledge documents in a standard format are segmented, each paragraph text is first obtained according to paragraph segmentation, and then each paragraph text is divided according to a preset number of lengths to obtain each text information.

[0053] In one possible implementation, before dividing each paragraph text according to a preset number of lengths to obtain each text information, first determine whether each paragraph text needs to be processed as an overlong paragraph. In determining whether each paragraph text needs to be processed as an overlong paragraph, the word unit length corresponding to each paragraph text is compared with a preset word unit length threshold. When the corresponding word unit length exceeds the word unit length threshold, the paragraph text corresponding to the current word unit needs to be processed as an overlong paragraph. Otherwise, the current paragraph text does not need to be processed as an overlong paragraph. It should be noted that the word unit threshold can be set according to actual conditions, and is preferably set to 5000.

[0054] It should be noted that when judging the size of the word-gram length and the word-gram length threshold, the word-gram length corresponding to each paragraph text is first calculated. Specifically, before calculating the word-gram length of each paragraph text, first obtain the word-gram corresponding to each paragraph text. In one possible implementation, when obtaining the word-gram corresponding to each paragraph text, the word-gram corresponding to each paragraph is directly generated according to the selected large model word segmenter. Based on the obtained word-gram, the word-gram length corresponding to each paragraph text can be obtained. It should be pointed out here that the word-gram refers to the basic unit of the paragraph text.

[0055] When it is determined that the word length exceeds the word length threshold, the corresponding paragraph text is processed as an overlong paragraph. In one possible implementation, when the paragraph text is processed as an overlong paragraph, it includes: segmenting the paragraph text to be processed as an overlong paragraph according to sentences, calculating the word length corresponding to each sentence based on the segmented sentences, and splicing each sentence into a new paragraph text according to the word length corresponding to each sentence according to the preset splicing requirements, that is, completing the overlong paragraph processing. Among them, the preset splicing requirement means that the sum of the word lengths corresponding to the spliced ​​sentences does not exceed the word length threshold.

[0056] Based on the new paragraph text obtained after the over-long paragraph processing and the paragraph text that does not need to be processed for over-long paragraphs, the text blocks are divided according to the preset number of lengths to obtain various text information. In a possible implementation method, when the paragraph texts are divided into text blocks according to the preset number of lengths, the number of divisions is pre-set according to the large language model selected above, and the large language model generates text information of each preset number of lengths under the premise of ensuring semantic coherence. Among them, the preset number of lengths can be set according to actual conditions, and is preferably set to 500.

[0057] Furthermore, when each knowledge document is segmented to obtain each text information, each standardized knowledge document can also be segmented according to sentences to obtain corresponding sentences. In this case, the text information includes at least one of a text block and a sentence.

[0058] It should be noted that when each knowledge document is segmented to obtain each text information, at least one of the above-mentioned text block segmentation and sentence segmentation is included, and no specific limitation is made.

[0059] Based on the obtained text information, information extraction is performed to obtain text triples. The method of extracting information from each text information to obtain corresponding text triples is consistent with the method of extracting question triples mentioned above, and will not be described in detail here.

[0060] Based on the obtained text triples, the corresponding knowledge graph is constructed, and the corresponding relationship between the text triples and each text information is constructed. The knowledge graph and the corresponding text information are stored in a one-to-one correspondence to obtain the final knowledge graph.

[0061] Among them, when constructing the corresponding knowledge graph based on each text triple, the entity category in the text triple is first determined. The same node is constructed for the same entity name and the same entity category, and different nodes are constructed for the same entity name but different entity categories. Then, the entities are connected according to the corresponding relationship to complete the knowledge graph.

[0062] In a possible implementation, when determining the entity category in a text triple, the selected large language model determines the category corresponding to the entity in the current triple based on a preset prompt word template and the sentence corresponding to the current triple. When determining the entity category of a triple, the sentence corresponding to the triple is first determined. Specifically, the corresponding sentence of the triple can be determined based on the correspondence between the text information and the triple.

[0063] Based on the determined sentence, it and the corresponding triple entity are embedded in the preset prompt template to obtain a complete prompt, and then the large language model determines the entity category based on the completed prompt. Among them, the preset prompt word template includes the entity category determination requirements, and the specific prompt word template can be set according to the actual situation.

[0064] For example, when the large language model is determining entity categories, the prompt word template can be set as:

[0065] #Please play the role of a category determination expert and judge what category of entities "{(subject entity)sbj x1}" and "{(object entity)obj x1}" are in the context of the following sentences.

[0066] """

[0067] {sentence st x1}

[0068] """

[0069] After obtaining the category corresponding to each entity, each entity is stored in correspondence with its category, and the knowledge graph can be constructed according to the entity category corresponding to the entity. Based on the above-mentioned knowledge graph construction process, the construction of the knowledge graph corresponding to each field is completed in the same way. Based on each knowledge graph obtained, the technical field identifier corresponding to each knowledge graph is constructed, and each knowledge graph and the technical field identifier are stored in the database in correspondence, and finally a pre-constructed graph database is obtained. It should be pointed out here that the database can be selected according to the actual situation, and the Neo4J graph database is preferred.

[0070] Furthermore, in order to enrich the content of the knowledge graph and improve the question-answer retrieval capability, when each text triple is obtained and the corresponding knowledge graph is constructed based on each text triple, it is also possible to expand the obtained text triples to obtain expanded text triples, and then store them in the database to obtain the final graph database. Among them, when expanding each text triple, it includes entity name expansion and coreference relationship construction.

[0071] In one possible implementation, when expanding the text triple entity name, first obtain the entity corresponding to the synonym entity set, expand the entity name based on each synonym entity, and obtain the expanded synonym entity set. It should be noted that synonym entities refer to entities with different names but the same essence.

[0072] When obtaining the synonymous entity corresponding to the entity, firstly, according to the corresponding relationship between the entity and the entity category, the entity is grouped according to the entity category to obtain each entity set. Based on each entity set, the corresponding synonymous entity set is obtained.

[0073] In one possible implementation, when obtaining corresponding synonymous entity sets based on various entity sets, it includes: calculating similar entity sets corresponding to each entity in various entity sets, determining synonymous entities corresponding to the entities based on the similar entity sets, and finally obtaining synonymous entity sets corresponding to the entities.

[0074] The following takes the process of obtaining the corresponding synonymous entity set for a single entity in a type of entity set as an example to illustrate the process of calculating the synonymous entity set corresponding to each entity in each type of entity set.

[0075] Specifically, when calculating the corresponding similar entity set of the same entity set, in order to improve the accuracy of the determination of entities with the same meaning but different names, when calculating the similar entity set, firstly, the corresponding similar entity set is calculated by vector to obtain the first similar entity set; at the same time, the corresponding similar entity set is calculated by the rearrangement model to obtain the second similar entity set; finally, the entities with the same meaning but different names are determined based on the first similar entity set and the second similar entity set.

[0076] In one possible implementation, when calculating the corresponding similar entity set through vectors, the entity vector corresponding to the entity is first calculated, and then the similarity between the current entity vector and each entity vector is calculated. Based on each similarity, similar entities corresponding to the current entity are screened, and the screened similar entities are combined to obtain a similar entity set corresponding to the current entity (i.e., a first similar entity set). When screening the corresponding similar entities based on each similarity, each similarity is compared with a similarity threshold, and the corresponding entity greater than the similarity threshold is the similar entity corresponding to the current entity. It should be noted that the similarity threshold can be set according to actual conditions, and is preferably set to 0.7.

[0077] At the same time, when the corresponding similar entity set is calculated by the reranking model to obtain the second similar entity set, the current entity is used as the model input, and the similarity threshold is set at the same time. The reranking model selects the corresponding similar entities greater than the similarity threshold according to the input entity and the similarity threshold, and combines them to obtain the similar entity set corresponding to the current entity (i.e., the second similar entity set). Among them, the reranking model can be selected according to the actual situation, and the 1024-dimensional Chinese version Rerank reranking model is preferred; at the same time, when setting the similarity threshold, it is determined according to the selected model. When the selected model is the 1024-dimensional Chinese version Rerank reranking model, the corresponding similarity threshold is preferably set to 0.5.

[0078] Based on the obtained first similar entity set and the second similar entity set, the synonymous entity corresponding to the current entity is determined (i.e., the corresponding synonymous entity set is obtained). Specifically, the entities in the first similar entity set and the entities in the second similar entity set are merged and deduplicated to obtain the third similar entity set corresponding to the current entity. At this time, the above-selected large language model is used to obtain the synonymous entity set corresponding to the current entity based on the third similar entity set. Specifically, the current entity and the corresponding third similar entity set are embedded in a pre-constructed prompt word template to obtain a complete prompt word, and the large language model obtains the synonymous entity set corresponding to the current entity based on the complete prompt word. Among them, the pre-constructed prompt word template includes the synonymous entity determination requirements, and the specific prompt word template is not limited.

[0079] For example, when determining entities with the same or different names, the prompt word template can be set as:

[0080] #Please determine which of the following entities "{(current entity)similar_enty xn}" is the same entity as. If there is no identical entity, return None.

[0081] """

[0082] {(third similar entity set) similar_set xn}

[0083] """

[0084] According to the process of obtaining the corresponding synonymous entity set for a single entity in the above-mentioned similar entity set, the synonymous entity set corresponding to each entity is obtained, and based on the obtained synonymous entity set, it is stored in one-to-one correspondence with the corresponding entity, so as to facilitate direct use in subsequent entity expansion.

[0085] Through the above operations, the synonymous entity sets corresponding to each entity in each entity set are obtained.

[0086] Based on the obtained entity sets corresponding to the entities with the same synonyms, corresponding entity names are expanded to obtain an expanded entity set. The following describes the detailed process of entity name expansion by taking an entity name expansion to obtain a corresponding expanded entity set.

[0087] Specifically, based on the synonymous entity set, each entity in the set is input into the big model, and the big model generates the corresponding expanded entity name set corresponding to the current synonymous entity, traverses each synonymous entity to obtain each expanded entity set, and then combines each expanded entity set to obtain the final expanded entity set corresponding to the current entity. When the traversal is completed, the entity name corresponding to the current entity is expanded.

[0088] The expanded entity set and the synonym set corresponding to the current entity are merged and deduplicated, and finally the expanded synonym set corresponding to the current entity is obtained, and it is stored in one-to-one correspondence with the current entity.

[0089] According to the above process of expanding the entity name of a single entity, the entity name expansion corresponding to each entity is completed, and the expanded synonym entity set corresponding to each entity is obtained. That is, the entity name expansion is completed.

[0090] After completing the entity name expansion and obtaining the expanded synonym entity set corresponding to each entity, it is necessary to construct the synonym relationship to complete the expansion of the text triple. Specifically, based on the expanded synonym entity set corresponding to each entity, the synonym relationship between two entities is constructed and stored, that is, the synonym relationship construction operation in the triple expansion is completed.

[0091] Based on the above entity expansion and coreference relationship construction, the text triple expansion is completed to obtain the final graph database.

[0092] According to the pre-constructed graph database, the corresponding text information of the question triple is searched from the graph database. Among them, when searching for the corresponding text information in the pre-constructed graph database, the question is first matched with the corresponding technical field (that is, the knowledge graph corresponding to the search is matched), and then the text information corresponding to the question triple is searched based on the matched knowledge graph. It should be noted here that the problem description contains the corresponding technical field identifier. In one possible implementation, when matching the corresponding knowledge graph according to the problem description, the corresponding knowledge graph is matched by the technical field identifier in the problem description. Specifically, the technical field identifier of the problem is first obtained from the problem description, and the obtained technical field identifier is matched with the technical field identifier in the pre-constructed database. After the technical field identifier is matched, the corresponding knowledge graph is searched according to the correspondence between the technical field identifier and the knowledge graph to obtain a knowledge graph that matches the current problem description.

[0093] Based on the matching knowledge graph, when searching for corresponding text information according to the question triple, the corresponding text information is searched by matching the question triple set with the text triples stored in the matching knowledge graph.

[0094] Before matching the question triple set with the text triple set, the question triple set is first obtained. Specifically, when obtaining the question triple set, the synonymous entity set corresponding to the question entity is first searched, and then the synonym relationship set corresponding to the question predicate relationship is searched, and the question triple set can be obtained by arranging and combining them.

[0095] When searching for the synonymous entity set corresponding to an entity, first obtain the entities in the question triple, including the subject entity and the object entity. Based on the obtained entities, search for the corresponding synonymous entity from the matching knowledge graph to obtain the synonymous entity set corresponding to the question entity.

[0096] Among them, when searching for a set of synonymous entities corresponding to a subject entity, first search for a preset number of similar entities from the matching knowledge graph, and then obtain the corresponding set of synonymous entities based on the similar entities. Specifically, when searching for a preset number of similar entities from the matching knowledge graph, calculate the similarity between the question subject entity and each entity in the graph database. Filter a preset number of question subject entities based on the calculated similarity.

[0097] It should be noted that when calculating the similarity between the subject entity and each entity in the matching knowledge graph, the subject entity is first vectorized to obtain the corresponding subject entity vector. Based on the obtained subject entity vector, it is also necessary to obtain the entity vector corresponding to each entity in the matching knowledge graph. Among them, in order to facilitate the search for the corresponding entity vector, when the entity is stored in the corresponding knowledge graph, a label corresponding to the entity is constructed, recorded as "Entity". At the same time, the vector corresponding to each entity is calculated, and a vector index is constructed, recorded as "VectorIndex 1". The entity vector corresponding to each entity is obtained according to the vector index.

[0098] Then, based on the obtained vectors, the similarity between the question subject vector and the entity vector in the matching knowledge graph is calculated. Based on the calculated similarity, each similarity is compared with the similarity threshold, and a preset number of entity vectors greater than the similarity threshold and arranged from high to low according to the similarity are selected. Then, based on the correspondence between the entity vector and the entity, the corresponding entity is searched from the entities labeled "Entity".

[0099] Based on the preset number of entities found, the synonymous entity of the question subject entity is determined to obtain a synonymous entity set corresponding to the question subject entity. The synonymous entity determination method for the question subject is consistent with the above-mentioned synonymous entity determination method, and will not be described in detail here.

[0100] Similarly, the same operation is performed for the question object entity to obtain the synonym entity set corresponding to the question object entity. It should be noted here that the preset number and similarity threshold can be set according to actual conditions, and the preset number is preferably set to 5, and the similarity threshold is preferably set to 0.5.

[0101] Based on the obtained synonymous and synonymous entity set corresponding to the subject entity, the synonymous and synonymous set corresponding to each entity in the set is searched according to each entity in the set and the pre-built synonymous relationship. Then the synonymous and synonymous sets corresponding to each entity found are merged and deduplicated to obtain the final synonymous and synonymous entity set corresponding to the question subject entity.

[0102] Similarly, the same method as above is used to obtain the final synonymous and cognatous set corresponding to the question object entity.

[0103] After obtaining the final synonymous set corresponding to the entity, based on the predicate relationship in the question triple, the synonymous relationship set corresponding to the predicate relationship of the current question is searched from the matching knowledge graph. Specifically, when searching for the synonymous relationship set corresponding to the predicate relationship of the question, first search for a preset number of similar relationship sets in the knowledge graph, and obtain the synonymous relationship set based on the similar relationship set. Specifically, first calculate the vector of the current predicate relationship, and then need to obtain the corresponding vectors of each relationship stored in the corresponding knowledge graph, wherein, in order to improve the efficiency of the query, when the relationship is stored in the knowledge graph, calculate the relationship vector corresponding to each relationship, and construct the vector index corresponding to the relationship vector, recorded as "VectorIndex 2". Based on the constructed vector index, search for the corresponding relationship vectors in the knowledge graph, calculate the similarity between the current predicate relationship and each relationship vector in the knowledge graph, and filter the preset number of corresponding relationships based on the calculated similarities. Its specific screening process is consistent with the above-mentioned entity screening corresponding similar entities process, and will not be repeated here. It should be pointed out here that the preset number of relationships can be set according to actual conditions, preferably set to 10.

[0104] After obtaining a preset number of similar relationships corresponding to the current question predicate relationship (recorded as the first similar relationship set), in order to improve the accuracy of the search, the synonymous synonymous sets corresponding to the obtained question subject entity and question object entity can be arranged and combined, and all the relationships between the current subject and object in the corresponding knowledge graph are searched based on the combined subject and object (recorded as the second similar relationship set). Then, based on the first similar relationship set, the synonymous relationships of each relationship in the first similar relationship set are searched from the second similar relationship set to obtain the synonymous relationship set corresponding to each relationship in the first similar relationship set, and the synonymous relationship sets corresponding to each relationship in the first similar relationship set are combined in the form of a set to obtain the final synonymous relationship set.

[0105] Based on the final entity corresponding synonym set and synonym relationship set, they are arranged and combined to obtain the question triple set. Then, based on the obtained question triple set, the corresponding stored triples are matched from the matching knowledge graph, and the text information set corresponding to the current question description is found according to the matched triples and the correspondence between the triples and the text information.

[0106] Based on the obtained text information set, the large model generates the final answer to the question according to the preset prompt words. Specifically, each text information in the searched text information set is embedded in the preset prompt word template to obtain a complete prompt word, and the large model generates the final answer to the question according to the complete prompt word. Among them, the preset prompt word template can be set according to the actual situation, and there is no specific limitation.

[0107] For example, when the large model generates the corresponding answer to a question based on the prompt word, the prompt word template can be set as:

[0108] "Please respond according to the following content (question description)}".

[0109] """

[0110] {(text information)Question_Chunk}

[0111] """

[0112] When obtaining the corresponding answer to the question based on the obtained text information, the question relevance judgment must be performed on each obtained text information. When it is judged that the current text information is relevant to the question description, the current text information matches the question description and the current text information is retained; when it is judged that the current text information is not relevant to the question, the text information does not match the question description and the current text information is discarded.

[0113] In a possible implementation, when judging the relevance of text information to a question, the current text information and the question description are embedded in a preset prompt word template to obtain a complete prompt word, and the large model judges the relevance of the text information based on the complete prompt word to obtain a judgment result. The preset prompt word template includes text relevance judgment requirements, and the specific prompt word template can be set according to actual conditions.

[0114] For example, when determining text relevance, the prompt word template can be set as:

[0115] Please judge whether "{text information qck}" can help answer the following questions. If not, answer "no"; if it can, please output the useful original content and remove the useless information.

[0116] """

[0117] {Question}

[0118] """

[0119] After determining the relevance of each piece of text information to the question, the question is answered based on the text information related to the question description, and finally an answer corresponding to the question description is obtained.

[0120] The present application provides an intelligent question-answering method, including obtaining a question description; extracting information from the question description to obtain a question triple; searching for corresponding text information in a pre-built graph database based on the question triple; and generating a corresponding answer to the question based on the text information. The present application extracts question triples, searches for the original description corresponding to the question based on the question triples, and then generates the final answer to the question based on the original description. The intelligent question-answering method adopted in the present application locates the answer to the question in the original description through the question triples, and then generates the final answer to the question based on the original description. This solves the technical problem that GraphRAG, due to its own overall technical complexity, summarizes the retrieved answers at each step in the intermediate level, resulting in the appearance of irrelevant information or fabricated information in the answers, and effectively improves the reliability of intelligent question-answering.

[0121] <Device Example>

[0122] Figure 2 FIG. 1 is a schematic block diagram of an intelligent question-answering device according to an embodiment of the present application. Figure 2 As shown, the device 100 includes: a question description acquisition module 110, an information extraction module 120, a text information search module 130, and a question answer generation module 140. The question description acquisition module 110 is used to obtain a question description; the information extraction module 120 is used to extract information from the question description to obtain a question triple; the text information search module 130 is used to search for corresponding text information in a pre-built graph database based on the question triple; and the question answer generation module 140 is used to generate a corresponding question answer based on the text information.

[0123] <Equipment Embodiment>

[0124] Figure 3 FIG. 1 is a schematic block diagram of an intelligent question-answering device according to an embodiment of the present application. Figure 3 As shown, the intelligent question-answering device 200 includes: a processor 210 and a memory 220 for storing executable instructions of the processor 210. The processor 210 is configured to implement any of the above intelligent question-answering methods when executing the executable instructions.

[0125] Here, it should be noted that the number of processors 210 can be one or more. At the same time, in the intelligent question-answering device 200 of the embodiment of the present application, an input device 230 and an output device 240 may also be included. Among them, the processor 210, the memory 220, the input device 230 and the output device 240 may be connected through a bus or in other ways, which are not specifically limited here.

[0126] The memory 220 is a computer-readable storage medium that can be used to store software programs, computer executable programs, and various modules, such as the program or module corresponding to the intelligent question-answering method of the embodiment of the present application. The processor 210 executes various functional applications and data processing of the intelligent question-answering device 200 by running the software programs or modules stored in the memory 220.

[0127] The input device 230 may be used to receive input numbers or signals. The signals may be key signals related to user settings and function control of the device / terminal / server. The output device 240 may include display devices such as display screens.

[0128] The embodiments of the present application have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for intelligent question answering, characterized in that: include: Get a description of the problem; Extract information from the problem description to obtain a problem triple; Based on the question triples, search for corresponding text information in a pre-built graph database; Generate a corresponding answer to the question according to the text information.

2. The method according to claim 1, characterized in that: The graph database contains more than two knowledge graphs, and different knowledge graphs correspond to different technical fields.

3. The method according to claim 1, characterized in that The construction of the graph database includes: Acquire knowledge documents; Segmenting the knowledge document to obtain segmented text information; Extracting information from each text information to obtain each text triple; The text information and the text triples are stored in a database to obtain the graph database.

4. The method according to claim 3, characterized in that When storing the texts and the text triples into a database to obtain the graph database, the process also includes expanding the text triples to obtain expanded text triples, and then storing the texts and the text triples into a database to obtain the graph database.

5. The method according to claim 4, characterized in that When searching for corresponding text information in a pre-constructed graph database, the corresponding text information is searched by matching the question triple set with the text triples.

6. The method according to claim 5, characterized in that When matching the question triple set with the text triple to search for the corresponding text information, the question triple set includes: Obtaining entities in the question triplet; Find a set of entities similar to the entity; Searching for a corresponding relationship set according to the similar entity set; The entity set and the relationship set are combined to obtain the question triple set.

7. The method according to any one of claims 1 to 6, characterized in that: When the corresponding answer to the question is generated according to the text information, it also includes judging the relevance of the text information to the question, and then executing the operation of generating the corresponding answer to the question according to the text information.

8. The method according to claim 7, characterized in that When the searched text information is subjected to the question relevance judgment, it includes: When it is determined that the text information is related to the problem, the text information matches the problem description, and the current text information is retained; When it is determined that the text information is not relevant to the problem, the text information does not match the problem description, and the current text information is discarded.

9. An intelligent question-answering device, characterized in that: include: A problem description acquisition module is used to acquire a problem description; An information extraction module, used to extract information from the problem description to obtain a problem triple; A text information search module, used to search for corresponding text information in a pre-built graph database based on the question triple; The question answer generation module is used to generate corresponding question answers based on the text information.

10. An intelligent question-answering device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 8 when executing the executable instructions.