Multi-source data intelligent question-answering system based on large language model and application method
By designing a multi-source data intelligent question and answer system, using a large language model to obtain and integrate information from different data sources, the problem of high computing cost of large language models and difficult design of model fusion strategy is solved, and efficient and accurate intelligent question and answer effects are achieved.
Patent Information
- Application Number
- CN202510140240.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to reduce the computing cost of large language models while ensuring performance. In scenarios where knowledge-intensive or real-time information demands, how to design effective model fusion strategies and intelligent triggering mechanisms to improve the overall performance and anti-interference capability of the system.
Design a multi-source data intelligent question and answer system based on large language model, including knowledge graph question and answer module, external interface, document question and answer module, and answer comprehensive generation module. The system can obtain information from different data sources, conduct semantic understanding, inference and answer generation through large models, and integrate and verify answers through answer comprehensive generation modules.
It realizes the efficient application of large language models in a multi-source data environment, can accurately understand user intentions, improve the breadth and accuracy of intelligent question-and-answer, and reduces the waste of computing resources, improves the overall performance and anti-interference ability of the system.
Smart Images

Figure CN120086331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to language model and information retrieval technologies, and particularly to a multi-source data intelligent question-answering system based on a large language model and an application method thereof. Background Art
[0002] In the wave of the intelligent era, the innovation and integration of technologies have continuously promoted the progress of human society. Among them, the combination of large language models and information retrieval technologies has undoubtedly become an important part of this wave. Currently, by integrating large language models and information retrieval technologies and giving full play to their powerful potential in improving information retrieval effects, the future development prospects of retrieval-enhanced large language models are very broad.
[0003] As a major breakthrough in the field of natural language processing, large language models have injected new vitality into information retrieval technologies with their powerful language understanding, reasoning, and generation capabilities. In traditional information retrieval technologies, keyword matching and semantic similarity calculation are the cores, but this is often limited to surface text matching and difficult to deeply understand the query intent and the complexity of the context. The introduction of large language models can not only understand the deep semantics of the query but also reason based on the context, thus providing more accurate retrieval results. For example, in the face of fuzzy or complex queries, large language models can convert the query into a more specific and accurate expression through semantic understanding, improving the efficiency and quality of retrieval.
[0004] The high efficiency of information retrieval technologies and the intelligence of large language models form a perfect complement. Information retrieval technologies can quickly locate relevant documents from a vast amount of information and provide rich external knowledge for large language models. The introduction of this knowledge can not only make up for the deficiencies of large language models in real-time information and domain knowledge but also enhance their performance in knowledge-intensive tasks. For example, when answering professional questions, large language models can provide the latest research results through retrieved up-to-date literature, thus improving the accuracy and timeliness of the question-answering.
[0005] In an information retrieval system, large language models can play a dual role of retrieval and re-ranking. As a retrieval base model, large language models can screen out the most relevant results from a vast amount of information by deeply understanding the query intent. At the same time, through re-ranking, ensure that users obtain the most satisfactory results. In addition, large language models can also provide training materials for information retrieval tasks by generating high-quality labeled data, further improving the performance of the model.
[0006] In some scenarios, large language models can meet task requirements with their internal knowledge without external retrieval. However, in knowledge-intensive or real-time information-demanding scenarios, actively triggered retrieval mechanisms are particularly important. How to design an intelligent trigger mechanism that can both ensure the necessity of retrieval and avoid unnecessary waste of computing resources is a hot topic in current research. For example, based on the complexity of the query and the scarcity of domain knowledge, it is possible to intelligently determine whether to initiate a retrieval, thereby achieving a balance between efficiency and performance.
[0007] Although the integration of large language models and information retrieval technology has shown broad application prospects, it also faces a series of technical challenges. First, large language models require high computing power. How to reduce computing costs while ensuring performance is an urgent problem to be solved. Secondly, how to design an effective model fusion strategy to achieve efficient collaboration between large language models and small retrieval models is the key to improving the overall performance of the system. In addition, in the face of noise information in the retrieval results, how to enhance the anti-interference ability of large language models is also a research focus. Summary of the invention
[0008] The purpose of the present invention is to provide a multi-source data intelligent question-answering system and application method based on a large language model, which can accurately understand the user's intention, obtain information from different data engines and then comprehensively organize it to form the answer required by the user, and can cite the evidence on which the answer depends to form a reasoning chain, wherein the data engine includes unstructured text, structured API interface, knowledge graph and other data sources with representative characteristics.
[0009] The technical solution to achieve the purpose of the present invention is: a multi-source data intelligent question-answering system based on a large language model, including a knowledge graph question-answering module based on the large model, an external interface and document question-answering module, and an answer comprehensive generation module;
[0010] The knowledge graph question-answering module is used to analyze the compiled knowledge graph and locate nodes and extract relevant information, including context completion of questions, identification of core entities, entity linking, and extraction of entity-related information;
[0011] The external interface and document question-and-answer module is used for unified access and question-and-answer of interface data sources and unstructured data sources. Based on the existing interface API application list, it can fully understand the task requirements through the big model, select the appropriate interface from the application list, and extract the parameter values required by the interface for calling; it can perform semantic similarity search based on the vector library constructed based on unstructured documents, intercept similar document fragments and generate answers through the big model;
[0012] The comprehensive answer generation module is used to comprehensively process answers from different sources. First, the questions are classified. If the question belongs to the judgment type, it will be displayed as yes / no. If the question belongs to the summation type, the returned results will be counted. Then, in the case where the answers from various sources are lengthy or conflicting, the answers are separated from the false based on their respective basis, and the answers are generated by comprehensive summarization and induction. Finally, the sensitivity test of the answer is used to confirm whether the final answer will be output.
[0013] An application method of a multi-source data intelligent question answering system based on a large language model provides a unified registration and publishing paradigm for multi-source data, and uses a large model to uniformly complete question sentences and classify questions based on context-based user question semantics. It is divided into query knowledge graphs, call service interface APIs, and document content understanding according to type. The results of the call are summarized and refined through a large language model, and ultimately the answer is automatically generated.
[0014] Compared with the existing technology, the significant advantages of the present invention are: the present invention has the ability to understand question semantics, decompose and call complex tasks. The advantage is that while being able to efficiently access the application system, the large model can accurately understand the user's intentions, drive various applications to obtain answers to questions, and can be integrated into the answer style that is most convenient for users to understand, effectively improving the breadth and accuracy of intelligent question and answering. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is the schematic diagram of the multi-source data intelligent question answering system based on a large language model. DETAILED DESCRIPTION
[0016] like Figure 1 As shown, the present invention proposes a multi-source data intelligent question-answering system based on a large language model, including a knowledge graph question-answering module based on a large model, an external interface and document question-answering module, and an answer comprehensive generation module;
[0017] The knowledge graph question-answering module is used to analyze the compiled knowledge graph and locate nodes and extract relevant information. Its main core functions include context completion, problem core entity recognition, entity linking, and entity-related information extraction.
[0018] The external interface and document question-and-answer module is used for unified access and question-and-answer of interface data sources and unstructured data sources. On the one hand, it can fully understand the task requirements through the big model based on the existing interface API application list, select the appropriate interface from the application list, and extract the parameter values required by the interface for calling; on the other hand, it can perform semantic similarity search based on the vector library constructed based on unstructured documents, intercept similar document fragments, and generate answers through the big model;
[0019] The answer synthesis generation module is used to comprehensively process answers from different sources. First, the questions are classified. If the question belongs to the judgment type, it will be displayed as yes / no. If the question belongs to the summation type, the returned results will be counted. Then, in the case where the answers from various sources are lengthy or conflicting, the answers are separated from the false based on their respective basis, and the answers are generated by comprehensive summarization and induction. Finally, the sensitivity test of the answer is used to confirm whether the final answer will be output.
[0020] The knowledge graph question-answering module includes a context completion tool, a question entity extraction tool, an entity linking tool, and a result compilation and organization tool. The context completion tool generates complete questions with reference completion based on historical question-answering records through a dedicated prompting project; the question entity extraction tool extracts the core entity names in the question through a dedicated prompting project; the entity linking tool uses the core entity as a keyword to perform an associated entity query in the knowledge graph, finds the top-k entities, and realizes the linking of key entities through a large model through the prompting project, and returns the specific attribute content of the key core entity after the link is accurate; the result compilation tool comprehensively organizes the specific attribute content, removes false and repeated content, and then comprehensively generates the answer to the question.
[0021] The external interface question and answer function in the external interface and document question and answer module mainly includes a problem interface mapping classification tool, an interface parameter extraction tool, and an interface calling tool. The problem interface mapping classification tool is to map the content of the existing API list through a deep analysis of the meaning of the problem through the prompt project; the interface parameter extraction tool is to extract the API parameter value and return value of the interface call through the interface and parameter information in the interface registered information after mapping the external interface to which it belongs, and then call the API interface to query the specific result; the interface calling tool standardizes and normalizes the corresponding natural language text, such as standardizing the location into the location name in the standard library, unifying the numerical value into a consistent unit, etc., and then calls the API interface to return the specific result, and replaces the field name of the result with Chinese to prompt the project to extract the core entity name in the problem; the entity linking tool uses the core entity as a keyword to perform an associated entity query in the knowledge graph, finds the top-k entity, and realizes the link of the key entity through the large model through the prompt project, and returns the specific attribute content of the key core entity after the link is accurate; the result compilation tool comprehensively organizes the specific attribute content, removes the pseudo and repeated content, and then generates the answer to the question.
[0022] The document question-and-answer function in the external interface and document question-and-answer module mainly includes a document vector library storage tool, a vectorized semantic retrieval tool, a candidate set and prompt generation tool, and a retrieval enhanced answer generation tool. The document vector library storage tool is to slice and vectorize the paragraphs in the document through the embedding model to obtain the maximum fine-grained understanding ability of the document; the vectorized semantic retrieval tool is to combine traditional semantic retrieval with a large model, first screen out the candidate set through vectorized semantic retrieval, then calculate the similarity between the document and the question through the large model to re-rank the similarity, and finally reduce the number of input words of the document through document summary or compression function to improve the calculation efficiency of subsequent steps; the candidate set and prompt generation tool generates corresponding prompt words by analyzing the characteristics of the candidate documents; the retrieval enhanced answer generation tool inputs the generated prompt words into the large model, matches the answers generated by the large model with the candidate documents, and comprehensively generates an answer structure containing references.
[0023] The answer comprehensive generation module includes an answer aggregation tool, an answer type classification tool, an answer source sorting tool, and a authenticity and sensitive word detection tool. The answer aggregation tool mainly plays the role of aggregating and organizing data from different sources; the answer classification model mainly determines the corresponding display method based on the representation style of the question and the answer; the answer source sorting tool displays different answers in order based on the sorting model; the authenticity and sensitive word detection tool makes the final judgment on the generated final answer, and avoids misleading the user if there is an authenticity problem or sensitive words are involved.
[0024] Based on the general inventive concept, the present invention also proposes an application method of a multi-source data intelligent question answering system based on a large language model, providing a unified registration and publishing paradigm for multi-source data, and using a large model to complete and classify user question semantics based on context, and according to the type, it is divided into query knowledge graph, call service interface API and document content understanding, and the call results are summarized and refined through a large language model, and finally the answer is automatically generated. The method comprises the following steps:
[0025] S1, the knowledge graph question-answering module analyzes the compiled knowledge graph and locates nodes and extracts relevant information. Its main core functions are context completion, problem core entity recognition, entity linking, and entity-related information extraction. The specific steps are as follows:
[0026] A. By taking the historical n rounds of question-answer records as input, and combining them with the prompt sentences exclusive to the context completion tool and inputting them into the big model, a complete question sentence with reference completion can be generated.
[0027] B. Use the completed question as input, and extract the core entity names in the question through the prompt engineering specific to the question entity extraction tool to form the core entity names extracted from the question.
[0028] C. According to the entity link interface called by the knowledge graph features, query the top-k candidate entities with different values and the key information query methods in the relevant candidate entities according to different types of entities.
[0029] D. After comprehensively re-ranking from different dimensions such as information granularity, information richness, and relevance to the question according to the core information related to the entity, return the entity information with the highest ranking.
[0030] E. Through the prompt engineering specific to the knowledge graph search result organization tool, organize the question results that can be answered in the entity information and return to generate the answer and basis.
[0031] S2. The external interface and document Q&A module uniformly accesses and answers interface data sources and unstructured data sources. On the one hand, based on the existing interface API application list, it can fully understand the task requirements through the large model, select appropriate interfaces from the application list, and extract the parameter values required by the interfaces for invocation; on the other hand, it can perform semantic similarity search based on the vector library constructed from unstructured documents, intercept similar document fragments, and generate answers through the large model. The specific steps are as follows:
[0032] A. Register information such as the function description, parameter description, and typical examples of the external interface in the registration center, where the function description needs to be unique and reduce ambiguity with other interfaces.
[0033] B. When the user enters a question, input the question and the external interface into the large model as a specific thought chain prompt engineering statement. The large model automatically understands the external interface and, based on the user's question and in the way of the thought chain, generates multiple possible API interface types. If there is no matching interface, it will go to F; if the match is successful, it will go to C.
[0034] C. According to the parameter description of the matching interface and the interface parameter extraction tool, organize and form a specific prompt engineering statement and input it into the large model to extract the parameter names, parameter values covered in the question, and the possible parameter return field list corresponding to the question.
[0035] D. According to the extracted parameter values, call the API interface, and generate a Chinese JSON style according to the return field list to organize and form the API return answer.
[0036] E. Organize the answer returned by the API and the question into a specific prompt engineering statement and input it into the large model to form the API service answer.
[0037] F. If the API interface cannot be matched, a fallback answer is formed jointly by the large model and the vector database. Candidate document slices are generated through vectorized search of the constructed document vector database.
[0038] G. Organize the candidate document slices, document marking information, and the question to form a dedicated prompt engineering statement and input it into the large model to form a document vector database and return the answer.
[0039] H. Aggregate the above answers and synthesize them as an alternative input to the large model.
[0040] S3. The answer synthesis generation module comprehensively processes answers from different sources. First, it classifies the questions. For example, if the question belongs to the judgment type, it presents "yes / no"; if the question belongs to the summation type, it calculates the returned results, etc. Then, for the situation where the answers from various sources are lengthy and there are conflicts between the answers, it sifts out the true from the false based on their respective answer bases, comprehensively summarizes and generates an abstract to form the answer. Finally, it confirms whether to output the final answer through the sensitivity detection of the answer. The specific steps are as follows:
[0041] A. Aggregate the question answers generated by different solution engines and perform filtering processing, leaving the available partial answers according to the filtering logic.
[0042] B. Organize the answer and the question to form a dedicated prompt engineering statement and input it into the large model to finally generate a concise question - answer - citation source.
[0043] C. Input the content of the question - answer into the answer type classification model to judge the display mode of the answer.
[0044] D. Input the answer - citation source into the data source sorting model to sort the content of the answer.
[0045] E. Organize the question - answer - citation source to form a dedicated prompt engineering statement for the answer authenticity and sensitivity judgment tool and input it into the large model. If the judgment is okay, the answer is directly returned; if there is a problem, it feedbacks that it cannot answer.
[0046] The present invention has the capabilities of question semantic understanding, complex task decomposition and invocation. The advantage is that while being able to efficiently access the application system, the large model can accurately understand the user's intention, drive various applications to obtain the question answer, and can integrate it into the answer style that is most convenient for users to understand, effectively improving the generality and accuracy of intelligent question - answering.
[0047] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0048] Embodiment
[0049] A multi-source data intelligent question answering system based on a large language model provided in an embodiment of the present invention is as follows Figure 1 As shown, the knowledge extraction platform includes a knowledge graph question and answer module, an external interface and document question and answer module, and a comprehensive answer generation module.
[0050] The knowledge graph question-answering module is used to analyze the compiled knowledge graph and locate nodes and extract relevant information. Its main core functions are context completion of questions, recognition of core entities, entity linking, and extraction of entity-related information.
[0051] In this embodiment, the compiled and improved knowledge graph is stored in the commonly used neo4j graph database, which saves the node information of common entities and the information between nodes.
[0052] The specific implementation examples are as follows:
[0053] A. By taking the historical n rounds of question-answer records as input, and combining them with the prompt sentences exclusive to the context completion tool as shown below and inputting them into the big model, a complete question sentence with reference completion can be generated.
[0054]
[0055] B. Take the completed question as input, and extract the core entity name in the question through the prompt project exclusive to the question entity extraction tool to form the core entity name extracted from the question.
[0056]
[0057]
[0058] C. The entity linking interface called according to the knowledge graph features selects the TOP-100 candidate entities according to the amount of data in this embodiment, as well as the key information query method in the related candidate entities.
[0059] D. According to the core information related to the entity, the entity is ranked by comprehensively calculating the information granularity, information richness, and relevance to the question. In this embodiment, the weights of the three are (0.1, 0.4, 0.5), and the entity information with the highest ranking is returned.
[0060] E. The present embodiment of the tool-specific prompting project through the knowledge graph search results organization organizes the answerable question results in the entity information and returns the generated answers and the basis.
[0061]
[0062] The external interface and document Q&A module is used to uniformly access and answer questions for interface-based data sources and unstructured data sources. On the one hand, it can, based on the existing interface API application list, fully understand the task requirements through a large model, select appropriate interfaces from the application list, and extract the parameter values required by the interfaces for invocation; on the other hand, it can perform semantic similarity search based on the vector library constructed from unstructured documents, intercept similar document fragments, and generate answers through the large model. The specific steps are as follows:
[0063] A. Register information such as the function description, parameter description, and typical examples of the external interface in the registration center. The function description needs to be unique and reduce ambiguity with other interfaces. The specific reference to this embodiment is as follows.
[0064]
[0065]
[0066] B. When the user inputs a question, input the question and the external interface to form a dedicated thought chain prompt engineering statement into the large model. The large model automatically understands the external interface and, based on the user's question and in a thought chain manner, generates multiple possible API interface types. If there is no matching interface, it goes to F; if the match is successful, it goes to C.
[0067]
[0068]
[0069] C. According to the parameter description of the matching interface and the interface parameter extraction tool, organize and form a dedicated prompt engineering statement and input it into the large model to extract the parameter names, parameter values, and the possible parameter return field list corresponding to the question covered in the question. The specific prompt statements can be as in the following embodiments.
[0070]
[0071]
[0072] D. According to the extracted parameter values, call the API interface, generate a Chinese JSON-style return value based on the return field list, and organize it to form an API return answer.
[0073] E. Organize the answer returned by the API and the question to form a dedicated prompt engineering statement and input it into the large model to form an API service answer. The specific prompt for this embodiment is as follows.
[0074]
[0075] F. If no API interface can be matched, a fallback answer is formed jointly by the large model and the vector database. Candidate document slices are generated through vectorized search of the constructed document vector database.
[0076] G. Organize the candidate document slices, document marking information, and the question to form a dedicated prompt engineering statement and input it into the large model to form an answer returned by the document vector database.
[0077] H. Aggregate the above answers and comprehensively input them as alternatives into the large model.
[0078] The specific prompts of this embodiment are as follows.
[0079]
[0080] The answer synthesis generation module is used to comprehensively process answers from different sources. First, it classifies the questions. For example, if the question belongs to the judgment type, it will display yes / no; if the question belongs to the summation type, it will count the returned results, etc. Then, for the situation where the answers from various sources are lengthy and there are conflicts between the answers, it discards the false and retains the true according to their respective answer bases, comprehensively summarizes and generates an answer. Finally, it confirms whether to output the final answer through sensitivity detection of the answer.
[0081] A. Aggregate the question answers generated by different solution engines and perform filtering processing. According to the filtering logic, retain the available partial answers. In this embodiment, the answers from the knowledge graph have the highest priority, followed by the content returned by the API interface, and the answers from document Q&A and the native content of the large model have the lowest priority.
[0082] B. Organize the answer and the question to form a dedicated prompt engineering statement and input it into the large model, and finally generate a concise question-answer-reference source. The specific prompts of this embodiment are as follows.
[0083]
[0084] C. Input the question-answer content into the answer type classification model to judge the display mode of the answer. The following is this embodiment.
[0085]
[0086]
[0087] D. Input the answer-reference source into the data source sorting model to sort the content of the answer.
[0088] E. Organize the question-answer-reference source to form a dedicated prompt engineering statement for the answer authenticity and sensitivity judgment tool and input it into the large model. If the judgment is okay, directly return the answer; if there is a problem, feedback that it cannot answer. The following is this embodiment.
[0089]
Claims
1. A multi-source data intelligent question answering system based on a large language model, characterized in that: It includes a knowledge graph question-answering module based on a large model, an external interface and document question-answering module, and a comprehensive answer generation module; The knowledge graph question-answering module is used to analyze the compiled knowledge graph and locate nodes and extract relevant information, including context completion of questions, identification of core entities, entity linking, and extraction of entity-related information; The external interface and document question-and-answer module is used for unified access and question-and-answer of interface data sources and unstructured data sources. Based on the existing interface API application list, it can fully understand the task requirements through the big model, select the appropriate interface from the application list, and extract the parameter values required by the interface for calling; it can perform semantic similarity search based on the vector library constructed based on unstructured documents, intercept similar document fragments and generate answers through the big model; The comprehensive answer generation module is used to comprehensively process answers from different sources. First, the questions are classified. If the question belongs to the judgment type, it will be displayed as yes / no. If the question belongs to the summation type, the returned results will be counted. Then, in the case where the answers from various sources are lengthy or conflicting, the answers are separated from the false based on their respective basis, and the answers are generated by comprehensive summarization and induction. Finally, the sensitivity test of the answer is used to confirm whether the final answer will be output.
2. The multi-source data intelligent question answering system based on a large language model according to claim 1 is characterized in that: The knowledge graph question-answering module includes a context completion tool, a question entity extraction tool, an entity linking tool, and a result compilation and organization tool; the context completion tool generates a complete question sentence for reference completion based on historical question-answering records through a dedicated prompting project; The question entity extraction tool extracts the core entity names in the question through a dedicated prompt project; the entity linking tool uses the core entities as keywords to perform related entity queries in the knowledge graph, finds the top-k entities, and uses the prompt project to link key entities through a large model. After the link is accurate, the specific attribute content of the key core entity is returned; the result compilation tool comprehensively organizes the specific attribute content, removes false and repeated content, and then comprehensively generates the answer to the question.
3. The multi-source data intelligent question answering system based on a large language model according to claim 1, characterized in that: The external interface question and answer function in the external interface and document question and answer module mainly includes a problem interface mapping classification tool, an interface parameter extraction tool, and an interface calling tool; the problem interface mapping classification tool parses the meaning of the problem and maps it with the content in the existing API list through the prompt project; the interface parameter extraction tool extracts the API parameter value and return value of the interface call through the interface and parameter information in the registered information of the interface after mapping the external interface to which it belongs, and then can call the API interface to query the specific result; the interface calling tool standardizes and normalizes the corresponding natural language text, then calls the API interface to return the specific result, and replaces the field name of the result with Chinese to prompt the project to extract the core entity name in the problem; the entity linking tool uses the core entity as a keyword to perform associated entity query in the knowledge graph, finds the top-k entities, and realizes the linking of key entities through the large model through the prompt project, and returns the specific attribute content of the key core entity after the link is accurate; the result compilation tool comprehensively organizes the specific attribute content, removes false and repeated content, and then comprehensively generates the answer to the question.
4. The multi-source data intelligent question answering system based on a large language model according to claim 1, characterized in that: The document question and answer function in the external interface and document question and answer module mainly includes a document vector library storage tool, a vectorized semantic retrieval tool, a candidate set and prompt generation tool, and a retrieval enhanced answer generation tool; the document vector library storage tool slices and vectorizes the paragraphs in the document through the embedding model to obtain the maximum fine-grained understanding ability of the document; the vectorized semantic retrieval tool combines traditional semantic retrieval with a large model, first screens out the candidate set through vectorized semantic retrieval, then calculates the similarity between the document and the question through the large model to re-rank the similarity, and finally reduces the number of input words in the document through document summary or compression function; the candidate set and prompt generation tool generates corresponding prompt words by analyzing the characteristics of the candidate documents; the retrieval enhanced answer generation tool inputs the generated prompt words into the large model, matches the answers generated by the large model with the candidate documents, and comprehensively generates an answer structure containing references.
5. The multi-source data intelligent question answering system based on a large language model according to claim 1, characterized in that: The comprehensive answer generation module includes an answer aggregation tool, an answer type classification tool, an answer source sorting tool, and an authenticity and sensitive word detection tool; the answer aggregation tool is used to aggregate and organize data from different sources; the answer type classification tool mainly determines the corresponding display method based on the representation style of the question and the answer; the answer source sorting tool displays different answers in order based on the sorting model; the authenticity and sensitive word detection tool makes the final judgment on the final answer generated, and if there are authenticity issues and sensitive words involved, avoids output that misleads users.
6. An application method of the multi-source data intelligent question answering system based on a large language model as claimed in claim 1, characterized in that: It provides a unified registration and publishing paradigm for multi-source data, and uses a large model to complete and classify user question semantics based on context. It is divided into query knowledge graph, call service interface API and document content understanding according to type. The results of the call are summarized and refined through a large language model, and finally the answer is automatically generated.
7. The method according to claim 6, characterized in that The method specifically comprises the following steps: The knowledge graph question-answering module analyzes the compiled knowledge graph and locates nodes and extracts relevant information, including question context completion, question core entity recognition, entity linking, and entity-related information extraction; The external interface and document question-and-answer module provides unified access and question-and-answer services for interface data sources and unstructured data sources. Based on the existing interface API application list, the module fully understands the task requirements through the big model, selects the appropriate interface from the application list, and extracts the parameter values required by the interface for calling. The module performs semantic similarity search based on the vector library built based on unstructured documents, extracts similar document fragments, and generates answers through the big model. The comprehensive answer generation module processes the answers from different sources comprehensively. First, it classifies the questions. If the question belongs to the judgment type, it will be displayed as yes / no. If the question belongs to the summation type, the result will be returned for statistics. Then, for the cases where the answers from various sources are lengthy or conflicting, the module will separate the true from the false based on their respective answers, and generate the answer through comprehensive summary and induction. Finally, the sensitivity test of the answer is used to confirm whether the final answer will be output.
8. The method according to claim 7, characterized in that The knowledge graph question-answering module analyzes the compiled knowledge graph and locates nodes and extracts relevant information, including question context completion, question core entity recognition, entity linking, and entity-related information extraction. The specific steps are as follows: A. By taking the historical n rounds of question-answer records as input, and combining them with the prompt sentences exclusive to the context completion tool and inputting them into the big model, a complete question sentence with reference completion can be generated; B. Take the completed question as input, and extract the core entity name in the question through the prompt engineering exclusive to the question entity extraction tool to form the core entity name extracted from the question; C. The entity linking interface called according to the knowledge graph features queries the top-k candidate entities with different values according to the different types of entities, as well as the key information query method in the relevant candidate entities; D. Based on the core information related to the entity, after comprehensive re-ranking based on the dimensions of information granularity, information richness, and relevance to the question, the entity information with the highest ranking is returned; E. Through the exclusive prompt project of the knowledge graph search result organization tool, the question results that can be answered in the entity information are organized and the generated answers and basis are returned.
9. The method according to claim 7, characterized in that: The external interface and document question-and-answer module provides unified access and question-and-answer services for interface data sources and unstructured data sources. Based on the existing interface API application list, the module fully understands the task requirements through the big model, selects the appropriate interface from the application list, and extracts the parameter values required by the interface for calling. The module performs semantic similarity search based on the vector library built based on unstructured documents, extracts similar document fragments, and generates answers through the big model. The specific steps are as follows: A. Register the function description, parameter description and typical example information of the external interface in the registration center; B. When the user inputs a question, the question and the external interface form a unique thinking chain prompt engineering statement and input it into the big model. The big model automatically understands the external interface and generates multiple possible API interface types based on the user's question and the thinking chain. If there is no matching interface, it will go to F. If a match is successful, it will go to C. C. According to the parameter description of the matching interface and the interface parameter extraction tool, a dedicated prompt engineering statement is formed and input into the big model to extract the parameter name, parameter value and possible parameter return field list corresponding to the problem covered in the question; D. Call the API interface based on the extracted parameter values, generate the return value in Chinese JSON format based on the return field list, and organize the API return answer; E. Organize the answers and questions returned by the API into a dedicated prompt engineering statement and input it into the big model to form the API service answer; F. If the API interface cannot be matched, a backup answer is formed through the big model and vector library; candidate document slices are generated through vectorized search of the constructed document vector library; G. Slice candidate documents, document tag information, and questions into exclusive prompt engineering sentences and input them into the big model to form a document vector library and return the answer; H. Gather the above answers to form alternative inputs to the large model.
10. The method according to claim 7, characterized in that The answer synthesis generation module is used to process the answers from different sources. First, the questions are classified. If the questions belong to the judgment type, they will be displayed as yes / no. If the questions belong to the sum type, the returned results will be counted. Then, if the answers from various sources are lengthy or conflicting, the answers are separated according to their respective answers, and the answers are generated by comprehensive summary and induction. Finally, the sensitivity test of the answers is used to confirm whether the final answer will be output, specifically: A. Gather answers generated by different problem-solving engines and filter them, leaving only some available answers based on the filtering logic; B. Organize the answers and questions into exclusive prompt engineering sentences and input them into the big model, and finally generate question-answer-reference source; C. Input the question-answer content into the answer type classification model to determine the display mode of the answer; D. Input the answer-reference source into the data source sorting model to sort the content of the answer; E. Organize the question-answer-reference source into prompt engineering statements exclusive to the answer authenticity and sensitivity judgment tool and input them into the big model. If it is judged that there is no problem, the answer is returned directly. If there is a problem, it is reported that it cannot be answered.
Citation Information
Cited By
Knowledge question-answering system based on intelligent agent
CN121029951A
Interactive progressive scene understanding and intelligent question answering method and system for power grid
CN121413721A