Intelligent question answering system and method based on multi-source mixed retrieval and dynamic self-optimization
By introducing multi-source hybrid retrieval and dynamic self-optimization technology into the intelligent question and answer system, integrating the document library and question and answer library, and using large language models and historical data, the problem of lack of knowledge source tracking and self-optimization in the existing system is solved, and dynamic optimization of the knowledge base and quality improvement of the question and answer library is achieved.
Patent Information
- Application Number
- CN202510468987.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing intelligent question and answer system lacks a knowledge source tracking mechanism for the generated answer text, and does not involve self-optimizing the generated content, making it difficult to achieve dynamic management and optimization of the knowledge base.
An intelligent question-answer system based on multi-source hybrid retrieval and dynamic self-optimization is proposed. By integrating the document library and question-answer library, it uses large language models and historical data to realize an efficient and intelligent question-answer system. Through the collaboration of multiple modules, the system generates the optimal response and tracks the knowledge source of the best response, ultimately achieving dynamic optimization of the knowledge base.
The synchronization and optimization of the Q&A library and the document library are realized, the quality and coverage of the Q&A library are improved, the dependence on the original document is reduced, and the response speed and the diversity and quality of answers are improved.
Smart Images

Figure CN120045682A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of natural language processing and machine learning, and particularly relates to an intelligent question-answering system and method based on multi-source hybrid retrieval and dynamic self-optimization. Background Art
[0002] Currently, intelligent document retrieval and question-answering systems based on retrieval-augmented generation technology have been widely applied in various fields. These systems usually combine retrieval and generation, and through natural language processing technology, improve the response speed and accuracy of user queries. Typical implementation methods include the following steps: First, collect documents from diverse data sources and perform text processing on them; then, use vectorization technology to index the text data; after the user queries, the system retrieves in the knowledge base to generate results that meet the user's needs; finally, the system will review and optimize the generated results according to a preset quality evaluation mechanism. This hybrid mode that combines retrieval and generation is a mainstream architecture of current intelligent question-answering systems.
[0003] The invention patent with the publication number CN117951274A discloses a knowledge question-answering method that combines vector retrieval and keyword retrieval. The system generates a vector representation of the question through a vector embedding model, retrieves semantically similar text paragraphs in the vector database, and simultaneously performs keyword retrieval in the search engine. Then, the system de-duplicates and cross-merges the vector retrieval and keyword retrieval results to form new prompt text, and inputs it into the large language model to generate answers.
[0004] The above patent combines vector retrieval and keyword retrieval, aiming to improve the relevance and accuracy of retrieval results through the complementarity of the two methods, but it lacks a mechanism for tracking the knowledge sources of the generated answer text and does not involve self-optimization of the generated content. Summary of the Invention
[0005] To solve the problems and deficiencies existing in the above-mentioned prior art, the present invention specifically proposes an intelligent question-answering system and method based on multi-source hybrid retrieval and dynamic self-optimization. The present invention integrates a document library and a question-answering library, and uses a large language model and historical data to implement an efficient and intelligent question-answering system. Through the cooperation of multiple modules, the system can ensure that when processing user queries, it can generate an optimal response, track the knowledge source of the best response, and finally realize the dynamic optimization of the knowledge base.
[0006] To achieve the above invention purpose, the technical solution of the present invention is as follows: On the one hand, the present invention proposes an intelligent question-answering system based on multi-source hybrid retrieval and dynamic self-optimization, and the system includes: an invalid question filtering module, a question rewriting module, a hybrid retrieval module, an answer generation module, a response evaluation module, and an output analysis module; wherein, The invalid question filtering module analyzes the question input by the user using a pre-trained language model in response to the user's input question to determine whether the user's current question is an invalid question; when the user's question is determined to be an invalid question, the system outputs a default answer; when the user's question is determined to be a valid question, it enters the question rewriting module for parsing; The question rewriting module analyzes and understands the question based on the user's input question and combines historical conversation data through a large language model, rewrites the user's input question, and finally sends the rewritten question to the hybrid retrieval module; The hybrid retrieval module vectorizes the rewritten user question and retrieves it in the vector database to obtain the text data most similar in semantics to the user question; at the same time, based on keyword retrieval technology, it retrieves the data matching the user question in the knowledge base; finally, it uses the reciprocal rank fusion technology to integrate and obtain the sorting result of the new retrieval data combining the information of both; The answer generation module uses a large language model to generate several different candidate answers based on the rewritten user question and the sorting result of the new retrieval data combining the information of both obtained by using the reciprocal rank fusion technology; The response evaluation module uses a pre-trained language model to score the generated candidate answers and evaluate their matching degree with the user question, and the candidate answer with the highest score will be selected as the best response and finally output the answer; The output analysis module analyzes the generated best response and determines its knowledge source; when the knowledge comes from the document library, it collects the user's question and the output best response as a question-answer pair into the question-answer library.
[0007] Preferably, the knowledge base includes a document library and a question-answer library.
[0008] Preferably, the vector database is converted from the knowledge base and includes a vector document library and a vector question-answer library.
[0009] Preferably, the system further includes a feedback module. After the system outputs an answer, the feedback module evaluates the answer output by the system through user feedback or manual review and determines whether to include the current question-answer pair in the question-answer library.
[0010] Preferably, the user feedback includes giving a like and scoring.
[0011] Preferably, the system determines whether the question-answer library is saturated based on the principle of maximum adoption number.
[0012] Preferably, the system preprocesses the question-answer pair and then collects it into the question-answer library.
[0013] Based on the same inventive concept, on the other hand, the present invention also proposes an intelligent question-answering method based on multi-source hybrid retrieval and dynamic self-optimization. The method is implemented based on the above system and mainly includes the following steps: First, construct a local knowledge base using knowledge data and form a vector database based on the constructed knowledge base; The user inputs a question, and based on the input question, it is judged whether the user's current question belongs to an invalid question; when the user's question is judged to be an invalid question, a default answer is output and the question-answering process is terminated, and the system returns to the initial interface; when the user's question is determined to be a valid question, the user's question is rewritten; The rewritten user question is vectorized and retrieved in the vector database to obtain the text data that is semantically most similar to the user question; at the same time, keyword retrieval is performed in the knowledge base to retrieve the data that matches the user question; finally, the reciprocal rank fusion technology is used to integrate and obtain the sorting result of the new retrieved data that combines the information of both; According to the rewritten user question and the sorting result of the new retrieved data that combines the information of both obtained by using the reciprocal rank fusion technology, different candidate answers are generated; Based on the generated different candidate answers, evaluate their matching degree with the user question, and select the candidate answer with the highest matching degree as the best response, and the system finally outputs this answer; The system analyzes the generated best response and judges its knowledge source; when the knowledge comes from the document library, the user's question and the output best response are formed into a question-answer pair and collected into the question-answer library.
[0014] Furthermore, on yet another aspect, the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable in the processor. When the processor executes the computer program, the above-mentioned knowledge-enhanced question-answering method is implemented.
[0015] Furthermore, on still another aspect, the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed in a computer processor, the above-mentioned knowledge-enhanced question-answering method is implemented.
[0016] Advantages of the present invention: 1. The present invention introduces a dynamic self-optimization mechanism based on user feedback. The system can continuously optimize the question-answer library and the document library through user feedback, and fine-tune the generation model, enabling the system to have the ability of self-evolution. The system continuously improves its own performance over time and adapts to the changing user needs.
[0017] 2. The present invention records the knowledge sources relied on by the best responses and, in combination with user feedback, dynamically optimizes the content of the Q&A library and the document library. The present invention realizes the synchronization and optimization of the Q&A library and the document library, enabling the knowledge in the document library to be effectively transformed into Q&A pairs, ultimately improving the quality and coverage of the Q&A library; and enabling the knowledge in the document library to be effectively transformed into Q&A pairs and enter the Q&A library. This mechanism improves the coverage and quality of the Q&A library, enabling the system to generate accurate answers more efficiently when processing queries, reducing the dependence on the original documents, and improving the response speed.
[0018] 3. The present invention records the knowledge sources relied on by the best responses and, in combination with user feedback, dynamically optimizes the content of the Q&A library and the document library. Only when the knowledge of the output Q&A pairs comes from the document library will the generated Q&A pairs be included in the Q&A library, avoiding the problem of duplicate Q&A library data.
[0019] 4. In the input processing stage, the present invention introduces an invalid question filtering mechanism and question rewriting, which can effectively improve the system's ability to understand user questions, ensuring that the questions input by users have been optimized before entering the retrieval and generation links. Therefore, the present invention significantly reduces the impact of invalid inputs on the system performance and improves the overall response quality.
[0020] 5. The answer generation module of the system of the present invention generates different responses based on different knowledge sources, evaluates and ranks these responses, and finally outputs the optimal answer. Compared with the current single answer generation method, the present invention significantly improves the diversity and quality of the generated answers. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The foregoing and following specific descriptions of the present invention become clearer when read in conjunction with the following drawings, in which: Figure 1 is the system structure diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will further illustrate the technical solutions for achieving the object of the present invention through several specific embodiments. It should be noted that the technical solutions claimed by the present invention include but are not limited to the following embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] Currently, the knowledge conversion efficiency between the document library and the Q&A library of the knowledge Q&A system is relatively low, resulting in a large amount of information in the document library not being effectively converted into efficient Q&A pairs. The utilization rate of the Q&A library is not high, and the response efficiency and accuracy of the system are relatively limited when dealing with repetitive or common questions. In this case, the system mostly relies on the document library for retrieval and generation, without fully utilizing the existing Q&A pair resources, ultimately affecting the response speed and quality of the system.
[0024] Moreover, existing systems generally lack a mechanism to track and record the knowledge sources of generated responses, and it is impossible to clarify the knowledge sources on which each response is based. This lack of knowledge source tracking makes it difficult for the system to evaluate the knowledge distribution between the document library and the Q&A library, and it is also unable to effectively perform dynamic management and optimization of the knowledge base. This limits the adaptive ability of the system during long-term operation and the effective utilization of knowledge resources.
[0025] Based on this, embodiments of the present invention propose an intelligent Q&A system and method based on multi-source hybrid retrieval and dynamic self-optimization. The present invention integrates the document library and the Q&A library, and uses large language models and historical data to implement an efficient and intelligent Q&A system. Through the cooperation of multiple modules, the system ensures that when processing user queries, it can generate the optimal response, track the knowledge source of the best response, and ultimately achieve the dynamic optimization of the knowledge base.
[0026] This embodiment discloses an intelligent Q&A system based on multi-source hybrid retrieval and dynamic self-optimization. Figure 1 This is the system architecture diagram of the present invention. Refer to the attached Figure 1 As shown in the figure, the Q&A system includes a question collection module, an invalid question filtering module, a question rewriting module, a hybrid retrieval module, an answer generation module, a response evaluation module, an output analysis module, and a feedback module. Below, each functional module will be specifically introduced and described in combination with the working process of the system.
[0027] First of all, system developers or managers will use knowledge data to build a local knowledge base in the system. The knowledge base consists of a document library and a Q&A library. It is worth mentioning that the document library is established based on various internal document files of the system owner, and it exists at the beginning of the system establishment. The Q&A library is formed as the system is used, and the feedback module continuously imports high-quality document data into this library. That is to say, the knowledge in the Q&A library comes from the document library.
[0028] Generally, the data in the Q&A library is composed of several Q&A pairs formed by the combination of high-quality user input questions and the best responses generated by the system. The document library generally stores various internal knowledge documents of the system owner, which can be doc, pdf, excel, etc., and are decomposed into paragraphs of knowledge through parsing. Both the document library and the Q&A library are composed of natural text.
[0029] Therefore, it can be understood that at the initial stage of the system, there is only a document library. As the system runs, the system gradually collects question-and-answer pairs and puts them into the question-and-answer library. That is, the system continuously imports high-quality user questions and the best responses generated based on the document library into the question-and-answer library, and finally forms a knowledge base in the form of a document library + question-and-answer library. Moreover, the question-and-answer library contains at least one question and at least one answer corresponding to each question, serving as the answer set for each question.
[0030] Then, based on the constructed knowledge base, the system performs vectorization processing on the knowledge base to form a vector database. The vector database is in the form of digital vectors, which corresponds one-to-one with the knowledge base and includes a vector document library and a vector question-and-answer library.
[0031] After the knowledge base is constructed, the system collects and obtains user questions, and then uses a pre-trained model and a large language model to filter, rewrite, and other processes on the user questions. Retrieve in the knowledge base according to the processed user questions and generate several different candidate answers. The system matches the best response as the answer to the user question and outputs it to achieve human-computer interaction between the system and the user. Finally, the system will also trace the knowledge source of the best response and then continuously optimize the knowledge base. The specific question-and-answer process of the system of the present invention is as follows: S1. User question collection Based on the question collection module, the system obtains the question input by the user and transmits the user question to the invalid question filtering module.
[0032] S2. Invalid question filtering In response to the question input by the user, the invalid question filtering module uses a pre-trained language model to analyze the question input by the user to determine whether the current user question belongs to an invalid question. When the user question is determined to be an invalid question, the system will output a default answer. When the user question is determined to be a valid question, it enters the question rewriting module to parse the question input by the user.
[0033] In the embodiment described in the present invention, whether the user's question is valid is judged by a standard established internally by the system developer or administrator. Among them, for the training of the model, the actual user questions within the system developer or administrator are used as inputs, and whether they are invalid questions are used as outputs, and an end-to-end invalid question filtering model can be obtained. In the actual scenario, only the user's question needs to be input into this model, and it can be directly judged whether the current question is an invalid question after calculation.
[0034] It is understandable that pre-trained language models such as BERT, ERNIE, and T5 can be used for filtering invalid questions. For the default answers output by the invalid question filtering module, their forms are diverse and depend on the answers set internally by the system administrator or developer. For example, they include "Hello, user. I don't quite understand your question. Can you please be more specific?" or "The question cannot be recognized. Please ask in another way." etc. It can even be set to no response from the system, directly terminate the Q&A process, and then return to the system initial interface.
[0035] S3. Question Rewriting Based on the question input by the user, the question rewriting module analyzes and rewrites the user's question through large language models (LLM models) such as LLM in combination with the historical conversation data with the user to standardize it.
[0036] Under normal circumstances, the questions input by users may contain dialects or colloquial expressions. Therefore, it is necessary to further process the user question text, rewrite it to conform to the written language norms. It is understandable that question rewriting usually includes replacing synonyms, changing sentence structures, etc. Finally, the question rewriting module sends the rewritten user question to the hybrid retrieval module.
[0037] For the rewriting of the user question, the system first extracts the user's current question and the historical data of the conversation with the user, designs a prompt (such as "Please rewrite the user's current question {specific question text} in combination with the user's conversation history {specific conversation history text}, and it is required to complete the parts omitted by the user and replace the pronouns with specific things."), combines the prompt with the extracted data, inputs it into the large model, the large model outputs the rewritten question, and finally sends the rewritten user question to the hybrid retrieval module.
[0038] S4. Hybrid Retrieval The hybrid retrieval module first vectorizes the rewritten user question, converts the text data into vector data, and then retrieves it in the vector database to obtain the text data that is most semantically similar to the user question (calculate the similarity of the two vectors and sort according to the similarity); at the same time, according to the rewritten user question, use the keyword retrieval technology to first extract keywords from the user's question, and then retrieve them in the knowledge base), retrieve the data that matches the user question (similarity ranking); finally, use the reciprocal ranking fusion technology to integrate and obtain the ranking result of the new retrieval data that combines the information of both.
[0039] In the implementation mode described in the present invention, the essence of keyword retrieval is to search for the text corresponding to the meaning of some keyword phrases in the user question in the knowledge base, which can be directly carried out through natural text. Therefore, keyword retrieval is directly carried out in the knowledge base.
[0040] Furthermore, due to the implementation of semantic retrieval, it is difficult to directly use deep learning methods for efficient and accurate semantic retrieval through natural text. Therefore, this system adopts a semantic retrieval method based on the embedding model to perform retrieval in the vector library.
[0041] S5. Answer Generation The answer generation module combines the user's question with the knowledge retrieved by the hybrid retrieval module to generate several different candidate answers. The candidate answers include document fragments and Q&A pairs, ensuring that each generated answer can maximize the match with the user's needs. Finally, the generated candidate answers are input into the response evaluation module.
[0042] When generating answers, the system will utilize the natural language generation ability of the large language model to ensure the fluency and naturalness of the output content. First, prompt words are designed (such as "Please answer the user's question {specific question text} in combination with relevant knowledge {specific knowledge text}"), and the prompt words are combined with the user's question and different knowledge (combined with the knowledge retrieved from the document library and the knowledge retrieved from the Q&A library respectively) and input into the large model. Different responses (i.e., candidate answers) are obtained through the large model, including responses generated from the knowledge in the document library and responses generated from the knowledge in the Q&A library.
[0043] S6. Response Evaluation The response evaluation module uses a pre-trained language model to score the generated responses and evaluate their matching degree with the user's question. The response with the highest score will be selected as the final output answer. Specifically, first, question-response pairs are prepared and quality scores in multiple dimensions such as accuracy, relevance, and fluency are labeled as tags. Then, the pre-trained model (such as BERT, T5) is fine-tuned using this data to enable it to have the ability to evaluate the quality of answers. The fine-tuned model will score in multiple dimensions such as accuracy, relevance, and fluency, and finally, the comprehensive score (taking the average value) is used to select the answer with the highest score as the final output answer.
[0044] S7. Output Analysis and Tracking The output analysis module of the system analyzes the best answer output, determines whether its knowledge source is the Q&A library or the document library of the system, and then records the analysis results. The system combines the answers whose knowledge source is the document library with the corresponding user questions to form Q&A pairs and includes them in the Q&A library.
[0045] It is worth mentioning that the Q&A library is continuously increasing during the use of the system. Its essence is high-quality responses generated from the knowledge in the document library with the assistance of the large model. However, the knowledge in the document library is ultimately limited. Therefore, new content will not continue to be added to the Q&A pairs in the Q&A library unless new documents are continuously added.
[0046] Therefore, the system needs to determine whether the Q&A database is saturated. Otherwise, the valid knowledge in the Q&A database will be redundant, resulting in unnecessary resource waste. The judgment method is that for each added Q&A pair, record which document knowledge it originates from. Here, a maximum adoption number can be set (for different document knowledge parsing methods, each piece of knowledge contains different information and needs to be determined internally by system developers or administrators. For example, at most two Q&A pairs are allowed to be extracted from the same piece of document knowledge). If a piece of document knowledge exceeds the maximum adoption number, no more Q&A pairs will be extracted from it (that is, it will no longer be included in the Q&A database); if each piece of knowledge in the document reaches or approaches the maximum adoption number, it can be determined that the current Q&A database is saturated.
[0047] S8. User feedback or manual review After the system outputs the best answer, the feedback module evaluates the output answer through user feedback and decides whether to include the current user question and the corresponding answer in the Q&A database. User feedback can be that the user scores or sets like and dislike buttons, and the user can evaluate according to the generated answer.
[0048] Furthermore, after the system runs for a certain period of time, the output answers of the system can also be evaluated through manual review in the background. The answers that pass the review and the corresponding user questions can be included in the Q&A database as Q&A pairs.
[0049] It can be understood that for the above user feedback or manual review, either one can be selected or both can be used simultaneously.
[0050] It should be noted that the system, device, model, or unit described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, in this specification, when describing the above devices, various units are described separately according to their functions. Of course, when implementing the present invention, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0051] Furthermore, in another aspect of this embodiment, a computer device is further provided. The computer device includes a processor, an input device, an output device, and a memory, and the processor, input device, output device, and memory are interconnected; wherein, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the steps in the above embodiments.
[0052] Even further, in yet another aspect of this embodiment, a computer-readable storage medium is further provided, characterized in that: the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the steps in the above embodiments.
[0053] In this embodiment, the processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.
[0054] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and units, such as the corresponding program units in the above method embodiments of the present invention. By running the non-transitory software programs, instructions, and modules stored in the memory, the processor executes various functional applications and work data processing of the processor, that is, implements the methods in the above method embodiments.
[0055] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0056] The one or more units are stored in the memory and, when executed by the processor, execute the methods in the above embodiments.
[0057] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0058] As described above, these are only the preferred embodiments of the present invention and do not pose any formal obstacles to the present invention. Any simple modifications and equivalent changes made to the above embodiments based on the technical essence of the present invention all fall within the protection scope of the present invention.
Claims
1. An intelligent question-answering system based on multi-source hybrid retrieval and dynamic self-optimization, characterized in that: The system comprises: The invalid question filtering module responds to the questions input by the user and uses the pre-trained language model to analyze the questions input by the user to determine whether the current question of the user is an invalid question; when the user question is judged to be an invalid question, the system outputs a default answer; when the user question is judged to be a valid question, it enters the question rewriting module for analysis; The question rewriting module analyzes and understands the user questions based on the historical conversation data through a large language model, rewrites the user questions, and finally sends the rewritten user questions to the hybrid retrieval module; The hybrid retrieval module vectorizes the rewritten user questions and searches in the vector database to obtain text data that is most similar to the user questions. At the same time, based on keyword retrieval technology, data matching the user questions is retrieved from the knowledge base. Finally, the inverse sorting fusion technology is used to integrate the new retrieval data and obtain the ranking results that combine the information of the two. The answer generation module uses a large language model to generate several different candidate answers based on the rewritten user question and the ranking results of the new search data obtained by integrating the information of the two using the inverse ranking fusion technology; The response evaluation module uses the pre-trained language model to score the generated candidate answers and evaluate their matching degree with the user's question. The candidate answer with the highest score will be selected as the best response and finally output. The output analysis module analyzes the best response generated and determines its knowledge source; when the knowledge comes from the document library, the user's question and the best output response are combined into question-answer pairs and collected in the question-answer library.
2. The intelligent question-answering system based on multi-source hybrid retrieval and dynamic self-optimization according to claim 1 is characterized in that: The knowledge base includes a document base and a question and answer base.
3. The intelligent question-answering system based on multi-source hybrid retrieval and dynamic self-optimization according to claim 1 is characterized in that: The vector database is converted from a knowledge base and includes a vector document library and a vector question and answer library.
4. The intelligent question-answering system based on multi-source hybrid retrieval and dynamic self-optimization according to claim 1 is characterized in that: The system also includes a feedback module. After the system outputs an answer, the feedback module evaluates the answer output by the system through user feedback and determines whether to include the current question and answer pair into the question and answer library.
5. The intelligent question-answering system based on multi-source hybrid retrieval and dynamic self-optimization according to claim 4 is characterized in that: The user feedback includes likes and ratings.
6. The intelligent question-answering system based on multi-source hybrid retrieval and dynamic self-optimization according to claim 4 is characterized in that: The feedback module evaluates the answers output by the system through manual review and determines whether to include the current question and answer pair into the question and answer library.
7. The intelligent question-answering system based on multi-source hybrid retrieval and dynamic self-optimization according to claim 1 is characterized in that: The system determines whether the question-answer database is saturated based on the maximum adoption number principle.
8. The intelligent question-answering system based on multi-source hybrid retrieval and dynamic self-optimization according to claim 1 is characterized in that: The system pre-processes the question and answer pairs and then collects them into the question and answer database.
9. An intelligent question-answering method based on multi-source hybrid retrieval and dynamic self-optimization, wherein the intelligent question-answering method based on multi-source hybrid retrieval and dynamic self-optimization is implemented based on the intelligent question-answering system based on multi-source hybrid retrieval and dynamic self-optimization described in any one of claims 1 to 8, and is characterized in that: The following steps are involved: Utilize knowledge data to build a local knowledge base, and form a vector database based on the built knowledge base; The user inputs a question, and based on the question input by the user, it is determined whether the user's current question is an invalid question; When a user's question is judged to be invalid, a default answer is output and the question-answering process is terminated, and the system returns to the initial interface; When a user's question is determined to be a valid question, the user's question is rewritten; The rewritten user questions are vectorized and searched in the vector database to obtain the text data that is most similar to the user questions. At the same time, keyword search is performed in the knowledge base to retrieve the data that matches the user questions. Finally, the reciprocal sorting fusion technology is used to integrate the new search data to obtain the ranking results that combine the information of the two. Generate different candidate answers based on the rewritten user questions and the ranking results of the new retrieval data obtained by integrating the information of the two using the inverse sorting fusion technology; Based on the different candidate answers generated, the matching degree between them and the user's question is evaluated, and the candidate answer with the highest matching degree is selected as the best response, which the system finally outputs; The system analyzes the best response generated and determines its knowledge source; when the knowledge comes from the document library, the user's question and the best output response are combined into question-answer pairs and collected in the question-answer library.
Citation Information
Patent Citations
RAG knowledge question-answering method and device based on fusion vector and keyword retrieval
CN117951274A
Intelligent operation and maintenance platform construction method and system based on artificial intelligence
CN119398144A
Intelligent question and answer method, device and equipment based on tax field and medium
CN119623647A
Cited By
Generative search illustration method and device, storage medium and program product
CN120892589A
Customer service on-duty intelligent response method and device, computer equipment and storage medium
CN121462539A