A multimodal RAG knowledge question answering method and device applied to vertical fields
Through the multimodal RAG knowledge question and answer method, the problem of insufficient quality of search recall and answer prediction of existing question and answer system in vertical fields is solved, and more efficient and professional solution data generation is achieved.
Patent Information
- Application Number
- CN202411660708.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-11-20
AI Technical Summary
The existing question-and-answer system has shortcomings in the quality of retrieval recall and answer prediction, especially in applications in vertical fields, and faces problems such as scarcity of information, weak multi-hop problem processing ability, and unstable generation results.
The multimodal RAG knowledge question and answer method is adopted, and by receiving multimodal type data, converting it into text type data, and using multiple preset search models to search related knowledge fragments from the knowledge base. Combining the reciprocal sorting fusion algorithm and the target reordering model, the knowledge fragments are fusion and sorted to generate the optimal solution data.
It improves the search recall rate and answer prediction quality of the Q&A system, enhances the understanding and answering ability of vertical fields, and provides more professional and high-quality answer data that meets actual needs.
Smart Images

Figure CN119150998B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal RAG knowledge question-answering method and device applied to a vertical field. Background Art
[0002] Question answering system (QA), as an advanced version of information retrieval system, is committed to using precise and concise natural language to respond to questions raised by users in natural language. The rise of this research field is mainly due to people's urgent need to obtain the required information quickly and accurately.
[0003] The current mainstream domain question-answering system technologies are mainly divided into types based on question-answer pairs, knowledge graphs, and LLM technology. Since the question-answering system needs to establish a database of question-answer pairs, and in reality, FAQ data is relatively small and mostly unstructured knowledge texts, the cost of obtaining and maintaining high-quality question-answer pairs is very high. In the question-answering system based on the knowledge graph, inaccurate triple extraction may affect the accuracy of the system; the system has weak processing capabilities for multi-hop questions and has difficulties in domain migration and expansion due to the lack of structured data. The question-answering system based on the large language model has the following disadvantages: first, there is a lack of domain knowledge and new content because the model training data is static; second, it may produce "hallucination" phenomena, generating false or wrong information without providing accurate sources; third, the stability of the generated results is poor, and different ways of asking the same question may get different answers. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a multimodal RAG knowledge question-answering method and device applied to a vertical field to solve the problems of low retrieval recall and poor answer prediction quality in the current mainstream question-answering system.
[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0006] The first aspect of the present invention discloses a multimodal RAG knowledge question answering method applied to a vertical field, the method comprising:
[0007] Receive multiple multimodal data types, convert all multimodal data types into text type data using a preset conversion model and summarize them to obtain the text to be answered;
[0008] Retrieving the knowledge fragments corresponding to the text to be answered from the knowledge base using multiple preset retrieval models to obtain multiple ranked lists containing the text to be answered and multiple knowledge fragments; the knowledge base is pre-constructed based on question-answer pair data and plain text data;
[0009] Merge the knowledge fragments in all the sorted lists according to the inverse sorting fusion algorithm;
[0010] The target re-ranking model is used to calculate the corresponding score of each knowledge fragment in the fused list, and each knowledge fragment in the fused list is ranked based on the score to obtain a first ranked list; the target re-ranking model is obtained by fine-tuning and training an open source re-ranking model in advance using a preset vertical field data set;
[0011] Calculate the weight of each knowledge fragment in the first sorted list according to the preset knowledge frequency information, calculate the target score of each knowledge fragment in the first sorted list by combining the score and the weight, and sort them to obtain a second sorted list;
[0012] Extract k pieces of the knowledge from the second sorted list and perform permutations and combinations to obtain permutation results, and construct a prompt combination with each permutation result and the text to be answered to obtain k prompt combinations;
[0013] The target reward model is used to calculate the scores of all prompt combinations, and n prompt combinations are extracted in descending order according to the scores, and marked as target prompt combinations; the target reward model is obtained by training the open source reward model in advance using the preference data set of the preset LLM question and answer;
[0014] A large language model is used to generate corresponding answer data for each target prompt combination, and a target scoring model is used to calculate the scores of all answer data, and the answer data with the largest score is output; the target scoring model is obtained by pre-training a Transformer encoder-based scoring model using a preset question-answer pair training set.
[0015] Preferably, the process of constructing a knowledge base based on question-answer pair data and plain text data in advance includes the method further comprising:
[0016] Acquire multiple raw data and parse and process all the raw data to obtain multiple data to be stored in the database;
[0017] For each of the data to be stored, when the data to be stored is question-answer pair data, use a pre-trained target embedding model to vectorize the question data in the data to be stored to obtain the vectorized question data;
[0018] Combining the vectorized question data and its corresponding answer data and storing them in a knowledge base;
[0019] For each of the data to be stored, when the data to be stored is pure document data, the data to be stored is segmented using a semantic segmentation model to obtain a plurality of segmentation segments, and the document title of the data to be stored is obtained;
[0020] Splicing the document title with each of the segmented fragments to obtain multiple knowledge fragments;
[0021] Generate target question-answer pair data using the large language model and all knowledge fragments, as well as summary data corresponding to each of the knowledge fragments;
[0022] Vectorizing the question data in the target question-answer pair data using the target embedding model to obtain target question data;
[0023] Combining the target question data and its corresponding answer data and storing them in a knowledge base;
[0024] splicing the summary data corresponding to each of the knowledge fragments with the knowledge fragments to obtain a new knowledge fragment;
[0025] The new knowledge fragment is vectorized using the target embedding model, and the vectorized new knowledge fragment is stored in the knowledge base.
[0026] Preferably, the acquiring of multiple original data and parsing and processing all the original data to obtain multiple data to be stored in the database includes:
[0027] Get multiple raw data;
[0028] For each of the original data, when the type of the original data is a multimodal type other than a text type, convert the original data into text type data using a preset conversion model and mark it as text data;
[0029] Generate a link placeholder according to the storage location of the original data;
[0030] Inserting the link placeholder into the text information of the text data to obtain new text data;
[0031] Detecting whether the text information of the new text data contains professional terms according to a preset professional term library;
[0032] If included, detailed information of the professional term is obtained from the preset professional term library, and the detailed information is stored in the text information of the new text data to obtain the data to be stored in the library.
[0033] Preferably, after outputting the answer data with the largest score, the method further comprises:
[0034] Determine whether the answer data with the largest score meets the user's needs;
[0035] If it does not meet the user's needs, calculate the similarity between all question-answer pair data in the knowledge base and the text to be answered;
[0036] Sort all similarities in descending order, obtain the top m questions and display them to the user;
[0037] When a question selection instruction is received, answer data of the selected question contained in the question selection instruction is obtained and displayed to the user.
[0038] Preferably, the receiving of multiple multimodal type data, converting all multimodal type data into text type data using a preset conversion model and summarizing the data to obtain the text to be answered, includes:
[0039] Receiving multiple multimodal type data;
[0040] For each of the multimodal type data, detecting whether the multimodal type data is text type data;
[0041] If not, convert the multimodal type data into text type data using a preset conversion model to obtain a plurality of text type data;
[0042] Summarize all text type data to obtain the text to be answered.
[0043] Preferably, the step of fusing the knowledge fragments in all sorted lists according to the inverse sorting fusion algorithm includes:
[0044] For each knowledge fragment in all the sorted lists, obtaining the ranking of the knowledge fragment in the plurality of sorted lists;
[0045] Calculate the reciprocal of the sum of each ranking and a constant, and accumulate all the reciprocals to obtain the corresponding score of the knowledge fragment;
[0046] All knowledge fragments are sorted based on the corresponding score of each knowledge fragment to obtain a fused list.
[0047] Preferably, the weight of each knowledge fragment in the first sorted list is calculated according to the preset knowledge frequency information, and the target score of each knowledge fragment in the first sorted list is calculated and sorted in combination with the score and the weight to obtain a second sorted list, including:
[0048] For each knowledge segment in the first sorted list, obtaining a frequency corresponding to the current knowledge segment and a maximum frequency from preset knowledge frequency information, wherein the preset knowledge frequency information includes knowledge segments and frequencies corresponding to the knowledge segments;
[0049] Calculate the ratio of the frequency corresponding to the current knowledge fragment to the maximum frequency and multiply it by the weight base to obtain the weight of the current knowledge fragment;
[0050] Calculate the sum of the weight of the current knowledge fragment plus 1, and calculate the product of the sum and the score to obtain the target score of the current knowledge fragment;
[0051] The target scores of all knowledge fragments are sorted in descending order to obtain a second sorted list.
[0052] A second aspect of the present invention discloses a multimodal RAG knowledge question-answering device applied to a vertical field, the device comprising:
[0053] A summarizing unit is used to receive multiple multimodal data, convert all multimodal data into text data using a preset conversion model, and summarize them to obtain a text to be answered;
[0054] A retrieval unit, used to retrieve the knowledge fragments corresponding to the text to be answered from the knowledge base using multiple preset retrieval models, and obtain multiple sorted lists including the text to be answered and multiple knowledge fragments; the knowledge base is pre-constructed based on question-answer pair data and plain text data;
[0055] A fusion unit, used for fusing the knowledge fragments in all the sorted lists according to a reciprocal sorting fusion algorithm;
[0056] A first sorting unit is used to calculate the corresponding score of each knowledge fragment in the fused list by using a target re-ranking model, and sort each knowledge fragment in the fused list based on the score to obtain a first sorted list; the target re-ranking model is obtained by fine-tuning and training an open source re-ranking model in advance using a preset vertical field data set;
[0057] A second sorting unit, configured to calculate the weight of each knowledge fragment in the first sorting list according to preset knowledge frequency information, calculate the target score of each knowledge fragment in the first sorting list in combination with the score and the weight, and sort the scores to obtain a second sorting list;
[0058] A construction unit is used to extract k pieces of the knowledge fragments from the second sorted list, arrange and combine them to obtain an arrangement result, and each of the arrangement result and the text to be answered is used to construct a prompt combination to obtain k prompt combinations;
[0059] An extraction unit, used to calculate the scores of all prompt combinations using a target reward model, extract n prompt combinations in descending order according to the scores, and mark them as target prompt combinations; the target reward model is obtained by training an open source reward model in advance using a preset LLM question-answering preference data set;
[0060] An output unit is used to generate answer data corresponding to each target prompt combination using a large language model, and calculate the scores of all answer data using a target scoring model, and output the answer data with the largest score; the target scoring model is obtained by pre-training a scoring model based on a Transformer encoder using a preset question-answer pair training set.
[0061] Preferably, the device further comprises:
[0062] An acquisition unit is used to acquire multiple original data and parse and process all the original data to obtain multiple data to be stored;
[0063] A first vectorization processing unit is used for, for each of the data to be stored, when the data to be stored is question-answer pair data, using a pre-trained target embedding model to perform vectorization processing on the question data in the data to be stored, so as to obtain the vectorized question data;
[0064] A first storage unit, used to store the vectorized question data and its corresponding answer data in a knowledge base;
[0065] A segmentation unit, for segmenting each of the data to be stored in the warehouse using a semantic segmentation model to obtain a plurality of segmentation segments when the data to be stored in the warehouse is pure document data, and obtaining a document title of the data to be stored in the warehouse;
[0066] A first concatenation unit is used to concatenate the document title with each of the segmented segments to obtain a plurality of knowledge segments;
[0067] A generation unit, used to generate target question-answer pair data using a large language model and all knowledge fragments, as well as summary data corresponding to each of the knowledge fragments;
[0068] A second vectorization processing unit, configured to perform vectorization processing on the question data in the target question-answer pair data using the target embedding model to obtain target question data;
[0069] A second storage unit, used to store the target question data and its corresponding answer data in combination in a knowledge base;
[0070] A second concatenation unit is used to concatenate the summary data corresponding to each of the knowledge fragments with the summary data to obtain a new knowledge fragment;
[0071] The third storage unit is used to use the target embedding model to vectorize the new knowledge fragment and store the vectorized new knowledge fragment in the knowledge base.
[0072] Preferably, the acquisition unit includes:
[0073] An acquisition module, used for acquiring multiple raw data;
[0074] A conversion module, for each of the original data, when the type of the original data is a multimodal type other than a text type, using a preset conversion model to convert the original data into text type data, and mark it as text data;
[0075] A generating module, used for generating a link placeholder according to the storage location of the original data;
[0076] An inserting module, used for inserting the link placeholder into the text information of the text data to obtain new text data;
[0077] A detection module, used to detect whether the text information of the new text data contains professional terms according to a preset professional term library;
[0078] The storage module is used to obtain detailed information of the professional term from the preset professional term library if included, and store the detailed information in the text information of the new text data to obtain the data to be stored in the library.
[0079] Based on the above-mentioned embodiment of the present invention, a multimodal RAG knowledge question-answering method and device applied to a vertical field is provided, which relates to the field of artificial intelligence technology. The present invention supports multimodal input of text to be answered, including text, pictures, audio and video, and makes full use of the semantic information of multimodal data. A plurality of preset retrieval models are used to retrieve the knowledge fragments corresponding to the text to be answered, and a plurality of retrieval results are obtained. The multiple retrieval results are fused by the RRF algorithm, and the scores are calculated by the target re-ranking model for re-ranking, thereby improving the recall rate. Based on the retrieval content and the preset knowledge frequency information, the sorting is further fine-tuned to make it more accurate and meet user expectations. Combining the natural language understanding ability of the large language model and the target reward model, as well as the relevant information of RAG recall, the optimal answer data is generated, and the target scoring model and domain expert knowledge are used to provide more professional and practical high-quality answer data. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0081] Figure 1 A flowchart of a multimodal RAG knowledge question-answering method applied to a vertical field provided by an embodiment of the present invention;
[0082] Figure 2 A flowchart of fine-tuning an open source model based on vertical field data provided in an embodiment of the present invention;
[0083] Figure 3 A schematic diagram of the contents of the overall solution provided by an embodiment of the present invention;
[0084] Figure 4 A schematic diagram of a multimodal input module provided by an embodiment of the present invention;
[0085] Figure 5 A flowchart of building a knowledge base based on question-answer pair data and plain text data provided by an embodiment of the present invention;
[0086] Figure 6 A schematic diagram of a reciprocal sorting fusion and rearrangement module provided in an embodiment of the present invention;
[0087] Figure 7 A schematic diagram of a weight updating and rearrangement module provided in an embodiment of the present invention;
[0088] Figure 8 A flowchart of determining a second sorting list according to preset knowledge frequency information and a first sorting list provided in an embodiment of the present invention;
[0089] Fig. 9 A flow chart of optimizing answer data provided by an embodiment of the present invention;
[0090] Fig.10 A schematic diagram of a knowledge base construction and knowledge enhancement module provided in an embodiment of the present invention;
[0091] Fig.11 A flowchart of parsing and processing raw data to obtain data to be stored provided by an embodiment of the present invention;
[0092] Fig.12 Another schematic diagram of the overall solution provided by an embodiment of the present invention;
[0093] Fig.13A structural block diagram of a multimodal RAG knowledge question-and-answer device applied to a vertical field provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0094] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0095] In this application, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.
[0096] From the background technology, we can see that the question-answering system based on knowledge graph has the following shortcomings: first, the triple extraction effect directly affects the accuracy of the system. If the relationship extraction is inaccurate, the quality of question answering may be reduced; second, the semantic information of the relationship between entities is insufficiently utilized, resulting in weak multi-hop problem processing capabilities; third, structured data is scarce in reality and most of it is unstructured text, which makes the extraction and knowledge fusion process complicated, making the system's domain portability and scalability poor.
[0097] Therefore, an embodiment of the present invention provides a multimodal RAG knowledge question-answering method and device applied to a vertical field, which relates to the field of artificial intelligence technology. The present invention supports multimodal input of text to be answered, including text, pictures, audio and video, and makes full use of the semantic information of multimodal data. A plurality of preset retrieval models are used to retrieve the knowledge fragments corresponding to the text to be answered, and a plurality of retrieval results are obtained. The multiple retrieval results are fused by the RRF algorithm, and the scores are calculated by the target re-ranking model for re-ranking, thereby improving the recall rate. Based on the retrieval content and the preset knowledge frequency information, the sorting is further fine-tuned to make it more accurate and meet user expectations. Combining the natural language understanding ability of the large language model and the target reward model, as well as the relevant information of RAG recall, the optimal answer data is generated, and the target scoring model and domain expert knowledge are used to provide more professional and practical high-quality answer data.
[0098] See also Figure 1 , showing a flowchart of a multimodal RAG knowledge question-answering method applied to a vertical field provided by an embodiment of the present invention.
[0099] It should be noted that multiple models will be used in this method. Before introducing this method, the training process of these models is introduced first:
[0100] It is understandable that in the field of knowledge question answering based on large language models, commonly used open source embedding models, reranker models, reward models and scoring models are usually trained on general domain data, which leads to insufficient understanding of their professional knowledge in vertical fields (such as the communication field), making it difficult to achieve the best results. Therefore, the present invention fine-tunes the open source model based on vertical domain data to obtain a target embedding model, target reranking model, target reward model and target scoring model in the vertical field.
[0101] See also Figure 2 , the specific fine-tuning training process includes:
[0102] Step S201: Obtain a preset vertical field dataset, a preset LLM question and answer preference dataset, and a preset question and answer pair training set.
[0103] It should be noted that the preset question-answer pair training set includes correct question-answer pair data and incorrect question-answer pair data (Frequently Asked Questions, FAQ data).
[0104] Step S202: Use the preset vertical field dataset to fine-tune the open source embedding model and the open source re-ranking model to obtain the target embedding model and the target re-ranking model.
[0105] First, the process of fine-tuning the open source embedding model using the preset vertical field dataset is as follows:
[0106] Preset vertical field datasets, for example: {"query": str, "pos": List[str], "neg": List[str]}. "query" is the question, "pos" is the correct answer, and "neg" is an irrelevant answer. Based on this, fine-tune the open source embedding model, for example, fine-tune the bge-large-zh-v1.5 model to obtain the target embedding model.
[0107] It is understandable that fine-tuning the open source embedding model using a preset vertical field dataset to obtain a target embedding model can help better capture the relationships and semantics of specific terms in the vertical field.
[0108] Secondly, the process of fine-tuning the open source reranking model using the preset vertical field dataset is to fine-tune the open source reranking model (such as the bge-reranker-large model) using the above-mentioned preset vertical field dataset to obtain the target reranking model.
[0109] It is understandable that fine-tuning the open source reranking model using a preset vertical field dataset to obtain a target reranking model can effectively improve the relevance and accuracy of knowledge retrieval.
[0110] Step S203: Use the preset LLM question-answering preference data set to train the open source reward model to obtain a target reward model based on the large language model.
[0111] Understandably, the preferred datasets for LLM question answering are preset, for example:
[0112] The input is: "Please answer the user's question based on the following knowledge snippet.
[0113] Knowledge fragment: {content1}{content2}{content3};
[0114] User question: {query}";
[0115] The output is: [correct answer, irrelevant answer].
[0116] Based on this, the open source reward model is trained to obtain a target reward model based on a large language model (LLM model).
[0117] It is understandable that training the open source reward model to obtain the target reward model can better capture the direct correlation between the problem and the relevant knowledge fragments.
[0118] Step S204: Use the preset question-answer pair training set to train the Transformer encoder-based scoring model to obtain a target scoring model.
[0119] In the specific implementation process of step S204, a preset question-answer pair training set including correct question-answer pair data and incorrect question-answer pair data is used to train a scoring model based on a Transformer encoder to obtain a target scoring model.
[0120] It can be understood that by using the preset question-answer pair training set to train the Transformer encoder-based scoring model, a target scoring model is obtained, which effectively improves the quality of the large language model's answers.
[0121] It is understandable that by fine-tuning open source embedding models, reranker models, reward models, and scoring models based on vertical domain data, the accuracy of the model can be significantly improved, allowing it to more accurately capture domain-specific knowledge and patterns, thereby improving the overall accuracy of the task. In addition, fine-tuning can enhance the relevance of the model in specific application scenarios, making it more effective in responding to domain-related issues. At the same time, this process can also reduce the generalization error of the model in specific tasks and improve its robustness in practical applications. Ultimately, fine-tuning not only speeds up the model's adaptation to actual needs, but also saves the time and resources required for training from scratch.
[0122] So far, the above content is Figure 3 All the contents of the "S1 vertical domain model training module" in the content diagram of the overall solution of this application are shown.
[0123] Combination Figure 2 As shown in the content, the embodiment of the present invention is trained on vertical field data, so that the model can more accurately understand the field-specific language and knowledge, thereby improving the accuracy and professionalism of the question-answering system. By combining the knowledge and evaluation criteria of field experts, the system can generate more professional answers that meet the needs of actual applications, further enhancing its professionalism and practicality. Through customized reward models and scoring models, adjustments can be made according to the needs of different fields and tasks to improve the applicability and scalability of the model.
[0124] The following is a detailed description of the specific content of the multimodal RAG knowledge question-answering method applied to vertical fields:
[0125] Step S101: receiving a plurality of multimodal type data, converting all the multimodal type data into text type data using a preset conversion model and summarizing the data to obtain a text to be answered.
[0126] It should be noted that multimodal data (such as text, pictures, audio, and video) can provide richer information and context. Receiving multiple multimodal types of data input by users can more accurately understand the user's intentions, and can make full use of the semantic information expression capabilities carried by multimodal data to achieve more complex and natural interaction methods, while also improving the user experience.
[0127] In step S101, multiple multimodal data are received, the multimodal data are parsed, and all multimodal data are converted into text data using a preset conversion model and summarized to obtain a text to be answered. Figure 4 The schematic diagram of the multimodal input module is shown, and the specific implementation process is as follows (process A1-process A4):
[0128] Process A1: receiving a plurality of multi-modal type data.
[0129] When specifically implementing process A1, multiple multimodal type data input by the user is received through the input interface.
[0130] Process A2: for each multimodal type data, detect whether the multimodal type data is text type data.
[0131] When specifically implementing process A2, for all received multimodal type data, determine whether each multimodal type data is text type data. If it is text type data, do not operate on it. If it is not text type data, execute process A3.
[0132] Process A3: If not, the multimodal type data is converted into text type data using a preset conversion model to obtain a plurality of text type data.
[0133] In the specific implementation process A3, if the current multimodal type data is not text type data, the current multimodal type data is converted into text type data using a preset conversion model corresponding to the current multimodal type data to obtain a plurality of text type data.
[0134] For example: If the current multimodal data type is audio data, then use Figure 4 The tool 1 shown (such as the Whisper model) converts the current audio type data into text type data; if the current multimodal type data is image type data, use Figure 4 The tool 2 shown (such as Clip model and OCR model) converts the current image type data into text type data; if the current multimodal type data is video type data, use Figure 4 The tool 3 shown (such as the VideoBert model) converts the current video type data into text type data.
[0135] Process A4: Summarize all text type data to obtain the text to be answered.
[0136] In the specific implementation process A4, all text type data are summarized to obtain the text to be answered, that is, Figure 4 "Question (summary text)" shown.
[0137] So far, the content of step S101 is Figure 3 All contents of the "S3 multimodal input module" in the content diagram of the overall solution of this application are shown.
[0138] Step S102: using a plurality of preset retrieval models to retrieve knowledge fragments corresponding to the text to be answered from the knowledge base, and obtaining a plurality of sorted lists including the text to be answered and a plurality of knowledge fragments.
[0139] It is understandable that the knowledge base is pre-built based on question-answer data and plain text data. For the specific construction process, see Figure 5 The flowchart of building a knowledge base based on question-answer pair data and plain text data is shown, wherein: Figure 5 The content shown is Figure 3 All the contents of the "S2 knowledge base construction and knowledge enhancement module" in the content diagram of the overall solution of this application are shown.
[0140] It should be noted that using multiple preset retrieval models to retrieve knowledge fragments corresponding to the text to be answered from the knowledge base can increase the retrieval coverage and ensure that more knowledge fragments can be retrieved, and different retrieval methods can complement each other to improve robustness.
[0141] In the specific implementation of step S102, multiple preset retrieval models are used to retrieve multiple knowledge fragments corresponding to the text to be answered from the knowledge base, and each preset retrieval model obtains a sorted list, each of which contains the text to be answered and multiple knowledge fragments corresponding to the text to be answered.
[0142] Combination Figure 6 The schematic diagram of the inverse sorting fusion and rearrangement module shown in FIG. 1 is as follows: (1) extracting keywords from the text to be answered, using the keyword retrieval model, retrieving the knowledge fragments corresponding to the keywords from the knowledge base, and obtaining a sorted list (such as Figure 6 The sorted list shown in 1); (2) using Figure 2 The trained target embedding model vectorizes the text to be answered, and then calculates the vector similarity in the corresponding vector knowledge base to obtain a sorted list from high to low similarity (such as Figure 6 (3) Use an open source embedding model (such as the bge-large-zh-v1.5 model) to vectorize the text to be answered, and then calculate the vector similarity in the corresponding vector knowledge base to obtain a sorted list from high to low similarity (such as Figure 6 Sorted list shown in 3).
[0143] Step S103: Merge all the knowledge fragments in the sorted list according to the inverse sorting fusion algorithm.
[0144] It should be noted that for the retrieval system, the best ranking is usually achieved by combining multiple search results. For example, the results of the vector search can be combined with the results of the scalar search. However, in actual operation, it is quite challenging to combine multiple search results into a final ranking. The main difficulty lies in the fact that each retrieval method uses different scoring criteria, and simply adding up the scores of each path often cannot get the ideal final ranking. Therefore, in the embodiment of the present invention, the reciprocal rank fusion algorithm (RRF) is applied to combine the rankings of knowledge fragments in different sorting lists into a final ranking.
[0145] It can be understood that the basic principle of the Reciprocal Rank Fusion (RRF) algorithm is that knowledge fragments that consistently appear at the top positions of different ranked lists are likely to be more relevant and should therefore receive a higher ranking in the merged result. The advantage of RRF is that it does not rely on the absolute scores assigned by a single retrieval method, but is based on relative rankings, which makes it very suitable for combining results from different retrieval methods that may have different score scales or distributions. By merging ranked lists from different retrieval methods, RRF increases the chances that the most relevant knowledge fragments will appear at the front of the final ranked list.
[0146] In step S103, Figure 6 The process of merging all sorted lists into one sorted list is as follows (process B1-process B3):
[0147] Process B1: For each knowledge fragment in all sorted lists, obtain the ranking of the knowledge fragment in multiple sorted lists.
[0148] For example: knowledge fragment 1 only appears in sorted list 1, and its ranking is 3; knowledge fragment 2 is in sorted list 1, and its ranking is 2, and knowledge fragment 2 is in sorted list 2, and its ranking is 3; knowledge fragment 3 is in sorted list 2, and its ranking is 1, and knowledge fragment 3 is in sorted list 3, and its ranking is 2.
[0149] Process B2: Calculate the reciprocal of the sum of each ranking and a constant, and accumulate all reciprocals to obtain the corresponding score of the knowledge fragment.
[0150] When implementing process B2, the corresponding score of each knowledge fragment is calculated according to formula (1) of the reciprocal ranking fusion (RRF) algorithm.
[0151] (1)
[0152] Wherein, d is a knowledge fragment; D is a set of knowledge fragments; R is a set of multiple sorted lists in step S102; k is a constant, usually set to 60 by default; and r(d) represents the ranking of the knowledge fragment.
[0153] Combined with the example in process B1: the score of knowledge fragment 1 is =0.016; the score of knowledge fragment 2 is =0.032; the score of knowledge fragment 3 is =0.0325.
[0154] Process B3: Sort all knowledge fragments based on the corresponding score of each knowledge fragment to obtain a fused list.
[0155] In the specific implementation process B3, all knowledge fragments are sorted according to the corresponding score of each knowledge fragment to obtain a fused list.
[0156] Step S104: Calculate the corresponding score of each knowledge fragment in the fused list using the target re-ranking model, and rank each knowledge fragment in the fused list based on the score to obtain a first ranked list.
[0157] It is understandable that the target re-ranking model is obtained by fine-tuning the open source re-ranking model using the preset vertical field dataset. The specific training process is described in Figure 2 The details are described in detail in , and will not be repeated here.
[0158] In the specific implementation of step S104, for each knowledge fragment in the fused list, the target re-ranking model is used to calculate the corresponding scores of each knowledge fragment and the text to be answered, and each knowledge fragment in the fused list is re-ranked based on the score to obtain a first ranked list.
[0159] For example, use the target reranking model or the open source reranker models (such as the bge-reranker-large model and the bge-reranker-v2-m3 model).
[0160] Another example: the fused list includes the text to be answered, knowledge fragment 1, knowledge fragment 2, and knowledge fragment 3. The text to be answered is combined with knowledge fragment 1, and the score 1 of the text to be answered-knowledge fragment 1 is calculated using the target re-ranking model; the text to be answered is combined with knowledge fragment 2, and the score 2 of the text to be answered-knowledge fragment 2 is calculated using the target re-ranking model; the text to be answered is combined with knowledge fragment 3, and the score 3 of the text to be answered-knowledge fragment 3 is calculated using the target re-ranking model; finally, knowledge fragment 1, knowledge fragment 2, and knowledge fragment 3 are re-ranked based on score 1, score 2, and score 3 to obtain the first ranked list.
[0161] It should be noted that the main advantage of using the target re-ranking model to re-rank the list after the RRF algorithm is that it can further optimize the ranking results. The target re-ranking model can deeply analyze the complex relationship between knowledge fragments and the text to be answered to more accurately evaluate the relevance of knowledge fragments, thereby improving the final ranking quality. By utilizing more complex features and models, this re-ranking can improve the accuracy of the first ranking list.
[0162] It can be understood that in the embodiment of the invention, a reciprocal ranking fusion (RRF) algorithm is used to fuse multiple search results, and then the scores are calculated using a rearrangement model to rearrange them, thereby further improving the recall rate.
[0163] So far, the contents of step S102 to step S104 are as follows: Figure 3 All the contents of the "S4 inverse sorting fusion retrieval and rearrangement module" in the content diagram of the overall solution of this application are shown.
[0164] Step S105: Calculate the weight of each knowledge fragment in the first sorted list according to the preset knowledge frequency information, calculate the target score of each knowledge fragment in the first sorted list by combining the score and the weight, and sort them to obtain a second sorted list.
[0165] It should be noted that in the process of actual business development, as time accumulates, the frequency of users' attention to each knowledge fragment can be learned. Generally speaking, if the user's attention frequency to a certain knowledge fragment is high, it means that the user pays more attention to this type of content, and if the user's attention frequency to a certain knowledge fragment is low, it means that the user is relatively uninterested in this type of content. Based on this, an embodiment of the present invention pre-acquires the user's attention frequency to multiple knowledge fragments to obtain preset knowledge frequency information.
[0166] In step S105, the weight of each knowledge fragment in the first sorting list is calculated in combination with the preset knowledge frequency information, and then the target score of each knowledge fragment in the first sorting list is calculated and sorted according to the score of each knowledge fragment in step S104 and the weight in step S105 to obtain a second sorting list. Figure 7 The weight update rearrangement module schematic diagram is shown in FIG. Figure 8 As shown:
[0167] Step S801: For each knowledge segment in the first sorting list, the frequency corresponding to the current knowledge segment and the maximum frequency are obtained from the preset knowledge frequency information.
[0168] It should be noted that the preset knowledge frequency information includes knowledge fragments and the frequencies corresponding to the knowledge fragments.
[0169] In the specific implementation of step S801, for each knowledge fragment in the first sorted list, the frequency of each knowledge fragment in the preset knowledge frequency information and the maximum frequency in the preset knowledge frequency information are obtained.
[0170] Step S802: Calculate the ratio of the frequency corresponding to the current knowledge fragment to the maximum frequency and multiply it by the weight base to obtain the weight of the current knowledge fragment.
[0171] In the specific implementation of step S802, the weight of the knowledge fragment is calculated according to formula (2). Specifically, for each knowledge fragment, the ratio of the frequency corresponding to the current knowledge fragment to the maximum frequency is calculated and multiplied by the weight base to obtain the weight of the current knowledge fragment.
[0172] (2)
[0173] Among them, w i is the weight, f i is the frequency corresponding to the current knowledge fragment, the denominator max is the maximum frequency of all knowledge fragments in the first sorted list, and α is an adjustable weight base.
[0174] Step S803: Calculate the sum of the weight of the current knowledge segment plus 1, and calculate the product of the sum and the score of the current knowledge segment to obtain the target score of the current knowledge segment.
[0175] In the specific implementation of step S803, the target score of the current knowledge segment is calculated according to formula (3). Specifically, for each knowledge segment, the sum of the weight of the current knowledge segment plus 1 is multiplied by the score of the current knowledge segment calculated in step S104 to obtain the target score of the current knowledge segment.
[0176] (3)
[0177] Among them, score old is the score calculated using the target re-ranking model in step S104; i is the weight; score new is the target score.
[0178] Step S804: sort the target scores of all knowledge fragments in descending order to obtain a second sorted list.
[0179] In the specific implementation of step S804, all knowledge fragments are sorted in descending order according to their target scores to obtain a second sorting list.
[0180] It is understandable that by fine-tuning the first sorting list in combination with the preset knowledge frequency information, the accuracy of the sorting can be significantly improved. This fine-tuning process optimizes the sorting algorithm by analyzing and utilizing user interaction data, so that the sorting results are more likely to hit content related to the text to be answered, which is more in line with user expectations. Ultimately, this adjustment can better meet user expectations and needs, provide more relevant and useful information, and thus enhance user experience.
[0181] So far, the content of step S105 is Figure 3 All the contents of the "S5 weight update and rearrangement module" in the content diagram of the overall solution of this application are shown.
[0182] Step S106: extract k knowledge fragments from the second sorted list and perform permutations and combinations to obtain permutation results, and construct a prompt combination with each permutation result and the text to be answered to obtain k prompt combinations.
[0183] In the process of implementing step S106, k knowledge fragments are extracted from the second sorted list for permutation and combination. For example, assuming that there are 5 knowledge fragments in the second sorted list: knowledge fragment 1, knowledge fragment 2, knowledge fragment 3, knowledge fragment 4 and knowledge fragment 5, the first 3 knowledge fragments are extracted from these 5 knowledge fragments and 2 knowledge fragments are respectively permuted and combined. The obtained permutation results are as follows: {knowledge fragment 1, knowledge fragment 2}; {knowledge fragment 1, knowledge fragment 3}; {knowledge fragment 2, knowledge fragment 1}; {knowledge fragment 2, knowledge fragment 3}; {knowledge fragment 3, knowledge fragment 1}; {knowledge fragment 3, knowledge fragment 2}.
[0184] The six permutation results are combined with the text to be answered to obtain six prompt combinations (such as prompt). For example, one of the prompt combinations prompt is:
[0185] “Please answer the user’s question based on the following knowledge snippet.
[0186] Knowledge fragment: {content1}{content2};
[0187] User question: {query}".
[0188] Another example:
[0189] “Please answer the user’s question based on the following knowledge snippet.
[0190] Knowledge fragment: {content2}{content3};
[0191] User question: {query}".
[0192] Step S107: Calculate the scores of all prompt combinations using the target reward model, extract n prompt combinations in descending order of scores, and mark them as target prompt combinations.
[0193] In the specific implementation of step S107, the target reward model is used to calculate the score of each prompt combination among all prompt combinations, and the prompt combinations are sorted in descending order according to the scores, and n (such as top 2) prompt combinations are extracted from the sorted prompt combinations and marked as target prompt combinations.
[0194] It is understandable that the target reward model is obtained by pre-training the open source reward model using the preset LLM question-answering preference dataset. The specific training process is described in Figure 2 The details are described in detail in , and will not be repeated here.
[0195] Step S108: Generate corresponding answer data for each target prompt combination using the large language model, calculate the scores of all answer data using the target scoring model, and output the answer data with the largest score.
[0196] It is understandable that the target scoring model is obtained by training the Transformer encoder-based scoring model using the preset question-answer pair training set. The specific training process is described in Figure 2 The details are described in detail in , and will not be repeated here.
[0197] In the specific implementation of step S108, a large language model (LLM) is used to generate corresponding answer data for each target prompt combination, and a target scoring model is used to calculate the scores of all answer data, and the answer data with the largest score is output.
[0198] It should be noted that each answer data is combined with the text to be answered, and the score of each combination of answer data and the text to be answered is calculated using the target scoring model to obtain the scores of all answer data.
[0199] For example: there are a total of 6 target prompt combinations, and the LLM model is used to generate the answer data of these 6 target prompt combinations. Each answer data is combined with the text to be answered, text to be answered-answer data 1, text to be answered-answer data 2, text to be answered-answer data 3, text to be answered-answer data 4, text to be answered-answer data 5 and text to be answered-answer data 6; then the target scoring model is used to calculate the score of text to be answered-answer data 1, the score of text to be answered-answer data 2, the score of text to be answered-answer data 3, the score of text to be answered-answer data 4, the score of text to be answered-answer data 5 and the score of text to be answered-answer data 6, and finally the answer data corresponding to the largest score among all the scores is output.
[0200] It is understandable that using the target scoring model to score the answer data generated by the large language model and selecting the answer data with the highest score provides an objective evaluation mechanism to ensure that subjective bias is reduced in the scoring process, making the quality assessment of the answer data more fair. Secondly, consistency is improved through unified scoring standards, so that the evaluation of different answer data is consistent in standards, thereby ensuring the reliability of the evaluation results. In addition, the efficient screening ability of the target scoring model can quickly find the best option when faced with a large amount of answer data, greatly saving the time and cost of manual review. Finally, the target scoring model comprehensively considers multiple factors such as the accuracy and relevance of the answer data, so that it can select the answer that best meets actual needs, optimize the decision-making process and improve overall efficiency.
[0201] Combination Fig. 9 The flowchart for optimizing answer data is shown. In some specific embodiments, in order to further improve user satisfaction, after outputting the answer data with the highest score to the user, the following operations are performed to improve user experience.
[0202] Step S901: Determine whether the answer data with the largest score meets the user's needs.
[0203] In the process of implementing step S901, a preset feedback collection method is used to determine whether the answer data with the highest score meets the user's needs. If the answer data with the highest score meets the user's needs, the process ends; if the answer data with the highest score does not meet the user's needs, step S902 is executed.
[0204] For example: the satisfaction of mobile phone users with the answer data can be measured through a questionnaire survey method or an online scoring system. If the satisfaction is higher than or equal to a threshold, it indicates that the answer data with the largest score meets the user's needs. If the satisfaction is lower than the threshold, it indicates that the answer data with the largest score does not meet the user's needs.
[0205] Step S902: If it does not meet the user's requirements, calculate the similarity between all question-answer pair data in the knowledge base and the text to be answered.
[0206] In the specific implementation of step S902, if the answer data with the largest score does not meet the user's needs, the similarity between all the question-answer pair data in the knowledge base and the text to be answered is calculated.
[0207] Step S903: sort all similarities in descending order, obtain the first m questions and display them to the user.
[0208] In the specific implementation of step S903, all question-answer pair data are sorted in descending order of similarity with the text to be answered, the questions corresponding to the top m similarities in the sorting result are obtained, and these m questions are displayed to the user.
[0209] Step S904: When a question selection instruction is received, the answer data of the selected question contained in the question selection instruction is obtained and displayed to the user.
[0210] In the specific implementation of step S904, when a question selection instruction triggered by a user is received, the answer data of the selected question (ie, any one of the m questions) contained in the question selection instruction is obtained from the question-answer pair data of the knowledge base and displayed to the user.
[0211] So far, the contents of step S106 to step S108 are as follows: Figure 3 All the contents of the "S6 reward model and scoring model assisted LLM reasoning module" in the content diagram of the overall solution of this application are shown.
[0212] It can be understood that the embodiment of the present invention introduces a target reward model based on the powerful natural language understanding ability of LLM, and combines the second ranking list of RAG recall in the previous steps to obtain the best knowledge fragments to associate the text to be answered. At the same time, the target scoring model is used to integrate the knowledge and evaluation criteria of domain experts, so that LLM can generate more professional and high-quality enlarged data that meets actual needs. In addition, this process can also update the question-answer data in the knowledge base and gradually improve its comprehensiveness.
[0213] In an embodiment of the present invention, the explanation of professional terms in vertical fields is expanded, and the system's understanding of professional knowledge is enhanced. It supports multimodal input, including text, pictures, audio and video, so as to make full use of the semantic information of multimodal data. The large language model is used to generate question-answer pair data for pure knowledge documents, thereby expanding the scarce question-answer pair data. The present invention also generates a summary of knowledge fragments to improve its expressive power. Multiple retrieval results are fused through the RRF algorithm, and the scores are calculated using the rearrangement model for rearrangement, thereby improving the recall rate of relevant content. Based on the retrieval content and the real-time knowledge frequency information of the system, the sorting is further fine-tuned to make it more accurate and in line with user expectations. Combined with the natural language understanding ability of the large language model and the target reward model, as well as the relevant information of RAG recall, the optimal answer data is generated, and the target scoring model and domain expert knowledge are used to provide more professional and practical high-quality answer data.
[0214] The above embodiments of the present invention Figure 1 For the specific implementation process of building a knowledge base based on question-answer data and plain text data, see Figure 5 The flowchart of building a knowledge base based on question-answer data and plain text data is shown, and the combination Fig.10 The schematic diagram of knowledge base construction and knowledge enhancement module is shown in FIG. The specific process of constructing the knowledge base is as follows:
[0215] Step S501: Acquire multiple original data and parse and process all the original data to obtain multiple data to be stored.
[0216] In step S501, the specific analysis process is as follows: Fig.11 shown.
[0217] Step S1101: Acquire multiple original data.
[0218] In the specific implementation of step S1101, a plurality of original data (including text type data, picture type data, audio type data, video type data, etc.) are obtained, for example, from various knowledge materials of a communication operator.
[0219] It should be noted that the data involved in this application (including but not limited to original data, data obtained from various knowledge materials of communication operators) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0220] Step S1102: for each piece of original data, when the type of the original data is a multimodal type other than a text type, the original data is converted into text type data using a preset conversion model and marked as text data.
[0221] In the specific implementation of step S1102, for each original data among all the original data, when the type of the original data is a multimodal type other than a text type, the original data is converted into text type data using a preset conversion model, for example:
[0222] Use the Whisper model to convert audio type data into text type data, use the Clip model and OCR model to convert image type data into text type data, and use the VideoBert model to convert video type data into text type data; finally, mark these data as text data.
[0223] Step S1103: Generate a link placeholder according to the storage location of the original data.
[0224] In the specific process of implementing step S1103, for each original data among all the original data, the storage location of the original data is obtained, and a plurality of link placeholders are generated.
[0225] Step S1104: inserting a link placeholder into the text information of the text data to obtain new text data.
[0226] In the specific implementation of step S1104, each link placeholder is inserted into the text information of the corresponding text data to obtain new text data.
[0227] It can be understood that by linking placeholders, each text data can be linked to the multimodal original data, which helps the system trace its origin.
[0228] Step S1105: Detect whether the text information of the new text data contains professional terms according to the preset professional term library.
[0229] It is understandable that the preset professional terminology library contains multiple professional terms and detailed information corresponding to each professional data.
[0230] In the process of implementing step S1105, it is detected whether the text information of the new text data contains the professional terms in the preset professional terminology library. If the text information of the new text data does not contain the professional terms in the preset professional terminology library, the process is terminated; if the text information of the new text data contains the professional terms in the preset professional terminology library, step S1106 is executed.
[0231] Step S1106: If included, detailed information of the professional term is obtained from a preset professional term library, and the detailed information is stored in the text information of the new text data to obtain the data to be stored.
[0232] In the specific implementation of step S1106, if the text information of the new text data contains professional terms in the preset professional terminology library, detailed information of the professional terms is obtained from the preset professional terminology library and stored in the text information of the new text data to obtain data to be stored.
[0233] It should be noted that after the original data is parsed and processed according to the above content to obtain the data to be stored, different storage processing is required according to the type of the data to be stored. The specific processing process is shown in the following content.
[0234] Step S502: For each data to be stored, when the data to be stored is question-answer pair data, the question data in the data to be stored is vectorized using a pre-trained target embedding model to obtain the vectorized question data.
[0235] In the process of implementing step S502, for each data to be stored in the warehouse, the type of the data to be stored is determined. When the data to be stored in the warehouse is question-answer pair data, the question data in the data to be stored is vectorized using a pre-trained target embedding model to obtain the vectorized question data.
[0236] It should be noted that in addition to using the pre-trained target embedding model to vectorize the problem data in the incoming data, you can also use the open source bge-large-zh-v1.5 model, bge-m3 model, multilingual-e5-large model, etc. to vectorize the problem data in the incoming data.
[0237] It should be noted that only the question data in the question-answer pair data is vectorized, and the answer data in the question-answer pair data is not vectorized. This is because the question data is the key to retrieving the relevant answer data. By vectorizing the question data, the most similar answer data can be found in the vector space.
[0238] Step S503: The vectorized question data and its corresponding answer data are combined and stored in the knowledge base.
[0239] In the specific implementation of step S503, the vectorized question data and its corresponding answer data are combined and stored in the knowledge base, that is, the question-answer pair data in the knowledge base.
[0240] Step S504: for each data to be stored, when the data to be stored is pure document data, the data to be stored is segmented using a semantic segmentation model to obtain a plurality of segmentation segments, and the document title of the data to be stored is obtained.
[0241] In the specific implementation of step S504, for each data to be stored, the type of the data to be stored is determined. When the data to be stored is pure document data, the semantic segmentation model is used to segment the data to be stored to obtain multiple segmented segments; at the same time, the document title of the data to be stored is obtained.
[0242] It should be noted that the purpose of using the semantic segmentation model to segment the incoming data is to prevent the text of the pure document data from being too long and exceeding the maximum length supported by the vector model.
[0243] Step S505: Connect the document title with each segmented fragment to obtain multiple knowledge fragments.
[0244] For example, the document title is concatenated in front of the segmented fragment to obtain multiple knowledge fragments, such as document title-segmented fragment.
[0245] It should be noted that such a splicing method helps to quickly identify the pure document data to which each knowledge fragment belongs through the document title.
[0246] Step S506: Generate target question-answer pair data and corresponding summary data for each knowledge fragment using the large language model and all knowledge fragments.
[0247] In the specific implementation of step S506, on the one hand, the large language model and each knowledge fragment are used to generate some target question-answer pair data, and on the other hand, the large language model is used to generate corresponding summary data for each knowledge fragment.
[0248] Step S507: Use the target embedding model to vectorize the question data in the target question-answer pair data to obtain target question data.
[0249] In the specific implementation process of step S507, for the target question-answer pair data generated by using the large language model and each knowledge fragment, the target embedding model is used to vectorize the question data in the target question-answer pair data to obtain the target question data.
[0250] Step S508: Combine the target question data and its corresponding answer data and store them in the knowledge base.
[0251] In the specific implementation of step S508, the target question data and its corresponding answer data are combined and stored in the knowledge base.
[0252] It should be noted that the specific implementation of step S507 and step S508 is the same as the implementation of step S502 and step S503, and will not be repeated here.
[0253] Step S509: concatenate the summary data corresponding to each knowledge fragment with it to obtain a new knowledge fragment.
[0254] For example, concatenate the summary data title in front of the knowledge fragment to obtain multiple new knowledge fragments, such as summary data-document title-segmented fragments.
[0255] Step S510: vectorize the new knowledge fragments using the target embedding model, and store the vectorized new knowledge fragments into the knowledge base.
[0256] In the specific implementation of step S510, the new knowledge fragments are vectorized using the target embedding model, and then the vectorized knowledge fragments are stored in the knowledge base.
[0257] It should be noted that since the knowledge base includes question-answer pair data and plain text data, it includes a question-answer pair library and a plain text library.
[0258] Therefore, combined with Fig.12 Another schematic diagram of the overall solution provided by an embodiment of the present invention is shown. In actual application, after executing step S101 to obtain the text to be answered, steps S102 to S107 are executed based on the question-answer pair library and the plain text library respectively using the text to be answered, and in step S108, a large language model (LLM) is used to generate corresponding answer data for each target prompt combination in the question-answer pair library and corresponding answer data for each target prompt combination in the plain text library, and the target scoring model is used to calculate the scores of all answer data, and the answer data with the largest score is output.
[0259] In the embodiments of the present invention, the problem that the existing system fails to efficiently utilize raw data information is solved by enhancing the knowledge processing method. Specifically, the present invention supports the integration of multimodal data and the collection of comprehensive information; expands the interpretation of professional terms in vertical fields and improves the system's ability to understand professional knowledge; generates question-answer pair data from pure knowledge documents using a large language model (LLM), which makes up for the problem of scarcity of question-answer pair data, thereby improving the difficulty of the existing FAQ system in obtaining question-answer pair data; and enhances the expressive power of knowledge fragments by generating summary abstracts of knowledge fragments. Compared with the knowledge graph question-answering system that relies on the quality of extracted triples, the method of the present invention is simpler and more practical.
[0260] Corresponding to the multimodal RAG knowledge question answering method applied to a vertical field provided by the above-mentioned embodiment of the present invention, see Fig.13, showing a structural block diagram of a multimodal RAG knowledge question-and-answer device applied to a vertical field provided by an embodiment of the present invention.
[0261] The multimodal RAG knowledge question-answering device applied to a vertical field includes: a summarization unit 1301, a retrieval unit 1302, a fusion unit 1303, a first sorting unit 1304, a second sorting unit 1305, a construction unit 1306, an extraction unit 1307 and an output unit 1308.
[0262] The summarizing unit 1301 is used to receive multiple multimodal type data, convert all multimodal type data into text type data using a preset conversion model and summarize them to obtain the text to be answered.
[0263] The retrieval unit 1302 is used to use multiple preset retrieval models to retrieve knowledge fragments corresponding to the text to be answered from the knowledge base, and obtain multiple sorted lists containing the text to be answered and multiple knowledge fragments; the knowledge base is pre-constructed based on question-answer data and plain text data.
[0264] The fusion unit 1303 is used to fuse the knowledge fragments in all the sorted lists according to the inverse sorting fusion algorithm.
[0265] The first sorting unit 1304 is used to calculate the corresponding score of each knowledge fragment in the fused list by using the target re-ranking model, and sort each knowledge fragment in the fused list based on the score to obtain a first sorted list; the target re-ranking model is obtained by fine-tuning and training the open source re-ranking model in advance using a preset vertical field data set.
[0266] The second sorting unit 1305 is used to calculate the weight of each knowledge fragment in the first sorting list according to the preset knowledge frequency information, calculate the target score of each knowledge fragment in the first sorting list by combining the score and the weight, and sort them to obtain a second sorting list.
[0267] The construction unit 1306 is used to extract k knowledge fragments from the second sorted list, arrange and combine them to obtain arrangement results, and construct a prompt combination with each arrangement result and the text to be answered to obtain k prompt combinations.
[0268] The extraction unit 1307 is used to calculate the scores of all prompt combinations using the target reward model, extract n prompt combinations in descending order of the scores, and mark them as target prompt combinations; the target reward model is obtained by training the open source reward model in advance using the preference data set of the preset LLM question and answer.
[0269] The output unit 1308 is used to generate corresponding answer data for each target prompt combination using the large language model, and calculate the scores of all answer data using the target scoring model, and output the answer data with the largest score; the target scoring model is obtained by pre-training a Transformer encoder-based scoring model using a preset question-answer pair training set.
[0270] In an embodiment of the present invention, the explanation of professional terms in vertical fields is expanded, and the system's understanding of professional knowledge is enhanced. It supports multimodal input, including text, pictures, audio and video, so as to make full use of the semantic information of multimodal data. The large language model is used to generate question-answer pair data for pure knowledge documents, thereby expanding the scarce question-answer pair data. The present invention also generates a summary of knowledge fragments to improve its expressive power. Multiple retrieval results are fused through the RRF algorithm, and the scores are calculated using the rearrangement model for rearrangement, thereby improving the recall rate of relevant content. Based on the retrieval content and the real-time knowledge frequency information of the system, the sorting is further fine-tuned to make it more accurate and in line with user expectations. Combined with the natural language understanding ability of the large language model and the target reward model, as well as the relevant information of RAG recall, the optimal answer data is generated, and the target scoring model and domain expert knowledge are used to provide more professional and practical high-quality answer data.
[0271] Combination Fig.13 The content shown, the device also includes: an acquisition unit, a first vectorization processing unit, a first storage unit, a segmentation unit, a first splicing unit, a generation unit, a second vectorization processing unit, a second storage unit, a second splicing unit and a third storage unit.
[0272] The acquisition unit is used to acquire multiple original data and parse and process all the original data to obtain multiple data to be stored.
[0273] The first vectorization processing unit is used to perform vectorization processing on the question data in each data to be stored, when the data to be stored is question-answer pair data, using a pre-trained target embedding model to obtain the vectorized question data.
[0274] The first storage unit is used to store the vectorized question data and its corresponding answer data in a knowledge base.
[0275] The segmentation unit is used to segment each data to be stored in the warehouse using a semantic segmentation model when the data to be stored in the warehouse is pure document data, obtain multiple segmentation segments, and obtain the document title of the data to be stored in the warehouse.
[0276] The first splicing unit is used to splice the document title with each segmented segment to obtain multiple knowledge segments.
[0277] The generation unit is used to generate target question-answer pair data and corresponding summary data for each knowledge fragment using the large language model and all knowledge fragments.
[0278] The second vectorization processing unit is used to use the target embedding model to vectorize the question data in the target question-answer pair data to obtain the target question data.
[0279] The second storage unit is used to store the target question data and its corresponding answer data in combination in the knowledge base.
[0280] The second concatenation unit is used to concatenate the corresponding summary data of each knowledge fragment with the knowledge fragment to obtain a new knowledge fragment.
[0281] The third storage unit is used to vectorize the new knowledge fragments using the target embedding model, and store the vectorized new knowledge fragments into the knowledge base.
[0282] Combination Fig.13 The content shown, the acquisition unit, includes: an acquisition module, a conversion module, a generation module, an insertion module, a detection module and a storage module.
[0283] The acquisition module is used to acquire multiple raw data.
[0284] The conversion module is used to convert each original data into text type data using a preset conversion model when the type of the original data is a multimodal type other than a text type, and mark it as text data.
[0285] A generation module is used to generate a link placeholder according to the storage location of the original data.
[0286] The inserting module is used to insert a link placeholder into the text information of the text data to obtain new text data.
[0287] The detection module is used to detect whether the text information of new text data contains professional terms based on a preset professional term library.
[0288] The storage module is used to obtain detailed information of the professional term from the preset professional term library if it is included, and store the detailed information in the text information of the new text data to obtain the data to be stored in the library.
[0289] Combination Fig.13 The device further includes: a judgment unit, a calculation unit, a first display unit and a second display unit.
[0290] The judgment unit is used to judge whether the answer data with the largest score meets the user's needs.
[0291] The calculation unit is used to calculate the similarity between all question-answer pair data in the knowledge base and the text to be answered if it does not meet the user's requirements.
[0292] The first display unit is used to sort all similarities in descending order, obtain the first m questions and display them to the user.
[0293] The second display unit is used to obtain the answer data of the selected question contained in the question selection instruction and display it to the user when receiving the question selection instruction.
[0294] Combination Fig.13 The content shown, the summarizing unit 1301, includes: a receiving module, a detection type module, a conversion type module and a summarizing module.
[0295] The receiving module is used to receive multiple multi-modal data.
[0296] The detection type module is used to detect, for each multimodal type data, whether the multimodal type data is text type data.
[0297] The conversion type module is used to convert the multimodal type data into text type data using a preset conversion model to obtain multiple text type data if not.
[0298] The summary module is used to summarize all text type data to obtain the text to be answered.
[0299] Combination Fig.13 The content shown, the fusion unit 1303, includes: a ranking acquisition module, a score calculation module and a sorting fusion module.
[0300] The ranking acquisition module is used to acquire the ranking of each knowledge fragment in all the ranked lists in multiple ranked lists.
[0301] The score calculation module is used to calculate the reciprocal of the sum of each ranking and a constant, and accumulate all the reciprocals to obtain the corresponding score of the knowledge fragment.
[0302] The sorting and fusion module is used to sort all knowledge fragments based on the corresponding score of each knowledge fragment to obtain a fused list.
[0303] Combination Fig.13 As shown in the content, the second sorting unit 1305 includes: a frequency acquisition module, a weight calculation module, a target score calculation module and a sorting module.
[0304] The frequency acquisition module is used to obtain the frequency corresponding to the current knowledge fragment and the maximum frequency from the preset knowledge frequency information for each knowledge fragment in the first sorted list, wherein the preset knowledge frequency information includes the knowledge fragment and the frequency corresponding to the knowledge fragment.
[0305] The weight calculation module is used to calculate the ratio of the frequency corresponding to the current knowledge fragment to the maximum frequency and multiply it by the weight base to obtain the weight of the current knowledge fragment.
[0306] The target score calculation module is used to calculate the sum of the weight of the current knowledge fragment plus 1, and calculate the product of the sum and the score to obtain the target score of the current knowledge fragment.
[0307] The sorting module is used to sort the target scores of all knowledge fragments in descending order to obtain a second sorting list.
[0308] In summary, the present invention covers the entire process from vertical domain model training, knowledge base construction, knowledge enhancement to multimodal input. By integrating RRF retrieval rearrangement, weight update rearrangement, reward model and scoring model to assist large language model reasoning, the input form of the question-answering system is enriched, the accuracy of retrieval recall is improved, and the overall quality of the answer is significantly improved.
[0309] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0310] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0311] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multimodal RAG knowledge question answering method applied to vertical fields, characterized in that: The method comprises: Receive multiple multimodal data types, convert all multimodal data types into text type data using a preset conversion model and summarize them to obtain the text to be answered; Retrieving the knowledge fragments corresponding to the text to be answered from the knowledge base using multiple preset retrieval models to obtain multiple ranked lists containing the text to be answered and multiple knowledge fragments; the knowledge base is pre-constructed based on question-answer pair data and plain text data; Merge the knowledge fragments in all the sorted lists according to the inverse sorting fusion algorithm; The target re-ranking model is used to calculate the corresponding score of each knowledge fragment in the fused list, and each knowledge fragment in the fused list is ranked based on the score to obtain a first ranked list; the target re-ranking model is obtained by fine-tuning and training an open source re-ranking model in advance using a preset vertical field data set; Calculate the weight of each knowledge fragment in the first sorted list according to the preset knowledge frequency information, calculate the target score of each knowledge fragment in the first sorted list by combining the score and the weight, and sort them to obtain a second sorted list; Extract k pieces of the knowledge from the second sorted list and perform permutations and combinations to obtain permutation results, and construct a prompt combination with each permutation result and the text to be answered to obtain k prompt combinations; The target reward model is used to calculate the scores of all prompt combinations, and n prompt combinations are extracted in descending order according to the scores, and marked as target prompt combinations; the target reward model is obtained by training the open source reward model in advance using the preference data set of the preset LLM question and answer; A large language model is used to generate corresponding answer data for each target prompt combination, and a target scoring model is used to calculate the scores of all answer data, and the answer data with the largest score is output; the target scoring model is obtained by pre-training a Transformer encoder-based scoring model using a preset question-answer pair training set.
2. The method according to claim 1, characterized in that The process of building a knowledge base based on question-answer pair data and plain text data in advance includes: Acquire multiple raw data and parse and process all the raw data to obtain multiple data to be stored in the database; For each of the data to be stored, when the data to be stored is question-answer pair data, use a pre-trained target embedding model to vectorize the question data in the data to be stored to obtain the vectorized question data; Combining the vectorized question data and its corresponding answer data and storing them in a knowledge base; For each of the data to be stored, when the data to be stored is pure document data, the data to be stored is segmented using a semantic segmentation model to obtain a plurality of segmentation segments, and the document title of the data to be stored is obtained; Splicing the document title with each of the segmented fragments to obtain multiple knowledge fragments; Generate target question-answer pair data using the large language model and all knowledge fragments, as well as summary data corresponding to each of the knowledge fragments; Vectorizing the question data in the target question-answer pair data using the target embedding model to obtain target question data; Combining the target question data and its corresponding answer data and storing them in a knowledge base; splicing the summary data corresponding to each of the knowledge fragments with the knowledge fragments to obtain a new knowledge fragment; The new knowledge fragment is vectorized using the target embedding model, and the vectorized new knowledge fragment is stored in the knowledge base.
3. The method according to claim 2, characterized in that The method of obtaining multiple original data and parsing and processing all the original data to obtain multiple data to be stored in the database includes: Get multiple raw data; For each of the original data, when the type of the original data is a multimodal type other than a text type, convert the original data into text type data using a preset conversion model and mark it as text data; Generate a link placeholder according to the storage location of the original data; Inserting the link placeholder into the text information of the text data to obtain new text data; Detecting whether the text information of the new text data contains professional terms according to a preset professional term library; If included, detailed information of the professional term is obtained from the preset professional term library, and the detailed information is stored in the text information of the new text data to obtain the data to be stored in the library.
4. The method according to claim 1, characterized in that: After outputting the answer data with the largest score, the method further includes: Determine whether the answer data with the largest score meets the user's needs; If it does not meet the user's needs, calculate the similarity between all question-answer pair data in the knowledge base and the text to be answered; Sort all similarities in descending order, obtain the top m questions and display them to the user; When a question selection instruction is received, answer data of the selected question contained in the question selection instruction is obtained and displayed to the user.
5. The method according to claim 1, characterized in that The receiving of a plurality of multimodal data types, converting all the multimodal data types into text type data using a preset conversion model and summarizing the data to obtain a text to be answered, includes: Receiving multiple multimodal type data; For each of the multimodal type data, detecting whether the multimodal type data is text type data; If not, convert the multimodal type data into text type data using a preset conversion model to obtain a plurality of text type data; Summarize all text type data to obtain the text to be answered.
6. The method according to claim 1, characterized in that The step of fusing the knowledge fragments in all the sorted lists according to the inverse sorting fusion algorithm includes: For each knowledge fragment in all the sorted lists, obtaining the ranking of the knowledge fragment in the plurality of sorted lists; Calculate the reciprocal of the sum of each ranking and a constant, and accumulate all the reciprocals to obtain the corresponding score of the knowledge fragment; All knowledge fragments are sorted based on the corresponding score of each knowledge fragment to obtain a fused list.
7. The method according to claim 1, characterized in that The step of calculating the weight of each knowledge fragment in the first sorted list according to the preset knowledge frequency information, calculating the target score of each knowledge fragment in the first sorted list in combination with the score and the weight, and sorting the scores to obtain a second sorted list includes: For each knowledge segment in the first sorted list, obtaining a frequency corresponding to the current knowledge segment and a maximum frequency from preset knowledge frequency information, wherein the preset knowledge frequency information includes knowledge segments and frequencies corresponding to the knowledge segments; Calculate the ratio of the frequency corresponding to the current knowledge fragment to the maximum frequency and multiply it by the weight base to obtain the weight of the current knowledge fragment; Calculate the sum of the weight of the current knowledge fragment plus 1, and calculate the product of the sum and the score to obtain the target score of the current knowledge fragment; The target scores of all knowledge fragments are sorted in descending order to obtain a second sorted list.
8. A multi-modal RAG knowledge question-answering device applied to vertical fields, characterized in that: The device comprises: A summarizing unit is used to receive multiple multimodal data, convert all multimodal data into text data using a preset conversion model, and summarize them to obtain a text to be answered; A retrieval unit, used to retrieve the knowledge fragments corresponding to the text to be answered from the knowledge base using multiple preset retrieval models, and obtain multiple sorted lists including the text to be answered and multiple knowledge fragments; the knowledge base is pre-constructed based on question-answer pair data and plain text data; A fusion unit, used for fusing the knowledge fragments in all the sorted lists according to a reciprocal sorting fusion algorithm; A first sorting unit is used to calculate the corresponding score of each knowledge fragment in the fused list by using the target re-ranking model, and sort each knowledge fragment in the fused list based on the score to obtain a first sorted list; the target re-ranking model is obtained by fine-tuning and training an open source re-ranking model in advance using a preset vertical field data set; A second sorting unit, configured to calculate the weight of each knowledge fragment in the first sorting list according to preset knowledge frequency information, calculate the target score of each knowledge fragment in the first sorting list in combination with the score and the weight, and sort the scores to obtain a second sorting list; A construction unit is used to extract k pieces of the knowledge fragments from the second sorted list, arrange and combine them to obtain an arrangement result, and each of the arrangement result and the text to be answered is used to construct a prompt combination to obtain k prompt combinations; An extraction unit, used to calculate the scores of all prompt combinations using a target reward model, extract n prompt combinations in descending order according to the scores, and mark them as target prompt combinations; the target reward model is obtained by training an open source reward model in advance using a preset LLM question-answering preference data set; An output unit is used to generate answer data corresponding to each target prompt combination using a large language model, and calculate the scores of all answer data using a target scoring model, and output the answer data with the largest score; the target scoring model is obtained by pre-training a scoring model based on a Transformer encoder using a preset question-answer pair training set.
9. The device according to claim 8, characterized in that The device also includes: An acquisition unit is used to acquire multiple original data and parse and process all the original data to obtain multiple data to be stored in the database; A first vectorization processing unit is used for, for each of the data to be stored, when the data to be stored is question-answer pair data, using a pre-trained target embedding model to perform vectorization processing on the question data in the data to be stored, so as to obtain the vectorized question data; A first storage unit, used to store the vectorized question data and its corresponding answer data in a knowledge base; A segmentation unit, for segmenting each of the data to be stored in the warehouse using a semantic segmentation model to obtain a plurality of segmentation segments when the data to be stored in the warehouse is pure document data, and obtaining a document title of the data to be stored in the warehouse; A first concatenation unit is used to concatenate the document title with each of the segmented fragments to obtain a plurality of knowledge fragments; A generation unit, used to generate target question-answer pair data using a large language model and all knowledge fragments, as well as summary data corresponding to each of the knowledge fragments; A second vectorization processing unit is used to perform vectorization processing on the question data in the target question-answer pair data using the target embedding model to obtain target question data; A second storage unit, used to store the target question data and its corresponding answer data in combination in a knowledge base; A second concatenation unit is used to concatenate the summary data corresponding to each of the knowledge fragments with the summary data to obtain a new knowledge fragment; The third storage unit is used to use the target embedding model to vectorize the new knowledge fragments, and store the vectorized new knowledge fragments in the knowledge base.
10. The device according to claim 9, characterized in that The acquisition unit comprises: An acquisition module, used for acquiring multiple raw data; A conversion module, for each of the original data, when the type of the original data is a multimodal type other than a text type, using a preset conversion model to convert the original data into text type data, and mark it as text data; A generating module, used for generating a link placeholder according to the storage location of the original data; An inserting module, used for inserting the link placeholder into the text information of the text data to obtain new text data; A detection module, used to detect whether the text information of the new text data contains professional terms according to a preset professional term library; The storage module is used to obtain detailed information of the professional term from the preset professional term library if included, and store the detailed information in the text information of the new text data to obtain the data to be stored in the library.
Citation Information
Patent Citations
Multi-mode interacting method applied to intelligent robot system and intelligent robots
CN106335058A
Personalized learning platform based on self-expansion knowledge base and multi-modal portrait
CN114372155A