RAG enhanced small-scale language model question and answer accuracy improving method and system

By adjusting the text sharding processing method in the knowledge base and optimizing the search process of the RAG model, the RAG enhanced small-scale language model is solved, and the problem of inaccurate answers and insufficient information coverage in the field of bidding Q&A is achieved, achieving higher Q&A accuracy and information coverage.

CN120162415AActive Publication Date: 2025-06-17JIANGSU BINGXIN TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510313450.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-17
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The existing RAG enhanced small-scale language model has problems such as inaccurate answers and insufficient information coverage in the field of tender question and answers, mainly due to the failure to retrieve sufficient content, the limited number of context parameters, and the short length of the question text is ignored.

Method used

By adjusting the text form of the answer basis in the knowledge base to text sharding, the search process of the RAG model is optimized, and the question is expanded and the initial answer basis is scored, ensuring the quality and simplicity of the final answer basis.

Benefits of technology

It improves the accuracy of question and answers and information coverage, adapts to the context length limitation of small-scale language models, reduces invalid information interference, and maximizes the accuracy of answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162415A_ABST
    Figure CN120162415A_ABST
Patent Text Reader

Abstract

The invention discloses an RAG enhanced small-scale language model question and answer accuracy improving method and system. The method comprises the following steps: adjusting a text form of an answer basis in a knowledge base according to a text fragmentation processing mode; obtaining and utilizing a first verification set to perform retrieval optimization on the RAG model; taking a question expanded by using a natural language technology as input, and performing retrieval in a knowledge base associated with a bid invitation scene by using the RAG model after retrieval optimization to obtain a plurality of initial answer bases; respectively outputting the plurality of initial answer bases to an initial answer base scoring model, screening and sorting the plurality of initial answer bases according to a scoring result, and completing de-duplication filtering processing; performing content simplification summarization on each processed initial answer basis by using a natural language technology to obtain a final answer basis; and according to the question and a final answer basis, generating an answer by using an RAG model. According to the method, the question and answer accuracy of the RAG enhanced small-scale language model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing and question-answering systems, and specifically relates to a method and system for improving the question-answering accuracy of a RAG-enhanced small-scale language model. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, especially the progress of natural language processing and knowledge base technology, small-scale language models enhanced by RAG (Retrieval-Augmented Generation) have received extensive attention in the application of question-answering systems in specific fields. These systems not only improve the information query efficiency but also, to a certain extent, improve the way of internal enterprise knowledge management. Especially in fields such as tendering and procurement, by entering tender documents into the knowledge base and combining RAG enhancement with language model technology, an automated question-answering function has been realized, greatly improving work efficiency and reducing labor cost input.

[0003] Currently, when solving problems in the tender question-answering field, the common practice is to adopt the technical means of knowledge base + RAG enhancement + language model. The specific process is as follows: First, the tender documents are pre-processed and entered into the knowledge base; when someone asks a question, the program will automatically match the top 5-10 relevant pieces of information from the knowledge base as the basis for answering, then provide these pieces of information together with the original question to the language model, and finally, the model generates the final answer and feedbacks it to the user.

[0004] Although the above method is more intelligent than the traditional method, there are still many defects, and there are still quite a few invalid or even incorrect answer situations. The main reasons for generating invalid answers are as follows: First, during the process of applying RAG enhancement, not enough rich content is retrieved, resulting in the language model lacking the necessary "information" for accurate answering; Second, this technology is usually applied in a local network environment and cannot connect to the Internet to call larger-scale language model resources. The maximum number of context parameters that can be supported is very limited, and it is impossible to simply rely on increasing additional parameters to improve the answering accuracy; Third, when the inquirer asks a question, the length of the question text is often too short and much smaller than the length of the text fragments already in the question-answering library, so the absolute similarity between the two will become extremely low, and key details are likely to be ignored because they do not reach the set text similarity threshold standard, further weakening the quality and credibility of the final answer. Therefore, a method and system for improving the question-answering accuracy of a RAG-enhanced small-scale language model are needed. Summary of the Invention

[0005] In order to improve the question-answering accuracy of a RAG-enhanced small-scale language model, this application provides a method and system for improving the question-answering accuracy of a RAG-enhanced small-scale language model.

[0006] In a first aspect, the present application provides a method for improving the accuracy of question answering of a RAG-enhanced small-scale language model, including: Adjust the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text sharding processing method; Obtain question-answer pairs containing answer bases stored in the form of text shards in the bidding scenario as the first validation set, and use the first validation set to optimize the retrieval of the RAG model; Using natural language technology, expand the content or quantity of the questions input into the RAG model, use the expanded questions as input, and use the optimized RAG model for retrieval in the knowledge base associated with the bidding scenario to obtain multiple initial answer bases; Based on the deep learning algorithm, construct an initial answer basis scoring model under the RAG model framework, output multiple initial answer bases to the initial answer basis scoring model respectively, screen and sort the multiple initial answer bases according to the RAG model token limit and the output scores, and complete the deduplication and filtering process to obtain multiple processed initial answer bases; use natural language technology to respectively streamline and summarize the content of each processed initial answer basis to obtain the final answer basis; Generate an answer using the RAG model based on the question and the final answer basis.

[0007] By adopting the above scheme, the RAG model is optimized using answer pairs containing answer bases in the form of text shards to improve the accuracy of retrieval output; the text sharding processing method is adopted for the knowledge base to replace the keyword retrieval and analysis of each word and sentence, thereby improving the information coverage and ensuring the retention of key answer bases; natural language technology is used to expand the questions and retrieve the initial answer bases in the optimized RAG model, and the scoring model is used to screen, sort and deduplicate the initial answer bases to remove invalid information and adapt to the context length limit of the small-scale language model.

[0008] Preferably, it further includes: On the basis of adjusting the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text sharding processing method, continue to complete the vectorization storage of each piece of text according to the vectors generated by dividing the large and small text blocks, and generate a knowledge base that allows the vector corresponding to the small text block to index the vector corresponding to the large text block; Obtain answer pairs containing answer bases stored in the form of vectors corresponding to large and small text blocks in the bidding scenario to replace the question-answer pairs containing answer bases stored in the form of text shards in the bidding scenario as the second validation set, and then use the second validation set to optimize the retrieval of the RAG model; During the retrieval process of the RAG model optimized by retrieval in the knowledge base associated with the bidding scenario, when the corresponding vector of a small piece of text is retrieved, the corresponding vector of the relevant large piece of text is indexed, so as to obtain multiple initial answer bases containing the large piece of text.

[0009] By adopting the above solution, in the bidding Q&A scenario, through the size-based fragmentation processing of the text form of the answer basis in the knowledge base and converting it into vector storage, an efficient knowledge base capable of indexing the corresponding vector of the large piece of text with the corresponding vector of the small piece of text is constructed, improving the information coverage in the retrieval process. Especially when facing short consulting questions, relevant complete large piece of text content can be obtained, avoiding the omission of effective information caused by text length differences.

[0010] Preferably, it further includes: Adjust the text form of the answer basis in the knowledge base associated with the bidding scenario according to different text fragmentation processing methods and correspondingly generate different types of knowledge bases, so that each type of knowledge base stores the answer basis with the text form adjusted according to a single text fragmentation processing method; the different text fragmentation processing methods include: clauses, paragraphs, and chapters; Obtain the Q&A pairs containing the answer basis stored in the text form generated by single text fragmentation processing in the bidding scenario as the third validation set, use the third validation set to optimize the retrieval of the RAG model, and correspondingly obtain several RAG models optimized by retrieval; For the expanded target question, use machine learning algorithms to evaluate the complexity of the expanded question and the importance of the information in the expanded question respectively, obtain the weighted scoring result value of the question complexity score and the question information importance score, and match the preset text fragmentation processing method corresponding to the scoring result value range where it is located; According to the matched preset text fragmentation processing method, correspondingly match the RAG model optimized by retrieval, and use the matched RAG model optimized by retrieval to retrieve in the knowledge base where the stored text form of the answer basis is the same as the text form generated by the preset text fragmentation processing method to obtain multiple initial answer bases.

[0011] By adopting the above solution, adjust the text form of the answer basis in the knowledge base according to different text fragmentation processing methods (clauses, paragraphs, chapters), and statistically generate knowledge bases corresponding to the text forms. Combining the evaluation results of the target question complexity and information importance, match the text fragmentation processing method, and adaptively match the RAG model optimized by retrieval to retrieve in the knowledge base in a specific format, effectively improving the accuracy and relevance of multiple initial answer bases.

[0012] Preferably, it further includes: When adjusting the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text slicing processing method, natural language analysis technology is used to analyze the logical coherence and information relevance of multiple text slices, and the logical coherence strength and information relevance strength of multiple text slices with logical coherence and / or information relevance are respectively evaluated; a marking mechanism is introduced to display and distinguish multiple text slices with different logical coherence strengths and information relevance strengths; After each initial answer basis after processing is concisely summarized, each initial answer basis with the content concisely summarized is sorted and reorganized according to the coherence strength and information relevance strength marked in the text content of the original initial answer basis, so that the answer basis text fragments belonging to the context content are sorted adjacent to each other and the answer basis text fragments with stronger information relevance content strength are sorted closer, and the final answer basis is obtained.

[0013] By adopting the above scheme, after obtaining the initial answer basis, the content of the initial basis is concisely summarized, and sorted and organized according to its logical coherence and information relevance, so that the final answer basis obtained not only is more concise and clear, but also ensures the reasonable arrangement order of the context and related content, and improves the accuracy and rationality of the small-scale language model in the bidding Q&A scenario.

[0014] Preferably, it further includes: Real-time receive the feedback data of the user on the generated answer. When the user's satisfaction with the feedback is lower than the preset satisfaction, the initial answer basis is correspondingly displayed and the user is prompted to evaluate the satisfaction of the initial answer basis, and the satisfaction data of the user on the initial answer basis is retained; When the frequency of the user's satisfaction being lower than the preset satisfaction within the preset time period is greater than the preset frequency, judge whether the frequency of the user's satisfaction with the initial answer basis retained within the preset time period being lower than the preset satisfaction is greater than the preset frequency. If so, select the RAG model optimized for retrieval for optimization training; otherwise, optimize the training of the initial answer basis scoring model.

[0015] By adopting the above scheme, the user feedback mechanism is used to determine that the root cause of the user's dissatisfaction with the generated answer lies in the problem of the RAG model optimized for retrieval or the scoring model, and adaptively select the optimization object according to the problem, so as to improve the user satisfaction of the generated answer.

[0016] Preferably, it further includes: Before adjusting the text form of the answer basis in the knowledge base associated with the tender scenario according to the text sharding processing method, information expansion is performed on the answer basis in the tender scenario knowledge base, including: supplementing information by collecting the associated tender text content using the knowledge graph in the tender scenario, or collecting the text content converted from the multimodal data in the tender documents for the tender scenario for information supplementation.

[0017] By adopting the above solution, the text content associated with each tender document or the text content converted from the multimodal data in the tender document in the tender scenario is collected and supplemented using the knowledge graph in the tender scenario, enriching the content of the answer basis at the source, thereby improving the answering quality.

[0018] Preferably, it further includes: The text form of the answer basis in the knowledge base associated with the tender scenario adjusted according to the text sharding processing method includes: sharding the text content of the answer basis in the knowledge base associated with the tender scenario according to the preset text sharding processing method, and adding the outline title content to which the shard belongs before the answer basis in each text form.

[0019] By adopting the above solution, adding the outline title content to which each piece of text belongs can effectively improve the information coverage in the subsequent retrieval process, help the language model more accurately locate relevant knowledge points, and thus improve the answering accuracy.

[0020] In the second aspect, the present application provides a system for improving the answering accuracy of a RAG-enhanced small-scale language model, including: A knowledge base sharding adjustment module for adjusting the text form of the answer basis in the knowledge base associated with the tender scenario according to the text sharding processing method; A RAG model retrieval optimization module for obtaining question-and-answer pairs containing answer bases stored in text shard form in the tender scenario as the first verification set, and using the first verification set to optimize the retrieval of the RAG model; A RAG model initial inspection result acquisition module for using natural language technology to expand the content or quantity of the question input into the RAG model, taking the expanded question as the input, and using the RAG model optimized by retrieval to retrieve in the knowledge base associated with the tender scenario to obtain multiple initial answer bases; A RAG model final inspection result acquisition module for constructing an initial answer basis scoring model under the RAG model framework based on a deep learning algorithm, outputting multiple initial answer bases to the initial answer basis scoring model respectively, screening and sorting the multiple initial answer bases according to the RAG model token limit and the output scores, and completing the duplicate filtering process to obtain multiple processed initial answer bases; using natural language technology to respectively perform content refinement and generalization on each processed initial answer basis to obtain the final answer basis; The RAG model answer generation acquisition module is used to generate an answer using the RAG model based on the question and the final answer basis.

[0021] By adopting the above solution, the text sharding processing method of the knowledge base associated with the bidding scenario is adjusted to cover more information; using the Q&A pairs containing the answer basis stored in the form of text shards as the validation set, the retrieval performance of the RAG model is optimized, and the matching accuracy is improved; the expansion of the input question, the screening, sorting and deduplication filtering of multiple initial answer bases ensure the quality and conciseness of the finally selected answer basis, jointly achieving the improvement of the Q&A accuracy of the small-scale language model in the bidding scenario.

[0022] In a third aspect, the present application provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method as described above.

[0023] In a fourth aspect, the present application provides a computer device, which includes a memory, a processor, and a program stored and executable on the memory. When the program is executed by the processor, it implements the steps of the method as described above.

[0024] In summary, the present application has the following beneficial effects: 1. By optimizing the text sharding processing method, the coverage rate of key information is improved; introducing an optimization method based on the validation set to complete the optimization of the RAG model retrieval, and improving the retrieval hit rate; using natural language technology to expand the question and streamline, summarize, score and sort the initial answer basis, improving the adaptability of the retrieved content, while removing redundant and invalid information, and maximizing the adaptation to the context parameter limitations of the small-scale language model, thereby enhancing the answer accuracy, and completing the optimization from multiple aspects such as questions and retrievals to improve the overall question answer generation accuracy; 2. Through the adaptive text sharding processing method, combined with the large and small block sharding strategy and vectorized storage, the coverage rate of key information and the retrieval hit rate are improved; 3. Introducing a user feedback mechanism, according to the user's satisfaction with the generated answer and the initial answer basis, correspondingly optimizing the RAG model or the scoring model, so as to obtain accurate answers using a better model. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a flowchart of the method for improving the Q&A accuracy of the RAG enhanced small-scale language model in the specific embodiment; Figure 2It is a diagram after text sharding processing of the text in the knowledge base in the method for improving the Q&A accuracy of the RAG-enhanced small-scale language model described in the specific embodiment; Figure 3 It is a schematic flow chart for obtaining the basis for the final answer in the method for improving the Q&A accuracy of the RAG-enhanced small-scale language model described in the specific embodiment; Figure 4 It is a schematic diagram of the retrieval chunking strategy in the method for improving the Q&A accuracy of the RAG-enhanced small-scale language model described in the specific embodiment; Figure 5 It is a schematic structural diagram of the RAG-enhanced small-scale language model Q&A accuracy improvement system described in the specific embodiment. Detailed implementation manner

[0026] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0027] Considering that it is difficult for small-scale language models used in the local area network environment to achieve ideal answering accuracy, especially affected by the context length limit and the knowledge base sharding granularity. For this reason, the present application mainly adopts the following solutions, significantly improving the Q&A accuracy and applicability of the RAG-enhanced small-scale language model. The following is a more detailed technical implementation plan.

[0028] Embodiment 1 As Figure 1 shown, the embodiment of the present application discloses a method for improving the Q&A accuracy of the RAG-enhanced small-scale language model, and the steps include: S1. Adjust the storage text form in the knowledge base associated with the bidding scenario according to the text sharding processing method.

[0029] Specifically, the RAG-enhanced small-scale language model, hereinafter simply referred to as the "RAG model", RAG enhancement is an expansion of the language model's ability to answer questions, retrieving a large amount of information in the knowledge base to enable the model to answer professional questions in the bidding Q&A field. Therefore, as the basis for obtaining the answer basis, in order to improve the information coverage as much as possible and ensure information integrity; as Figure 2 shown, the text of the bidding documents, texts related to bidding projects and other answer basis texts stored in the knowledge base associated with the bidding scenario are sharded, and when sharding the text, it is no longer subdivided into sentences, but sharded into the form of clauses, paragraphs or chapters to ensure that when retrieving fragments related to the question, as much information as possible is covered.

[0030] S2. Introduce validation set data to complete the retrieval optimization of the RAG model.

[0031] Specifically, obtain the answer pairs in the bidding scenario as the first validation set. Among them, the answer pairs include typical questions in the bidding scenario and the answer basis generated by experts, and the answer basis is stored in the form of text shards.

[0032] The RAG model mainly includes two basic models: a retrieval model and an answer generation model. The retrieval model is used to retrieve the corresponding answer basis from the knowledge base according to the semantics of the question, and the answer generation model is used to answer the question based on the retrieved basis. Thus, traverse and verify the first data set, output the questions in the first validation set into the RAG model, compare the retrieval results (answer basis related to the question) with the reference data in the first validation set (answer basis generated by experts), and confirm their similarity. Since there are also multiple records in the first validation set, the overall judgment of whether to optimize can be judged from two perspectives. One is whether all multiple bases are matched. When any one of the retrieval results has a similarity greater than 95% with the verification basis, it is marked as a match, and the retrieval result is considered excellent. The other is to evaluate the similarity of each basis, take the highest value of the similarity between each reference basis and the retrieval result or take the average value of multiple records. If the finally obtained similarity is greater than the preset retrieval similarity value, it indicates that the current retrieval result of the RAG is excellent, and no retrieval optimization is required; otherwise, the data in the first validation set is used as the training set for retrieval optimization of the RAG model.

[0033] Among them, the similarity calculation method: vectorize the two texts, and calculate a value between 0% and 100% through the cosine algorithm. The higher the value, the greater the similarity. The embedding model used for vectorization is bge-large-zh-v1.5.

[0034] S3. Use the RAG model after retrieval optimization to retrieve and obtain multiple initial answer bases.

[0035] Specifically, in order to avoid the problem that the length gap between the consultation question and the text shards in the knowledge base is too large to be hit, the question expansion method is adopted to retrieve more sufficient answer bases. In this embodiment, natural language technology is used to expand the content or quantity of the question input into the RAG model.

[0036] Take the expanded question as the input, and use the RAG model after retrieval optimization to retrieve in the knowledge base associated with the bidding scenario to obtain multiple initial answer bases.

[0037] S4. Screen and filter the multiple initial answer bases obtained to obtain the final answer basis.

[0038] While considering covering as much information as possible, the scale of the RAG model also needs to be taken into account. It is necessary to maximize the effective information within the limited context parameters. Therefore, some invalid information needs to be removed to adapt to the context length of the small model.

[0039] As Figure 3 shown, each initial answer basis retrieved is scored, and multiple initial answer bases are re - sorted according to the scores. Combined with the model token limit for screening, answer bases with higher relevance to the question are obtained. Specifically, based on the deep - learning algorithm, an initial answer basis scoring model is constructed within the RAG model framework (such as between the retrieval model and the answer generation model), and is trained and generated through a training set formed by multiple initially retrieved answers marked with scores according to historical questions. Each of the multiple initial answer bases is output to the initial answer basis scoring model. According to the RAG model token limit and the output scores, the multiple initial answer bases are screened and sorted, and the initial answer bases with higher scores are retained. For example, in the consulting scenario of tendering questions and answers, the top 4 data are selected.

[0040] While performing the scoring and screening, duplicate - filtering processing also needs to be completed to obtain multiple processed initial answer bases.

[0041] In addition, to further improve the accuracy of the answers generated by the input answer generation model, each of the initially processed answer bases is either concisely summarized or sorted based on the logic and information relevance of the content within the initial answer basis.

[0042] Specifically, using natural language technology, each of the processed initial answer bases is concisely summarized to obtain the final answer basis.

[0043] After each of the processed initial answer bases is concisely summarized, each of the concisely summarized initial answer bases is sorted and re - organized according to the marked coherence strength and information association strength in the text content of the original initial answer basis, so that the answer basis text fragments belonging to the context content are sorted adjacent to each other and the answer basis text fragments with stronger information association strength are sorted closer, to obtain the final answer basis.

[0044] Among them, the specific coherence strength and information correlation strength marked in the text content on which the initial answer is based include: when adjusting the text form of the answer basis in the knowledge base related to the bidding scenario according to the text sharding processing method, natural language analysis technology is used to analyze the logical coherence and information relevance of multiple text shards, and the logical coherence strength and information correlation strength of multiple text shards with logical coherence and / or information relevance are respectively evaluated; a marking mechanism is introduced to display and distinguish multiple text shards with different logical coherence strengths and information correlation strengths.

[0045] S5. Generate an answer using the RAG model according to the question and the final answer basis.

[0046] Specifically, input the question and the final answer basis into the RAG model, and use the answer generation model to complete the generation of the final answer.

[0047] By adopting the above method, reasonable text sharding processing is carried out on the knowledge base, improving the information coverage rate; the retrieval effect is enhanced by introducing natural language extension technology and optimizing the validation set; the interference of invalid information is reduced by the scoring model and content simplification, so as to achieve the purpose of maximizing the effective information under limited resources.

[0048] Embodiment 2 The difference from the above Embodiment 1 is that the large and small block sharding strategy and vectorized storage are combined to further improve the coverage rate and retrieval hit rate of key information. The method further includes: As Figure 4 shown, on the basis of adjusting the text form of the answer basis in the knowledge base related to the bidding scenario according to the text sharding processing method, continue to complete the vectorized storage of each piece of text according to the vectors generated by the divided large and small block texts, and generate a knowledge base that allows the vector of the small block text to index the vector of the large block text.

[0049] That is, for each divided text shard, large and small block texts are divided and the small block texts are generated by dividing the large block. For example, large block A is divided into small blocks A_1, A_2, A_3, and large block B is divided into B_1, B_2, B_3, and corresponding small block indexes are generated to find the associated large block through the small blocks.

[0050] In order to find the corresponding large and small block texts in the knowledge base, correspondingly, after slicing the texts of the answer bases stored in the knowledge base associated with the bidding scenario, they are divided into large block texts, and then further divided into small block texts, and correspondingly converted into the vector form of large and small block texts for storage. Correspondingly, in order to improve the retrieval in the knowledge base in the form of large and small block texts, obtain the answer pairs containing the answer bases stored in the vector form of large and small block texts corresponding to the bidding scenario to replace the Q&A pairs containing the answer bases stored in the text slicing form in the bidding scenario as the second validation set, and then use the second validation set to optimize the retrieval of the RAG model.

[0051] During the process of retrieving in the knowledge base associated with the bidding scenario using the RAG model optimized by retrieval, in order to improve the coverage rate of key information, when retrieving the vector corresponding to the small block text, index the vector corresponding to the relevant large block text, so as to obtain multiple initial answer bases containing the information of the large block text.

[0052] Embodiment 3 The difference from Embodiment 1 is that different text slicing processing methods need to be adopted according to the complexity and information importance of different questions. For example: for simple and generally important questions, answers with high accuracy can be obtained without searching for answer bases with high information coverage. For complex or relatively important questions, answer bases with high information coverage need to be searched as much as possible, and an adaptive slicing method is adopted to meet the needs of different questions. The method further includes: Adjust the text form of the answer bases in the knowledge base associated with the bidding scenario according to different text slicing processing methods and correspondingly generate different types of knowledge bases, so that each type of knowledge base stores the answer bases whose text forms are adjusted according to a single text slicing processing method; the different text slicing processing methods include: clauses, paragraphs, and chapters. The first knowledge base stores the answer bases whose text forms are adjusted according to the clause text slicing processing method, the second knowledge base stores the answer bases whose text forms are adjusted according to the paragraph text slicing processing method, and the third knowledge base stores the answer bases whose text forms are adjusted according to the chapter text slicing processing method. Of course, different text processing methods can also be mixed and correspondingly generate different types of knowledge bases, so that each type of knowledge base stores the answer bases whose text forms are adjusted according to the mixed text slicing processing method.

[0053] In addition, in order to better assist in retrieving the corresponding sliced texts, while slicing the text content of the answer bases in the knowledge base associated with the bidding scenario according to the preset text slicing processing method, add the outline title content to which the slice belongs before each text form of the answer base.

[0054] For adaptive retrieval to obtain the content of a specific shard, obtain the question-answer pairs containing the answer basis stored in the form of text generated by processing a single text shard in the bidding scenario as the third validation set, and use the third validation set to optimize the retrieval of the RAG model, and correspondingly obtain several RAG models optimized for retrieval. For the expanded target question, use machine learning algorithms to evaluate the complexity of the expanded question and the importance of the information in the expanded question respectively, obtain the weighted score result value of the question complexity score and the question information importance score, and match the corresponding preset text shard processing method according to the weighted score result value range; for example: the first score result value range corresponds to the preset text shard processing method which is the clause text shard processing method, the second score result value range corresponds to the preset text shard processing method which is the paragraph text shard processing method, and the third score result value range corresponds to the preset text shard processing method which is the chapter text shard processing method.

[0055] According to the matched preset text shard processing method, correspondingly match the RAG model optimized for retrieval. For example: if the current question matches the paragraph text shard processing method, correspondingly match the RAG model optimized for retrieval related to the paragraph text shard processing method; use the matched RAG model optimized for retrieval to retrieve in the knowledge base where the stored answer basis text form is the same as the text form generated by the preset text shard processing method, and obtain multiple initial answer bases. For example: use the matched RAG model optimized for retrieval to retrieve in the second knowledge base with the same text form in the form of paragraphs.

[0056] Embodiment 4 The difference from the above Embodiment 1 is that a feedback mechanism is introduced to improve the accuracy of question answering. The method further includes: Receive the feedback data of the user on the generated answer in real time, including: user satisfaction. When the user's satisfaction with the feedback (such as: 60 points) is lower than the preset satisfaction (such as: 80 points), then correspondingly display the initial answer basis and prompt the user to evaluate the satisfaction of the initial answer basis; in this embodiment, set to evaluate the satisfaction of all initial answer bases one by one or as a whole, and retain the satisfaction data of the user on the initial answer basis, such as: the average value of the satisfaction evaluated one by one or the satisfaction value evaluated as a whole.

[0057] When the frequency of the user satisfaction being lower than the preset satisfaction within the preset time period is greater than the preset frequency, it indicates that the user is dissatisfied with the generated answers for most of the time. Then, it is necessary to determine whether the frequency of the user satisfaction with the basis of the initial answer being lower than the preset satisfaction within the corresponding statistical preset time period is greater than the preset frequency. If so, it means that the user is dissatisfied with the basis of the initial answer for most of the time, and it is necessary to further improve the search accuracy of the basis of the initial answer. Then, select the RAG model optimized for retrieval and perform optimization training again; otherwise, it means that the user is satisfied with the basis of the initial answer for most of the time, but dissatisfied with the finally generated answer, indicating that there may be a large error in the process of obtaining the basis of the final answer. Therefore, select to optimize the training of the scoring model for the basis of the initial answer.

[0058] Embodiment 5 Different from the above Embodiment 1, the process of building the knowledge base is further optimized. In addition to collecting static materials, a dynamic monitoring function is added, which can update the content of the knowledge base in a timely manner to maintain its timeliness and advancement. The method further includes: Before adjusting the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text sharding processing method, information expansion is performed on the answer basis in the knowledge base for the bidding scenario, specifically including: Using the knowledge graph in the bidding scenario to collect the associated bidding text content for information supplementation, or collecting the text content converted from the multi-modal data in the bidding documents for the bidding scenario for information supplementation, such as: the text content converted from videos and audios.

[0059] Such as Figure 5 As shown, this application embodiment discloses a system for improving the accuracy of answering questions of a RAG enhanced small-scale language model, specifically including: A knowledge base sharding adjustment module 101, configured to adjust the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text sharding processing method; A RAG model retrieval optimization module 102, configured to obtain the question-answer pairs containing the answer basis stored in the form of text shards in the bidding scenario as the first verification set, and use the first verification set to perform retrieval optimization on the RAG model; A RAG model initial inspection result acquisition module 103, configured to use natural language technology to expand the content or quantity of the question input into the RAG model, use the expanded question as the input, and perform retrieval in the knowledge base associated with the bidding scenario by using the RAG model optimized for retrieval to obtain multiple initial answer bases; The RAG model final inspection result acquisition module 104 is used to construct an initial answer basis scoring model under the RAG model framework based on deep learning algorithms, output multiple initial answer bases to the initial answer basis scoring model respectively, screen and sort the multiple initial answer bases according to the RAG model token limit and the output scores, and complete the deduplication and filtering process to obtain multiple processed initial answer bases; use natural language technology to respectively streamline and summarize the content of each processed initial answer basis to obtain the final answer basis; The RAG model answer generation acquisition module 105 is used to generate an answer using the RAG model according to the question and the final answer basis.

[0060] In a specific embodiment, the knowledge base sharding adjustment module 101 is further used to, on the basis of adjusting the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text sharding processing method, continue to complete the vectorization storage of each piece of text according to the vectors generated by dividing the large and small text blocks, and generate a knowledge base that allows indexing the vector of the large text block corresponding to the vector of the small text block; The RAG model retrieval optimization module 102 is further used to obtain the question-and-answer pairs containing the answer basis stored in the form of the corresponding vectors of the large and small text blocks in the bidding scenario to replace the question-and-answer pairs containing the answer basis stored in the text sharding form in the bidding scenario as the second verification set, and then use the second verification set to perform retrieval optimization on the RAG model; The RAG model initial inspection result acquisition module 103 is further used to, during the retrieval process of the RAG model optimized by retrieval in the knowledge base associated with the bidding scenario, when the corresponding vector of the small text block is retrieved, index the corresponding vector of the relevant large text block, so as to obtain multiple initial answer bases containing the large text block.

[0061] In a specific embodiment, the knowledge base sharding adjustment module 101 is further used to adjust the text form of the answer basis in the knowledge base associated with the bidding scenario according to different text sharding processing methods and correspondingly generate different types of knowledge bases so that each type of knowledge base stores the answer basis whose text form is adjusted according to a single text sharding processing method; the different text sharding processing methods include: clauses, paragraphs, and chapters; The RAG model retrieval optimization module 102 is further used to obtain the question-and-answer pairs containing the answer basis stored in the text form generated by single text sharding processing in the bidding scenario as the third verification set, use the third verification set to perform retrieval optimization on the RAG model, and correspondingly obtain several RAG models optimized by retrieval; The RAG model initial inspection result acquisition module 103 is further configured to, for the expanded target question, use machine learning algorithms to evaluate the complexity of the expanded question and the importance of the information in the expanded question respectively, obtain the weighted score result value of the question complexity score and the question information importance score, and match the corresponding preset text sharding processing method according to the range of the weighted score result value; match the retrieved and optimized RAG model according to the matched preset text sharding processing method, and use the matched retrieved and optimized RAG model to retrieve in the knowledge base whose stored answer basis text form is the same as the text form generated by the preset text sharding processing method to obtain multiple initial answer bases.

[0062] In a specific embodiment, the RAG model final inspection result acquisition module 104 is further configured to, when adjusting the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text sharding processing method, use natural language analysis technology to analyze the logical coherence and information relevance of multiple text shards, and evaluate the logical coherence strength and information relevance strength of multiple text shards with logical coherence and / or information relevance respectively; introduce a marking mechanism to display and distinguish multiple text shards with different logical coherence strengths and information relevance strengths; after content simplification and generalization for each processed initial answer basis, sort and reorganize each content-simplified initial answer basis according to the coherence strength and information relevance strength marked in the original text content of the initial answer basis, so that the answer basis text fragments belonging to the context content are sorted adjacent to each other and the answer basis text fragments with stronger information relevance content strength are sorted closer, to obtain the final answer basis.

[0063] In a specific embodiment, it further includes: the RAG model generated answer verification module 106, which is configured to receive the feedback data of the user on the generated answer in real time. When the user's satisfaction with the feedback is lower than the preset satisfaction, the initial answer basis is correspondingly displayed and the user is prompted to evaluate the satisfaction with the initial answer basis, and the satisfaction data of the user with the initial answer basis is retained; when the frequency of the user's satisfaction being lower than the preset satisfaction within the preset time period is greater than the preset frequency, it is judged whether the corresponding statistical frequency of the user's satisfaction with the initial answer basis being lower than the preset satisfaction retained within the preset time period is greater than the preset frequency. If so, the retrieved and optimized RAG model is selected for optimization training; otherwise, the initial answer basis scoring model is optimized for training.

[0064] A specific embodiment further includes: a knowledge base content supplement module 107, configured to expand information for the answer basis in the knowledge base associated with the bidding scenario before adjusting the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text slicing processing method, including: supplementing information by using the bidding text content collected by the knowledge graph in the bidding scenario, or supplementing information by collecting the text content converted from the multimodal data in the bidding documents in the bidding scenario.

[0065] An embodiment of the present application also discloses a computer-readable storage medium.

[0066] Specifically, the computer-readable storage medium stores a computer program that can be loaded and executed by a processor, such as the method for improving the question-answering accuracy of the RAG-enhanced small-scale language model as described above. The computer-readable storage medium includes, for example, various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0067] An embodiment of the present application also discloses a computer device.

[0068] Specifically, the computer device includes a memory and a processor, and the memory stores a computer program that can be loaded and executed by the processor, such as the method for improving the question-answering accuracy of the RAG-enhanced small-scale language model as described above.

[0069] The above are all preferred embodiments of the present application. Without limiting the protection scope of the present application accordingly, any feature disclosed in this specification (including the abstract and drawings), unless specifically described, can be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically described, each feature is only an example of a series of equivalent or similar features.

Claims

1. A RAG-enhanced small-scale language model question answering accuracy improvement method, characterized in that: include: Adjust the text format of the answer basis in the knowledge base associated with the bidding scenario according to the text segmentation processing method; Obtain question-answer pairs containing answer evidence stored in the form of text fragments in the bidding scenario as the first verification set, and use the first verification set to perform retrieval optimization on the RAG model; Using natural language technology, the content or quantity of the questions input into the RAG model is expanded, and the expanded questions are used as input. The RAG model after retrieval optimization is used to search in the knowledge base associated with the bidding scenario to obtain multiple initial answer bases; Based on the deep learning algorithm, an initial answer basis scoring model is constructed under the RAG model framework, and multiple initial answer bases are output to the initial answer basis scoring model respectively. The multiple initial answer bases are screened and sorted according to the RAG model token restriction and the output score, and the deduplication and filtering process is completed to obtain the processed multiple initial answer bases; Using natural language technology, the content of each processed initial answer basis is succinctly summarized to obtain the final answer basis; Based on the questions and the basis for the final answer, the RAG model is used to generate answers.

2. The RAG enhanced small-scale language model question answering accuracy improvement method according to claim 1, characterized in that: Also includes: On the basis of adjusting the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text segmentation processing method, the vectorized storage of each piece of text is completed according to the vector generated by converting the divided large and small blocks of text, and a knowledge base is generated that allows the vector corresponding to the small block of text to index the vector corresponding to the large block of text; Obtain answer pairs containing answer bases stored in the form of large and small block text corresponding vectors in the bidding scenario to replace the question-answer pairs containing answer bases stored in the form of text fragments in the bidding scenario as the second verification set, and then use the second verification set to perform retrieval optimization on the RAG model; In the process of searching in the knowledge base associated with the bidding scenario using the optimized RAG model, when a vector corresponding to a small block of text is retrieved, the vector corresponding to the related large block of text is indexed, thereby obtaining multiple initial answer bases containing the large block of text.

3. The RAG enhanced small-scale language model question answering accuracy improvement method according to claim 1, characterized in that: Also includes: The text form of the answer basis in the knowledge base associated with the bidding scenario is adjusted according to different text segmentation processing methods, and different types of knowledge bases are generated accordingly so that each type of knowledge base stores the answer basis in the text form adjusted according to a single text segmentation processing method; the different text segmentation processing methods include: clauses, paragraphs, and chapters; Obtain a question-answer pair containing answer evidence stored in a text format generated by processing a single text fragment in a bidding scenario as a third validation set, use the third validation set to perform retrieval optimization on the RAG model, and correspondingly obtain a RAG model using several retrieval optimizations; For the expanded target question, the machine learning algorithm is used to evaluate the complexity of the expanded question and the importance of the expanded question information, and the weighted scoring result values ​​of the question complexity score and the question information importance score are obtained. The preset text segmentation processing method corresponding to the scoring result value range is matched according to the weighted scoring result value; According to the matching preset text segmentation processing method, a RAG model optimized for retrieval is matched. The RAG model optimized for retrieval is used to search in a knowledge base whose stored answer basis text form is the same as the text form generated by the preset text segmentation processing method to obtain multiple initial answer bases.

4. The RAG enhanced small-scale language model question answering accuracy improvement method according to claim 1, characterized in that: Also includes: When adjusting the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text segmentation processing method, the natural language analysis technology is used to analyze the logical coherence and information relevance of multiple text segments, and the logical coherence strength and information relevance strength of multiple text segments with logical coherence and / or information relevance are evaluated respectively; a marking mechanism is introduced to display and distinguish multiple text segments with different logical coherence strengths and information relevance strengths; After concisely summarizing the content of each processed initial answer basis, sort and reorganize the context content and information-related content of each simplified initial answer basis according to the coherence strength and information-related strength marked in the text content of the original initial answer basis, so that the answer basis text fragments belonging to the context content are sorted adjacently and the answer basis text fragments with stronger information-related content are sorted closer, so as to obtain the final answer basis.

5. The RAG enhanced small-scale language model question answering accuracy improvement method according to claim 1, characterized in that: Also includes: Receive user feedback data on the generated answers in real time. When the user's satisfaction with the feedback is lower than the preset satisfaction, the initial answer basis is displayed accordingly and the user is prompted to evaluate the satisfaction with the initial answer basis, and the user's satisfaction data on the initial answer basis is retained; When the frequency of user satisfaction being lower than the preset satisfaction within the preset time period is greater than the preset frequency, determine whether the frequency of user satisfaction with the basis for the initial answer retained within the corresponding statistical preset time period is lower than the preset satisfaction is greater than the preset frequency. If so, select the RAG model after retrieval optimization for optimization training; otherwise, optimize the initial answer based on the scoring model.

6. The RAG enhanced small-scale language model question answering accuracy improvement method according to claim 1, characterized in that: Also includes: Before adjusting the text form of the basis for the answer in the knowledge base associated with the bidding scenario according to the text segmentation processing method, the information of the basis for the answer in the bidding scenario knowledge base is expanded, including: using the knowledge graph in the bidding scenario to collect related bidding text content to supplement the information, or collecting text content converted from multimodal data in the bidding documents in the bidding scenario to supplement the information.

7. The RAG enhanced small-scale language model question answering accuracy improvement method according to claim 1, characterized in that: Also includes: Adjusting the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text segmentation processing method includes: segmenting the text content of the answer basis in the knowledge base associated with the bidding scenario according to the preset text segmentation processing method, and adding the outline title content to which the segment belongs before each piece of text form of the answer basis.

8. A RAG-enhanced small-scale language model question answering accuracy improvement system, characterized in that: include: A knowledge base sharding adjustment module is used to adjust the text form of the answer basis in the knowledge base associated with the bidding scenario according to the text sharding processing method; The RAG model retrieval optimization module is used to obtain the question-answer pairs containing the answer basis stored in the form of text fragments in the bidding scenario as the first verification set, and use the first verification set to perform retrieval optimization on the RAG model; The RAG model initial inspection result acquisition module is used to use natural language technology to expand the content or quantity of the questions input into the RAG model, take the expanded questions as input, and use the optimized RAG model to search in the knowledge base associated with the bidding scenario to obtain multiple initial answer bases; The RAG model final inspection result acquisition module is used to build an initial answer basis scoring model under the RAG model framework based on a deep learning algorithm, output multiple initial answer bases to the initial answer basis scoring model respectively, filter and sort the multiple initial answer bases according to the RAG model token restriction and the output score, and complete the deduplication and filtering process to obtain the processed multiple initial answer bases; Using natural language technology, the content of each processed initial answer basis is succinctly summarized to obtain the final answer basis; The RAG model generates an answer acquisition module, which is used to generate answers using the RAG model based on the questions and the final answer basis.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

10. A computer device, characterized in that: The computer device comprises a memory, a processor and a program stored and executable on the memory, and the program implements the steps of the method according to any one of claims 1 to 7 when executed by the processor.

Citation Information

Patent Citations

  • Retrieval method and device of knowledge base and storage medium

    CN116756295A

  • Online intelligent question answering method and device based on instruction fine tuning and retrieval enhancement generation

    CN117688163A

  • Retrieval enhancement method and device, equipment and storage medium

    CN118394793A

  • RAG-based vertical domain knowledge multi-round question and answer method

    CN118964556A

  • Urban rail transit emergency field question and answer method and system based on optimized RAG

    CN119311793A