Method for improving recall accuracy of article document by using RAG technology
Through the knowledge construction and recall process of RAG technology, combined with large models and vectorization technology, the general big model has solved the shortcomings in the recall accuracy and data security of article-based documents, achieved higher recall accuracy and data security, and reduced training costs.
Patent Information
- Application Number
- CN202510301392.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The existing general model has low recall accuracy when processing article documents, which is difficult to meet the needs of real-time data and private data security. At the same time, the training cost is high and the data timeliness is poor.
Using RAG technology, we use knowledge construction process and knowledge recall process, including knowledge preprocessing, sharding, keyword indexing, text vectorization, big model QA enhancement and vectorization, and combine rerank and big model problem rewriting technology to improve the recall accuracy of article-based documents.
It significantly improves the recall accuracy of article-based documents, supports the accuracy of upper-level intelligent applications in answering user questions, has strong interpretability, data security and adaptability, and reduces training costs.
Smart Images

Figure CN120216708A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of RAG retrieval augmented generation technology and large model technology, and particularly relates to a method for improving the recall accuracy of regulatory document using RAG technology. Background Art
[0002] Currently, general large models cannot meet the needs of actual business, and mainly have the following defects: 1) Limitation of knowledge: The knowledge of the model itself completely comes from its training data, and the training sets of existing mainstream large models are basically constructed from publicly available data on the Internet. For some real-time, non-public or offline data, it is impossible to obtain, and this part of knowledge cannot be possessed; 2) Hallucination problem: The underlying principle of all AI models is based on mathematical probability, and the model output is essentially a series of numerical operations. The same is true for large models. Therefore, it sometimes talks nonsense seriously, especially in scenarios where the large model itself does not have knowledge in a certain aspect or is not good at. And it is relatively difficult to distinguish this hallucination problem because it requires the user to have knowledge in the corresponding field; 3) Data security: For enterprises, data security is crucial. No enterprise is willing to take the risk of data leakage and upload its private domain data to a third-party platform for training. This also leads to the application solutions that completely rely on the capabilities of general large models having to make a trade-off between data security and effect; 4) Training cost problem: Small and medium-sized enterprises or large models deployed in the private domain lack professional large model personnel and are difficult to bear the high training costs; 5) Data timeliness problem: Limited by the lag of training texts and the training cycle, large models are difficult to give accurate answers based on the latest data.
[0003] Facing the above problems, the RAG retrieval augmented generation technology is the current mainstream solution in the industry. Retrieval-augmented Generation, abbreviated as RAG, is one of the current popular cutting-edge large model technologies. It generates answers or content by referring to information in an external knowledge base, and has strong interpretability and customization capabilities. It has various advantages such as effectively utilizing the external knowledge base, good timeliness of data update, interpretability and traceability of responses, protecting the security and privacy of private data, reducing training costs, and being suitable for domain customization.
[0004] Regulations, laws, enterprise regulations, etc. all have clear regulatory formats, but the general RAG recall method has poor effects for such scenarios. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a method for improving the recall accuracy of regulatory document using RAG technology, which can significantly improve the recall accuracy of regulatory document and further support the accuracy of the upper-layer intelligent agent application to answer user questions.
[0006] The technical solution of the present invention is as follows:
[0007] A method for improving the recall accuracy of regulatory text documents using the RAG technology, which is used to improve the recall accuracy of regulatory text documents, includes the following steps: 1) The knowledge construction process, including knowledge preprocessing, knowledge sharding, keyword index extraction, text vectorization, QA enhancement based on large models, and vectorization; 2) The knowledge recall process, including question preprocessing, keyword index, vector index, rerank, large model question rewriting & re - recall, and post - recall processing; 3) Provide an interface for operating the knowledge base.
[0008] Furthermore,
[0009] In the document preprocessing step of the knowledge construction process, a document format conversion module is provided to convert pdf and doc documents into the docx format of word.
[0010] In the knowledge sharding step, a knowledge sharding function is adopted. By parsing the content of the word document and combining configurable recognition rules, the accuracy of knowledge recall is improved.
[0011] In the knowledge sharding step, exception enhancement processing is provided. For fragments with a length greater than a specific threshold after sharding, rule matching and slicing processing are performed.
[0012] In the knowledge sharding step, a degree identification method and a method for chapter merging according to the degree are provided.
[0013] In the keyword extraction step, a keyword extraction process is provided, and the keywords are associated with the fragments through unique fragment identifiers.
[0014] In the text vectorization step, a vectorization enhancement process is provided. The shards after cleaning are vectorized, and the vector values are stored. The cleaning strategies include but are not limited to the following strategies: r"[^\u4e00-\u9fff\A - Za - z0 - 9]". Support but are not limited to using the following vector libraries for vector storage and lookup: Weaviate, Qdrant, milvus, pgvector.
[0015] In the QA enhancement and vectorization step based on large models, the large model service is called using a formatted prompt template to return questions related to the fragments. A prompt template with the following content is provided:
[0016] System:Please help me predict the three most likely questions that human would ask,and keeping each question under 30 characters.
[0017] The output must be an array in JSON format following the specified schema:
[0018] ["question1","question2","question3"]
[0019] User:{Article fragment}
[0020] A question preprocessing method in the knowledge recall process.
[0021] Including the following cleaning rules:
[0022] r"[’·°–!"#$%&\'()*+,-. / :;<=>?@,。?★、…【】()《》?“”‘’![\\]^_`{|}~]+?"
[0023] An optional keyword indexing step in the knowledge recall process, including tokenizing the cleaned question and calculating keyword similarity using the BM25 algorithm.
[0024] An optional vector indexing step in the knowledge recall process, including vectorizing the cleaned question and performing vector lookup.
[0025] The rerank enhanced indexing step in the knowledge recall process, merging and rearranging the results of keyword indexing and vector indexing; using the rerank model of the bge series for rearrangement.
[0026] The large model question rewriting step. If the knowledge fragment fails to be recalled successfully, the large model is called to rewrite the question. For the question rewritten by the large model, the process from question preprocessing to rerank enhanced indexing is executed again.
[0027] The post-recall processing step in the knowledge recall process, including obtaining the original format of the fragment according to the fragment identifier, performing regular matching and recombining processing on the original format, and returning the technology.
[0028] Furthermore,
[0029] A knowledge sharding method for improving the accuracy of document knowledge recall, characterized by including the following steps: 1) Parse the word document file name, remove the file extension and special symbols; parse the XML structure of the word document to obtain the document chapter shard set D, and each shard is denoted as D i . 2) Provide a set of configurable recognition rule combinations RG, and each rule in it is denoted as RG xy , where x is the rule group number and y is a certain rule matched in the rule group. 3) Process each obtained document shard D i in turn in a loop, and obtain the first regular rule RG i matched by the document shard D xy , and the rule group RG xy where the rule RG x is located. Calculate the matching times of each rule group if_it(D i , RG x ) indicates whether the document shard D i can match the rule group RG x . If it matches, its value is 1, otherwise it is 0. Select the rule group with the highest matching times H(x) of the rule group as the best matching rule group RG max of the document structure. 4) Process each obtained document shard D i in a loop, and obtain the document shard D i , and a rule RG max matched in the rule group RG xy . Take y as the chapter level of this fragment, that is, for each fragment D i , its level Y i = y. 5) Provide exception enhancement processing. For the fragments with a length greater than the threshold α after sharding, call the rules in the rule group R2 to match them in turn. 6) Process each shard as follows: Slice according to the first matched rule. After slicing, the level of the fragment is increased by 1 compared to the level of the parent shard, and an empty fragment with the level of the parent shard is inserted after this fragment; loop to judge the fragments after sharding until all fragments are less than the threshold α or all rules cannot match the paragraph. 7) Provide exception enhancement processing. For the fragments with a length still greater than the threshold β after the above sharding, split them according to the carriage return character "\n". If the length after splitting is less than the threshold γ, merge the shards into the previous shard. 8) Shard aggregation processing. Process each document shard D i in a loop, and merge several fragments with the same level into one fragment. Finally, obtain the shard set S. 9) Assign a unique shard identifier S i to the fragments S i after sharding; the fragment S iSimultaneously store the fragment identifier of its previous adjacent fragment, marked as S i _ID_PREV; Fragment S i Simultaneously store the fragment identifier of its next adjacent fragment, marked as S i _ID_NEXT. 10) For the fragmented fragment S i , perform rewriting and write the file name to the first line of fragment Si
[0030] The aforementioned recognition rule combination RG xy , where x is the rule group number and y is a certain rule matched in the rule group, including but not limited to the following rule combinations:
[0031]
[0032]
[0033] The aforementioned rule group R2 includes the following rules:
[0034]
[0035] A set of interfaces for operating the knowledge base, including a file upload interface and a knowledge recall call interface
[0036] Among them, the file upload interface is implemented in the Python language and includes the following call methods:
[0037]
[0038] The knowledge recall interface is implemented in the Python language and includes the following call methods:
[0039]
[0040]
[0041] The beneficial effects of the present invention are
[0042] 1. Higher knowledge recall accuracy: For clause-based documents, the recalled paragraphs have a higher relevance to the user's question and are more accurate; 2. More complete knowledge recall semantics: For clause-based documents, the entire question content can be recalled more completely; 3. Strong explanatory power of answers: The answers are based on complete evidence and can clearly prompt the user to obtain corresponding answers based on XX regulations and XX articles; 4. Stronger customizability and adaptability: Through configurable rule groups, it can better adapt to the document characteristics in different scenarios and improve the segmentation accuracy; 5. Obvious advantages compared with the pure large model Q&A solution: It has various advantages such as effectively utilizing external knowledge bases, good timeliness of data update, interpretability and traceability of replies, protecting the security and privacy of private data, reducing training costs, and being suitable for domain customization Brief Description of the Drawings
[0043] Figure 1 is the flowchart of the knowledge construction process;
[0044] Figure 2 is the flowchart of the knowledge recall process. Detailed Implementation Manner
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] The present invention proposes a method for improving the recall accuracy of document-based documents using the RAG technology, which integrates multiple technologies such as traditional keyword indexing technology, vectorization technology, rerank technology, and large model technology, greatly improving the accuracy of knowledge Q&A.
[0047] The present invention is divided into a knowledge construction process and a knowledge recall process. The knowledge construction process is as follows: knowledge preprocessing → knowledge sharding → keyword index extraction → text vectorization → QA enhancement and vectorization based on a large model; the knowledge recall process is as follows: question preprocessing → keyword index → vector index → rerank → large model question rewriting & re-recall → post-recall processing.
[0048] S1. Knowledge Construction
[0049] S101. Knowledge Preprocessing
[0050] Perform document preprocessing, and provide a document format conversion module to convert pdf and doc documents into the docx format of word.
[0051] S102. Knowledge Sharding
[0052] Traditional sharding methods, such as: using the token length splitting method, often result in incomplete paragraph splitting; using the carriage return and line feed character splitting method, the semantic integrity of the context paragraphs is not good; using the docx document chapter splitting method, the requirements for word documents are relatively high, etc. In the scenario of document-based documents, the practical effects are not good.
[0053] This example provides a knowledge sharding function, which can effectively improve the recall accuracy of knowledge. Its core processing logic is as follows:
[0054] S10201. Analyze the file name of the word document, remove the file extension and special symbols; analyze the XML structure of the word document to obtain the document chapter fragment set D, and each fragment is denoted as D i .
[0055] S10202. Provide a set of configurable recognition rule combinations RG, and each rule in it is denoted as RG xy , where x is the rule group number and y is a certain rule matched in the rule group. By default, the following regular rule combinations are provided but not limited to:
[0056]
[0057]
[0058] S10203. Process each obtained document fragment D in turn i , obtain the document fragment D i The first regular rule RG matched xy , and the rule group RG where the rule RG xy is located x . Calculate the matching times of each rule group if_it(D i , RG x ) indicates whether the document fragment D i , can match the rule group RG x . If it matches, its value is 1, otherwise it is 0. Select the rule group with the highest matching times H(x) of the rule group as the best matching rule group RG of the document structure max .
[0059] Process each obtained document fragment D in turn i , obtain the document fragment D i , and match a rule RG max matched in the rule group RG xy , and take y as the chapter level of this fragment, that is, for each fragment D i , its level Y i =y
[0060] S10204. Provide exception enhancement processing. For fragments with a length greater than the threshold α after fragmentation, call the rules in the rule group R2 to match in turn. By default, the following regular rule combinations are provided but not limited to in R2:
[0061]
[0062] Each shard is processed in sequence as follows: according to the first matched rule, perform slicing. After slicing, the level of the shard is increased by 1 compared to the level of the parent shard. Meanwhile, an empty shard with the level of the parent shard is inserted after this shard. Continuously judge the shards after slicing until all shards are smaller than the threshold α or all rules cannot match the paragraph.
[0063] S10205. Provide exception enhancement processing. For the shards that are still longer than the threshold β after the above slicing, split them according to the carriage return character "\n". If the length after splitting is smaller than the threshold γ, then merge the shards into the previous shard.
[0064] S10206. Shard aggregation processing. Process each document shard D in a loop i , and merge several shards with the same level into one shard. Finally, obtain the shard set S.
[0065] S10207. For the shards S after slicing i , assign a unique shard identifier S i _ID; for the shard S i simultaneously store the shard identifier of its previous adjacent shard, marked as S i _ID_PREV; for the shard S i simultaneously store the shard identifier of its next adjacent shard, marked as S i _ID_NEXT.
[0066] S10208. For the shards S after slicing i , perform rewriting and write the file name to the first line of the shard S i .
[0067] S103. Extract keywords
[0068] S10301. Provide a keyword extraction process. Perform word segmentation on the shard S i , and the word segmentation is associated with the unique shard identifier S i _ID and the shard S i .
[0069] S104. Text vectorization
[0070] S10401. Provide a vectorization enhancement process. For the sharded paragraph S i , perform cleaning to obtain the cleaned shard SC i , and the cleaning strategies include but are not limited to the following strategies:
[0071] r"[^\u4e00-\u9fff\A-Za-z0-9]"
[0072] For the shard SC i and the shard S iAssociated by the unique fragment identifier S i _ID.
[0073] S10402. Provide a vectorization processing module to vectorize the cleaned shard SC i and store the vector values. Provide a multi-vector library support module that supports, but is not limited to, the following vector libraries: Weaviate, Qdrant, milvus, pgvector. The vectors are associated with the fragment S i by the unique fragment identifier S i _ID.
[0074] S105. QA Enhancement and Vectorization Based on Large Models
[0075] S10501. Use a formatted prompt template to call the large model service and require the large model to return 3 questions related to the provided fragment S i . The provided prompt includes the following content:
[0076]
[0077] S10502. Vectorize the relevant questions returned by the large model. The vectors are associated with the fragment S i by the unique fragment identifier S i _ID.
[0078] S106. An Example of Sharding Effect
[0079]
[0080]
[0081] S2. Knowledge Recall
[0082] In this case, through the previously constructed keywords and vector indexes, combined with the rerank recall technology and the large model question rewriting technology, a more accurate recall method is provided.
[0083] S201. Question Preprocessing
[0084] Provide a question preprocessing method, including the following cleaning rules:
[0085]
[0086] S202. Keyword Index
[0087] Optionally, if the user requests keyword indexing, the cleaned question is tokenized. The BM25 algorithm is used to calculate the keyword similarity. First, morpheme parsing is performed to generate morphemes qi. Then, for each search result D, the relevance score of each morpheme qi to D is calculated. Finally, the relevance scores of qi relative to D are weighted and summed to obtain the relevance score of the Query to D. According to the user-specified index topK parameter and threshold score_threshold, the top K items with the highest scores whose relevance scores are greater than score_threshold are returned.
[0088] S203. Vector Indexing
[0089] Optionally, if the user requests vector indexing, the cleaned question is vectorized. Vector search is performed, and according to the user-specified index topK parameter and threshold score_threshold, the top K items with the highest scores whose relevance scores are greater than score_threshold are returned.
[0090] S204. Rerank Enhanced Indexing
[0091] After merging the results returned by keyword indexing and vector indexing, the rerank model is called for rearrangement. According to the user-specified index topK parameter and threshold score_threshold, the top K items with the highest scores whose relevance scores are greater than score_threshold are returned. In the present invention, the bge series of rerank models are used by default.
[0092] S205. Large Model Question Rewriting
[0093] If the above processing fails to successfully recall the knowledge fragments, the large model is called to require the large model to rewrite the question. For the question rewritten by the large model, the above process is executed again to attempt to recall the knowledge.
[0094] S206. Post-Processing after Recall
[0095] For each recalled fragment, according to the fragment identifier S i _ID, the original format fragment S i is obtained. For the obtained fragment S i , it is rematched according to the regular rules in S i . If no match is found, one associated fragment before and after it is obtained and combined into a new fragment for return.
[0096] S207. An Example Recall Effect
[0097]
[0098]
[0099] S3. An interface for operating a knowledge base
[0100] S301. The present invention is implemented in the Python language.
[0101] S302. The present invention includes an interface for file upload, and its calling methods include but are not limited to the following:
[0102]
[0103] S303. The present invention includes an interface for knowledge recall call, and its calling methods include the following:
[0104]
[0105]
[0106] The above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for improving the recall accuracy of clause documents using RAG technology, characterized in that: The following steps are involved: 1) Knowledge construction process, including document preprocessing, knowledge segmentation, keyword index extraction, text vectorization, QA enhancement and vectorization based on large models; 2) Knowledge recall process, including question preprocessing, keyword indexing, vector indexing, reranking, large model question rewriting & re-recall, and post-recall processing; 3) Provide an interface for operating the knowledge base.
2. The method according to claim 1, characterized in that The document preprocessing step in the knowledge construction process provides a document format conversion module to convert PDF and doc documents into word's docx format.
3. The method according to claim 1, characterized in that In the knowledge segmentation step, the knowledge segmentation function is used to improve the accuracy of knowledge recall by parsing the content of the word document and combining configurable recognition rules. Provides exception enhancement processing, performs rule matching and slicing processing for segments whose length after segmentation is greater than a specific threshold; provides a level identification method, and merges chapters according to the level.
4. The method according to claim 3, characterized in that Knowledge sharding includes the following steps: 1) Parse the word document file name, remove the file extension and special symbols; parse the XML structure of the word document, obtain the document chapter fragment set D, each fragment is denoted as D i ; 2) Provide a set of configurable identification rule combinations RG, where each rule is denoted as RG xy ,x is the rule group number, y is a rule matched in the rule group; 3) Loop through each document fragment D obtained i , get document fragment D i The first matched regular rule RG xy , and rule RG xy Rule group RG x ; Calculate the number of matches for each rule group if_it(D i , R.G. x ) represents document fragment D i , and the rule group RG x Whether it can be matched, if it is matched, its value is 1, otherwise it is 0; the rule group with the highest number of rule group matches H(x) is taken as the best matching rule group RG for the document structure max ; 4) Loop through each document fragment D obtained i , get document fragment D i , in the rule group RG max , a rule RG is matched in xy , taking y as the chapter level of the segment, that is, for each segment D i , its level Y i =y; 5) Provide exception enhancement processing. For fragments whose length is greater than the threshold α after fragmentation, call the rules in the rule group R2 to match them in sequence; 6) Each fragment is processed in turn as follows: according to the first matched rule, it is sliced, and the degree of the sliced fragment is the degree of the parent fragment increased by 1. At the same time, an empty fragment is inserted after the fragment, and the degree of the empty fragment is the degree of the parent fragment; the fragments after the slice are judged cyclically until all fragments are smaller than the threshold α or all rules cannot match the paragraph; 7) Provide exception enhancement processing. For the segments whose length is still greater than the threshold β after the above segmentation, segment them according to the carriage return character "\n". If the length after segmentation is less than the threshold γ, merge the segment into the previous segment; 8) Shard aggregation processing, loop processing each document shard D i , merge several fragments of the same level into one fragment; finally, the fragment set S is obtained; 9) After the fragmentation, i , assign a unique fragment identifier S i _ID; fragment S i At the same time, the segment identifier of the immediately preceding segment is stored, marked as S i _ID_PREV; fragment S i At the same time, the segment identifier of the next segment is stored, marked as S i _ID_NEXT. 10) For the fragmented segment S i , rewrite and write the file name into fragment S i The first line of .
5. The method according to claim 1, characterized in that In the keyword extraction step, a keyword extraction process is provided, and keywords are associated with segments through unique segment identifiers.
6. The method according to claim 1, characterized in that In the text vectorization step, a vectorization enhancement process is provided to vectorize the cleaned fragments and store the vector values; the cleaning strategies include: r"[^\u4e00-\u9fff\A-Za-z0-9]"; the following vector libraries are used for vector storage and search: Weaviate, Qdrant, milvus, and pgvector.
7. The method according to claim 1, characterized in that In the QA enhancement and vectorization step based on the big model, a formatted prompt template is used to call the big model service to return questions related to the fragment.
8. The method according to claim 1, characterized in that In the knowledge recall process, The keyword indexing step includes word segmentation of the cleaned questions and calculating keyword similarity using the BM25 algorithm; The vector indexing step includes vectorizing the cleaned problem and performing vector lookup; Rerank enhances the indexing step, merging and re-ranking the results of keyword indexing and vector indexing; using the bge series rerank model for re-ranking; The post-recall processing step includes obtaining the fragment in the original format according to the fragment identifier, performing regular matching on the original format, and recombining processing and returning the technology.
9. The method according to claim 1 or 8, characterized in that: In the big model question rewriting step, if the knowledge fragment cannot be successfully recalled, the big model is called to rewrite the question; for the question rewritten by the big model, the process from question preprocessing to rerank enhanced indexing is re-executed.
10. The method according to claim 1, characterized in that The interfaces used to operate the knowledge base, including the file upload interface and the knowledge recall call interface, are all implemented in Python.
Citation Information
Patent Citations
Intelligent extraction method and system for key knowledge points of document based on docx format
CN117764051A
Retrieval enhancement method based on semantic comprehension and semantic generation model
CN118733715A
Large model retrieval enhancement generation method and system
CN119294516A