RAG retrieval generation system oriented to intensive data scene
Through the collaborative work of the block database module, context construction module and related information aggregation module, the semantic relationship capture and multi-document information aggregation problems of the RAG search and generation system in intensive data scenarios are solved, efficient and accurate retrieval and generation are achieved, and user experience and system reliability are improved.
Patent Information
- Application Number
- CN202510783502.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing RAG search and generation system is difficult to accurately capture the complex semantic relationship between queries and documents in intensive data scenarios, resulting in lack of coherence and correlation of search results, inaccurate content of generation, and ineffective multi-document information, affecting user experience and system reliability.
The block database module is used to automatically generate the title, summary and paragraph semantic summary of the document, and store these metadata information in the database. The structured context construction module of query results is enhanced through the context construction module. The relevant information aggregation module extracts relevant text fragments from multiple query results after searching, and combines the optimization algorithm to adjust the fragment correlation and length to provide accurate metadata answers.
It significantly improves retrieval efficiency and accuracy, enhances the adaptability and intelligence of the system in intensive data scenarios, provides accurate and comprehensive answers, reduces the illusion that the generated content deviates from the topic, and improves user experience and system reliability.
Smart Images

Figure CN120296138A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of RAG retrieval and generation systems, and more particularly to a RAG retrieval and generation system for intensive data scenarios. Background Art
[0002] In intensive data scenarios, existing RAG (Retrieval-Augmented Generation) technologies face numerous challenges. First, its retrieval module typically relies on traditional vector retrieval methods, such as BERT-based dense retrieval, which is difficult to accurately capture the complex semantic relationships between queries and documents. As a result, the retrieval results often do not match the user's intent, leading to returned segments lacking coherence or relevance, and even context loss.
[0003] Second, when the retrieval module fails to hit relevant documents, the generator may produce inaccurate or "hallucination" content, which seriously affects the user experience and system reliability.
[0004] In addition, when dealing with multi-document and multi-structured data, existing RAG technologies often cannot effectively aggregate relevant information from different documents, resulting in data bias and inaccurate generated content.
[0005] To address these issues, the present invention proposes a RAG retrieval and generation system for intensive data scenarios. The system automatically generates the title, abstract, and paragraph semantic summaries of documents through the block database module, and stores these metadata information together with the documents in the database, supporting multi-way queries for retrieval, thereby improving retrieval efficiency and accuracy. At the same time, through the Relevant Information Aggregation module, the most relevant text segments are extracted from multiple query results during the post-retrieval generation stage and combined into a meta-document to provide users with accurate and comprehensive answers. These innovations enable the present invention to significantly improve the accuracy of knowledge base Q&A data in intensive data scenarios such as financial reports and operation and maintenance documents, addressing the deficiencies existing in the prior art. Therefore, those skilled in the art have provided a RAG retrieval and generation system for intensive data scenarios to solve the problems raised in the above background art. Summary of the Invention
[0006] The present invention provides the following technical solutions: A RAG retrieval and generation system for intensive data scenarios, comprising: a block database module, a context construction module, a relevant information aggregation module, and an optimization algorithm module; The block database module, which is the SegmentDB module, is used to automatically generate the title, abstract, and paragraph semantic summaries of documents, and store the metadata information of the documents and their respective semantic segments in the database, providing multi-way queries for retrieval; Context construction module, which constructs a structured context for the document and its semantic segments before vectorization or question answering generation to enhance the accuracy of query results; Relevant information aggregation module, which extracts relevant text fragments from multiple query results during the post-retrieval generation stage, combines these fragments into a meta-document, and provides accurate answers; Optimization algorithm module, which adjusts the relevance and length of the generated fragments based on the association value to further optimize the accuracy of the answers.
[0007] Furthermore, the block database module generates the title, abstract, and chapter abstract of the document through the Enhance_context function and adds them to each document block to achieve the storage of metadata information of the document and multi-way query; The context construction module constructs a structured context for the document and its semantic segments before vectorization or question answering generation to enhance the accuracy of query results.
[0008] Furthermore, the context construction module adopts the method of semantic chunking plus hierarchical segmentation to divide the document into chunks, combines the chunk header and chunk content to construct structured context information, constructs a structured context for the document and its semantic segments before vectorization or question answering generation, automatically generates the document title, document abstract, paragraph semantic abstract, and constructs the structure prompt word header, uses the LLM output result as the document title and abstract, performs summary generation on each semantic segment in the document, concatenates the document title, document abstract, paragraph title, and abstract as the context prompt header, combines it with the chunk body as the final input, significantly improves the representation quality of the embedding model, enables the generation model to have a complete structural awareness, reduces off-topic hallucinations, enhances the adaptability to the private domain operation and maintenance scenario, and automatically establishes a hierarchical index and context for complex documents; The relevant information aggregation module extracts relevant text fragments from multiple query results during the post-retrieval generation stage, combines these fragments into a meta-document, provides accurate answers, extracts the most relevant text fragments from multiple queries during the post-retrieval generation stage, combines them into a meta-document, calculates the relevance value of the fragments using an exponential decay function, controls the length and relevance of the fragments through parameter configuration, traverses the ranking results of each query, calculates the relevance value of each fragment, adjusts the relevance value according to the fragment length, selects the optimal non-overlapping fragments that meet the length limit using a greedy algorithm, divides them into highly relevant fragments and weakly relevant fragments, ensures data quality while meeting the context of the large model, selects text fragments that are both relevant and meet the length limit, and adapts to different application scenarios through parameter configuration, such as short text in FAQ, long operation and maintenance documents, multi-turn question answering, and structured response generation.
[0009] Furthermore, the relevant information aggregation module calculates the relevance value of each text segment through the RelevantInformationAggregation algorithm, obtains the optimal segment combination, and divides them into highly relevant segments and weakly relevant segments according to the calculated values. The highly relevant segments are used as the main answers, and the weakly relevant segments are used as auxiliary answers. The optimization algorithm module adjusts the relevance and length of the generated segments based on the association value to optimize the accuracy of the answers. During the generation stage, the RelevantInformationAggregation algorithm is used for text segment aggregation and selection. Combining with large model reasoning, the optimal segments are classified and summarized to improve the retrieval generation accuracy. Especially in dense data scenarios, through multi-channel aggregation and parameter configuration, the data quality and relevance in different application scenarios are optimized.
[0010] Furthermore, the relevant information aggregation module also includes a parameter configuration function. The parameters of the parameter configuration function include the maximum number of segments, the total length limit, and the minimum score threshold configured according to different application scenarios.
[0011] Furthermore, the block database module includes a user interface module. The user interface module integrates a feedback mechanism to optimize the retrieval algorithm by collecting users' feedback on the retrieval results. The user interface module is used to receive user queries, display retrieval results, and support multiple query methods such as keyword queries and natural language queries. The user interface module is responsible for receiving user input, presenting retrieval results, and generating answers, supporting multiple input and output formats, providing a user-friendly interaction interface, enhancing the user experience, supporting multiple input and output formats, and meeting the needs of different users.
[0012] Furthermore, the feedback mechanism includes a scoring system. Users can rate the accuracy and relevance of the retrieval results, and the system automatically adjusts the retrieval strategy according to the scoring results. The system dynamically adjusts the strategies of each module by collecting users' ratings of the accuracy and relevance of the retrieval results: optimizing the query weights and metadata of the block database, adjusting the chunking granularity and weight distribution of context construction, tuning the segment screening threshold and combination logic of information aggregation, and continuously iterating algorithm parameters based on reinforcement learning to form a data-driven closed-loop optimization mechanism.
[0013] Furthermore, the block database module also includes a security module, which is used to protect the security of the document's metadata information and user query information, including functions such as data encryption, access control, and audit logs. The security module is used to protect the security of the document's metadata information and user query information, including functions such as data encryption, access control, and audit logs. The security module is responsible for functions such as data encryption, access control, and privacy protection. It uses advanced encryption technologies to protect the security of stored and transmitted data, implements strict access control policies to ensure that only authorized users can access sensitive data, safeguards the security and privacy of data, prevents data leakage and abuse, enhances the reliability and credibility of the system, and meets the requirements of laws, regulations, and industry standards.
[0014] Furthermore, the security module supports multi-level permission management, assigns different access permissions according to user roles, and ensures the security of data.
[0015] In summary, compared with the prior art, the present invention provides a RAG retrieval and generation system for intensive data scenarios, which has the following beneficial effects: Through the collaborative work of multiple modules, significant technical and beneficial effects are achieved. The block database module automatically generates the title, abstract, and paragraph semantic abstract of the document, stores these metadata information together with the document in the database, supports multi-way queries for retrieval, improves the retrieval efficiency and accuracy, and at the same time provides a structured knowledge input unit, enhancing the system's knowledge processing ability and adaptability. The context construction module performs structured context construction on the document and its semantic segments before vectorization or question answering generation, significantly improving the accuracy of query results, enabling the generation model to have a complete structural awareness, reducing the hallucination of deviating from the topic, and enhancing the applicability of the system in the private domain operation and maintenance scenario. The RelevantInformationAggregation module extracts relevant text fragments from multiple query results during the post-retrieval generation stage, combines them into a meta-document, provides accurate and comprehensive answers for users, and adjusts the length and relevance of the fragments based on the correlation value by calculating the correlation value of the text fragments, further optimizing the accuracy of the answers, and supporting parameter configuration to adapt to different application scenarios. The optimization algorithm module adjusts the relevance and length of the generated fragments based on the correlation value, classifies and summarizes the optimal fragments in combination with large model reasoning, and improves the intelligence and accuracy of the system. Description of the Drawings
[0016] Figure 1 is the architecture diagram of the RAG retrieval and generation system of the present invention for intensive data scenarios.
[0017] Figure 2 is the flowchart of the RAG retrieval and generation system of the present invention for intensive data scenarios. Detailed Embodiments
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0019] Please refer to Figure 1 、 Figure 2 , the present invention provides a technical solution, a RAG retrieval and generation system for intensive data scenarios, including: a block database module, a context construction module, a RIA module, and an optimization algorithm module. The system stores the metadata information of documents through the block database module, provides multi-way queries for retrieval, constructs a structured context for the document and its semantic segments before vectorization or question-and-answer generation through the context construction module to enhance the accuracy of query results, extracts relevant text fragments from multiple query results during the retrieval and generation stage through the RIA module, combines these fragments into a meta-document to provide accurate answers, and finally adjusts the relevance and length of the generated fragments based on the association value through the optimization algorithm module to further optimize the accuracy of the answers.
[0020] The goal of the block database module is to save the document units after semantic chunking, block-level metadata, structural hierarchies, and their reference relationships, and provide a knowledge input unit with complete context and controllable structure for subsequent retrieval enhancement and generation.
[0021] In terms of document block division, SegmentDB uses semantic large block recognition, such as chapters and logical segments, and performs sub-block refinement, such as code blocks, list items, warning messages, etc. At the same time, each block is bound with its structural features such as the chapter, paragraph position, and type it belongs to. SegmentDB designs each document block as a combination of a block header and a block body. The block header contains information such as the document title, document abstract, paragraph abstract, and block type label, and the block body contains the actual content, such as command lines or descriptive text; In terms of data storage form, SegmentDB supports data structure formats such as JSON or vectorized databases. The storage backend supports JSON / MongoDB / SQLite for text parts, and connects to vector libraries such as FAISS / Weaviate / Elasticsearch for embedded vector parts, and supports the "structural retrieval + vector retrieval" joint index mode.
[0022] The goal of the RIA module is to extract the most relevant text fragments from multiple queries and combine these fragments into a meta-document for final generation. The core functions of the RIA module include obtaining the meta-document, calculating the relevance values of text fragments, optimizing the association values, and providing parameter configurations, etc. Among them, obtaining the meta-document is to generate a rerank list set through the rerank sorting of the retrieved content. Calculating the relevance values of text fragments is based on the cosine similarity of vector inner products. Optimizing the association values is to adjust the relevance values according to the fragment lengths. Parameter configuration provides different preset parameter combinations to control the fragment aggregation behavior in different retrieval scenarios.
[0023] In terms of algorithm principle, the fragment relevance calculation is based on the cosine similarity of vector inner products, and the paragraph combination score calculation is that the overall score of the candidate combination is the sum of the relevance scores of all fragments in the combination; In terms of parameter configuration and strategy, the RIA module controls the fragment aggregation behavior in different retrieval scenarios through a parameter dictionary and selects different parameter combinations according to the scenarios.
[0024] In the scenario of operation and maintenance log analysis, the RAG retrieval and generation system proposed by the present invention also demonstrates powerful application capabilities. The system uses the SegmentDB technology to perform semantic chunking and structural feature binding on the operation and maintenance logs, and stores the processed log chunks and their metadata into the SegmentDB database. At the same time, configure the parameters of the RIA algorithm to adapt to the characteristics and analysis requirements of the operation and maintenance logs. When the operation and maintenance personnel input a query problem related to the logs, the system retrieves the log chunks related to the query from the SegmentDB and uses the RIA algorithm to aggregate and optimize the retrieved log chunks to generate a meta-document containing key information. Finally, input the meta-document into the large model for reasoning and generation, and finally return accurate solutions or suggestions to the operation and maintenance personnel.
[0025] The user interface module is integrated with a feedback mechanism, and users can rate the accuracy and relevance of the retrieval results through this interface. Specifically, in the block database module, there is a user interface module, and the user interface module provides a simple and clear scoring interface, and users can rate the retrieval results according to their own needs. After the system collects these scoring data, through the built-in feedback analysis algorithm, it automatically adjusts the retrieval strategy to optimize the accuracy and relevance of subsequent retrieval results.
[0026] Through the user feedback mechanism, the system can continuously learn and optimize the retrieval algorithm, improve the accuracy of the retrieval results and user satisfaction. The scoring system provides a basis for the quantitative evaluation of the retrieval results, which helps the system to continuously improve and optimize.
[0027] Security Module and Multi-Level Permission Management The system also includes a security module, which adopts advanced data encryption technology to ensure the security of the document's metadata information and user query information during transmission and storage. At the same time, the security module also implements strict access control policies to prevent unauthorized access. In addition, the security module records detailed audit logs for tracking and reviewing all access and operation behaviors. The security module also supports multi-level permission management. The system assigns different access permissions according to user roles. For example, ordinary users can only access and retrieve public documents, while administrator users have the permission to access and retrieve all documents.
[0028] This multi-level permission management mechanism ensures data security and prevents data leakage and abuse.
[0029] Data encryption technology protects the security of user data and document metadata, preventing data leakage and illegal access. The access control policy ensures that only authorized users can access sensitive information, improving the security of the system. Audit logs record all access and operation behaviors, providing strong support for system security audits and troubleshooting. The multi-level permission management mechanism flexibly meets the needs of different users, ensuring data security and compliance.
[0030] The present invention discloses a RAG retrieval and generation system for intensive data scenarios, including the following steps; Overall Technical Solution We implemented a class named "KnowledgeBase" to manage and operate a knowledge base. The core functions of the knowledge base include adding, deleting, querying documents, and storing and retrieving related metadata and vector data. It can be explained in two steps: offline and online. The offline part mainly includes document processing and storing it in vectorDB and segmentDB; The document addition function defines the Enhance_context function, whose main function is to generate the title, abstract, and chapter abstract of the document and add this information to each document chunk. The main logic is as follows: Document Title Generation: If the document title is not provided and the configuration allows title generation, call the function to generate the title. Otherwise, use the document ID as the title.
[0031] Document Abstract Generation: If the configuration allows generating a document abstract, call the function to generate the abstract.
[0032] Chapter Abstract Generation: If the configuration allows generating chapter abstracts, traverse each chapter and call the function to generate the abstract.
[0033] Custom term mapping: If custom term mapping is allowed in the configuration, call the `annotate_chunks` function to annotate document chunks. Prepare embedding chunks: Generate a header for each document chunk and combine the header with the chunk content to generate the final embedding chunk. Add the processed document chunks and their metadata to the SegmentDB (Segment Database). Add the generated embedding vectors and their metadata to the VectorDB (Vector Database). Enhance_context function Overview Used to construct structured context for documents and their semantic segments before vectorization or Q&A generation, including: 1. Automatically generate the document title; 2. Automatically generate the document abstract; 3. Automatically generate the semantic abstract of paragraphs; 4. Construct the structure prompt header (chunkheader) and embed it in each shard text.
[0034] Detailed description of the technical implementation principle Document title generation Construct the title generation prompt (DOCUMENT_TITLE_PROMPT) as follows: `What is the title of the following document? Your response MUST be the title of the document, and nothing else. DOCUMENT {document_text}` Call the LLM to output the result as the unique title of the document Document abstract generation Use `DOCUMENT_SUMMARIZATION_PROMPT` to construct a semantic summary request: `What is this document about? Respond in the format: "This document is about: X" DOCUMENT Document name: {generated_title} {document_text}` The returned result is used to record the semantic gist of the entire document and to generate stage prompts for guidance Paragraph summary generation (structural fragment level) Execute for each semantic segment in the document (e.g., segmented by the SemanticSectioning module): `What is this section about? Respond in the format: "This section is about: X" Document name: Section name: Section text: The result is used as the structural semantic tag for this segment (such as "Abnormality Analysis Segment", "Recovery Suggestion Segment").
[0035] Chunk Header Construction and Fusion Concatenate the document title, document summary, paragraph title, and summary to form a context hint header: `[Document Title]: MySQL Tuning Practice Guide [Document Summary]: This document is about: [Section Title]: Connection Pool Parameter Configuration Suggestions [Section Summary]: This section is about: [Content]: Original content sharding` The overall constructed content is input into the embedding model or used as a prompt for the generation model to process Innovation Points Significantly enhanced context ability: Not only the content itself is passed in, but also the context semantics, greatly improving the representation quality of the embedding model; Improved generation consistency: Enables the generation model to have a complete structural awareness and reduces off-topic hallucinations; High adaptability to the private domain operation and maintenance scenario: - Automatically establish hierarchical indexes and contexts for "untitled work orders", "multi-section configuration instructions", and "mixed-structure knowledge"; - The output supports structuring, such as Markdown / JSON format, facilitating docking with UI / interfaces; 4. Strong controllability: The prompt structure is fixed and traceable, and it adapts to multiple languages (supports automatic switching between Chinese and English).
[0036] annotate_chunks function Function: Automatically assign semantic tags to each fragment (chunk) in the document to generate more context-aware prompts (prompts) to improve accuracy during subsequent embedding or Q&A calls.
[0037] Detailed Explanation: Input Data Structure - `chunks` A list of segmented document content - `custom_term_mapping` keyword → tag mapping provided manually or by rules `{"Alarm processing":["Alarm","alarm","Alarm"],"Restart operation":["restart","Restart service"]}` Core steps - Call `find_all_term_variations` -Batch splice multiple chunks; - Perform "Fuzzy Keyword Spotting" using LLM (calling `get_response()); -Capture semantically close words (different spellings but similar meanings) to improve recall.
[0038] - Execute `annotate_chunk` -Insert tags one by one in the chunk text, such as: high CPU usage (performance bottleneck) appears` -Use regular reverse order to insert and ensure that the index is correct.
[0039] - Returns a list of annotated chunks that can be used for embedding or context display.
[0040] SegmentDB database Objective: This paper proposes a structured segmented information storage system (SegmentDB) to store semantically segmented document units, segment-level metadata, structural hierarchies, and their reference relationships, aiming to provide context-complete and structure-controllable knowledge input units for subsequent retrieval enhancement and generation.
[0041] Technical implementation details and structural design: Document block division method This system does not use fixed token length sharding, but is built using semantic chunking + hierarchical segmentation: Block strategy: 1. Semantic chunk recognition (e.g. chapters, logical paragraphs) →Use the LLM model to perform "semantic paragraph recognition" on the document to obtain a semantically complete "large block" (mother block) 2. Sub-block refinement (such as code blocks, list items, warnings, etc.) →Sub-dividing (sub-blocks) within the parent block according to sentence structure, breakpoints, and field types 3. Structural feature binding →Each block is bound to its chapter, paragraph position, and type (such as table block, command block, description block) Example structure: Combined form of chunk header and chunk content Each chunk consists of two parts: ChunkHeader (chunk header) - Generated by the Enhance_Context module; - Contains document title, document abstract, paragraph abstract, chunk type label, etc.; - Format example: `[Document Title]: Kubernetes Operations Manual [Paragraph Abstract]: This paragraph mainly describes the container restart operation command [Chunk Type]: Command Operation Chunk` ChunkBody (chunk body) Actual content, such as command line or descriptive text: `kubectl rollout restart deployment my-app` Combined form: The chunk header + chunk body are concatenated as the final embedded content or prompt text SegmentDB data storage form Data structure format (JSON or vector database), each record contains: `{"chunk_id":"B3.1","doc_id":"DOC_123","header":{"document_title":"Application Exception Manual","section_title":"Service Restart Process","chunk_type":"Operation Command","chunk_summary":"This paragraph describes restarting the service via systemctl"},"content":"systemctl restart app.service","position":"Chapter 2 - Troubleshooting Process - Paragraph 2","vector":[0.123,0.004,...,0.991], optional, embedding vector "refs":["B2.4","B3"], cross-chunk semantic reference "is_anchor":false}` Supported storage backends: - The text part can be stored in JSON / MongoDB / SQLite - The embedding vector part can be connected to vector databases such as FAISS / Weaviate / Elasticsearch - Supports the combined index mode of "structured retrieval + vector retrieval" Usage strategies during retrieval / generation 1. Structured recall: Can be filtered by chunk type, e.g., only recall chunks of the "Alarm Handling" type; 2. Multi-block aggregation generation: Aggregate semantically related blocks through the `refs` field to form a combined context input; 3. Structured answer output: Combine block header information to guide the LLM to output structured answers such as Markdown and JSON; 4. Controllable context length compression: Only the block header summary can be retained for lightweight prompting.
[0042] In the online part, it mainly includes the user input query, retrieves from the knowledge base, and finally returns to the user through the generation optimization algorithm combined with the large model inference.
[0043] The main core function is to use the RelevantInformationAggregation (RIA) algorithm in the generation stage to extract the most relevant text segments (segments) from multiple queries. In fact, it aggregates the relevant knowledge scattered in multiple places, combines them into a meta-document, and ensures that its segments do not overlap and meet the length limit through optimization, and provides it to the large model for inference.
[0044] The core functions of RIA are as follows: Obtain the meta-document The retrieved content will first be sorted by rerank to generate a rerank list set all_ranked_results, whose content includes metadata, summary, and specific segment text. Extract the IDs from the top-k documents of each query, calculate the ranking results, and generate the split points and start points of the meta-document.
[0045] Obtain the calculated value of the block Function: Calculate the relevance value of each text segment Implementation: Use the exponential decay function to calculate the relevance value of the segment, and the length and relevance of the segment can be controlled by adjusting `irrelevant_chunk_penalty` and `decay_rate` Obtain the association value of the block Function: Calculate the relevance value of each segment in the meta-document Implementation: Traverse the ranking results of each query and calculate the relevance value of each segment.
[0046] If `chunk_length_adjustment` is `True`, adjust the relevance value according to the length of the segment Optimize the association value Function: Adjust the relevance value according to the length of the segment Implementation: Multiply the relevance value by the ratio of the segment length to the reference length to adjust the relevance value Parameter dictionary Function: Provide different preset parameter configurations to control the behavior of segment selection Presets: - `balanced` balance mode, applicable to general situations.
[0047] - `precision` precision mode, applicable to scenarios that require high relevance.
[0048] - `find_all` find all mode, applicable to scenarios that require as many segments as possible.
[0049] Obtain the optimal segment Function: Find the optimal text segment from the relevance values of multiple queries Implementation: - By looping through the relevance values of each query, find the optimal segment that does not overlap and meets the length limit.
[0050] - Use the greedy algorithm to select the segment with the highest relevance value in the current query each time.
[0051] - Ensure that the segment does not cross the document boundary and the total length does not exceed `overall_max_length` Classify the optimal segment, and classify it into highly relevant segments and weakly relevant segments according to the calculated value - Highly relevant segments are used as the main answer - Weakly relevant segments are used as auxiliary answers Principle of RIA algorithm: Goal: Among multiple text segments (chunks) obtained from the preliminary retrieval, select several adjacent segment combinations (segments) based on the semantic relevance of the segments to the query, so that their content is semantically coherent, most matching the query and meeting the length limit constraint, for final generation Core calculation formula: Formula 1 Segment relevance calculation, the segment relevance is calculated by the following method, based on the vector inner product: R(Ci,Q)=cosine_similarity(vCi,vQ) where vCi and vQ are the embedding vectors of the chunk and the query respectively Formula 2 Paragraph combination score calculation For any candidate paragraph combination S(i,j)={Ci,Ci + 1,Cj}, the score is: Score(S(i,j))=∑k = i to jR(Ck,Q) That is, the sum of the relevance scores of all segments in this combination Parameter configuration and strategy description: Parameter dictionary (RSE_PARAMS_PRESETS) The RIA uses a parameter dictionary to control the segment aggregation behavior in different retrieval scenarios. Typical parameter examples are as follows: Parameter selection method Optimal segment combination process: Classification strategy: The RIA combination type is divided into three categories Final selection algorithm logic: foreach query in all_queries: while total_length < overall_max: foreach (i, j) in candidate_segments: if segment not overlapping others and not crossing documents: computeScore(S_i_j) select segment with max score and add to results Summary: The core function of this solution is to extract the most relevant text segments from multiple query results and combine these segments into a meta-document. By adjusting the parameters, the length, relevance, and total length limit of the segments can be controlled to adapt to different application scenarios. Ensure that the selected segments are both relevant and meet the length limit.
[0052] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0053] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A RAG retrieval and generation system for dense data scenarios, characterized in that It includes: a block database module, a context construction module, a relevant information aggregation module, and an optimization algorithm module; The block database module is used to automatically generate the title, abstract, and paragraph semantic abstract of a document, store the metadata information of the document and its semantic segments in a database, and provide multiple-way queries for retrieval; The context construction module constructs a structured context for the document and its semantic segments before vectorization or question answering generation to enhance the accuracy of query results; The relevant information aggregation module extracts relevant text fragments from multiple query results in the post-retrieval generation stage, combines these fragments into a meta-document, and provides an accurate answer; The optimization algorithm module adjusts the relevance and length of the generated fragments based on the association value to further optimize the accuracy of the answer.
2. The RAG retrieval and generation system for dense data scenarios according to claim 1, wherein: The block database module generates the title, abstract, and chapter abstract of the document through the Enhance_context function and adds them to each document block to achieve the storage and multiple-way query of the metadata information of the document.
3. The RAG retrieval and generation system for dense data scenarios according to claim 1, wherein: The context construction module uses the method of semantic chunking plus hierarchical segmentation to divide the document into chunks, and combines the chunk headers and chunk contents to construct structured context information.
4. The RAG retrieval and generation system for dense data scenarios according to claim 1, wherein: The relevant information aggregation module calculates the relevance value of each text fragment through the RelevantInformationAggregation algorithm, obtains the optimal fragment combination, and divides them into highly relevant fragments and weakly relevant fragments according to the calculated value. The highly relevant fragments are used as the main answer, and the weakly relevant fragments are used as the auxiliary answer.
5. The RAG retrieval and generation system for dense data scenarios according to claim 1, characterized in that: The relevant information aggregation module also includes a parameter configuration function, and the parameters of the parameter configuration function include parameters such as the maximum number of fragments, the total length limit, and the minimum score threshold configured according to different application scenarios.
6. The RAG retrieval and generation system for dense data scenarios according to claim 1, characterized in that: The block database module includes a user interface module, and the user interface module is used to receive user queries and display retrieval results.
7. The RAG retrieval and generation system for intensive data scenarios according to claim 6, wherein: The user interface module integrates a feedback mechanism to optimize the retrieval algorithm by collecting user feedback on retrieval results.
8. The RAG retrieval and generation system for dense data scenarios according to claim 7, characterized in that: The feedback mechanism includes a scoring system. Users can score according to the accuracy and relevance of the retrieval results, and the system adjusts the retrieval strategy according to the scoring results.
9. The RAG retrieval and generation system for dense data scenarios according to claim 1, characterized in that: The block database module also includes a security module, which is used to protect the security of the metadata information of the document and user query information.
10. The RAG retrieval and generation system for dense data scenarios according to claim 9, characterized in that: The security module supports multi-level permission management, assigns different access permissions according to user roles, and ensures the security of data.
Citation Information
Patent Citations
RAG system optimization method and system, electronic equipment and storage medium
CN118917305A
Enhanced question and answer method based on large language model and electronic equipment
CN119202216A
Retrieval method and system based on semantic partitioning, electronic equipment and storage medium
CN119782460A
Multi-modal document retrieval enhancement generation method based on large model
CN119988588A
Wounded information management and sharing method and system based on cloud computing
CN120032832A
Cited By
Hardware isolation security processing system and method for generating RAG (Retrieval Enhanced Generation) in oriented way
CN121959590A