Retrieval enhancement generation method and device, electronic equipment and storage medium

By receiving multi-source heterogeneous data in real time through a streaming processing engine and performing vectorized encoding, combined with hierarchical storage and hybrid retrieval mechanisms, the problem of knowledge update delay in existing technologies is solved, achieving second-level dynamic updates of the knowledge base and improving the timeliness and accuracy of generated content.

CN120950651APending Publication Date: 2025-11-14JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511071172.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing search-enhanced generation methods have failed to effectively achieve real-time collaborative processing of multi-source heterogeneous data, resulting in delays in knowledge updates and affecting the timeliness and accuracy of generated content. In particular, the requirements for timely reflection and accuracy of real-time data in the financial and judicial fields have not been met.

Method used

The streaming engine receives multi-source heterogeneous data in real time, performs vectorization encoding, generates dynamic knowledge vectors, triggers incremental index updates based on a predetermined time window, utilizes a hierarchical dynamic knowledge base for hybrid retrieval, and optimizes the output by combining streaming context management and a dual verification mechanism.

Benefits of technology

It achieves second-level dynamic updates of the knowledge base, significantly improving the timeliness and accuracy of generated content and solving the generation deviation problem caused by knowledge lag in traditional RAG systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950651A_ABST
    Figure CN120950651A_ABST
Patent Text Reader

Abstract

The invention provides a retrieval enhancement generation method and device, electronic equipment and a storage medium, and relates to the technical field of data processing.The method comprises the steps that multi-source heterogeneous data are received in real time through a streaming processing engine, vectorization coding is conducted on the multi-source heterogeneous data, and dynamic knowledge vectors are generated; incremental index updating is triggered based on a preset time window, the dynamic knowledge vector is utilized to update a layered storage dynamic knowledge base, and the dynamic knowledge base comprises a real-time data storage layer and a historical knowledge storage layer; selecting a mixed retrieval strategy according to the query type, and jointly retrieving real-time data in the dynamic knowledge base and a historical knowledge base to obtain related data; and related data obtained by retrieval is combined with streaming context management to generate an initial result, and the initial result is optimized and output through a dual-check mechanism. Compared with the prior art, the second-level dynamic updating of the knowledge base can be realized, the timeliness and accuracy of the generated content are remarkably improved, and the problem of generation deviation caused by knowledge lag is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a method and apparatus for retrieval enhancement generation, an electronic device, and a storage medium. Background Technology

[0002] With the deep integration of artificial intelligence and big data processing technologies, Retrieval-Augmented Generation (RAG) has become a key technology for improving the accuracy and relevance of generative models, and is widely used in professional fields such as finance, healthcare, and law, where the timeliness of knowledge is extremely important. RAG extracts relevant information from external knowledge bases through a retrieval module and works collaboratively with the generative model to overcome the shortcomings of traditional Large Language Models (LLMs), which rely on closed training data and struggle to update knowledge in real time. Specifically, this technology system covers the entire process from data acquisition, vectorization processing, dynamic index construction to hybrid retrieval and generation feedback. The collaborative mechanism between the streaming engine, vector database, and generative model is particularly crucial, providing a technological foundation for the real-time fusion of multi-source heterogeneous data and dynamic knowledge injection.

[0003] However, existing RAG methods directly utilize static knowledge bases for information retrieval without establishing an efficient collaborative mechanism with real-time data streams. This can lead to significant delays in knowledge updates, affecting the timeliness and accuracy of generated content. For example, in the financial sector, if second-level fluctuations in market conditions are not reflected in the knowledge base in a timely manner, it will directly impact the reliability of risk assessments and decision-making recommendations. In judicial scenarios, if the continuous input of case evidence cannot be updated in real time, legal advice may become inconsistent with the latest facts. Furthermore, traditional RAG lacks a unified vector encoding system when processing structured, unstructured, and time-series data, resulting in low efficiency in multi-source data fusion, which in turn affects the relevance of retrieval results and the logical consistency of generated content. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, and storage medium for retrieval enhancement generation. It aims to at least partially address the technical problems in related technologies.

[0005] According to a first aspect of this disclosure, a method for retrieval enhancement generation is provided, comprising:

[0006] The system receives multi-source heterogeneous data in real time through a streaming processing engine, and performs vectorization encoding on the multi-source heterogeneous data to generate dynamic knowledge vectors.

[0007] Incremental index updates are triggered based on a predetermined time window, and the dynamic knowledge base stored in a hierarchical manner is updated using the dynamic knowledge vector. The dynamic knowledge base includes a real-time data storage layer and a historical knowledge storage layer.

[0008] Select a hybrid retrieval strategy based on the query type, and jointly retrieve real-time data and historical knowledge from the dynamic knowledge base to obtain relevant data;

[0009] The retrieved relevant data is combined with streaming context management to generate initial results, and the output is optimized through a dual verification mechanism.

[0010] According to a second aspect of this disclosure, an apparatus for retrieval enhancement generation is provided, comprising:

[0011] The generation unit is used to receive multi-source heterogeneous data in real time through a streaming processing engine, and to vectorize and encode the multi-source heterogeneous data to generate dynamic knowledge vectors.

[0012] An update unit is used to trigger incremental index updates based on a predetermined time window and update the hierarchically stored dynamic knowledge base using the dynamic knowledge vector. The dynamic knowledge base includes a real-time data storage layer and a historical knowledge storage layer.

[0013] The retrieval unit is used to select a hybrid retrieval strategy based on the query type, and jointly retrieve real-time data and historical knowledge from the dynamic knowledge base to obtain relevant data.

[0014] The optimization unit is used to generate initial results from the retrieved relevant data by combining streaming context management, and optimize the output through a dual verification mechanism.

[0015] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the retrieval enhancement generation method described in the first aspect above.

[0019] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the retrieval enhancement generation method described in the first aspect above.

[0020] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the retrieval enhancement generation method as described in the first aspect above.

[0021] This disclosure provides a method, apparatus, electronic device, and storage medium for enhanced retrieval generation, relating to the field of data processing technology. It involves receiving multi-source heterogeneous data in real time via a streaming processing engine, vectorizing and encoding the data to generate dynamic knowledge vectors; triggering incremental index updates based on a predetermined time window, and updating a hierarchically stored dynamic knowledge base using the dynamic knowledge vectors. The dynamic knowledge base includes a real-time data storage layer and a historical knowledge storage layer; selecting a hybrid retrieval strategy based on the query type, and jointly retrieving real-time data and historical knowledge from the dynamic knowledge base to obtain relevant data; combining the retrieved relevant data with streaming context management to generate initial results, and optimizing the output through a dual verification mechanism. Compared with related technologies, this disclosure can achieve second-level dynamic updates of the knowledge base, significantly improving the timeliness and accuracy of generated content, and effectively solving the generation deviation problem caused by knowledge lag in traditional RAG systems.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0023] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0024] Figure 1 A flowchart illustrating a method for generating enhanced search results according to an embodiment of this disclosure;

[0025] Figure 2 This is a schematic diagram of a retrieval enhancement generation device provided in an embodiment of the present disclosure. Detailed Implementation

[0026] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0027] The following description, with reference to the accompanying drawings, outlines a method, apparatus, electronic device, and storage medium for generating search enhancements according to embodiments of this disclosure.

[0028] Figure 1 This is a schematic flowchart illustrating a method for generating enhanced search results according to an embodiment of this disclosure.

[0029] like Figure 1 As shown, the method includes the following steps:

[0030] Step 101: Receive multi-source heterogeneous data in real time through a streaming processing engine, and vectorize the multi-source heterogeneous data to generate dynamic knowledge vectors.

[0031] In this embodiment, the streaming processing engine is a computing framework for real-time processing of continuous data streams, capable of efficiently receiving and processing multi-source heterogeneous data. This multi-source heterogeneous data encompasses various types, including structured database logs, unstructured video streams, and IoT sensor time-series data. These data originate from different data sources and possess different formats and characteristics. For example, structured data has a clear table structure and data type, unstructured data lacks a fixed format and is complex in content, while time-series data is continuously generated in chronological order and carries time tags.

[0032] The streaming engine aggregates the aforementioned multi-source data streams using the Apache Flink real-time computing framework. Specifically, it employs a 30-second sliding window mechanism, which defines a continuously moving time interval within a continuous data stream. This mechanism then aggregates and processes the multi-source data within that interval, thereby achieving real-time data aggregation and ensuring that the latest data information can be captured and processed in a timely manner.

[0033] After receiving and aggregating heterogeneous data from multiple sources, the streaming engine also uses a complex event processing rule engine to define event triggering logic. For example, when a specific event is detected, such as an industrial equipment temperature exceeding 100°C for 10 seconds, knowledge retrieval will be triggered to achieve timely response to key events in the data.

[0034] Subsequently, for the aggregated multi-source heterogeneous data, a hybrid model of Bidirectional Encoder-Representation Transformer (BERT), Long Short-Term Memory (LSTM) network, and Residual Network (ResNet) is employed for vectorization encoding. The BERT is suitable for processing text data, capturing contextual semantic information; LSTM excels at processing time-series data, effectively extracting temporal features; and ResNet is adept at processing image data, extracting key visual features. This hybrid model transforms different types of multi-source heterogeneous data, such as text, images, and time-series data, into a unified vector representation, generating dynamic knowledge vectors. These vectors accurately reflect the features and semantic information of the original data, laying the foundation for subsequent processing and applications.

[0035] Step 102: Trigger incremental index update based on a predetermined time window, and update the hierarchically stored dynamic knowledge base using the dynamic knowledge vector. The dynamic knowledge base includes a real-time data storage layer and a historical knowledge storage layer.

[0036] In this embodiment, the predetermined time window is a fixed time interval used to trigger incremental index updates, specifically set to 5 seconds. This time window allows for periodic initiation of updates to the dynamic knowledge base, ensuring that dynamic knowledge vectors are promptly incorporated into the storage system. Incremental index updates refer to indexing and updating only newly added or changed dynamic knowledge vectors, rather than re-indexing the entire knowledge base. This approach significantly improves update efficiency, reduces resource consumption, and meets the needs of real-time processing. The hierarchical storage structure of the dynamic knowledge base is designed to balance data real-time performance and query efficiency. The real-time data storage layer is specifically designed to store recent dynamic data, employing a graphics processor-accelerated index. This index combines a hierarchical-navigable small-world approach with an inverted file-product quantization algorithm, supporting tens of thousands of vector write operations per second. It is specifically used to store dynamic data from the past 24 hours, such as real-time case progress and the latest judicial interpretations. To facilitate management and control of the data lifecycle, the index of the real-time data storage layer is sharded hourly. When data exceeds the 24-hour storage period, it is automatically archived to the historical knowledge storage layer, ensuring that the real-time data storage layer always maintains high query performance.

[0037] The historical knowledge storage layer stores long-accumulated static knowledge. It employs a hybrid index from an elastic search database, combining the BM25 algorithm with dense vector technology. This index simultaneously supports semantic and keyword retrieval, thereby improving the recall rate of historical static knowledge. The stored content includes historical case databases, full texts of laws and regulations, and other information that doesn't change frequently. This layered storage architecture, combined with an incremental index update mechanism based on predetermined time windows, ensures timely updates to dynamic knowledge vectors while accommodating the storage needs and query efficiency of different data types. This allows the dynamic knowledge base to quickly respond to real-time data changes while efficiently supporting the retrieval and access of historical knowledge.

[0038] Step 103: Select a hybrid retrieval strategy based on the query type, and jointly retrieve real-time data and historical knowledge from the dynamic knowledge base to obtain relevant data.

[0039] In this embodiment, the query types include explicit entity queries, open-ended question queries, and complex scenario queries. The hybrid retrieval strategy dynamically selects the matching retrieval method based on different query types and combines the real-time data storage layer and historical knowledge storage layer in the dynamic knowledge base to retrieve relevant data. Explicit entity queries refer to queries targeting specific, definite entity information, such as specific legal clause numbers; open-ended question queries refer to queries that are not limited to fixed answers and require semantic understanding, such as sentencing standards for a certain type of case; complex scenario queries refer to queries involving multiple entities or conditions and requiring comprehensive analysis, such as the handling of specific crimes in cases involving special groups. For explicit entity queries, the hybrid retrieval strategy adopts a precise mode, directly calling the keyword retrieval function of the elastic search database in the historical knowledge storage layer. This function uses the BM25 algorithm to quickly locate structured entity information, ensuring a response time of less than 50 milliseconds, thereby efficiently acquiring the corresponding static knowledge data.

[0040] For open-ended queries, the hybrid retrieval strategy adopts a semantic model. It uses the vector retrieval function of Milvis version 2.3 in the real-time data storage layer to obtain the top 10 dynamic data related to the query semantics. Then, it uses a bidirectional encoder representation converter to reorder the retrieval results and finally selects the top 3 most relevant data. At the same time, it combines the semantic retrieval results of the elastic search database in the historical knowledge storage layer to supplement static knowledge as support and improve the comprehensiveness of the retrieval results.

[0041] For complex query scenarios, the hybrid retrieval strategy adopts a hybrid mode. First, static knowledge related to the core keywords in the query is extracted from the historical knowledge storage layer through keyword retrieval. At the same time, dynamic data that matches the query semantics is obtained from the real-time data storage layer through semantic retrieval. Then, the two types of retrieval results are cross-validated to eliminate information ambiguity and finally integrated to form comprehensive relevant data.

[0042] By employing the hybrid retrieval strategy selected based on query type, it is possible to jointly retrieve dynamic data from the past 24 hours in the real-time data storage layer and static knowledge from the historical knowledge storage layer. This ensures timely access to the latest information while fully utilizing historically accumulated knowledge, thereby improving the relevance and accuracy of retrieval results and meeting query needs in different scenarios.

[0043] Step 104: Generate initial results from the retrieved relevant data using streaming context management, and optimize the output through a dual verification mechanism.

[0044] In this embodiment, the streaming context management refers to the dynamic maintenance and semantic integration of the dialogue history generated during a user session. Specifically, it maintains the semantic vectors of the most recent 10 rounds of dialogue and uses a multi-head attention mechanism to dynamically weight the contextual information of different rounds, giving higher weight to important dialogue content and lower weight to secondary information. This allows for accurate capture of the core logic and contextual relationships of the dialogue when generating the initial result. The semantic vector is a 768-dimensional vector transformed from the dialogue text by a sentence-level bidirectional encoder-representation converter, which quantifies the semantic information of the text. The multi-head attention mechanism uses multiple parallel attention heads to focus on different dimensions of information in the dialogue, and then integrates the results of each dimension to capture multi-level contextual relationships.

[0045] When generating the initial result, the system concatenates the retrieved relevant data (including dynamic data from the real-time data storage layer and static knowledge from the historical knowledge storage layer), the original query content, and the semantic vector of the dialogue history processed by streaming context management to form structured prompts. These prompts are then input into the large language model to generate an initial answer. During this process, the prompts are combined with domain-specific instruction templates, such as "Please answer based on the judicial interpretation of the Supreme People's Court" in the legal field, to constrain the format and basis of the generated content and ensure that the initial result conforms to professional standards.

[0046] The dual verification mechanism optimizes the initial results and includes two parts: rule engine verification and lightweight bidirectional encoder-representation converter-classifier filtering. The rule engine uses preset verification rules such as numerical ranges and logical relationships to perform compliance checks on key data in the initial results. For example, in legal scenarios, it verifies whether the calculation of sentence duration conforms to the legal scope and whether the citation of clauses is accurate. The lightweight bidirectional encoder-representation converter-classifier filters out content with a confidence level below a threshold by calculating the semantic similarity between the initial results and the retrieved data, avoiding the output of "illusory" information. After dual verification, if the results are abnormal (such as rule verification failing or insufficient confidence), the system will trigger a secondary search or manual review, ultimately outputting optimized and accurate results. Streaming context management ensures consistency between the generated content and the dialogue context, and the dual verification mechanism enhances the reliability of the results, ensuring that the output content is both user-friendly and meets the accuracy requirements of the professional field.

[0047] This disclosure provides a method for enhanced knowledge generation (RAG) through a streaming engine. It receives multi-source heterogeneous data in real time and performs vectorization encoding on the data to generate dynamic knowledge vectors. Incremental index updates are triggered based on a predetermined time window, and the dynamic knowledge vectors are used to update a hierarchically stored dynamic knowledge base, which includes a real-time data storage layer and a historical knowledge storage layer. A hybrid retrieval strategy is selected based on the query type to jointly retrieve real-time data from the dynamic knowledge base and the historical knowledge base, obtaining relevant data. The retrieved relevant data is then combined with streaming context management to generate initial results, and the output is optimized through a dual-validation mechanism. Compared with related technologies, this disclosure enables second-level dynamic updates of the knowledge base, significantly improving the timeliness and accuracy of generated content and effectively solving the generation deviation problem caused by knowledge lag in traditional RAG systems.

[0048] Furthermore, in the embodiments of this disclosure, the specific implementation of the operation of "receiving multi-source heterogeneous data in real time through a streaming processing engine and vectorizing the multi-source heterogeneous data" is diverse. For clarity, the following lists, but is not limited to, some implementation methods: using a distributed streaming processing framework to aggregate structured logs, unstructured text, and time-series sensor data in a sliding window manner; and using a complex event processing rule engine to define event triggering logic to trigger a knowledge retrieval request when a preset event pattern is detected.

[0049] Specifically, this framework possesses the ability to efficiently process continuous data streams, simultaneously handling multi-source heterogeneous data from different sources, including but not limited to structured database change logs (such as financial transaction records and case information entry logs), unstructured text files and video streams (such as on-site law enforcement recording videos and user consultation texts), and time-series sensor data (such as industrial equipment operating parameters and real-time data from medical monitors). To achieve real-time aggregation and preliminary processing of this dynamic data, this distributed stream processing framework employs a sliding window mechanism with a window duration set at 30 seconds. Through this continuously shifting time interval division, multi-source data within a unit of time can be centrally aggregated, ensuring real-time data processing while avoiding the inefficiency caused by overly fragmented data, thus laying the foundation for subsequent unified processing.

[0050] Simultaneously, a complex event processing rule engine is integrated into the streaming process to achieve accurate identification and response to preset event patterns. This rule engine can predefine various event triggering logics according to actual business needs. For example, in industrial monitoring scenarios, when the equipment temperature is detected to exceed 100°C for 10 consecutive seconds, a knowledge retrieval request is automatically triggered to promptly retrieve relevant equipment maintenance knowledge. In public safety scenarios, if a 300% surge in pedestrian traffic in a certain area during the night is detected, an anomaly warning-related knowledge retrieval is triggered to provide information support for rapid decision-making. Through this event-driven mechanism, the streaming processing engine can proactively identify key information and trigger subsequent operations while receiving data, improving the overall system's response speed to dynamic events.

[0051] After receiving and aggregating the data, the multi-source heterogeneous data needs to be vectorized to generate unified dynamic knowledge vectors. This process employs a hybrid model of Bidirectional Encoder-Representation Transformer (BERT), Long Short-Term Memory (LSTM) network, and Residual Network. The BERT primarily processes textual data, deeply capturing contextual semantic information and transforming unstructured text into vectors with semantic representation capabilities. The LSTM network focuses on processing temporal sensor data, effectively extracting dynamic features from time-series data and converting them into vectors by leveraging its ability to capture time-series dependencies. The Residual Network (ResNet) handles image data (such as keyframes in video streams), extracting visual features from images through multi-layer convolutional operations and converting them into vectors. Through the collaborative work of these three models, multi-source heterogeneous data of varying formats and types can be uniformly encoded into a consistent vector form—dynamic knowledge vectors. This allows the data to be efficiently processed and utilized by the subsequent dynamic indexing engine and hybrid retrieval module, ensuring the smooth operation of the entire real-time data-driven retrieval enhancement generation architecture.

[0052] Furthermore, in this embodiment of the disclosure, the sliding window time of the distributed stream processing framework is a first duration; the event triggering logic includes maintaining short-term and long-term states based on a business entity state management model, wherein the short-term state stores real-time data within a second duration, and the long-term state achieves historical data backtracking through persistent storage.

[0053] Specifically, in this embodiment of the disclosure, regarding the specific implementation of "receiving multi-source heterogeneous data in real time through a streaming processing engine," the sliding window time (i.e., the first duration) adopted by the distributed streaming processing framework is set to 30 seconds. This duration ensures both the real-time aggregation efficiency of multi-source heterogeneous data and avoids data fragmentation caused by an excessively short window or processing delays caused by an excessively long window. Through this 30-second sliding window, the distributed streaming processing framework can centrally aggregate structured logs (such as financial transaction logs), unstructured text (such as user consultation content), and time-series sensor data (such as equipment operating temperature data) within a unit of time, achieving segmented processing of continuous data streams and providing a regularized data source for subsequent vectorized encoding.

[0054] Meanwhile, the specific implementation of the event triggering logic based on the business entity state management model to maintain short-term and long-term states is as follows: The business entity state management model is built on the state backend of the distributed stream processing framework (such as FlinkStateBackend), with business entities (such as cases, devices, users, etc.) as the core management unit, and associates multi-source data through unique identifiers (such as case number, device ID). The short-term state is used to store real-time data within a second time period, which is set to 1 hour. This is implemented through the ValueState storage structure. For example, in a public security law enforcement scenario, it can store the movement trajectory points of the suspects and the real-time operating parameters of the equipment within the past hour, supporting real-time heat map rendering, dynamic trajectory tracking, and other immediate business needs. The long-term state is implemented through a persistent storage mechanism (such as RocksDBStateBackend) to save historical data of the entire lifecycle of the business entity (such as complete information of a case from filing to closing, and the operating records of equipment from activation to scrapping). This persistent storage not only supports long-term data retention but also enables historical data backtracking and cross-entity correlation analysis. For example, in case investigation, the processing flow of past cases can be backtracked to assist in the current case analysis, and in equipment maintenance, historical fault data can be analyzed to predict potential problems. By managing short-term and long-term states in a hierarchical manner, the business entity state management model can provide comprehensive business context support for event triggering logic while ensuring fast access to real-time data. This enables the complex event processing rule engine to make accurate judgments by combining short-term real-time data with long-term historical data when detecting preset event patterns (such as abnormal trajectory of involved personnel or excessive equipment parameters), thereby triggering corresponding knowledge retrieval requests.

[0055] Furthermore, in the embodiments of this disclosure, the specific implementation of the operation of "triggering incremental index updates based on a predetermined time window" is diverse. For clarity, the following lists, but is not limited to, some implementation methods: updating the vector database every third time interval and using a time-to-live strategy to manage data storage; the hierarchical storage dynamic knowledge base includes a hot storage layer and a cold storage layer. The hot storage layer uses a graphics processor to accelerate indexing and achieve high-speed vector writing, while the cold storage layer uses a hybrid index combined with semantic retrieval and keyword retrieval.

[0056] Specifically, in this embodiment, for the operation of "triggering incremental index updates based on a predetermined time window," a third duration (i.e., 5 seconds) is used as a fixed time interval to periodically trigger incremental update operations on the vector database. The vector database used here is Milvus. The incremental update mechanism only indexes and updates newly added or changed dynamic knowledge vectors, rather than rebuilding the entire database, thereby significantly improving update efficiency, reducing computational resource consumption, and ensuring that real-time data can be quickly incorporated into the storage system. Simultaneously, to achieve lifecycle management for different types of data, a Time-to-Live (TTL) strategy is adopted, setting different retention periods based on the data's business attributes: for highly time-sensitive data such as financial market data, the retention period is set to 1 minute, after which it is automatically cleaned up to release storage space; for data requiring longer retention periods, such as case evidence, the retention period is set to 24 hours to ensure effective retrieval within the case processing cycle.

[0057] In the hierarchical dynamic knowledge base, the hot storage layer and cold storage layer are implemented as follows: The hot storage layer is based on the Milvis 2.3 version GPU-accelerated index. This index uses an algorithm combining Hierarchical-Navigable Small World (HNSW) and Inverted File-Product Quantization (IVF-PQ), which can support high-speed write operations of tens of thousands of vectors per second. It is specifically used to store dynamic data from the past 24 hours, such as real-time case progress and the latest judicial interpretations. For ease of management, the index of the hot storage layer is sharded hourly. When the data storage time exceeds 24 hours, it is automatically archived to the cold storage layer to ensure that the hot storage layer always maintains high query performance.

[0058] The cold storage layer employs a hybrid index from Elasticsearch, which combines the BM25 keyword retrieval algorithm with dense vector semantic retrieval technology. This allows for precise keyword matching of structured static knowledge (such as case numbers in historical case databases and clause numbers in full-text laws and regulations) and the mining of deep textual connections through semantic similarity calculations, thereby improving the retrieval recall rate for historical static knowledge. The cold storage layer primarily stores information that does not change frequently, such as historical case databases and full-text laws and regulations. Through collaboration with the hot storage layer, it forms a complete knowledge storage system that balances real-time performance with historical depth.

[0059] Through the aforementioned incremental index update mechanism based on a predetermined time window (5 seconds), time-to-live strategy, and hierarchical storage architecture, it is possible to ensure timely updates of dynamic knowledge vectors while achieving differentiated storage and management based on data characteristics. This not only meets the high responsiveness requirements of real-time data in scenarios such as finance and law enforcement, but also ensures the effective retention and efficient retrieval of historical knowledge.

[0060] Furthermore, in this embodiment of the disclosure, the hot storage layer index is sharded by hour, and expired data is automatically archived to the cold storage layer; the cold storage layer stores historical case databases and full texts of laws and regulations, and performs logical paragraph segmentation of the text through a semantic segmentation algorithm, including basic segmentation, related segmentation and dynamic segmentation.

[0061] Specifically, in this embodiment, regarding the implementation of the hot storage layer and the cold storage layer, the hot storage layer adopts an hourly sharding index management mechanism. A new index shard is automatically generated every hour to store the dynamic knowledge vectors added within that hour. This sharding method enables fine-grained time-dimensional management of data, facilitating rapid data location and retrieval by time period. When the storage duration corresponding to a certain index shard reaches 24 hours (i.e., exceeding the storage period of the hot storage layer), the system automatically triggers an archiving process, transferring the shard data completely to the cold storage layer through a preset migration mechanism. Simultaneously, the corresponding index pointer is retained in the hot storage layer to support cross-layer retrieval, ensuring both the timeliness of the data in the hot storage layer and achieving automated management of the data lifecycle. When storing static knowledge such as historical case databases and full texts of laws and regulations, the cold storage layer uses a semantic block segmentation algorithm to logically segment the text into paragraphs to improve retrieval accuracy and efficiency. The basic segmentation is based on the natural paragraph structure of the text. For example, each complete clause in laws and regulations is treated as an independent segment, ensuring that each segment contains complete semantic information and legal text content. Related segments target supplementary materials related to the main legal provisions, such as judicial interpretations and precedents. By extracting citation relationships in the text (e.g., "refer to Article X" or "according to the spirit of precedent X"), a three-in-one hyperlink structure of "main legal provision - judicial interpretation - related precedent" is constructed, making related legal knowledge an organic whole and facilitating cross-text related queries during retrieval. Dynamic segments are mainly used to process historical case description texts in the cold storage layer. A sliding window of 512 tokens combined with entity recognition technology (e.g., identifying key entities such as involved persons, time, place, and events) is used to ensure that each segment contains complete event elements (Who-When-Where-What), avoiding semantic breaks due to overly fine segmentation or information redundancy due to overly coarse segmentation. Through the application of the above semantic segmentation algorithm, the cold storage layer can achieve structured organization of static knowledge while preserving the original semantics of the text, providing accurate retrieval units for subsequent hybrid retrieval.

[0062] Furthermore, in the embodiments of this disclosure, there are various specific implementation methods for the operation of "selecting a hybrid retrieval strategy according to the query type". For clarity, the following lists some implementation methods, including but not limited to: using keyword retrieval mode for explicit entity queries; using semantic retrieval mode to obtain relevant data for open questions and taking the top three results after reordering; using a hybrid mode to integrate keyword retrieval and semantic retrieval results for complex scenarios and eliminating ambiguity through cross-validation.

[0063] Specifically, in this embodiment, regarding the operation of "selecting a hybrid retrieval strategy based on query type," the query type needs to be accurately determined first, and then the corresponding retrieval mode needs to be matched to obtain relevant data by combining the real-time data storage layer and the historical knowledge storage layer of the dynamic knowledge base. Specifically, entity queries refer to queries targeting specific information with unique identifiers, such as "the content of Article 263 of the Criminal Law" or "the filing time of a certain case number." Such queries require rapid and accurate result locating; therefore, a keyword retrieval mode is adopted, directly calling the keyword retrieval function of the elastic search database in the historical knowledge storage layer. This function, based on the BM25 algorithm, achieves accurate matching of structured entity information and can return the corresponding static knowledge data within 50 milliseconds, ensuring the timeliness and accuracy of the query results. For open-ended queries, such as "What are the sentencing standards for robbery?" or "How is liquidated damages calculated in contract disputes?", which lack fixed answers and require semantic understanding, a semantic retrieval model is adopted. First, the semantic retrieval function of the Milvis vector database in the real-time data storage layer is used to obtain the top 10 dynamic data (such as the latest case law and judicial interpretations) related to the query semantics based on vector similarity calculation. Then, the bidirectional encoder representation converter is used to reorder these 10 results, and the top 3 most relevant dynamic data are selected based on semantic relevance. At the same time, the semantic retrieval results of the elastic search database in the historical knowledge storage layer are combined to supplement the corresponding static knowledge such as laws and regulations and historical cases, forming a retrieval result that is both timely and comprehensive.

[0064] For complex query scenarios, such as "special handling procedures for minors committing theft" or "tariffs and legal clauses involved in cross-border trade disputes," which involve multiple entities, conditions, or overlapping fields, a hybrid approach is adopted. First, keyword retrieval extracts static knowledge (such as corresponding legal provisions and similar historical cases) related to the core keywords (e.g., "theft" and "minors") from the historical knowledge storage layer. Then, semantic retrieval retrieves dynamic data matching the query's semantics from the real-time data storage layer (e.g., recent case outcomes and the latest judicial guidance). Subsequently, the two types of search results are cross-validated, eliminating information ambiguity by comparing semantic consistency and clause applicability, ultimately integrating comprehensive and accurate relevant data. This hybrid retrieval strategy, dynamically adjusted according to query type, leverages the precision of keyword retrieval while capturing deep connections through semantic retrieval, and simultaneously considers both real-time data and historical knowledge, effectively improving the relevance and reliability of search results across different scenarios.

[0065] Furthermore, in the embodiments of this disclosure, the specific implementation of the operation of "generating initial results by combining the retrieved relevant data with streaming context management and optimizing the output through a dual verification mechanism" is diverse. For clarity, the following lists, but is not limited to, some implementation methods: maintaining the semantic vector of historical dialogues as streaming context and dynamically weighting it through a multi-head attention mechanism; using a rule engine to perform numerical range verification and a lightweight classification model to filter low-confidence generated content, and updating the prompt word template.

[0066] Specifically, in this embodiment, regarding the operation of "generating initial results by combining the retrieved relevant data with streaming context management and optimizing the output through a dual verification mechanism," the implementation first involves maintaining historical dialogue information through streaming context management to ensure consistency between the generated content and the dialogue context. This streaming context management uses Sentence-BERT to dynamically encode the text of the most recent 10 rounds of dialogue into 768-dimensional semantic vectors. These vectors accurately capture the core semantics of each round of dialogue. Simultaneously, a multi-head attention mechanism is used to dynamically weight the semantic vectors of these 10 rounds of dialogue—giving higher weights to dialogue rounds with high relevance to the current query (such as directly mentioned entities or conditions) and lower weights to rounds with low relevance (such as early small talk or minor information), thereby highlighting key contextual information and preventing irrelevant content from interfering with the generation process.

[0067] In the initial result generation stage, the system structurally concatenates the retrieved relevant data (including dynamic data from the real-time data storage layer and static knowledge from the historical knowledge storage layer), the user's original query, and the weighted streaming context semantic vector to form prompts containing complete background information. These prompts are then combined with domain-specific instruction templates (such as "Please answer based on the latest judicial interpretations and relevant precedents" in the legal field) to constrain the format and basis of the generated content. This content is then input into a large language model such as Llama-2-13B to generate a preliminary answer. Prior to this, the search results are pre-summarized using a T5 model to extract core facts and legal basis, ensuring that the information input to the large model is concise and crucial, thus improving generation efficiency and accuracy.

[0068] A dual verification mechanism is used to optimize the initial results to reduce the risk of outputting "illusory" information. The rule engine is responsible for verifying numerical ranges and logical relationships. For example, in financial scenarios, it verifies whether the rate of return on investment advice is within a reasonable range; in legal scenarios, it verifies whether the calculation of prison sentences conforms to the legal upper and lower limits and whether the citation of clauses is accurate. The lightweight BERT classifier calculates the semantic similarity between the initial results and the retrieved data, filtering out content with a confidence level below a preset threshold (e.g., 0.6) to ensure a high degree of match between the generated content and the retrieval basis. If the initial results are found to be abnormal (e.g., the rule engine fails verification, or the classifier determines the confidence level is too low), the system will automatically trigger a secondary search to retrieve relevant data from the dynamic knowledge base to supplement the generation basis, or, if the secondary search still cannot correct the issue, it will initiate a manual review process, with domain experts intervening to confirm the accuracy of the results.

[0069] Furthermore, the prompt word templates are dynamically updated during the generation process through human feedback and reinforcement learning mechanisms. The system records user feedback on the generated results and review feedback from domain experts. It then uses reinforcement learning algorithms to evaluate the effectiveness of different prompt word templates, selecting and updating the templates that perform best in specific scenarios (such as complex case consultations and real-time financial analysis). This allows the prompt words to more accurately guide the large model in generating content that meets business needs, further improving the reliability and professionalism of the output results. Through the synergy of streaming context management, multi-stage generation, and dual verification, this operation ensures the coherence of the generated content and dialogue context, while also improving the accuracy and compliance of the results through multi-layered verification and template optimization.

[0070] It should be noted that the embodiments of this disclosure may include multiple steps. For ease of description, these steps are numbered, but these numbers are not a limitation on the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of this disclosure do not limit this.

[0071] Corresponding to the search enhancement generation method described above, this disclosure also proposes a search enhancement generation apparatus. Since the apparatus embodiments of this disclosure correspond to the method embodiments described above, details not disclosed in the apparatus embodiments can be referred to the method embodiments described above, and will not be repeated here.

[0072] Figure 2 This is a schematic diagram of the structure of a retrieval enhancement generation device provided in an embodiment of this disclosure, as shown below. Figure 2 As shown, it includes:

[0073] The generation unit 21 is used to receive multi-source heterogeneous data in real time through a streaming processing engine, and to perform vectorization encoding on the multi-source heterogeneous data to generate dynamic knowledge vectors.

[0074] Update unit 22 is used to trigger incremental index updates based on a predetermined time window and update the hierarchically stored dynamic knowledge base using the dynamic knowledge vector. The dynamic knowledge base includes a real-time data storage layer and a historical knowledge storage layer.

[0075] Retrieval unit 23 is used to select a hybrid retrieval strategy according to the query type, and jointly retrieve real-time data and historical knowledge base in the dynamic knowledge base to obtain relevant data;

[0076] The optimization unit 24 is used to generate initial results by combining the retrieved relevant data with streaming context management, and optimize the output through a dual verification mechanism.

[0077] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and the principle is the same, so it is not limited in this embodiment.

[0078] For a description of the features in the embodiment corresponding to the retrieval enhancement generation device, please refer to the relevant description in the embodiment corresponding to the retrieval enhancement generation method, which will not be repeated here.

[0079] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described methods for search enhancement generation.

[0080] Embodiments of this application also provide a computer-readable storage medium storing a computer program configured to execute the steps in any of the above-described methods for search enhancement generation.

[0081] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0082] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described methods for search enhancement generation.

[0083] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above-described methods for search enhancement generation.

[0084] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0085] The foregoing has provided a detailed description of the search enhancement generation method, apparatus, electronic device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for enhancing retrieval generation, characterized in that, include: The system receives multi-source heterogeneous data in real time through a streaming processing engine, and performs vectorization encoding on the multi-source heterogeneous data to generate dynamic knowledge vectors. Incremental index updates are triggered based on a predetermined time window, and the dynamic knowledge base stored in a hierarchical manner is updated using the dynamic knowledge vector. The dynamic knowledge base includes a real-time data storage layer and a historical knowledge storage layer. Select a hybrid retrieval strategy based on the query type, and jointly retrieve real-time data and historical knowledge from the dynamic knowledge base to obtain relevant data; The retrieved relevant data is combined with streaming context management to generate initial results, and the output is optimized through a dual verification mechanism.

2. The method for generating enhanced search results according to claim 1, characterized in that, The step of receiving multi-source heterogeneous data in real time through a streaming processing engine and vectorizing the multi-source heterogeneous data includes: A distributed stream processing framework is used to aggregate structured logs, unstructured text, and time-series sensor data in a sliding window manner. A complex event processing rule engine is used to define event triggering logic to trigger a knowledge retrieval request when a preset event pattern is detected.

3. The method for generating enhanced search results according to claim 2, characterized in that, The sliding window time of the distributed stream processing framework is a first duration; The event triggering logic includes maintaining short-term and long-term states based on a business entity state management model. The short-term state stores real-time data within a second time period, and the long-term state enables historical data backtracking through persistent storage.

4. The method for generating enhanced search results according to claim 1, characterized in that, The incremental index update triggered based on a predetermined time window includes: The vector database is updated every third time interval, and a time-to-live strategy is used to manage data storage. The hierarchical dynamic knowledge base includes a hot storage layer and a cold storage layer. The hot storage layer uses a graphics processor to accelerate indexing and achieve high-speed vector writing, while the cold storage layer uses a hybrid index that combines semantic retrieval and keyword retrieval.

5. The method for generating enhanced search results according to claim 4, characterized in that, The hot storage layer index is sharded by hour, and expired data is automatically archived to the cold storage layer; the cold storage layer stores the historical case database and the full text of laws and regulations, and performs logical paragraph segmentation of the text through a semantic segmentation algorithm, including basic segmentation, related segmentation and dynamic segmentation.

6. The method for generating enhanced search results according to claim 1, characterized in that, The selection of a hybrid retrieval strategy based on query type includes: For queries targeting specific entities, a keyword search mode is used. For open-ended questions, a semantic retrieval model is used to obtain relevant data, which is then reordered and the top three results are selected. For complex scenarios, a hybrid approach is adopted to integrate keyword retrieval and semantic retrieval results, and ambiguity is eliminated through cross-validation.

7. The method for generating enhanced search results according to claim 1, characterized in that, The process of generating initial results from the retrieved relevant data using streaming context management and optimizing the output through a dual-validation mechanism includes: The semantic vectors of historical dialogues are maintained as streaming context and dynamically weighted through a multi-head attention mechanism; A rule engine is used to verify the numerical range and a lightweight classification model is used to filter low-confidence generated content, and the prompt word template is updated.

8. A retrieval enhancement generation apparatus, characterized in that, include: The generation unit is used to receive multi-source heterogeneous data in real time through a streaming processing engine, and to vectorize and encode the multi-source heterogeneous data to generate dynamic knowledge vectors. An update unit is used to trigger incremental index updates based on a predetermined time window and update the hierarchically stored dynamic knowledge base using the dynamic knowledge vector. The dynamic knowledge base includes a real-time data storage layer and a historical knowledge storage layer. The retrieval unit is used to select a hybrid retrieval strategy based on the query type, and jointly retrieve real-time data and historical knowledge from the dynamic knowledge base to obtain relevant data. The optimization unit is used to generate initial results from the retrieved relevant data by combining streaming context management, and optimize the output through a dual verification mechanism.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the retrieval enhancement generation method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method of retrieval enhancement generation according to any one of claims 1-7.

Citation Information

Cited By

  • Retrieval system and method based on retrieval enhancement generation

    CN121935302A