Time-sensitive double-link hybrid retrieval enhancement method
Through the time-sensitive dual-link hybrid search method, combined with the improved inverted index and HNSW algorithm, the problem that the existing technology is difficult to effectively search prompts in a limited time is solved, and efficient retrieval and response in a time-critical scenario is achieved.
Patent Information
- Application Number
- CN202510090173.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing search enhancement generation technology is difficult to effectively retrieve and return multiple related prompts within a limited time, resulting in large language models being unable to respond in a time-critical scenario.
The time-sensitive dual-link hybrid search method is adopted, and the sparse search module and intensive search module are run in parallel, combined with improved inverted indexing and HNSW algorithms, the search list is quickly generated, and the list is optimized through the dual-link hybrid sorting module to ensure that as many effective search results can be returned within a limited time.
Retrieval efficiency and response timeliness are significantly improved within limited search time, ensuring that large language models can quickly acquire external knowledge and respond in a timely manner in a scenario where time is tight.
Smart Images

Figure CN119938856A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to retrieval enhancement generation of a large language model under time constraints, and belongs to the technical field of retrieval enhancement generation. Background Art
[0002] Retrieval-augmented generation (RAG) technology can integrate external knowledge to generate prompts, improve the accuracy and relevance of the answers of large language models, and is a key technology to alleviate the "hallucination" problem of large language models. The core link of RAG technology is to retrieve the most relevant information from the external knowledge base to the input question in order to generate effective prompts for the large model. Specifically, RAG retrieval methods can be divided into two categories: sparse retrieval and dense retrieval. Among them: sparse retrieval uses discrete representation for data indexing and retrieval, with TF-IDF and BM25 algorithms as typical representatives; dense retrieval uses deep learning technology to map data to vector space to establish indexes, and uses algorithms such as cosine distance, Euclidean distance, and inner product to recall data. These retrieval methods focus more on improving the retrieval recall rate, and rarely consider the impact of retrieval efficiency on the number of recalls under limited time. This is not conducive to recalling effective prompt information under limited retrieval time, and even fails to prompt large language models in time. In time-critical scenarios such as high-frequency trading and disaster relief, large models often need to quickly acquire external knowledge to process perceived information and respond promptly within a limited time. For this purpose, the present invention is proposed. The main idea is to set a fixed retrieval time parameter for the retrieval process, during which as many valid retrieval results as possible are returned to prompt the large language model. Summary of the invention
[0003] The purpose of the present invention is to retrieve as much valid information as possible within a limited retrieval time. Figure 1 As shown, the method mainly includes four modules: sparse retrieval module, dense retrieval module, dual-link hybrid sorting module, and time-sensitive list recall module, among which: the sparse retrieval module and the dense retrieval module use a single algorithm to retrieve data related to the question in the database and return a sorted list; the dual-link hybrid sorting module removes duplicates and re-sorts the data based on the output of the sparse retrieval module and the dense retrieval module; the time-sensitive list recall module uses certain rules to control the source of the list used to generate prompts. The following steps are used to carry out the retrieval:
[0004] First, run the sparse retrieval module and the dense retrieval module in parallel, where:
[0005] (1) The sparse retrieval module implements full-text retrieval based on the improved inverted index and BM25 algorithm. The inverted index table generally contains keywords and their corresponding data numbers, word frequencies, and positions. The improved inverted index ( Figure 2 ) will additionally include the keyword idf value and the keyword tf value in the document, calculated as:
[0006] idf i =ln(1+(mn i +0.5) / (n i +0.5))(1)
[0007] tf i =freq i ×(1+k1) / (freq i +k1×(1-b+b×dl / avgdl))(2)
[0008] In formula (1), n i is the number of documents containing keyword i, m is the total number of documents in the database; in formula (2), freq i is the number of keyword i in the document; k1 and b are two constants, 1.2 and 0.75 respectively; dl is the total number of word segments in the document, and avgdl is the average number of word segments in all documents in the database.
[0009] In sparse search, we first segment the question to get keywords, and then calculate the BM25 score of each data by querying the inverted file information corresponding to the keywords, that is, BM25 j =Σtf i ×idf i , and sort and generate list L1 accordingly, and at the same time determine whether the retrieval time has timed out. If not, take the first k data of L1 and add them to the list to be sorted L s .
[0010] (2) The dense retrieval module builds an index and performs retrieval based on the HNSW algorithm. When building the index, the database documents are converted into 1024-dimensional vectors using the embedding model, the maximum number of "neighbors" of each node is set to 1.5k (rounded down if not an integer), the Euclidean distance is selected to build the HNSW graph index, and the vector distance between the insertion point and the "neighbor" is recorded (only once if two vectors are "neighbors" of each other).
[0011] Before searching, the question is first converted into a 1024-dimensional vector, and a dynamic list is created to record document information and distance information; a "starting point" is randomly selected from the top layer of the HNSW graph index, and the point closest to the question vector among this point and all its "neighboring points" is found. This point is used as the next "starting point" to repeat the search. When the closest point no longer changes, the single-layer search is completed; when searching the lower layer, starting from the last "starting point" of the upper layer search, the above single-layer search steps are repeated until the bottom layer; when searching the bottom layer, the documents in the dynamic list are used as the order, and the distance between the document and all its "neighboring points" and the question vector is calculated respectively, and compared, sorted, and updated with the existing information in the list. When the dynamic list no longer changes, the bottom-level search is completed.
[0012] The vector distance in the retrieval process uses the Euclidean distance metric, that is, Where Q and C are the question vector and the document vector in the database respectively. When searching, set the 15% quantile of all distances in the graph index as the distance threshold, return the documents below the threshold in the dynamic list to generate list L2, and determine whether it has timed out. If not, add list L2 to the list to be sorted L s .
[0013] If both the sparse retrieval module and the dense retrieval module have finished running and have not exceeded the time window, the dual-link hybrid sorting module will be run. Otherwise, the L s Generate hints. Dual-link hybrid ranking is based on the sequence value (rs) of the document under the two retrieval methods. j , rdj), by calculating the comprehensive score RRF j =1 / (60+rs j )+1 / (60+rd j ), reorder the data to generate list L s ', and determine whether it has timed out. If it has not timed out, use L s 'Generate a prompt, otherwise use L s Generate prompts.
[0014] The advantages of this invention are:
[0015] (1) Setting up time-sensitive recall rules. Compared with the general mixed search process, the time to return the search list is shorter, and the response timeliness is stronger within a fixed time.
[0016] (2) When building the inverted index, the keyword idf value and its tf value in each document are calculated. Through the pre-calculation work, the amount of calculation of the BM25 algorithm in the retrieval stage is reduced, making the sparse retrieval process faster and more likely to return the retrieval results in time within the time window.
[0017] (3) When using the HNSW algorithm to build a graph index, the distance information is recorded. During retrieval, the 15% quantile of all distances in the graph is used as the accuracy threshold. The vector process list is dynamically returned to shorten the initial response time of vector retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solution of the embodiment of the present invention, the following is a brief introduction to the drawings required for use in the embodiment of the present invention. Obviously, the drawings described below are only some examples of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 : A time-sensitive dual-link hybrid retrieval enhancement method structure diagram;
[0020] Figure 2 : Improved inverted index structure;
[0021] Figure 3 : Flowchart of a time-sensitive dual-link hybrid retrieval enhancement method. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0023] The time-sensitive dual-link hybrid retrieval enhancement method of the present invention is as follows: Figure 3 As shown, the structure is Figure 1 shown.
[0024] First, the sparse retrieval module and the dense retrieval module are run in parallel. The user question input time is recorded as T0, and the retrieval time window is recorded as t, where:
[0025] (1) The steps of sparse retrieval are as follows:
[0026] ① Use the word segmenter to extract keywords from user questions to form a set W;
[0027] ② Calculate the BM25 score of each data in the database. Improved inverted index (the improved inverted index structure is as follows Figure 2 The keyword set in the inverted index is denoted as I, and the question keywords and all their corresponding inverted lists are searched from the inverted index. The BM25 scores of all documents in the inverted list are calculated by the following formula:
[0028]
[0029] ③According to BM25 j Sort from high to low to generate list L1, record time T1. If T1-T0>t, do not output the list, otherwise take the first k data of L1 and add them to the list to be sorted L s .
[0030] (2) The steps of intensive retrieval are as follows:
[0031] ① Use the m3e-large model to convert the user question into a 1024-dimensional vector, denoted as Q. Calculate the 15% quantile of all distances in the HNSW graph index and denoted as S0.
[0032] ② Create a dynamic list of length 1.5k, which contains the document number and the distance S between the document and Q j , using the Euclidean distance metric, the calculation formula is Where: C is the vector corresponding to the document.
[0033] ③ Randomly select a “starting point” at the top level of the HNSW graph index, and log the ids and S of the “starting point” and all its “neighboring points”. j The score is added to the dynamic list. Press S j Sorted in ascending order.
[0034] ④ When searching above the bottom layer, select S j The smallest point is used as the next "starting point" and the S of the "starting point" and all its "neighboring points" is calculated. j The scores (the points that have been calculated will not be calculated again) are compared, sorted, and updated with the existing data in the dynamic list, and the search steps are repeated. When the "starting point" no longer changes, the single-layer search is completed. The last "starting point" is used as the initial "starting point" of the next layer, and the above search steps are repeated.
[0035] ⑤ When searching at the bottom level, find the "neighboring points" of each point in the dynamic list and calculate their S j The score is compared, sorted, and updated with the existing data in the list. When the dynamic list no longer changes, the underlying search stops.
[0036] ⑥When the previous step is executed, the score S of the first k data is returned along with the dynamic list j <S0 data, add to list L2, record time T2. If T2-T0>t, do not output the list, otherwise add list L2 to the list to be sorted L s .
[0037] If both the sparse retrieval module and the dense retrieval module have finished running and max(T1,T2)-T0<t, run the dual-link hybrid sorting module, otherwise directly use L s Generate hints. The dual-link hybrid sorting module is designed to s Rearrange, for j∈L s Calculate the composite score RRF j , the calculation formula is:
[0038] RRF j =1 / (60+rs j )+1 / (60+rd j )(4)
[0039] In formula (4), rs j is the sequence value of j in L1, rd j is the sequence value of j in L2. Then according to RRF jRearrange from high to low to generate a sorted list L s ', record time T3. If T3-T0>t, then L s Generate a prompt, otherwise L s 'Generate prompts.
[0040] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A time-sensitive dual-link hybrid retrieval enhancement method, characterized in that It includes sparse retrieval module, dense retrieval module, dual-link hybrid sorting module, time-sensitive list recall module, and uses the following steps to carry out retrieval: S1: Run the sparse search module and the dense search module in parallel. The sparse search steps are as follows: ①Use the word segmenter to extract keywords from user questions; ② Calculate the BM25 score of the document containing the question keyword, search for the question keyword and its corresponding inverted list from the improved inverted index, and summarize and calculate the BM25 score of the document in the inverted list. The calculation formula is: BM25 = ∑tf × idf; ③ Generate list L1 by sorting from high to low according to the BM25 score, and determine whether the search time has timed out. If not, take the first k data in L1 and add them to the list to be sorted L s ; S2: Run the sparse search module and the dense search module in parallel. The dense search steps are as follows: ① Use the embedding model to convert user questions into high-dimensional vectors and create a fixed-length dynamic list; ② Randomly select a "starting point" at the top level of the HNSW graph index and calculate the Euclidean distance between this point and all its "neighboring points" and the problem vector And add it to the dynamic list, and use the point with the smallest S as the next "starting point" to repeat the search. When the search point no longer changes, the single-layer search is completed; ③ When searching the lower layer, start from the last "starting point" of the upper layer search and repeat step ② until the bottom layer; ④ During the bottom-level search, the documents in the dynamic list are ordered, and the distance between the document and all its "neighbors" and the question vector is calculated respectively. The information in the list is compared, sorted, and updated. When the dynamic list does not change, the bottom-level search is completed. ⑤ When the previous step is running, the data that meets the distance threshold is returned with the dynamic list, added to the list L2, and it is determined whether it has timed out. If it has not timed out, the first k data in L2 are added to the list to be sorted L s ; S3: Determine whether to perform mixed sorting. If steps S1 and S2 are completed and have not timed out, perform dual-link mixed sorting. Otherwise, directly perform L sorting. s Generate prompts; S4: RRF with comprehensive score j =1 / (60+rs j )+1 / (60+rd j ) for list L s Reorder the data in to generate a sorted list L s ′; S5: Determine whether steps S1-S4 have timed out. If not, s 'Generate a prompt, otherwise use L s Generate prompts.
2. The time-sensitive dual-link hybrid retrieval enhancement method according to claim 1 is characterized in that The improved inverted index in step S1 includes not only the keyword and its corresponding data number, word frequency and position, but also the keyword idf value and the keyword tf value in the document.
Citation Information
Patent Citations
Semantic similarity vector re-sparse coding indexing and retrieval method
CN114860868A
Retrieval enhancement generation system and method based on knowledge graph
CN117973540A
Research report question and answer method, device and equipment based on research report knowledge base and storage medium
CN118838995A
Inspection report generation method and device for cloud network networking scene, medium and product
CN119271802A
Retrieval aware embedding
US20220253435A1