Retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering

By using dynamic weight allocation and DeepSeek model reordering, the problem of fixed weights in mixed search results in the RAG system is solved, improving the accuracy and adaptability of search results and optimizing the sorting ability for complex queries.

CN120892553APending Publication Date: 2025-11-04CHINA UNIV OF MINING & TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510931184.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

In existing RAG system retrieval methods, the fixed weights of mixed retrieval results lead to insufficient adaptability to query intent, making it difficult to balance precise term matching and semantic relevance. Furthermore, the Rerank model has limited ranking capabilities for complex queries.

Method used

By employing dynamic weight allocation and multi-dimensional calibration mechanisms, combined with BM25 and vector retrieval, and utilizing the DeepSeek model for re-ranking, the query feature set and score fusion are optimized to achieve dynamic adaptation to different query types.

Benefits of technology

It significantly improves the accuracy and relevance of search results, enhances the performance of the RAG system in complex query scenarios, and provides more efficient and reliable search support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892553A_ABST
    Figure CN120892553A_ABST
Patent Text Reader

Abstract

The invention provides a retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering, and belongs to the technical field of artificial intelligence. Constructing a domain query feature set, training a simple linear regression model, and dynamically predicting weights of keyword retrieval and vector retrieval according to user query; through quantile calibration, calculating a weight calibration factor, and adjusting a prediction weight; removing abnormal scores of keyword retrieval and vector retrieval by adopting a truncation normalization method, and generating score distribution of a unified dimension; and fusing the normalized scores of the query and candidate documents through convex combination, sorting to obtain TopK most relevant documents, calling a DeepSeek model to carry out correlation scoring on the query and candidate documents, and carrying out weighted fusion and resorting with the mixed scores to obtain TopN most relevant documents. According to the method, the problems of limitation of a fixed weight, insufficient query diversity adaptation and the like in RAG mixed retrieval are effectively solved, and the retrieval precision and stability in a complex query scene are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application provides a retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering, and belongs to the technical field of artificial intelligence. BACKGROUND

[0002] Large language models (LLMs) have made significant progress in natural language processing, text generation, and question-answering systems, thanks to their advanced pre-training neural network architecture. Through the pre-training of massive amounts of data, these models successfully capture and simulate the complexity of natural language, driving the development of machine translation, text summarization, and dialogue systems. Although LLMs have demonstrated superior performance in numerous downstream tasks, their application in specific vertical domains still faces many challenges. For example, outdated domain knowledge updates may lead to the generation of outdated information, while the illusion of generated content may result in inaccurate or fabricated answers. To address these challenges, researchers have proposed retrieval-augmented generation systems (RAG), which combine retrieval techniques to assist LLMs, aiming to provide more accurate and timely answers while avoiding the high cost of LLM fine-tuning.

[0003] RAG systems are simple to implement, easy to quickly deploy and use, but they also face some non-negligible challenges. Currently, the search method of RAG systems is relatively single, mainly relying on vector search, which may limit the matching accuracy of the retrieval results and keywords, and thus affect the recall rate of the documents. To improve the performance of RAG systems, researchers are exploring various improvement methods. For example, combining sparse retrieval and dense retrieval can focus on both the central theme and global features of the text segment, thereby improving retrieval accuracy; in the post-retrieval stage, fine-tuning reordering models (such as Cohere Rerank, BAAI / bge-reranker) are widely used to further improve the ranking of relevant documents by fine-tuning the initial retrieval results, optimizing the overall performance of the RAG system.

[0004] Although the above methods have good performance in the retrieval system of RAG systems, there are still several key defects: existing hybrid retrieval methods usually use fixed weights to fuse different retrieval results, which lacks dynamic adaptability to query intent, making it difficult to balance between term accurate matching and semantic relevance. The Rerank model is limited by fixed parameter size, and its sorting ability for complex queries is limited, making it difficult to adapt to different query types. Therefore, the present application improves the retrieval method based on the original shortcomings, characterized by more comprehensively capturing the semantics and context information of the query, and reordering based on the DeepSeek model, thereby improving the accuracy of the retrieval results. SUMMARY

[0005] The application aims at the deficiencies of the prior art, and provides a retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering.

[0006] To achieve the above technical purposes, the retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering of the application has the following steps:

[0007] Step 1: Analyze a plurality of heterogeneous field data sources, extract unstructured data, save as a JSON file format, perform inverted indexing based on a BM25 algorithm, and use the IndexFlatIP of FAISS to construct a dense vector index;

[0008] Step 2: Construct a field query feature set, train a simple linear regression model, dynamically predict the weights of keyword retrieval and vector retrieval according to user queries, and adapt to different query types;

[0009] Step 3: Use quantile statistics, Wasserstein distance and entropy ratio to realize the standardization alignment of different score distributions, and calculate the weight calibration factor of the scores of keyword retrieval and vector retrieval;

[0010] Step 4: Adopt adaptive truncated normalization technology based on median and standard deviation to process abnormal scores of keyword retrieval and vector retrieval, and generate a unified dimension score distribution;

[0011] Step 5: Merge the keyword retrieval and vector retrieval results through key-value mapping, combine the predicted weights and the weight calibration factor, perform convex combination linear weighting on the normalized scores, sort and return TopK documents;

[0012] Step 6: Call the DeepSeek model to score the relevance of queries and candidate documents, reweight and fuse with a fixed weight to reorder TopN relevant documents, and obtain the final retrieval result.

[0013] Beneficial Effects: This invention combines an optimized hybrid search engine with a re-ranking mechanism based on a large language model. Through adaptive weight estimation and quantile calibration, it dynamically integrates the advantages of keyword retrieval and vector retrieval, effectively addressing the diverse needs of short keyword queries and long context queries. This significantly improves the relevance and robustness of search results, overcoming the inconsistency issues of traditional methods across different query types and document distributions. Simultaneously, leveraging the structured output capabilities of the DeepSeekAPI, it performs refined scoring and re-ranking of the initial search results, further optimizing the accuracy of the results. This method significantly improves the performance of the RAG system in complex query scenarios, providing more efficient and accurate search results for RAG applications. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering of the present invention.

[0015] Figure 2 This is the LLMRerankingPrompt template in this embodiment of the invention. Detailed Implementation

[0016] The embodiments of the present invention will be further described below with reference to the accompanying drawings:

[0017] like Figure 1 As shown, this invention proposes a retrieval enhancement method based on dynamic hybrid retrieval and LLM re-ranking, belonging to the field of artificial intelligence technology. It constructs a domain query feature set, trains a simple linear regression model, and dynamically predicts the weights of keyword retrieval and vector retrieval based on user queries. Through quantile calibration, a weight calibration factor is calculated to adjust the predicted weights. A truncation normalization method is used to remove abnormal scores from keyword and vector retrieval, generating a score distribution with uniform dimensions. The normalized scores are fused through convex combination to obtain the Top K most relevant documents. The DeepSeek model is then used to score the relevance of the query and candidate documents, and the results are weighted and fused with the hybrid scores to re-rank the Top N most relevant documents. This invention effectively solves the limitations of fixed weights and insufficient adaptation to query diversity in RAG hybrid retrieval, significantly improving retrieval accuracy and stability in complex query scenarios.

[0018] Specifically, the following steps are included:

[0019] Step 1: Parse multiple heterogeneous domain data sources, extract unstructured data, save it as a JSON file, perform inverted indexing based on the BM25 algorithm, and build a dense vector index using FAISS's IndexFlatIP.

[0020] Step 2: Construct a domain query feature set, train a simple linear regression model, and dynamically predict the weights of keyword retrieval and vector retrieval based on user queries to adapt to different query types;

[0021] Step 3: Using quantile statistics, Wasserstein distance, and entropy ratio, standardize and align different score distributions, and calculate the weight calibration factors for keyword retrieval and vector retrieval scores;

[0022] Step 4: Adaptive truncation normalization based on median and standard deviation is used to process abnormal scores in keyword retrieval and vector retrieval, and generate a score distribution with uniform dimensions.

[0023] Step 5: Merge keyword search and vector search results by deduplication through key-value mapping, combine predicted weights and weight calibration factors, perform convex combination linear weighting on the normalized scores, sort and return the TopK documents;

[0024] Step 6: Call the DeepSeek model to score the relevance of the query and candidate documents, then re-weight and merge them with fixed weights to re-rank them to obtain the Top N relevant documents, thus obtaining the final search results.

[0025] Further, in step 1, high-quality domain data is downloaded from the network, and the unstructured data is parsed and cleaned using MinerU as the raw corpus. LangChain is used for recursive segmentation, and the structured data is stored in standard JSON format. Keyword retrieval is implemented based on the BM25 algorithm, and a dense vector index is constructed using FAISS's IndexFlatIP.

[0026] Further, step 2 constructs a diverse query feature dataset and trains a linear regression model. The feature dataset includes word count. , number of characters Keyword density Query entropy It covers queries ranging from short keyword searches to long, high-entropy searches; based on the principle of linear regression, the prediction function formula for the BM25 weights is as follows:

[0027]

[0028] in These are the regression coefficients learned by minimizing the mean squared error. The model optimizes the following objective function using the least squares method:

[0029]

[0030] Vector weights The training data contains 7 representative samples, and the model is optimized by simulating user search scenarios to ensure its generalization ability to unseen queries.

[0031] Furthermore, step 3 involves analyzing the BM25 retrieval score. And vector retrieval score Based on quantiles, the scores are cropped to the 5th to 95th percentiles to remove extreme values ​​and reduce the interference of outliers on the distribution statistics. and This facilitates subsequent quantile statistics and Wasserstein distance calculation.

[0032] Further, step 4 extracts the distribution features (10%, 25%, 50%, 75%, and 90% quantiles) from the cropped retrieval scores, denoted as the target score scores_n and the reference score ref_n, respectively. The distribution features are as follows:

[0033]

[0034] The distribution difference between scores_n and ref_n is quantified using Wasserstein distance; when calibrating BM25 retrieval scores, scores_n is the BM25 retrieval score and ref_n is the vector retrieval score; when calibrating vector retrieval scores, scores_n is the vector retrieval score and ref_n is the BM25 retrieval score, with Wasserstein distance used for quantification. As shown below:

[0035]

[0036] Add a mean of 0 and a standard deviation of 0 to scores_n and ref_n. tiny Gaussian noise Smoothing is performed to prevent numerical instability during entropy calculation; the entropy of the smoothed probability distribution is calculated using the Shannon entropy formula, yielding the entropies of scores_n and ref_n:

[0037]

[0038] The entropy ratio is obtained by taking the maximum value of the two distribution entropies. .

[0039] A dynamic weight is calculated based on the Wasserstein distance. The larger the Wasserstein distance, the greater the difference between the two distributions, thus reducing the contribution of the calibration term based on the distribution distance; where... and These are the 50th percentiles of the reference score and the target score, respectively. The median ratio formula is as follows:

[0040]

[0041] Wasserstein distance decay formula:

[0042]

[0043] Interquartile range ratio:

[0044]

[0045] The reciprocal of entropy:

[0046]

[0047] The weight calibration factor is obtained by combining the median ratio, Wasserstein distance decay, interquartile range ratio, and inverse entropy ratio, as shown in the following formula:

[0048]

[0049] Furthermore, step 5 uses the median and standard deviation to crop the raw scores. and The truncated scores were limited to the range of median ± 2.5 standard deviations. StandardScaler was used to perform Z-score normalization on the truncated scores, making their mean 0 and standard deviation 1. Finally, linear normalization was applied to the range [0,1] to obtain... and Combining the initial predicted weights obtained in step 2 and the weight calibration factors in step 4, the final hybrid retrieval score is calculated. The BM25 and vector retrieval results are deduplicated using key-value mapping, and the final hybrid retrieval score is calculated using the following formula. The TopK documents with the scores are then returned.

[0050]

[0051] Further, step 6 calls the DeepSeek model to score the relevance of the query and candidate documents, creates a RerankingPrompt template, and uses the DeepSeek API response for structured parsing, providing a score inference process based on the prompt words. This score is then fused with the mixed score using fixed weights to re-rank the documents, resulting in the Top N relevant documents and the final search results.

[0052] This invention may also have many other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various appropriate modifications and variations based on this invention, but all such modifications and variations should fall within the protection scope defined by the claims of this invention.

Claims

1. A retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering, characterized in that, The steps are as follows: Step 1: Parse multiple heterogeneous domain data sources, extract unstructured data, save it as a JSON file, perform inverted indexing based on the BM25 algorithm, and build a dense vector index using FAISS's IndexFlatIP. Step 2: Construct a domain query feature set, train a simple linear regression model, and dynamically predict the weights of keyword retrieval and vector retrieval based on user queries to adapt to different query types; Step 3: Using quantile statistics, Wasserstein distance, and entropy ratio, standardize and align different score distributions, and calculate the weight calibration factors for keyword retrieval and vector retrieval scores; Step 4: Adaptive truncation normalization based on median and standard deviation is used to process abnormal scores in keyword retrieval and vector retrieval, and generate a score distribution with uniform dimensions. Step 5: Merge keyword retrieval and vector retrieval results by deduplication through key-value mapping, combine predicted weights and weight calibration factors, perform convex combination linear weighting on the normalized scores, sort and return the TopK documents; Step 6: Call the DeepSeek model to score the relevance of the query and candidate documents, then re-weight and merge them with fixed weights to re-rank them to obtain the Top N relevant documents, thus obtaining the final search results.

2. The retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering according to claim 1, characterized in that, High-quality domain data is downloaded from the network, and unstructured data is parsed and cleaned using MinerU as the raw corpus. LangChain is used for recursive segmentation, and structured data is stored in standard JSON format. Keyword retrieval is implemented based on the BM25 algorithm, and dense vector indexes are built using FAISS's IndexFlatIP.

3. The retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering according to claim 1, characterized in that, Construct a diverse query feature dataset and train a linear regression model. The feature dataset includes word count. , number of characters Keyword density Query entropy It covers queries ranging from short keyword searches to long, high-entropy searches; based on the principle of linear regression, the prediction function formula for the BM25 weights is as follows: in These are the regression coefficients learned by minimizing the mean squared error. The model optimizes the following objective function using the least squares method: Vector weights The training data contains 7 representative samples, and the model is optimized by simulating user search scenarios to ensure its generalization ability to unseen queries.

4. The retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering according to claim 1, characterized in that, First, let's analyze the BM25 search score. And vector retrieval score Based on quantiles, the scores are cropped to the 5th to 95th percentiles to remove extreme values ​​and reduce the interference of outliers on the distribution statistics. and This facilitates subsequent quantile statistics and Wasserstein distance calculation.

5. The retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering according to claim 4, characterized in that, The distribution features (10%, 25%, 50%, 75%, and 90% quantiles) of the cropped retrieval scores are extracted and denoted as the target score scores_n and the reference score ref_n, respectively. The distribution features are as follows: The distribution difference between scores_n and ref_n is quantified using Wasserstein distance. When calibrating BM25 retrieval scores, scores_n represents the BM25 retrieval score, and ref_n represents the vector retrieval score; when calibrating vector retrieval scores, scores_n represents the vector retrieval score, and ref_n represents the BM25 retrieval score. The Wasserstein distance... as follows:

6. The retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering according to claim 5, characterized in that, Add a mean of 0 and a standard deviation of 0 to scores_n and ref_n. tiny Gaussian noise Smoothing is performed to prevent numerical instability during entropy calculation; the entropy of the smoothed probability distribution is calculated using the Shannon entropy formula, yielding the entropies of scores_n and ref_n: The entropy ratio is obtained by taking the maximum value of the two distribution entropies. .

7. The retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering according to claims 5 and 6, characterized in that, A dynamic weight is calculated based on the Wasserstein distance. The greater the Wasserstein distance, The smaller the value, the greater the difference between the two distributions, thus reducing the contribution of the calibration term based on distribution distance; where... and These are the 50th percentiles of the reference score and the target score, respectively. The median ratio formula is as follows: Wasserstein distance decay formula: Interquartile range ratio: The reciprocal of entropy: Combining the median ratio, Wasserstein distance decay, interquartile range ratio, and inverse entropy ratio, the weight calibration factor is obtained as follows:

8. The retrieval enhancement method based on dynamic hybrid retrieval and LLM reordering according to claim 1, characterized in that, Raw scores were cropped using the median and standard deviation. and The truncated scores were limited to the range of median ± 2.5 standard deviations. StandardScaler was used to perform Z-score normalization on the truncated scores, making their mean 0 and standard deviation 1. Finally, linear normalization was applied to the range [0,1] to obtain... and Combining the initial predicted weights obtained in claim 2 and the weight calibration factors in claim 7, the final hybrid retrieval score is calculated. The BM25 and vector retrieval results are deduplicated using key-value mapping, and the final hybrid retrieval score is calculated using the formula shown below. The TopK documents by score are then returned.

9. The hybrid retrieval optimization method based on dynamic calibration and intelligent fusion according to claim 1, characterized in that, The DeepSeek model is invoked to score the relevance of the query and candidate documents, a RerankingPrompt template is created, and the response of the DeepSeek API is used for structured parsing. The score inference process is given based on the prompt words. The results are then merged with the mixed score with a fixed weight and re-ranked to obtain the Top N relevant documents, resulting in the final search results.

Citation Information

Cited By

  • Knowledge document search engine based on word vectors and contextual summary

    CN121579655A

  • Cross-modal fashion commodity retrieval method and system based on multi-granularity feature fusion and storage medium

    CN121614632A

  • A cross-modal fashion commodity retrieval method and system based on multi-granularity feature fusion and a storage medium

    CN121614632B

  • Data retrieval method and system based on ES large model scoring

    CN122153036A

  • Intelligent agent end-to-end data asset structuring method of data guide enabling UGC

    CN122220410A