A document retrieval method and system based on reinforcement learning and large language model

CN122673255APending Publication Date: 2026-09-01HUNAN EQUATORIAL GALAXY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610819914.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0004]本发明提出了一种基于强化学习与大语言模型的文献检索方法及系统,旨在解决现有技术中存在的如何在有限的大语言模型调用预算下,自适应地调度多源检索结果的评估顺序,以同时保证文献检索的高相关性和低成本问题

Benefits of technology

第一,通过大语言模型智能体实现自然语言检索请求到结构化检索对象的转换,能够统一处理主题、标题、DOI、引用、作者、元数据等多种请求;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122673255A_ABST
    Figure CN122673255A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of information retrieval and artificial intelligence technology, specifically a document retrieval method and system based on reinforcement learning and a large language model. The method includes: using a large language model agent to parse user natural language queries into structured objects containing retrieval intent and content evaluation criteria; triggering multi-source recall based on intent to obtain a candidate document set and generating prior scores through reordering; modeling candidate documents from different recall sources as arms of a multi-armed gambling machine, further subdividing them into sub-arms, and initializing the expected return of each arm with the average reordered score; in iterative evaluation, employing the Thompson sampling algorithm, combining the expected return of the arm with historical rewards, to dynamically allocate the large language model evaluation budget. This invention significantly reduces the cost of calling the large language model while improving the relevance and recall rate of document retrieval and enhancing the interpretability of the retrieval results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of information retrieval and artificial intelligence, and in particular to a document retrieval method and system based on reinforcement learning and large language models. Background Technology

[0002] In existing retrieval systems, even with the introduction of vector retrieval or semantic reordering, users typically still need to pre-select a search entry point, or simply expand the user input into several keywords before uniformly calling the same search source. This approach lacks the ability to proactively orchestrate the retrieval task. For example, when a user asks for cited papers of a particular paper, the system should first locate the title or DOI, and then obtain the citation relationships; when a user asks for papers by a particular author in a specific area, the system should first disambiguate the author, and then perform relevance filtering based on topic constraints. The retrieval paths for different tasks differ, and without agent-level intent recognition and process orchestration, simple keyword retrieval or general question-answering models cannot reliably cover these differences.

[0003] On the other hand, the relevance of academic literature is not always reflected in the title. Some papers have broad titles, but the text excerpts or citations are highly relevant to the user's topic; other papers are repeatedly mentioned in the context of citations in other papers, reflecting their actual contributions to a particular research direction. Relying solely on titles, abstracts, or a single external search interface can easily lead to insufficient recall; on the other hand, having all candidate documents evaluated one by one by a large language model can result in high costs, long processing times, and unstable responses. Summary of the Invention

[0004] This invention proposes a document retrieval method and system based on reinforcement learning and a large language model. It aims to solve the problem in the prior art of how to adaptively schedule the evaluation order of multi-source retrieval results under a limited large language model calling budget, so as to simultaneously ensure high relevance and low cost of document retrieval.

[0005] In a first aspect, the present invention provides a document retrieval method based on reinforcement learning and a large language model, comprising: The large language model agent is invoked to parse the user's natural language retrieval request into a structured query object, which includes at least the retrieval intent and a set of content evaluation criteria. According to the search intent, the corresponding search pipeline is executed to obtain an initial candidate document set from multiple heterogeneous recall channels; Calculate the prior relevance score between each candidate document and the core topic query; Candidate documents from different recall channels are constructed into different retrieval arms, and the expected return of each retrieval arm is initialized using the mean of the prior relevance scores of the candidate documents within each retrieval arm. For each retrieval arm that still has unevaluated literature, the evaluation budget for the large language model in this round is allocated through Thompson sampling based on historical rewards and expected returns, and a corresponding number of candidate literatures are extracted. The large language model agent is invoked to perform criterion-by-criterion binary judgment on the extracted documents according to the content evaluation criterion set, and the reward value of each document is calculated based on the judgment result; the historical reward and expected revenue of the corresponding retrieval arm are updated using the reward value, and the documents are selected according to the reward value to form a retrieval result set.

[0006] Furthermore, the structured query object also includes at least one of the following: rewritten query, synonymous topic set, Boolean query, metadata constraint, and title list; the search intent is one of the following seven types: searching for papers by title or numeric object identifier, searching for references, searching for cited references, searching by topic, searching by topic with exclusion conditions, searching by author, and searching by metadata.

[0007] Furthermore, when the search intent is a topic-based search or a topic-based search with exclusion conditions, the search pipeline is executed in parallel: local index title search, full-text fragment search, and citation expansion search; the citation expansion search expands the recall of cited documents by parsing the citation marks in the full-text fragment search results.

[0008] Furthermore, each candidate document within a search arm is divided into a sub-arm according to the sorting order and a preset number of documents. Each sub-arm independently participates in the budget allocation and status update in the iterative evaluation.

[0009] Furthermore, regarding candidate literature Suppose that its corresponding set of content evaluation criteria includes Based on the following criteria, the binary decision vector output by the large language model agent is [ , ,..., ],in The reward value of this document is... ;when When =1, it is marked as a complete match. The time marker is a partial match, where This is a preset partial matching threshold.

[0010] Furthermore, the partial matching threshold Based on the number of criteria Setting: When hour, ;when hour, .

[0011] Furthermore, for any search arm Let the number of evaluated documents be... The cumulative reward is Expected return is ; Convert cumulative rewards into success counts at the criteria level. and number of failures Construct a Beta distribution And sampling to obtain exploration factors ; This round of budget allocation ,in For the total budget of this round, It is a preset minimum positive number.

[0012] Furthermore, when constructing the search results set, the total number of criteria hits for each document is considered. The scores are then superimposed on the prior relevance scores to generate the final ranking scores, and the data is sorted according to these scores.

[0013] Furthermore, the iterative evaluation ends when any of the following conditions are met: the number of fully matched and partially matched documents both reach the target value; no new valid documents are added for three consecutive rounds; the total number of evaluated documents reaches the maximum budget; or the number of iteration rounds reaches the upper limit.

[0014] Secondly, the present invention provides a document retrieval system based on reinforcement learning and a large language model, the system being used to execute the method, comprising: The query parsing module is configured to call a large language model agent to parse the user's natural language retrieval request into a structured query object; The multi-source recall module is configured to obtain an initial candidate document set in parallel from multiple heterogeneous recall channels based on the retrieval intent in the structured query object. The reordering module is configured to calculate a prior relevance score between each candidate document and the core topic query. The arm modeling module is configured to construct retrieval arms from documents in different recall channels and initialize the expected return of the arm using the mean of the prior relevance scores of documents within each retrieval arm. The reinforcement learning evaluation module has a built-in Thompson sampling unit and is configured to dynamically allocate the evaluation budget of the large language model based on the historical rewards and expected benefits of each retrieval arm during iterative evaluation. It also calls the large language model agent to perform criterion-by-criterion binary judgment on the documents according to the content evaluation criteria to obtain the reward value, and then update the status of the corresponding retrieval arm. The results filtering and output module is configured to filter literature based on reward values ​​and generate the final search results.

[0015] The technical effects of this invention are: First, it uses a large language model intelligent agent to convert natural language retrieval requests into structured retrieval objects, and can uniformly handle various requests such as topics, titles, DOIs, citations, authors, and metadata; Second, by constructing a multi-source candidate set through title retrieval, full-text fragment retrieval, and citation-extended retrieval, the recall coverage of topic retrieval is improved; Third, the reordering score is used as a cold start prior for the multi-armed gambling machine, the evaluation results of the large language model are used as rewards, and the evaluation budget is allocated through Thompson sampling, so that the limited large language model calls are prioritized to more valuable retrieval arms and sub-arms, which significantly reduces the model call cost while ensuring retrieval quality. Fourth, the reward function is directly determined by the content evaluation criteria generated during the query understanding phase. The reason for retaining each document can be traced back to the 0 / 1 judgment result of the specific criterion, making it easy for users to understand and verify. Fifth, the system supports streaming event output, which can display the search progress and evaluation iteration in real time, improving the interactive experience. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating a document retrieval method based on reinforcement learning and a large language model proposed in an embodiment of the present invention. Figure 2 A comparison chart of retrieval metrics on a self-built dataset of 5000 academic documents provided in this embodiment of the invention; Figure 3 A comparison chart of LLM evaluation budget and the number of valid results provided for embodiments of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] The technical problem to be solved by this invention is: under complex natural language literature retrieval requests, how to organically combine large language model agents, multi-source literature recall, reordering priors and reinforcement learning-based evaluation scheduling, so that the system can not only improve the coverage of relevant literature recall, but also control the evaluation cost of large language models, and output interpretable and traceable search results to users.

[0019] To address the aforementioned technical problems, this invention provides a literature retrieval method based on reinforcement learning and a large language model. (Reference) Figures 1 to 3As shown, this method first uses a large language model agent to parse the user's natural language query into a structured object containing search intent and content evaluation criteria; based on the intent, it triggers multi-source recall (including local index title retrieval, full-text fragment retrieval, and citation expansion retrieval) to obtain a candidate document set, which is then reordered to generate prior scores; the candidate documents from different recall sources are modeled as arms in a multi-armed gambling machine, and further subdivided into sub-arms, with the expected payoff of the arm initialized using the average of the reordered scores; in the iterative evaluation, the Thompson sampling algorithm is used, combining the current expected payoff and historical reward of each arm (calculated by the large language model based on the satisfaction rate of the document according to the content criteria), to dynamically allocate a limited large language model evaluation budget; the system prioritizes the evaluation of documents in high-potential arms and updates the arm status in real time according to the evaluation results until the termination condition is met.

[0020] The initialization and updating of the Beta distribution parameters include: In this embodiment, each search arm Maintain a Beta distribution for Thompson sampling.

[0021] (1) Initial parameters: Before the first round of evaluation, there was no historical data, so no prior information was used, i.e. , , corresponding to the uniform distribution Beta(1,1).

[0022] (2) Parameter Update: After each round of evaluation, the parameters are updated based on the evaluated literature. Assume an arm... The number of evaluated documents is The cumulative reward is Let the number of criteria be K. Then: ; ; in The total number of hits (real value) is rounded to the nearest integer; the denominator is... This represents the total number of criteria; adding 1 enables Laplace smoothing.

[0023] (3) Example: An arm evaluates two papers (K=3), with reward values ​​of 2 / 3 and 1 respectively. =1.6667, =5.0, total number of criteria is 2×3=6, therefore... =5+1=6, =(6 5)+1=2, the next round will sample from Beta(6,2).

[0024] (4) Sampling and Budget Allocation: Before the start of each round, from Beta ( , ) Sampling to obtain exploration factors Combined with current expected returns Calculate the allocation of this round to the arm The assessment budget for a.

[0025] In this embodiment, the literature retrieval system is deployed as a single server, providing a RESTful query interface. Users input natural language search requests through the front-end interface, such as: "Please search for highly cited review papers on graph neural networks from the past five years, which must discuss graph convolutional networks and attention mechanisms, and exclude spectral domain-based methods."

[0026] This embodiment deploys the literature retrieval system as a single server, providing a RESTful query interface. The server's hardware configuration must meet the computational requirements of large language model inference and reordering models. Taking a typical academic literature retrieval scenario (concurrent users ≤ 10, literature database size 5000~100,000 articles) as an example, the recommended configuration is as follows: Central Processing Unit (CPU): 16 cores or more, clock speed ≥ 2.5GHz, used to handle non-model inference tasks such as retrieval scheduling, request routing, and text preprocessing.

[0027] Graphics Processing Unit (GPU): At least one GPU with ≥24GB of video memory (such as NVIDIA A10, RTX 4090, V100, etc.) is required to load and run large language models (7B~13B parameters) and reordering cross-encoder models. If using a commercial LLM API (such as GPT-4o), a local GPU is not required; only a regular computing server is needed.

[0028] Memory (RAM): 64GB or more, used to cache local indexes, candidate document sets, metadata and intermediate results.

[0029] Storage: 500GB or larger SSD hard drive for storing local document indexes, fragment databases and system logs.

[0030] Network: Gigabit Ethernet to ensure stable communication with LLM API services (if used) or external databases.

[0031] If the database is larger (e.g., in the millions) or has higher concurrency requirements, the number of GPUs and memory capacity can be increased accordingly, and a distributed deployment (e.g., multi-node, load balancing) can be adopted. Those skilled in the art can select appropriate hardware configurations based on the actual data scale and performance requirements. The above parameters are for illustrative purposes only and do not constitute a limitation on the scope of protection.

[0032] In this embodiment, the large language model agent uses a general large language model (such as GPT-4o, LLaMA3-70B, etc.) as a base, and ensures stability through prompt word templates, low temperature parameters (0.2) and output verification mechanisms.

[0033] (a) Query parsing template: The output is limited to JSON format and includes intent (one of seven enumeration values), rewritten query, list of synonyms, Boolean search query, content evaluation criteria (3-4, each with name and description), and metadata constraints. The template explicitly requires "output only JSON, do not add any additional explanations".

[0034] (ii) Literature Evaluation Template: Input the set of criteria and the titles, abstracts, text excerpts, and citations of candidate documents, and output a binary array (e.g., [1,0,1]) equal to the number of criteria. The template emphasizes the priority of evidence: titles and abstracts are primary evidence, text excerpts are secondary, and citations are only supplementary.

[0035] (III) Output Validation and Error Tolerance: Parse the returned JSON, check the field integrity, the validity of the intent enumeration values, and the length of the criterion array (add default values ​​if there are fewer than 3). If the evaluation task outputs an incorrect format (such as returning a text description), automatically retry once and append "Please only output binary arrays". If the second retry still fails, mark the document as failed and do not participate in the reward calculation.

[0036] (iv) Model selection and degradation: High-precision models (such as GPT-4o) should be used first for query parsing tasks, while low-cost models (such as GPT-3.5-Turbo) can be used for literature evaluation tasks. If the LLM service is unavailable or times out (30 seconds), query parsing will be downgraded to regular rule extraction, and literature evaluation will be downgraded to relying solely on re-ranking scores.

[0037] Query parsing and structure generation: After receiving the above request, the system first calls the query parsing module. This module integrates a Large Language Model agent (hereinafter referred to as the LLM agent), which constrains the output format of the LLM through carefully designed prompts. The prompts explicitly define seven types of search intents and require the output to be a structured object conforming to JSONSchema.

[0038] In this embodiment, the structured query object output by the LLM agent It can be represented as: ; in, Indicates the search intent. This indicates the rewritten core topic query. Represents a set of synonymous topics. This represents a Boolean search query. Represents a set of content evaluation criteria, such as "Discussion on graph convolutional networks" "Discussing attention mechanisms" "This is a review article"; Represents metadata constraints, This indicates the title or list of DOIs. The content evaluation criteria are not supplementary information, but rather a direct source of the subsequent reward function.

[0039] The LLM agent also generates evaluation descriptions for each criterion, which guide model decisions during subsequent evaluations. For example, for the criterion... The evaluation criteria are: "Whether the paper explicitly introduces or uses graph convolutional networks."

[0040] Multi-source candidate retrieval: Based on search intent For "Search by Topic", the system calls the multi-source recall module to execute three search pipelines in parallel: Local index title retrieval: Using a rewritten query q and metadata constraints M, the title field is retrieved from the local Elasticsearch index, recalling approximately 300 candidate documents, denoted as . .

[0041] Full-text fragment retrieval: based on rewritten query q and synonym set Each word in the query sequence is used to retrieve data from a pre-built full-text fragment database. This database stores the paragraph segments of each paper. The system calculates the frequency, fragment score, and fragment position of each paper in relevant fragments. This approach retrieves approximately 800 candidate documents, denoted as [missing information]. .

[0042] Citation-based extended search: Search results for full-text segments For each segment in the document, the citation annotations (such as "[1]", "(Author, 2020)") are parsed, and the unique identifier of the cited paper (such as DOI or custom ID) is extracted. The frequency of each cited paper in the relevant segment and the context score are counted. This approach recalls approximately 200 candidate documents, denoted as .

[0043] The final initial candidate document set is the union of the three-way recall: ; in, Search by corresponding title Corresponding to full-text segment retrieval, Corresponding references expand the search. The system extends the search for collections. The system performs deduplication and priority merging: records with the highest overall source priority are retained, with the priority order being: full-text fragment search results > citation extended search results > title search results. Simultaneously, the system supplements each document's metadata from the local metadata database, including title, abstract, author, year, citation count, journal / conference name, etc., prioritizing the retention of richer information such as fragments and citation context.

[0044] For each candidate paper, the system constructs a standardized evaluation text: ; in, and This can be empty. The evaluation prompts for the large language model specify the priority of evidence as title, abstract, paragraph, and quotation, and the quotation is only used as supplementary evidence and cannot determine the criterion satisfaction on its own.

[0045] The formation of reordered priors: Before candidate documents enter SmartBaSE, the system calls the re-ranking model based on the query. A relevance score is calculated for the candidate titles, and this score is then written into the candidate documents. Field. Let the first... The title of the candidate document is represented as The reordering model is Then the prior relevance score is: ; The reordering score is not directly used as the final relevance indicator, but rather as a cold-start prior for multi-armed gambling machines. By... By transforming it into an arm-level prior, the system can prioritize more likely effective candidate sources in the first round of evaluation, while still retaining the exploration of other recall channels.

[0046] The specific implementation of the reordering model is as follows: In this embodiment, the reranking model employs a cross-encoder architecture, specifically using a pre-trained "BAAI / bge-reranker-v2-m3" model (or an equivalent model, such as ms-marco-MiniLM-L-6-v2). This model will query q and candidate document titles. After concatenation, the data is input into a Transformer encoder, which outputs a correlation score between 0 and 1. .

[0047] Model architecture: Input format is [CLS]+q+[SEP]+ +[SEP], after passing through 12 Transformer encoder layers, the vector at position [CLS] is taken and then passed through a linear layer and a Sigmoid activation function to output the correlation score.

[0048] Training method: Fine-tuning was performed using publicly available academic literature relevance annotation datasets (such as the MS MARCO document ranking dataset and TREC Deep Learning Track data). A binary classification cross-entropy loss function was used during training, with positive samples representing document title pairs relevant to the query and negative samples representing irrelevant document title pairs.

[0049] Parameter source: This embodiment directly uses the officially released weights of the pre-trained model (such as Hugging FaceModel Hub), without the need for self-training. If it is necessary to adapt to a specific academic field, further fine-tuning can be performed on a small-scale self-built domain dataset with a small number of steps (about 1000 steps).

[0050] Cold start processing: When a newly deployed system has no evaluation data, the re-ranking model directly loads and uses the pre-trained weights, and its output score serves as an unbiased prior. After deployment, the parameters of the re-ranking model itself are not updated; only the expected return of the retrieval arm is updated by the subsequent reinforcement learning module.

[0051] Arm modeling and initialization: First, construct search arms for candidate documents based on their source. For example, a topic search scenario should include at least a title search arm. (Corresponding source) Full-text fragment search arm (Corresponding source) ) and reference extension arm (Corresponding source) Since each source may contain a large number of candidate documents (e.g.) (There are 800 articles), and the system further divides each search arm into several sub-arms by a fixed number: ; in, Indicates the first One search source, This indicates the source is the next There are 100 candidate segments. In the source code implementation, each subarm is segmented every 100 articles according to the candidate order. This design enables the system to distinguish between different recall sources, as well as between high-confidence candidates at the beginning and low-confidence candidates at the end of the same source. This refers to the number of sub-arms. For example, after sorting candidate documents for the full-text fragment retrieval arm by reordering score from highest to lowest, the first 100 documents would be sub-arms. Articles 101-200 are for the subarm. And so on.

[0052] For each child arm The system initializes the expected return with the mean reordering score of candidate documents within that subarm: ; in, Indicates the subarm The candidate literature set in Indicate candidate documents The reordering score is calculated using this formula. This formula transforms the judgments of the fast reordering model into prior knowledge for reinforcement learning scheduling, ensuring that the multi-armed gambling machine does not blindly explore randomly in the initial stage. In this embodiment, each sub-arm has at least 100 articles, while the last sub-arm may have fewer than 100 articles; the average is still calculated based on the actual number.

[0053] LLM criterion reward function: The system analyzes candidate documents. The evaluation was performed using the large language model criteria, and the results were obtained. A binary result The reward value is defined as the criterion satisfaction rate: ; when When, the document is marked as an exact match; when If a document is found to be a partial match, it is marked as such; otherwise, it is not included in the valid result set. Threshold The number of criteria can be set according to the number of criteria. In the preferred embodiment, when the number of criteria is 3, the following is taken: Otherwise, it is acceptable. The key to this reward function is that the reward comes from content criteria automatically generated from user queries, rather than from click-through rates or simple keyword hits, thus more closely reflecting genuine academic search intent. Adaptive reinforcement learning iterative evaluation: In the During the evaluation round, the system maintains the number of evaluated documents for each active subarm. Cumulative rewards and average return To model the uncertainty of criterion hits using a Beta distribution, the system converts the cumulative reward into the number of successes at the criterion level: ; ; Then, the exploration factor for this round is obtained by sampling from the Beta distribution: ; Let the first The round's estimable budget is The system allocates the budget jointly based on the expected return and the Thompson sampling value: ; ; in, Indicates the first The round still has a set of active child arms that have not been evaluated. To prevent extremely small constants from being zero, This indicates that the current round starts from the subarm. The number of candidate documents extracted. This budget allocation formula embodies the combinatorial innovation of this invention: the re-ranking model provides... Large Language Model Criterion Evaluation ThompsonSampling generates based on success / failure uncertainty. The three parties jointly decide on the assessment budget.

[0054] The above budget allocation is not based on fixed-proportion sampling, nor is it evaluated solely according to reordering scores from highest to lowest. If a subarm initially has a high reordering score but consistently fails the LLM criterion evaluation, its cumulative reward will be... and average return The initial score of a subarm will decrease, and the subsequent budget will automatically decrease. However, if a subarm's initial score is not outstanding but it continuously produces perfectly or partially matched literature during evaluation, its Beta sampling and average return will increase, and the subsequent budget will automatically increase. Thus, the system can continuously revise the reordering prior during the retrieval process, gradually shifting the evaluated resources from candidate sources that "appear relevant" to candidate sources that are "verified as relevant by criteria."

[0055] State update and sort writeback: For the candidate literature evaluated in this round If it belongs to the subarm Then the system will update: ; ; ; Simultaneously, the system records the identifiers of evaluated documents to avoid duplicate evaluations. The evaluation results of candidate documents are then written to... The fields include a binary criterion vector and a reward value. During final sorting, the system can add the criterion hit count to the re-sorting score. ; This write-back mechanism allows the final sort to consider both the fast reordering score and the LLM criterion-by-criterion judgment result.

[0056] Termination conditions: SmartBaSE terminates when any of the following conditions are met: the number of fully matched and partially matched documents both reach the target value; no new valid documents are added for several consecutive rounds; the number of evaluated documents reaches the maximum budget; or the number of iteration rounds reaches the upper limit. Let the set of fully matched documents be... The partial matching set is The target quantities are respectively and The maximum budget is Then the termination condition of the target can be expressed as: ; The budget termination condition can be expressed as: ; in, This indicates that the literature collection has been evaluated. Through the above termination conditions, the system can achieve a controllable balance between search quality and model invocation cost.

[0057] When the system runs, the user first submits a natural language search request. The system calls a large language model agent to parse the request, obtaining the intent, topic, synonyms, Boolean expressions, content criteria, and metadata constraints. Subsequently, the intent routing module selects whether to search for a topic, title DOI, citations, author, or metadata based on the intent.

[0058] Under the streaming response mechanism, the system can sequentially push query decomposition results, extracted search descriptions, start and end statuses of each search channel, metadata filtering results, reordering status, the number of complete and partial matches added in each round by the adaptive evaluation module, and the final document list to the front end. In this way, users can see that the agent is not providing static answers all at once, but rather performing an observable search task. For researchers, these intermediate states can be used to determine whether the search query deviates from expectations; for system debugging, these events can also serve as log evidence for analyzing the quality of the recall sources and the quality of the LLM evaluation.

[0059] In topic retrieval, the system simultaneously performs local index title retrieval, full-text fragment retrieval, and citation expansion retrieval, and performs deduplication, completion, cleaning, and standardization on candidate documents. Then, the system calls a re-ranking model to generate prior scores for candidate documents and inputs the three candidate sources into an adaptive reinforcement learning evaluation module. This module allocates a large language model evaluation budget over multiple iterations, progressively filtering for fully matching and partially matching documents. Finally, the system outputs retrieval progress, evaluation status, and the final document list in a streaming event manner, and can generate review abstracts with citation tags based on the first few documents.

[0060] The closed-loop effect of large language model agents: The large language model agent in this invention can be understood as a retrieval controller composed of four sub-capabilities. The first is semantic parsing, used to convert user natural language requests into structured fields and determine retrieval intent. The second is retrieval plan generation, used to select paths such as title retrieval, fragment retrieval, citation expansion, author retrieval, or metadata retrieval based on intent. The third is criteria-based evaluation, used to break down the user's research topic into several indispensable content criteria and evaluate them one by one during the candidate literature evaluation stage. The fourth is result presentation, used to generate review-style abstracts with citation identifiers based on candidate literature after the final result set is formed.

[0061] The four capabilities mentioned above are not isolated from each other. The intent generated by semantic parsing determines the retrieval route; the candidate set generated by the retrieval route enters the reordering and adaptive reinforcement learning evaluation module; the reward generated by the criterionized evaluation is used to update the multi-armed gambling machine state; and the final result can be generated by the agent and fed back to the user. Therefore, in this invention, the LLM agent is not only the query understander at the entry point, but also the evaluator in the intermediate stage and the result interpreter at the exit point.

[0062] To prevent large language models from exhibiting unchecked behavior and leading to unstable evaluations, this invention imposes structured constraints on the agent's output. The query understanding phase requires outputting fixed fields, while the evaluation phase requires outputting a binary array consistent with the number of criteria, which is then validated through the model structure. This constraint allows LLM capabilities to be incorporated into a deterministic engineering process, avoiding the uncontrollable narratives and unverifiable conclusions common in traditional chatbot-style question-answering systems.

[0063] Key technical parameters: The key technical parameters of this invention can be adjusted according to deployment conditions. In a preferred embodiment, title retrieval in topic search recalls approximately 300 candidates by default, while full-text fragment retrieval searches for rewritten queries and synonymous topics separately. Before the system enters SmartBaSE, the title retrieval candidates can be truncated to the top 50, the fragment paper candidates to the top 800, and the citation expansion candidates to the top 200; each search arm is divided into sub-arms of 100 articles.

[0064] SmartBaSE's default maximum evaluation budget can be set to 1000 articles, the maximum number of iteration rounds to 5, the LLM concurrent evaluation count to 4, and the LLM batch evaluation size to 5. The exact match target number can be set to the smaller of the number of returned items and 40, and the partial match target number can be set to the smaller of the number of returned items and 60. When the number of content criteria is 3, the partial match threshold can be set to... When the number of content criteria is not 3, the partial matching threshold can be set to 3. .

[0065] In one embodiment, the system is deployed as a literature retrieval service, providing a query interface. When a user enters a search request for a specific research topic, the system first generates a structured query object using a large language model agent. The object's intent is topic retrieval, the query is rewritten as an English topic sentence, and several synonymous topics, Boolean search expressions, and three to four content evaluation criteria are generated.

[0066] The system then performs multi-source retrieval. Local index title retrieval yields a batch of title-related papers; full-text fragment retrieval uses rewrite queries and synonym searches to retrieve text fragments, and calculates the fragment frequency and score for each paper; the citation extension module parses citation annotations in the fragments to obtain papers cited in relevant paragraphs. The system deduplicates the three types of results and completes metadata such as title, abstract, author, year, and citation count.

[0067] Then, the system calls the re-ranking model for each of the three candidate categories to obtain... The candidate texts are then formatted into a text format suitable for LLM evaluation. SmartBaSE uses title search results, fragment search results, and citation expansion results as three initial search arms, further dividing them into sub-arms of 100 articles each. During each round of evaluation, SmartBaSE evaluates each sub-arm based on its performance. , , The number of LLM evaluations is assigned based on Thompson Sampling values. For each candidate output criterion, LLM generates a binary vector, which the system uses to calculate rewards and update subarm states.

[0068] If a paper meets all the criteria, it is included in the complete match set; if it meets some of the criteria and reaches the threshold, it is included in the partial match set. The system ends the evaluation after reaching the target number or budget limit, merges and sorts the complete and partial matches, and returns the results, which can generate review abstracts.

[0069] Experimental verification: (a) Experimental Dataset: To verify the effectiveness of the literature retrieval system described in this invention, a self-built academic literature dataset was constructed. This dataset contains 5000 records, each including the paper title, abstract, author, year, journal or conference name, citation count, subject area, external identifier, and optional full-text excerpts or citation context. The dataset covers multiple disciplines such as artificial intelligence, medical informatics, materials science, educational technology, and social science computing, simulating real-world research retrieval scenarios characterized by broad subject scope, diverse expression methods, and complex metadata constraints.

[0070] Several retrieval tasks were constructed in the experiment, each corresponding to a natural language retrieval request. Request types included topic retrieval, topic retrieval with exclusion criteria, author-topic retrieval, citation relationship retrieval, and metadata-constrained retrieval. For topic-based tasks, relevant literature sets were manually labeled, distinguishing between fully relevant and partially relevant documents. Fully relevant documents indicate that the paper simultaneously meets all core content criteria in the query; partially relevant documents indicate that the paper meets the main topic but lacks a certain limiting condition. Evaluation primarily used four metrics: Precision@20, Recall@20, MAP@20, and nDCG@20.

[0071] (II) Comparison Method: The following baseline methods were selected for comparison in the experiment. BM25 represents the traditional term matching retrieval method, used to measure the baseline performance of keyword retrieval. Vectorized retrieval represents a dense retrieval method that encodes queries and documents into vectors and then recalls them based on similarity. DPR represents a typical dual-tower dense retrieval method, used to compare the performance of semantic vector-based recall. ColBERT represents a post-interactive neural retrieval method, which improves relevance judgment through fine-grained interactions between queries and document tokens. SPLADE represents a neural sparse retrieval method, which combines sparse term representation and neural semantic expansion capabilities. HyDE represents a retrieval enhancement method based on hypothetical document generation, which generates hypothetical answers or hypothetical documents before performing vector retrieval. LLM-based Rerank represents a method that first recalls candidate documents and then re-ranks them using a large language model or cross-encoder.

[0072] The difference between this invention and the aforementioned baseline lies in the fact that this invention does not employ single-path recall, nor does it perform a single reordering at the end of the candidate set. Instead, it integrates the query understanding of the LLM agent, multi-source recall, reordering priors, criterionized semantic evaluation, and multi-armed gambler budget scheduling into a closed loop. In particular, the adaptive evaluation module reallocates the subsequent evaluation budget based on the reward results of each round of LLM criterion evaluation, enabling the system to prioritize the evaluation of retrieval arms more likely to produce relevant literature within a limited number of model calls.

[0073] (III) Evaluation Indicators Precision@20 represents the proportion of relevant literature among the first 20 returned results, reflecting the quality of the results users see first; Recall@20 represents the proportion of the first 20 results covering the labeled relevant literature, reflecting recall capability; MAP@20 represents the mean precision of the first 20 results, emphasizing the position of relevant literature in the ranking; nDCG@20 represents the cumulative gain after deduction, which can reflect the ranking quality of perfectly relevant and partially relevant literature. (The last part is a partial translation and doesn't need a direct translation.) The number of relevant documents in the results is ,but: ; Let the set of annotated relevant documents be . ,forward The result set is ,but: ; The average precision can be expressed as: ; in, Indicates the first Whether the results are related. nDCG@k can be represented as: .

[0074] from Figure 2 It can be seen that BM25 is effective for queries with specific keywords, but its Precision@20, Recall@20, MAP@20, and nDCG@20 are all low for tasks involving synonyms, cross-domain expressions, and text fragments. Vectorized retrieval and DPR can improve semantic recall, but they tend to rank papers with similar semantics but not fully satisfying task constraints at the top. ColBERT and SPLADE have better ranking performance than ordinary vector retrieval, but they still lack explicit judgment on multiple content criteria in user queries. HyDE can extend recall on some open topics, but the generated hypothetical documents may introduce additional topic shifts. LLM-based Rerank can improve ranking quality, but if all candidates are evaluated, the model call cost is high; if only a small number of candidates are reranked, it is easily affected by initial recall bias.

[0075] The method of this invention achieves the highest results across all four metrics. This is because the LLM agent first decomposes natural language queries into structured intents and content criteria, enabling the system to accurately select retrieval paths; multi-source recall simultaneously covers titles, full-text segments, and citation contexts, reducing omissions from single retrieval sources; the re-ranking model provides initial priors for candidates; and the adaptive evaluation module uses the LLM criterion satisfaction rate as a reward to dynamically adjust the evaluation budgets of different retrieval arms. Therefore, this invention can simultaneously improve the accuracy of top-ranked results and the coverage of relevant literature.

[0076] Figure 3This further illustrates the advantages of this invention in terms of model call costs. When using a conventional LLMRerank strategy, achieving high search quality typically requires evaluating all or a large number of candidates. However, this invention, through ThompsonSampling, concentrates the evaluation budget on high-yield search arms, achieving more relevant results even with a significantly reduced number of evaluated documents. This result demonstrates that this invention not only pursues single-shot ranking accuracy but also improves the effective document output per unit of LLM call in cost-constrained scenarios.

[0077] Experimental results show that on a self-built dataset of 5000 academic documents, the method of this invention outperforms BM25, vectorized retrieval, DPR, ColBERT, SPLADE, HyDE, and the LLM-based Rerank method in terms of Precision@20, Recall@20, MAP@20, and nDCG@20. This advantage does not stem from the capabilities of a single model, but rather from the synergistic effect of LLM agent-based structured query understanding, multi-source recall, re-ranking priors, and multi-armed gambling machine budget scheduling. Especially when user queries contain multiple topic constraints or require text fragment evidence, this invention can exclude documents that are only semantically similar but do not meet the criteria through criterion-based evaluation, thereby improving the accuracy and interpretability of the final returned results.

[0078] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.

Claims

1. A document retrieval method based on reinforcement learning and a large language model, characterized in that, include: The large language model agent is invoked to parse the user's natural language retrieval request into a structured query object, which includes at least the retrieval intent and a set of content evaluation criteria. According to the search intent, the corresponding search pipeline is executed to obtain an initial candidate document set from multiple heterogeneous recall channels; Calculate the prior relevance score between each candidate document and the core topic query; Candidate documents from different recall channels are constructed into different retrieval arms, and the expected return of each retrieval arm is initialized using the mean of the prior relevance scores of the candidate documents within each retrieval arm. For each retrieval arm that still has unevaluated literature, the evaluation budget for the large language model in this round is allocated through Thompson sampling based on historical rewards and expected returns, and a corresponding number of candidate literatures are extracted. The large language model agent is invoked to perform criterion-by-criterion binary judgment on the extracted documents according to the content evaluation criterion set, and the reward value of each document is calculated based on the judgment result; the historical reward and expected revenue of the corresponding retrieval arm are updated using the reward value, and the documents are selected according to the reward value to form a retrieval result set.

2. The method according to claim 1, characterized in that, The structured query object also includes at least one of the following: rewritten query, synonym set, Boolean query, metadata constraint, and title list; the search intent is one of the following seven types: search for papers by title or numeric object identifier, search for references, search for cited references, search by topic, search by topic with exclusion conditions, search by author, and search by metadata.

3. The method according to claim 1, characterized in that, When the search intent is a topic-based search or a topic-based search with exclusion conditions, the search pipeline is executed in parallel: local index title search, full-text fragment search, and citation expansion search; the citation expansion search expands the recall of cited documents by parsing the citation marks in the full-text fragment search results.

4. The method according to claim 1, characterized in that, Candidate documents within each search arm are divided into subarms according to the sorting order and a preset number of documents. Each subarm independently participates in the budget allocation and status update in the iterative evaluation.

5. The method according to claim 1, characterized in that, For candidate documents Suppose that its corresponding set of content evaluation criteria includes Based on the following criteria, the binary decision vector output by the large language model agent for candidate documents is [ , ,..., ],in , If a candidate document explicitly satisfies the i-th criterion, then the reward value for that document is... ;when When =1, it is marked as a complete match. The time marker is a partial match, where This is a preset partial matching threshold.

6. The method according to claim 5, characterized in that, The partial matching threshold Based on the number of criteria Setting: When hour, ;when hour, .

7. The method according to claim 5, characterized in that, For any search arm Let the number of evaluated documents be... The cumulative reward is Expected return is ; Convert cumulative rewards into success counts at the criteria level. and number of failures Construct a Beta distribution And sampling to obtain exploration factors ; This round of budget allocation ,in For the total budget of this round, It is a preset minimum positive number.

8. The method according to claim 5, characterized in that, When constructing the search results set, the total number of criteria hits for each document is considered. The scores are then superimposed on the prior relevance scores to generate the final ranking scores, and the data is sorted according to these scores.

9. The method according to claim 1, characterized in that, The iterative evaluation ends when any of the following conditions are met: the number of fully matched and partially matched documents both reach the target value; no new valid documents are added for three consecutive rounds; the total number of evaluated documents reaches the maximum budget; or the number of iteration rounds reaches the upper limit.

10. A document retrieval system based on reinforcement learning and a large language model, characterized in that, The system is used to perform the method according to any one of claims 1-9, comprising: The query parsing module is configured to call a large language model agent to parse the user's natural language retrieval request into a structured query object; The multi-source recall module is configured to obtain an initial candidate document set in parallel from multiple heterogeneous recall channels based on the retrieval intent in the structured query object. The reordering module is configured to calculate a prior relevance score between each candidate document and the core topic query. The arm modeling module is configured to construct retrieval arms from documents in different recall channels and initialize the expected return of the arm using the mean of the prior relevance scores of documents within each retrieval arm. The reinforcement learning evaluation module has a built-in Thompson sampling unit and is configured to dynamically allocate the evaluation budget of the large language model based on the historical rewards and expected benefits of each retrieval arm during iterative evaluation. It also calls the large language model agent to perform criterion-by-criterion binary judgment on the documents according to the content evaluation criteria to obtain the reward value, and then update the status of the corresponding retrieval arm. The results filtering and output module is configured to filter literature based on reward values ​​and generate the final search results.