Document retrieval query rewriting method and system based on large language model
Through the document search query rewriting method based on the large language model, the inaccuracy and noise problems of existing systems when dealing with long-tail queries are solved, more accurate query understanding and rewriting are achieved, and retrieval relevance and user satisfaction are improved.
Patent Information
- Application Number
- CN202411939591.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-26
Smart Images

Figure CN120011482A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to natural language processing, and in particular to a method and system for improving query rewriting in a document retrieval system to enhance the ability to understand and process long-tail queries. Background Art
[0002] Query rewriting is an important tool to improve user search efficiency. Traditional query rewriting methods focus on queries with higher frequency, but for those long-tail queries with lower frequency, it is difficult to produce good recommendation effects due to the sparsity of their information. Long-tail keywords are more specific, longer search phrases (usually three or more words) that can clearly convey the searcher's intention. Short-tail keywords are phrases with general meanings, such as "mobile phone", "weight loss". Generally speaking, long-tail keyword phrases are longer and more specific search terms that can better describe specific things or concepts, such as "the latest Android phone recommendations in 2024", "the fastest way to lose weight and recipes".
[0003] Since long-tail queries involve complex and diverse phrase combinations, traditional methods rely on manually constructed dictionaries or rules and cannot effectively handle a large number of long-tail queries. Existing document retrieval systems often face problems such as inaccurate query understanding, poor rewriting effects, and poor relevance of retrieval results when processing long-tail queries. These problems mainly stem from the fact that existing query rewriting models lack the ability to fully understand and effectively handle long-tail queries. Traditional methods, such as pseudo-relevance feedback mechanisms, expand queries by adding keywords to the original query, but these methods rely on the quality of the initial retrieval results. Although query rewriting methods based on term replacement improve retrieval to a certain extent, they have limited effects when handling document ranking tasks. Doc2Query and Query2Doc use neural models to predict queries related to documents to enhance document representation, but these methods are inefficient in reasoning and are prone to deviations from the original meaning, resulting in poor retrieval relevance. Summary of the invention
[0004] The purpose of the present invention is to propose a document retrieval query rewriting method and system based on a large language model, which is intended to solve the problems that the existing document retrieval system is prone to introduce inaccuracy and noise when processing long-tail queries, which often leads to inaccurate query understanding, poor rewriting effect and poor relevance of retrieval results.
[0005] The technical solution to achieve the purpose of the present invention is: a document retrieval query rewriting method based on a large language model, comprising:
[0006] Step 1: Collect the query rewriting data set, perform relevance filtering and retrieval increments, and retain data samples that are highly relevant to the query rewriting task;
[0007] Step 2: Collect auxiliary task datasets that are highly relevant to the query rewriting task and construct a multi-task SFT dataset.
[0008] Step 3: Build a self-supervised fine-tuning model based on GPT-2, with query as input and query rewriting as output, and use the multi-task SFT dataset to train the self-supervised fine-tuning model;
[0009] Step 4: Treat the self-supervised fine-tuning model as an intelligent agent and perform goal alignment based on reinforcement learning to make it more consistent with user intent when generating query rewrites;
[0010] In step 5, a beam search algorithm is used to generate multiple candidate rewrites for each query, which are input into the target-aligned self-supervised fine-tuning model to retrieve a set of relevant documents.
[0011] Furthermore, in step 1, the query rewriting data set is collected, correlation filtering and retrieval increments are performed, and data samples highly relevant to the query rewriting task are retained. The specific method is as follows:
[0012] Step 1.1, obtain the initial query rewriting dataset from the document retrieval system;
[0013] User initiated query x i ∈X, the document retrieval system generates a series of query rewrites Y={y1,y2,y3,...,y n}, select the highest ranking y1 as the first choice, and construct an initial query rewriting dataset D containing N samples:
[0014]
[0015] Where p(x) is the query distribution in the document search system, π origin and θ origin They represent the rewriting strategies and parameters of the document retrieval system respectively;
[0016] Step 1.2, in order to ensure the semantic relevance between the query and the rewriting in the document retrieval query rewriting, the initial query rewriting dataset D is subjected to relevance filtering;
[0017] Rewrite the initial query data set D and perform rejection sampling according to formula (2) to obtain the data set D rele :
[0018]
[0019] Among them, rele(·) and τ rele They represent the correlation method and its threshold respectively. The calculation rule of rele(·) is as follows:
[0020]
[0021] in, is the indicator function, f(·) is the function for evaluating the relevance between the document title and the query, τ' is the semantic relevance threshold of the query-document title pair, is the list of offline retrieved documents for the query rewrite y;
[0022] Step 1.3: In order to alleviate the “small recall” problem caused by long-tail queries, we use rejection sampling again to select the dataset D rele Relevance filtering, while considering the retrieval increment, retrieve the document titles that have recently interacted with the query x as supplementary information, and obtain the final query rewriting dataset D final :
[0023]
[0024] Among them, ε x is the list of interactive document titles for query x, incr(·) and τ incr They represent the increment function and its threshold respectively. The calculation formula of incr(·) is as follows:
[0025]
[0026] in, represents the offline retrieved text set of the query, Refers to a curated set of documents.
[0027] Furthermore, in step 2, we collect auxiliary task datasets that are highly relevant to the query rewriting task and construct a multi-task SFT dataset, where the auxiliary tasks include:
[0028] a) Quality classification task: After extracting query pairs from the initial query rewriting dataset D, they are manually annotated to evaluate the quality of the query rewriting;
[0029] b) Product title prediction task: select the document that the user recently queried and clicked as a reference to form a (query, document title) pair;
[0030] c) Thought Chaining Task: Using the original query building prompt, provide an improved query rephrasing and explain their thought process.
[0031] Furthermore, in step 3, a self-supervised fine-tuning model is constructed based on GPT-2, with query as input and query rewriting as output, and the self-supervised fine-tuning model is trained using the multi-task SFT dataset. The specific method is:
[0032] Given an original query x and its query rewrite y, the training objective of the self-supervised fine-tuning model is to minimize L SFT (θ SFT ), the calculation formula is as follows:
[0033]
[0034] Among them, D multi is a multi-task SFT dataset, (x, y) ~ D multi Indicates that from D multi In the sampling, E means taking the average of the loss of all samples, that is, calculating the expected loss, π SFT (·) and θ SFT Represents the GPT-2 model and its parameters.
[0035] Furthermore, in step 4, the self-supervised fine-tuning model is regarded as an intelligent agent, and target alignment is performed based on reinforcement learning, so that it is more in line with user intention when generating query rewriting. The specific method is as follows:
[0036] Define state s, action a, and reward r as follows:
[0037] s={x,t}(7)
[0038] a=SFTModel(s)(8)
[0039]
[0040] Where x is the original query rewritten, t is the number of rewrites, rele(·) represents the relevance function, incr(·) and τ incr represents the incremental function; SFTModel represents the self-supervised fine-tuning model, β1 and β2 represent the weights of the correlation function and the incremental function respectively;
[0041] The DQN algorithm is used to train the agent and update the parameters of the self-supervised fine-tuning model.
[0042] A document retrieval query rewriting system based on a large language model implements the document retrieval query rewriting method based on a large language model, realizes document retrieval query rewriting based on a large language model, and is divided into five modules to respectively execute steps 1 to 5.
[0043] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the document retrieval query rewriting method based on a large language model is implemented to achieve document retrieval query rewriting based on a large language model.
[0044] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the document retrieval query rewriting method based on a large language model is implemented to achieve document retrieval query rewriting based on a large language model.
[0045] Compared with the prior art, the present invention has the following significant advantages: it adopts a document retrieval query rewriting method based on a large language model to achieve understanding and rewriting of long-tail queries, and can effectively reduce the noise that may be introduced in the retrieval through fine-tuning and multi-task learning based on the large language model, as well as target alignment of query rewriting and retrieval results, which is beneficial to subsequent retrieval recall, improves retrieval relevance, and improves user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Rewrite the flow chart for document retrieval query;
[0047] Figure 2 This is a schematic diagram of the data set processing flow;
[0048] Figure 3 Schematic diagram of the interactive process of the reinforcement learning environment. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0050] The present invention discloses a document retrieval query rewriting method based on a large language model, the method flow is as follows: Figure 1 As shown, the following steps are included:
[0051] Step 1: Dataset preparation and processing
[0052] Perform relevance filtering and retrieval increments on the collected query rewriting data set to retain data samples that are highly relevant to the query rewriting task;
[0053] First, we obtain the initial query rewriting dataset from the existing document retrieval system, such as Figure 2 As shown. User initiates query x i ∈X, a series of query rewrites Y = {y1, y2, y3, ..., y n}, select the highest ranking y1 as the first choice, and construct an initial query rewriting dataset D containing N samples:
[0054]
[0055] Where p(x) is the query distribution in the document search system, π origin and θ originRespectively represent the rewriting strategy and parameters of the document retrieval system. Then, in order to ensure the semantic relevance of the query and the rewriting in the document retrieval query rewriting, rejection sampling is performed according to formula (2) to filter the relevance of D, and the dataset D is obtained. rele :
[0056]
[0057] Among them, rele(·) and τ rele They represent the correlation method and its threshold respectively. The calculation rule of rele(·) is as follows:
[0058]
[0059] in, is the indicator function, f(·) is the function for evaluating the relevance between the document title and the query, τ' is the semantic relevance threshold of the query-document title pair, is the offline retrieved document list of query rewriting y. Finally, in order to alleviate the “small recall” problem caused by long-tail queries, we again use rejection sampling to select D rele Filtering, while considering the retrieval increment, takes the document titles that have recently interacted with query x as supplementary information to better solve the long-tail query problem, and obtains the final query rewriting dataset D final :
[0060]
[0061] Among them, ε x is the list of interactive document titles for query x, incr(·) and τ incr They represent the increment function and its threshold respectively. The calculation formula of incr(·) is as follows:
[0062]
[0063] in, represents the offline retrieved text set of the query, Refers to the selected document set. At the same time, rele(·) and incr(·) are also used to evaluate the subsequent query rewriting results.
[0064] Step 2: Auxiliary task dataset preparation
[0065] Collect auxiliary task datasets that are highly relevant to the query rewriting task and combine them to form a multi-task SFT dataset;
[0066] In order to enhance the understanding ability of large language models (LLMs) for long-tail queries, three task datasets highly related to query rewriting were collected. The three auxiliary tasks are:
[0067] a) Quality classification task: extract query pairs from the initial query rewriting dataset D and manually annotate them to evaluate their quality;
[0068] b) Product title prediction task: select the document that the user recently queried and clicked as a reference to form a (query, document title) pair;
[0069] c) Chain of Thought: Using the original query building prompt, provide an improved query rewrite and explain the thought process;
[0070] The detailed prompts of these auxiliary tasks will be explained in the subsequent experimental section. The collected data and query rewriting dataset D final Mix to form the final multi-task SFT dataset D multi , and randomly shuffled in the subsequent model training phase.
[0071] Step 3: Model training
[0072] Model training can be divided into two parts: fine-tuning and target alignment.
[0073] a) Self-supervised fine-tuning (SFTModel): First, a base model is selected. The present invention uses the pre-trained GPT-2 model with 117M parameters and 12 layers. Then the multi-task SFT dataset D obtained in step 2 is used. multi The dataset is divided into training set, validation set and test set in a ratio of 8:1:1.
[0074] Given an original query x and its query rewrite y, assume the conditional probability p(y|x) = Π i=1 p(y i |y 0:i-1 ,x;θ SFT ), then the model training goal is to minimize L SFT (θ SFT ), the calculation formula is as follows:
[0075]
[0076] Among them, D multi is a multi-task SFT dataset, (x, y) ~ D multi Indicates that from D multi In the sampling, E means taking the average of the loss of all samples, that is, calculating the expected loss, π SFT (·) and θ SFT Represents the GPT-2 model and its parameters.
[0077] b) Goal alignment: Goal alignment based on reinforcement learning (RL)
[0078] During the fine-tuning process, SFTModel may learn preferences for the training dataset, which may not be completely consistent with the goals of the document retrieval system. It can generate multiple query rewriting candidates, but the quality of these query rewritings and the impression of the retrieval results need further evaluation and optimization. By considering indicators such as the relevance and increment of the retrieval, target alignment can ensure that the generated query rewriting can effectively expand the semantics of the original query and that the output matches the user's intention or task goal.
[0079] Goal alignment can force the model to learn the partial order of query rewrite results provided by offline feedback. Partial order means that in a set of query rewrite results, some rewrite results are considered better or worse than other rewrites. For example, if offline feedback tells the model that rewrite A is better than rewrite B, then the goal alignment process will guide the model to prefer rewrites similar to A in future generation. Reinforcement Learning (RL) is a machine learning method that learns optimal policies by continuously interacting with the environment. RL can clarify the goal by defining a reward function and optimize the policy through multiple trials and errors.
[0080] like Figure 3 As shown in the document retrieval system definition environment, SFTModel is regarded as an intelligent agent, the current state is s, the intelligent agent selects action a according to the current state, obtains reward r, and jumps to the next state s', which is defined as follows:
[0081] s={x,t} (7)
[0082] a=SFTModel(s) (8)
[0083]
[0084] Among them, x is the original query rewrite, t is the number of rewrites, and a is the query rewrite output by SFTModel, that is, the action. Each time a query is rewritten, it is input into the document retrieval system, and the retrieved document is obtained. The reward is calculated according to formula (9). When the cumulative reward reaches 10, the loop ends, the original query rewrite is initialized, and a new round of loop begins, setting β1 and β2 to 0.5.
[0085] The classic DQN algorithm in reinforcement learning is used to train the agent and update the SFTModel model parameters. DQN is an algorithm that combines deep learning and reinforcement learning. It contains two networks, one is the current network Q used to generate predictions, that is, SFTModel, and the other is the target network Q target , whose parameters are regularly copied from the Q network. The parameter update formula of the Q network is as follows:
[0086] LDQN (θ SFT )=E[(y i -Q target (s i ,a i θ target )) 2 ] (10)
[0087]
[0088] where θ target is the parameter of the target Q network, α is the learning rate, y i The calculation formula is shown in formula (11), where γ is the attenuation rate, which is generally 0.99:
[0089] y i =r i +γmax a Q target (s',a) (12)
[0090] Step 4: Query Rewriting Generation: Beam Search
[0091] Beam Search is a heuristic search algorithm commonly used in sequence generation problems, especially in the field of natural language processing, such as machine translation, text summarization and other tasks. Beam Search is used to generate multiple candidate rewrites for each query, and a set of relevant documents is retrieved using the trained SFTModel.
[0092] The query is input into the SFTModel to obtain query rewriting candidates, and these query rewriting candidates are rewritten and input into the retrieval system to retrieve a set of relevant documents.
[0093] The present invention also proposes a document retrieval query rewriting system based on a large language model, implements the document retrieval query rewriting method based on a large language model, realizes document retrieval query rewriting based on a large language model, and executes steps 1 to 5 respectively in five modules.
[0094] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the document retrieval query rewriting method based on a large language model is implemented to achieve document retrieval query rewriting based on a large language model.
[0095] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the document retrieval query rewriting method based on a large language model is implemented to achieve document retrieval query rewriting based on a large language model.
[0096] In summary, through reinforcement learning, the present invention can dynamically optimize query rewriting strategies, gradually adjust and adapt to the needs of different users, and thus achieve adaptive query rewriting. At the same time, reinforcement learning can continuously align user goals across multiple rounds of interactions through an optimization mechanism based on long-term rewards, thereby improving the global optimality of the rewriting strategy. Compared with traditional static methods, query rewriting can be continuously optimized through implicit feedback, further improving the relevance and response efficiency of query rewriting generation.
[0097] Example
[0098] In order to verify the effectiveness of the scheme of the present invention, the following experiment was carried out.
[0099] Step 1: Training set preparation and processing
[0100] Step 1.1: Collect 10 million records from the online rewrite log of the existing document retrieval system. Then perform two rounds of rejection sampling: relevance and increment, and obtain 342,071 (query, query rewrite) pairs. In addition, 100,032 manual query rewrite data are included in the dataset to ensure that the query rewrite of SFTModel meets human preferences.
[0101] Step 1.2: Collect (query, query rewrite) pairs from online logs, construct "prompt: Is this a qualified query rewrite?" and then have humans annotate to determine whether it is relevant, and construct a (quality classification) instruction dataset;
[0102] Step 1.3: Construct (query, document title) pairs using the query and the document title with the highest posterior under the query, and organize them into “prompt: Does this document title match the query content?”
[0103] Step 1.4: Have human annotations generate a rewrite of the query and supplement the thought chain process, and finally form the following "prompt: Your task is to rewrite the input query into a query that is easier to search for related products, and you need to give the thought process and query rewriting results. Use a semicolon to separate the thought process and query rewriting results."
[0104] Step 1.5: 35,000 samples were selected from each of the quality classification, document title prediction, and thought chain tasks, and compared with D final Combined to construct D multi Dataset.
[0105] Step 2: Model training
[0106] Step 2.1: SFT stage: preprocess the input data, remove stop words and perform word segmentation, etc. The hyperparameters are set as follows: learning rate is 1e-5, weight decay is 0.01, batch size is 16, training rounds are 5, and optimizer is AdamW.
[0107] Next, input the current original query into GPT-2, and give GPT-2 the prompt: For an information-seeking dialog, please help reformulate the question into rewrite that can fully express the user's information needs without the need of context, so that GPT-2 can rewrite the original query based on these contents.
[0108] Step 2.2: Target alignment stage: set α = 1 × 10 -6 , γ=0.99, maximum number of iterations episode=1×10 6 , batch_size=32, dataset D final
[0109] Step 2.2.1: Initialization: For any state s and a, Q(s,a) = 0, k = 0, t = 0
[0110] Step 2.2.2: From D multi Randomly extract batchB
[0111] Step 2.2.3: Input s into SFTModel to get action a
[0112] Step 2.2.4: Calculate r(s,a) according to formula (9); update Q(s,a)
[0113] Step 2.2.5: The cumulative reward does not exceed 10, t←t+1, s←(a,t), repeat step 2.2.3
[0114] Step 2.2.6: Every 20 steps, update the target network parameters according to formula (11)
[0115] Step 2.2.7: k←k+1, repeat step 2.2.2
[0116] Step 2.2.8: k>1×10 6 , and obtain the converged SFTModel
[0117] Step 3: Model testing: Based on the trained SFTModel, select an input query from the test set: “please tell me the climate change impact on agriculture.”
[0118] Step 3.1: Preprocess the query, including word segmentation and removal of stop words. The result is:
[0119] “climate / change / impact / agriculture”
[0120] Step 3.2: Use the trained SFTModel to generate large queries to rewrite the candidate probability distribution. The result list is:
[0121] "Cropyielddecline duetoclimate change":0.35
[0122] "Impactofclimate change onagriculturalwaterresources":0.25
[0123] "Impactofclimate change onagriculturalpests anddiseases":0.20
[0124] "effects ofglobalwarmingonagriculture":0.10
[0125] "climate change and its impact on farming":0.10
[0126] Step 4: Input the query rewriting candidate list into Beam Search. The specific steps are as follows:
[0127] Step 4.1: Select the k query rewriting candidates with the highest probability from the probability distribution generated in step 1.2 as the beam width, that is, the number of candidate sequences kept at each step, and select k=3;
[0128] Step 4.2: Use these query rewrite candidates as the initial state of Beam Search;
[0129] Step 4.3: For each candidate, use SFTModel to generate the next possible rewrite. Calculate the probability of each newly generated rewrite and multiply it with the probability of the current candidate (probability propagation); then select several rewrites with the highest probability as candidates for the next step, and the number is still k (maintaining the beam width); repeat the above process until the predetermined number of iterations is reached or a rewrite sequence that meets the conditions is generated, and the output result is as follows:
[0130] "Crop yield decline due to climate change"
[0131] "Impact ofclimate change on agricultural pests and diseases"
[0132] "Impact of climate change on agricultural water resources"
[0133] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0134] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A document retrieval query rewriting method based on a large language model, characterized in that: include: Step 1: Collect the query rewriting data set, perform relevance filtering and retrieval increments, and retain data samples that are highly relevant to the query rewriting task; Step 2: Collect auxiliary task datasets that are highly relevant to the query rewriting task and construct a multi-task SFT dataset. Step 3: Build a self-supervised fine-tuning model based on GPT-2, with query as input and query rewriting as output, and use the multi-task SFT dataset to train the self-supervised fine-tuning model; Step 4: Treat the self-supervised fine-tuning model as an intelligent agent and perform goal alignment based on reinforcement learning to make it more consistent with user intent when generating query rewrites; In step 5, a beam search algorithm is used to generate multiple candidate rewrites for each query, which are input into the target-aligned self-supervised fine-tuning model to retrieve a set of relevant documents.
2. The document retrieval query rewriting method based on a large language model according to claim 1, characterized in that: Step 1: Collect the query rewriting data set, perform relevance filtering and retrieval increments, and retain data samples that are highly relevant to the query rewriting task. The specific method is as follows: Step 1.1, obtain the initial query rewriting dataset from the document retrieval system; User initiated query x i ∈X, the document retrieval system generates a series of query rewrites Y={y1,y2,y3,...,y n }, select the highest ranked y1 as the first choice, and construct an initial query rewriting dataset D containing N samples: where p(x) is the query distribution in the document search system, π origin and θ origin They represent the rewriting strategies and parameters of the document retrieval system respectively; Step 1.2, in order to ensure the semantic relevance between the query and the rewriting in the document retrieval query rewriting, the initial query rewriting dataset D is subjected to relevance filtering; Rewrite the initial query data set D and perform rejection sampling according to formula (2) to obtain the data set D rele : Among them, rele(·) and τ rele They represent the correlation method and its threshold respectively. The calculation rule of rele(·) is as follows: in, is the indicator function, f(·) is the function for evaluating the relevance between the document title and the query, τ' is the semantic relevance threshold of the query-document title pair, is the list of offline retrieved documents for query rewriting y; Step 1.3, in order to alleviate the "small recall" problem caused by long-tail queries, rejection sampling is used again to select the dataset D rele Relevance filtering, while considering the retrieval increment, retrieve the document titles that have recently interacted with the query x as supplementary information, and obtain the final query rewriting dataset D final : Among them, ε x is the list of interactive document titles for query x, incr(·) and τ incr They represent the increment function and its threshold respectively. The calculation formula of incr(·) is as follows: in, represents the offline retrieved text set of the query, Refers to a curated set of documents.
3. The document retrieval query rewriting method based on a large language model according to claim 1, characterized in that: Step 2: Collect auxiliary task datasets that are highly relevant to the query rewriting task and construct a multi-task SFT dataset. The auxiliary tasks include: a) Quality classification task: After extracting query pairs from the initial query rewriting dataset D, they are manually annotated to evaluate the quality of the query rewriting; b) Product title prediction task: select the document that the user recently queried and clicked as a reference to form a (query, document title) pair; c) Thought Chaining Task: Using the original query building prompt, provide an improved query rephrasing and explain their thought process.
4. The document retrieval query rewriting method based on a large language model according to claim 1, characterized in that: Step 3: Build a self-supervised fine-tuning model based on GPT-2, with query as input and query rewriting as output. Use the multi-task SFT dataset to train the self-supervised fine-tuning model. The specific method is: Given an original query x and its query rewrite y, the training objective of the self-supervised fine-tuning model is to minimize L SFT (θ SFT ), the calculation formula is as follows: Among them, D multi is a multi-task SFT dataset, (x, y) ~ D multi Indicates that from D multi In the sampling, E means taking the average of the loss of all samples, that is, calculating the expected loss, π SFT (·) and θ SFT Represents the GPT-2 model and its parameters.
5. The document retrieval query rewriting method based on a large language model according to claim 1, characterized in that: Step 4: Treat the self-supervised fine-tuning model as an intelligent agent and perform goal alignment based on reinforcement learning to make it more consistent with user intent when generating query rewrites. The specific method is as follows: Define state s, action a, and reward r as follows: s={x,t} (7) a=SFTModel(s) (8) Where x is the original query rewritten, t is the number of rewrites, rele(·) represents the relevance function, incr(·) and τ incr represents the incremental function; SFTModel represents the self-supervised fine-tuning model, β1 and β2 represent the weights of the correlation function and the incremental function respectively; The DQN algorithm is used to train the agent and update the parameters of the self-supervised fine-tuning model.
6. A document retrieval query rewriting system based on a large language model, characterized in that: Implement the document retrieval query rewriting method based on a large language model as described in any one of claims 1-5 to achieve document retrieval query rewriting based on a large language model, and perform steps 1 to 5 respectively in five modules.
7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the document retrieval query rewriting method based on a large language model as described in any one of claims 1 to 5 is implemented to realize document retrieval query rewriting based on a large language model.
8. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the document retrieval query rewriting method based on a large language model according to any one of claims 1 to 5 is implemented to achieve document retrieval query rewriting based on a large language model.
Citation Information
Patent Citations
Query search method, query information processing method, equipment and storage medium
CN117520477A
Method for rewriting query text based on large language model
CN118626686A
Intelligent patent retrieval method and system based on Agent
CN118760761A
Intelligent search engine construction method based on large language model
CN118964589A
Systems and methods for prompt-based query generation for diverse retrieval
WO2024064249A1
Cited By
Tobacco business information query rewriting method and system based on direct preference optimization
CN120849437A
Construction method and use method of self-questioning generation model
CN122088719A