Query rewriting model training method, retrieval method and related device

By training a query rewriting model and combining it with multi-dimensional reward optimization, the problem of insufficient generalization ability of rule-based rewriting schemes is solved, generating high-quality, diverse and semantically consistent rewriting results, thereby improving retrieval recall.

CN121029947APending Publication Date: 2025-11-28IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511230061.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing rule-based query rewriting solutions rely on manual design, have poor generalization ability, and cannot effectively improve the semantic clarity and retrieval recall of query questions.

Method used

By training a query rewriting model, a reinforcement learning strategy is used to rewrite query question samples. The model is then optimized by combining retrieval performance rewards, diversity rewards, and semantic consistency rewards to generate diverse rewritten results that meet retrieval requirements.

Benefits of technology

It improves the rewriting effect of the query rewriting model, and the generated rewritten results are semantically consistent with the original query question. It also has strong robustness and generalization ability on multiple retrieval models, thereby improving the retrieval recall rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029947A_ABST
    Figure CN121029947A_ABST
Patent Text Reader

Abstract

The invention discloses a query rewriting model training method, a retrieval method and a related device, and relates to the technical field of information retrieval, the query rewriting model training method comprises the following steps: rewriting a query problem sample based on a query rewriting model to obtain a rewriting result set of the query problem sample, performing knowledge retrieval on rewriting results in the rewriting result set, determining retrieval effect rewards according to retrieval results of the rewriting results in the rewriting result set, determining diversity rewards according to the rewriting result set, and determining semantic consistency rewards according to query problem samples and the rewriting result set; and optimizing the query rewriting model according to the retrieval effect reward, the diversity reward and the semantic consistency reward. Through the query rewriting model training method disclosed by the invention, the query rewriting model with a relatively good rewriting effect can be obtained through training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information retrieval technology, and in particular to a query rewriting model training method, retrieval method and related apparatus. Background Technology

[0002] With the rapid development of artificial intelligence technology, the deep integration of natural language processing and information retrieval is driving significant changes in intelligent search and question-answering systems. In scenarios such as search engines, intelligent customer service, and knowledge base question answering, user-inputted queries often suffer from problems such as vague wording, redundant information, or semantic ambiguity. To improve retrieval recall, it is usually necessary to rewrite the query and then perform knowledge retrieval on the rewritten query.

[0003] Currently, the main solutions for rewriting query questions are rule-based, which use manually defined rewriting rules (such as synonym replacement and syntactic template matching) to rewrite the query question. However, rule-based rewriting solutions rely on manual design and have poor generalization ability.

[0004] To address the problems with rule-based query rewriting schemes, one current approach is to train a query rewriting model and then use this model to rewrite queries. However, how to train a query rewriting model with good rewriting performance is a problem that urgently needs to be solved. Summary of the Invention

[0005] In view of this, this application provides a query rewriting model training method, retrieval method, and related apparatus to obtain a query rewriting model with good rewriting effect. The technical solution is as follows:

[0006] A query rewriting model training method includes:

[0007] Based on the query rewriting model, the query question sample is rewritten to obtain the rewritten result set of the query question sample;

[0008] Knowledge retrieval is performed on the rewriting results in the rewriting result set to obtain the retrieval results of the rewriting results in the rewriting result set. Based on the retrieval results of the rewriting results in the rewriting result set, a retrieval performance reward is determined.

[0009] Based on the rewriting result set, a diversity reward is determined, wherein the diversity reward reflects the diversity of the rewriting results in the rewriting result set;

[0010] The query rewriting model is optimized based on the retrieval performance reward and the diversity reward.

[0011] In one possible implementation, the query rewriting model training method further includes:

[0012] Based on the query question sample and the rewritten result set, a semantic consistency reward is determined, wherein the semantic consistency reward reflects the semantic consistency between the query question sample and the rewritten results in the rewritten result set.

[0013] In one possible implementation, the query rewriting model training method further includes:

[0014] Noise is added to the query question sample to obtain the noisy query question sample;

[0015] Based on the query rewriting model, the noisy query question sample is rewritten to obtain the rewritten result of the noisy query question sample;

[0016] The rewritten results of the noisy query problem samples are added to the rewritten result set.

[0017] In one possible implementation, determining the search performance reward based on the search results of the rewritten results in the rewritten result set includes:

[0018] Based on the position of the standard answer to the query question sample in the search results of each rewritten result in the rewritten result set, determine the first search effect reward corresponding to each rewritten result in the rewritten result set;

[0019] The first search effect rewards corresponding to each rewritten result in the rewritten result set are merged to obtain the first merged search effect reward, which is used as the final search effect reward.

[0020] In one possible implementation, determining the first search performance reward corresponding to each rewritten result in the rewritten result set based on the position of the standard answer to the query question sample within the search results of each rewritten result in the rewritten result set includes:

[0021] For each rewrite result in the rewrite result set:

[0022] The search results of the rewritten result are divided into categories according to the preset categorization method. Each category is assigned a corresponding weight. The higher the relevance of the search result to the query question, the higher the category it belongs to, and the higher the category, the greater the weight.

[0023] Determine the category of the standard answer to the query question sample;

[0024] The first search performance reward corresponding to the rewritten result is determined based on the weight corresponding to the tier of the standard answer to the query question sample.

[0025] In one possible implementation, determining the search performance reward based on the search results of the rewritten results in the rewritten result set further includes:

[0026] The retrieval results of each rewritten result in the rewritten result set are rearranged based on the reordering model to obtain the confidence level of the retrieval results of each rewritten result in the rewritten result set.

[0027] Based on the confidence level of the retrieval results for each rewritten result in the rewritten result set, determine the second retrieval performance reward corresponding to each rewritten result in the rewritten result set;

[0028] The second search effect reward corresponding to each rewrite result in the rewrite result set is merged to obtain the second merged search effect reward.

[0029] The first fused search effect reward and the second fused search effect reward are combined to obtain the third fused search effect reward, which is used as the final search effect reward.

[0030] In one possible implementation, the step of performing knowledge retrieval on the rewrite results in the rewrite result set to obtain retrieval results for the rewrite results in the rewrite result set includes:

[0031] Based on multiple different retrieval models, knowledge retrieval is performed on each rewritten result in the rewritten result set to obtain the retrieval results of each rewritten result in the rewritten result set on each retrieval model.

[0032] The search performance rewards corresponding to each rewritten result in the rewritten result set are merged, including:

[0033] For each retrieval model, the retrieval performance rewards corresponding to each rewritten result in the rewritten result set on that retrieval model are merged to obtain the retrieval performance reward on that retrieval model.

[0034] The search performance rewards across different search models will be combined.

[0035] In one possible implementation, determining the semantic consistency reward based on the query question sample and the rewritten result set includes:

[0036] Calculate the semantic similarity between the query question sample and each rewritten result in the rewritten result set;

[0037] Based on the semantic similarity between the query question sample and each rewritten result in the rewritten result set, the semantic consistency reward corresponding to each rewritten result in the rewritten result set is determined.

[0038] The semantic consistency rewards corresponding to each rewrite result in the rewrite result set are merged to obtain the final semantic consistency reward.

[0039] In one possible implementation, determining the diversity reward based on the rewritten result set includes:

[0040] The similarity between each pair of rewritten results in the rewritten result set is calculated to obtain several similarity scores;

[0041] Based on the aforementioned similarities, diversity rewards are determined.

[0042] In one possible implementation, determining the diversity reward based on the plurality of similarities includes:

[0043] Several similarity intervals and corresponding rewards are obtained, wherein the several similarity intervals are obtained by dividing the similarity range based on a preset similarity, and the rewards corresponding to the several similarity intervals are different.

[0044] From the plurality of similarity intervals, determine the intervals in which the plurality of similarities respectively lie;

[0045] The rewards corresponding to the intervals where the similarities are located are merged to obtain the first diversity reward, which serves as the final diversity reward.

[0046] In one possible implementation, determining the diversity reward based on the plurality of similarities further includes:

[0047] The variance of the aforementioned similarities is calculated to obtain the similarity variance;

[0048] The second diversity reward is determined based on the similarity variance.

[0049] The first diversity reward and the second diversity reward are merged, and the merged diversity reward is used as the final diversity reward.

[0050] A second aspect of this application provides a retrieval method, comprising:

[0051] To obtain the target query question;

[0052] The target query problem is rewritten based on the query rewriting model to obtain the rewritten target query problem, wherein the query rewriting model is trained using any of the above-mentioned query rewriting model training methods;

[0053] A knowledge retrieval was performed on the rewritten target query question to obtain the retrieval results.

[0054] A third aspect of this application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:

[0055] The memory is used to store computer programs;

[0056] The processor is used to execute the computer program so that the electronic device can implement the steps of any of the above-described query rewriting model training methods, or implement the steps of the above-described retrieval method.

[0057] A fourth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the steps of any of the above-described query rewriting model training methods, or to implement the steps of the above-described retrieval method.

[0058] By employing the above technical solution, the query rewriting model training method provided in this application, after obtaining a query question sample, first rewrites the query question sample based on the query rewriting model. After obtaining the rewritten result set of the query question sample, on the one hand, knowledge retrieval is performed on the rewritten results in the rewritten result set. Based on the retrieval results of the rewritten results in the rewritten result set, the reward for the query rewriting model in the retrieval performance dimension is determined. On the other hand, based on the rewritten result set, the reward for the query rewriting model in the diversity dimension is determined. Then, based on the rewards in the retrieval performance dimension and the diversity dimension, the query rewriting model is optimized. The optimization of the query rewriting model based on the rewards in the retrieval performance dimension and the diversity dimension enables the query rewriting model to generate diverse rewritten results that meet retrieval requirements. In summary, the query rewriting model training method provided in this application can train a query rewriting model with good rewriting performance. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0060] Figure 1 A flowchart illustrating the query rewriting model training method provided in this application embodiment;

[0061] Figure 2 This is a schematic diagram illustrating the query rewriting model training process provided in an embodiment of this application;

[0062] Figure 3This is a schematic diagram of the structure of the query rewriting model training device provided in the embodiments of this application. Detailed Implementation

[0063] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0064] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0065] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0066] Considering that users' queries often have problems such as vague expression, redundant information or semantic ambiguity, it is usually necessary to rewrite the query in order to improve the retrieval recall rate. At present, the main rewriting solution for query is rule-based rewriting solution. However, rule-based rewriting solution has many problems such as reliance on manual design and poor generalization ability.

[0067] Given the numerous problems with rule-based rewriting schemes, the inventors of this case conducted research. The initial idea was to improve the retrieval model and enhance its robustness to low-quality query problems. However, improving the retrieval model has high computational costs, is prone to performance bottlenecks in retrieval scenarios that require large-scale recall, and cannot fundamentally improve the semantic clarity of query problems.

[0068] In view of this, the inventors of this case continued their research and came up with the idea of ​​using a reinforcement learning strategy to train a query rewriting model. Then, they used the trained query rewriting model to rewrite the user's query. When training the query rewriting model using a reinforcement learning strategy, they determined the retrieval performance reward for the rewriting result of the query rewriting model on the query question sample, and optimized the query rewriting model based on the retrieval performance reward.

[0069] However, focusing solely on retrieval performance when training a query rewriting model can lead to semantic distortion in the rewritten results or the generation of numerous highly repetitive query variants. To address this issue, the inventors of this case continued their research and ultimately proposed a query rewriting model training method. The query rewriting model trained using this method exhibits superior rewriting performance.

[0070] The query rewriting model training method provided in this application will be described below through the following embodiments.

[0071] Please see Figure 1 The diagram illustrates a flowchart of a query rewriting model training method provided in an embodiment of this application. This method may include:

[0072] Step S101: Based on the query rewriting model, rewrite the query question sample to obtain the rewritten result set of the query question sample.

[0073] like Figure 2 As shown, the query question can be input into the query rewriting model multiple times to rewrite it, so as to obtain multiple rewritten results of the query question sample. The multiple rewritten results of the query question sample constitute the rewritten result set of the query question sample.

[0074] Step S102a: Perform knowledge retrieval on the rewriting results in the rewriting result set to obtain the retrieval results of the rewriting results in the rewriting result set, and determine the retrieval effect reward based on the retrieval results of the rewriting results in the rewriting result set.

[0075] After obtaining the rewritten result set of the query question sample, knowledge retrieval can be performed on each rewritten result in the rewritten result set to obtain the retrieval result of each rewritten result in the rewritten result set. Then, based on the retrieval result of each rewritten result in the rewritten result set, the reward for the query rewriting model in the retrieval effect dimension can be determined, that is, the retrieval effect reward.

[0076] The retrieval performance reward is used to encourage the query rewriting model to generate rewritten results that achieve better retrieval performance.

[0077] Step S102b: Determine the diversity reward based on the rewritten result set.

[0078] To avoid the query rewriting model generating highly repetitive query variants, this application introduces a rewriting diversity dimension reward, namely the diversity reward. The diversity reward reflects the diversity of rewriting results in the rewriting result set and is used to encourage the query rewriting model to generate diverse rewriting results.

[0079] Step S103: Optimize the query rewriting model based on the search performance reward and diversity reward.

[0080] Optionally, to ensure that the semantics of the rewritten results generated by the query rewriting model are consistent with the semantics of the original query question sample, this application may introduce a reward for semantic consistency. That is, the query rewriting model trainer may further include:

[0081] Step S102c: Determine the semantic consistency reward based on the query problem sample and the rewritten result set.

[0082] The semantic consistency reward reflects the semantic consistency between the query question sample and the rewritten results in the rewritten result set. It is used to encourage the query rewriting model to generate rewritten results that are semantically consistent with the input query question sample.

[0083] Furthermore, the query rewriting model can be optimized based on retrieval performance rewards, diversity rewards, and semantic consistency rewards.

[0084] The query rewriting model is optimized based on multi-dimensional rewards, enabling it to generate high-quality rewriting results.

[0085] The query rewriting model training method provided in this application, after obtaining a query question sample, first rewrites the query question sample based on the query rewriting model. After obtaining the rewritten result set of the query question sample, it retrieves the rewritten results in the result set and determines the reward for the query rewriting model in the retrieval performance dimension based on the retrieval results. It also determines the reward for the query rewriting model in the diversity dimension based on the result set. Furthermore, it optimizes the query rewriting model based on the rewards for retrieval performance and diversity, enabling the model to generate diverse rewritten results that meet retrieval requirements. Additionally, a semantic consistency reward is introduced to ensure the model generates rewritten results that are semantically consistent with the original query question. In summary, the query rewriting model training method provided in this application can train a query rewriting model with good rewriting performance.

[0086] In some embodiments of this application, the query rewriting model training method may further include: adding noise to a query question sample to obtain a noisy query question sample; rewriting the noisy query question sample based on the query rewriting model to obtain a rewritten result of the noisy query question sample; and adding the rewritten result of the noisy query question sample to the rewritten result set. It should be noted that one rewritten result of the noisy query question sample can be obtained based on the query rewriting model, or multiple rewritten results of the noisy query question sample can be obtained based on the query rewriting model (multiple inputs of the noisy query question sample into the query rewriting model for rewriting can yield multiple rewritten results of the noisy query question sample).

[0087] Considering that user queries may contain incomplete or redundant information, this application proposes, to enable the query rewriting model to generate high-quality rewritten results, that after obtaining a sample query query, if... Figure 2 As shown, noise can be added to the query question sample (e.g., arbitrarily adding or deleting characters). Then, the noisy query question sample is input into the query rewriting model for rewriting. After obtaining the rewritten result of the noisy query question sample, the rewritten result of the noisy query question sample is added to the rewriting result set.

[0088] In some embodiments of this application, the implementation process of "step S102: performing knowledge retrieval on the rewriting results in the rewriting result set to obtain the retrieval results of the rewriting results in the rewriting result set, and determining the retrieval effect reward based on the retrieval results of the rewriting results in the rewriting result set" is described.

[0089] The process of performing knowledge retrieval on the rewritten results in the rewritten result set to obtain the retrieval results of the rewritten results in the rewritten result set, and determining the retrieval performance reward based on the retrieval results of the rewritten results in the rewritten result set, may include: performing knowledge retrieval on each rewritten result in the rewritten result set based on a retrieval model, and determining the retrieval performance reward based on the retrieval results of each rewritten result in the rewritten result set.

[0090] There are multiple ways to determine the retrieval performance reward based on the retrieval results of each rewritten result in the rewritten result set. In one possible implementation method, the process may include:

[0091] Step a1: Based on the position of the standard answer to the query question in the search results of each rewritten result in the rewritten result set, determine the search performance reward corresponding to each rewritten result in the rewritten result set.

[0092] For the i-th rewrite result q in the rewrite result set i Rewrite the result q based on the standard answer to the query question. i The position in the search results determines the rewritten result q. i The corresponding search performance reward process may include:

[0093] Step a11: According to the preset grading method, rewrite the result q i The search results are categorized.

[0094] Each tier has a corresponding weight. The higher the relevance of the search result to the query question sample, the higher the tier it occupies, and the higher the tier, the greater the weight.

[0095] For example, rewrite the result qi The search results are 10. (The last part, "q", appears to be a typo and can be omitted.) i The search results are represented as R i ={r1, r2, r3,r4, r5, r6, …, r 10}, where 10 search results are compared with the rewritten results q i The relevance of the results is sorted from highest to lowest, and can be categorized by the position of the first, third, and fifth search results. Thus, q i Search results R i It is divided into 4 levels, namely {r1, r2, r3, r4, r5, r6, …, r 10 That is, r1 is in the first gear, r2 and r3 are in the second gear, r4 and r5 are in the third gear, and r6~r 10 Currently in the 4th gear, there are preset weights for each of the 4 gears. The higher the gear, the greater the weight. The 1st gear is the highest gear, the 2nd gear is the second highest gear, and so on, with the 3rd and 4th gears in that order. The weight of the 1st gear is greater than that of the 2nd gear, which is greater than that of the 3rd gear, which is greater than that of the 4th gear.

[0096] Step a12: Determine the level of the standard answer to the query question sample.

[0097] In the example above, assume the standard answer to the query question sample appears in the search result R. i The second position, i.e., q i Search results R i If r2 represents the standard answer, then the standard answer to the query question sample is located in the second tier. Assume the standard answer to the query question sample appears in search result R. i The 5th position, i.e., q i Search results R i If r5 is the standard answer, then the standard answer of the query question sample is located in the 3rd tier.

[0098] Step a13: Determine the rewritten result q based on the weight corresponding to the rank of the standard answer to the query question sample. i Corresponding search performance rewards.

[0099] In one possible implementation, a base score for the retrieval effect dimension can be preset. The weight corresponding to the level of the standard answer to the query question sample is multiplied by the base score, and the result is used as the rewritten result q. i Corresponding search performance rewards.

[0100] Of course, this embodiment is not limited to this. For example, the weight corresponding to the grade of the standard answer to the query question sample can also be directly used as the rewritten result q. i Corresponding search performance rewards.

[0101] It should be noted that the above-mentioned grading method ({r1, / r2, r3, / r4, r5, / r6, …, r 10}) This is merely an example; this embodiment is not limited to using the above-described grading method to rewrite the result q. i The search results can be categorized, for example, the categorization method can also be {r1, r2, r3, r4, r5, r6, …, r 10}, {r1, r2, / r3, r4, / r5, / r6, / …, r 10} etc. As long as the rewritten result q i The search results are divided into multiple tiers, and different weights are set for different tiers (the higher the tier, the greater the weight). The search effect rewards are determined based on the tier of the standard answer and are all within the scope of protection of this application.

[0102] Step a2: Merge the search performance rewards corresponding to each rewritten result in the rewritten result set, and use the merged search performance reward as the final search performance reward.

[0103] In one possible implementation, the average search performance reward corresponding to each rewritten result in the rewritten result set can be calculated, and the average search performance reward obtained can be used as the final search performance reward. It should be noted that this embodiment is not limited to averaging; other fusion methods are also possible, such as direct summation.

[0104] This embodiment also provides another way to determine the retrieval performance reward based on the retrieval results of each rewritten result in the rewritten result set:

[0105] Step b1-a: Based on the position of the standard answer to the query question in the search results of each rewritten result in the rewritten result set, determine the first search performance reward corresponding to each rewritten result in the rewritten result set.

[0106] Step b2-a: Merge the first search effect rewards corresponding to each rewritten result in the rewritten result set to obtain the first merged search effect reward.

[0107] The specific implementation process of steps b1-a and b2-a can be found in the specific implementation process of steps a1 and a2, which will not be repeated here in this embodiment.

[0108] Step b1-b: Based on the reordering model, rearrange the retrieval results of each rewritten result in the rewritten result set to obtain the confidence level of the retrieval results of each rewritten result in the rewritten result set. Based on the confidence level of the retrieval results of each rewritten result in the rewritten result set, determine the second retrieval performance reward corresponding to each rewritten result in the rewritten result set.

[0109] For each rewritten result in the rewritten result set, the search results of that rewritten result can be rearranged based on a rearrangement model to obtain the confidence level of the search results of that rewritten result. Assuming there are 10 search results for that rewritten result, the confidence level of each of the 10 search results can be obtained based on the rearrangement model. After obtaining the confidence level of the search results of that rewritten result, the second search effect reward corresponding to that rewritten result can be determined based on the confidence level of that search result. For example, a base score can be preset, and the confidence level of each search result can be multiplied by the base score and then summed to obtain the second search effect reward corresponding to that rewritten result. Alternatively, the sum of the confidence levels of each search result can be directly used as the second search effect reward corresponding to that rewritten result.

[0110] Step b2-b: Merge the second search effect rewards corresponding to each rewritten result in the rewritten result set to obtain the second merged search effect reward.

[0111] In one possible implementation, the second search effect rewards corresponding to each rewritten result in the rewritten result set can be merged by averaging. Of course, this embodiment does not limit the merging method to averaging; the merging method can also be other, such as direct summation.

[0112] Step b3: Combine the first fusion search effect reward with the second fusion search effect reward to obtain the third fusion search effect reward, which will be used as the final search effect reward.

[0113] Optionally, the first fused retrieval effect reward and the second fused retrieval effect reward can be summed, and the summed reward can be used as the final retrieval effect reward. Of course, this embodiment does not limit the fusion method of the first fused retrieval effect reward and the second fused retrieval effect reward to direct summation; the fusion method can also be other, such as weighted summation.

[0114] It should be noted that, in another possible implementation, the above-mentioned second fusion search performance reward can also be used as the final search performance reward.

[0115] To ensure that the rewriting results of the query rewriting model perform well across multiple different retrieval models, this application provides an alternative implementation of "Step S102: Perform knowledge retrieval on the rewriting results in the rewriting result set to obtain the retrieval results of the rewriting results in the rewriting result set, and determine the retrieval performance reward based on the retrieval results of the rewriting results in the rewriting result set," including:

[0116] Step S1021: Based on multiple different retrieval models (e.g., bge-m3, qwen3, etc.), perform knowledge retrieval on each rewritten result in the rewritten result set to obtain the retrieval results of each rewritten result in the rewritten result set on each retrieval model.

[0117] Assuming there are N retrieval models, let's consider the i-th rewritten result q in the rewritten result set. i Based on N retrieval models, the rewritten result q is processed respectively. i A knowledge retrieval was performed, and the rewritten result q was obtained. i Search results R on N retrieval models i 1 R i 2 ...R i N .

[0118] Step S1022: Determine the retrieval performance reward based on the retrieval results of each rewritten result in the rewritten result set on each retrieval model.

[0119] Specifically, for the j-th retrieval model M out of N retrieval models... j Based on each rewrite result in the rewrite result set, in the retrieval model M j The search results are used to determine the search performance reward, and the result is obtained in the search model M. j The search performance reward is calculated by combining the search performance rewards obtained from N search models to obtain the final search performance reward.

[0120] Among them, based on each rewrite result in the rewrite result set, in the retrieval model M j The search results on the search model M determine the search results. j The specific implementation process of the search performance reward can be found in the above embodiment of the implementation process of "determining the search performance reward based on the search results of each rewritten result in the rewritten result set" (i.e., see steps a1~a2, steps b1~b3).

[0121] In one possible implementation, after obtaining the search performance rewards on N search models, the search performance rewards on the N search models can be averaged to obtain the final search performance reward. This embodiment is not limited to averaging for the fusion method of the N search performance rewards; other fusion methods are also possible, such as direct summation.

[0122] The retrieval performance reward obtained through the above method is used to encourage query rewriting models to generate rewritten results that are robust and generalizable across different retrieval models.

[0123] As mentioned in the above embodiments, in order to make the semantics of the rewritten results generated by the query rewriting model consistent with the semantics of the original query question sample, a semantic consistency reward can be introduced, that is, the semantic consistency reward is determined based on the query question sample and the rewritten result set.

[0124] In some embodiments of this application, the process of determining semantic consistency rewards based on query question samples and rewritten result sets is described.

[0125] The process of determining the semantic consistency reward based on the query question sample and the rewritten result set may include:

[0126] Step c1: Determine the semantic similarity between the query question and each rewritten result in the rewritten result set.

[0127] In one possible implementation, a vector representation of the query question can be obtained, and a vector representation of each rewritten result in the rewritten result set can be obtained. Then, the cosine similarity between the vector representation of the query question and the vector representation of each rewritten result in the rewritten result set can be calculated to obtain the semantic similarity between the query question and each rewritten result in the rewritten result set.

[0128] Optionally, the vector representation of the query question and the vector representation of each rewritten result in the rewritten result set can be obtained through an NLI model (Natural Language Inference Model).

[0129] It should be noted that this embodiment is not limited to using the above method to determine the semantic similarity between the query question and the rewritten result. Any method that can determine the semantic similarity between two texts is applicable to this application.

[0130] Step c2: Based on the semantic similarity between the query question and each rewritten result in the rewritten result set, determine the semantic consistency reward corresponding to each rewritten result in the rewritten result set.

[0131] In one possible implementation, a base score for semantic consistency can be preset. For each rewritten result in the rewritten result set, the semantic similarity between the query question and the rewritten result can be multiplied by the base score, and the result is used as the semantic consistency reward for that rewritten result. In another possible implementation, for each rewritten result in the rewritten result set, the semantic similarity between the query question and the rewritten result can be used as the semantic consistency reward for that rewritten result.

[0132] Step c3: Merge the semantic consistency rewards corresponding to each rewrite result in the rewrite result set to obtain the final semantic consistency reward.

[0133] Optionally, the semantic consistency rewards corresponding to each rewrite result in the rewrite result set can be merged by averaging (i.e., averaging the semantic consistency rewards corresponding to each rewrite result in the rewrite result set). Of course, this embodiment is not limited to averaging; other fusion methods are also possible, such as direct summation.

[0134] The above embodiments provide that, in order for the query rewriting model to generate diverse rewriting results, a diversity reward can be introduced, that is, the diversity reward is determined based on the rewriting result set.

[0135] Some embodiments of this application describe the process of "step S103: determining diversity rewards based on the rewritten result set".

[0136] In one possible implementation, the process of determining the diversity reward based on the rewritten result set includes:

[0137] Step S1031: Calculate the similarity of each pair of rewritten results in the rewritten result set to obtain several similarity scores.

[0138] Specifically, the rewrite results in the rewrite result set can be combined in pairs to obtain several sets of rewrite results. The similarity of each set of rewrite results can be calculated to obtain several similarity scores.

[0139] For example, the rewrite result set includes 6 rewrite results. Combining the 6 rewrite results in pairs yields 15 sets of rewrite results. Calculating the similarity of each of the 15 sets of rewrite results yields 15 similarity scores.

[0140] Optionally, when calculating the similarity between two rewritten results, the vector representation of each rewritten result can be obtained, and then the cosine similarity between the vector representations of the two rewritten results can be calculated. It should be noted that this embodiment does not limit the method of calculating the similarity between two rewritten results; any method that can determine the similarity between two rewritten results is applicable to this application.

[0141] Step S1032: Determine the diversity reward based on several similarities.

[0142] Based on several similarities, there are multiple ways to implement diversity rewards. This embodiment provides the following two optional implementation methods.

[0143] Based on several similarities, the first method for implementing diversity rewards is determined:

[0144] Step d1: Obtain several similarity intervals and the preset reward values ​​corresponding to each of the several similarity intervals.

[0145] Among them, several similarity intervals are obtained by dividing the similarity range based on preset similarity, and each of the several similarity intervals corresponds to a different preset reward.

[0146] For example, the similarity range is [0,1], and the maximum similarity s can be preset. h and minimum similarity s l Based on maximum similarity s h and minimum similarity s l The similarity interval [0,1] is divided into three similarity intervals [0, s]. l ), [s l , s h ] and (s h [,1], and set rewards for three similarity intervals, for example, for the similarity interval [s l , s h Set a maximum reward (e.g., 1) for the similarity interval (s). h Set a low reward (e.g., 0.5) for the similarity range [0, s]. l Set a minimum reward (e.g., 0).

[0147] It should be noted that the similarity between the two rewritten results lies in the interval [0, s]. l If the similarity between the two rewritten results is within the range [s], it indicates a significant difference, suggesting the rewritten result may be incorrect, and a minimum reward (e.g., 0) may be given. l , s h Within the range of ], the highest reward (e.g., 1) is given, and the similarity between the two rewritten results lies in the interval (s). h Within [1], it indicates that the two rewritten results are quite similar, and a lower reward (e.g., 0.5) is given.

[0148] Step d2: From several similarity intervals, determine the intervals in which several similarities are located.

[0149] For example, several similarity intervals are [0, 0.1), [0.1, 0.9] and (0.9, 1]. The reward corresponding to [0, 0.1] is 0, the reward corresponding to [0.1, 0.9] is 1, and the reward corresponding to (0.9, 1] is 0.5. If one of the similarity intervals is 0.75, its interval can be determined to be [0.1, 0.9].

[0150] Step d3: Merge the rewards corresponding to the intervals where the similarity is located to obtain a variety of rewards.

[0151] Optionally, the rewards corresponding to each similarity interval can be merged by averaging (i.e., averaging the rewards corresponding to each similarity interval). Of course, this embodiment is not limited to averaging; other fusion methods are also possible, such as direct summation.

[0152] Based on several similarities, a second method for implementing diversity rewards is determined:

[0153] Steps e1-a1: Obtain several similarity intervals and the preset reward values ​​corresponding to each of the several similarity intervals.

[0154] Among them, several similarity intervals are obtained by dividing the similarity range based on preset similarity, and each of the several similarity intervals corresponds to a different preset reward.

[0155] Steps e1-a2: From several similarity intervals, determine the intervals in which several similarities are located.

[0156] Steps e1-a3: Merge the rewards corresponding to the intervals where several similarities are located to obtain the first diversity reward.

[0157] The specific implementation process of steps e1-a1 to e1-a3 can be found in the specific implementation process of steps d1 to d3, which will not be repeated here in this embodiment.

[0158] Steps e1-b1: Calculate the variance for several similarities to obtain the similarity variance.

[0159] Steps e1-b2: Determine the second diversity reward based on the similarity variance.

[0160] In one possible implementation, the process of determining the second diversity reward based on the similarity variance includes: obtaining several similarity variance intervals and preset rewards corresponding to each similarity variance interval; determining the similarity variance interval containing the similarity variance from the several similarity variance intervals; and determining the reward corresponding to the similarity variance interval containing the similarity variance as the second diversity reward. Specifically, the several similarity variance intervals are obtained by dividing the similarity variance range based on the preset similarity variance, and each of the several similarity variance intervals corresponds to a different reward.

[0161] For example, the similarity variance range is [sv min , sv max The maximum similarity variance sv can be preset. h and minimum similarity variance sv l Based on maximum similarity variance sv h and minimum similarity variance sv l The similarity variance interval [sv min ,sv max Divided into three similarity variance intervals [sv] min , sv l ), [sv l , sv h ] and (sv h , sv max ], and set rewards for each of the three similarity variance intervals. For example, for the similarity variance interval [sv l , sv h Set a maximum reward (e.g., 1) for the similarity variance interval (sv). h [1] Set a low reward (e.g., 0.5) for the similarity variance interval [sv min , sv l Set a minimum reward (e.g., 0), calculate the variance for several similarities, and then use [sv] to calculate the similarity variance. min , sv l ), [sv l ,sv h ] and (sv h , sv max The similarity variance interval is determined in the [ ], assuming the similarity variance interval is (sv h , sv max ], then the second diversity reward is determined to be (sv h , sv max The corresponding reward (e.g., 0.5).

[0162] Step e2: Combine the first diversity reward with the second diversity reward to obtain the final diversity reward.

[0163] Optionally, the first diversity reward and the second diversity reward can be merged by direct summation. Of course, this embodiment is not limited to direct summation; other methods can also be used, such as weighted summation.

[0164] It should be noted that, in another possible implementation, the second diversity reward can be used as the final diversity reward.

[0165] After obtaining the search performance reward, diversity reward, and semantic consistency reward, these rewards can be combined to obtain a comprehensive reward. Then, the query rewriting model can be optimized based on the comprehensive reward.

[0166] In one possible implementation, weights can be set for the search performance reward, diversity reward, and semantic consistency reward, respectively. Then, the search performance reward, diversity reward, and semantic consistency reward are weighted and summed according to the set weights to obtain the comprehensive reward.

[0167] It should be noted that this embodiment is not limited to using a weighted summation method to fuse the search performance reward, diversity reward, and semantic consistency reward. Other fusion methods can also be used, such as direct summation.

[0168] Optionally, when optimizing the query rewriting model based on the comprehensive reward, the GRPO algorithm (Group Relative Policy Optimization Algorithm, which is an innovative reinforcement learning algorithm) can be used to optimize the query rewriting model based on the comprehensive reward.

[0169] The query rewriting model training method provided in this application uses training data (query question samples and standard answers to query questions) from the training dataset. It trains the query rewriting model based on a reinforcement learning strategy. During training, multi-dimensional rewards (such as retrieval performance rewards, diversity rewards, and semantic consistency rewards) are determined based on the rewriting results of the query rewriting model for the query question samples. The query rewriting model is then optimized based on these multi-dimensional rewards. The query rewriting model training method provided in this application balances retrieval performance, semantic fidelity, and generation diversity, thereby improving the overall performance of query rewriting and effectively enhancing retrieval results.

[0170] Based on the query rewriting model training method provided in the above embodiments, this application also provides a retrieval method, which may include:

[0171] Step f1: Obtain the target query question.

[0172] Step f2: Rewrite the target query problem based on the query rewriting model to obtain the rewritten target query problem.

[0173] The query rewriting model was trained using the query rewriting model training method provided in the above embodiments.

[0174] Step f3: Perform knowledge retrieval on the rewritten target query question to obtain the retrieval results.

[0175] The retrieval method provided in this application, after obtaining the target query question, first rewrites it based on a query rewriting model, and then performs knowledge retrieval on the rewritten target query question. The retrieval method provided in this application has a high retrieval recall rate and a good user experience.

[0176] This application also provides a query rewriting model training device, such as... Figure 3 As shown, the query rewriting model training device may include: a query question rewriting unit 301, a retrieval unit 302, a retrieval effect reward determination unit 303a, a diversity reward determination unit 303b, and a model optimization unit 304.

[0177] The query question rewriting unit 301 is used to rewrite the query question sample based on the query rewriting model to obtain the rewritten result set of the query question sample.

[0178] The retrieval unit 302 is used to perform knowledge retrieval on the rewriting results in the rewriting result set to obtain the retrieval results of the rewriting results in the rewriting result set.

[0179] The retrieval performance reward determination unit 303a is used to determine the retrieval performance reward based on the retrieval results of the rewritten results in the rewritten result set.

[0180] The diversity reward determination unit 303b is used to determine the diversity reward based on the rewriting result set, wherein the diversity reward reflects the diversity of the rewriting results in the rewriting result set.

[0181] Model optimization unit 304 is used to optimize the query rewriting model based on retrieval performance rewards and diversity rewards.

[0182] In one possible implementation, the query rewriting model training apparatus provided in this application embodiment may further include: a semantic consistency reward determination unit 303c.

[0183] The semantic consistency reward determination unit 303c is used to determine the semantic consistency reward based on the query question sample and the rewritten result set, wherein the semantic consistency reward reflects the semantic consistency between the query question sample and the rewritten results in the rewritten result set.

[0184] In one possible implementation, the query rewriting model training apparatus provided in this application embodiment may further include: a noise-adding unit.

[0185] The noise-adding unit is used to add noise to the query question sample to obtain the noisy query question sample.

[0186] The query question rewriting unit 301 is also used to rewrite the noisy query question sample based on the query rewriting model, obtain the rewritten result of the noisy query question sample, and add the rewritten result of the noisy query question sample to the rewriting result set.

[0187] In one possible implementation, the process by which the retrieval performance reward determination unit 303a determines the retrieval performance reward based on the retrieval results of the rewritten results in the rewritten result set includes:

[0188] Based on the position of the standard answer to the query question sample in the search results of each rewritten result in the rewritten result set, determine the first search performance reward corresponding to each rewritten result in the rewritten result set;

[0189] The first search performance reward corresponding to each rewritten result in the rewritten result set is merged to obtain the first merged search performance reward, which is used as the final search performance reward.

[0190] In one possible implementation, the process by which the retrieval performance reward determination unit 303a determines the first retrieval performance reward for each rewritten result in the rewritten result set based on the position of the standard answer to the query question sample in the retrieval results of each rewritten result in the rewritten result set includes:

[0191] For each rewrite result in the rewrite result set:

[0192] The search results of the rewritten result are divided into categories according to the preset categorization method. Each category is assigned a corresponding weight. The higher the relevance of the search result to the query question, the higher the category it belongs to, and the higher the category, the greater the weight.

[0193] Determine the category of the standard answer to the query question sample;

[0194] The first search performance reward for the rewritten result is determined based on the weight corresponding to the tier of the standard answer to the query question sample.

[0195] In one possible implementation, the process of determining the retrieval performance reward by the retrieval performance reward unit 303a based on the retrieval results of the rewritten results in the rewritten result set further includes:

[0196] The retrieval results of each rewritten result in the rewritten result set are rearranged based on the reordering model in order to obtain the confidence of the retrieval results of each rewritten result in the rewritten result set.

[0197] The second retrieval performance reward corresponding to each rewritten result in the rewritten result set is determined based on the confidence level of the retrieval results for each rewritten result in the rewritten result set.

[0198] The second search performance reward corresponding to each rewritten result in the rewritten result set is merged to obtain the second merged search performance reward.

[0199] The first fusion search effect reward and the second fusion search effect reward are combined to obtain the third fusion search effect reward, which is used as the final search effect reward.

[0200] In one possible implementation, the process by which the retrieval unit 302 performs knowledge retrieval on the rewrite results in the rewrite result set to obtain the retrieval results of the rewrite results in the rewrite result set includes:

[0201] Based on multiple different retrieval models, knowledge retrieval is performed on each rewritten result in the rewritten result set to obtain the retrieval results for each rewritten result in the rewritten result set on each retrieval model.

[0202] In one possible implementation, the process by which the retrieval performance reward determination unit 303a merges the retrieval performance rewards corresponding to each rewritten result in the rewritten result set includes:

[0203] For each retrieval model, the retrieval performance rewards corresponding to each rewritten result in the rewritten result set on that retrieval model are merged to obtain the retrieval performance reward on that retrieval model.

[0204] The search performance rewards across different search models will be combined.

[0205] In one possible implementation, the semantic consistency reward determination unit 303c determines the semantic consistency reward based on the query question sample and the rewritten result set, including the following process:

[0206] Calculate the semantic similarity between the query question sample and each rewritten result in the rewritten result set;

[0207] Based on the semantic similarity between the query question sample and each rewritten result in the rewritten result set, determine the semantic consistency reward corresponding to each rewritten result in the rewritten result set;

[0208] The semantic consistency rewards corresponding to each rewrite result in the rewrite result set are merged to obtain the final semantic consistency reward.

[0209] In one possible implementation, the process by which the diversity reward determination unit 303b determines the diversity reward based on the rewritten result set includes:

[0210] Calculate the similarity between each pair of rewritten results in the rewritten result set to obtain several similarity scores;

[0211] Diversity rewards are determined based on several similarities.

[0212] In one possible implementation, the process by which the diversity reward determination unit 303b determines the diversity reward based on several similarities includes:

[0213] Obtain several similarity intervals and the corresponding rewards for each of the several similarity intervals. The several similarity intervals are obtained by dividing the similarity range based on a preset similarity, and the rewards for each of the several similarity intervals are different.

[0214] From a number of similarity intervals, determine the intervals in which each similarity level belongs;

[0215] The rewards corresponding to the intervals where several similarities are located are merged to obtain the first diversity reward, which serves as the final diversity reward.

[0216] In one possible implementation, the process by which the diversity reward determination unit 303b determines the diversity reward based on the aforementioned similarities further includes:

[0217] The variance of several similarities is calculated to obtain the similarity variance;

[0218] The second diversity reward is determined based on the similarity variance;

[0219] The first diversity reward and the second diversity reward are merged, and the merged diversity reward is used as the final diversity reward.

[0220] The modules in the aforementioned query rewriting model training device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware within or independently of the computer device's processing unit, or stored in software within the computer device's memory, so that the processor can call and execute the operations corresponding to each module.

[0221] The query rewriting model training apparatus provided in this application, after obtaining a query question sample, first rewrites the query question sample based on the query rewriting model. After obtaining the rewritten result set of the query question sample, it retrieves the rewritten results in the result set and determines the reward for the query rewriting model in the retrieval performance dimension based on the retrieval results. Simultaneously, it determines the reward for the query rewriting model in the diversity dimension based on the result set. Then, it optimizes the query rewriting model based on the rewards for retrieval performance and diversity, enabling the model to generate diverse rewritten results that meet retrieval requirements. Furthermore, a semantic consistency reward is introduced to ensure the model generates rewritten results that are semantically consistent with the original query question. In summary, the query rewriting model training apparatus provided in this application can train a query rewriting model with good rewriting performance.

[0222] This application also provides a retrieval device, which may include: a query question acquisition unit, a query question rewriting unit, and a retrieval unit.

[0223] The query question retrieval unit is used to retrieve the target query question.

[0224] The query rewriting unit is used to rewrite the target query question based on the query rewriting model to obtain the rewritten target query question.

[0225] The query rewriting model is obtained using the query rewriting model training device provided in the above embodiments.

[0226] The retrieval unit is used to perform knowledge retrieval on the rewritten target query question and obtain retrieval results.

[0227] The retrieval device provided in this application embodiment has a high retrieval recall rate and a good user experience.

[0228] This application embodiment also provides an electronic device, including at least one processor and a memory connected to the processor, wherein:

[0229] The memory is used to store computer programs;

[0230] The processor is used to execute the computer program so that the electronic device can implement the steps of the query rewriting model training method provided in the above embodiments, or implement the steps of the retrieval method provided in the above embodiments.

[0231] This disclosure also provides a computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the steps of the query rewriting model training method provided in the above embodiments, or implement the steps of the retrieval method provided in the above embodiments.

[0232] This application also provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device causes the electronic device to implement the steps of the query rewriting model training method provided in the above embodiments, or to implement the steps of the retrieval method provided in the above embodiments.

[0233] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0234] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0235] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0236] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A method for training a query rewriting model, characterized in that, include: Based on the query rewriting model, the query question sample is rewritten to obtain the rewritten result set of the query question sample; Knowledge retrieval is performed on the rewriting results in the rewriting result set to obtain the retrieval results of the rewriting results in the rewriting result set. Based on the retrieval results of the rewriting results in the rewriting result set, a retrieval performance reward is determined. Based on the rewriting result set, a diversity reward is determined, wherein the diversity reward reflects the diversity of the rewriting results in the rewriting result set; The query rewriting model is optimized based on the retrieval performance reward and the diversity reward.

2. The query rewriting model training method according to claim 1, characterized in that, Also includes: Based on the query question sample and the rewritten result set, a semantic consistency reward is determined, wherein the semantic consistency reward reflects the semantic consistency between the query question sample and the rewritten results in the rewritten result set.

3. The query rewriting model training method according to claim 1 or 2, characterized in that, Also includes: Noise is added to the query question sample to obtain the noisy query question sample; Based on the query rewriting model, the noisy query question sample is rewritten to obtain the rewritten result of the noisy query question sample; The rewritten results of the noisy query problem samples are added to the rewritten result set.

4. The query rewriting model training method according to claim 1, characterized in that, The step of determining the search performance reward based on the search results of the rewritten results in the rewritten result set includes: Based on the position of the standard answer to the query question sample in the search results of each rewritten result in the rewritten result set, determine the first search effect reward corresponding to each rewritten result in the rewritten result set; The first search effect rewards corresponding to each rewritten result in the rewritten result set are merged to obtain the first merged search effect reward, which is used as the final search effect reward.

5. The query rewriting model training method according to claim 4, characterized in that, The step of determining the first search performance reward corresponding to each rewritten result in the rewritten result set based on the position of the standard answer of the query question sample in the search results of each rewritten result in the rewritten result set includes: For each rewrite result in the rewrite result set: The search results of the rewritten result are divided into categories according to the preset categorization method. Each category is assigned a corresponding weight. The higher the relevance of the search result to the query question, the higher the category it belongs to, and the higher the category, the greater the weight. Determine the category of the standard answer to the query question sample; The first search performance reward corresponding to the rewritten result is determined based on the weight corresponding to the tier of the standard answer to the query question sample.

6. The query rewriting model training method according to claim 4, characterized in that, The step of determining the search performance reward based on the search results of the rewritten results in the rewritten result set also includes: The retrieval results of each rewritten result in the rewritten result set are rearranged based on the reordering model to obtain the confidence level of the retrieval results of each rewritten result in the rewritten result set. Based on the confidence level of the retrieval results for each rewritten result in the rewritten result set, determine the second retrieval performance reward corresponding to each rewritten result in the rewritten result set; The second search effect reward corresponding to each rewrite result in the rewrite result set is merged to obtain the second merged search effect reward. The first fused search effect reward and the second fused search effect reward are combined to obtain the third fused search effect reward, which is used as the final search effect reward.

7. The query rewriting model training method according to any one of claims 4 to 6, characterized in that, The step of performing knowledge retrieval on the rewriting results in the rewriting result set to obtain retrieval results for the rewriting results in the rewriting result set includes: Based on multiple different retrieval models, knowledge retrieval is performed on each rewritten result in the rewritten result set to obtain the retrieval results of each rewritten result in the rewritten result set on each retrieval model. The search performance rewards corresponding to each rewritten result in the rewritten result set are merged, including: For each retrieval model, the retrieval performance rewards corresponding to each rewritten result in the rewritten result set on that retrieval model are merged to obtain the retrieval performance reward on that retrieval model. The search performance rewards across different search models will be combined.

8. The query rewriting model training method according to claim 2, characterized in that, The step of determining the semantic consistency reward based on the query question sample and the rewritten result set includes: Calculate the semantic similarity between the query question sample and each rewritten result in the rewritten result set; Based on the semantic similarity between the query question sample and each rewritten result in the rewritten result set, the semantic consistency reward corresponding to each rewritten result in the rewritten result set is determined. The semantic consistency rewards corresponding to each rewrite result in the rewrite result set are merged to obtain the final semantic consistency reward.

9. The query rewriting model training method according to any one of claims 1 to 3, characterized in that, The step of determining the diversity reward based on the rewritten result set includes: The similarity between each pair of rewritten results in the rewritten result set is calculated to obtain several similarity scores; Based on the aforementioned similarities, diversity rewards are determined.

10. The query rewriting model training method according to claim 9, characterized in that, The determination of diversity rewards based on the aforementioned similarities includes: Obtain several similarity intervals and the corresponding rewards for each of the several similarity intervals, wherein the several similarity intervals are obtained by dividing the similarity range based on a preset similarity, and the rewards corresponding to each of the several similarity intervals are different; From the plurality of similarity intervals, determine the intervals in which the plurality of similarities respectively lie; The rewards corresponding to the intervals where the similarities are located are merged to obtain the first diversity reward, which serves as the final diversity reward.

11. The query rewriting model training method according to claim 10, characterized in that, The determination of diversity rewards based on the aforementioned similarities also includes: The variance of the aforementioned similarities is calculated to obtain the similarity variance; The second diversity reward is determined based on the similarity variance. The first diversity reward and the second diversity reward are merged, and the merged diversity reward is used as the final diversity reward.

12. A retrieval method, characterized in that, include: To obtain the target query question; The target query problem is rewritten based on the query rewriting model to obtain the rewritten target query problem, wherein the query rewriting model is trained using the query rewriting model training method as described in any one of claims 1 to 11; A knowledge retrieval was performed on the rewritten target query question to obtain the retrieval results.

13. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is configured to execute the computer program to enable the electronic device to implement the steps of the query rewriting model training method as described in any one of claims 1 to 11, or to implement the steps of the retrieval method as described in claim 12.

14. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the steps of the query rewriting model training method as described in any one of claims 1 to 11, or the steps of the retrieval method as described in claim 12.