Query rewriting method and device based on multi-aspect feedback, electronic equipment and medium
By training the neural network model and using the sample set to optimize the neural network parameters, the problem of query data being significantly different from user intent after modification by the rule engine is solved, achieving more efficient query data rewriting and reducing the waste of computing resources and repeated modifications.
Patent Information
- Application Number
- CN202510806318.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
In the prior art, the query data input by users is complex and changeable, and after modification by the rule engine, it is significantly different from the user's actual intention, resulting in a waste of computing resources and repeated modifications.
By obtaining a sample set, training a neural network model, using sample conversation history data and current query data, generating reward loss values, optimizing neural network parameters, and generating a query rewriting model, the accuracy of query data rewriting is improved.
It reduces the waste of computing resources, improves the accuracy of query data rewriting, makes the rewritten query data closer to the user's true intention, and reduces the need for multiple rewrites.
Smart Images

Figure CN120705287A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a query rewriting method, apparatus, electronic device, and medium based on multi-faceted feedback. Background Art
[0002] Query rewriting methods can rewrite query data entered by users into a large language model to make it easier for the large language model to capture the user's true intent. Currently, when rewriting user-entered query data, the common method is to use a pre-defined rule engine to complete or correct the user-entered query data.
[0003] However, when rewriting the query data input by the user in the above manner, the following technical problems often occur: The query data input by users is complex and changeable. The rule engine only modifies the query data according to pre-set rules, which can easily lead to the modified query data being significantly different from the user's true intentions, and in turn, the answers generated based on the modified query data being significantly different from the user's actual needs. This can easily lead to the need to modify the same query data multiple times to meet the user's actual needs, and can easily lead to a waste of computing resources during repeated modifications. Summary of the Invention
[0004] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] Some embodiments of the present disclosure provide a query rewriting method, apparatus, electronic device, and computer-readable medium based on multi-faceted feedback to solve one or more of the technical problems mentioned in the above background technology section.
[0006] In a first aspect, some embodiments of the present disclosure provide a query rewriting method based on multi-faceted feedback, the method comprising: obtaining a sample set, wherein each sample in the sample set comprises a sample conversation history data sequence, sample current query data and sample rewriting query data, and each sample conversation history data in the sample conversation history data sequence comprises sample history query data and sample history answer data; based on the sample set, performing the following training steps: updating the number of iterations based on a preset value; inputting the sample conversation history data sequence and sample current query data included in each sample of at least one sample in the sample set into an initial neural network, obtaining the sample rewriting query data and pre-trained rewriting query data corresponding to each sample of the at least one sample; and initializing the initial neural network based on the sample rewriting query data included in each sample of the at least one sample and the pre-trained rewriting query data corresponding to each sample of the at least one sample. Initialization to obtain an initialized neural network; generating reward loss value data corresponding to each sample in the at least one sample based on the sample rewriting query data included in each sample in the at least one sample, the pre-trained rewriting query data corresponding to each sample in the at least one sample, and the sampling rewriting query data; generating a target neural network based on the initialized neural network, the sample rewriting query data included in each sample in the at least one sample, the pre-trained rewriting query data corresponding to each sample in the at least one sample, and the reward loss value data; in response to determining that the updated number of iterations meets a preset iteration condition, determining the target neural network as a query rewriting model; in response to receiving current query data sent by a user terminal, obtaining a conversation history data sequence corresponding to the user terminal; generating rewriting answer data based on the conversation history data sequence, the current query data, and the query rewriting model; and sending the rewriting answer data to the user terminal.
[0007] In a second aspect, some embodiments of the present disclosure provide a query rewriting device based on multi-faceted feedback, the device comprising: a first acquisition unit, configured to acquire a sample set, wherein each sample in the sample set comprises a sample conversation history data sequence, sample current query data and sample rewriting query data, and each sample conversation history data in the sample conversation history data sequence comprises sample history query data and sample history answer data; an execution unit, configured to perform the following training steps based on the sample set: updating the number of iterations based on a preset value; inputting the sample conversation history data sequence and sample current query data included in each sample of at least one sample in the sample set into an initial neural network, obtaining the sample rewriting query data and pre-trained rewriting query data corresponding to each sample of the at least one sample; initializing the initial neural network based on the sample rewriting query data included in each sample of the at least one sample and the pre-trained rewriting query data corresponding to each sample of the at least one sample, and obtaining Initialize the neural network; generate reward loss value data corresponding to each sample in the at least one sample based on the sample rewrite query data included in each sample in the at least one sample, the pre-trained rewrite query data corresponding to each sample in the at least one sample, and the sampling rewrite query data; generate a target neural network based on the initialized neural network, the sample rewrite query data included in each sample in the at least one sample, the pre-trained rewrite query data corresponding to each sample in the at least one sample, and the reward loss value data; in response to determining that the updated number of iterations meets the preset iteration condition, determine the target neural network as a query rewriting model; a second acquisition unit is configured to acquire a conversation history data sequence corresponding to the above-mentioned user terminal in response to receiving the current query data sent by the user terminal; a generation unit is configured to generate rewrite answer data based on the above-mentioned conversation history data sequence, the above-mentioned current query data, and the above-mentioned query rewriting model; a sending unit is configured to send the above-mentioned rewrite answer data to the above-mentioned user terminal.
[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation of the first aspect is implemented.
[0010] The above-described various embodiments of the present disclosure have the following beneficial effects: The query rewriting method based on multi-faceted feedback in some embodiments of the present disclosure can reduce computing resource waste. Specifically, the reason for this computing resource waste is that the query data input by the user is complex and changeable. The rule engine only modifies the query data according to pre-set rules, which can easily lead to the modified query data being significantly different from the user's true intent. In turn, the answer generated based on the modified query data is significantly different from the user's actual needs. This can easily lead to the same query data being modified multiple times to meet the user's actual needs, which can easily lead to computing resource waste due to repeated modifications. Based on this, the query rewriting method based on multi-faceted feedback in some embodiments of the present disclosure first obtains a sample set, wherein each sample in the sample set includes a sample conversation history data sequence, a sample current query data, and a sample rewritten query data, and each sample conversation history data in the sample conversation history data sequence includes sample historical query data and sample historical answer data. Thus, a sample set is obtained. Next, based on the sample set, the following training steps are performed: First, the number of iterations is updated based on a preset value. Thus, the number of iterations during training can be determined. Next, the sample conversation history data sequence and the sample current query data included in each of at least one sample in the sample set are input into an initial neural network to obtain sampled rewritten query data and pre-trained rewritten query data corresponding to each of the at least one sample. Thus, by inputting both the sample conversation history data sequence and the sample current query data into the initial neural network, the initial neural network's ability to reference contextual semantics can be trained, improving the accuracy of query data rewriting by the trained initial neural network, and ensuring that the rewritten query data is more closely aligned with the user's true intent. Then, based on the sample rewritten query data included in each of the at least one sample and the pre-trained rewritten query data corresponding to each of the at least one sample, the initial neural network is initialized to obtain an initialized neural network. Thus, network parameters of the initial neural network can be adjusted based on the pre-trained rewritten query data output by the initial neural network to obtain an initialized neural network. Then, based on the sample rewritten query data included in each of the at least one sample, the pre-trained rewritten query data corresponding to each of the at least one sample, and the sampled rewritten query data, reward loss value data corresponding to each of the at least one sample is generated. Thus, based on the pre-trained rewritten query data, the sampled rewritten query data, and the sample rewritten query data, a loss value can be obtained between the query data output by the initial neural network and the actual query data to be obtained. Then, based on the initialized neural network, the sample rewritten query data included in each of the at least one sample, the pre-trained rewritten query data corresponding to each of the at least one sample, and the reward loss value data, a target neural network is generated. Thus, the target neural network can be obtained.Then, in response to determining that the updated number of iterations satisfies a preset iteration condition, the target neural network is determined as a query rewriting model. Thus, a query rewriting model can be obtained. Then, in response to receiving the current query data sent by the user terminal, a sequence of conversation history data corresponding to the user terminal is obtained. Thus, the original data required for rewriting can be obtained. Then, based on the sequence of conversation history data, the current query data, and the query rewriting model, rewritten answer data is generated. Thus, answer data corresponding to the rewritten query data can be obtained. Finally, the rewritten answer data is sent to the user terminal. Thus, the rewritten answer data can be sent to the user terminal. Because the network parameters of the initial neural network can be continuously optimized using the reward loss value data and the initialized neural network, the accuracy of the query rewriting model obtained after training when rewriting query data can be improved, making the rewritten query data more closely aligned with the user's true intentions. Also, when rewriting the query data input by the user, the context semantics corresponding to the query data can be referred to first, and then the query data can be rewritten. Therefore, the accuracy of the rewritten query data can be improved, and the rewritten query data can be closer to the user's true intention, thereby reducing the probability of having to rewrite the same query data multiple times to meet user needs, thereby reducing the waste of computing resources during multiple rewrites. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0012] Figure 1 is a flowchart of some embodiments of a query rewriting method based on multi-faceted feedback according to the present disclosure; Figure 2 is a schematic structural diagram of some embodiments of a query rewriting apparatus based on multi-aspect feedback according to the present disclosure; Figure 3 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0014] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0015] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0016] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0017] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0018] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0019] Figure 1 The flowchart 100 of some embodiments of the query rewriting method based on multi-faceted feedback according to the present disclosure is shown. The query rewriting method based on multi-faceted feedback includes the following steps: Step 101: Obtain a sample set.
[0020] In some embodiments, an executing entity (e.g., a computing device) of a query rewriting method based on multi-faceted feedback may obtain a sample set. Each sample in the sample set includes a sample conversation history data sequence, sample current query data, and sample rewritten query data. The sample conversation history data sequence may be a conversation history data sequence used for model training. Each conversation history data in the conversation history data sequence may be text data from a user's interaction with a large language model. Each sample conversation history data in the sample conversation history data sequence may include sample historical query data and sample historical answer data. The sample historical query data may be historical query data used for model training. The historical query data may be text data previously input by a user into the large language model. The sample historical answer data may be historical answer data used for model training. The historical answer data may be responses output by the large language model based on the text data input by the user. The sample current query data may be current query data used for model training. The current query data may be text data currently input by the user into the large language model. The sample current query data corresponds to a sample target document and sample answer data. The sample target document may be the document most relevant to the sample current query data after searching the sample current query data. The sample answer data may be a preset standard answer to the sample current query data. The sample rewritten query data may be standard rewritten query data corresponding to the sample current query data. The rewritten query data may be text data obtained by rewriting the current query data to better express the user's intention. For example, when the sample current query data input by the user is "How is the weather today?" and the user's actual need is to ask about the weather, the corresponding sample rewritten query data may be "How is the weather today?" The sample rewritten query data may include a sample byte data sequence. Each sample byte data in the sample byte data sequence may be a word in the sample rewritten query data. Each sample byte data in the sample byte data sequence has corresponding sample byte distribution data. The sample byte distribution data may be a one-hot vector corresponding to the sample byte data. The sample byte distribution data may include individual element values. For example, when the sample rewriting query data is "What's the weather like today?", "today" can be a sample byte data in the sample rewriting query data, and the sample byte distribution data corresponding to "today" can be "[1,0,0]", and "1", "0", and "0" can be the individual element values included in the sample byte distribution data. The aforementioned large language model may include, but is not limited to: GPT-4, DeepSeek. The aforementioned execution entity may be a server. In practice, the aforementioned execution entity may obtain a sample set from a preset sample database. The aforementioned sample database may be a database for storing sample sets.
[0021] Step 102: Based on the sample set, perform the following training steps: Step 1021: Update the number of iterations based on a preset value.
[0022] In some embodiments, the execution entity may update the number of iterations based on a preset value. The preset value may be a pre-set value. The specific setting of the preset value is not limited. The number of iterations may be a value representing the number of times the training step is executed. In practice, the execution entity may determine the sum of the number of iterations and the preset value as the number of iterations to update the number of iterations.
[0023] Step 1022: Input the sample conversation history data sequence and sample current query data included in each sample of at least one sample in the sample set into the initial neural network to obtain the sample rewriting query data and pre-training rewriting query data corresponding to each sample of at least one sample.
[0024] In some embodiments, the execution entity may input the sample conversation history data sequence and sample current query data included in each sample of at least one sample in the sample set into the initial neural network to obtain the sample rewriting query data and pre-training rewriting query data corresponding to each sample of the at least one sample.
[0025] In some optional implementations of some embodiments, the execution entity may input the sample conversation history data sequence and sample current query data included in each of at least one sample in the sample set into the initial neural network through the following steps to obtain the sample rewritten query data and pre-trained rewritten query data corresponding to each of the at least one sample: The first step is to concatenate the sample conversation history data sequence and the sample current query data included in each of at least one sample in the sample set to obtain a concatenated data sequence corresponding to each of the at least one sample. The concatenated data sequence may be a data sequence concatenated from the sample conversation history data sequence and the sample current query data. In practice, first, for each of the at least one sample and each sample conversation history data in the sample conversation history data sequence included in the sample, the execution entity may concatenate a preset delimiter character, the sample historical query data, and the sample historical answer data included in the sample conversation history data using a concatenation operator to obtain a concatenated data subsequence corresponding to the sample conversation history data. The delimiter character may be a character used to delimit character strings. For example, the delimiter character may be [SEP]. As an example, if the sample conversation history data is "{sample historical query data, sample historical answer data}", the concatenated data subsequence is "{sample historical answer data [SEP] sample historical query data}". Then, the execution entity may concatenate the preset start character, the obtained concatenated data subsequences, and the separator character using the concatenation operator to obtain a concatenated data sequence corresponding to the sample. The start character may be used to indicate the beginning of a string. For example, the start character may be [CLS]. The concatenation operator may be a plus operator.
[0026] As an example, when the above-mentioned spliced data subsequences are respectively "{sample historical answer data 1[SEP]sample historical query data 1}" and "{sample historical answer data 2[SEP]sample historical query data 2}", the corresponding spliced data sequence can be "{[CLS]sample historical answer data 2[SEP]sample historical query data 2[SEP]sample historical answer data 1[SEP]sample historical query data 1}".
[0027] In the second step, the concatenated data sequence corresponding to each sample in the at least one sample is input into the initial neural network to obtain sampled rewritten query data and pretrained rewritten query data corresponding to each sample in the at least one sample. The pretrained rewritten query data may be the rewritten query data output by the initial neural network after the concatenated data sequence is input into the initial neural network. The pretrained rewritten query data includes a pretrained rewritten query byte data sequence. Each pretrained rewritten query byte data in the pretrained rewritten query byte data sequence may be a word in the pretrained rewritten query data. Each pretrained rewritten query byte data in the pretrained rewritten query byte data sequence has corresponding byte distribution data. The byte distribution data may be a probability distribution corresponding to the pretrained rewritten query byte data. The sampled rewritten query data may be the text data randomly output by the initial neural network after the concatenated data sequence is input into the initial neural network. The initial neural network may be a neural network that takes the concatenated data sequence as input and outputs the rewritten query data corresponding to the concatenated data sequence. When performing the training step for the first time, the initial neural network used may be a preset initial neural network. When the training step is performed a second time, the initial neural network used may be the target neural network obtained in the first training step. When the training step is performed a third time, the initial neural network used may be the target neural network obtained in the second training step. Similarly, each subsequent training step uses the target neural network obtained in the previous initial neural network until a query rewriting model is obtained.
[0028] The initial neural network may include four layers.
[0029] The first layer may be a word segmentation processing layer. The word segmentation processing layer may be a word segmenter that takes the concatenated data sequence as input and outputs a word vector corresponding to the concatenated data sequence. The word segmentation processing layer may be a SentencePiece.
[0030] The second layer may be an encoding layer. The encoding layer may be a neural network that takes word vectors as input and outputs context vectors corresponding to the word vectors. The encoding layer may be a Transformer.
[0031] The third layer may be a decoding layer. The decoding layer may be a neural network that takes a context vector as input and outputs a target word vector corresponding to the context vector. The decoding layer may be a Transformer. The target word vector may be the word vector corresponding to the target text to be output.
[0032] The fourth layer can be an output layer. The above output layer can include a mapping layer, a normalization function, and a search algorithm. The above mapping layer can be a linear layer that takes the target word vector as input and outputs the mapped word vector. Among them, the above mapped word vector can be the word vector obtained by mapping the target word vector to the size of the vocabulary. The above normalization function can be a softmax function that takes the mapped word vector as input and outputs the word vector probability distribution data sequence corresponding to the mapped word vector. Among them, each word vector probability distribution data in the above word vector probability distribution data sequence can be the probability distribution of the corresponding word in the mapped word vector. Each word vector probability distribution data in the above word vector probability distribution data sequence can include a word and a probability. For example, the above word vector probability distribution data can be "{word: today, probability: 90%; word: tomorrow, probability: 7%; word: day, probability: 3%}". The above search algorithm can be an algorithm that takes the word vector probability distribution data sequence as input and outputs the rewritten query data. Among them, the above search algorithm can be a greedy algorithm. The rewritten query data output by the above search algorithm can be the text data composed of the words with the largest corresponding probabilities in each word vector probability distribution data of the word vector probability distribution data sequence. The above rewritten query data includes a word sequence. Each word in the above word sequence corresponds to word distribution data. The above word distribution data can be the probability distribution corresponding to the word. The above word distribution data can include each probability value.
[0033] As an example, when a word in the word sequence is "today", and the word vector probability distribution data where "today" is located is "{word: today, probability: 90%; word: tomorrow, probability: 7%; word: day, probability: 3%}", then the word distribution data corresponding to the above word "today" is [0.9, 0.07, 0.03], and "0.9", "0.07", and "0.03" are each probability value included in the word distribution data.
[0034] In practice, for each of the above at least one sample, first, the above execution entity can input the concatenated data sequence corresponding to the sample into the initial neural network to obtain the rewritten query data corresponding to the sample. Then, the word sequence included in the above rewritten query data can be determined as the pre-trained rewritten query byte data sequence. Then, the respective word distribution data corresponding to the above word sequence can be determined as the respective byte distribution data corresponding to the above pre-trained rewritten query byte data sequence. Finally, the above pre-trained rewritten query byte data sequence and the above respective byte distribution data can be combined into pre-trained rewritten query data.
[0035] In practice, for each of the at least one sample, first, the concatenated data sequence corresponding to the sample can be input into the initial neural network to obtain the word vector probability distribution data sequence output by the normalization function in the fourth layer of the initial neural network. Then, for each word vector probability distribution data in the word vector probability distribution data sequence, the execution entity can randomly select a word from the word vector probability distribution data as a sampling word by random sampling. Then, the selected sampling words can be combined into a sampling word sequence as the sampled rewriting query data in the order of selection.
[0036] Step 1023: Initialize the initial neural network based on the sample rewriting query data included in each sample in the at least one sample and the pre-trained rewriting query data corresponding to each sample in the at least one sample to obtain an initialized neural network.
[0037] In some embodiments, the execution entity may initialize the initial neural network based on the sample rewriting query data included in each sample in the at least one sample and the pre-trained rewriting query data corresponding to each sample in the at least one sample to obtain an initialized neural network.
[0038] In some optional implementations of some embodiments, the execution entity may initialize the initial neural network based on the sample rewriting query data included in each sample in the at least one sample and the pre-trained rewriting query data corresponding to each sample in the at least one sample through the following steps to obtain the initialized neural network: In the first step, for each of the at least one sample, perform the following steps: In a first sub-step, based on the pre-trained rewriting query data corresponding to the sample, each byte distribution data corresponding to the pre-trained rewriting query data is determined as each pre-trained byte distribution data.
[0039] The second sub-step is to perform the following steps for each sample byte distribution data in each sample byte distribution data corresponding to the above sample: Sub-step 1: Determine the pre-trained byte distribution data corresponding to the sample byte distribution data among the respective pre-trained byte distribution data as the target pre-trained byte distribution data. For example, when the sample byte distribution data is the sample byte distribution data corresponding to the first sample byte data in the sample byte data sequence, the pre-trained byte distribution data corresponding to the sample byte distribution data is the pre-trained byte distribution data corresponding to the first pre-trained rewriting query byte data in the pre-training rewriting query byte data sequence.
[0040] Sub-step 2: Determine target logarithmic distribution data based on the target pre-trained byte distribution data. In practice, for each probability value included in the target pre-trained byte distribution data, the execution entity may determine the probability value as the target probability value. Then, the logarithm of the target probability value may be determined as the target logarithmic distribution data. The logarithm may be a base-2 logarithm.
[0041] Sub-step three: for each target logarithmic distribution data in the above target logarithmic distribution data, perform the following steps: First, based on the above-mentioned sample byte distribution data, the target sample element data corresponding to the above-mentioned target logarithmic distribution data is determined. The above-mentioned target sample element data can be the element value corresponding to the above-mentioned target logarithmic distribution data among the various element values included in the above-mentioned sample byte distribution data. As an example, when the above-mentioned target pre-trained byte distribution data is [0.1, 0.85, 0.05], the above-mentioned sample byte distribution data is [1, 0, 0], and the target logarithmic distribution data is the logarithm of 0.85, which ranks second in the above-mentioned target pre-trained byte distribution data [0.1, 0.85, 0.05], then the corresponding target sample element data is the second-ranked "0" in the above-mentioned [1, 0, 0].
[0042] Then, the product of the target logarithmic distribution data and the target sample element data is determined as byte distribution product data.
[0043] Sub-step four: determining the sum of the determined byte distribution product data as the added byte distribution data.
[0044] Sub-step five: determining the negative number of the above-mentioned summed byte distribution data as the rewrite loss value data.
[0045] In the third sub-step, the sum of the determined rewriting loss value data is determined as the comprehensive rewriting loss data.
[0046] In the second step, the average value of the determined comprehensive rewrite loss data is determined as the average rewrite loss data.
[0047] The third step is to adjust the network parameters of the initial neural network based on the average rewrite loss data to initialize the initial neural network and obtain an initialized neural network. In practice, in response to determining that the average rewrite loss data meets a preset initialization condition, the execution entity may adjust the network parameters of the initial neural network using a gradient descent method. The initialization condition may be that the average rewrite loss data is greater than a preset initialization threshold. The specific setting of the initialization threshold is not limited here.
[0048] Step 1024: Generate reward loss value data corresponding to each sample in at least one sample based on the sample rewriting query data included in each sample in at least one sample, the pre-training rewriting query data corresponding to each sample in at least one sample, and the sampling rewriting query data.
[0049] In some embodiments, the execution entity may generate reward loss value data corresponding to each of the at least one sample based on the sample rewriting query data included in each of the at least one sample, the pre-trained rewriting query data corresponding to each of the at least one sample, and the sampling rewriting query data. The reward loss value data may be a loss value between the sample rewriting query data, the pre-trained rewriting query data, and the sampling rewriting query data.
[0050] In some optional implementations of some embodiments, the execution entity may generate reward loss value data corresponding to each sample in the at least one sample based on the sample rewriting query data included in each sample in the at least one sample, the pre-trained rewriting query data corresponding to each sample in the at least one sample, and the sampling rewriting query data through the following steps: For each of the at least one sample above, perform the following steps: The first step is to perform feature extraction processing on the sample rewriting query data included in the above sample, the pre-trained rewriting query data corresponding to the above sample, and the sampling rewriting query data to obtain sample rewriting dense feature information, pre-trained rewriting dense feature information, and sampling rewriting dense feature information. Among them, the above sample rewriting dense feature information can be a dense vector corresponding to the sample rewriting query data. The above pre-trained rewriting dense feature information can be a dense vector corresponding to the pre-trained rewriting query data. The above sampling rewriting dense feature information can be a dense vector corresponding to the sampling rewriting query data. In practice, the above execution entity can input the sample rewriting query data included in the above sample, the pre-trained rewriting query data corresponding to the above sample, and the sampling rewriting query data into a pre-trained dense vector generation model respectively, so as to perform feature extraction processing on the above sample rewriting query data, the pre-trained rewriting query data, and the sampling rewriting query data to obtain sample rewriting dense feature information, pre-trained rewriting dense feature information, and sampling rewriting dense feature information. Among them, the above dense vector generation model can be a neural network model that takes text data as input and outputs dense vectors corresponding to text data. The dense vector generation model can be a pre-trained BERT (Bidirectional Encoder Representations from Transformers) model. The pre-training can be a process of fine-tuning the BERT model using annotated text data and a cross-entropy loss function.
[0051] In the second step, feature extraction is performed on the sample target document corresponding to the sample to obtain document-intensive feature information. The document-intensive feature information may be a dense vector corresponding to the sample target document. In practice, the execution entity may input the sample target document into the dense vector generation model to perform feature extraction on the sample target document to obtain document-intensive feature information.
[0052] The third step is to determine the similarity between the sample rewriting dense feature information and the document dense feature information as the sample rewriting similarity. In practice, the execution entity may determine the cosine similarity between the sample rewriting dense feature information and the document dense feature information as the sample rewriting similarity.
[0053] In the fourth step, the similarity between the pre-trained rewritten dense feature information and the document dense feature information is determined as the pre-trained rewriting similarity. In practice, the execution entity may determine the cosine similarity between the pre-trained rewritten dense feature information and the document dense feature information as the pre-trained rewriting similarity.
[0054] In step 5, the similarity between the sample rewriting dense feature information and the document dense feature information is determined as the sample rewriting similarity. In practice, the execution entity may determine the cosine similarity between the sample rewriting dense feature information and the document dense feature information as the sample rewriting similarity.
[0055] In the sixth step, the sample rewriting similarity, the pre-trained rewriting similarity and the sampled rewriting similarity are combined into document loss value data.
[0056] In the seventh step, feature extraction is performed on the sample answer data corresponding to the sample to obtain dense feature information of the sample answer. The dense feature information of the sample answer may be a dense vector corresponding to the sample answer data. In practice, the execution entity may input the sample answer data into the dense vector generation model to perform feature extraction on the sample answer data to obtain dense feature information of the sample answer.
[0057] In the eighth step, based on the above-mentioned sample answer intensive feature information, the sample rewriting query data included in the above-mentioned sample, the pre-training rewriting query data and the sampling rewriting query data corresponding to the above-mentioned sample, weight loss value data is generated. Among them, the above-mentioned weight loss value data can characterize the similarity between the above-mentioned sample rewriting query data, the above-mentioned pre-training rewriting query data, the above-mentioned sampling rewriting query data and the above-mentioned sample answer intensive feature information. The above-mentioned weight loss value data can include sample weight intensive data, pre-training weight intensive data and sampling weight intensive data. The above-mentioned sample weight intensive data can be a numerical value used to characterize the similarity between the above-mentioned sample rewriting query data and the above-mentioned sample answer intensive feature information. The above-mentioned pre-training weight intensive data can be a numerical value used to characterize the similarity between the above-mentioned pre-training rewriting query data and the above-mentioned sample answer intensive feature information. The above-mentioned sampling weight intensive data can be a numerical value used to characterize the similarity between the above-mentioned sampling rewriting query data and the above-mentioned sample answer intensive feature information.
[0058] The ninth step is to generate sample rewritten answer data, pre-trained rewritten answer data and sample rewritten answer data based on the sample rewritten query data included in the above sample, the pre-trained rewritten query data corresponding to the above sample and the sample rewritten query data. The sample rewritten answer data can be the text data output by the above large language model after the sample rewritten query data is input into the above large language model. The pre-trained rewritten answer data can be the text data output by the above large language model after the pre-trained rewritten query data is input into the above large language model. The sample rewritten answer data can be the text data output by the above large language model after the sample rewritten query data is input into the above large language model. In practice, the execution entity can input the sample rewritten query data included in the above sample, the pre-trained rewritten query data corresponding to the above sample and the sample rewritten query data into the above large language model respectively to obtain sample rewritten answer data, pre-trained rewritten answer data and sample rewritten answer data.
[0059] In the tenth step, based on the sample rewritten answer data, the pre-trained rewritten answer data, the sampled rewritten answer data, and the sample answer data corresponding to the sample, generate sample rewritten answer score data, pre-trained rewritten answer score data, and sampled rewritten answer score data. The sample rewritten answer score data may be the ROUGE (Recall-Oriented Understudy for Gisting Evaluation) metric between the sample rewritten answer data and the sample answer data. The pre-trained rewritten answer score data may be the ROUGE metric between the pre-trained rewritten answer data and the sample answer data. The sampled rewritten answer score data may be the ROUGE metric between the sample rewritten answer data and the sample answer data. In practice, first, the execution entity may use a classification algorithm to perform a similarity comparison between the sample rewritten answer data and the sample answer data to obtain the sample rewritten answer score data. Second, the classification algorithm may be used to perform a similarity comparison between the pre-trained rewritten answer data and the sample answer data to obtain the pre-trained rewritten answer score data. Then, the sampled rewritten answer data and the sample answer data can be compared for similarity using the classification algorithm to obtain the sampled rewritten answer score data. The classification algorithm can be an algorithm capable of comparing the similarity of two texts. For example, the classification algorithm can be a Rocchio algorithm.
[0060] In the eleventh step, the sample rewriting answer score data, the pre-training rewriting answer score data and the sampling rewriting answer score data are combined into score loss value data.
[0061] In the twelfth step, the document loss value data, the weight loss value data and the score loss value data are combined into the reward loss value data corresponding to the sample.
[0062] In some optional implementations of some embodiments, the execution entity may generate weight loss value data based on the sample answer dense feature information, the sample rewritten query data included in the sample, the pre-trained rewritten query data corresponding to the sample, and the sampled rewritten query data through the following steps: The first step is to generate a sample rewritten document sequence, a pretrained rewritten document sequence, and a sample rewritten document sequence based on the sample rewritten query data included in the sample, the pretrained rewritten query data corresponding to the sample, and the sample rewritten query data. Each sample rewritten document in the sample rewritten document sequence may be a document related to the sample rewritten query data. For example, when the semantic similarity between the sample rewritten document and the sample rewritten query data is greater than a preset similarity threshold, it can be determined that the sample rewritten document is related to the sample rewritten query data. The similarity threshold may be a pre-set value. The specific setting of the similarity threshold is not limited here. Each pretrained rewritten document in the pretrained rewritten document sequence may be a document related to the pretrained rewritten query data. For example, when the semantic similarity between the pretrained rewritten document and the pretrained rewritten query data is greater than the similarity threshold, it can be determined that the pretrained rewritten document is related to the pretrained rewritten query data. Each sample rewritten document in the sample rewritten document sequence may be a document related to the sample rewritten query data. For example, when the semantic similarity between the sampled rewritten document and the sampled rewritten query data is greater than the aforementioned similarity threshold, it can be determined that the sampled rewritten document is related to the sampled rewritten query data.
[0063] In practice, the execution entity may input the sample rewritten query data, the pretrained rewritten query data, and the sampled rewritten query data into a dense searcher, respectively, to obtain a sample rewritten document sequence, a pretrained rewritten document sequence, and a sampled rewritten document sequence. The dense searcher may retrieve documents related to the text data based on the text data. For example, the dense searcher may be msmarco-roberta-base-ance-firstp.
[0064] In the second step, for each sample rewritten document in the above sample rewritten document sequence, perform the following steps: In a first sub-step, feature extraction is performed on the sample rewritten document to obtain dense feature information of the sample document. The dense feature information of the sample document may be a dense vector corresponding to the sample rewritten document. In practice, the execution entity may input the sample rewritten document into the dense vector generation model to perform feature extraction on the sample rewritten document to obtain dense feature information of the sample document.
[0065] In the second sub-step, the similarity between the sample answer dense feature information and the sample document dense feature information is determined as the sample document similarity. In practice, the execution entity may determine the sample document similarity as the cosine similarity between the sample answer dense feature information and the sample document dense feature information.
[0066] The third sub-step is to determine the sample ranking data corresponding to the sample rewritten document based on the sample rewritten document, wherein the sample ranking data may be the reciprocal of the position of the sample rewritten document in the sample rewritten document sequence.
[0067] In the fourth sub-step, the product of the sample ranking data and the sample document similarity is determined as the sample document weight data.
[0068] In the third step, the sum of the determined weight data of each sample document is determined as sample weight dense data.
[0069] The fourth step is to generate pre-trained weight-intensive data and sampled weight-intensive data based on the pre-trained rewritten document sequence, the sampled rewritten document sequence, and the sample answer-intensive feature information. In practice, the method for generating pre-trained weight-intensive data and sampled weight-intensive data can refer to the specific implementation method for generating the sample weight-intensive data, and will not be repeated here.
[0070] In the fifth step, the sample weight dense data, the pre-training weight dense data and the sampling weight dense data are combined into weight loss value data.
[0071] Step 1025: Generate a target neural network based on the initialized neural network, the sample rewrite query data included in each sample in the at least one sample, the pre-trained rewrite query data corresponding to each sample in the at least one sample, and the reward loss value data.
[0072] In some embodiments, the execution entity may generate a target neural network based on the initialized neural network, the sample rewrite query data included in each of the at least one sample, the pre-trained rewrite query data and the reward loss value data corresponding to each of the at least one sample.
[0073] In the process of adopting technical solutions to solve the above technical problems, the following problems often arise: When using a rule engine to rewrite query data, the rules in the rule engine must be traversed one by one until a matching rule is found. This results in low processing efficiency when rewriting query data. Furthermore, each time different query data is rewritten, computing resources must be used to match the query data with the rules in the rule engine one by one, resulting in a high consumption of computing resources when rewriting query data.
[0074] Faced with the above technical problems, we decided to adopt the following solutions: In some optional implementations of some embodiments, the execution entity may generate a target neural network based on the initialized neural network, the sample rewrite query data included in each of the at least one sample, the pre-trained rewrite query data corresponding to each of the at least one sample, and the reward loss value data through the following steps: In the first step, for each of the at least one sample, perform the following steps: In a first sub-step, the sample conversation history data sequence and the sample current query data included in the sample are concatenated to obtain sample concatenated data. The sample concatenated data may be text data obtained by concatenating the sample conversation history data sequence and the sample current query data included in the sample. In practice, the execution entity may use a concatenation function to concatenate the sample conversation history data sequence and the sample current query data included in the sample into the sample concatenated data. The concatenation function may be a function capable of concatenating two text data sets. For example, the concatenation function may be a concat function.
[0075] The second sub-step is to generate a document data pair group, a weight data pair group and a score data pair group based on the above-mentioned sample splicing data, the above-mentioned reward loss value data, the sample rewritten query data included in the above-mentioned sample, the pre-trained rewritten query data corresponding to the above-mentioned sample and the sampled rewritten query data. Among them, each document data pair in the above-mentioned document data pair group can be a data pair obtained by combining the above-mentioned sample splicing data, the above-mentioned sample rewritten query data, the above-mentioned pre-trained rewritten query data and the above-mentioned sampled rewritten query data according to the document loss value data included in the reward loss value data. Each document data pair in the above-mentioned document data pair group can include document selection data and document rejection data. The above-mentioned document selection data can represent the data that the query rewriting model should prioritize under the constraints of the above-mentioned document loss value data. The above-mentioned document matrix data can represent the data that the query rewriting model should secondarily select relative to the above-mentioned document selection data. The above-mentioned query rewriting model can be a trained initial neural network. The network structure of the above-mentioned query rewriting model can refer to the initial neural network and will not be repeated here.
[0076] Each weight data pair in the above-mentioned weight data pair group can be a data pair obtained by combining the above-mentioned sample splicing data, the above-mentioned sample rewrite query data, the above-mentioned pre-trained rewrite query data and the above-mentioned sample rewrite query data based on the weight loss value data included in the above-mentioned reward loss value data. Each weight data pair in the above-mentioned weight data pair group may include weight selection data and weight rejection data. The above-mentioned weight selection data can represent the data that the above-mentioned query rewrite model should preferentially select under the constraints of the above-mentioned weight loss value data. The above-mentioned weight rejection data can represent the data that the above-mentioned query rewrite model should secondarily select relative to the above-mentioned weight selection data.
[0077] Each score data pair in the above-mentioned score data pair group can be a data pair obtained by combining the above-mentioned sample splicing data, the above-mentioned sample rewritten query data, the above-mentioned pre-trained rewritten query data, and the above-mentioned sample rewritten query data based on the score loss value data included in the above-mentioned reward loss value data. Each score data pair in the above-mentioned score data pair group can include score selection data and score rejection data. The above-mentioned score selection data can represent the data that the above-mentioned query rewriting model should preferentially select under the constraints of the above-mentioned score loss value data. The above-mentioned score rejection data can represent the data that the above-mentioned query rewriting model should secondary select relative to the above-mentioned score selection data.
[0078] In practice, for the document loss value data included in the above-mentioned reward loss value data, and for the sample rewriting similarity, pre-training rewriting similarity and sampling rewriting similarity included in the above-mentioned document loss value data, first, the above-mentioned execution entity can arrange the above-mentioned sample rewriting similarity, the above-mentioned pre-training rewriting similarity and the above-mentioned sampling rewriting similarity in order from large to small, and obtain a sequence composed of the above-mentioned sample rewriting similarity, the above-mentioned pre-training rewriting similarity and the above-mentioned sampling rewriting similarity as a similarity sequence. Then, the above-mentioned sample rewriting query data, the above-mentioned pre-training rewriting query data and the above-mentioned sampling rewriting query data can be sorted according to the order in the above-mentioned similarity sequence to obtain a sequence as a document rewriting data sequence. Among them, the above-mentioned sample rewriting similarity corresponds to the above-mentioned sample rewriting query data. The above-mentioned pre-training rewriting similarity corresponds to the above-mentioned pre-training rewriting query data. The above-mentioned sampling rewriting similarity corresponds to the above-mentioned sampling rewriting query data. For example, when the above-mentioned similarity sequence is "{sample rewrite similarity, pre-trained rewrite similarity, sampling rewrite similarity}", the document rewrite data sequence can be "{sample rewrite query data, pre-trained rewrite query data, sampling rewrite query data}". Then, the document rewrite data ranked first in the above-mentioned document rewrite data sequence can be determined as the first document rewrite data. The document rewrite data ranked second in the above-mentioned document rewrite data sequence can be determined as the second document rewrite data. The document rewrite data ranked third in the above-mentioned document rewrite data sequence can be determined as the third document rewrite data. Then, the above-mentioned sample splicing data and the above-mentioned first document rewrite data can be combined into document selection data. The above-mentioned sample splicing data and the above-mentioned second document rewrite data are combined into document rejection data corresponding to the above-mentioned document selection data. Then, the above-mentioned document selection data and the above-mentioned document rejection data can be combined into a document data pair as the first document data pair. When combining document data pairs, ensure that the document rewrite data in the document selection data is ranked before the document rewrite data in the document rejection data in the above-mentioned document rewrite data sequence. Similarly, using the first document rewrite data, the second document rewrite data, and the third document rewrite data, a second document data pair and a third document data pair can be generated. The method for generating the second document data pair and the third document data pair can refer to the specific implementation method for generating the first document data pair and will not be repeated here. The first document data pair, the second document data pair, and the third document data pair can then be combined into a document data pair group.
[0079] In practice, for the weight loss value data included in the above-mentioned reward loss value data, and for the sample weight dense data, pre-trained weight dense data, and sampled weight dense data included in the above-mentioned weight loss value data, first, the above-mentioned sample weight dense data, the above-mentioned pre-trained weight dense data, and the above-mentioned sampled weight dense data can be sorted in order from large to small, and a sequence consisting of the above-mentioned sample weight dense data, the above-mentioned pre-trained weight dense data, and the above-mentioned sampled weight dense data is obtained as a dense data sequence. Then, according to the order in the above-mentioned dense data sequence, the above-mentioned sample rewrite query data, the above-mentioned pre-trained rewrite query data, and the above-mentioned sampled rewrite query data are sorted to obtain a sequence as a weight rewrite data sequence. Among them, the above-mentioned sample weight dense data can correspond to the above-mentioned sample rewrite query data. The above-mentioned pre-trained weight dense data can correspond to the above-mentioned pre-trained rewrite query data. The above-mentioned sample weight dense data can correspond to the above-mentioned sample rewrite query data. Then, the weight rewrite data ranked first in the above-mentioned weight rewrite data sequence can be determined as the first weight rewrite data. The weight rewrite data ranked second in the above-mentioned weight rewrite data sequence can be determined as the second weight rewrite data. The weight rewrite data ranked third in the above-mentioned weight rewrite data sequence can be determined as the third weight rewrite data. Then, the above-mentioned sample splicing data and the above-mentioned first weight rewrite data can be combined into weight selection data. The above-mentioned sample splicing data and the above-mentioned second weight rewrite data are combined into weight rejection data corresponding to the above-mentioned weight selection data. Then, the above-mentioned weight selection data and the above-mentioned weight rejection data can be combined into a weight data pair as the first weight data pair. When combining the weight data pairs, ensure that the ranking of the weight rewrite data in the weight selection data in the above-mentioned weight rewrite data sequence is before the weight rewrite data in the weight rejection data. By analogy, the second weight data pair and the third weight data pair can be generated through the above-mentioned first weight rewrite data, the above-mentioned second weight rewrite data and the above-mentioned third weight rewrite data. The method for generating the above-mentioned second weight data pair and the above-mentioned third weight data pair can refer to the specific implementation method for generating the above-mentioned first weight data pair, which will not be repeated here. Then, the above-mentioned first weight data pair, the above-mentioned second weight data pair and the above-mentioned third weight data pair can be combined into a weight data pair group.
[0080] In practice, for the score loss value data included in the reward loss value data, and for the sample rewritten answer score data, pre-trained rewritten answer score data, and sample rewritten answer score data included in the score loss value data, first, the execution entity may sort the sample rewritten answer score data, pre-trained rewritten answer score data, and sample rewritten answer score data in descending order, obtaining a sequence consisting of the sample rewritten answer score data, pre-trained rewritten answer score data, and sample rewritten answer score data as a score data sequence. Then, the sample rewritten query data, pre-trained rewritten query data, and sample rewritten query data may be sorted according to the order of precedence in the score data sequence, obtaining a sequence as a score rewritten data sequence. The sample rewritten answer score data corresponds to the sample rewritten query data. The pre-trained rewritten answer score data corresponds to the pre-trained rewritten query data. The sample rewritten answer score data corresponds to the sample rewritten query data. Then, the score rewritten data ranked first in the score rewritten data sequence may be determined as the first score rewritten data. The fractional rewrite data ranked second in the fractional rewrite data sequence can be determined as the second fractional rewrite data. The fractional rewrite data ranked third in the fractional rewrite data sequence can be determined as the third fractional rewrite data. Then, the sample splicing data and the first fractional rewrite data can be combined into fractional selection data. The sample splicing data and the second fractional rewrite data can be combined into fractional rejection data corresponding to the fractional selection data. Then, the fractional selection data and the fractional rejection data can be combined into a fractional data pair as the first fractional data pair. When combining the fractional data pairs, it is ensured that the fractional rewrite data in the fractional selection data is ranked before the fractional rewrite data in the fractional rejection data in the fractional rewrite data sequence. Similarly, a second fractional data pair and a third fractional data pair can be generated using the first fractional rewrite data, the second fractional rewrite data, and the third fractional rewrite data. The method for generating the second fractional data pair and the third fractional data pair can refer to the specific implementation method for generating the first fractional data pair, which will not be repeated here. Then, the first score data pair, the second score data pair, and the third score data pair may be combined into a score data pair group.
[0081] In the third sub-step, for each document data pair in the above document data pair group, perform the following steps: Sub-step one: Generate document selection score data and document rejection score data based on the document selection data and document rejection data included in the above-mentioned document data pair. The above-mentioned document selection score data may be a numerical value used to characterize the quality of the above-mentioned document selection data. The above-mentioned document rejection score data may be a numerical value used to characterize the quality of the above-mentioned document rejection data. In practice, first, the above-mentioned execution entity may input the above-mentioned document selection score data into a preset data reward model to obtain document selection score data. Then, the above-mentioned document rejection data may be input into the above-mentioned data reward model to obtain document rejection score data. The above-mentioned data reward model may be a neural network model that takes text data as input and outputs a numerical value used to characterize the data quality corresponding to the text data. The above-mentioned data reward model may be a pre-trained T5 model (Text-to-Text Transfer Transformer). The above-mentioned pre-training may be a process of fine-tuning the T5 model using an annotated text dataset and a Bradley-Terry model as a loss function.
[0082] Sub-step 2: determining the difference between the document selection score data and the document rejection score data as document difference data.
[0083] Sub-step three: Generate document loss data based on the document difference data. The document loss data may be document difference data mapped to the range [0, 1]. In practice, the execution entity may input the document difference data into a nonlinear activation function to obtain the document loss data. The nonlinear activation function may be a Sigmoid function.
[0084] Sub-step 4: determining the logarithm of the document loss data as the document logarithmic loss data, wherein the logarithm may be a logarithm with base 2.
[0085] The fourth sub-step is to determine the average value of the determined logarithmic loss data of each document as the document average value data.
[0086] In the fifth sub-step, the negative of the above document average value data is determined as the document negative logarithmic loss data.
[0087] The sixth sub-step is to generate weighted negative logarithmic loss data based on the above-mentioned weighted data pair. The above-mentioned weighted negative logarithmic loss data can be the loss value corresponding to the weighted data pair. In practice, the method of generating weighted negative logarithmic loss data can refer to the specific implementation method of generating the above-mentioned document negative logarithmic loss data, which will not be repeated here.
[0088] The seventh sub-step is to generate fractional negative logarithmic loss data based on the above-mentioned score data pair group. The above-mentioned fractional negative logarithmic loss data can be the loss value corresponding to the score data pair group. In practice, the method for generating the above-mentioned fractional negative logarithmic loss data can refer to the specific implementation method of generating the above-mentioned document negative logarithmic loss data, and will not be repeated here.
[0089] In an eighth sub-step, the product of the preset first parameter data and the document negative logarithmic loss data is determined as the document parameter data. The first parameter data may be a value used to control the weight of the document negative logarithmic loss data. For example, the first parameter data may be 0.6.
[0090] In a ninth sub-step, the product of the preset second parameter data and the weighted negative logarithmic loss data is determined as the weighted parameter data. The second parameter data may be a numerical value for controlling the weight of the weighted negative logarithmic loss data. For example, the second parameter data may be 0.3.
[0091] In a tenth sub-step, the product of the preset third parameter data and the fractional negative logarithmic loss data is determined as the fractional parameter data. The third parameter data may be a value used to control the weight of the fractional negative logarithmic loss data. For example, the third parameter data may be 0.1.
[0092] In the eleventh sub-step, rewritten quantized data is generated based on the sample rewritten query data and the pre-trained rewritten query data corresponding to the sample. The rewritten quantized data may be a ROUGE metric between the sample rewritten query data and the pre-trained rewritten query data. In practice, the execution entity may perform a similarity comparison between the sample rewritten query data and the pre-trained rewritten query data using the classification algorithm to obtain the rewritten quantized data.
[0093] In the twelfth sub-step, the product of the preset weight parameter and the rewritten quantization data is determined as parameter quantization data.
[0094] In the thirteenth sub-step, the sum of the document parameter data, the weight parameter data, the score parameter data and the rewritten quantization data is determined as pre-training feedback data.
[0095] In the second step, the average value of each determined pre-training feedback data is determined as the average feedback data.
[0096] The third step is to generate a target neural network based on the average feedback data, the initialized neural network, and a preset penalty coefficient. The target neural network can be an optimized initial neural network. In practice, first, the execution entity can determine the spliced data sequence corresponding to any sample in the sample set as the target spliced data sequence. The method for generating the spliced data sequence corresponding to the sample can refer to the specific implementation of step 1022 and will not be described in detail here. Then, the target spliced data sequence can be input into the initial neural network to obtain the rewritten query data output by the initial neural network as the initial rewritten query data. Then, the target spliced data sequence can be input into the initialized neural network to obtain the rewritten query data output by the initialized neural network as the initial rewritten query data. Then, the relative entropy between the initial rewritten query data and the initialized rewritten query data can be determined as relative entropy data. Next, the product of the preset penalty coefficient and the relative entropy data can be determined as penalty term data. The penalty coefficient can be a numerical value representing the weight corresponding to the relative entropy data. Then, an equation can be constructed as the target training equation by setting the difference between the average feedback data and the penalty term data equal to the target training value. The target training value can be a dependent variable. Then, the network parameters of the initial neural network can be continuously adjusted through a policy optimization algorithm until the target training value converges, and the adjusted initial neural network is obtained as the target neural network. The above-mentioned policy optimization algorithm can be a proximal policy optimization algorithm (PPO).
[0097] The above technical solution and its related contents, combined with steps 101 to 105, serve as an inventive feature of an embodiment of the present disclosure, solving the problem of "high computing resource consumption." Factors that often lead to storage resource waste are as follows: When using a rule engine to rewrite query data, it is necessary to traverse the rules in the rule engine one by one until a rule matching the query data is found before the query data can be rewritten, resulting in low processing efficiency when rewriting the query data. Moreover, each time different query data is rewritten, computing resources need to be called upon to match the query data with the rules in the rule engine one by one, resulting in high computing resource consumption when rewriting the query data. If the above factors are resolved, computing resource consumption can be reduced. To achieve this effect, the present disclosure first performs the following steps for each sample in the at least one sample: First, the sample conversation history data sequence and the sample current query data included in the sample are spliced together to obtain sample spliced data. In this way, the sample conversation history data sequence and the sample current query data can be spliced together. Next, based on the sample concatenation data, the reward loss value data, the sample rewritten query data included in the sample, the pre-trained rewritten query data corresponding to the sample, and the sample rewritten query data, document data pair groups, weight data pair groups, and score data pair groups are generated. Each document data pair in the document data pair group includes document selection data and document rejection data, each weight data pair in the weight data pair group includes weight selection data and weight rejection data, and each score data pair in the score data pair group includes score selection data and score rejection data. Thus, document data pair groups, weight data pair groups, and score data pair groups are obtained. Then, for each document data pair in the document data pair group, the following steps are performed: First, based on the document selection data and document rejection data included in the document data pair, document selection score data and document rejection score data are generated. Thus, score indicators corresponding to the document selection data and document rejection data are obtained. Second, the difference between the document selection score data and the document rejection score data is determined as document difference data. Then, based on the document difference data, document loss data is generated. Then, the logarithm of the above-mentioned document loss data is determined as the document logarithmic loss data. Then, the average value of the determined document logarithmic loss data is determined as the document average data. Then, the negative number of the above-mentioned document average data is determined as the document negative logarithmic loss data. Thus, the document negative logarithmic loss data of the document data pair group can be obtained. Then, based on the above-mentioned weight data pair group, the weighted negative logarithmic loss data is generated. Thus, the weighted negative logarithmic loss data can be obtained. Then, based on the above-mentioned score data pair group, the score negative logarithmic loss data is generated. Thus, the score negative logarithmic loss data can be obtained. Then, the product of the preset first parameter data and the above-mentioned document negative logarithmic loss data is determined as the document parameter data.Thus, the weight of the document negative logarithmic loss data can be determined. Then, the product of the preset second parameter data and the above-mentioned weighted negative logarithmic loss data is determined as weight parameter data. Thus, the weight of the weighted negative logarithmic loss data can be determined. Then, the product of the preset third parameter data and the above-mentioned score negative logarithmic loss data is determined as score parameter data. Thus, the weight of the score negative logarithmic loss data can be determined. Then, based on the sample rewrite query data and pre-trained rewrite query data corresponding to the above-mentioned sample, rewrite quantization data is generated. Then, the product of the preset weight parameter and the above-mentioned rewrite quantization data is determined as parameter quantization data. Thus, the weight of the rewrite quantization data can be determined. Then, the sum of the above-mentioned document parameter data, the above-mentioned weight parameter data, the above-mentioned score parameter data and the above-mentioned rewrite quantization data is determined as pre-training feedback data. Then, the average value of the determined individual pre-training feedback data is determined as average feedback data. Finally, based on the above-mentioned average feedback data, the initialized neural network and the preset penalty coefficient, a target neural network is generated. Thus, the target neural network can be obtained. Furthermore, because the query rewriting model can be obtained by iteratively training the initial neural network, the query rewriting model can then be used to directly rewrite the query data without requiring matching or other processing on the query data. This improves the efficiency of query rewriting. Furthermore, because the trained query rewriting model can be directly used to rewrite the query data without requiring computing resources for matching or other processing, the computing resources consumed when rewriting the query data can be reduced.
[0098] Step 1026: In response to determining that the updated number of iterations satisfies a preset iteration condition, the target neural network is determined to be a query rewriting model.
[0099] In some embodiments, the execution entity may determine the target neural network as the query rewriting model in response to determining that the updated number of iterations satisfies a preset iteration condition. The iteration condition may be that the updated number of iterations is equal to a preset target iteration value. The target iteration value may be a pre-set value. The specific setting of the target iteration value is not limited herein.
[0100] Optionally, after step 1026, the execution entity may further perform the following steps: In response to determining that the updated number of iterations does not satisfy the above iteration condition, unused samples are used to form a sample set, and the target neural network is used as the initial neural network to perform the above training step again.
[0101] Step 103: in response to receiving the current query data sent by the user terminal, obtaining a conversation history data sequence corresponding to the user terminal.
[0102] In some embodiments, the execution entity may, in response to receiving current query data sent by a user terminal, obtain a sequence of conversation history data corresponding to the user terminal. Each conversation history data in the sequence of conversation history data includes historical query data and historical answer data. In practice, the execution entity may obtain the sequence of conversation history data from a pre-set conversation history database. The conversation history database may be a database for storing conversation history data sequences.
[0103] Step 104: Generate rewritten answer data based on the conversation history data sequence, current query data, and the query rewriting model.
[0104] In some embodiments, the execution entity may generate rewritten answer data based on the conversation history data sequence, the current query data, and the query rewriting model.
[0105] In some optional implementations of some embodiments, the execution entity may generate rewritten answer data based on the conversation history data sequence, the current query data, and the query rewriting model through the following steps: In the first step, among the individual conversation history data included in the conversation history data sequence, conversation history data that meets a preset conversation condition is determined as target conversation history data. The preset conversation condition may be that the conversation history data is generated most recently to the current query data. The target conversation history data may include target query data and target answer data. The target query data may be historical query data corresponding to the target conversation history data. The target answer data may be historical answer data corresponding to the target conversation history data.
[0106] In the second step, the conversation history data sequence and the current query data are input into the query rewriting model to obtain rewritten query data.
[0107] The third step is to generate a query document data sequence corresponding to the rewritten query data based on the rewritten query data. Each query document data in the query document data sequence may be a document related to the content of the rewritten query data. In practice, the execution entity may input the rewritten query data into the dense search engine to obtain the query document data sequence.
[0108] The fourth step is to generate rewritten answer data corresponding to the rewritten query data based on the rewritten query data, the query document data sequence, and the target conversation history data. In practice, the execution entity may input the rewritten query data, the query document data sequence, and the target conversation history data into the large language model to generate the rewritten answer data.
[0109] Step 105: Send the rewritten answer data to the user terminal.
[0110] In some embodiments, the execution entity may send the rewritten answer data to the user terminal.
[0111] The above-described various embodiments of the present disclosure have the following beneficial effects: The query rewriting method based on multi-faceted feedback in some embodiments of the present disclosure can reduce computing resource waste. Specifically, the reason for this computing resource waste is that the query data input by the user is complex and changeable. The rule engine only modifies the query data according to pre-set rules, which can easily lead to the modified query data being significantly different from the user's true intent. In turn, the answer generated based on the modified query data is significantly different from the user's actual needs. This can easily lead to the same query data being modified multiple times to meet the user's actual needs, which can easily lead to computing resource waste due to repeated modifications. Based on this, the query rewriting method based on multi-faceted feedback in some embodiments of the present disclosure first obtains a sample set, wherein each sample in the sample set includes a sample conversation history data sequence, a sample current query data, and a sample rewritten query data, and each sample conversation history data in the sample conversation history data sequence includes sample historical query data and sample historical answer data. Thus, a sample set is obtained. Next, based on the sample set, the following training steps are performed: First, the number of iterations is updated based on a preset value. Thus, the number of iterations during training can be determined. Next, the sample conversation history data sequence and the sample current query data included in each of at least one sample in the sample set are input into an initial neural network to obtain sampled rewritten query data and pre-trained rewritten query data corresponding to each of the at least one sample. Thus, by inputting both the sample conversation history data sequence and the sample current query data into the initial neural network, the initial neural network's ability to reference contextual semantics can be trained, improving the accuracy of query data rewriting by the trained initial neural network, and ensuring that the rewritten query data is more closely aligned with the user's true intent. Then, based on the sample rewritten query data included in each of the at least one sample and the pre-trained rewritten query data corresponding to each of the at least one sample, the initial neural network is initialized to obtain an initialized neural network. Thus, network parameters of the initial neural network can be adjusted based on the pre-trained rewritten query data output by the initial neural network to obtain an initialized neural network. Then, based on the sample rewritten query data included in each of the at least one sample, the pre-trained rewritten query data corresponding to each of the at least one sample, and the sampled rewritten query data, reward loss value data corresponding to each of the at least one sample is generated. Thus, based on the pre-trained rewritten query data, the sampled rewritten query data, and the sample rewritten query data, a loss value can be obtained between the query data output by the initial neural network and the actual query data to be obtained. Then, based on the initialized neural network, the sample rewritten query data included in each of the at least one sample, the pre-trained rewritten query data corresponding to each of the at least one sample, and the reward loss value data, a target neural network is generated. Thus, the target neural network can be obtained.Then, in response to determining that the updated number of iterations satisfies a preset iteration condition, the target neural network is determined as a query rewriting model. Thus, a query rewriting model can be obtained. Then, in response to receiving the current query data sent by the user terminal, a sequence of conversation history data corresponding to the user terminal is obtained. Thus, the original data required for rewriting can be obtained. Then, based on the sequence of conversation history data, the current query data, and the query rewriting model, rewritten answer data is generated. Thus, answer data corresponding to the rewritten query data can be obtained. Finally, the rewritten answer data is sent to the user terminal. Thus, the rewritten answer data can be sent to the user terminal. Because the network parameters of the initial neural network can be continuously optimized using the reward loss value data and the initialized neural network, the accuracy of the query rewriting model obtained after training when rewriting query data can be improved, making the rewritten query data more closely aligned with the user's true intentions. Also, when rewriting the query data input by the user, the context semantics corresponding to the query data can be referred to first, and then the query data can be rewritten. Therefore, the accuracy of the rewritten query data can be improved, and the rewritten query data can be closer to the user's true intention, thereby reducing the probability of having to rewrite the same query data multiple times to meet user needs, thereby reducing the waste of computing resources during multiple rewrites.
[0112] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a query rewriting device based on multi-faceted feedback. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0113] like Figure 1As shown, some embodiments of the query rewriting apparatus 200 based on multi-faceted feedback include: a first acquisition unit 201, an execution unit 202, a second acquisition unit 203, a generation unit 204, and a sending unit 205. The first acquisition unit 201 is configured to acquire a sample set, wherein each sample in the sample set includes a sample conversation history data sequence, a sample current query data, and a sample rewritten query data, and each sample conversation history data in the sample conversation history data sequence includes sample history query data and sample history answer data; the execution unit 202 is configured to perform the following training steps based on the sample set: based on a preset value, update the number of iterations; input the sample conversation history data sequence and the sample current query data included in each sample of at least one sample in the sample set into the initial neural network to obtain the sample rewritten query data and pre-trained rewritten query data corresponding to each sample in the at least one sample; initialize the initial neural network based on the sample rewritten query data included in each sample in the at least one sample and the pre-trained rewritten query data corresponding to each sample in the at least one sample to obtain an initialized neural network; based on the at least one sample The reward loss value data corresponding to each sample in the at least one sample is generated based on the sample rewrite query data included in each sample, the pre-trained rewrite query data corresponding to each sample in the at least one sample, and the sampling rewrite query data; the target neural network is generated based on the initialized neural network, the sample rewrite query data included in each sample in the at least one sample, the pre-trained rewrite query data corresponding to each sample in the at least one sample, and the reward loss value data; in response to determining that the updated number of iterations meets the preset iteration condition, the target neural network is determined as a query rewriting model; the second acquisition unit 203 is configured to acquire a conversation history data sequence corresponding to the above user terminal in response to receiving the current query data sent by the user terminal; the generation unit 204 is configured to generate rewrite answer data based on the above conversation history data sequence, the above current query data, and the above query rewriting model; the sending unit 205 is configured to send the above rewrite answer data to the above user terminal.
[0114] It is understandable that the various units recorded in the query rewriting device 200 based on multi-faceted feedback are similar to the reference Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features and beneficial effects described above for the method are also applicable to the query rewriting device 200 based on multi-aspect feedback and the units included therein, and will not be described in detail here.
[0115] Reference below Figure 3 , which shows a structural diagram of an electronic device (such as a computing device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0116] like Figure 3 As shown, electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 302 or programs loaded from a storage device 308 into a random access memory (RAM) 303. RAM 303 also stores various programs and data required for the operation of electronic device 300. Processing device 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to bus 304.
[0117] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0118] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.
[0119] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. Furthermore, in some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.
[0120] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0121] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently and not be assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: obtains a sample set, wherein each sample in the above-mentioned sample set includes a sample conversation history data sequence, sample current query data and sample rewrite query data, and each sample conversation history data in the above-mentioned sample conversation history data sequence includes sample history query data and sample history answer data; based on the sample set, performs the following training steps: based on a preset value, updates the number of iterations; inputs the sample conversation history data sequence and sample current query data included in each sample of at least one sample in the sample set into the initial neural network, and obtains the sample rewrite query data and pre-trained rewrite query data corresponding to each sample of the above-mentioned at least one sample; based on the sample rewrite query data included in each sample of the above-mentioned at least one sample and the pre-trained rewrite query data corresponding to each sample of the above-mentioned at least one sample, the initial neural network is trained. The network is initialized to obtain an initialized neural network; based on the sample rewriting query data included in each sample of the at least one sample, the pre-trained rewriting query data corresponding to each sample of the at least one sample, and the sampling rewriting query data, the reward loss value data corresponding to each sample of the at least one sample is generated; based on the initialized neural network, the sample rewriting query data included in each sample of the at least one sample, the pre-trained rewriting query data corresponding to each sample of the at least one sample, and the reward loss value data, a target neural network is generated; in response to determining that the updated number of iterations meets a preset iteration condition, the target neural network is determined as a query rewriting model; in response to receiving the current query data sent by the user terminal, a conversation history data sequence corresponding to the user terminal is obtained; based on the conversation history data sequence, the current query data, and the query rewriting model, rewriting answer data is generated; and the rewriting answer data is sent to the user terminal.
[0122] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0124] The units described in some embodiments of the present disclosure may be implemented in software or hardware. The units described may also be provided in a processor. For example, they may be described as follows: a processor includes a first acquisition unit, an execution unit, a second acquisition unit, a generation unit, and a sending unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the first acquisition unit may also be described as a "unit for acquiring a sample set."
[0125] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0126] The above descriptions are merely some preferred embodiments of the present disclosure and illustrate the underlying technical principles. Those skilled in the art should understand that the scope of the invention encompassed by the embodiments of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned inventive concept. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A query rewriting method based on multi-faceted feedback, comprising: Acquire a sample set, wherein each sample in the sample set includes a sample conversation history data sequence, sample current query data, and sample rewritten query data, and each sample conversation history data in the sample conversation history data sequence includes sample history query data and sample history answer data; Based on the sample set, the following training steps are performed: Based on the preset value, the number of iterations is updated; Inputting a sample conversation history data sequence and sample current query data included in each of at least one sample in the sample set into an initial neural network to obtain sample rewritten query data and pre-trained rewritten query data corresponding to each of the at least one sample; Initializing an initial neural network based on the sample rewriting query data included in each sample of the at least one sample and the pre-trained rewriting query data corresponding to each sample of the at least one sample to obtain an initialized neural network; Generate reward loss value data corresponding to each sample in the at least one sample based on the sample rewriting query data included in each sample in the at least one sample, the pre-training rewriting query data corresponding to each sample in the at least one sample, and the sampling rewriting query data; Generate a target neural network based on the initialized neural network, the sample rewrite query data included in each sample of the at least one sample, the pre-trained rewrite query data corresponding to each sample of the at least one sample, and the reward loss value data; In response to determining that the updated number of iterations satisfies a preset iteration condition, determining the target neural network as a query rewriting model; In response to receiving current query data sent by a user terminal, obtaining a conversation history data sequence corresponding to the user terminal; generating rewritten answer data based on the conversation history data sequence, the current query data, and the query rewriting model; The rewritten answer data is sent to the user terminal.
2. The method according to claim 1, wherein After determining the target neural network as the query rewriting model in response to determining that the updated number of iterations satisfies a preset iteration condition, the method further includes: In response to determining that the updated number of iterations does not satisfy the iteration condition, unused samples are used to form a sample set, and the target neural network is used as an initial neural network to perform the training step again.
3. The method according to claim 1, wherein Each conversation history data in the conversation history data sequence includes historical query data and historical answer data; And generating rewritten answer data based on the conversation history data sequence, the current query data and the query rewriting model includes: determining, among the individual conversation history data included in the conversation history data sequence, conversation history data that meets a preset conversation condition as target conversation history data; Inputting the conversation history data sequence and the current query data into the query rewriting model to obtain rewritten query data; Based on the rewritten query data, generating a query document data sequence corresponding to the rewritten query data; Based on the rewritten query data, the query document data sequence, and the target conversation history data, rewritten answer data corresponding to the rewritten query data is generated.
4. The method according to claim 1, wherein The step of inputting the sample conversation history data sequence and the sample current query data included in each of at least one sample in the sample set into the initial neural network to obtain the sample rewriting query data and the pre-trained rewriting query data corresponding to each of the at least one sample comprises: Splicing the sample conversation history data sequence and the sample current query data included in each of at least one sample in the sample set to obtain a spliced data sequence corresponding to each of the at least one sample; The spliced data sequence corresponding to each sample in the at least one sample is input into the initial neural network to obtain sampling rewrite query data and pretrained rewrite query data corresponding to each sample in the at least one sample, wherein the pretrained rewrite query data includes a pretrained rewrite query byte data sequence, and each pretrained rewrite query byte data in the pretrained rewrite query byte data sequence corresponds to byte distribution data.
5. The method according to claim 4, wherein The sample rewriting query data includes a sample byte data sequence, each sample byte data in the sample byte data sequence corresponds to sample byte distribution data; and initializing the initial neural network based on the sample rewriting query data included in each sample of the at least one sample and the pre-trained rewriting query data corresponding to each sample of the at least one sample to obtain the initialized neural network, including: For each sample of the at least one sample, performing the following steps: Based on the pre-trained rewriting query data corresponding to the sample, determining each byte distribution data corresponding to the pre-trained rewriting query data as each pre-trained byte distribution data; For each sample byte distribution data in each sample byte distribution data corresponding to the sample, the following steps are performed: Determining the pre-training byte distribution data corresponding to the sample byte distribution data in each of the pre-training byte distribution data as the target pre-training byte distribution data; Determining target logarithmic distribution data based on the target pre-trained byte distribution data; For each target logarithmic distribution data in the target logarithmic distribution data, the following steps are performed: Determining target sample element data corresponding to the target logarithmic distribution data based on the sample byte distribution data; Determine the product of the target logarithmic distribution data and the target sample element data as byte distribution product data; Determining the sum of the determined byte distribution product data as summed byte distribution data; determining a negative number of the summed byte distribution data as rewriting loss value data; determining a sum of the determined respective rewriting loss value data as comprehensive rewriting loss data; determining an average value of the determined respective integrated rewrite loss data as average rewrite loss data; Based on the average rewriting loss data, the network parameters of the initial neural network are adjusted to initialize the initial neural network to obtain an initialized neural network.
6. The method according to claim 1, wherein The sample current query data included in each of the at least one sample corresponds to a sample target document and sample answer data; and generating reward loss value data corresponding to each of the at least one sample based on the sample rewritten query data included in each of the at least one sample, the pre-trained rewritten query data corresponding to each of the at least one sample, and the sampled rewritten query data, comprises: For each sample of the at least one sample, performing the following steps: Performing feature extraction processing on the sample rewriting query data included in the sample, the pre-trained rewriting query data corresponding to the sample, and the sampled rewriting query data to obtain sample rewriting dense feature information, pre-trained rewriting dense feature information, and sampled rewriting dense feature information; Performing feature extraction processing on the sample target document corresponding to the sample to obtain document dense feature information; Determining the similarity between the sample rewriting intensive feature information and the document intensive feature information as the sample rewriting similarity; Determining the similarity between the pre-trained rewriting dense feature information and the document dense feature information as a pre-trained rewriting similarity; Determining the similarity between the sample rewriting dense feature information and the document dense feature information as the sample rewriting similarity; combining the sample rewriting similarity, the pre-trained rewriting similarity, and the sampled rewriting similarity into document loss value data; Performing feature extraction processing on the sample answer data corresponding to the sample to obtain dense feature information of the sample answer; Generate weight loss value data based on the sample answer dense feature information, the sample rewriting query data included in the sample, the pre-trained rewriting query data corresponding to the sample, and the sampled rewriting query data; Generate sample rewriting answer data, pretrained rewriting answer data, and sampled rewriting answer data based on the sample rewriting query data included in the sample, the pretrained rewriting query data corresponding to the sample, and the sampled rewriting query data; Generate sample rewritten answer score data, pretrained rewritten answer score data, and sample rewritten answer score data based on the sample rewritten answer data, the pretrained rewritten answer data, the sampled rewritten answer data, and the sample answer data corresponding to the sample; combining the sample rewritten answer score data, the pre-trained rewritten answer score data, and the sampled rewritten answer score data into score loss value data; The document loss value data, the weight loss value data, and the score loss value data are combined into reward loss value data corresponding to the sample.
7. The method according to claim 6, wherein: The generating weight loss value data based on the sample answer dense feature information, the sample rewriting query data included in the sample, the pre-trained rewriting query data corresponding to the sample, and the sampling rewriting query data includes: Generate a sample rewritten document sequence, a pretrained rewritten document sequence, and a sampled rewritten document sequence based on the sample rewritten query data included in the sample, the pretrained rewritten query data corresponding to the sample, and the sampled rewritten query data; For each sample rewritten document in the sequence of sample rewritten documents, perform the following steps: Performing feature extraction processing on the sample rewritten document to obtain dense feature information of the sample document; Determining the similarity between the sample answer intensive feature information and the sample document intensive feature information as the sample document similarity; Based on the sample rewritten document, determining sample ranking data corresponding to the sample rewritten document; Determine the product of the sample ranking data and the sample document similarity as sample document weight data; Determine the sum of the determined weight data of each sample document as sample weight intensive data; Generate pre-trained weight-intensive data and sampled weight-intensive data based on the pre-trained rewritten document sequence, the sampled rewritten document sequence, and the sample answer-intensive feature information; The sample weight dense data, the pre-trained weight dense data and the sampled weight dense data are combined into weight loss value data.
8. A query rewriting device based on multi-faceted feedback, comprising: a first acquisition unit configured to acquire a sample set, wherein each sample in the sample set includes a sample conversation history data sequence, sample current query data, and sample rewritten query data, and each sample conversation history data in the sample conversation history data sequence includes sample history query data and sample history answer data; The execution unit is configured to perform the following training steps based on the sample set: updating the number of iterations based on a preset value; inputting the sample conversation history data sequence and sample current query data included in each of at least one sample in the sample set into the initial neural network to obtain sampling rewrite query data and pre-training rewrite query data corresponding to each of the at least one sample; initializing the initial neural network based on the sample rewrite query data included in each of the at least one sample and the pre-training rewrite query data corresponding to each of the at least one sample to obtain an initialized neural network; generating reward loss value data corresponding to each of the at least one sample based on the sample rewrite query data included in each of the at least one sample, the pre-training rewrite query data corresponding to each of the at least one sample, and the sampling rewrite query data; generating a target neural network based on the initialized neural network, the sample rewrite query data included in each of the at least one sample, the pre-training rewrite query data corresponding to each of the at least one sample, and the reward loss value data; and determining the target neural network as a query rewriting model in response to determining that the updated number of iterations meets a preset iteration condition; a second acquiring unit configured to acquire a conversation history data sequence corresponding to the user terminal in response to receiving current query data sent by the user terminal; a generating unit configured to generate rewritten answer data based on the conversation history data sequence, the current query data, and the query rewriting model; A sending unit is configured to send the rewritten answer data to the user terminal.
9. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.