Question and answer data processing method and device and storage medium

By using a generative scoring model in the question-and-answer pair output by the generative language model, and using components such as feature transformation module and pooling layer for scoring calculation, the problem of insufficient performance of the existing understanding scoring model is solved, and a more efficient question-and-answer pair scoring is achieved.

CN120218169APending Publication Date: 2025-06-27HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311789068.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing comprehension scoring model lacks performance when dealing with question-and-answer matching outputs from generative language models and cannot meet the current scoring needs.

Method used

A generative scoring model is used to score the Q&A pairs output by the generative language model. By obtaining the query instructions and reply text in the Q&A pair, feature extraction and scoring calculation are used using components such as feature transformation module, masked average pooling layer and fully connected layer.

Benefits of technology

It improves the scoring performance of Q&A, makes the scoring results more reliable and accurate, and makes full use of the text generation ability of the generative large language model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218169A_ABST
    Figure CN120218169A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a question and answer data processing method and device and a storage medium. In the question and answer data processing method, after a reply text generated by a target language model for an inquiry instruction is obtained, a target scoring model can be adopted to score question and answer pairs formed by the inquiry instruction and the reply text. Wherein a base model of the target language model is a pre-trained generative large language model, and a feature transformation module in the target scoring model is obtained by fine tuning according to an encoder and / or a decoder in the pre-trained generative large language model. In the embodiment, partial model structures and model parameters of the target scoring model and the target language model are from the same base model, so that knowledge migration and knowledge sharing can be realized on the basis of the structures and the parameters of the same base model, the scoring performance of the target scoring model is improved, and the scoring result is more reliable. And on the other hand, the scoring model of the target language model is generated by utilizing the base model of the target language model, so that the existing model structure and model parameters can be fully reused, and the construction cost of the target scoring model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, and storage medium for processing question-and-answer data. Background Art

[0002] In the field of Natural Language Processing (NLP), a scoring model refers to a model that uses machine learning algorithms to classify or score text. In the NLP scoring model, some evaluation metrics are usually used to evaluate the performance of the model, such as accuracy, recall, etc. These metrics can be used to measure the performance of the NLP model in classification or scoring tasks, so as to better optimize and improve the model. After inputting a text pair (such as a question and an answer) into the scoring model, the scoring model can output a score to represent the matching degree between the answer and the question in the text pair.

[0003] In the NLP field, traditional scoring models are usually implemented based on small understanding-based models, such as BERT (Bidirectional Encoder Representations from Transformers) and its derivative models. In recent years, with the continuous development of generative large models, the number of model parameters and the amount of pre-trained data have increased extremely rapidly, and the understanding-based scoring models can no longer meet the scoring requirements. Therefore, there is a need to propose a new solution. Summary of the Invention

[0004] Multiple aspects of this application provide a method, device, and storage medium for processing question-and-answer data, which are used to score question-and-answer pairs output by a generative language model based on a generative scoring model, and improve the scoring performance.

[0005] An embodiment of this application provides a method for processing question-and-answer data, including: obtaining at least one question-and-answer pair, where any question-and-answer pair consists of an inquiry instruction and a reply text; the reply text is generated by a target language model according to the inquiry instruction; the base model of the target language model is a pre-trained generative large language model; the pre-trained generative large language model includes at least: an encoder and / or a decoder; inputting the at least one question-and-answer pair into a target scoring model, so that the target scoring model scores the at least one question-and-answer pair to obtain an evaluation score for each of the at least one question-and-answer pair; where the target scoring model includes at least a feature transformation module, and the feature transformation module is fine-tuned according to the encoder and / or decoder in the pre-trained generative large language model.

[0006] Optionally, the target scoring model further includes: a masked average pooling layer and a fully connected layer; inputting the at least one question-and-answer pair into the target scoring model to enable the target scoring model to score the at least one question-and-answer pair, obtaining the evaluation scores of the at least one question-and-answer pair respectively, including: obtaining a first text sequence of the query instruction in the question-and-answer pair and a second text sequence of the answer text; inputting the first text sequence and the second text sequence into a feature transformation module in the target scoring model, so that the feature transformation module performs feature transformation according to the first text sequence and the second text sequence to obtain a vector sequence of the question-and-answer pair; inputting the vector sequence of the question-and-answer pair into the masked average pooling layer, so that the average pooling layer performs average pooling calculation on the vector sequence of the question-and-answer pair according to preset weight parameters to obtain an average pooling calculation result; using the fully connected layer to perform a fully connected calculation on the average pooling calculation result according to a pre-trained weight matrix and bias parameters to obtain the evaluation score of the question-and-answer pair.

[0007] Optionally, inputting the first text sequence and the second text sequence into the feature transformation module in the target scoring model includes: concatenating the first text sequence and the second text sequence to obtain a third text sequence; inputting the third text sequence into an encoder in the scoring model, so that the encoder performs encoding processing on multiple text elements in the third text sequence to obtain a vector sequence of the question-and-answer pair.

[0008] Optionally, inputting the first text sequence and the second text sequence into the feature transformation module in the target scoring model includes: concatenating the first text sequence and the second text sequence to obtain a concatenated third text sequence; inputting the third text sequence into a decoder in the scoring model, so that the decoder performs decoding processing on multiple text elements in the third text sequence to obtain feature vectors of the multiple text elements in the third text sequence respectively; obtaining a vector sequence of the question-and-answer pair according to the feature vectors of the multiple text elements in the third text sequence.

[0009] Optionally, inputting the first text sequence and the second text sequence into the feature transformation module in the target scoring model includes: inputting the first text sequence into an encoder in the scoring model, so that the encoder performs encoding processing on multiple text elements in the first text sequence to obtain a vector sequence corresponding to the first text sequence; inputting the vector sequence corresponding to the first text sequence and the second text sequence into a decoder in the scoring model, so that the decoder performs decoding processing on the vector sequence corresponding to the first text sequence and the second text sequence to obtain a vector sequence of the question-and-answer pair.

[0010] Optionally, inputting the first text sequence and the second text sequence into a feature transformation module in the target scoring model includes: inputting the first text sequence into an encoder in the scoring model, so that the encoder encodes multiple text elements in the first text sequence to obtain a first vector sequence corresponding to the first text sequence; inputting the first vector sequence and the second text sequence into a decoder in the scoring model, so that the decoder decodes the first vector sequence and the second text sequence to obtain a second vector sequence; and obtaining a vector sequence of the question-and-answer pair according to the first vector sequence and the second vector sequence.

[0011] Optionally, the method further includes: obtaining training data; the training data includes multiple groups of question-and-answer sample pairs; any group of question-and-answer sample pairs includes: an instruction sample, a positive answer sample of the instruction sample, and multiple negative answer samples; inputting the multiple groups of question-and-answer sample pairs into an initial model to obtain prediction evaluation scores of the multiple groups of question-and-answer sample pairs respectively; the initial model includes: an encoder and / or a decoder in the pre-trained generative large language model; the prediction evaluation score of any group of question-and-answer sample pairs includes: a prediction evaluation score of a positive sample pair formed by the instruction sample and its positive answer sample, and prediction evaluation scores of multiple negative sample pairs formed by the instruction sample and its multiple negative answer samples respectively; for any group of question-and-answer sample pairs, obtaining a difference between the prediction evaluation score of the positive sample pair in the group of question-and-answer sample pairs and the maximum value of the prediction evaluation scores of the multiple negative sample pairs respectively; determining a main loss according to a first comparison threshold and the difference; obtaining a first average value of the prediction evaluation scores of the positive sample pairs included in the multiple groups of question-and-answer sample pairs respectively, and a second average value of the prediction evaluation scores of the negative sample pairs included in the multiple groups of question-and-answer sample pairs respectively; determining an auxiliary loss according to a difference between the first average value and the second average value and a second comparison threshold; determining a scoring loss of the initial model according to the main loss and the auxiliary loss; and fine-tuning the initial model with the goal of converging the scoring loss until the scoring loss converges to a specified range, and outputting the converged initial model as the target scoring model.

[0012] Optionally, the inquiry instructions of the at least one question-and-answer pair are the same; after obtaining the evaluation scores of the at least one question-and-answer pair respectively, it further includes: screening out a target question-and-answer pair from the at least one question-and-answer pair according to the evaluation scores of the at least one question-and-answer pair respectively, and providing a question-and-answer service according to the target question-and-answer pair; or determining a scoring loss of the scoring model according to the evaluation scores of the at least one question-and-answer pair respectively and the sample type of the answer text in the at least one question-and-answer pair; and optimizing the target scoring model according to the scoring loss; the sample type includes: a positive sample type or a negative sample type.

[0013] An embodiment of the present application further provides a server, including: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions for: executing the steps in the method provided by the embodiment of the present application.

[0014] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it can implement the steps in the method provided by the embodiment of the present application.

[0015] In the question-and-answer data processing method provided by the embodiment of the present application, after obtaining the response text generated by the target language model for the query instruction, a target scoring model can be used to score the question-and-answer pair formed by the query instruction and the response text. Among them, the base model of the target language model is a pre-trained generative large language model, and the feature transformation module in the target scoring model is fine-tuned according to the encoder and / or decoder in the pre-trained generative large language model. In this implementation manner, part of the model structure and model parameters of the target scoring model and the target language model come from the same base model, which is conducive to realizing knowledge transfer and knowledge sharing based on the structure and parameters of the same base model, thereby improving the scoring performance of the target scoring model and making the scoring result more reliable. On the other hand, using the base model of the target language model to generate the scoring model of the target language model can fully reuse the existing model structure and model parameters, reducing the construction cost of the target scoring model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0017] Figure 1 It is a schematic flowchart of a question-and-answer data processing method provided by an exemplary embodiment of the present application;

[0018] Figure 2 It is a schematic structural diagram of a scoring model provided by an exemplary embodiment of the present application;

[0019] Figure 3 It is a schematic structural diagram of a scoring model provided by another exemplary embodiment of the present application;

[0020] Figure 4 It is a schematic structural diagram of a scoring model provided by yet another exemplary embodiment of the present application;

[0021] Figure 5 It is a schematic structural diagram of a scoring model provided by yet another exemplary embodiment of the present application;

[0022] Figure 6 The structural schematic diagram of the server provided for an exemplary embodiment of the present application. Detailed implementation manners

[0023] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0024] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plural" generally includes at least two, but does not exclude the case of including at least one.

[0025] It should be understood that the term "and / or" used herein is only a kind of association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0026] It should also be noted that the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a commodity or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such commodity or system. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the commodity or system including the said element.

[0027] In the field of NLP, traditional scoring models are usually implemented based on understanding-based small models, such as BERT and its derivative models. In recent years, with the continuous development of generative large models, the number of model parameters and the amount of pre-trained data have increased extremely rapidly, and the understanding-based scoring models can no longer meet the scoring requirements.

[0028] In view of the above technical problems, in some embodiments of the present application, a solution is provided. The technical solutions provided in each embodiment of the present application will be described in detail below with reference to the drawings.

[0029] Figure 1It is a schematic flowchart of a Q&A data processing method provided by an exemplary embodiment of the present application. The method may include the following steps as Figure 1 shown:

[0030] Step 101: Obtain at least one Q&A pair. Any Q&A pair consists of an inquiry instruction and a reply text. The reply text is generated by a target language model according to the inquiry instruction. The base model of the target language model is a pre-trained generative large language model. The pre-trained generative large language model includes at least: an encoder and / or a decoder.

[0031] Step 102: Input the at least one Q&A pair into a target scoring model, so that the target scoring model scores the at least one Q&A pair to obtain the evaluation scores of the at least one Q&A pair respectively. Among them, the target scoring model includes at least a feature transformation module, and the feature transformation module is fine-tuned according to the encoder and / or decoder in the pre-trained generative large language model.

[0032] Among them, the base model of the target language model can be a generative large language model. A generative language model is a natural language processing model based on a neural network, used to generate new language content according to the input text prompt. A generative large language model refers to a generative language model with the number of parameters greater than a set threshold. In this embodiment, the base model of the target language model can be: GPT (Generative Pre-trained Transformer) or T5 (Text-to-Text Transfer Transformer), and this embodiment does not limit this. A generative large language model is obtained by pre-training with a large amount of training data. On the basis of pre-training, the pre-trained model is fine-tuned with the instruction datasets corresponding to different industry fields to obtain the language models corresponding to different industry fields. The target language model in this embodiment can be any language model corresponding to an industry field obtained by fine-tuning.

[0033] Among them, the target language model outputs a reply text according to the inquiry instruction. The inquiry instruction may include: a question or a prompt word input by the user, or may include an instruction or a prompt word provided by other application programs. After obtaining the inquiry instruction, the target language model can generate an answer related to the question by combining its own training data and knowledge base, and return the answer to the user. Among them, the inquiry instruction and the answer can form a Q&A pair.

[0034] After obtaining the question-answer pair, a scoring model can be used to score the question-answer pair to obtain an evaluation score for the question-answer pair, so as to evaluate the matching degree between the answer generated by the language model and the query instruction according to the evaluation score. Among them, the higher the evaluation score of the question-answer pair, the higher the matching degree between the query instruction and the answer in the question-answer pair, that is, the better the fine-tuning effect of the target language model.

[0035] In this embodiment, the target scoring model is fine-tuned according to the encoder and / or decoder in the pre-trained generative large language model, and can be called a generative scoring model. In the generative large language model, the encoder is used to encode the input sequence into a vector representation. It takes each element of the input sequence as input, extracts and integrates the information of the input sequence through a series of calculations (such as the encoding layer of a recurrent neural network or Transformer), and finally obtains a fixed-length vector representation, also called a context vector or encoding vector. This vector can be regarded as an abstract representation of the input sequence, containing all the information in the sequence. The decoder is used to decode the input data into an output sequence. It takes the encoded vector or other data as input, and then uses a series of calculations to decode the input data into an output sequence.

[0036] Among them, the goal of the generative pre-training model is usually to learn the text generation probability distribution, and it is more suitable for performing tasks related to language generation. Compared with the comprehension scoring model, the generative scoring model makes full use of the text generation ability learned by the base model, and thus can better evaluate the text generated by the target language model.

[0037] After obtaining the pre-trained generative large language model, the structure and related parameters of the encoder and / or decoder can be obtained from the generative large language model, and an initial model can be constructed according to the structure and related parameters of the encoder and / or decoder. By fine-tuning the initial model according to the training data in the target industry field, the target scoring model can be obtained. The training process of the target scoring model will be exemplarily described below.

[0038] In some alternative embodiments, a supervised learning method can be used to train an initial model to obtain a simulated scoring model. Optionally, positive sample question-and-answer pairs and their true evaluation scores, as well as negative sample question-and-answer pairs and their true evaluation scores can be obtained. Among them, the positive sample question-and-answer pairs include instruction samples and answer samples with a high matching degree with the instruction samples, and the negative sample question-and-answer pairs include instruction samples and answer samples with a low matching degree with the instruction samples. After inputting the positive sample question-and-answer pairs into the initial model, the initial model can output the predicted evaluation scores of the positive sample question-and-answer pairs. After inputting the negative sample question-and-answer pairs into the initial model, the initial model can output the predicted evaluation scores of the negative sample question-and-answer pairs. According to the true evaluation scores and predicted evaluation scores of the positive sample question-and-answer pairs and the true evaluation scores and predicted evaluation scores of the negative sample question-and-answer pairs, the scoring loss of the initial model can be calculated. According to the scoring loss, the initial model can be fine-tuned until the loss of the initial model converges, and the converged initial model is output as the target scoring model.

[0039] In some other alternative embodiments, a contrastive learning method can be used to train an initial model to obtain a simulated scoring model. In this implementation, a batch of training data can be obtained, and the training data includes multiple groups of question-and-answer sample pairs. Among them, any group of question-and-answer sample pairs includes: an instruction sample, a positive answer sample of the instruction sample, and multiple negative answer samples. For example, a group of question-and-answer sample pairs can be expressed as: <Instruction: Q; Positive answer sample: R pos ; Negative answer sample: R neg1 (1), R neg (2), …, R neg (k)>, where k represents the total number of negative answer samples in the question-and-answer sample pair. Inputting the multiple groups of question-and-answer sample pairs into the initial model, the predicted evaluation scores of each group of question-and-answer sample pairs are obtained. Among them, the predicted evaluation score of any group of question-and-answer sample pairs includes: the predicted evaluation score of the positive sample pair formed by the instruction sample and its positive answer sample, for example, the evaluation score of the question-and-answer pair <Q, R pos >; and the predicted evaluation scores of multiple negative sample pairs formed by the instruction sample and its multiple negative answer samples, for example, the evaluation score of the question-and-answer pair <Q, R neg1 (1)>, the evaluation score of the question-and-answer pair <Q, R neg (2)>, …, and the evaluation score of the question-and-answer pair <Q, R neg(k). For any pair of question-answer samples, the difference between the predicted evaluation score of the positive sample pair in the question-answer sample pair and the maximum value of the predicted evaluation scores of multiple negative sample pairs can be obtained, and the main loss can be determined according to the first comparison threshold and the difference. The first average value of the predicted evaluation scores of the positive sample pairs included in each of the multiple groups of question-answer sample pairs and the second average value of the predicted evaluation scores of the negative sample pairs included in each of the multiple groups of question-answer sample pairs can be obtained, and the auxiliary loss can be determined according to the difference between the first average value and the second average value and the second comparison threshold. According to the main loss and the auxiliary loss, the scoring loss of the initial model can be determined. According to the scoring loss, the initial model can be fine-tuned until the loss of the initial model converges, and the converged initial model is output as the target scoring model.

[0040] In some embodiments, the following loss function can be used to calculate the scoring loss of the initial model:

[0041] L loss = l main + l auxiliary

[0042]

[0043] l auxiliary = max(0, m2 - (s a - s b ))

[0044] where L loss represents the scoring loss of the initial model, l main represents the main loss, and l auxiliary represents the auxiliary loss. In l main , s pos is the predicted evaluation score of the positive sample pair in a pair of question-answer samples, represents the set of negative sample pairs in a pair of question-answer samples, and s j is the predicted evaluation score of the jth negative sample pair. Among them, m1 represents the first comparison threshold. In a pair of question-answer samples, when the difference between the predicted evaluation score of the positive sample pair and the maximum value of the predicted evaluation scores of the negative samples is greater than m1, the value of l main is 0, and the main loss reaches convergence. Usually, m1 can be set to 0.5 or other set values. That is, when the difference between the predicted evaluation score of the positive sample pair and the maximum value of the predicted evaluation scores of the negative samples is greater than 0.5 or other set values, the main loss reaches convergence. In l auxiliary , s a is the average value of the predicted evaluation scores of all positive sample pairs in a batch of multiple groups of question-answer sample pairs, and s bIt is the average of the predicted evaluation scores of all negative sample pairs in the multiple groups of Q&A sample pairs. m2 represents the second comparison threshold. In a training batch, when the difference between the average of the predicted evaluation scores of all positive sample pairs and the average of the predicted evaluation scores of all negative sample pairs is greater than m2, the value of l auxiliary is 0, and the auxiliary loss reaches convergence. Usually, m2 can be set to 0.5 or other set values. That is, within a batch, when the difference between the average of the predicted evaluation scores of positive sample pairs and the average of the predicted evaluation scores of negative sample pairs is greater than 0.5 or other set values, the auxiliary loss reaches convergence.

[0045] After obtaining the scoring loss of the initial model for the training data of a batch based on the above implementation manner, aiming at the converged scoring loss, fine-tune the initial model until the scoring loss converges to the specified range, and output the converged initial model as the target scoring model.

[0046] After obtaining the Q&A pair, the Q&A pair can be input into the target scoring model so that the target scoring model scores the Q&A pair to obtain the evaluation score of the Q&A pair.

[0047] Optionally, the query instruction and the reply text in the Q&A pair are input into the target scoring model in the form of a text sequence. Among them, the text sequence is composed of multiple ordered text elements (Tokens). A text element refers to the most basic meaningful element obtained by decomposing the text, also called a token. These tokens can be words, punctuation marks, numbers, etc. The multiple text elements obtained by decomposing the query instruction can form a first text sequence, and the multiple text elements obtained by decomposing the reply text can form a second text sequence. Among them, "first" and "second" are only used to distinguish the same or similar description objects, and do not limit the number and order of text elements included in the text sequence.

[0048] Among them, the target scoring model may at least include: a feature transformation module, a masked average pooling layer, and a fully connected layer. The first text sequence and the second text sequence can be input into the feature transformation module in the target scoring model. The feature transformation module can perform feature transformation according to the first text sequence and the second text sequence to obtain the vector sequence of the Q&A pair. Among them, for each input text sequence, the feature transformation module can perform feature transformation on the text sequence to obtain the feature vector corresponding to the text sequence. When multiple text elements are input, the feature transformation module can output the feature vectors corresponding to each of the multiple text elements. The multiple feature vectors are arranged in sequence to form the vector sequence of the Q&A pair.

[0049] After obtaining the vector sequence of the question-answer pair, a pooling operation can be performed on the vector sequence. During the training of the scoring model, the training data is input into the scoring model in batches (batches), and the text lengths within the same batch are not exactly the same. In this case, meaningless values (such as filling with 0) can be supplemented in some texts to make the lengths of the text sequences within a batch the same. That is, there are some meaningless characters in the text sequence input into the scoring model. Therefore, during the pooling operation, it is not necessary to calculate the feature vectors of all text elements in the text sequence. In this embodiment, a masked average pooling layer can be used to perform the pooling operation to eliminate or reduce the influence of the feature vectors of the supplemented values on the calculation results.

[0050] Among them, the masked average pooling layer is used to perform a pooling operation on the sequence data. During the pooling process, a masking weight can be assigned to each input element according to the importance of the element, and these weights are used to pool the input sequence. Among them, important elements will be assigned larger weights, while less important elements will be assigned smaller weights. Based on different weights, the pooling operation can better capture the important information in the input sequence.

[0051] In this embodiment, after the vector sequence of the question-answer pair is input into the masked average pooling layer, the masked average pooling layer can obtain the weight of each input feature vector by performing a convolution operation on the input vector sequence. Then, the average pooling calculation is performed on multiple feature vectors using the weight of each feature vector to obtain the average pooling calculation result.

[0052] After obtaining the average pooling calculation result, a fully connected layer can be used to perform a fully connected calculation on the average pooling calculation result according to the pre-trained weight matrix and bias parameters to obtain the evaluation score of the question-answer pair. Among them, the fully connected layer can be implemented based on an MLP (Multi-Layer Perceptron). In the fully connected layer, the process of performing a fully connected calculation on the input data can be expressed as y = W * x + b, where x is the input average pooling calculation result, and W and b are the pre-trained weight matrix and bias parameters respectively.

[0053] Among them, when the base model of the target language model is different, the structure of the target scoring model can also be different. The optional structures of the scoring model and the scoring process will be exemplarily described below.

[0054] In some optional Example A cases, the base model of the target language model includes an encoder. For example, this base model can be a T5 model. Correspondingly, the feature transformation module in the target scoring model can be fine-tuned according to the encoder in the generative large language model. AsFigure 2 As shown, the target scoring model may at least include: an encoder, a masked average pooling layer, and a fully connected layer connected in sequence.

[0055] In this implementation, the first text sequence and the second text sequence may be concatenated to obtain a third text sequence. Optionally, the method of concatenating the first text sequence and the second text sequence may include: adding a separator between the first text sequence and the second text sequence, or adding a prompt word indicating separation. For example, a prompt word "question" or "query (asking)" indicating a question may be added before the first text sequence, and a prompt word "answer" or "document" indicating an answer may be added between the second text sequence and the first text sequence.

[0056] After the third text sequence is obtained by concatenation, the third text sequence may be input into the encoder in the scoring model, so that the encoder encodes multiple text elements in the third text sequence to obtain feature vectors of each of the multiple text elements in the third text sequence. Among them, the encoder may include multiple hidden layers. A hidden layer refers to each computing layer located between the input layer and the output layer of the encoder and not in contact with external signals. Any hidden layer may include multiple artificial neuron nodes for performing non-linear calculations on the input data to encode the features of the input data into another dimensional space. The encoder may encode multiple text elements in the third text sequence based on multiple hidden layers. Among them, the feature vectors of each of the multiple text elements in the third text sequence may be output by the last hidden layer among the multiple hidden layers of the encoder. Among these multiple hidden layers, the last hidden layer has stronger feature expression ability and can transform the feature information extracted by the previous multiple basic layers into meaningful adjusted expressions, thus facilitating the improvement of subsequent task processing performance.

[0057] According to the feature vectors of each of the multiple text elements in the third text sequence, a vector sequence of the question-answer pair can be obtained. The vector sequence of the question-answer pair may be input into the masked average pooling layer connected to the encoder for subsequent calculations.

[0058] As Figure 2 shown, q and r are respectively used to represent the query instruction and the answer text in the text pair to be scored. For example, q is used to represent the user's question, and r is used to represent the answer text of the target language model; or, q is used to represent the user's question, and r is used to represent the retrieved text segment; or, q is used to represent the user's question, and r is used to represent the question with a known standard answer in the high-frequency question-answer library. q contains n text elements and can be expressed as q = <q1, q2,..., q n >; r contains m text elements and can be expressed as r = <r1, r2,..., r m >. Among them, h represents the feature vector of the last hidden layer output by the feature transformation module (i.e., the encoder).

[0059] As Figure 2 shown, SEP (separator) can be used to splice q and r together. The third text sequence obtained after splicing is: {Question: q1, q2, …, q n ; Answer: r1, r2, …, r m , EoS}, where EoS is the end symbol of the sentence (i.e., End of Sentence), which is used to help the model determine when a sentence ends, so as to better understand and analyze the text content. After inputting the above third text sequence into the encoder, the encoder can encode each text element in the third text sequence to obtain the hidden layer vector (i.e., feature vector) of each text element. According to the hidden layer vectors of multiple text elements, the vector sequence h = <h1, h2…, h N > can be obtained, where N is the total number of elements input to the encoder, and N = n + m.

[0060] In some optional Example B ones, the base model of the target language model includes a decoder. For example, the base model can be a T5 model or a GPT model. Correspondingly, the feature transformation module in the target scoring model can be fine-tuned according to the decoder in the generative large language model. As Figure 3 shown, the target scoring model can at least include: a decoder, a masked average pooling layer, and a fully connected layer connected in sequence.

[0061] In this implementation, the first text sequence and the second text sequence can be spliced to obtain the third text sequence. Optionally, the method of splicing the first text sequence and the second text sequence can include: adding a separator between the first text sequence and the second text sequence, or adding a prompt word indicating separation.

[0062] After splicing to obtain the third text sequence, the third text sequence can be input into the decoder in the scoring model, so that the decoder decodes multiple text elements in the third text sequence to obtain the feature vectors of each of the multiple text elements in the third text sequence. Among them, the decoder can include an input layer, multiple hidden layers, and an output layer. The decoder can decode multiple text elements in the third text sequence based on multiple hidden layers. Among them, the feature vectors of each of the multiple text elements in the third text sequence can be output by the last hidden layer among the multiple hidden layers of the decoder. According to the feature vectors of each of the multiple text elements in the third text sequence, the vector sequence of the question-answer pair can be obtained. The vector sequence of the question-answer pair can be input into the masked average pooling layer connected to the decoder for subsequent calculations.

[0063] As Figure 3As shown, let q and r represent the query instruction and the answer text in the text pair to be scored respectively. SEP can be used to splice q and r together. q contains n text elements, and r contains m text elements. The third text sequence obtained after splicing is: {Question: q1, q2, …, q n ; Answer: r1, r2, …, r m , EoS}. After inputting the above third text sequence into the decoder, the decoder can decode each text element in the third text sequence to obtain the hidden layer vector (i.e., feature vector) of each text element. According to the hidden layer vectors of multiple text elements, a vector sequence h = <h1, h1, h2…, h N > can be obtained, where N is the total number of elements input into the decoder, and N = n + m.

[0064] In some optional Example C ones, the base model of the target language model includes an encoder and a decoder. For example, this base model can be a T5 model. The feature transformation module in the target scoring model is fine-tuned according to the encoder and decoder in the generative large language model. As Figure 4 shown, the target scoring model can at least include: an encoder, a decoder, a masked average pooling layer, and a fully connected layer connected in sequence.

[0065] In this implementation, the first text sequence can be input into the encoder in the scoring model, so that the encoder encodes multiple text elements in the first text sequence to obtain the feature vectors of each of the multiple text elements in the first text sequence. According to these multiple feature vectors, a vector sequence corresponding to the first text sequence is obtained. The second text sequence and the vector sequence corresponding to the first text sequence are input into the decoder in the scoring model, so that the decoder decodes the second text sequence and the vector sequence corresponding to the first text sequence to obtain the vector sequence of the question-answer pair.

[0066] As Figure 4 shown, the first text sequence q = <q1, q2, …, q n > can be input into the encoder to obtain a vector sequence h = <h1, h2…, h n >. The vector sequence h = <h1, h2…, h n > and the second text sequence r = <r1, r2, …, r m , EoS> can be input into the decoder. The decoder can perform decoding calculations according to the input data to obtain the vector sequence h = <h n+1 , h n+2 …, h N > of the question-answer pair <q, r>.

[0067] In some optional Example DAmong them, the base model of the target language model includes an encoder and a decoder. For example, the base model can be a T5 model. The feature transformation module in the target scoring model is fine-tuned according to the encoder and decoder in the generative large language model. As Figure 5 shown, the target scoring model may at least include: an encoder, a decoder, a masked average pooling layer, and a fully connected layer. Among them, the encoder is connected to the decoder and the masked average pooling layer, the decoder is connected to the masked average pooling layer, and the masked average pooling layer is connected to the fully connected layer.

[0068] In this implementation, the first text sequence can be input into the encoder in the scoring model, so that the encoder encodes multiple text elements in the first text sequence to obtain the feature vectors of each of the multiple text elements in the first text sequence. The feature vectors of each of the multiple text elements form a vector sequence corresponding to the first text sequence, hereinafter referred to as the first vector sequence. The second text sequence and the first vector sequence are input into the decoder in the scoring model, so that the decoder decodes the first vector sequence and the second text sequence to obtain a second vector sequence; according to the first vector sequence and the second vector sequence, the vector sequence of the question-answer pair can be obtained. For example, the first vector sequence and the second vector sequence can be concatenated to obtain the vector sequence of the question-answer pair.

[0069] As Figure 5 shown, the first text sequence q = <q1, q2,..., q n > can be input into the encoder to obtain the first vector sequence h = <h1, h2..., h n >. The first vector sequence h = <h1, h2..., h n > and the second text sequence r = <r1, r2,..., r m , EoS> can be input into the decoder to obtain the second vector sequence h = <h n+1 , h n+ 2..., h N >. The first vector sequence h = <h1, h2..., h n > and the second vector sequence h = <h n+1 , h n+ 2..., h N > can be used as the vector sequence of the question-answer pair <q, r> and input into the masked average pooling layer for average pooling calculation.

[0070] In the above embodiments, some model structures and model parameters of the target scoring model and the target language model are from the same base model, which is conducive to realizing knowledge transfer and knowledge sharing based on the structure and parameters of the same base model, thereby improving the scoring performance of the target scoring model and making the scoring results more reliable. On the other hand, using the base model of the target language model to generate the scoring model of the target language model can fully reuse the existing model structures and model parameters, reducing the construction cost of the target scoring model.

[0071] The question-and-answer data processing method provided in this embodiment is mainly used to score the question-and-answer pairs output by the language model to obtain the evaluation scores of the question-and-answer pairs. This question-and-answer data processing method can be used in a variety of different application scenarios, which will be exemplarily described below.

[0072] In some scenarios, this question-and-answer data processing method can be used to screen the text output by the target language model to provide higher-quality question-and-answer results. When using the language model, by setting the parameters of the language model, the language model can output multiple answers to the same question. Based on this question-and-answer data processing method, the text pairs corresponding to the multiple answers can be scored, and the multiple text pairs can be sorted according to the scoring results. Furthermore, according to the scoring results, an answer that is more matched to the question can be selected from the multiple answers.

[0073] Optionally, in this scenario, the inquiry instructions of the at least one question-and-answer pair input to the target scoring model are the same, that is, the at least one question-and-answer pair consists of an inquiry instruction and different reply texts output by the target language model according to the inquiry instruction. After obtaining the respective evaluation scores of the at least one question-and-answer pair based on the target scoring model, the target question-and-answer pair can be screened out from the at least one question-and-answer pair according to the respective evaluation scores of the at least one question-and-answer pair, and a question-and-answer service can be provided according to the target question-and-answer pair. For example, the target question-and-answer pair can be returned to the user, or the target question-and-answer pair can be provided to a downstream application for use.

[0074] In some scenarios, this question-and-answer data processing method can be used to optimize the target scoring model. In this scenario, the scoring loss of the scoring model can be determined according to the respective evaluation scores of the at least one question-and-answer pair and the sample type of the reply text in the at least one question-and-answer pair. The sample type includes: positive sample type or negative sample type. The target scoring model can be optimized according to the scoring loss. The optional implementation manners for determining the scoring loss can refer to the descriptions of the foregoing embodiments and will not be elaborated here.

[0075] In some other scenarios, this Q&A data processing method can be used to evaluate the model performance of a language model. By scoring the text output by the language model, the generation quality and performance of the language model can be evaluated. The higher the quality of the text generated by the language model, the higher the score of the generated text usually is. When the text output by the language model has a high score, it can be considered that the language model has learned how to generate high-quality text.

[0076] In still some other scenarios, this Q&A data processing method can be used to optimize a language model. During the training process of the language model, the evaluation score of the text output by the language model can be obtained. If the evaluation score is low, the performance of the model can be improved by adjusting the model parameters or retraining the model.

[0077] Of course, in addition to the above application scenarios, this Q&A data processing method can also be applied to the model selection scenario. For example, when there are multiple different language models, the Q&A pairs output by different language models can be scored to compare the matching relationship between the language models and different industries and fields.

[0078] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can also be executed by different devices as the execution subject. For example, the execution subject of steps 101 to 104 can be device A; for another example, the execution subject of steps 101 and 102 can be device A, and the execution subject of step 103 can be device B; and so on.

[0079] In addition, in some processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as 101, 102, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0080] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0081] Figure 6 Schematically shows a structural diagram of a server provided by an exemplary embodiment of the present application. This server is applicable to the question-and-answer data processing method provided by the foregoing embodiment. As Figure 6 shown, the server includes: a memory 601, a processor 602, and a communication component 603.

[0082] The memory 601 is used to store computer programs and can be configured to store various other data to support operations on the server. Examples of these data include instructions for any application or method operating on the server.

[0083] The processor 602 is coupled to the memory 601 and is used to execute the computer programs in the memory 601 for: obtaining at least one question-and-answer pair, where any question-and-answer pair consists of an interrogation instruction and a reply text; the reply text is generated by a target language model according to the interrogation instruction; the base model of the target language model is a pre-trained generative large language model; the pre-trained generative large language model at least includes: an encoder and / or a decoder; inputting the at least one question-and-answer pair into a target scoring model to enable the target scoring model to score the at least one question-and-answer pair, obtaining the evaluation scores of the at least one question-and-answer pair respectively; where the target scoring model at least includes a feature transformation module, and the feature transformation module is fine-tuned according to the encoder and / or decoder in the pre-trained generative large language model.

[0084] Optionally, the target scoring model further includes: a masked average pooling layer and a fully connected layer. When the processor 602 inputs the at least one question-and-answer pair into the target scoring model to enable the target scoring model to score the at least one question-and-answer pair, obtaining the evaluation scores of the at least one question-and-answer pair respectively, it includes: obtaining a first text sequence of the interrogation instruction in the question-and-answer pair and a second text sequence of the reply text; inputting the first text sequence and the second text sequence into the feature transformation module in the target scoring model to enable the feature transformation module to perform feature transformation according to the first text sequence and the second text sequence, obtaining a vector sequence of the question-and-answer pair; inputting the vector sequence of the question-and-answer pair into the masked average pooling layer to enable the average pooling layer to perform average pooling calculation on the vector sequence of the question-and-answer pair according to preset weight parameters, obtaining an average pooling calculation result; using the fully connected layer to perform a fully connected calculation on the average pooling calculation result according to a pre-trained weight matrix and bias parameters, obtaining the evaluation score of the question-and-answer pair.

[0085] Optionally, when the processor 602 inputs the first text sequence and the second text sequence into the feature transformation module in the target scoring model, it includes: concatenating the first text sequence and the second text sequence to obtain a third text sequence; inputting the third text sequence into the encoder in the scoring model, so that the encoder performs encoding processing on multiple text elements in the third text sequence based on multiple hidden layers to obtain a vector sequence of the question-and-answer pair.

[0086] Optionally, when the processor 602 inputs the first text sequence and the second text sequence into the feature transformation module in the target scoring model, it includes: concatenating the first text sequence and the second text sequence to obtain a concatenated third text sequence; inputting the third text sequence into the decoder in the scoring model, so that the decoder performs decoding processing on multiple text elements in the third text sequence based on multiple hidden layers to obtain feature vectors of the multiple text elements in the third text sequence respectively; obtaining a vector sequence of the question-and-answer pair according to the feature vectors of the multiple text elements in the third text sequence respectively.

[0087] Optionally, when the processor 602 inputs the first text sequence and the second text sequence into the feature transformation module in the target scoring model, it includes: inputting the first text sequence into the encoder in the scoring model, so that the encoder performs encoding processing on multiple text elements in the first text sequence to obtain a vector sequence corresponding to the first text sequence; inputting the vector sequence corresponding to the first text sequence and the second text sequence into the decoder in the scoring model, so that the decoder performs decoding processing on the vector sequence corresponding to the first text sequence and the second text sequence to obtain a vector sequence of the question-and-answer pair.

[0088] Optionally, when the processor 602 inputs the first text sequence and the second text sequence into the feature transformation module in the target scoring model, it includes: inputting the first text sequence into the encoder in the scoring model, so that the encoder performs encoding processing on multiple text elements in the first text sequence to obtain a first vector sequence corresponding to the first text sequence; inputting the first vector sequence and the second text sequence into the decoder in the scoring model, so that the decoder performs decoding processing on the first vector sequence and the second text sequence to obtain a second vector sequence; obtaining a vector sequence of the question-and-answer pair according to the first vector sequence and the second vector sequence.

[0089] Optionally, the processor 602 is further configured to: obtain training data; the training data includes multiple groups of question-and-answer sample pairs; any group of question-and-answer sample pairs includes: an instruction sample, a positive answer sample of the instruction sample, and multiple negative answer samples; input the multiple groups of question-and-answer sample pairs into an initial model to obtain the predicted evaluation scores of the multiple groups of question-and-answer sample pairs respectively; the initial model includes: an encoder and / or a decoder in the pre-trained generative large language model; the predicted evaluation score of any group of question-and-answer sample pairs includes: the predicted evaluation score of a positive sample pair formed by the instruction sample and its positive answer sample, and the predicted evaluation scores of multiple negative sample pairs formed by the instruction sample and its multiple negative answer samples respectively; for any group of question-and-answer sample pairs, obtain the difference between the predicted evaluation score of the positive sample pair in the group of question-and-answer sample pairs and the maximum value among the predicted evaluation scores of the multiple negative sample pairs; determine a main loss according to a first comparison threshold and the difference; obtain a first average value of the predicted evaluation scores of the positive sample pairs included in the multiple groups of question-and-answer sample pairs respectively, and a second average value of the predicted evaluation scores of the negative sample pairs included in the multiple groups of question-and-answer sample pairs respectively; determine an auxiliary loss according to the difference between the first average value and the second average value and a second comparison threshold; determine a scoring loss of the initial model according to the main loss and the auxiliary loss; and fine-tune the initial model with the goal of converging the scoring loss until the scoring loss converges to a specified range, and output the converged initial model as the target scoring model.

[0090] Optionally, the inquiry instructions of the at least one question-and-answer pair are the same; after obtaining the evaluation scores of the at least one question-and-answer pair respectively, the processor 602 is further configured to: screen out a target question-and-answer pair from the at least one question-and-answer pair according to the evaluation scores of the at least one question-and-answer pair, and provide a question-and-answer service according to the target question-and-answer pair; or determine a scoring loss of the scoring model according to the evaluation scores of the at least one question-and-answer pair and the sample type of the answer text in the at least one question-and-answer pair; and optimize the target scoring model according to the scoring loss; the sample type includes: a positive sample type or a negative sample type.

[0091] Further, as Figure 6 shown, the server further includes: a power supply component 604 and other components. Figure 6 Only some components are schematically shown, and it does not mean that the server only includes Figure 6 the components shown.

[0092] Among them, the memory 601 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0093] Among them, the communication component 603 is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as Wi-Fi (wireless network communication technology), 2G (such as Global System for Mobile Communications (GSM), etc.), 3G (such as Wideband Code Division Multiple Access (WCDMA)), 4G (such as Long Term Evolution (LTE), etc.), 4G+ (such as LTE-Advanced (LTE-A), etc.) or 5G (5th Generation Mobile Communication Technology), or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can be implemented based on near field communication (NFC) technology, radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra wide band (UWB) technology, Bluetooth (BT) technology and other technologies.

[0094] Among them, the power supply component 604 is used to provide power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0095] In this embodiment, after obtaining the response text generated by the target language model for the query instruction, the target scoring model can be used to score the question-and-answer pair formed by the query instruction and the response text. Among them, the base model of the target language model is a pre-trained generative large language model, and the feature transformation module in the target scoring model is fine-tuned according to the encoder and / or decoder in the pre-trained generative large language model. In this implementation, some model structures and model parameters of the target scoring model and the target language model come from the same base model, which is conducive to realizing knowledge transfer and knowledge sharing based on the structure and parameters of the same base model, thereby improving the scoring performance of the target scoring model and making the scoring result more reliable. On the other hand, using the base model of the target language model to generate the scoring model of the target language model can fully reuse the existing model structures and model parameters, reducing the construction cost of the target scoring model.

[0096] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement the steps executable by the server in the above method embodiment.

[0097] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM (Compact Disc Read-Only Memory), optical storage, etc.) containing computer-usable program code.

[0098] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0099] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the functions specified in one or more of the procedures Figure 1 one or more of the procedures and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0100] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the procedures Figure 1 one or more of the procedures and / or blocks Figure 1 specified in one or more of the blocks or blocks.

[0101] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), an input / output interface, a network interface, and memory.

[0102] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium.

[0103] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (Parallel Random Access Machine, PRAM), static random access memory (SRAM), dynamic random access memory (Dynamic Random Access Memory, DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (Digital Video Disc, DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0104] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the said element.

[0105] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for processing question-and-answer data, characterized in that, Including: Obtaining at least one question-and-answer pair, where any question-and-answer pair consists of an inquiry instruction and a response text; the response text is generated by a target language model according to the inquiry instruction; The base model of the target language model is a pre-trained generative large language model; the pre-trained generative large language model at least includes: an encoder and / or a decoder; Inputting the at least one question-and-answer pair into a target scoring model, so that the target scoring model scores the at least one question-and-answer pair to obtain the evaluation scores of the at least one question-and-answer pair respectively; wherein, the target scoring model at least includes a feature transformation module, and the feature transformation module is fine-tuned according to the encoder and / or decoder in the pre-trained generative large language model.

2. The method according to claim 1, wherein The target scoring model further includes: a masked average pooling layer and a fully connected layer; inputting the at least one question-and-answer pair into the target scoring model, so that the target scoring model scores the at least one question-and-answer pair to obtain the evaluation scores of the at least one question-and-answer pair respectively, including: Obtaining a first text sequence of the inquiry instruction in the question-and-answer pair and a second text sequence of the response text; Inputting the first text sequence and the second text sequence into the feature transformation module in the target scoring model, so that the feature transformation module performs feature transformation according to the first text sequence and the second text sequence to obtain a vector sequence of the question-and-answer pair; Inputting the vector sequence of the question-and-answer pair into the average pooling layer, so that the average pooling layer performs average pooling calculation on the vector sequence of the question-and-answer pair according to preset weight parameters to obtain an average pooling calculation result; Using the fully connected layer to perform a fully connected calculation on the average pooling calculation result according to a pre-trained weight matrix and bias parameters to obtain the evaluation score of the question-and-answer pair.

3. The method according to claim 2, characterized in that, Inputting the first text sequence and the second text sequence into the feature transformation module in the target scoring model, including: Concatenating the first text sequence and the second text sequence to obtain a third text sequence; Inputting the third text sequence into the encoder in the scoring model, so that the encoder encodes multiple text elements in the third text sequence to obtain a vector sequence of the question-and-answer pair.

4. The method according to claim 2, characterized in that, Inputting the first text sequence and the second text sequence into the feature transformation module in the target scoring model, including: Concatenating the first text sequence and the second text sequence to obtain a concatenated third text sequence; Inputting the third text sequence into the decoder in the scoring model, so that the decoder decodes multiple text elements in the third text sequence to obtain the feature vectors of the multiple text elements in the third text sequence respectively; Obtaining a vector sequence of the question-and-answer pair according to the feature vectors of the multiple text elements in the third text sequence.

5. The method according to claim 2, wherein Inputting the first text sequence and the second text sequence into the feature transformation module in the target scoring model, including: Input the first text sequence into the encoder in the scoring model, so that the encoder encodes multiple text elements in the first text sequence to obtain a vector sequence corresponding to the first text sequence; Input the vector sequence corresponding to the first text sequence and the second text sequence into the decoder in the scoring model, so that the decoder decodes the vector sequence corresponding to the first text sequence and the second text sequence to obtain a vector sequence of the question-answer pair.

6. The method according to claim 2, characterized in that, Input the first text sequence and the second text sequence into the feature transformation module in the target scoring model, including: Input the first text sequence into the encoder in the scoring model, so that the encoder encodes multiple text elements in the first text sequence to obtain a first vector sequence corresponding to the first text sequence; Input the first vector sequence and the second text sequence into the decoder in the scoring model, so that the decoder decodes the first vector sequence and the second text sequence to obtain a second vector sequence; Obtain the vector sequence of the question-answer pair according to the first vector sequence and the second vector sequence.

7. The method according to any one of claims 1 to 6, characterized in that, Obtain training data; the training data includes multiple groups of question-answer sample pairs; any group of question-answer sample pairs includes: an instruction sample, a positive answer sample of the instruction sample, and multiple negative answer samples; Input the multiple groups of question-answer sample pairs into the initial model to obtain the predicted evaluation scores of the multiple groups of question-answer sample pairs respectively; the initial model includes: the encoder and / or decoder in the pre-trained generative large language model; the predicted evaluation score of any group of question-answer sample pairs includes: the predicted evaluation score of the positive sample pair formed by the instruction sample and its positive answer sample, and the predicted evaluation scores of the multiple negative sample pairs formed by the instruction sample and its multiple negative answer samples respectively; For any group of question-answer sample pairs, obtain the difference between the predicted evaluation score of the positive sample pair in the group of question-answer sample pairs and the maximum value of the predicted evaluation scores of the multiple negative sample pairs respectively; Determine the main loss according to the first comparison threshold and the difference; Obtain the first average value of the predicted evaluation scores of the positive sample pairs included in the multiple groups of question-answer sample pairs respectively, and the second average value of the predicted evaluation scores of the negative sample pairs included in the multiple groups of question-answer sample pairs respectively; Determine the auxiliary loss according to the difference between the first average value and the second average value and the second comparison threshold; Determine the scoring loss of the initial model according to the main loss and the auxiliary loss; Taking the convergence of the scoring loss as the goal, fine-tune the initial model until the scoring loss converges to the specified range, and output the converged initial model as the target scoring model.

8. The method according to any one of claims 1 to 6, characterized in that The inquiry instructions of the at least one question-answer pair are the same; After obtaining the evaluation scores of the at least one question-answer pair respectively, it further includes: According to the evaluation scores of the at least one question-answer pair respectively, screen out the target question-answer pair from the at least one question-answer pair and provide question-answer services according to the target question-answer pair; or, Determine the scoring loss of the scoring model according to the respective evaluation scores of the at least one question-and-answer pair and the sample type of the response text in the at least one question-and-answer pair; and optimize the target scoring model according to the scoring loss; the sample type includes: positive sample type or negative sample type.

9. A server, characterized in that, Including: A memory and a processor; The memory is used to store one or more computer instructions; The processor is used to execute the one or more computer instructions for: executing the steps in the method according to any one of claims 1-8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement the question-and-answer data processing method according to any one of claims 1-8.