Question and answer quality evaluation model training method and question and answer quality evaluation method
By constructing a question-and-answer quality assessment model and using multiple pre-set review models for answer preference labeling and parameter optimization, the problem of balancing performance and transparency in question-and-answer quality assessment models is solved, the robustness and generalization ability of the model are improved, and more accurate question-and-answer quality assessment is achieved.
Patent Information
- Application Number
- CN202510973117.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-07
AI Technical Summary
Existing question-answering quality assessment models face challenges in balancing performance and training transparency. Traditional automated metrics struggle to accurately reflect the linguistic coherence and sentiment alignment of multi-turn dialogues. Closed-source models lack transparency and repeatability, while open-source models fall short in terms of robustness and scalability.
By constructing a question-answering quality assessment model, multiple pre-defined review models are used to label the question-answer pair dataset with answer preferences. Reliability parameters to be estimated are introduced, and noise probability and preference probability functions of the review model are constructed. A joint likelihood function is constructed, and the model parameters are optimized to improve the robustness and generalization ability of the model.
It improves the accuracy and transparency of the question-answering quality assessment model in the absence of absolute labels, reduces the impact of noise annotation, and enhances the credibility of the model output and the true reflection of dialogue quality.
Smart Images

Figure CN120911635A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data evaluation, in particular to a question and answer quality evaluation model training method and a question and answer quality evaluation method. BACKGROUND
[0002] In recent years, with the wide application of large language models (LLM) in dialogue systems, how to effectively evaluate the quality of the generated replies has become a research hotspot. Traditional automatic indicators (such as BLEU, ROUGE, and BERTScore) are limited by semantic understanding ability and are difficult to accurately reflect key dimensions such as language coherence, instruction compliance, and sentiment alignment in multi-round dialogue. To overcome this limitation, related technologies use advanced language models for generative scoring or preference judgment, including closed-source models and open-source models. However, closed-source models lack transparency in training data, making it difficult to guarantee fairness and reproducibility, and there are many problems in controllability and cost. Open-source models have shortcomings in robustness and scalability, and these models mainly focus on single-round dialogue evaluation.
[0003] There is currently no effective solution to the problem of quality evaluation models in related technologies that cannot balance performance and training transparency. SUMMARY
[0004] Therefore, it is necessary to provide a question and answer quality evaluation model training method and a question and answer quality evaluation method to solve the above technical problems.
[0005] In a first aspect, the present application provides a question and answer quality evaluation model training method, which comprises:
[0006] determining a question and answer quality evaluation model to be trained and a question and answer pair preference data set; wherein the question and answer pair preference data set is determined by marking the answers of a question and answer pair data set using multiple preset review models; and each preset review model has a corresponding to-be-estimated reliability parameter;
[0007] For each question and answer pair in the question and answer pair preference data set, a review model noise probability function corresponding to each question and answer pair is constructed according to the to-be-estimated reliability parameters of the multiple preset review models and the marking results of each question and answer pair;
[0008] According to the predicted quality score relationship of each question and answer pair determined by the question and answer quality evaluation model to be trained, a preference probability function corresponding to each question and answer pair is constructed;
[0009] construct a joint likelihood function according to the review model noise probability function and the preference probability function corresponding to each of the question-answer pairs; the joint likelihood function is used to quantify a probability of observing the question-answer pair preference data set under the trainable model parameters corresponding to the question-answer quality evaluation model to be trained and the to-be-estimated reliability parameters corresponding to each of the preset review models;
[0010] construct a target loss function corresponding to the question-answer quality evaluation model to be trained according to the joint likelihood function;
[0011] train the question-answer quality evaluation model to be trained according to the question-answer pair preference data set, and iteratively update the trainable model parameters and the to-be-estimated reliability parameters corresponding to each of the preset review models in the training process until the target loss function converges, to obtain a question-answer quality evaluation model.
[0012] In one of the embodiments, the to-be-estimated reliability parameters include to-be-estimated true positive rates; the annotation results of each of the question-answer pairs include a plurality of preference labels; the plurality of preference labels correspond one-to-one to the plurality of preset review models; the review model noise probability function includes a first noise probability function; and the constructing of the review model noise probability function corresponding to each of the question-answer pairs according to the to-be-estimated reliability parameters corresponding to the plurality of preset review models and the annotation results of each of the question-answer pairs includes:
[0013] if the latent true preference label corresponding to the question-answer pair is a first preference label, constructing a first judgment likelihood function of each of the preset review models under the first preference label according to the to-be-estimated true positive rate corresponding to each of the preset review models and the preference label;
[0014] multiplying the first judgment likelihood functions corresponding to all the preset review models to obtain a first noise probability function corresponding to each of the question-answer pairs; the first noise probability function is used to quantify a probability of observing the plurality of preference labels corresponding to the question-answer pair under the first preference label.
[0015] In one of the embodiments, the to-be-estimated reliability parameters include to-be-estimated true negative rates; the review model noise probability function includes a second noise probability function; and the constructing of the review model noise probability function corresponding to each of the question-answer pairs according to the to-be-estimated reliability parameters corresponding to the plurality of preset review models and the annotation results of each of the question-answer pairs includes:
[0016] if the potential real preference label corresponding to the question-answer pair is a second preference label, constructing a second judging likelihood function of each of the preset review models under the second preference label according to the respective to-be-estimated true negative rate and the preference label of each of the preset review models;
[0017] multiplying the second judging likelihood functions corresponding to all the preset review models to obtain a second noise probability function corresponding to each of the question-answer pairs; the second noise probability function is used to quantify a probability of observing multiple preference labels corresponding to the question-answer pair under the second preference label.
[0018] In one of the embodiments, each of the question-answer pairs includes a first to-be-compared question-answer pair and a second to-be-compared question-answer pair; the preference probability function includes a first preference probability function and a second preference probability function; and the constructing of the preference probability function corresponding to each of the question-answer pairs according to the predicted quality score relationship determined by the to-be-trained question-answer quality evaluation model includes:
[0019] for each of the question-answer pairs, determining a predicted quality difference relationship of the question-answer pair according to a first predicted quality score relationship of the first to-be-compared question-answer pair and a second predicted quality score relationship of the second to-be-compared question-answer pair determined by the to-be-trained question-answer quality evaluation model;
[0020] if the potential real preference label corresponding to the question-answer pair is a first preference label, constructing a first preference probability function corresponding to each of the question-answer pairs according to the respective predicted quality difference relationship of each of the question-answer pairs; the first preference probability function is used to quantify a probability that the quality of the second to-be-compared question-answer pair is higher than that of the first to-be-compared question-answer pair;
[0021] if the potential real preference label corresponding to the question-answer pair is a second preference label, constructing a second preference probability function corresponding to each of the question-answer pairs according to the respective predicted quality difference relationship of each of the question-answer pairs; the second preference probability function is used to quantify a probability that the quality of the first to-be-compared question-answer pair is higher than that of the second to-be-compared question-answer pair.
[0022] In one of the embodiments, the review model noise probability function includes a first noise probability function and a second noise probability function; the preference probability function includes a first preference probability function and a second preference probability function; and the constructing of the joint likelihood function according to the review model noise probability function and the preference probability function corresponding to all the question-answer pairs includes:
[0023] for each of the question-answer pairs, multiplying the first noise probability function and the first preference probability function corresponding to the question-answer pair to construct a first joint probability function;
[0024] multiplying the second noise probability function corresponding to the question-answer pair and the second preference probability function to construct a second joint probability function;
[0025] determining a sum of the first joint probability function and the second joint probability function as an intermediate joint likelihood function corresponding to the question-answer pair;
[0026] multiplying the intermediate joint likelihood functions corresponding to all groups of question-answer pairs to construct a joint likelihood function.
[0027] In one embodiment, constructing, according to the joint likelihood function, a target loss function corresponding to the to-be-trained question-answer quality evaluation model comprises:
[0028] taking a logarithm of the joint likelihood function and then taking a negative value to obtain a negative log-likelihood function;
[0029] determining the negative log-likelihood function as the target loss function corresponding to the to-be-trained question-answer quality evaluation model.
[0030] In one embodiment, the to-be-trained question-answer quality evaluation model comprises a pre-trained embedding module and a to-be-trained quality scorer; the trainable model parameters comprise model parameters corresponding to the to-be-trained quality scorer; training the to-be-trained question-answer quality evaluation model according to the question-answer pair preference data set and iteratively updating the trainable model parameters and the to-be-estimated reliability parameters corresponding to each of the preset review models in the training process until the target loss function converges to obtain a question-answer quality evaluation model, comprising:
[0031] training the to-be-trained quality scorer according to the question-answer pair preference data set while keeping the model parameters of the pre-trained embedding module unchanged, and iteratively updating the model parameters corresponding to the to-be-trained quality scorer and the to-be-estimated reliability parameters corresponding to each of the preset review models in the training process until the target loss function converges to obtain a trained quality scorer;
[0032] obtaining a question-answer quality evaluation model according to the pre-trained embedding module and the trained quality scorer.
[0033] In a second aspect, the present application provides a question-answer quality evaluation method, comprising:
[0034] obtaining a first target question-answer pair and a second target question-answer pair;
[0035] The first target question and answer pair and the second target question and answer pair are respectively evaluated in quality by using the question and answer quality evaluation model, and a first target quality score corresponding to the first target question and answer pair and a second target quality score corresponding to the second target question and answer pair are obtained.
[0036] In a third aspect, the present application provides a question and answer quality evaluation model training device, the device comprising:
[0037] A determination module is configured to determine a question and answer quality evaluation model to be trained and a question and answer pair preference data set. The question and answer pair preference data set is determined by marking the answer preference of a question and answer pair data set through a plurality of preset review models. Each preset review model has a corresponding to-be-estimated reliability parameter.
[0038] A first construction module is configured to, for each group of question and answer pairs in the question and answer pair preference data set, construct a review model noise probability function corresponding to each group of question and answer pairs according to a plurality of to-be-estimated reliability parameters corresponding to a plurality of preset review models and the marking result of each group of question and answer pairs.
[0039] A second construction module is configured to construct a preference probability function corresponding to each group of question and answer pairs according to a predicted quality score relationship of each group of question and answer pairs determined by the question and answer quality evaluation model to be trained.
[0040] A third construction module is configured to construct a joint likelihood function according to the review model noise probability function and the preference probability function corresponding to all groups of question and answer pairs. The joint likelihood function is used to quantify the probability of observing the question and answer pair preference data set under the trainable model parameter corresponding to the question and answer quality evaluation model to be trained and the to-be-estimated reliability parameter corresponding to each preset review model.
[0041] A fourth construction module is configured to construct a target loss function corresponding to the question and answer quality evaluation model to be trained according to the joint likelihood function.
[0042] A training module is configured to train the question and answer quality evaluation model to be trained according to the question and answer pair preference data set, and iteratively update the trainable model parameter and the to-be-estimated reliability parameter corresponding to each preset review model in the training process until the target loss function converges, so as to obtain a question and answer quality evaluation model.
[0043] In a fourth aspect, the present application provides a computer device comprising a memory and a processor. The memory stores a computer program, and the processor implements the steps of the method described above when executing the computer program.
[0044] In a fifth aspect, the present application provides a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the steps of the method as described above.
[0045] In a sixth aspect, the present application provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method as described above.
[0046] The above question and answer quality evaluation model training method and question and answer quality evaluation method determine a to-be-trained question and answer quality evaluation model and a question and answer pair preference data set; wherein the question and answer pair preference data set is determined by marking the answer preference of a question and answer pair data set through a plurality of preset review models; each preset review model has a corresponding to-be-estimated reliability parameter; for each group of question and answer pairs in the question and answer pair preference data set, a review model noise probability function corresponding to each group of question and answer pairs is constructed according to a plurality of to-be-estimated reliability parameters corresponding to a plurality of preset review models and the marking result of each group of question and answer pairs, which realizes the display modeling of the preference marking deviation of the preset review model, effectively reduces the influence of noise marking on model training, and helps to more accurately depict the observed question and answer pair preference data set; a preference probability function corresponding to each group of question and answer pairs is constructed according to a predicted quality score relationship of each group of question and answer pairs determined by the to-be-trained question and answer quality evaluation model, which effectively enhances the credibility of the model output, so that the model can better reflect the real difference in dialogue quality; a joint likelihood function is constructed according to the review model noise probability function and the preference probability function corresponding to all groups of question and answer pairs; the joint likelihood function is used to quantify the probability of observing the question and answer pair preference data set under the trainable model parameters corresponding to the to-be-trained question and answer quality evaluation model and the to-be-estimated reliability parameters corresponding to each preset review model; a target loss function corresponding to the to-be-trained question and answer quality evaluation model is constructed according to the joint likelihood function, which can ensure that the model parameters better fit the real observation distribution during training; the to-be-trained question and answer quality evaluation model is trained according to the question and answer pair preference data set, and the trainable model parameters and the to-be-estimated reliability parameters corresponding to each preset review model are iteratively updated during the training process until the target loss function converges, obtaining the question and answer quality evaluation model, which realizes the joint optimization of the trainable model parameters and the to-be-estimated reliability parameters of the preset review model in the absence of absolute labels, and improves the robustness, generalization ability and training transparency of the model. BRIEF DESCRIPTION OF DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description only represent some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0048] Figure 1 An application environment diagram of the question and answer quality evaluation model training method in one embodiment;
[0049] Figure 2 A flowchart of the question and answer quality evaluation model training method in one embodiment;
[0050] Figure 3 A flowchart of the first noise probability function construction step in one embodiment;
[0051] Figure 4 A flowchart of the second noise probability function construction step in one embodiment;
[0052] Figure 5 A flowchart of the preference probability function construction step in one embodiment;
[0053] Figure 6 A flowchart of the joint likelihood function construction step in one embodiment;
[0054] Figure 7 A model performance comparison diagram of one specific embodiment;
[0055] Figure 8 A structural block diagram of the question and answer quality evaluation model training device in one embodiment;
[0056] Figure 9 An internal structure diagram of the computer device in one embodiment. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0058] Due to the complexity and multidimensionality of dialogue interactions, evaluating the quality of dialogues generated by Large Language Models (LLMs) faces numerous challenges. While LLMs have made significant progress in evaluating single-turn dialogues in recent years, assessing their performance in multi-turn dialogue scenarios remains extremely difficult, particularly in evaluating key capabilities such as instruction following, self-consistency, and sentiment alignment. Traditional automated evaluation metrics (such as BLEU, ROUGE, and BERTScore) rely on fixed vocabulary or semantic overlap, often failing to effectively reflect the flexibility of human-perceived natural language and the rich semantics inherent in multi-turn dialogues. Furthermore, these metrics typically require external reference answers to evaluate knowledge-based responses, limiting their application in scenarios lacking sufficient reference information.
[0059] In recent years, the evaluation paradigm of "LLM as reviewer" has received widespread attention. State-of-the-art LLMs are used as generative evaluators to assess answer quality through individual scoring and pairwise comparisons. Proprietary LLMs excel in evaluation speed and consistency with human evaluation, but due to the lack of transparency in training data, fairness and repeatability are difficult to guarantee, and there are also issues with controllability and cost. As an alternative, open, transparent, and controllable LLM evaluators have recently emerged, but they still have shortcomings in robustness and scalability. Furthermore, when using LLMs as reviewers, problems such as self-bias bias, score compression, high variance, cue sensitivity, and tolerance bias are still prevalent. Based on these issues, this application proposes a training method for a question-answering quality assessment model.
[0060] The question-answering quality assessment model training method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located on a cloud or other network server. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, and tablets. Server 104 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services.
[0061] In one embodiment, such as Figure 2 As shown, Figure 2 This is a flowchart illustrating a question-answering quality assessment model training method in one embodiment. This embodiment uses the method applied to a terminal as an example; it is understood that the method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0062] Step S201, determine the question and answer quality evaluation model to be trained and the question and answer pair preference data set.
[0063] The question and answer quality evaluation model to be trained includes a pre-trained embedding module and a quality scorer to be trained; the pre-trained embedding module refers to an embedding model that is pre-trained and whose parameters are not updated again, and is used to map an input question and answer pair to an embedding space and output a corresponding embedding vector; the quality scorer to be trained is used to further map the embedding vector to a quality space and output a corresponding quality score of the question and answer pair.
[0064] In an exemplary embodiment, the pre-trained embedding module can be but is not limited to the Llama-3 pre-trained RLHFlow / ArmoRM-Llama3-8B-v0.1 model, which is not specifically limited here. The quality scorer to be trained can be but is not limited to a trainable multi-layer perceptron quality scorer, which is not specifically limited here.
[0065] The question and answer pair preference data set is determined by a plurality of preset review models performing answer preference annotation on a question and answer pair data set; each preset review model has a corresponding to-be-estimated reliability parameter.
[0066] The preset review model refers to a large language model that has strong dialogue quality judgment ability and is fully pre-trained, and has high accuracy and consistency in generative evaluation tasks. The preset review model can include but is not limited to GPT, DeepSeek, Grok, Claude, Gemin, etc., which are not specifically limited here. It can be understood that the preset review model participates in the preference annotation process of the question and answer pair data set as an automatic reviewer.
[0067] It should be noted that different preset review models may exhibit different judgment tendencies in the answer preference annotation process. Therefore, the embodiment introduces a corresponding to-be-estimated reliability parameter for each preset review model to model its accuracy and consistency in preference judgment, thereby effectively controlling the annotation noise.
[0068] Exemplarily, let the question and answer pair data set be , containing N groups of question and answer pairs , is the first to-be-compared question and answer pair in the ith group of question and answer pairs; is the second to-be-compared question and answer pair in the ith group of question and answer pairs; each group of question and answer pairs is annotated by M preset review models. Each preset review model j outputs a relative answer preference score , The results can be categorized into the following three cases: The pre-defined evaluation model j considers the first question-and-answer pair A to be superior to the second question-and-answer pair B, meaning the first question-and-answer pair A wins. The pre-defined evaluation model j considers the second question-and-answer pair B to be superior to the first question-and-answer pair A, meaning the second question-and-answer pair B wins. The pre-defined evaluation model j indicates that the first question-and-answer pair A and the second question-and-answer pair B have the same performance.
[0069] The first and second question-and-answer pairs to be compared can be combinations of questions and answers in a single-turn dialogue or complete interaction sequences in a multi-turn dialogue context; no specific limitations are imposed here.
[0070] It is understood that the goal of this embodiment in labeling the question-answer pair dataset with answer preferences through multiple preset review models is to learn a lightweight question-answer quality assessment model on the question-answer pair preference dataset to evaluate the quality of individual question-answer pairs.
[0071] Step S202: For each question-answer pair in the question-answer pair preference dataset, construct the noise probability function of the corresponding review model for each question-answer pair based on the multiple reliability parameters to be estimated corresponding to multiple preset review models and the annotation results of each question-answer pair.
[0072] Each question-and-answer pair is labeled with multiple preference tags; each preference tag corresponds one-to-one with a pre-set review model.
[0073] The preference labels include at least a first preference label and a second preference label; wherein, the first preference label is used to indicate that the quality of the second question-and-answer pair to be compared is higher than that of the first question-and-answer pair to be compared; the second preference label is used to indicate that the quality of the first question-and-answer pair to be compared is higher than that of the second question-and-answer pair to be compared; for example, the first preference label is... The second preference label is .
[0074] In one embodiment, the preference label further includes a third preference label; the third preference label is used to characterize that the quality of the second question-and-answer pair to be compared is equal to that of the first question-and-answer pair to be compared, i.e. It should be noted that since the third preference label cannot provide effective positive and negative sample information, the case where the preference label is the third preference label should be excluded during the training phase to avoid introducing noise and uncertainty.
[0075] The to-be-estimated reliability parameter can include, but is not limited to, a to-be-estimated true positive rate and a to-be-estimated true negative rate; the to-be-estimated true positive rate, that is, a hit rate, refers to a probability that the preference label given by the preset review model is the first preference label on the premise that the latent true preference label corresponding to the question-answer pair is the first preference label; and the to-be-estimated true negative rate, that is, a correct rejection rate, refers to a probability that the preference label given by the preset review model is the second preference label on the premise that the latent true preference label corresponding to the question-answer pair is the second preference label.
[0076] The latent true preference label corresponding to the question-answer pair refers to a latent variable reflecting the quality relationship between the first to-be-compared question-answer pair and the second to-be-compared question-answer pair; it can be understood that the latent true preference label cannot be directly observed.
[0077] For example, the to-be-estimated true positive rate of the jth preset review model is denoted as The to-be-estimated true positive rate is defined as: ; wherein represents that the latent true preference label corresponding to the question-answer pair is the first preference label, represents that the preference label given by the jth preset review model is the first preference label; the to-be-estimated true negative rate of the jth preset review model is denoted as The to-be-estimated true negative rate is defined as: ; wherein represents that the latent true preference label corresponding to the question-answer pair is the second preference label; represents that the preference label given by the jth preset review model is the second preference label.
[0078] The review model noise probability function is used to quantify a probability of observing the labeled result corresponding to the question-answer pair under different preference labels; it can be understood that the review model noise probability function is used to describe noise or uncertainty in the labeling behavior of the preset review model.
[0079] The review model noise probability function includes a first noise probability function and a second noise probability function; wherein; the first noise probability function is used to quantify a probability of observing the plurality of preference labels corresponding to the question-answer pair under the first preference label; and the second noise probability function is used to quantify a probability of observing the plurality of preference labels corresponding to the question-answer pair under the second preference label.
[0080] For example, the first noise probability function is defined as The second noise probability function is defined as ; wherein , indicating that the corresponding latent true preference label of the ith set of question and answer pairs is the first preference label; , indicating that the corresponding latent true preference label of the ith set of question and answer pairs is the second preference label; , indicating the preference label given by the first preset review model when annotating the ith set of question and answer pairs; the same applies to the rest, which will not be repeated here. , representing the true positive rate to be estimated corresponding to the preset review model; , representing the true negative rate to be estimated corresponding to the preset review model.
[0081] Step S203, according to the prediction quality score relationship of each set of question and answer pairs determined by the question and answer quality evaluation model to be trained, the corresponding preference probability function of each set of question and answer pairs is constructed.
[0082] , the prediction quality score relationship is used to represent the prediction quality score output by the question and answer quality evaluation model to be trained for the first to-be-compared question and answer pair and the second to-be-compared question and answer pair in the same set of question and answer pairs. The prediction quality score relationship includes a first prediction quality score relationship and a second prediction quality score relationship; the first prediction quality score relationship corresponds to the first to-be-compared question and answer pair; and the second prediction quality score relationship corresponds to the second to-be-compared question and answer pair.
[0083] , the preference probability function includes a first preference probability function and a second preference probability function; the first preference probability function is used to quantify the probability that the quality of the second to-be-compared question and answer pair is higher than that of the first to-be-compared question and answer pair; and the second preference probability function is used to quantify the probability that the quality of the first to-be-compared question and answer pair is higher than that of the second to-be-compared question and answer pair. In an exemplary embodiment, the preference probability function can be a cumulative distribution function of a standard normal distribution.
[0084] Exemplarily, it is assumed that , indicating the quality score of the to-be-compared question and answer pair x; in order to model the uncertainty in the quality of the question and answer pair, this embodiment adopts Thurstone’s Case V model for preference probability modeling; specifically, it is assumed that obeys a Gaussian distribution with as the mean and a constant standard deviation, i.e. ; under this assumption, according to the first prediction quality score relationship corresponding to the first to-be-compared question and answer pair A and the second prediction quality score relationship corresponding to the second to-be-compared question and answer pair B, the corresponding preference probability function of each set of question and answer pairs, i.e., the first preference probability function and the second preference probability function, can be constructed; wherein, , indicating the pre-trained embedding module; , indicating the quality scorer to be trained.
[0085] Step S204, constructing a joint likelihood function according to the review model noise probability function and the preference probability function corresponding to all group question-answer pairs.
[0086] wherein the joint likelihood function is used to quantify the probability of observing the question-answer pair preference data set under the trainable model parameters of the question-answer quality evaluation model to be trained and the respective reliability parameters of each of the preset review models.
[0087] wherein the trainable model parameters include the model parameters corresponding to the quality scorer to be trained.
[0088] Step S205, constructing a target loss function corresponding to the question-answer quality evaluation model to be trained according to the joint likelihood function.
[0089] Step S206, training the question-answer quality evaluation model to be trained according to the question-answer pair preference data set, and iteratively updating the trainable model parameters and the respective reliability parameters of each of the preset review models in the training process until the target loss function converges, obtaining the question-answer quality evaluation model.
[0090] In an exemplary embodiment, constructing a target loss function corresponding to the question-answer quality evaluation model to be trained according to the joint likelihood function includes: taking the logarithm of the joint likelihood function and then taking the negative value to obtain a negative log-likelihood function; and determining the negative log-likelihood function as the target loss function corresponding to the question-answer quality evaluation model to be trained.
[0091] Specifically, the joint likelihood function is denoted as wherein is the question-answer pair preference data set; is the model parameter corresponding to the quality scorer to be trained; is the reliability parameter to be estimated corresponding to the preset review model; the negative log-likelihood function obtained by taking the logarithm of the joint likelihood function and then taking the negative value is ; further, the target loss function corresponding to the question-answer quality evaluation model to be trained is According to the question-answer pair preference data set, the question-answer quality evaluation model to be trained is trained, and in the training process, the target loss function is optimized by stochastic gradient descent until the target loss function converges, obtaining the optimal parameters and the question-answer quality evaluation model.
[0092] In this embodiment, the question-answering quality assessment model to be trained and the question-answering pair preference dataset are determined. The question-answering pair preference dataset is determined by labeling the question-answering pair dataset with answer preferences using multiple pre-defined review models. Each pre-defined review model has corresponding reliability parameters to be estimated. For each question-answering pair in the question-answering pair preference dataset, based on the multiple reliability parameters to be estimated corresponding to the multiple pre-defined review models and the labeling results of each question-answering pair, a noise probability function corresponding to the review model for each question-answering pair is constructed. This achieves explicit modeling of the preference labeling bias of the pre-defined review models, effectively reducing the impact of noise labeling on model training and helping to more accurately characterize the observed question-answering pair preference dataset. Based on the predicted quality score relationship of each question-answering pair determined by the question-answering quality assessment model to be trained, a preference probability function corresponding to each question-answering pair is constructed, effectively enhancing the credibility of the model output and enabling the model to better reflect the real differences in dialogue quality. Based on the noise probability function and preference probability function of the review models corresponding to all question-answer pairs, a joint likelihood function is constructed. This joint likelihood function quantifies the probability of observing the question-answer pair preference dataset under the trainable model parameters corresponding to the question-answer quality assessment model to be trained and the estimated reliability parameters corresponding to each pre-defined review model. Based on the joint likelihood function, a target loss function corresponding to the question-answer quality assessment model to be trained is constructed, ensuring that the model parameters better fit the real observation distribution during training. Based on the question-answer pair preference dataset, the question-answer quality assessment model to be trained is trained, and the trainable model parameters and the estimated reliability parameters corresponding to each pre-defined review model are iteratively updated during training until the target loss function converges, resulting in the question-answer quality assessment model. This achieves joint optimization of the trainable model parameters and the estimated reliability parameters of the pre-defined review models in the absence of absolute labels, improving the model's robustness, generalization ability, and training transparency.
[0093] In one embodiment, such as Figure 3 As shown, Figure 3 This is a flowchart illustrating the first noise probability function construction step in one embodiment. Based on multiple reliability parameters to be estimated corresponding to multiple preset review models, and the annotation results of each question-answer pair, the noise probability function of the review model corresponding to each question-answer pair is constructed, including the following steps:
[0094] Step S301: If the potential true preference label corresponding to the question-answer pair is the first preference label, then according to the estimated true positive rate and preference label of each preset review model, construct the first judgment likelihood function of each preset review model under the first preference label.
[0095] The first preference label is used to characterize that the quality of the second question-and-answer pair to be compared is higher than that of the first question-and-answer pair to be compared.
[0096] wherein the first judging likelihood function is used to quantify a probability that the preset evaluation model generates the preference label corresponding to the current question-answer pair under the first preference label.
[0097] It should be noted that under the first preference label, each preset evaluation model is only dependent on the true positive rate to be estimated for the preference label of the current question-answer pair.
[0098] Step S302, multiplying the first judging likelihood functions corresponding to all preset evaluation models to obtain the first noise probability function corresponding to each group of question-answer pairs.
[0099] wherein the first noise probability function is used to quantify a probability that the multiple preference labels corresponding to the question-answer pair are observed under the first preference label.
[0100] It can be understood that each preset evaluation model is independent of each other, the labeling behavior of the multiple preset evaluation models is a Bernoulli trial, and the first noise probability function is a continuous multiplication of the first judging likelihood functions corresponding to the preset evaluation models.
[0101] Exemplarily, if the latent true preference label corresponding to the question-answer pair is the first preference label, i.e. then the first judging likelihood function of each preset evaluation model under the first preference label is constructed according to the true positive rate to be estimated and the preference label corresponding to each preset evaluation model, and is specifically as follows:
[0102] ;
[0103] wherein denotes the preference label given by the jth preset evaluation model for the ith group of question-answer pairs; denotes the true positive rate to be estimated corresponding to the jth preset evaluation model; denotes that the latent true preference label corresponding to the ith group of question-answer pairs is the first preference label. Each term is essentially a Bernoulli distribution: when , takes , otherwise takes .
[0104] Further, the first judging likelihood functions corresponding to all preset evaluation models are multiplied to obtain the first noise probability function corresponding to each group of question-answer pairs , and is specifically as follows:
[0105] .
[0106] In the embodiment, by constructing the first judgment likelihood function of each preset review model under the first preference label based on the estimated true positive rate of each preset review model and the actual output preference label thereof under the premise that the real preference label is the first preference label, and multiplying the first judgment likelihood functions of all preset review models to obtain the first noise probability function, explicit modeling of the annotation uncertainty of the preset review model is realized, the reliability difference of different preset review models in preference judgment is effectively described, the influence of noise annotation on model training is reduced, and a foundation is laid for improving the robustness and generalization ability of the question and answer quality evaluation model.
[0107] In one embodiment, as shown in Figure 4 , Figure 4 is a flowchart of the second noise probability function construction step in one embodiment; the review model noise probability function corresponding to each group of question and answer pairs is constructed according to the plurality of estimated reliability parameters corresponding to the plurality of preset review models and the annotation results of each group of question and answer pairs, including the following steps:
[0108] Step S401, if the potential real preference label of the question and answer pair is the second preference label, then the second judgment likelihood function of each preset review model under the second preference label is constructed according to the estimated true negative rate and the preference label corresponding to each preset review model.
[0109] The second preference label is used to represent that the quality of the first to-be-compared question and answer pair is higher than that of the second to-be-compared question and answer pair.
[0110] The second judgment likelihood function is used to quantify the probability of generating the preference label corresponding to the current question and answer pair by the preset review model under the second preference label.
[0111] It should be noted that under the second preference label, the preference label of the current question and answer pair for each preset review model only depends on the estimated true negative rate thereof.
[0112] Step S402, multiplying the second judgment likelihood functions corresponding to all preset review models to obtain the second noise probability function corresponding to each group of question and answer pairs.
[0113] The second noise probability function is used to quantify the probability of observing the plurality of preference labels corresponding to the question and answer pair under the second preference label.
[0114] For example, if the potential real preference label of the question and answer pair is the second preference label, i.e. , the second judgment likelihood function of each preset review model under the second preference label is constructed according to the estimated true negative rate and the preference label corresponding to each preset review model. The details are as follows:
[0115] ;
[0116] in, , representing the preference label given by the j-th preset review model for the i-th question-answer pair; , representing the true negative rate to be estimated corresponding to the j-th preset review model; The potential true preference label corresponding to the i-th question-answer pair is the second preference label.
[0117] Furthermore, by multiplying the second likelihood functions corresponding to all preset review models, the second noise probability function corresponding to each question-answer pair is obtained. The details are as follows:
[0118] .
[0119] In this embodiment, under the premise that the true preference label is the second preference label, a second judgment likelihood function is constructed for each preset review model under the second preference label based on the estimated true negative rate of each preset review model and its actual output preference label. The second noise probability function is obtained by multiplying the second judgment likelihood functions of all preset review models. This realizes explicit modeling of the uncertainty of the preset review model labeling, effectively characterizes the reliability difference of different preset review models in preference judgment, reduces the impact of noise labeling on model training, and lays the foundation for improving the robustness and generalization ability of the question answering quality assessment model.
[0120] In one embodiment, such as Figure 5 As shown, Figure 5 This is a flowchart illustrating the steps for constructing the preference probability function in one embodiment. Based on the predicted quality score relationship for each question-answer pair determined by the question-answer quality assessment model to be trained, the preference probability function corresponding to each question-answer pair is constructed, including the following steps:
[0121] Step S501: For each question-answer pair, determine the prediction quality difference relationship of the question-answer pair based on the first prediction quality score relationship of the first question-answer pair to be compared and the second prediction quality score relationship of the second question-answer pair to be compared, as determined by the question-answer quality assessment model to be trained.
[0122] The prediction quality difference formula is determined by subtracting the first prediction quality score formula from the second prediction quality score formula.
[0123] Step S502: If the potential true preference label corresponding to the question-answer pair is the first preference label, then construct the first preference probability function corresponding to each question-answer pair according to the prediction quality difference relationship corresponding to each question-answer pair.
[0124] wherein the first preference probability function is used to quantify the probability that the quality of the second to-be-compared question-answer pair is higher than that of the first to-be-compared question-answer pair.
[0125] Exemplarily, it is assumed that denotes the quality score of the to-be-compared question-answer pair x; in order to model the uncertainty in the quality of the question-answer pair, the embodiment adopts a Thurstone’s Case V model to model the preference probability; specifically, it is assumed that obeys a Gaussian distribution with as the mean value and a constant as the standard deviation, that is, Under this assumption, the difference between the quality score of the first to-be-compared question-answer pair A and the quality score of the second to-be-compared question-answer pair B also obeys a Gaussian distribution, with the mean value and the variance 2; wherein is the first predicted quality score relationship formula corresponding to the first to-be-compared question-answer pair A, is the second predicted quality score relationship formula corresponding to the second to-be-compared question-answer pair B, wherein denotes a pre-trained embedding module; denotes a to-be-trained quality scorer.
[0126] If the latent true preference label corresponding to the question-answer pair is the first preference label, that is, then according to the second predicted quality score relationship formula and the first predicted quality score relationship formula corresponding to the first to-be-compared question-answer pair A and the second to-be-compared question-answer pair B, a predicted quality difference relationship formula is constructed, and each first preference probability function corresponding to each group of question-answer pairs is constructed, specifically as follows:
[0127] ;
[0128] wherein denotes the cumulative distribution function of the standard normal distribution.
[0129] In step S503, if the latent true preference label corresponding to the question-answer pair is the second preference label, then according to the predicted quality difference relationship formula corresponding to each group of question-answer pairs, a second preference probability function corresponding to each group of question-answer pairs is constructed.
[0130] wherein the second preference probability function is used to quantify the probability that the quality of the first to-be-compared question-answer pair is higher than that of the second to-be-compared question-answer pair.
[0131] It should be noted that the construction principle of the second preference probability function is the same as that of the first preference probability function described above, and will not be described herein again.
[0132] It should be noted that the sum of the first preference probability function and the second preference probability function is 1; that is, the second preference probability function... .
[0133] In this embodiment, by constructing a first preference probability function or a second preference probability function based on the potential true preference labels corresponding to the question-answer pair, the uncertainty of the question-answer pair quality is modeled, which provides a basis for constructing the joint likelihood function, so as to ensure that the model can still accurately infer the true quality distribution of the question-answer pair in the absence of absolute labels.
[0134] In one embodiment, such as Figure 6 As shown, Figure 6 This is a flowchart illustrating the steps for constructing the joint likelihood function in one embodiment. Based on the noise probability function and preference probability function of the review model corresponding to all question-answer pairs, the joint likelihood function is constructed, including the following steps:
[0135] Step S601: For each question-answer pair, multiply the first noise probability function and the first preference probability function corresponding to the question-answer pair to construct the first joint probability function.
[0136] Step S602: Multiply the second noise probability function and the second preference probability function corresponding to the question-answer pair to construct the second joint probability function.
[0137] Step S603: The sum of the first joint probability function and the second joint probability function is determined as the intermediate joint likelihood function corresponding to the question-answer pair.
[0138] Step S604: Multiply the intermediate joint likelihood functions corresponding to all question-answer pairs to construct the joint likelihood function.
[0139] For example, based on the implementation method in the above embodiments, a first noise probability function, a first preference probability function, a second noise probability function, and a second preference probability function are constructed for each question-answer pair. The first noise probability function is denoted as... The first preference probability function is denoted as The second noise probability function is denoted as The second preference probability function is denoted as The first noise probability function corresponding to the question-answer pair. and the first preference probability function Multiplying them together, we obtain the first joint probability function as follows: The second noise probability function corresponding to the question-answer pair. Second preference probability function Multiplying them together, we obtain the second joint probability function as follows: Then, by adding the first joint probability function to the second joint probability function, we obtain the intermediate joint likelihood function corresponding to the question-answer pair. Further, the intermediate joint likelihood functions corresponding to all the group question-answer pairs are multiplied to construct a joint likelihood function Specifically as follows:
[0140]
[0141] Further, the joint likelihood function is taken logarithm and then negative value to obtain a negative log-likelihood function; the negative log-likelihood function is determined as a target loss function corresponding to the question-answer quality evaluation model to be trained, and the target loss function is as follows:
[0142]
[0143] In this embodiment, by constructing the joint likelihood function on the question-answer pair preference data set and designing the target loss function based thereon, the parameter optimization problem of the question-answer quality evaluation model is formalized as a maximum likelihood estimation task. As a classical statistical parameter estimation method, the core of maximum likelihood estimation is to find the model parameter configuration that maximizes the probability of the observed data, i.e., the question-answer pair preference data set, so as to improve the robustness and generalization ability of model training, and lay a foundation for realizing high-quality question-answer pair quality evaluation in the absence of absolute labels.
[0144] In one embodiment, the question-answer quality evaluation model to be trained includes a pre-trained embedding module and a quality scorer to be trained; wherein the pre-trained embedding module refers to an embedding model whose parameters are pre-trained and not updated again, and is used to map the input question-answer pair to an embedding space and output a corresponding embedding vector; the quality scorer to be trained is used to further map the embedding vector to a quality space and output a quality score corresponding to the question-answer pair. The trainable model parameters include the model parameters corresponding to the quality scorer to be trained.
[0145] According to the question-answer pair preference data set, the question-answer quality evaluation model to be trained is trained, and the trainable model parameters and the to-be-estimated reliability parameters corresponding to each pre-set review model are iteratively updated during the training process until the target loss function converges, to obtain the question-answer quality evaluation model, including the following steps:
[0146] Step 1, under the condition that the model parameters of the pre-trained embedding module remain unchanged, the quality scorer to be trained is trained according to the question-answer pair preference data set, and the model parameters corresponding to the quality scorer to be trained and the to-be-estimated reliability parameters corresponding to each pre-set review model are iteratively updated during the training process until the target loss function converges, to obtain the trained quality scorer;
[0147] Step 2, according to the pre-trained embedding module and the trained quality scorer, the question-answer quality evaluation model is obtained.
[0148] Exemplarily, the embedding module g pre-trained in the training process remains frozen, mapping the input question-answer pair into a high-dimensional latent representation, which is then input into the question-answer quality evaluation model to be trained (the input dimension of which is consistent with the output dimension of the embedding), outputting the corresponding quality score. The model parameters ω of the question-answer quality evaluation model to be trained and the reliability parameters of the preset review model to be estimated are jointly optimized by maximum likelihood estimation until the target loss function converges, obtaining the trained quality scorer. According to the pre-trained embedding module and the trained quality scorer, the question-answer quality evaluation model is obtained.
[0149] It should be noted that the training phase needs to exclude the case where the preference label is the third preference label (i.e., the quality of the second to-be-compared question-answer pair in the question-answer pair is equal to the first to-be-compared question-answer pair, ), because the third preference label cannot provide effective positive and negative sample information, and therefore cannot be used to estimate the model parameters ω of the question-answer quality evaluation model to be trained and the reliability parameters of the preset review model to be estimated , avoiding the introduction of noise and uncertainty.
[0150] In this embodiment, under the premise of freezing the model parameters of the pre-trained embedding module, only the question-answer quality evaluation model to be trained is trained, realizing the collaborative learning of the quality scorer and the reliability of the preset review model, ensuring that the finally formed question-answer quality evaluation model has higher discrimination ability and stronger anti-noise performance, and is suitable for automatic quality evaluation tasks in complex dialogue scenarios.
[0151] In one embodiment, a question-answer quality evaluation method is also provided, comprising the following steps:
[0152] Step 1, obtaining a first target question-answer pair and a second target question-answer pair.
[0153] Among them, the first target question-answer pair and the second target question-answer pair constitute a group of question-answer pairs to be evaluated for quality comparison and preference judgment; wherein the first target question-answer pair and the second target question-answer pair can be a question and answer combination in a single round of dialogue, or a complete interaction sequence in a multi-round dialogue context.
[0154] It can be understood that the first target question-answer pair and the second target question-answer pair are generated by different large language models.
[0155] Step 2, using the question-answer quality evaluation model in any one of the above embodiments to respectively evaluate the quality of the first target question-answer pair and the second target question-answer pair, obtaining a first target quality score corresponding to the first target question-answer pair and a second target quality score corresponding to the second target question-answer pair.
[0156] For example, taking a first target question-answer pair as an example, the first target question-answer pair is input into the pre-trained embedding module of the question-answer quality assessment model to obtain the corresponding embedding vector; the embedding vector corresponding to the first target question-answer pair is input into the trained quality scorer to obtain the first target quality score corresponding to the first target question-answer pair; similarly, the second target quality score corresponding to the second target question-answer pair can be obtained. Furthermore, preference judgment can be performed based on the first target quality score corresponding to the first target question-answer pair and the second target question-answer pair corresponding to the second target question-answer pair to determine the question-answer pair with better quality from the first target question-answer pair and the second target question-answer pair.
[0157] In this embodiment, based on the question-and-answer quality assessment model, the quality quantification assessment of the first target question-and-answer pair and the second target question-and-answer pair is realized, which improves the accuracy and reliability of the assessment results. At the same time, by directly outputting the corresponding target quality scores, the efficiency of quality assessment is effectively improved and the complexity of quality assessment calculation is reduced.
[0158] In one specific embodiment, such as Figure 7 As shown, Figure 7 The following is a performance comparison chart of a specific embodiment model; the question-answering quality assessment model of this application is denoted as MTDEval. In order to evaluate the performance of the question-answering quality assessment model MTDEval of this application, two types of benchmark evaluation are adopted: individual scoring at the overall level and pairwise comparison.
[0159] The individual scoring benchmarks include xDial-IEval, MT-Bench, and Daily-MTD, and the Pearson and Spearman rank correlation coefficients are used as evaluation metrics to assess the correlation between the model scores and the reference answers.
[0160] The paired comparison benchmarks include xDial-IEval-Pair, MT-Bench-Human, Chatbot-Arena, and Daily-MTD-Pair. Accuracy is used as the evaluation metric, and two evaluation methods are adopted: one is to exclude the "tie" case (denoted as "w / o tie"), and the other is to include the "tie" response in the evaluation (denoted as "w / tie").
[0161] The embodiments of the present application select advanced LLM baseline models, including GPT-4o, Grok-3, Claude-3.7-Sonnet, and Deepseek-R1. In addition, open source LLM baseline models with strong instruction following capabilities are also used, including Llama-3-8B-Instruct and Llama-3.1-8B-Instruct, as well as Qwen-2-7B-Instruct and Qwen2.5-7B-Instruct. For the current most advanced evaluator baseline, the embodiments of the present application use models such as AutoJ-13B, Samer, Prometheus-7B, Prometheus2-7B, and ArmoRM-8B.
[0162] As can be seen from Figure 7 , in the single scoring task, the MTDEval of the present application has achieved significant improvement on all three single scoring benchmarks, far exceeding existing open source baselines. In the xDial-IEval task, the improvement of the MTDEval of the present application is particularly prominent, even surpassing most proprietary models. In addition, the correlation coefficient of the MTDEval of the present application on all benchmarks is improved by more than 0.1 compared with its ArmoRM backbone model, which effectively verifies the effectiveness of the proposed framework. However, it is worth noting that the performance of all open source models on MT-Bench is still significantly lower than that of SOTA proprietary models, highlighting the continuous challenge of this benchmark to open source evaluators and reflecting the performance gap between open source LLM and proprietary LLM.
[0163] In the paired comparison task, the MTDEval of the present application shows a strong leading position among open source models, achieving the best performance in seven of the eight tasks of the four benchmarks, and ranking second in the remaining one. Although its ArmoRM backbone is already competitive among open source models, the MTDEval of the present application still achieves at least 5% improvement on this basis, and on high difficulty benchmarks such as MT-Bench-Human and Chatbot Arena, the improvement is even close to 15%. The MTDEval of the present application exceeds almost all proprietary LLMs on multiple paired comparison datasets, fully verifying its excellent cross-task generalization ability.
[0164] It can be understood that the MTDEval model of the present application performs far better than the open source model benchmark, close to or even exceeding the performance of closed source SOTA models, showing excellent generalization ability and robustness, while also having the characteristics of light weight and high efficiency.
[0165] It should be understood that although the steps in the flowcharts involved in the embodiments described above are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the embodiments described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.
[0166] Based on the same inventive concept, the embodiments of the present application also provide a question and answer quality evaluation model training device for implementing the question and answer quality evaluation model training method described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more question and answer quality evaluation model training device embodiments provided below can refer to the limitations of the question and answer quality evaluation model training method in the above text, which will not be repeated here.
[0167] In one exemplary embodiment, as shown in Figure 8 a question and answer quality evaluation model training device is provided, comprising: a determination module 801, a first construction module 802, a second construction module 803, a third construction module 804, a fourth construction module 805 and a training module 806, wherein:
[0168] The determination module 801 is configured to determine a question and answer quality evaluation model to be trained and a question and answer pair preference data set; wherein the question and answer pair preference data set is determined by marking the answer preference of a question and answer pair data set through a plurality of preset review models; each preset review model has a corresponding to-be-estimated reliability parameter;
[0169] The first construction module 802 is configured to, for each set of question and answer pairs in the question and answer pair preference data set, construct a review model noise probability function corresponding to each set of question and answer pairs according to a plurality of to-be-estimated reliability parameters corresponding to a plurality of preset review models and the marking result of each set of question and answer pairs;
[0170] The second construction module 803 is configured to construct a preference probability function corresponding to each set of question and answer pairs according to a predicted quality score relationship of each set of question and answer pairs determined by the question and answer quality evaluation model to be trained;
[0171] The third construction module 804 is configured to construct a joint likelihood function according to the review model noise probability functions and the preference probability functions corresponding to all groups of question-answer pairs; the joint likelihood function is used to quantify a probability of observing the question-answer pair preference data set under the trainable model parameters corresponding to the question-answer quality evaluation model to be trained and the to-be-estimated reliability parameters corresponding to each of the preset review models.
[0172] The fourth construction module 805 is configured to construct a target loss function corresponding to the question-answer quality evaluation model to be trained according to the joint likelihood function.
[0173] The training module 806 is configured to train the question-answer quality evaluation model to be trained according to the question-answer pair preference data set, and iteratively update the trainable model parameters and the to-be-estimated reliability parameters corresponding to each of the preset review models in the training process until the target loss function converges, so as to obtain the question-answer quality evaluation model.
[0174] The question-answer quality evaluation model training apparatus determines the question-answer quality evaluation model to be trained and the question-answer pair preference data set; the question-answer pair preference data set is determined by performing answer preference labeling on the question-answer pair data set by using the plurality of preset review models; each of the preset review models has a corresponding to-be-estimated reliability parameter; for each group of question-answer pairs in the question-answer pair preference data set, a review model noise probability function corresponding to each group of question-answer pairs is constructed according to the plurality of to-be-estimated reliability parameters corresponding to the plurality of preset review models and the labeling result of each group of question-answer pairs, which realizes display modeling of the preference labeling deviation of the preset review model, effectively reduces the influence of noise labeling on model training, and helps to more accurately depict the observed question-answer pair preference data set; a preference probability function corresponding to each group of question-answer pairs is constructed according to a predicted quality score relationship of each group of question-answer pairs determined by the question-answer quality evaluation model to be trained, which effectively enhances the credibility of the model output, so that the model can better reflect the real difference in dialogue quality; a joint likelihood function is constructed according to the review model noise probability functions and the preference probability functions corresponding to all groups of question-answer pairs; the joint likelihood function is used to quantify a probability of observing the question-answer pair preference data set under the trainable model parameters corresponding to the question-answer quality evaluation model to be trained and the to-be-estimated reliability parameters corresponding to each of the preset review models; a target loss function corresponding to the question-answer quality evaluation model to be trained is constructed according to the joint likelihood function, which can ensure that the model parameters are better fitted to the real observation distribution during training; the question-answer quality evaluation model to be trained is trained according to the question-answer pair preference data set, and the trainable model parameters and the to-be-estimated reliability parameters corresponding to each of the preset review models are iteratively updated in the training process until the target loss function converges, so as to obtain the question-answer quality evaluation model, which realizes joint optimization of the trainable model parameters and the to-be-estimated reliability parameters of the preset review model in the absence of absolute labels, and improves the robustness, generalization ability and training transparency of the model.
[0175] In an embodiment, the reliability parameter to be estimated comprises a true positive rate to be estimated; the annotation result of each set of question-answer pairs comprises a plurality of preference labels; the plurality of preference labels correspond to a plurality of preset review models one by one; the review model noise probability function comprises a first noise probability function; the first construction module 802 is further configured to:
[0176] If the latent true preference label corresponding to the question-answer pair is a first preference label, then according to the true positive rate to be estimated corresponding to each preset review model and the preference label, a first judgment likelihood function of each preset review model under the first preference label is constructed;
[0177] The first judgment likelihood functions corresponding to all preset review models are multiplied to obtain a first noise probability function corresponding to each set of question-answer pairs; the first noise probability function is used to quantify the probability of observing the plurality of preference labels corresponding to the question-answer pair under the first preference label.
[0178] In an embodiment, the reliability parameter to be estimated comprises a true negative rate to be estimated; the review model noise probability function comprises a second noise probability function; the first construction module 802 is further configured to:
[0179] If the latent true preference label corresponding to the question-answer pair is a second preference label, then according to the true negative rate to be estimated corresponding to each preset review model and the preference label, a second judgment likelihood function of each preset review model under the second preference label is constructed;
[0180] The second judgment likelihood functions corresponding to all preset review models are multiplied to obtain a second noise probability function corresponding to each set of question-answer pairs; the second noise probability function is used to quantify the probability of observing the plurality of preference labels corresponding to the question-answer pair under the second preference label.
[0181] In an embodiment, each set of question-answer pairs comprises a first to-be-compared question-answer pair and a second to-be-compared question-answer pair; the preference probability function comprises a first preference probability function and a second preference probability function; the second construction module 803 is further configured to:
[0182] For each set of question-answer pairs, according to a first predicted quality score relationship of the first to-be-compared question-answer pair and a second predicted quality score relationship of the second to-be-compared question-answer pair determined by the question-answer quality evaluation model to be trained, a predicted quality difference relationship of the question-answer pair is determined;
[0183] If the latent true preference label corresponding to the question-answer pair is a first preference label, then according to the predicted quality difference relationship corresponding to each set of question-answer pairs, a first preference probability function corresponding to each set of question-answer pairs is constructed; the first preference probability function is used to quantify the probability that the quality of the second to-be-compared question-answer pair is higher than that of the first to-be-compared question-answer pair;
[0184] If the corresponding potential true preference label of the question-answer pair is the second preference label, a second preference probability function corresponding to each group of question-answer pairs is constructed according to the respective prediction quality difference relationship of each group of question-answer pairs; the second preference probability function is used to quantify the probability that the quality of the first to-be-compared question-answer pair is higher than that of the second to-be-compared question-answer pair.
[0185] In an embodiment, the review model noise probability function includes a first noise probability function and a second noise probability function; the preference probability function includes a first preference probability function and a second preference probability function; and the third construction module 804 is further configured to:
[0186] For each group of question-answer pairs, the first noise probability function and the first preference probability function corresponding to the question-answer pair are multiplied to construct a first joint probability function;
[0187] The second noise probability function and the second preference probability function corresponding to the question-answer pair are multiplied to construct a second joint probability function;
[0188] The sum of the first joint probability function and the second joint probability function is determined as an intermediate joint likelihood function corresponding to the question-answer pair;
[0189] The intermediate joint likelihood functions corresponding to all groups of question-answer pairs are multiplied to construct a joint likelihood function.
[0190] In an embodiment, the fourth construction module 805 is further configured to:
[0191] The joint likelihood function is taken as a logarithm and then taken as a negative value to obtain a negative log-likelihood function;
[0192] The negative log-likelihood function is determined as a target loss function corresponding to the to-be-trained question-answer quality evaluation model.
[0193] In an embodiment, the to-be-trained question-answer quality evaluation model includes a pre-trained embedding module and a to-be-trained quality scorer; the trainable model parameter includes a model parameter corresponding to the to-be-trained quality scorer; and the training module 806 is further configured to:
[0194] Under the condition that the model parameter of the pre-trained embedding module is kept unchanged, the to-be-trained quality scorer is trained according to the question-answer pair preference data set, and the model parameter corresponding to the to-be-trained quality scorer and the to-be-estimated reliability parameter corresponding to each pre-set review model are iteratively updated during the training process until the target loss function converges, so as to obtain a trained quality scorer;
[0195] The question-answer quality evaluation model is obtained according to the pre-trained embedding module and the trained quality scorer.
[0196] The various modules in the above question and answer quality evaluation model training apparatus can be implemented by software, hardware, or a combination thereof. The various modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform the operations corresponding to the various modules.
[0197] In an exemplary embodiment, a computer device, which can be a server, has an internal structure diagram as shown in Figure 9 The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store data related to question and answer quality evaluation model training and question and answer quality evaluation. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement a question and answer quality evaluation model training method and a question and answer quality evaluation method.
[0198] Those skilled in the art can understand that Figure 9 The structure shown in the above embodiment is only a block diagram of part of the structure related to the scheme of the present application, and does not limit the computer device to which the scheme of the present application is applied. Specifically, the computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0199] In an embodiment, a computer device is also provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0200] In an embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0201] In an embodiment, a computer program product is provided, which includes a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0202] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0203] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments of each method. In the embodiments provided in the present application, any reference to memory, database or other medium can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.
[0204] Any technical features in the above embodiments can be combined, and for the sake of brevity, not all possible combinations are described above, however, any combination of these technical features is deemed to be within the scope of the present application.
[0205] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be pointed out that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for training a question-answering quality assessment model, characterized in that, The method comprises: determining a question and answer quality evaluation model to be trained and a question and answer pair preference data set; wherein the question and answer pair preference data set is determined by multiple preset review models performing answer preference labeling on a question and answer pair data set; each of the preset review models has a corresponding to-be-estimated reliability parameter; for each group of question and answer pairs in the question and answer pair preference data set, constructing a review model noise probability function corresponding to each group of question and answer pairs according to multiple to-be-estimated reliability parameters corresponding to multiple preset review models and the labeling result of each group of question and answer pairs; constructing a preference probability function corresponding to each group of question and answer pairs according to a predicted quality score relationship of each group of question and answer pairs determined by the question and answer quality evaluation model to be trained; constructing a joint likelihood function according to the review model noise probability functions and the preference probability functions corresponding to all groups of question and answer pairs; the joint likelihood function is used to quantify the probability of observing the question and answer pair preference data set under the trainable model parameters corresponding to the question and answer quality evaluation model to be trained and the to-be-estimated reliability parameters corresponding to each of the preset review models; constructing a target loss function corresponding to the question and answer quality evaluation model to be trained according to the joint likelihood function; training the question and answer quality evaluation model to be trained according to the question and answer pair preference data set, and iteratively updating the trainable model parameters and the to-be-estimated reliability parameters corresponding to each of the preset review models in the training process until the target loss function converges, to obtain a question and answer quality evaluation model.
2. The method of claim 1, wherein, The to-be-estimated reliability parameters include a to-be-estimated true positive rate; the labeling result of each group of question and answer pairs includes multiple preference labels; the multiple preference labels correspond one-to-one to the multiple preset review models; and the review model noise probability function includes a first noise probability function. The method comprises: if the latent true preference label corresponding to the question and answer pair is a first preference label, constructing a first judgment likelihood function of each of the preset review models under the first preference label according to the to-be-estimated true positive rate corresponding to each of the preset review models and the preference label; multiplying the first judgment likelihood functions corresponding to all the preset review models to obtain a first noise probability function corresponding to each group of question and answer pairs; the first noise probability function is used to quantify the probability of observing the multiple preference labels corresponding to the question and answer pair under the first preference label.
3. The method of claim 2, wherein, The to-be-estimated reliability parameters include a to-be-estimated true negative rate; the review model noise probability function includes a second noise probability function; and the method comprises: if the latent true preference label corresponding to the question and answer pair is a first preference label, constructing a first judgment likelihood function of each of the preset review models under the first preference label according to the to-be-estimated true positive rate corresponding to each of the preset review models and the preference label; if the potential real preference label corresponding to the question-answer pair is a second preference label, then a second judging likelihood function of each of the preset review models under the second preference label is constructed according to the respective to-be-estimated true negative rate and the preference label of each of the preset review models; all the second judging likelihood functions corresponding to the preset review models are multiplied to obtain a second noise probability function corresponding to each of the question-answer pairs; the second noise probability function is used to quantify a probability of observing multiple preference labels corresponding to the question-answer pair under the second preference label.
4. The method of claim 3, wherein, Each of the question-answer pairs includes a first to-be-compared question-answer pair and a second to-be-compared question-answer pair; the preference probability function includes a first preference probability function and a second preference probability function; and the constructing of the preference probability function corresponding to each of the question-answer pairs according to the predicted quality score relationship of each of the question-answer pairs determined by the to-be-trained question-answer quality evaluation model includes: For each of the question-answer pairs, a predicted quality difference relationship of the question-answer pair is determined according to a first predicted quality score relationship of the first to-be-compared question-answer pair and a second predicted quality score relationship of the second to-be-compared question-answer pair determined by the to-be-trained question-answer quality evaluation model; if the potential real preference label corresponding to the question-answer pair is a first preference label, then a first preference probability function corresponding to each of the question-answer pairs is constructed according to the predicted quality difference relationship of each of the question-answer pairs; the first preference probability function is used to quantify a probability that the quality of the second to-be-compared question-answer pair is higher than that of the first to-be-compared question-answer pair; if the potential real preference label corresponding to the question-answer pair is a second preference label, then a second preference probability function corresponding to each of the question-answer pairs is constructed according to the predicted quality difference relationship of each of the question-answer pairs; the second preference probability function is used to quantify a probability that the quality of the first to-be-compared question-answer pair is higher than that of the second to-be-compared question-answer pair.
5. The method of claim 1 or claim 4, wherein, The review model noise probability function includes a first noise probability function and a second noise probability function; and the preference probability function includes a first preference probability function and a second preference probability function. The constructing of the joint likelihood function according to the review model noise probability function and the preference probability function corresponding to all the question-answer pairs includes: for each of the question-answer pairs, a first joint probability function is constructed by multiplying the first noise probability function and the first preference probability function corresponding to the question-answer pair; a second joint probability function is constructed by multiplying the second noise probability function and the second preference probability function corresponding to the question-answer pair; a sum of the first joint probability function and the second joint probability function is determined as an intermediate joint likelihood function corresponding to the question-answer pair; all the intermediate joint likelihood functions corresponding to the question-answer pairs are multiplied to construct a joint likelihood function.
6. The method of claim 1, wherein, The constructing of the target loss function corresponding to the to-be-trained question-answer quality evaluation model according to the joint likelihood function includes: a negative log-likelihood function is obtained by taking a logarithm of the joint likelihood function and then taking a negative value. The negative log-likelihood function is determined as a target loss function corresponding to the to-be-trained question and answer quality evaluation model.
7. The method of claim 1, wherein, The to-be-trained question and answer quality evaluation model comprises a pre-trained embedding module and a to-be-trained quality scorer; the trainable model parameters comprise model parameters corresponding to the to-be-trained quality scorer; the to-be-trained question and answer quality evaluation model is trained according to the question and answer pair preference data set, and the trainable model parameters and the to-be-estimated reliability parameters corresponding to each of the preset review models are iteratively updated during the training process until the target loss function converges, to obtain a question and answer quality evaluation model, comprising: Under the condition that the model parameters of the pre-trained embedding module remain unchanged, the to-be-trained quality scorer is trained according to the question and answer pair preference data set, and the model parameters corresponding to the to-be-trained quality scorer and the to-be-estimated reliability parameters corresponding to each of the preset review models are iteratively updated during the training process until the target loss function converges, to obtain a trained quality scorer. The question and answer quality evaluation model is obtained according to the pre-trained embedding module and the trained quality scorer.
8. A question-and-answer quality assessment method, characterized in that, The method comprises: obtaining a first target question and answer pair and a second target question and answer pair; using the question and answer quality evaluation model of any one of claims 1 to 7 to perform quality evaluation on the first target question and answer pair and the second target question and answer pair respectively, to obtain a first target quality score corresponding to the first target question and answer pair and a second target quality score corresponding to the second target question and answer pair. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 8.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 8.
11. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 8. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 8.