Model quality evaluation method and device, electronic equipment and storage medium
By constructing an evaluation framework that combines objective and subjective evaluation indicators, the problem of insufficient comprehensiveness and objectivity in the evaluation results of existing technologies is solved, and the comprehensiveness and scientific nature of the quality evaluation of large voice interaction models are realized, thereby improving the user experience.
Patent Information
- Application Number
- CN202511204998.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing methods for evaluating the quality of large-scale voice interaction models rely too heavily on accuracy and neglect user experience, resulting in evaluation results that are not objective or comprehensive enough.
An evaluation framework combining objective and subjective evaluation indicators is constructed. Dialogue data is input into a large voice interaction model to determine the evaluation results under both objective and subjective evaluation indicators. The two are then integrated to generate a comprehensive and scientific quality evaluation result.
It achieves comprehensiveness and objectivity in the quality evaluation of large-scale voice interaction models, provides clear and reliable basis for iterative optimization, and significantly improves user experience.
Smart Images

Figure CN120748372B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a model quality evaluation method and device, electronic equipment and a storage medium. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, the voice interaction large model has been widely used in intelligent customer service, intelligent assistants and many other fields due to its powerful functions. However, in the actual application process, the quality of the model directly affects the user experience and product reputation, so a scientific and effective model quality evaluation method is crucial.
[0003] Currently, the quality evaluation of the voice interaction large model mainly relies on a single indicator, showing a trend of only focusing on "intelligence quotient" and ignoring "emotional quotient". Specifically, most evaluations only focus on the correct answer rate, taking the question "how high is Mount Everest" as an example, as long as the model gives the correct answer, it is determined to be qualified. This overemphasis on accuracy will result in a single evaluation dimension of the model, which cannot fully reflect the performance of the model in actual application, and thus the comprehensiveness, accuracy and reliability of the evaluation results are questionable. SUMMARY
[0004] The present application provides a model quality evaluation method, device, electronic equipment and storage medium, to solve the problem that the quality evaluation of the model in the prior art relies too much on accuracy and ignores user experience, resulting in evaluation results that are not objective, comprehensive and accurate. The evaluation framework combining objective evaluation indicators and subjective evaluation indicators is constructed to deeply integrate technical indicators and user experience, thereby achieving objective and comprehensive quality evaluation.
[0005] The present application provides a model quality evaluation method, comprising:
[0006] determining dialogue data for model quality evaluation;
[0007] inputting the dialogue data into a voice interaction large model to be evaluated, and obtaining an answer result of the dialogue data by answering the dialogue data by the voice interaction large model;
[0008] based on the answer result, determining an indicator evaluation result of the voice interaction large model under objective evaluation indicators and subjective evaluation indicators, respectively;
[0009] based on the indicator evaluation results of the voice interaction large model under the objective evaluation indicators and the subjective evaluation indicators, determining a quality evaluation result of the voice interaction large model.
[0010] According to the model quality evaluation method provided by the application, the subjective evaluation index includes empathy degree; and the index evaluation result of the voice interaction large model under the empathy degree is determined based on the following steps.
[0011] An initial interaction state is determined, and the initial interaction state includes a user emotional state.
[0012] Based on the initial interaction state and the dialogue data, a user and the voice interaction large model are simulated to perform multi-round interactions to obtain dialogue data of the multi-round interactions and emotional changes and emotional scores of each round of interaction in the multi-round interactions. The dialogue data of each round of interaction includes a response result output by the voice interaction large model. The input of each round of interaction in the multi-round interactions is determined based on the emotional changes and emotional scores of the previous round of interaction.
[0013] Based on the dialogue data of the multi-round interactions and the emotional changes and emotional scores of each round of interaction in the multi-round interactions, the index evaluation result of the voice interaction large model under the empathy degree is determined.
[0014] According to the model quality evaluation method provided by the application, the determination of the index evaluation result of the voice interaction large model under the empathy degree includes:
[0015] Based on the dialogue data of the multi-round interactions and the emotional changes and emotional scores of each round of interaction in the multi-round interactions, an interaction empathy degree score is determined.
[0016] An empathy degree evaluation result of an evaluation expert under the empathy degree for the dialogue data of the multi-round interactions and the emotional changes and emotional scores of each round of interaction in the multi-round interactions is obtained.
[0017] Based on the interaction empathy degree score and the empathy degree evaluation result, the index evaluation result of the voice interaction large model under the empathy degree is determined.
[0018] According to the model quality evaluation method provided by the application, the determination of the initial interaction state includes:
[0019] Based on the dialogue data, an interaction element is determined. The interaction element includes an interaction role, an interaction background, an interaction target, and a potential intention.
[0020] Based on the interaction element, an initial interaction state including a user emotional state is constructed.
[0021] According to the model quality evaluation method provided by the application, the subjective evaluation index includes anthropomorphism degree.
[0022] The index evaluation result of the voice interaction large model under the anthropomorphism degree is determined based on the following steps.
[0023] perform content redundancy detection based on the response result, to obtain a redundancy detection result;
[0024] obtain a personification evaluation result of the evaluation expert evaluating the response result under the personification;
[0025] determine an index evaluation result of the voice interaction large model under the personification based on the content redundancy detection and the personification evaluation result.
[0026] According to the model quality evaluation method provided by the application, the content redundancy detection based on the response result comprises the following steps:
[0027] perform sentence naturalness detection based on the response result, to obtain a naturalness detection result;
[0028] perform word and sentence similarity detection based on the response result, to obtain a similarity detection result;
[0029] perform content redundancy detection based on the response result, to obtain a redundancy detection result;
[0030] determine the redundancy detection result based on at least one of the naturalness detection result, the similarity detection result and the redundancy detection result.
[0031] According to the model quality evaluation method provided by the application, the subjective evaluation index comprises richness;
[0032] The index evaluation result of the voice interaction large model under the richness is determined based on the following steps:
[0033] perform content expansibility detection on the response result based on the dialogue data, and perform content appropriateness detection based on the expanded content obtained through content expansibility detection and context content corresponding to the expanded content in the response result, to obtain an appropriateness detection result;
[0034] perform content practicality detection based on the response result, to obtain a practicality detection result;
[0035] perform content richness detection on the response result based on the scene type corresponding to the dialogue data, to obtain a richness detection result;
[0036] determine the index evaluation result of the voice interaction large model under the richness based on at least one of the appropriateness detection result, the practicality detection result and the richness detection result.
[0037] According to the model quality evaluation method provided by the application, the objective evaluation index comprises a correlation degree;
[0038] The index evaluation result of the voice interaction large model under the correlation degree is determined based on the following steps;
[0039] Single-turn correlation detection is performed based on the dialogue data and the response result, and a single-turn correlation detection result is obtained;
[0040] Multi-turn correlation detection is performed based on the dialogue data and the response result, and a multi-turn correlation detection result is obtained;
[0041] Viewpoint consistency detection is performed based on the dialogue data and the response result, and a viewpoint consistency detection result is obtained;
[0042] Instruction following detection is performed based on the dialogue data and the response result, and an instruction following detection result is obtained;
[0043] At least one of the multi-turn correlation detection result, the viewpoint consistency detection result and the instruction following detection result, and the single-turn correlation detection result are used to determine the index evaluation result of the voice interaction large model under the correlation degree.
[0044] According to the model quality evaluation method provided by the application, the objective evaluation index comprises accuracy; and the index evaluation result of the voice interaction large model under the accuracy is determined based on the following steps;
[0045] Content accuracy detection is performed on the response result based on the dialogue data, and an accuracy detection result is obtained;
[0046] Content fiction degree detection is performed on the response result based on the dialogue data, and a fiction degree detection result is obtained;
[0047] Sentence fluency detection is performed on the response result based on the dialogue data, and a fluency detection result is obtained;
[0048] At least one of the accuracy detection result, the fiction degree detection result and the fluency detection result is used to determine the index evaluation result of the voice interaction large model under the accuracy.
[0049] According to the model quality evaluation method provided by the application, the dialogue data used for model quality evaluation is determined by the following steps:
[0050] Dialogue data used for model quality evaluation is selected from a multi-scene dialogue data set;
[0051] The multi-scenario dialogue dataset includes at least two of the following: task-oriented dialogue data, game-oriented dialogue data, emotional companionship dialogue data, casual conversation dialogue data, and multi-turn challenge dialogue data.
[0052] The present invention also provides a model quality evaluation device, comprising:
[0053] The data determination unit is used to determine the dialogue data for model quality evaluation.
[0054] The data processing unit is used to input the dialogue data into the large voice interaction model to be evaluated, and the large voice interaction model responds to the dialogue data to obtain the response result of the dialogue data.
[0055] The result determination unit is used to determine the indicator evaluation results of the voice interaction model under objective evaluation indicators and subjective evaluation indicators based on the response results.
[0056] The quality evaluation unit is used to determine the quality evaluation result of the voice interaction model based on the evaluation results of the objective evaluation indicators and subjective evaluation indicators.
[0057] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the model quality evaluation method as described above.
[0058] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the model quality evaluation method as described above.
[0059] The model quality evaluation method, apparatus, electronic device, and storage medium provided by this invention input dialogue data into a large-scale voice interaction model to be evaluated, and obtain the response results output by the large-scale voice interaction model. Based on the response results, the evaluation results of the large-scale voice interaction model under objective evaluation indicators and subjective evaluation indicators are determined respectively. Based on the evaluation results under objective evaluation indicators and subjective evaluation indicators, the quality evaluation result of the large-scale voice interaction model is determined. This overcomes the shortcomings of traditional solutions, such as single evaluation dimensions and insufficient objectivity and comprehensiveness of evaluation results. By combining objective evaluation indicators and subjective evaluation indicators for quality evaluation, the evaluation process can be expanded from a single dimension that focuses on machine performance to a comprehensive dimension that takes into account objective facts and subjective experience. This makes the final quality evaluation result more comprehensive, scientific, and objective, and can provide a clear and reliable basis for subsequent iterative optimization of the model, significantly improving the user experience of the final product. Attached Figure Description
[0060] In order to make the technical solutions in the present application or the prior art clearer, the accompanying drawings needed in the embodiments or the prior art description will be briefly described below. Obviously, the accompanying drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0061] Figure 1 is a flowchart of the model quality evaluation method provided by the present application;
[0062] Figure 2 is a whole framework diagram of the quality evaluation process provided by the present application;
[0063] Figure 3 is a schematic diagram of the evaluation index and evaluation content provided by the present application;
[0064] Figure 4 is a structural schematic diagram of the model quality evaluation device provided by the present application;
[0065] Figure 5 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0066] In order to make the technical solutions in the present application or the prior art clearer, the accompanying drawings needed in the embodiments or the prior art description will be briefly described below. Obviously, the accompanying drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0067] At present, with the vigorous development of artificial intelligence, voice interaction large models have been widely applied to intelligent customer service, smart home, vehicle-mounted assistant and other scenarios, and the interaction quality directly affects the user experience and product value. However, the current quality evaluation method often excessively focuses on the accuracy rate, and ignores the user's feelings. For example, the existing evaluation method excessively relies on the objective answer accuracy rate, and pays insufficient attention to the personification, empathy ability and other subjective feelings in the interaction process, thereby causing the evaluation result to be unable to comprehensively and objectively reflect the real performance of the model, and it is also difficult to guide the effective optimization of the model.
[0068] To this end, the present application provides a model quality evaluation method, which aims to build a multi-dimensional evaluation system covering technical performance and user experience, so as to provide a set of comprehensive, scientific and objective evaluation standard for the quality of voice interaction large models, so as to solve the problem of single evaluation dimension, ignoring the subjective feelings of users, thereby causing the evaluation result to be not comprehensive, accurate and objective. Figure 1is a flowchart of a model quality evaluation method provided by the present application, as shown in Figure 1 The method comprises the following steps:
[0069] In step 110, dialogue data for model quality evaluation is determined.
[0070] In step 120, the dialogue data is input into the speech interaction large model to be evaluated, and the speech interaction large model responds to the dialogue data to obtain a response result of the dialogue data.
[0071] In step 130, based on the response result, index evaluation results of the speech interaction large model under objective evaluation indicators and subjective evaluation indicators are respectively determined.
[0072] In step 140, based on the index evaluation results of the speech interaction large model under the objective evaluation indicators and the subjective evaluation indicators, a quality evaluation result of the speech interaction large model is determined.
[0073] Specifically, before the model quality evaluation is performed, the "test questions", i.e. the dialogue data, for quality evaluation need to be prepared first. The dialogue data here can come from multiple channels. For example, it can be collected, desensitized and labeled from real online products (such as smart speakers, mobile phone voice assistants, etc.); it can also be artificially constructed by evaluation experts according to the preset evaluation target, i.e. dialogue data; it can also be generated by using automatic tools to generate large-scale dialogue data covering multiple interactive scenarios.
[0074] Here, the dialogue data can be automatically selected, such as automatically selected according to the evaluation target, randomly selected, etc., or manually selected or configured by evaluation experts according to the specific evaluation target. For example, a group of dialogue data can be selected from a large dialogue data corpus, a pre-constructed multi-scene dialogue data set for this evaluation task.
[0075] Among them, the dialogue data for model quality evaluation can be in the form of text, voice, or a combination of the two. The dialogue data can be a single-turn dialogue (i.e. one question and one answer), or a multi-turn dialogue containing context information, which is not specifically limited in the embodiments of the present application.
[0076] After determining the dialogue data for model quality evaluation, the speech interaction large model to be evaluated can be applied to process the dialogue data in the embodiments of the present application to obtain a processing result, so that the final quality evaluation can be performed based on the processing result.
[0077] Here, the voice interaction large model is a target model to be evaluated, which can be a large-scale pre-training model based on deep learning (such as a Transformer architecture) that can understand and generate natural language, such as the Xinghuo cognitive large model. The model has the ability to interact with users (such as continuous dialogue) through voice, text, etc.
[0078] Specifically, here the dialogue data can be input into the voice interaction large model so that the voice interaction large model processes it, i.e., understands and analyzes the input dialogue data, and responds to the questions involved and the questions raised, and outputs the response result. The response result here can include analysis of the dialogue data, answers to questions, etc., which can be in the form of text, synthesized audio files, or text and audio. The specific output form can be adaptively selected according to the actual situation, or can be pre-set, and the embodiments of the present application do not make specific limitations.
[0079] Further, after obtaining the response result, in the embodiments of the present application, quality evaluation can be performed based on the response result to obtain a quality evaluation result. However, considering that the evaluation based on a single index in the traditional scheme mainly focuses on accuracy and ignores the user's feelings, resulting in inaccurate and comprehensive evaluation results, in the embodiments of the present application, to improve the comprehensiveness and objectivity of model quality evaluation, subjective and objective evaluation dimensions are set to evaluate respectively, and finally the evaluation results of the two dimensions are fused to obtain a final comprehensive, accurate and reliable quality evaluation result.
[0080] Specifically, after obtaining the response result output by the voice interaction large model, the response result can be analyzed and evaluated from the objective and subjective aspects respectively. That is, in the subjective aspect, the response model is analyzed according to the pre-set subjective evaluation index to determine its performance corresponding to the subjective evaluation index, thereby obtaining the index evaluation result under the subjective evaluation index. Correspondingly, in the objective aspect, the evaluation index is also pre-set, and the response result can be analyzed under the objective evaluation index to analyze its performance, thereby obtaining the index evaluation result under the objective evaluation index. Thus, the index evaluation results of the voice interaction large model under the objective evaluation index and the subjective evaluation index are determined respectively. The index evaluation result here can be a specific score (such as 1-5 points), a grade (such as excellent, good, and poor), or a conclusion (such as pass or fail), and the embodiments of the present application do not make specific limitations.
[0081] Among them, the subjective evaluation index refers to the evaluation index that focuses more on evaluating the subjective feelings and experiences of users in the interaction process. The evaluation of such index usually needs to judge whether the model output conforms to human language habits, emotional needs and social norms. For example, the subjective evaluation index can be used to measure whether the response of the model is natural and fluent, has practical value, has the tone and emotion of a real person, and can give users a sense of empathy, etc.
[0082] The objective evaluation index refers to the evaluation index that can be calculated or judged by an automatic program or based on facts. Such index usually focuses on the logicality, factual accuracy and relevance of the content of the model output, and the performance of the machine performance level. For example, the objective evaluation index can be used to measure whether the response content of the model is correct, whether it is related to the user's question, whether it maintains logical consistency in multiple rounds of dialogue, etc.
[0083] Here, it should be noted that the subjective evaluation index and the objective evaluation index can be one or more. In the case of multiple, the index evaluation results of the voice interaction large model under each subjective evaluation index and each objective evaluation index need to be determined one by one, so as to fuse and summarize the final quality evaluation result based on this.
[0084] After that, the final quality evaluation result can be determined according to the respective index evaluation results of the model under the objective evaluation index and the subjective evaluation index. That is, the specific index evaluation results of the model under each objective evaluation index and subjective evaluation index obtained in the previous step can be synthesized, or the subjective dimension and the objective dimension can be summarized respectively, and then the synthesis results of the two dimensions are synthesized, so as to obtain a general and comprehensive quality evaluation result. The quality evaluation result is the final output of this evaluation task, which comprehensively and accurately reflects the comprehensive ability of the voice interaction large model in each dimension.
[0085] Specifically, when determining the final quality evaluation result, all index evaluation results can be directly summed, weighted summed, etc. to obtain a final comprehensive score as the quality evaluation result. According to the index evaluation results of the model under each evaluation index, a multi-dimensional capability radar chart can be generated to intuitively show the strong and weak performance of the model in different evaluation indexes. According to the index evaluation results of the model under each evaluation index, a detailed analysis report can be generated, which not only contains the scores, but also may include typical error case analysis and optimization suggestions of the model. Or other forms, the embodiments of the present application do not make specific limitation.
[0086] In the embodiment of the present application, by constructing an evaluation framework combining objective evaluation indicators and subjective evaluation indicators, the technical indicators and user experience are combined, comprehensive and objective quality evaluation can be realized, and the quality detection and optimization of various dialogue systems can be widely applied.
[0087] The model quality evaluation method provided by the present application inputs the dialogue data into the speech interaction large model to be evaluated to obtain a response result output by the speech interaction large model; based on the response result, index evaluation results of the speech interaction large model under objective evaluation indicators and subjective evaluation indicators are determined respectively, and based on the index evaluation results under the objective evaluation indicators and the subjective evaluation indicators, a quality evaluation result of the speech interaction large model is determined, which overcomes the defects of single evaluation dimension, insufficient objective and comprehensive evaluation result in the traditional scheme, and combines the objective evaluation indicators and the subjective evaluation indicators for quality evaluation, so that the evaluation process is expanded from a single dimension focusing on machine performance to a comprehensive dimension considering objective facts and subjective experience, so that the finally generated quality evaluation result is more comprehensive, scientific and objective, and then a clear and reliable basis can be provided for subsequent iteration optimization of the model, and the user experience of the final product is significantly improved.
[0088] Based on the above embodiment, the subjective evaluation indicators include empathy degree;
[0089] The index evaluation result of the speech interaction large model under the empathy degree is determined based on the following steps;
[0090] An initial interaction state is determined, and the initial interaction state includes a user emotional state;
[0091] Based on the initial interaction state and the dialogue data, the user and the speech interaction large model are simulated to perform multi-round interaction to obtain dialogue data of the multi-round interaction, and emotional changes and emotional scores of each round of interaction in the multi-round interaction; the dialogue data of each round of interaction includes a response result output by the speech interaction large model; the input of each round of interaction in the multi-round interaction is determined based on the emotional changes and the emotional scores of the previous round of interaction;
[0092] Based on the dialogue data of the multi-round interaction and the emotional changes and the emotional scores of each round of interaction in the multi-round interaction, the index evaluation result of the speech interaction large model under the empathy degree is determined.
[0093] Specifically, considering that the current quality evaluation scheme only evaluates from a single dimension, excessively relies on the correct answer rate, and ignores the user experience, especially fails to fully consider the user's intention and emotional appeal, and the change of user emotion in the interaction process, resulting in the problem that the evaluation result is not objective and comprehensive, and the user experience is not good.
[0094] Based on this, in the embodiment of the present application, in order to improve the comprehensiveness of the evaluation, realize a more in-depth evaluation closer to the real feelings of the user, the evaluation index under the subjective dimension can be set to empathy, that is, the subjective evaluation dimension can include empathy. The empathy evaluates the high-order social cognitive ability of the model, specifically measures whether the voice interaction large model can show the ability of understanding, listening and resonance to the user's emotion when interacting with the user, which is a key indicator to judge whether a voice interaction large model is "intelligent" and "warm", and is particularly important for application scenarios such as emotional companionship and psychological counseling.
[0095] In detail, in the embodiment of the present application, the determination process of the index evaluation result of the voice interaction large model under the empathy is actually an automatic evaluation process simulating real interaction. The process quantitatively evaluates the empathy ability of the voice interaction large model by dynamically simulating an interaction between a user with emotional changes and the voice interaction large model.
[0096] Specifically, the determination process of the index evaluation result under the empathy can include: first, the initial state of the interaction, that is, the initial interaction state, which contains the user emotional state. Here, the initial interaction state can be understood as a basic scenario or background story set for this empathy evaluation. This state provides a starting point and basic constraints for subsequent simulated interactions. A core element is the preset user emotional state, which can be diverse, such as "feeling depressed due to exam failure", "feeling anxious due to heavy work pressure", "feeling surprised due to receiving unexpected gifts", etc. This user emotional state lays the foundation for simulating the first round of user interaction behavior.
[0097] Subsequently, the initial interaction state and the dialogue data can be used to simulate multiple rounds of interaction between the user and the voice interaction large model to obtain the emotional changes and emotional scores in the current interaction round in the multiple rounds of interaction. Figure 2 is the overall framework of the quality evaluation process provided by the present application, as Figure 2 shown, the multiple rounds of interaction are not performed by the user, but by an automated evaluation system to "play" or simulate the user. The system generates the first round of user input (for example, if the user emotional state is "depressed", the first round of user input can be "I feel terrible today") according to the determined initial interaction state, and sends the dialogue data and the first round of user input to the voice interaction large model to be evaluated to make the model respond.
[0098] Next, the system can receive the response result of the voice interaction large model, analyze the response result, and evaluate its impact on the simulated user's emotion, thereby obtaining two key outputs, i.e., the emotional change and the emotional score of the first round of interaction. For example, if the answer of the model is "Don't be sad, cheer up", the system may determine that this answer is relatively perfunctory and fails to effectively empathize, so the emotional change may be "from depressed to irritable", and the emotional score of this round will be lower. Conversely, if the answer of the model is "It sounds like you had a very bad day today, would you like to tell me what happened?", the system may determine that this answer shows listening and care, the emotional change may be "from depressed to slightly calm", and the emotional score will be higher.
[0099] Then, the user input of the second round of interaction can be determined according to the emotional change and the emotional score of the first round of interaction, and the dialogue data of the first round of interaction. This is essentially a process of dynamic feedback, i.e., the automated evaluation system will determine what the simulated user should say in the next round according to the emotional change and the emotional score generated after the last round of interaction. This makes the entire dialogue process more realistic and coherent. For example, if the user's emotion changes to "irritable", the user input of the next round of interaction may be "You simply don't understand, stop making empty promises"; if the emotion changes to "slightly calm", the user input of the next round of interaction may be "Well, I messed up an important project", thereby leading the dialogue to a deeper level.
[0100] After that, the generated user input of the second round of interaction, together with the previous dialogue history (i.e., the dialogue data of the first round of interaction), can be given to the voice interaction large model to obtain the response result of the model output in the second round of interaction. By analyzing the response result, the emotional change and the emotional score of the second round of interaction can be obtained, and the user input of the third round of interaction can be generated based on this, and so on. The automated evaluation system can simulate the user to interact with the voice interaction large model for multiple rounds, and can obtain the dialogue data of multiple rounds of interaction, and the emotional change and the emotional score of each round of interaction.
[0101] Finally, the index evaluation result of the voice interaction large model under the empathy degree can be evaluated according to the dialogue data of the multi-round interaction and the emotional change and emotional score of each round of interaction in the multi-round interaction. Here, specifically, after a preset number of rounds (for example, 5 rounds or 10 rounds) of simulated interaction, the automatic evaluation system collects the data of the entire interaction process, including the user input of each round, the response result of the model, the emotional change and the emotional score, to obtain the dialogue data of the multi-round interaction and the emotional change and the emotional score of each round of interaction in the multi-round interaction. Based on the dialogue data and the emotional change and the emotional score of each round of interaction, the system finally gives a comprehensive index evaluation result under the empathy degree. The result can be a weighted average of the emotional scores of each round, or a grade or score determined according to the final emotional state and the interaction trajectory. For example, if the emotional state of the simulated user finally changes from negative to positive after multiple rounds of interaction, the empathy score of the model will be higher.
[0102] In the embodiments of the present application, the performance of the voice interaction large model in the emotional interaction scene is dynamically evaluated in an automatic and reproducible manner, overcoming the defects of the current evaluation scheme that cannot simulate real dialogue flow and the high cost, strong subjectivity and inconsistent standards of pure manual evaluation, making the evaluation of the empathy ability of the model more objective, quantitative and efficient, and providing a scientific evaluation method for developing high-quality voice interaction products that can truly understand and care for users.
[0103] Based on the above embodiments, the index evaluation result of the voice interaction large model under the empathy degree is determined, including:
[0104] Based on the dialogue data of the multi-round interaction and the emotional change and emotional score of each round of interaction in the multi-round interaction, an interaction empathy score is determined;
[0105] An empathy evaluation result of the evaluation expert under the empathy degree for the dialogue data of the multi-round interaction and the emotional change and emotional score of each round of interaction in the multi-round interaction is obtained;
[0106] Based on the interaction empathy score and the empathy evaluation result, the index evaluation result of the voice interaction large model under the empathy degree is determined.
[0107] Specifically, to further improve the accuracy and reliability of the index evaluation result of the voice interaction large model under the empathy degree, the present application provides a man-machine collaborative evaluation mechanism, that is, a comprehensive evaluation mechanism combining automatic evaluation and manual evaluation. Based on this evaluation mechanism, the process of determining the index evaluation result of the voice interaction large model under the empathy degree can specifically include the following steps:
[0108] First, the interactive empathy score can be determined according to the dialogue data of the multi-turn interaction and the emotional changes and emotional scores of each turn of the multi-turn interaction. This process is essentially the machine evaluation in the human-machine collaborative evaluation mechanism, that is, automatic evaluation. The process of this automatic evaluation specifically includes: after the automatic evaluation system simulates the user and the voice interaction large model to complete the multi-turn interaction, the system can calculate a quantitative score, that is, the interactive empathy score, according to the recorded dialogue data of the multi-turn interaction and the emotional changes and emotional scores of each turn of the multi-turn interaction. This interactive empathy score provides a preliminary and objective quantitative benchmark for the empathy ability of the model.
[0109] Here, the interactive empathy score can be determined based on various strategies. For example, it can be the cumulative value or weighted average value of the emotional score of each turn, and the result of the calculation is the final interactive empathy score; it can also be the difference value of the emotional state change before and after the interaction, or the calculation result of a complex function considering multiple factors such as the number of interaction rounds and emotional trajectory, and the present embodiment does not make specific limitations.
[0110] At the same time, the empathy evaluation result of the evaluation expert under the empathy degree for the dialogue data of the multi-turn interaction and the emotional changes and emotional scores of each turn of the multi-turn interaction can be obtained. That is, the human evaluation link is introduced as a supplement and calibration to the automatic evaluation process. Specifically, here the dialogue data of the multi-turn interaction and the emotional changes and emotional scores of each turn of the multi-turn interaction can be presented to one or more professional evaluation experts after completing the multi-turn interaction.
[0111] The evaluation expert is a person who is trained and familiar with the empathy evaluation standard. He / she will evaluate the multi-turn interaction process from the perspective of human and judge whether the model's performance truly achieves the effect of empathy. For example, the evaluation expert will make judgments according to a series of more detailed evaluation standards, which can include: (1) whether to empathize and the accuracy of empathy, that is, whether the model accurately identifies the simulated user's emotions and gives appropriate responses; (2) whether there are offensive, irony and other impolite language in the response result; (3) whether there is a conflict with the pre-set personality (such as personality, tone, values, etc.) of the model.
[0112] After completing the evaluation, the evaluation expert will give an empathy evaluation result. The result can be a specific score, such as scoring on a Likert scale of 1-5; or a qualitative rating, such as "excellent", "good", "poor", etc., and the present embodiment does not make specific limitations.
[0113] After that, the index evaluation result of the voice interaction large model under the empathy can be determined according to the interactive empathy score and the empathy evaluation result; that is, the interactive empathy score obtained by the above-mentioned automatic evaluation and the empathy evaluation result obtained by manual evaluation can be fused according to the set fusion strategy, so that the final and more reliable index evaluation result under the empathy is obtained.
[0114] Here, the fusion strategy can be weighted average of the interactive empathy score (specific score) and the empathy evaluation result, such as final score = automatic score 0.5 + manual score 0.5; or the empathy evaluation result fed back by the evaluation expert can be used as a correction factor of the interactive empathy score, such as if the interactive empathy score is 4 points, but the expert rating is "poor", the final result may be greatly reduced. This fusion mechanism ensures that the evaluation result is efficient (benefiting from automation) and consistent with human perception (benefiting from expert evaluation).
[0115] In the embodiment of the application, the efficiency and reproducibility of automatic evaluation are combined with the depth and accuracy of manual evaluation to form a complementary evaluation mode. This "combination of automatic indicators and manual evaluation" effectively overcomes the limitations of a single evaluation method, making the index evaluation result of the empathy ability of the voice interaction large model objective and quantitative, and deep and reliable, so as to more accurately reflect the real level of the model and provide a more valuable reference for model optimization.
[0116] Based on the above embodiment, the initial state of the interaction is determined, including:
[0117] Based on the dialogue data, the interaction elements are determined; the interaction elements include interaction roles, interaction backgrounds, interaction targets and potential intentions;
[0118] Based on the interaction elements, the initial state of the interaction containing the user emotional state is constructed.
[0119] Specifically, to make the simulated interaction in automatic evaluation more realistic and diverse, so as to more comprehensively test the empathy ability of the model, in the embodiment of the application, when determining the initial state of the interaction, the four elements in the set basic scene or background story, i.e. interaction roles, interaction backgrounds, interaction targets and potential intentions, can be first determined by using the dialogue data. These elements are collectively referred to as interaction elements, based on which the basic framework of the simulated interaction can be defined.
[0120] Here, the interaction role defines the basic character information of the simulated user, which can include age, occupation, personality characteristics (such as introverted, optimistic), values, etc. For example, an interaction role can be set as "a young designer who is sensitive in character and just entered the job market".
[0121] The interaction context describes the specific situation or event cause of the dialogue, which provides a macro environment for the dialogue. For example, the interaction context can be "the designer feels frustrated because an important design proposal is completely denied by the customer".
[0122] The interaction goal clearly defines the purpose of simulating the user initiating this dialogue, which determines the core direction of the dialogue. For example, the interaction goal can be "complain to the model about the trouble, hoping to get some comfort and encouragement".
[0123] The underlying intention reveals the deeper psychological needs of the simulated user below the surface goal, which is crucial for the evaluation of the model's empathy ability. For example, in the above scenario, the underlying intention of the simulated user can be "hoping that the model can recognize his efforts and provide some specific and constructive suggestions to help him get out of trouble, rather than just empty comfort".
[0124] These interaction elements can be determined according to specific dialogue data. The automated evaluation system can generate the first round of user input based on these interaction elements when the simulated user interacts with the large-scale voice interaction model, so as to construct a basic scene or background story.
[0125] After that, the interaction initial state containing the user's emotional state can be constructed according to the interaction elements. That is, the system can comprehensively construct a specific and executable interaction initial state according to the above structured interaction elements. The core of this construction process is to deduce a user emotional state that is highly matched with the current scene or background story, which provides a starting point and basic constraint for subsequent simulated interaction.
[0126] For example, after determining the interaction elements (such as interaction role: sensitive designer; interaction context: proposal is denied; interaction goal: seek comfort; underlying intention: seek recognition and suggestions), the system can comprehensively judge the emotional state of the simulated user at the beginning of the dialogue, which is likely to be "depressed, self-doubting and slightly anxious". This interaction initial state containing the specific user emotional state will serve as the starting point of the simulated interaction story for the automated evaluation system. The system will generate the first round of user input (for example, "I feel terrible, the proposal I spent a month on is worthless") based on this interaction initial state, thereby officially starting the simulated interaction process with the large-scale voice interaction model.
[0127] In the embodiment of the present application, by defining the four interactive elements of interactive role, interactive background, interactive target and potential intention, the constructed scene or story is no longer scattered, random single sentence, but complete story line with cause and effect, which can greatly enrich the comprehensiveness of the scene, so as to more accurately test the comprehensive response ability of the model in the face of different people, different difficulties.
[0128] Based on the above embodiment, the subjective evaluation index includes personification degree;
[0129] The index evaluation result of the voice interaction large model under the personification degree is determined based on the following steps;
[0130] Content repetition degree detection is performed based on the response result to obtain a repetition degree detection result;
[0131] The personification degree evaluation result of the evaluation expert under the personification degree for the response result is obtained;
[0132] Based on the content repetition degree detection and the personification degree evaluation result, the index evaluation result of the voice interaction large model under the personification degree is determined.
[0133] Specifically, considering that the current evaluation scheme is difficult to quantify and standardize, resulting in inaccurate and unreliable evaluation results, in the embodiment of the present application, to further improve the evaluation effect of the subjective dimension and optimize the evaluation quality, the subjective evaluation index can include the personification degree. The personification degree here is an index for judging whether the response result of the voice interaction large model has similar human behavior characteristics, making the user feel like talking to a real person, rather than a cold machine. It focuses on the naturalness and fluency of the interaction, and is an important dimension for measuring whether the model can integrate into human daily communication habits. A high personification degree model should have the thought, tone and interaction habit of a real person.
[0134] In detail, to ensure the comprehensiveness and accuracy of the personification degree evaluation, in the embodiment of the present application, an automatic evaluation combined with manual evaluation is used to determine the index evaluation result under the personification degree. Specifically, here, automatic evaluation can be performed first to evaluate the personification degree of the model from a key and quantifiable angle, that is, content repetition degree detection is performed on the response result to obtain a repetition degree detection result. Considering that in natural conversation, language and sentence repetition is usually avoided, therefore, the content repetition degree can be an important index for measuring whether the response result of the model output is "mechanical" or "templated".
[0135] Specifically, the content repetition detection here can be an analysis and detection on the response results output by the model in one or more interactions, such as detecting whether the sentences in the response results are natural, and detecting whether there is unnecessary content in the response results; or, for example, detecting whether there is repetition of words and sentences within a single round of response results, and detecting whether the same words or sentence patterns appear multiple times in the context of a multi-round dialogue. For example, if the model starts with "My answer to your question is..." in the output of three consecutive interactions, the content repetition degree may be high.
[0136] After the detection is completed, a quantitative repetition detection result can be obtained. The result can be a specific repetition rate value, or a repetition degree level (such as "high", "medium", "low") based on a preset threshold, or other values, which are not specifically limited in the embodiments of the present application.
[0137] At the same time, manual evaluation can be performed. That is, since anthropomorphism is a comprehensive subjective concept, repetition detection alone cannot completely cover it. Therefore, in the embodiments of the present application, manual evaluation is introduced to make up for the shortcomings of automatic detection.
[0138] Specifically, the response results output by the model can be presented to one or more professional evaluation experts. The evaluation experts will comprehensively judge the anthropomorphism of the model from the perspective of human intuition. Here, the evaluation criteria followed by the evaluation experts are usually more comprehensive than automatic detection, and can include but are not limited to: (1) sentence smoothness, coherence, etc.; (2) whether there is too much written expression that does not conform to daily oral communication habits; (3) the conciseness of the response, whether the response result is lengthy. For example, a response that is completely grammatically correct but full of academic jargon may score very high in automatic detection, but the anthropomorphism may be very low in the eyes of experts. After completing the evaluation, the evaluation experts will give a subjective anthropomorphism evaluation result, for example, a score of 1-5.
[0139] Then, the content repetition detection result obtained by automatic evaluation and the anthropomorphism evaluation result obtained by manual evaluation can be fused to obtain a final, comprehensive anthropomorphism index evaluation result. Here, the fusion method can be diverse, for example, the two can be combined by weighting to obtain a final anthropomorphism score.
[0140] In the embodiments of the present application, the anthropomorphism of the model is comprehensively evaluated by combining the objective quantification ability of the machine and the comprehensive judgment ability of the human, which not only guarantees the efficiency and consistency of the evaluation, but also makes the evaluation result more accurate and reliable.
[0141] Based on the above embodiments, the content repetition degree is detected based on the response results to obtain a repetition detection result, which includes:
[0142] performing sentence naturalness detection based on the response result to obtain a naturalness detection result;
[0143] performing word-sentence similarity detection based on the response result to obtain a similarity detection result;
[0144] performing content redundancy detection based on the response result to obtain a redundancy detection result;
[0145] determining a repetition detection result based on at least one of the naturalness detection result, the similarity detection result, and the redundancy detection result.
[0146] Specifically, the process of performing content repetition detection based on the response result to obtain a repetition detection result can specifically include:
[0147] First, the response result can be detected from multiple aspects to obtain multi-dimensional detection results. That is, the response result can be detected from one or more dimensions of sentence naturalness, word-sentence similarity, and content redundancy, so as to obtain at least one of the naturalness detection result, the similarity detection result, and the redundancy detection result.
[0148] Here, the sentence naturalness detection mainly focuses on whether the response result contains an expression mode that does not conform to human language habits. For example, a language model can be used to calculate the perplexity of the text content of the response result. The lower the perplexity, the more fluent and natural the sentence is. It can also be detected whether there are too many written language or professional terms in the response result, which are not common in daily conversations.
[0149] The word-sentence similarity detection is mainly used to identify the repetition in the content. For example, the semantic similarity between different clauses in the single-round response result can be calculated to determine whether the same word or sentence pattern is repeated. More importantly, the similarity between the response result of the current round and the historical multi-round response results can be calculated to determine whether the model tends to use the same or similar sentence patterns and phrases to answer different questions. For example, if the model frequently uses a fixed sentence pattern such as “My answer to your question is...” in multiple rounds of dialogue, the word-sentence similarity detection will find this pattern.
[0150] The content redundancy detection focuses on whether the information amount of the response result meets the requirement of conciseness. For example, it can be analyzed whether there are unnecessary modifiers, catchphrases, or redundant information irrelevant to the question in the response result. For example, when the user only needs a simple “yes” or “no” answer, the model provides a long explanatory text, and at this time, the content redundancy can be determined.
[0151] Then, the one or more detection results can be synthesized to determine a final redundancy detection result. For example, the naturalness detection result, the similarity detection result, and the redundancy detection result can be weighted and averaged to obtain a total redundancy score.
[0152] In the embodiments of the present application, through multi-dimensional refinement detection, various repetitive problems affecting anthropomorphism can be more accurately captured from different angles, so that the results of automatic evaluation are more targeted and interpretable.
[0153] Based on the above embodiments, the subjective evaluation indicators include richness;
[0154] The index evaluation result of the voice interaction large model under the richness is determined based on the following steps;
[0155] Based on the dialogue data, content expandability detection is performed on the response result, and based on the expanded content obtained by the content expandability detection and the context content corresponding to the expanded content in the response result, content appropriateness detection is performed to obtain an appropriateness detection result;
[0156] Based on the response result, content practicality detection is performed to obtain a practicality detection result;
[0157] Based on the scene type corresponding to the dialogue data, content richness detection is performed on the response result to obtain a richness detection result;
[0158] Based on at least one of the appropriateness detection result, the practicality detection result, and the richness detection result, an index evaluation result of the voice interaction large model under the richness is determined.
[0159] Specifically, the subjective evaluation indicators also include richness. Richness is an indicator for measuring whether the response result of the voice interaction large model has sufficient information, practical value, and the ability to reasonably expand and actively guide interaction. A high-quality voice interaction large model should not have a short and perfunctory answer, but should provide valuable and rich information on the basis of meeting the user's core needs, and even guide the dialogue to a more interesting and deeper direction. Therefore, in the embodiments of the present application, a subjective evaluation indicator of richness is also provided to evaluate the quality of the model.
[0160] In detail, the determination process of the index evaluation result of the voice interaction large model under the richness can specifically include:
[0161] Firstly, content extensibility detection can be performed, i.e., content extensibility detection is performed on the response result according to the dialogue data, to identify whether the model response provides additional relevant information after answering the core question. For example, when asking "What is the weather like in A city today?", a simple answer is "Sunny, 25 degrees", while an extended answer can be "Sunny, 25 degrees. However, the ultraviolet rays are strong, please pay attention to sun protection when going out. In addition, it may rain in the next two days." Through content extensibility detection, the part of the response result that exceeds the core answer can be automatically detected, and the content of this part is the extension content. Then, content appropriateness detection can be performed. That is, since only extension is not enough, the extended content must also be appropriate. Therefore, after obtaining the extension content, content appropriateness detection is performed to analyze the correlation degree between the extension content and the context content. Inappropriate extension may confuse or disturb the user. For example, when the user asks a serious academic question, the model extends an unrelated joke after answering, which is inappropriate. Based on this correlation analysis, an appropriateness detection result can be obtained.
[0162] At the same time, content usefulness detection can be performed, i.e., content usefulness detection is performed on the response result, to detect whether the model output response result has practical value, rather than being broad, empty or superficial. For example, when the user asks "What do you recommend for fun?", if the model answers "There are many fun things", the usefulness of this response result is low. If the model answers "You can go for a walk in the nearby park, or watch a recently released movie xxx, which has a good rating", the usefulness of this response result is high. In the process of content usefulness detection, the usefulness can be quantified by information entropy, keyword density, whether executable suggestions are included, etc., to obtain a usefulness detection result.
[0163] Similarly, content richness detection can be performed, i.e., content richness detection is performed on the response result according to the scene type corresponding to the dialogue data, to obtain a richness detection result. For example, when performing an explicit task (such as "set an alarm for 7 am tomorrow"), the response result should be concise, and excessive richness will reduce efficiency. In a casual conversation scene (such as "let's chat"), the response result is expected to be more rich and interesting. Therefore, in the embodiment of the present application, when performing richness detection, the scene type corresponding to the dialogue data can be identified first, and then whether the length and information amount of the current response result are appropriate according to the scene type. For example, in a non-special scene (such as casual conversation), if the response result of the model is always too short (such as only answering "um", "ok", etc.), the richness is low.
[0164] Afterwards, the one or more detection results above can be synthesized to determine the final index evaluation result under the richness. For example, the suitability detection result, the practicality detection result and the richness detection result can be weighted and summed to obtain a total richness score.
[0165] In the embodiments of the present application, the richness of the model is evaluated from multiple angles such as the suitability of the extension, the practicality of the content and the scene matching degree of the response length, objective and quantifiable model quality evaluation is realized, the defect that the value of the answer information is difficult to quantify in the traditional evaluation scheme is overcome, and effective data support is provided for model optimization.
[0166] Based on the above embodiments, the objective evaluation index includes the correlation degree;
[0167] The index evaluation result of the voice interaction large model under the correlation degree is determined based on the following steps;
[0168] Single-round correlation detection is performed based on the dialogue data and the response result to obtain a single-round correlation detection result;
[0169] Multi-round correlation detection is performed based on the dialogue data and the response result to obtain a multi-round correlation detection result;
[0170] Viewpoint consistency detection is performed based on the dialogue data and the response result to obtain a viewpoint consistency detection result;
[0171] Instruction following detection is performed based on the dialogue data and the response result to obtain an instruction following detection result;
[0172] The index evaluation result of the voice interaction large model under the correlation degree is determined based on at least one of the multi-round correlation detection result, the viewpoint consistency detection result and the instruction following detection result, and the single-round correlation detection result.
[0173] Specifically, considering that the current evaluation scheme is excessively accurate, the correctness of the answer is ensured, but in the multi-round dialogue process, it ignores the "coherence" of the dialogue, that is, whether the model can remember the historical dialogue, whether it will forget the historical dialogue, and whether it can keep logical consistency in the multi-round dialogue. These high-order abilities are often ignored in the past evaluation process.
[0174] Therefore, in the embodiments of the present application, the evaluation index also sets the correlation degree, Figure 3 is a schematic diagram of the evaluation index and evaluation content provided by the present application, as Figure 3 shown, the correlation degree aims to measure the matching degree of the response result of the voice interaction large model and the user's intention, and the ability to keep logical coherence and memory in continuous multi-round dialogue. The model with low correlation degree often has problems such as answering irrelevant questions, contradictions, forgetting user instructions, etc.
[0175] In detail, the index evaluation results of the voice interaction large model under the relevance degree can be obtained through the following steps: first, attention can be paid to whether the response result of the model is directly related to the current round of user input (query) in a single interaction. For example, when the user asks "How high is Mount Everest?", the model's answer should be about the height of the mountain, not about the weather nearby. By calculating the semantic similarity or correlation score between the user input and the response result of the model in a single round of dialogue data, the single-round relevance detection result can be obtained. If the correlation score is lower than the preset threshold, it can be determined as irrelevant; otherwise, it is relevant.
[0176] At the same time, the context understanding and memory ability of the model in continuous dialogue, i.e. multi-round relevance, can be investigated. Unlike single-round relevance, multi-round relevance detection requires the model to understand the user's intent in the current round in combination with historical dialogue information. For example, the dialogue history is "User: Help me find a nearby Sichuan restaurant. Model: OK, I found 'CCC' for you. User: How much is the average consumption there?" In this case, the user's "it" clearly refers to "CCC", and the model needs to understand this reference relationship and answer the consumption information, rather than asking "What is it?". By analyzing whether the current response result in the dialogue data utilizes the historical information, the multi-round relevance can be evaluated, and the multi-round relevance detection result can be obtained.
[0177] The consistency of the model's output of viewpoints, facts or positions in the entire dialogue session can also be investigated, unless the user explicitly guides it to correct it. For example, if the model says "A is the capital of B" at the beginning of the dialogue, and "C is the capital of B" at the end of the dialogue, there is a discrepancy in the viewpoints. By tracking the key information in the dialogue data and the response result, and comparing whether they are consistent or whether there are logical contradictions, viewpoint consistency detection can be achieved, and the viewpoint consistency detection result can be obtained. That is, if the key information is contradictory, it can be determined that the viewpoints are inconsistent, and vice versa.
[0178] The model's ability to accurately and completely execute the user's multi-step instructions or follow the user's set rules can also be investigated, which is particularly important in task-based and challenging dialogues. That is, instruction following detection can be performed according to the dialogue data and the response result, and the instruction following detection result can be obtained. For example, the user's instruction is "Help me write a five-character quatrains about spring, and do not contain the word 'flower'", at this time, it is necessary to detect whether the response result output by the model satisfies the three conditions of "five-character quatrains", "about spring" and "without the word 'flower'". The omission or error of any one condition means that the instruction following fails.
[0179] After that, the one or more detection results can be synthesized to determine the final index evaluation result under the correlation degree. It is worth noting that single-turn correlation is the basis for effective dialogue, so it is a must-check item. Multi-turn correlation, viewpoint consistency and instruction following are higher-order correlation capabilities, and one or more detections can be selectively performed according to the evaluation requirements. As a preferred, the final index evaluation result under the correlation degree in the embodiment of the application can be obtained by weighted summation of the scores of these detection results.
[0180] In the embodiment of the application, the dialogue correlation capability of the model is evaluated from multiple aspects such as single-turn correlation, context memory, logical consistency and instruction execution, which can accurately locate the short board of the model in dialogue logic and coherence, and provides strong technical support for improving the interaction effectiveness and reliability of the model.
[0181] Based on the above embodiment, the objective evaluation index includes accuracy;
[0182] The index evaluation result of the voice interaction large model under the accuracy is determined based on the following steps;
[0183] Based on the dialogue data, the content accuracy of the response result is detected to obtain an accuracy detection result;
[0184] Based on the dialogue data, the content accuracy of the response result is detected to obtain an accuracy detection result;
[0185] Based on the dialogue data, the content accuracy of the response result is detected to obtain an accuracy detection result;
[0186] Based on at least one of the accuracy detection result, the fabrication degree detection result and the fluency detection result, the index evaluation result of the voice interaction large model under the accuracy is determined.
[0187] Specifically, referring to Figure 2 and Figure 3 It can be seen that the objective evaluation index also includes accuracy. The accuracy aims to measure the performance of the response result of the voice interaction large model in terms of factual correctness and language standardization. The model with low accuracy often provides incorrect information, which is misleading, and even outputs sentences that are not fluent and difficult to understand, which seriously affects its usability and credibility.
[0188] In detail, the index evaluation result of the voice interaction large model under the accuracy can be determined by the following steps:
[0189] Firstly, the accuracy of the model reply can be investigated, that is, whether the key information contained in the response result of the model meets the objective fact. For example, when the user asks "What year is this year?", the correct information "2025" must be contained in the answer of the model. In the embodiment of the present application, when the content accuracy is detected, the key entity and relationship in the response result can be compared with an existing authoritative knowledge base to automatically determine its correctness. If the comparison is consistent, the content is accurate, and the content accuracy detection result is obtained, which can be a binary judgment of correct or incorrect, or a quantitative score based on the correct information coverage rate.
[0190] At the same time, the content fabrication degree of the response result can be detected to identify the "hallucination" phenomenon commonly seen in large models, that is, whether the model fabricates non-existent facts, characters, events or data, which is more serious than simply answering incorrectly because it is very misleading. For example, when asked about the author of a paper, the model may "seriously" fabricate a non-existent author name. Through cross-validation, traceability detection and other methods, the fabrication degree of the content in the response result can be accurately detected to obtain the content fabrication degree detection result.
[0191] The response result can also be subjected to sentence fluency detection to evaluate whether the response result output by the model is in compliance with the language level, that is, whether there are syntax errors, sentence fluency, improper word usage and other problems. An accurate response result not only has correct content, but also clear and fluent expression. In the embodiment of the present application, natural language processing tools (such as grammar checkers, language models, etc.) can be used to automatically analyze the response result and identify possible syntax problems. For example, it can be detected whether the subject and predicate are consistent, whether there are component defects, whether there are misspelled words, etc. to obtain the sentence fluency detection result, which can be a score calculated based on the number of syntax errors, or other content, which is not limited in the embodiment of the present application.
[0192] Then, the above one or more detection results can be integrated to determine the index evaluation result under the accuracy. In the embodiment of the present application, the content accuracy focuses on "whether correct", the content fabrication degree focuses on "whether real", and the sentence fluency focuses on "whether the expression is standard". According to the focus of the evaluation, one of the detection results can be used alone, or they can be combined with weights to obtain a total accuracy score.
[0193] In the embodiment of the present application, the accuracy of the model is evaluated from the aspects of fact correctness, content authenticity and language standardization, which can effectively identify various errors of the model in knowledge mastery and language generation, and can provide accurate measurement and optimization direction for the improvement of model reliability and professionalism.
[0194] Based on the above embodiments, step 110 comprises:
[0195] From the multi-scene dialogue data set, dialogue data for model quality evaluation is selected;
[0196] The multi-scene dialogue data set contains at least two of task-oriented dialogue data, game-oriented dialogue data, emotional companion dialogue data, casual dialogue data, and multi-round challenge dialogue data.
[0197] Specifically, the dialogue data for model quality evaluation can be selected from the multi-scene dialogue data set. Here, specifically, a large-scale, high-quality multi-scene dialogue data set can be pre-constructed. The data set contains dialogue data of multiple scenes. When performing a specific evaluation task, appropriate dialogue data can be selected from the multi-scene dialogue data set as input for this evaluation. This selection can be random sampling to ensure the generalization of the evaluation; or targeted selection, i.e. selecting according to the evaluation goal of the evaluation task to deeply test a certain specific ability of the model.
[0198] Among them, the dialogue data contained in the multi-scene dialogue data set can be at least two of task-oriented dialogue data, game-oriented dialogue data, emotional companion dialogue data, casual dialogue data, and multi-round challenge dialogue data.
[0199] Here, task-oriented dialogue data is dialogue data in an interactive scene with a clear intention and goal. For example, querying the weather, setting reminders, searching for information, booking tickets, etc. This type of dialogue data is mainly used to test the model's search, induction, and instruction execution capabilities.
[0200] Game-oriented dialogue data is targeted at special interactive scenarios with interesting and rule-based characteristics. For example, Chinese proverb connection, poetry connection, telling jokes, riddle solving, role-playing games, etc. This type of dialogue data is mainly used to test the model's knowledge reserve, understanding ability, and multi-round dialogue memory ability.
[0201] Emotional companion dialogue data is targeted at scenarios of emotional support and companionship. For example, users confiding their troubles to the model, sharing joy, seeking comfort, etc. This type of dialogue data is used to test the model's high-level social cognitive abilities, such as empathy and anthropomorphism.
[0202] Casual dialogue data corresponds to scenarios without specific purposes and open-ended conversations. For example, "How was your day?" "Let's talk about movies," etc. This type of dialogue data is mainly used to test the model's dialogue coherence, breadth of knowledge, and naturalness and interest of language.
[0203] Multi-round challenge type includes complex context logic, multi-step continuous instructions, viewpoint consistency challenge, role-playing consistency challenge, etc. For example, the model is required to ''first help me plan a route from S to M, then find all Chinese restaurants on the route with a score higher than 4.5, exclude spicy cuisine, and finally sort by price from low to high''. Such data is mainly used to deeply test the context logic reasoning ability, long-term memory ability, instruction following ability and self-consistency of the model.
[0204] In the embodiment of the present application, a multi-scene dialogue data set is constructed, and quality evaluation is performed based on the multi-scene dialogue data set, so that the comprehensiveness of model quality evaluation can be ensured, and the quality evaluation is no longer limited to a single question and answer scene, but covers various real application scenarios from simple tasks to complex challenges, from information acquisition to emotional communication. The evaluation result can more comprehensively reflect the comprehensive ability of the model, avoid the phenomenon of ''poor performance in some subjects'', and provide a solid data and evaluation foundation for creating an ''all-around'' voice interaction large model.
[0205] The model quality evaluation device provided by the present application is described below. The model quality evaluation device described below can be correspondingly referred to the model quality evaluation method described above.
[0206] Figure 4 is a structural schematic diagram of the model quality evaluation device provided by the present application, as Figure 4 shown, the device comprises:
[0207] The data determination unit 410 is configured to determine dialogue data for model quality evaluation.
[0208] The data processing unit 420 is configured to input the dialogue data into a voice interaction large model to be evaluated, and obtain a response result of the dialogue data by responding to the dialogue data by the voice interaction large model.
[0209] The result determination unit 430 is configured to determine an index evaluation result of the voice interaction large model under objective evaluation indexes and subjective evaluation indexes based on the response result.
[0210] The quality evaluation unit 440 is configured to determine a quality evaluation result of the voice interaction large model based on the index evaluation result of the voice interaction large model under the objective evaluation indexes and the subjective evaluation indexes.
[0211] The model quality evaluation device provided by the application inputs dialogue data into a speech interaction large model to be evaluated to obtain a response result output by the speech interaction large model; based on the response result, index evaluation results of the speech interaction large model under objective evaluation indexes and subjective evaluation indexes are determined respectively, and based on the index evaluation results under the objective evaluation indexes and the subjective evaluation indexes, a quality evaluation result of the speech interaction large model is determined, thereby overcoming the defects of single evaluation dimension, insufficient objectivity and comprehensiveness of the evaluation result in the traditional scheme, and the quality evaluation is combined with the objective evaluation indexes and the subjective evaluation indexes, so that the evaluation process is expanded from a single dimension focusing on machine performance to a comprehensive dimension taking into account objective facts and subjective experience, thereby making the finally generated quality evaluation result more comprehensive, scientific and objective, and further providing an explicit and reliable basis for subsequent iteration optimization of the model, and significantly improving the user experience of the final product.
[0212] Based on the above embodiment, the subjective evaluation index includes empathy degree;
[0213] The result determination unit 430 is configured to:
[0214] Determine an initial state of interaction, wherein the initial state of interaction includes a user emotional state;
[0215] Based on the initial state of interaction and the dialogue data, simulate a plurality of rounds of interaction between the user and the speech interaction large model to obtain dialogue data of the plurality of rounds of interaction and emotional changes and emotional scores of each round of interaction in the plurality of rounds of interaction; the dialogue data of each round of interaction includes a response result output by the speech interaction large model; the input of each round of interaction in the plurality of rounds of interaction is determined based on the emotional changes and the emotional scores of the previous round of interaction;
[0216] Based on the dialogue data of the plurality of rounds of interaction and the emotional changes and the emotional scores of each round of interaction in the plurality of rounds of interaction, determine an index evaluation result of the speech interaction large model under the empathy degree.
[0217] Based on the above embodiment, the result determination unit 430 is configured to:
[0218] Based on the dialogue data of the plurality of rounds of interaction and the emotional changes and the emotional scores of each round of interaction in the plurality of rounds of interaction, determine an interaction empathy degree score;
[0219] Obtain an empathy degree evaluation result of an evaluation expert in evaluating the dialogue data of the plurality of rounds of interaction and the emotional changes and the emotional scores of each round of interaction in the plurality of rounds of interaction under the empathy degree;
[0220] Based on the interaction empathy degree score and the empathy degree evaluation result, determine an index evaluation result of the speech interaction large model under the empathy degree.
[0221] Based on the above embodiment, the result determination unit 430 is configured to:
[0222] Based on the dialogue data, determine an interaction element; the interaction element includes an interaction role, an interaction background, an interaction target, and a potential intention;
[0223] Based on the interaction element, an initial interaction state containing a user emotional state is constructed.
[0224] Based on the above embodiment, the subjective evaluation index includes anthropomorphism;
[0225] The result determination unit 430 is configured to:
[0226] Based on the response result, content repetition detection is performed to obtain a repetition detection result;
[0227] An anthropomorphism evaluation result of the evaluation expert under the anthropomorphism for the response result is obtained;
[0228] Based on the content repetition detection and the anthropomorphism evaluation result, an index evaluation result of the voice interaction large model under the anthropomorphism is determined.
[0229] Based on the above embodiment, the result determination unit 430 is configured to:
[0230] Based on the response result, sentence naturalness detection is performed to obtain a naturalness detection result;
[0231] Based on the response result, word and sentence similarity detection is performed to obtain a similarity detection result;
[0232] Based on the response result, content redundancy detection is performed to obtain a redundancy detection result;
[0233] Based on at least one of the naturalness detection result, the similarity detection result, and the redundancy detection result, the repetition detection result is determined.
[0234] Based on the above embodiment, the subjective evaluation index includes richness;
[0235] The result determination unit 430 is configured to:
[0236] Based on the dialogue data, the response result is subjected to content expansibility detection, and based on the expanded content obtained by content expansibility detection and the context content corresponding to the expanded content in the response result, content appropriateness detection is performed to obtain an appropriateness detection result;
[0237] Based on the response result, content practicality detection is performed to obtain a practicality detection result;
[0238] The content richness of the response result is detected based on a scene type corresponding to the dialogue data, to obtain a richness detection result.
[0239] The index evaluation result of the voice interaction large model under the richness is determined based on at least one of the suitability detection result, the practicability detection result, and the richness detection result.
[0240] Based on the above embodiment, the objective evaluation index includes relevance;
[0241] The result determination unit 430 is configured to:
[0242] Single-turn relevance detection is performed based on the dialogue data and the response result, to obtain a single-turn relevance detection result.
[0243] Multi-turn relevance detection is performed based on the dialogue data and the response result, to obtain a multi-turn relevance detection result.
[0244] Viewpoint consistency detection is performed based on the dialogue data and the response result, to obtain a viewpoint consistency detection result.
[0245] Instruction following detection is performed based on the dialogue data and the response result, to obtain an instruction following detection result.
[0246] The index evaluation result of the voice interaction large model under the relevance is determined based on at least one of the multi-turn relevance detection result, the viewpoint consistency detection result, and the instruction following detection result, and the single-turn relevance detection result.
[0247] Based on the above embodiment, the objective evaluation index includes accuracy;
[0248] The result determination unit 430 is configured to:
[0249] Content accuracy detection is performed on the response result based on the dialogue data, to obtain an accuracy detection result.
[0250] Content fabrication degree detection is performed on the response result based on the dialogue data, to obtain a fabrication degree detection result.
[0251] Sentence fluency detection is performed on the response result based on the dialogue data, to obtain a fluency detection result.
[0252] The index evaluation result of the voice interaction large model under the accuracy is determined based on at least one of the accuracy detection result, the fabrication degree detection result, and the fluency detection result.
[0253] Based on the above embodiment, the data determination unit 410 is configured to:
[0254] dialogue data for model quality evaluation is selected from the multi-scene dialogue data set;
[0255] The multi-scene dialogue data set contains at least two of task-oriented dialogue data, game-oriented dialogue data, emotional companion dialogue data, casual dialogue data, and multi-round challenge dialogue data.
[0256] Figure 5 An example of an entity structure diagram of an electronic device is shown in Figure 5 As shown, the electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communications bus 540. The processor 510 can invoke the logic instructions in the memory 530 to execute the model quality evaluation method, which includes determining dialogue data for model quality evaluation; inputting the dialogue data into a voice interaction large model to be evaluated, and obtaining a response result of the dialogue data by the voice interaction large model responding to the dialogue data; based on the response result, determining index evaluation results of the voice interaction large model under objective evaluation indicators and subjective evaluation indicators; and based on the index evaluation results of the voice interaction large model under the objective evaluation indicators and the subjective evaluation indicators, determining a quality evaluation result of the voice interaction large model.
[0257] In addition, the logic instructions in the memory 530 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium, includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0258] In another aspect, the present application also provides a computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions that, when executed by a computer, enable the computer to perform the model quality evaluation method provided by any of the above methods, the method comprising: determining dialogue data for model quality evaluation; inputting the dialogue data into a voice interaction large model to be evaluated, and obtaining a response result of the dialogue data by responding to the dialogue data by the voice interaction large model; determining index evaluation results of the voice interaction large model under objective evaluation indicators and subjective evaluation indicators based on the response result; and determining a quality evaluation result of the voice interaction large model based on the index evaluation results of the voice interaction large model under the objective evaluation indicators and the subjective evaluation indicators.
[0259] In another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement a model quality evaluation method provided by any of the above methods, the method comprising: determining dialogue data for model quality evaluation; inputting the dialogue data into a voice interaction large model to be evaluated, and obtaining a response result of the dialogue data by responding to the dialogue data by the voice interaction large model; determining index evaluation results of the voice interaction large model under objective evaluation indicators and subjective evaluation indicators based on the response result; and determining a quality evaluation result of the voice interaction large model based on the index evaluation results of the voice interaction large model under the objective evaluation indicators and the subjective evaluation indicators.
[0260] The device embodiments described above are merely illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0261] From the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in terms of the contribution to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0262] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the same; although the present application has been described in detail with reference to the foregoing examples, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A model quality evaluation method, characterized by, The method comprises the following steps: determining dialogue data for model quality evaluation; inputting the dialogue data into a voice interaction large model to be evaluated, and obtaining a response result of the dialogue data by responding to the dialogue data by the voice interaction large model; determining an index evaluation result of the voice interaction large model under an objective evaluation index and a subjective evaluation index based on the response result; determining a quality evaluation result of the voice interaction large model based on the index evaluation result of the voice interaction large model under the objective evaluation index and the subjective evaluation index; the subjective evaluation index comprises empathy degree; the index evaluation result of the voice interaction large model under the empathy degree is determined based on the following steps: determining an initial state of interaction, wherein the initial state of interaction comprises a user emotional state; based on the initial state of interaction and the dialogue data, simulating a user and the voice interaction large model to perform a plurality of rounds of interaction, obtaining dialogue data of the plurality of rounds of interaction, and emotional changes and emotional scores of each round of interaction in the plurality of rounds of interaction; the dialogue data of each round of interaction comprises a response result output by the voice interaction large model; the input of each round of interaction in the plurality of rounds of interaction is determined based on the emotional changes and emotional scores of the previous round of interaction; based on the dialogue data of the plurality of rounds of interaction, and the emotional changes and emotional scores of each round of interaction in the plurality of rounds of interaction, determining the index evaluation result of the voice interaction large model under the empathy degree.
2. The model quality evaluation method of claim 1, wherein, The determination of the index evaluation result of the voice interaction large model under the empathy degree comprises: based on the dialogue data of the plurality of rounds of interaction, and the emotional changes and emotional scores of each round of interaction in the plurality of rounds of interaction, determining an interaction empathy degree score; obtaining an empathy degree evaluation result of an evaluation expert for evaluating the dialogue data of the plurality of rounds of interaction, and the emotional changes and emotional scores of each round of interaction in the plurality of rounds of interaction under the empathy degree; based on the interaction empathy degree score and the empathy degree evaluation result, determining the index evaluation result of the voice interaction large model under the empathy degree.
3. The model quality evaluation method of claim 1, wherein, The determination of the initial state of interaction comprises: based on the dialogue data, determining an interaction element; the interaction element comprises an interaction role, an interaction background, an interaction target, and a potential intention; based on the interaction element, constructing an initial state of interaction comprising a user emotional state.
4. The model quality evaluation method according to any one of claims 1 to 3, characterized in that, The subjective evaluation index comprises anthropomorphism degree; the index evaluation result of the voice interaction large model under the anthropomorphism degree is determined based on the following steps: performing content repetition degree detection based on the response result to obtain a repetition degree detection result; obtaining an anthropomorphism degree evaluation result of an evaluation expert for evaluating the response result under the anthropomorphism degree; based on the content repetition degree detection and the anthropomorphism degree evaluation result, determining the index evaluation result of the voice interaction large model under the anthropomorphism degree.
5. The model quality evaluation method of claim 4, wherein, The content repetition degree detection based on the response result to obtain a repetition degree detection result comprises: performing sentence naturalness detection based on the response result to obtain a naturalness detection result; performing word and sentence similarity detection based on the response result to obtain a similarity detection result; Perform content redundancy detection based on the response result to obtain a redundancy detection result; Determine the repetition detection result based on at least one of the naturalness detection result, the similarity detection result, and the redundancy detection result.
6. The model quality evaluation method according to any one of claims 1 to 3, characterized in that, The subjective evaluation index includes richness; The index evaluation result of the voice interaction large model under the richness is determined based on the following steps; Perform content expansibility detection on the response result based on the dialogue data, and perform content appropriateness detection based on the expanded content obtained through content expansibility detection and the context content corresponding to the expanded content in the response result to obtain an appropriateness detection result; Perform content practicality detection on the response result based on the dialogue data to obtain a practicality detection result; Perform content richness detection on the response result based on the scene type corresponding to the dialogue data to obtain a richness detection result; Determine the index evaluation result of the voice interaction large model under the richness based on at least one of the appropriateness detection result, the practicality detection result, and the richness detection result.
7. The model quality evaluation method according to any one of claims 1 to 3, characterized in that, The objective evaluation index includes relevance; The index evaluation result of the voice interaction large model under the relevance is determined based on the following steps; Perform single-round relevance detection based on the dialogue data and the response result to obtain a single-round relevance detection result; Perform multi-round relevance detection based on the dialogue data and the response result to obtain a multi-round relevance detection result; Perform viewpoint consistency detection based on the dialogue data and the response result to obtain a viewpoint consistency detection result; Perform instruction followability detection based on the dialogue data and the response result to obtain an instruction followability detection result; Determine the index evaluation result of the voice interaction large model under the relevance based on at least one of the multi-round relevance detection result, the viewpoint consistency detection result, and the instruction followability detection result, and the single-round relevance detection result.
8. The model quality evaluation method according to any one of claims 1 to 3, characterized in that, The objective evaluation index includes accuracy; The index evaluation result of the voice interaction large model under the accuracy is determined based on the following steps; Perform content accuracy detection on the response result based on the dialogue data to obtain an accuracy detection result; Perform content fictionality detection on the response result based on the dialogue data to obtain a fictionality detection result; Perform sentence fluency detection on the response result based on the dialogue data to obtain a fluency detection result; Determine the index evaluation result of the voice interaction large model under the accuracy based on at least one of the accuracy detection result, the fictionality detection result, and the fluency detection result.
9. The model quality evaluation method according to any one of claims 1 to 3, characterized in that, The determination of the dialogue data for model quality evaluation includes: Select dialogue data for model quality evaluation from a multi-scene dialogue data set; The multi-scene dialogue data set includes at least two of task-type dialogue data, game-type dialogue data, emotional companion-type dialogue data, casual conversation-type dialogue data, and multi-round challenge-type dialogue data.
10. A model quality evaluation device characterized by comprising: It includes: A data determination unit is configured to determine dialogue data for model quality evaluation. a data processing unit, configured to input the dialogue data to a voice interaction large model to be evaluated, and obtain a response result of the dialogue data by responding to the dialogue data by the voice interaction large model; a result determining unit, configured to determine an index evaluation result of the voice interaction large model under an objective evaluation index and a subjective evaluation index based on the response result; a quality evaluation unit, configured to determine a quality evaluation result of the voice interaction large model based on the index evaluation result of the voice interaction large model under the objective evaluation index and the subjective evaluation index; the subjective evaluation index comprises empathy degree; the index evaluation result of the voice interaction large model under the empathy degree is determined based on the following steps; determining an initial interaction state, wherein the initial interaction state comprises a user emotional state; based on the initial interaction state and the dialogue data, simulating a user and the voice interaction large model to perform a plurality of rounds of interaction, obtaining dialogue data of the plurality of rounds of interaction, and emotional change and emotional score of each round of interaction in the plurality of rounds of interaction; the dialogue data of each round of interaction comprises a response result output by the voice interaction large model; input of each round of interaction in the plurality of rounds of interaction is determined based on emotional change and emotional score of a previous round of interaction; based on the dialogue data of the plurality of rounds of interaction, and the emotional change and emotional score of each round of interaction in the plurality of rounds of interaction, determining the index evaluation result of the voice interaction large model under the empathy degree.
11. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, the processor executes the computer program to implement the model quality evaluation method in any one of claims 1 to 9. 12.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, the computer program is executed by the processor to implement the model quality evaluation method in any one of claims 1 to 9.
Citation Information
Patent Citations
Evaluation method and device based on multiple rounds of dialogues, equipment and storage medium
CN120046719A