Construction method and device of evaluation data set of large language model, computer device and readable storage medium

CN120705527BActive Publication Date: 2026-10-09GUANGZHOU QUWAN NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510815123.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2026-10-09
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

[0003]相关技术中,在构建对大语言模型进行对话状态感知能力评估的评估数据集时,对话状态分类较为粗糙,无法全面评估模型在真实对话中的对话状态感知能力

Benefits of technology

[0040] Thus, by acquiring dialogue state topic sentences corresponding to different evaluation dimensions, which contain dialogue state information and topic scene information, the dialogue state perception capability of large language models in complex and varied dialogue scenarios can be evaluated more objectively. By expanding the dialogue state topic sentences, the expanded dialogue state topic sentences include preset key elements for analyzing dialogue states. Then, by optimizing the expanded dialogue state topic sentences, the optimized sentences to be analyzed contain richer dialogue state information, thereby effectively ensuring the diversity and complexity of the sentences to be analyzed. Based on the sentences to be analyzed and the corresponding dialogue state analysis result set, an evaluation dataset is constructed to evaluate the dialogue state perception capability of the pre-trained large language model. The dialogue state analysis result set includes correct analysis results and at least one analysis result that differs from the correct analysis result, thereby increasing the testing difficulty of the pre-trained large language model, comprehensively evaluating the model's dialogue state perception capability in real dialogues, and thus more accurately evaluating the dialogue state perception capability of the large language model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705527B_ABST
    Figure CN120705527B_ABST
Patent Text Reader

Abstract

The application relates to a construction method and device of an evaluation data set of a large language model, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: obtaining dialogue state topic sentences corresponding to different evaluation dimensions; expanding the dialogue state topic sentences to obtain expanded dialogue state topic sentences; optimizing the expanded dialogue state topic sentences to obtain to-be-analyzed sentences; the richness of dialogue state information of the to-be-analyzed sentences is higher than that of the expanded dialogue state topic sentences; taking the to-be-analyzed sentences and a corresponding dialogue state analysis result set as an evaluation data set; the dialogue state analysis result set comprises correct analysis results and at least one analysis result different from the correct analysis results; and the evaluation data set is used for evaluating the dialogue state perception ability of a pre-trained large language model. The dialogue state perception ability of the large language model can be more accurately evaluated by using the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for constructing an evaluation dataset for a large language model. Background Technology

[0002] With the widespread application of large language models in dialogue systems, there is an increasing demand for evaluating the dialogue state perception capabilities of large language models. Dialogue state analysis (such as dialogue state recognition and dialogue state understanding) has become an important indicator for measuring model performance.

[0003] In related technologies, when constructing an evaluation dataset to assess the dialogue state perception ability of large language models, the dialogue state classification is relatively coarse and cannot comprehensively evaluate the model's dialogue state perception ability in real dialogues.

[0004] Therefore, there is a problem in the related technologies that the assessment of the dialogue state perception capability of large language models is not accurate enough. Summary of the Invention

[0005] Based on this, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for constructing an evaluation dataset for a large language model that can more accurately assess the dialogue state awareness capability of a large language model, in order to address the aforementioned technical problems.

[0006] Firstly, this application provides a method for constructing an evaluation dataset for a large language model, including:

[0007] Obtain dialogue state topic sentences corresponding to different evaluation dimensions; the dialogue state topic sentences contain dialogue state information and topic scene information; the evaluation dimensions are dimensions used to evaluate the dialogue state perception ability of the pre-trained large language model.

[0008] The dialogue state topic statement is expanded to obtain an expanded dialogue state topic statement; the expanded dialogue state topic statement includes preset key elements for analyzing the dialogue state.

[0009] The expanded dialogue state topic statement is optimized to obtain the statement to be analyzed; the richness of the dialogue state information of the statement to be analyzed is higher than that of the expanded dialogue state topic statement.

[0010] Obtain the dialogue state analysis result set corresponding to the statement to be analyzed, and use the statement to be analyzed and the corresponding dialogue state analysis result set as the evaluation dataset; the dialogue state analysis result set includes correct analysis results and at least one analysis result that is different from the correct analysis results; the evaluation dataset is used to evaluate the dialogue state awareness capability of the pre-trained large language model.

[0011] In one embodiment, the correct analysis result includes a first correct sub-result and a second correct sub-result, and each analysis result different from the correct analysis result includes a first sub-result and a second sub-result. The step of obtaining the dialogue state analysis result set corresponding to the statement to be analyzed includes:

[0012] Obtain the first correct sub-result and the corresponding second correct sub-result for the statement to be analyzed; the first correct sub-result represents the correct dialogue state category corresponding to the statement to be analyzed; the second correct sub-result represents the reason for the generation of the correct dialogue state category.

[0013] Obtain at least one first sub-result that is different from the first correct sub-result; the first sub-result represents a dialogue state category that is different from the correct dialogue state category;

[0014] Obtain at least one second sub-result that is different from the second correct sub-result; the second sub-result represents the reason for the occurrence of a dialogue state category that is different from the correct dialogue state category.

[0015] In one embodiment, the method further includes:

[0016] Obtain a first analysis result set corresponding to the statement to be analyzed; the first analysis result set includes a first category label and at least one second category label that is different from the first category label; the first category label represents the correct dialogue state category corresponding to the statement to be analyzed; the second category label represents a dialogue state category that is different from the correct dialogue state category.

[0017] The first category label is used as the first correct sub-result, and the second category label is used as the first sub-result.

[0018] In one embodiment, the method further includes:

[0019] Obtain the second analysis result set corresponding to the statement to be analyzed; the second analysis result set includes first reason information and at least one second reason information that is different from the first reason information; the first reason information represents the reason for the generation of the correct dialogue state category; the second reason information represents the reason for the generation of a dialogue state category that is different from the correct dialogue state category;

[0020] The first reason information is used as the second correct sub-result, and the second reason information is used as the second sub-result.

[0021] In one embodiment, obtaining the dialogue state topic statements corresponding to different evaluation dimensions includes:

[0022] Obtain a set of dialogue statements from a pre-set multi-turn dialogue dataset; the dialogue statements in the set of dialogue statements contain the dialogue state information.

[0023] The set of dialogue statements is filtered to obtain a target set of dialogue statements; the target dialogue statements in the target set of dialogue statements meet the preset dialogue quality conditions.

[0024] Obtain the expanded target dialogue statement and use the expanded target dialogue statement as the dialogue state topic statement; the expanded target dialogue statement is obtained by expanding the topic scenario of the target dialogue statement according to the definition of the evaluation dimension.

[0025] In one embodiment, optimizing the expanded dialogue state topic statement to obtain the statement to be analyzed includes:

[0026] The expanded dialogue state topic statement is input into a pre-trained statement optimization model, which outputs the statement to be analyzed.

[0027] The pre-trained statement optimization model uses the following training method:

[0028] Obtain a training sample set; the training sample set includes the original statement and the corresponding tag-optimized statement; the richness of the dialogue state information of the tag-optimized statement is higher than that of the original statement.

[0029] The original statement is input into the statement optimization model to be trained, and the predicted optimized statement corresponding to the original statement is output.

[0030] Based on the difference between the predicted optimized statement and the corresponding labeled optimized statement corresponding to the original statement, the model parameters of the statement optimization model to be trained are adjusted and iterative training continues until the trained statement optimization model meets the training termination condition, at which point training stops, and the pre-trained statement optimization model is obtained.

[0031] Secondly, this application also provides an apparatus for constructing an evaluation dataset for a large language model, comprising:

[0032] The statement acquisition module is used to acquire dialogue state topic statements corresponding to different evaluation dimensions; the dialogue state topic statements contain dialogue state information and topic scene information; the evaluation dimensions are dimensions used to evaluate the dialogue state perception ability of the pre-trained large language model.

[0033] An extension module is used to expand the dialogue state topic statement to obtain an expanded dialogue state topic statement; the expanded dialogue state topic statement includes preset key elements for analyzing the dialogue state.

[0034] An optimization module is used to optimize the expanded dialogue state topic statement to obtain a statement to be analyzed; the richness of the dialogue state information of the statement to be analyzed is higher than the richness of the dialogue state information of the expanded dialogue state topic statement.

[0035] A construction module is used to obtain the dialogue state analysis result set corresponding to the statement to be analyzed, and to use the statement to be analyzed and the corresponding dialogue state analysis result set as an evaluation dataset; the dialogue state analysis result set includes correct analysis results and at least one analysis result that is different from the correct analysis results; the evaluation dataset is used to evaluate the dialogue state perception capability of the pre-trained large language model.

[0036] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, implements the steps of the method described above.

[0037] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.

[0038] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the above-described method.

[0039] The aforementioned method, apparatus, computer device, computer-readable storage medium, and computer program product for constructing the evaluation dataset of the large language model involve: acquiring dialogue state topic sentences corresponding to different evaluation dimensions; the dialogue state topic sentences contain dialogue state information and topic scene information; the evaluation dimensions are dimensions used to evaluate the dialogue state perception ability of the pre-trained large language model; expanding the dialogue state topic sentences to obtain expanded dialogue state topic sentences; the expanded dialogue state topic sentences include preset key elements for analyzing the dialogue state; optimizing the expanded dialogue state topic sentences to obtain the sentences to be analyzed; the richness of dialogue state information in the sentences to be analyzed is higher than that in the expanded dialogue state topic sentences; acquiring the dialogue state analysis result set corresponding to the sentences to be analyzed, and using the sentences to be analyzed and the corresponding dialogue state analysis result set as the evaluation dataset; the dialogue state analysis result set includes correct analysis results and at least one analysis result different from the correct analysis results; the evaluation dataset is used to evaluate the dialogue state perception ability of the pre-trained large language model.

[0040] Thus, by acquiring dialogue state topic sentences corresponding to different evaluation dimensions, which contain dialogue state information and topic scene information, the dialogue state perception capability of large language models in complex and varied dialogue scenarios can be evaluated more objectively. By expanding the dialogue state topic sentences, the expanded dialogue state topic sentences include preset key elements for analyzing dialogue states. Then, by optimizing the expanded dialogue state topic sentences, the optimized sentences to be analyzed contain richer dialogue state information, thereby effectively ensuring the diversity and complexity of the sentences to be analyzed. Based on the sentences to be analyzed and the corresponding dialogue state analysis result set, an evaluation dataset is constructed to evaluate the dialogue state perception capability of the pre-trained large language model. The dialogue state analysis result set includes correct analysis results and at least one analysis result that differs from the correct analysis result, thereby increasing the testing difficulty of the pre-trained large language model, comprehensively evaluating the model's dialogue state perception capability in real dialogues, and thus more accurately evaluating the dialogue state perception capability of the large language model. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart illustrating a method for outputting evaluation information of a large language model in one embodiment;

[0043] Figure 2 This is a flowchart illustrating a method for constructing an evaluation dataset for a large language model in one embodiment.

[0044] Figure 3 This is a flowchart illustrating a method for constructing an evaluation dataset for another large language model in one embodiment.

[0045] Figure 4 This is a flowchart illustrating a method for constructing an evaluation dataset for a large language model, as described in another embodiment.

[0046] Figure 5 This is a flowchart illustrating a dialogue state response in one embodiment.

[0047] Figure 6 This is a flowchart illustrating an evaluation information output method for a large language model in another embodiment;

[0048] Figure 7 A structural block diagram of a device for constructing an evaluation dataset for a large language model in one embodiment;

[0049] Figure 8 This is a structural block diagram of an evaluation information output device for a large language model in one embodiment;

[0050] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0052] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0053] In one embodiment, such as Figure 1As shown, a method for outputting evaluation information of a large language model is provided. This embodiment illustrates the method by applying it to a computer device. It is understood that the computer device can be a terminal, a server, or a system including both a terminal and a server. In this embodiment, the method includes the following steps:

[0054] Step S110: Obtain the evaluation dataset corresponding to each evaluation stage of the preset dialogue state evaluation process.

[0055] The dialogue state assessment process is a process for evaluating the dialogue state perception capability of a pre-trained large language model.

[0056] Among them, dialogue state awareness refers to the ability of a large language model to understand the user's dialogue state and generate responses that conform to the dialogue state during the dialogue process.

[0057] Dialogue state refers to the user's state during a conversation, such as emotional state. In practical applications, the ability to perceive dialogue state can be termed emotional perception ability.

[0058] Each evaluation stage has different evaluation dimensions.

[0059] The evaluation dataset for each evaluation stage includes the statements to be analyzed under different evaluation dimensions and the corresponding dialogue state analysis result set.

[0060] The statement to be analyzed can refer to a dialogue statement selected from an existing multi-turn dialogue dataset.

[0061] Among them, the statement to be analyzed contains dialogue state information, which refers to information related to the dialogue state.

[0062] The dialogue state analysis result set corresponding to each statement to be analyzed includes the correct analysis result and at least one analysis result that is different from the correct analysis result.

[0063] The analysis result refers to the result obtained by analyzing the statement to be analyzed. The analysis result may include, but is not limited to, the results of dialogue state identification and understanding, dialogue decision results, and dialogue response results.

[0064] The dialogue state recognition and understanding result refers to the recognition result of the dialogue state category (in practical applications, it can refer to the emotion category, such as happy, angry, expectant, etc.) corresponding to the statement to be analyzed, as well as the analysis result of the reason for the generation of this dialogue state category.

[0065] Among them, the dialogue decision outcome refers to the decision of which response mode (such as the appeasement mode, the companionship mode, the casual conversation mode, etc.) to advance the next step of the dialogue.

[0066] The dialogue response result refers to the text information used to respond to the statement to be analyzed.

[0067] In this context, "correct analysis result" can refer to a manually generated correct analysis result for the statement to be analyzed. In practical applications, multiple experts in psychology or sociology can collaborate to design corresponding correct analysis results for the statement to be analyzed.

[0068] In this process, analytical results that differ from the correct results can be generated by the large language model, or generated by the large language model and then corrected by data labelers to ensure logical coherence and a degree of confusion. This semi-automated process (large model generation and manual correction) effectively reduces data construction costs.

[0069] In specific implementation, when evaluating the dialogue state perception capability of the pre-trained large language model, the evaluation dataset corresponding to each evaluation step of the preset dialogue state evaluation process can be obtained first. The evaluation dataset corresponding to each evaluation step includes the statement to be analyzed under different evaluation dimensions and the corresponding dialogue state analysis result set. The dialogue state analysis result set corresponding to each statement to be analyzed includes the correct analysis result and at least one analysis result that is different from the correct analysis result.

[0070] In practical applications, the dialogue state analysis result set corresponding to each statement to be analyzed can include one correct analysis result and three analysis results that are different from the correct analysis result. The number of analysis results that are different from the correct analysis result can be set according to actual needs, and no specific limit is made here.

[0071] Step S120: In each evaluation stage, the evaluation dataset corresponding to the evaluation stage is input into the pre-trained large language model, and the target prediction result corresponding to the sentence to be analyzed in the evaluation dataset is output.

[0072] The target prediction result is the analysis result selected by the pre-trained large language model from the set of dialogue state analysis results corresponding to the input statement to be analyzed.

[0073] In practice, in each evaluation stage, the evaluation dataset corresponding to the evaluation stage can be input into the pre-trained large language model, instructing the pre-trained large language model to select the analysis result that matches the statement to be analyzed from the dialogue state analysis result set corresponding to the statement to be analyzed, and output it as the target prediction result.

[0074] Step S130: Compare the correct analysis results and the corresponding target prediction results of the sentences to be analyzed in the evaluation dataset to obtain the evaluation information of the pre-trained large language model in each evaluation stage.

[0075] The evaluation information is used to characterize the dialogue state awareness capability of the pre-trained large language model.

[0076] In practice, computer equipment can use automated scripts to compare the correct analysis results and the corresponding target prediction results for the statements to be analyzed in the evaluation dataset. This determines whether the target prediction results output by the pre-trained large language model for the statements to be analyzed are the same as the corresponding correct analysis results. By counting the number of correct analysis results selected by the pre-trained large language model in the evaluation dataset for each evaluation stage, the dialogue state perception capability of the pre-trained large language model in that evaluation stage is quantified, thus outputting the evaluation information of the pre-trained large language model in that evaluation stage. In this way, by counting the number of correct analysis results selected by the pre-trained large language model in the evaluation dataset for each evaluation stage, the evaluation information of the pre-trained large language model in each evaluation stage is obtained.

[0077] In the above-mentioned method for outputting evaluation information of a large language model, the evaluation dataset corresponding to each evaluation stage of a pre-defined dialogue state evaluation process is obtained. The dialogue state evaluation process is a process for evaluating the dialogue state perception capability of a pre-trained large language model. The evaluation dataset corresponding to each evaluation stage includes the statement to be analyzed under different evaluation dimensions and the corresponding dialogue state analysis result set. The statement to be analyzed contains dialogue state information. The dialogue state analysis result set corresponding to each statement to be analyzed includes the correct analysis result and at least one analysis result different from the correct analysis result. In each evaluation stage, the evaluation dataset corresponding to the evaluation stage is input into the pre-trained large language model, and the target prediction result corresponding to the statement to be analyzed in the evaluation dataset is output. The target prediction result is the analysis result selected by the pre-trained large language model in the dialogue state analysis result set corresponding to the input statement to be analyzed. The correct analysis result and the corresponding target prediction result corresponding to the statement to be analyzed in the evaluation dataset are compared to obtain the evaluation information of the pre-trained large language model in each evaluation stage. The evaluation information is used to characterize the dialogue state perception capability of the pre-trained large language model.

[0078] Thus, by acquiring the evaluation dataset corresponding to each evaluation stage of the pre-defined dialogue state evaluation process, this application, compared to related technologies that only evaluate the dialogue state perception capability of a large language model at a single stage, can obtain a more comprehensive evaluation dataset for evaluating the large language model. Simultaneously, the evaluation dataset corresponding to each evaluation stage includes the statement to be analyzed under different evaluation dimensions and the corresponding dialogue state analysis result set, which can more objectively evaluate the dialogue state perception capability of the large language model in complex and varied dialogue scenarios, and can more comprehensively evaluate the performance of the large language model, improving the accuracy and reliability of the large language model evaluation. Furthermore, the dialogue state analysis result set corresponding to each statement to be analyzed includes the correct analysis result and at least one analysis result different from the correct analysis result. In the evaluation phase, the evaluation dataset corresponding to the evaluation phase is input into the large language model. The output is the target prediction result selected by the large language model in the dialogue state analysis result set corresponding to the statement to be analyzed. The correct analysis result corresponding to the statement to be analyzed is compared with the corresponding target prediction result to determine whether the large language model has selected the correct analysis result. Compared with simply inputting the statement to be analyzed into the large language model and scoring the output analysis result, this application can directly evaluate whether the large language model has the ability to understand and select the correct dialogue state analysis result, resulting in higher evaluation accuracy. Furthermore, it can efficiently evaluate the dialogue state perception ability of the large language model under different evaluation phases and dimensions, improving evaluation efficiency and reducing evaluation costs. In summary, this application establishes a multi-phase, multi-dimensional, unified, and standardized evaluation system to evaluate the large language model from multiple aspects, ensuring the accuracy and reliability of the evaluation of the large language model's dialogue state perception ability.

[0079] In practical applications, the evaluation phase of the dialogue state evaluation process can include three stages: dialogue state recognition and understanding, dialogue state management, and general dialogue state response. In some embodiments, the evaluation phase of the dialogue state evaluation process can further include a dialogue state response stage for specific scenarios. The dialogue state perception capability includes dialogue state recognition and understanding capability, dialogue state management capability, and dialogue state response capability. The evaluation dataset corresponding to the dialogue state recognition and understanding stage is used to evaluate the dialogue state recognition and understanding capability of the large language model; the analysis results in the evaluation dataset corresponding to the dialogue state recognition and understanding stage can be the dialogue state recognition and understanding results. The evaluation dataset corresponding to the dialogue state management stage is used to evaluate the dialogue state management capability of the large language model; the analysis results in the evaluation dataset corresponding to the dialogue state management stage can be the dialogue decision results. The evaluation dataset corresponding to the dialogue state response stage is used to evaluate the dialogue state response capability of the large language model; the analysis results in the evaluation dataset corresponding to the dialogue state response stage can be the dialogue response results.

[0080] In practical applications, dialogue state recognition and understanding can be termed emotion recognition and understanding; dialogue state management can be termed emotion management; and dialogue state response can be termed emotion response. Below, we will specifically introduce the different evaluation dimensions included in each evaluation stage and how to construct the dialogue state analysis result sets corresponding to different evaluation stages. Specifically:

[0081] 1. Emotion Recognition and Understanding

[0082] The emotion recognition and understanding benchmark primarily evaluates the dialogue state perception capability of large language models from the following four assessment dimensions:

[0083] (1) Sources of Emotion (also known as sources of dialogue state): These include internal emotions and external emotions. These sources of emotion can refer to emotions generated internally by an individual or emotions triggered by external events.

[0084] (2) Modes of Emotional Expression (also known as Dialogue Expression Modes): These include direct expression and indirect expression. Direct expression involves describing one's inner feelings, while indirect expression involves expressing inner feelings through non-emotional language dialogue.

[0085] (3) Context of Emotion (also known as Dialogue State Context): This covers everyday contexts, social interaction contexts, work and study contexts, and special event contexts. These contexts correspond to different social or personal activity scenarios, and the emotional expressions involved may change with the context. Special event contexts include emotional reactions in unconventional situations.

[0086] (4) Dynamic Changes in Emotion (also known as Dynamic Changes in Dialogue State): This includes Emotion Fluctuations and Emotion Transitions. Emotion Fluctuations refer to the changes in the intensity of emotions in different situations, such as from pleasure to sadness or from anxiety to relaxation; while Emotion Transitions emphasize the process of gradually transitioning from one emotion to another.

[0087] Each evaluation dimension is further subdivided into several subcategories, covering most possible emotional scenarios in life. This classification system not only meticulously categorizes emotions from multiple dimensions but also combines direct and indirect emotional expressions, internal and external emotional sources, and complex, multi-layered emotional change contexts. This refined emotion classification structure will help comprehensively evaluate the recognition and understanding capabilities of large language models in diverse emotional scenarios, providing an effective benchmark and guidance for further improving the emotional intelligence of dialogue systems.

[0088] After determining the baseline classification and subclasses, the statements to be analyzed and the analysis results in each subclass can be designed to ensure the quality and scientific rigor of the test.

[0089] In practical applications, the evaluation dataset for the emotion recognition and understanding stage includes sentences to be analyzed that correspond to different emotion sources, different emotion expression methods, different emotion contexts, and different dynamic changes in emotion. This is to evaluate whether the large model can accurately understand the emotion category and the reason for the emotion in the sentences to be analyzed across different evaluation dimensions.

[0090] Specifically, such as Figure 2 As shown, a method for constructing an evaluation dataset for a large language model is provided to explain how to construct the evaluation dataset corresponding to the emotion recognition and understanding stages. The method includes the following steps:

[0091] Step S210: Obtain the dialogue state topic statements corresponding to different evaluation dimensions.

[0092] The dialogue state topic statement contains dialogue state information and topic scenario information.

[0093] Thematic scenario information refers to information related to the thematic scenario of the dialogue, such as the thematic scenario of a dialogue being buying airline tickets, medical consultation, or financial consultation.

[0094] In this embodiment, the evaluation dimension is used to evaluate the dialogue state recognition and understanding ability of the pre-trained large language model. In the emotion recognition and understanding stage, the evaluation dimension may include the sources of emotion, the modes of emotion expression, the context of emotion, and the dynamic changes in emotion.

[0095] In practice, computer devices can acquire dialogue state topic statements corresponding to different evaluation dimensions. For example, a dialogue state topic statement could be "the sense of accomplishment after successfully completing a challenging task," with the corresponding theme scenario being task completion and the dialogue state including pride. Alternatively, a dialogue state topic statement could be "the sadness of permanently losing an important memento," with the corresponding theme scenario being item loss and the dialogue state including sadness.

[0096] In some embodiments, the computer device can acquire a multi-turn dialogue dataset and, according to the definition of the evaluation dimension, for example, according to the definition of each dialogue state information subclass under the evaluation dimension, expand the dialogue statements containing dialogue state information into topic scenarios to obtain dialogue state topic statements.

[0097] The definitions of the various dialogue state information subcategories (such as internal emotion, external emotion, direct expression, indirect expression, etc.) under the evaluation dimension of the emotion recognition and understanding process have been explained above and will not be repeated here.

[0098] In practical applications, the dialogue status statement can be named the emotion statement.

[0099] Step S220: Expand the dialog state topic statement to obtain the expanded dialog state topic statement.

[0100] The expanded dialogue state topic statement includes preset key elements for analyzing the dialogue state.

[0101] Among these, the key preset elements may include the character, the dialogue state (emotion), and the context.

[0102] In practice, after obtaining the dialogue state topic statement, the dialogue state topic statement can be expanded to obtain the expanded dialogue state topic statement. The expanded dialogue state topic statement includes preset key elements for analyzing the dialogue state. The preset key elements may include role, dialogue state (emotion) and context.

[0103] In practical applications, users (e.g., emotion experts) can use large model tools to assist in constructing expanded dialogue state topic statements. At this stage, each expanded dialogue state topic statement needs to include three key elements: character, emotion, and context. The character element must include at least one protagonist, with one or two supporting characters, or none at all; the context element describes the specific situation or event, requiring clear logic and realism. Finally, the emotion element requires that the outcome of the event indirectly reflects the protagonist's emotions.

[0104] For example, the dialogue status theme is "The sadness of permanently losing an important memento." An expanded dialogue status theme could be "After discovering that all the negatives he had treasured throughout his life had turned to mold, the old photographer silently put the empty album back on the bookshelf and never picked up a camera again." Here, the character element is the old photographer; the situational element is that all the negatives had turned to mold; and the emotional elements are "silently put it back" and "never picked it up again." The restrained actions replace wailing, and the abandonment of professional habits is used as a ritual of mourning, showing a sense of powerlessness.

[0105] Step S230: Optimize the expanded dialogue state topic statement to obtain the statement to be analyzed.

[0106] Among them, the richness of the dialogue state information of the statement to be analyzed is higher than that of the dialogue state information of the expanded dialogue state topic statement.

[0107] For example, the statement to be analyzed can have more emotional fluctuations and shifts compared to the expanded dialogue state statement.

[0108] In practice, computer devices can optimize the expanded dialogue state topic statements to obtain the statements to be analyzed in the emotion recognition and understanding stage, so that the richness of the dialogue state information of the statements to be analyzed is higher than the richness of the dialogue state information of the expanded dialogue state topic statements.

[0109] In practical applications, the design requirements for the statement to be analyzed can be: "Construct a misleading or pivotal scenario that may create a potential conflict in the emotional expression of the preceding and following information." For example, the statement to be analyzed could be: "After discovering that all the negatives he had treasured throughout his life had become moldy, the old photographer suddenly smiled and ordered the latest digital camera—until his family sorted through his belongings and discovered that all the memory cards in the new cameras were completely blank." Here, the surface behavior (misleading): "Smiling and ordering a new camera" conveys a signal of encouragement, suggesting that he has emerged from the shadows; the truth (pivotal): "The blank memory cards" reveals that he never actually pressed the shutter, using false positivity to cover up despair; the emotional conflict: the outward nonchalance and the inner desolation create an emotional conflict, reinforcing the end of his creative life due to the destruction of the memento.

[0110] Step S240: Obtain the dialogue state analysis result set corresponding to the statement to be analyzed, and use the statement to be analyzed and the corresponding dialogue state analysis result set as the evaluation dataset.

[0111] The dialogue state analysis result set includes correct analysis results and at least one analysis result that differs from the correct analysis results.

[0112] The evaluation dataset is used to assess the dialogue state awareness capabilities of pre-trained large language models.

[0113] In practice, the computer device can obtain the dialogue state analysis result set corresponding to the statement to be analyzed, and use the statement to be analyzed and the corresponding dialogue state analysis result set as the evaluation dataset corresponding to the emotion recognition and understanding stage. The dialogue state perception capability of the pre-trained large language model is evaluated through the evaluation dataset.

[0114] In this embodiment, the dialogue state recognition and understanding capabilities of a pre-trained large language model are specifically evaluated using an evaluation dataset. These capabilities include two aspects: dialogue state recognition capability and dialogue state understanding capability. Dialogue state recognition capability refers to the pre-trained large language model's ability to accurately identify the dialogue state category corresponding to the statement to be analyzed. Dialogue state understanding capability refers to the pre-trained large language model's ability to accurately understand the reasons for the generation of dialogue state categories. Correspondingly, dialogue state understanding capability can also be named the ability to understand the reasons for the generation of dialogue state categories. Further, correct analysis results are used to characterize the correct dialogue state category and the reasons for its generation. Analysis results different from the correct analysis results are used to characterize dialogue state categories different from the correct ones, and the reasons for their generation.

[0115] In the method for constructing the evaluation dataset of the aforementioned large language model, the following steps are taken: First, dialogue state topic sentences corresponding to different evaluation dimensions are obtained. These dialogue state topic sentences contain both dialogue state information and topic scene information. The evaluation dimensions are used to evaluate the dialogue state perception ability of the pre-trained large language model. Second, the dialogue state topic sentences are expanded to obtain expanded dialogue state topic sentences. These expanded sentences include preset key elements for analyzing the dialogue state. Third, the expanded sentences are optimized to obtain sentences to be analyzed. The richness of dialogue state information in the sentences to be analyzed is higher than that in the expanded sentences. Fourth, the dialogue state analysis result set corresponding to the sentences to be analyzed is obtained, and the sentences to be analyzed and the corresponding dialogue state analysis result set are used as the evaluation dataset. The dialogue state analysis result set includes correct analysis results and at least one analysis result that differs from the correct analysis results. The evaluation dataset is used to evaluate the dialogue state perception ability of the pre-trained large language model.

[0116] Thus, by acquiring dialogue state topic sentences corresponding to different evaluation dimensions, which contain dialogue state information and topic scene information, the dialogue state perception capability of large language models in complex and varied dialogue scenarios can be evaluated more objectively. By expanding the dialogue state topic sentences, the expanded dialogue state topic sentences include preset key elements for analyzing dialogue states. Then, by optimizing the expanded dialogue state topic sentences, the optimized sentences to be analyzed contain richer dialogue state information, thereby effectively ensuring the diversity and complexity of the sentences to be analyzed. Based on the sentences to be analyzed and the corresponding dialogue state analysis result set, an evaluation dataset is constructed to evaluate the dialogue state perception capability of the pre-trained large language model. The dialogue state analysis result set includes correct analysis results and at least one analysis result that differs from the correct analysis result, thereby increasing the testing difficulty of the pre-trained large language model, comprehensively evaluating the model's dialogue state perception capability in real dialogues, and thus more accurately evaluating the dialogue state perception capability of the large language model.

[0117] In one embodiment, obtaining the dialogue state analysis result set corresponding to the statement to be analyzed includes: obtaining a first correct sub-result and a corresponding second correct sub-result corresponding to the statement to be analyzed; the first correct sub-result represents the correct dialogue state category corresponding to the statement to be analyzed; the second correct sub-result represents the reason for the generation of the correct dialogue state category; obtaining at least one first sub-result that is different from the first correct sub-result; the first sub-result represents a dialogue state category that is different from the correct dialogue state category; obtaining at least one second sub-result that is different from the second correct sub-result; the second sub-result represents the reason for the generation of the dialogue state category that is different from the correct dialogue state category.

[0118] The correct analysis result includes the first correct sub-result and the second correct sub-result. Each analysis result that differs from the correct analysis result also includes the first sub-result and the second sub-result.

[0119] The dialogue state categories can include at least one of happiness, trust, fear, surprise, sadness, disgust, anger, and anticipation, as well as combinations of them and their derived emotions.

[0120] In specific implementation, the correct analysis result includes a first correct sub-result and a second correct sub-result. Each analysis result that differs from the correct analysis result also includes a first sub-result and a second sub-result. During the process of obtaining the dialogue state analysis result set corresponding to the statement to be analyzed, the computer device can obtain the first correct sub-result and the corresponding second correct sub-result corresponding to the statement to be analyzed. Among them, the first correct sub-result represents the correct dialogue state category corresponding to the statement to be analyzed; the second correct sub-result represents the reason for the generation of the correct dialogue state category; and at least one first sub-result that differs from the first correct sub-result is obtained. The first sub-result represents the dialogue state category that differs from the correct dialogue state category; and at least one second sub-result that differs from the second correct sub-result is obtained. The second sub-result represents the reason for the generation of the dialogue state category that differs from the correct dialogue state category.

[0121] For example, the statement to be analyzed is: "Going to meet friends after get off work." The first correct sub-result is: anticipation, and the second correct sub-result is: meeting friends. Each analysis result that differs from the correct analysis result includes a first sub-result and a second sub-result. For instance, there are three analysis results that differ from the correct analysis result: Analysis Result 1, Analysis Result 2, and Analysis Result 3. In Analysis Result 1, the first sub-result is: anticipation, and the second sub-result is: after get off work; in Analysis Result 2, the first sub-result is: happy, and the second sub-result is: going to meet friends after get off work; in Analysis Result 3, the first sub-result is: surprised, and the second sub-result is: meeting friends.

[0122] In this embodiment, the correct analysis result includes a first correct sub-result and a second correct sub-result. Each analysis result that differs from the correct analysis result also includes a first sub-result and a second sub-result. The process involves obtaining the first correct sub-result and the corresponding second correct sub-result corresponding to the statement to be analyzed. The first correct sub-result represents the correct dialogue state category corresponding to the statement to be analyzed. The second correct sub-result represents the reason for the generation of the correct dialogue state category. At least one first sub-result that differs from the first correct sub-result is obtained. The first sub-result represents a dialogue state category different from the correct dialogue state category. At least one second sub-result that differs from the second correct sub-result is obtained. The second sub-result represents the reason for the generation of the dialogue state category different from the correct dialogue state category.

[0123] Thus, the dialogue state analysis result set corresponding to the statement to be analyzed includes the dialogue state category corresponding to the statement to be analyzed and the reason for the generation of the dialogue state category. By evaluating the pre-trained large language model through the statement to be analyzed and the corresponding dialogue state analysis result set, we can not only evaluate the dialogue state category recognition ability of the pre-trained large language model, but also evaluate the understanding ability of the reason for the generation of the dialogue state category of the pre-trained large language model. This can reflect the deep dialogue state perception intelligence of the model and provide a benchmark for improving the dialogue state perception intelligence of the dialogue system.

[0124] In one embodiment, the method further includes: obtaining a first analysis result set corresponding to the statement to be analyzed; the first analysis result set includes a first category label and at least one second category label different from the first category label; the first category label represents the correct dialogue state category corresponding to the statement to be analyzed; the second category label represents the dialogue state category different from the correct dialogue state category; the first category label is used as a first correct sub-result, and the second category label is used as a first sub-result.

[0125] In a specific implementation, the computer device can obtain a first analysis result set corresponding to the statement to be analyzed; the first analysis result set includes a first category label and at least one second category label that is different from the first category label; wherein, the first category label represents the correct dialogue state category corresponding to the statement to be analyzed; the second category label represents the dialogue state category that is different from the correct dialogue state category; the first category label is used as the first correct sub-result, and the second category label is used as the first sub-result.

[0126] In practical applications, the first-category labels can be designed manually, such as by sentiment experts and data labelers who can design the most suitable first-category labels for each statement to be analyzed. Specifically, four category labels can be designed for each statement to be analyzed: one first-category label representing the correct dialogue state category, and the other three second-category labels representing dialogue state categories different from the correct dialogue state category. These second-category labels can be generated by a large language model and then modified by data labelers to ensure logical coherence and a certain degree of obfuscation. The number of second-category labels can be set according to actual needs and is not specifically limited here.

[0127] The technical solution of this embodiment obtains a first analysis result set corresponding to the statement to be analyzed. The first analysis result set includes a first category label and at least one second category label different from the first category label. The first category label represents the correct dialogue state category corresponding to the statement to be analyzed. The second category label represents a dialogue state category different from the correct dialogue state category. The first category label is used as the first correct sub-result, and the second category label is used as the first sub-result. In this way, the correct dialogue state category is determined for the statement to be analyzed, and a dialogue state category different from the correct dialogue state category is designed. The dialogue state recognition capability of the pre-trained large language model is evaluated to determine whether the pre-trained large language model can select the correct state category from multiple dialogue state category options, thereby increasing the difficulty of the dialogue state recognition capability test.

[0128] In one embodiment, the method further includes: obtaining a second analysis result set corresponding to the statement to be analyzed; the second analysis result set includes first reason information and at least one second reason information different from the first reason information; the first reason information represents the reason for the generation of the correct dialogue state category; the second reason information represents the reason for the generation of the dialogue state category different from the correct dialogue state category; the first reason information is used as the second correct sub-result, and the second reason information is used as the second sub-result.

[0129] In a specific implementation, the computer device can also obtain a second analysis result set corresponding to the statement to be analyzed; the second analysis result set includes first reason information and at least one second reason information that is different from the first reason information; the first reason information represents the reason for the generation of the correct dialogue state category; the second reason information represents the reason for the generation of the dialogue state category that is different from the correct dialogue state category; the first reason information is used as the second correct sub-result, and the second reason information is used as the second sub-result.

[0130] In practical applications, when determining the correct dialogue state category corresponding to the statement to be analyzed, the reason for the correct dialogue state category can be identified simultaneously, serving as a second correct sub-result. Data annotators design three second sub-results based on the large language model, representing the reasons for the generation of dialogue state categories that differ from the correct dialogue state category, to obfuscate the second correct sub-result. These second sub-results not only need to be logically sound and clearly expressed, but also must have some connection to the second correct sub-result, although they are not actually the reason for the generation of the correct dialogue state category. The number of second sub-results is determined based on actual needs and is not specifically limited here.

[0131] The technical solution of this embodiment obtains a second analysis result set corresponding to the statement to be analyzed. The second analysis result set includes first reason information and at least one second reason information different from the first reason information. The first reason information represents the reason for the generation of the correct dialogue state category. The second reason information represents the reason for the generation of the dialogue state category different from the correct dialogue state category. The first reason information is used as the second correct sub-result, and the second reason information is used as the second sub-result. In this way, by determining the reason for the generation of the correct dialogue state category for the statement to be analyzed and designing the reason for the generation of the dialogue state category different from the correct dialogue state category, it can be used to evaluate the understanding ability of the pre-trained large language model of the reason for the generation of dialogue state categories, and to determine whether the pre-trained large language model can select the correct dialogue state category reason option from multiple dialogue state category reason options, thereby increasing the difficulty of the test of the understanding ability of the reason for the generation of dialogue state categories.

[0132] In some embodiments, obtaining dialogue state topic statements corresponding to different evaluation dimensions includes: obtaining a dialogue statement set from a preset multi-turn dialogue dataset; the dialogue statements in the dialogue statement set contain dialogue state information; filtering the dialogue statement set to obtain a target dialogue statement set; the target dialogue statements in the target dialogue statement set satisfy preset dialogue quality conditions; obtaining the expanded target dialogue statements, and using the expanded target dialogue statements as dialogue state topic statements.

[0133] The expanded target dialogue statement is obtained by expanding the target dialogue statement with thematic scenarios according to the definition of the evaluation dimensions.

[0134] In practice, when acquiring dialogue state topic statements corresponding to different evaluation dimensions, the computer device can extract dialogue data containing dialogue state information from a pre-set multi-turn dialogue dataset as dialogue statements to form a dialogue statement set. Then, the dialogue statement set is filtered according to pre-set dialogue quality conditions, and dialogue statements that meet the pre-set dialogue quality conditions are selected as target dialogue statements. Multiple target dialogue statements form a target dialogue statement set. The pre-set dialogue quality conditions are used to filter out dialogue statements from the dialogue statement set that do not contain sensitive information and / or are logically sound.

[0135] In this way, dialogue state topic statements corresponding to different evaluation dimensions can be obtained based on the target dialogue statement set. Specifically, the target dialogue statements can be expanded into topic scenarios according to the definitions of each dialogue state information subclassification under each evaluation dimension corresponding to the emotion recognition and understanding stage, thereby obtaining dialogue state topic statements corresponding to different evaluation dimensions.

[0136] In practical applications, after extracting a large amount of dialogue data containing dialogue state information from a pre-set multi-turn dialogue dataset as basic material, computer equipment can use semi-automatic tools to filter dialogue quality conditions, eliminating logically illogical or sensitive dialogue statements to obtain the target dialogue statements. Alternatively, it can use rule-based regular expressions to filter logically illogical or sensitive dialogue statements to obtain the target dialogue statements. Subsequently, data labelers can expand the target dialogue statements with thematic scenarios based on the definitions of each dialogue state information subclassification in the emotion recognition and understanding process, designing dialogue state theme statements related to the dialogue state information subclassification. These dialogue state theme statements retain the dialogue state information in the original dialogue statements, providing a foundation for subsequent expansion.

[0137] The technical solution of this embodiment involves obtaining a set of dialogue statements from a pre-set multi-turn dialogue dataset; the dialogue statements in the set contain dialogue state information and topic / scene information; filtering the set of dialogue statements to obtain a target set of dialogue statements; the target dialogue statements in the target set of dialogue statements satisfy pre-set dialogue quality conditions; obtaining expanded target dialogue statements, and using the expanded target dialogue statements as dialogue state / topic statements, wherein the expanded target dialogue statements are obtained by expanding the target dialogue statements according to the definition of the evaluation dimension. Thus, after filtering out a set of dialogue statements containing dialogue state information and topic / scene information from the multi-turn dialogue dataset, and then filtering out a set of target dialogue statements that meet the pre-set dialogue quality conditions, the target dialogue statements are expanded according to the definition of the evaluation dimension, effectively improving the quality of the obtained dialogue state / topic statements. This makes the dialogue state / topic statements more suitable for evaluating dialogue state awareness capabilities, and thus effectively improving the evaluation credibility during the process of evaluating the dialogue state awareness capabilities of a pre-trained large language model based on the dialogue state / topic statements.

[0138] In one embodiment, optimizing the expanded dialogue state topic statement to obtain the statement to be analyzed includes: inputting the expanded dialogue state topic statement into a pre-trained statement optimization model and outputting the statement to be analyzed; the pre-trained statement optimization model adopts the following training method: obtaining a training sample set; the training sample set includes the original statement and the corresponding label-optimized statement; the richness of dialogue state information in the label-optimized statement is higher than that in the original statement; inputting the original statement into the statement optimization model to be trained and outputting the predicted optimized statement corresponding to the original statement; adjusting the model parameters of the statement optimization model to be trained based on the difference between the predicted optimized statement corresponding to the original statement and the corresponding label-optimized statement and continuing iterative training until the trained statement optimization model meets the training termination condition and training stops, thus obtaining the pre-trained statement optimization model.

[0139] The pre-trained sentence optimization model can be the large language model that needs to be pre-trained for dialogue state awareness assessment, or other pre-trained large language models, without specific limitations here.

[0140] In practice, when the computer device optimizes the expanded dialogue state topic statement to obtain the statement to be analyzed, it can input the expanded dialogue state topic statement into a pre-trained statement optimization model. The pre-trained statement optimization model optimizes the expanded dialogue state topic statement and outputs the statement to be analyzed.

[0141] The pre-trained statement optimization model uses the following training method: the computer device can first obtain a training sample set; the training sample set includes the original statement and the corresponding tag-optimized statement; the richness of the dialogue state information of the tag-optimized statement is higher than that of the original statement.

[0142] The tagged and optimized sentences can be obtained by emotion experts optimizing the original sentences. The design requirements for the tagged and optimized sentences can be: "Construct a misleading or transitional situation so that the preceding and following information may have a potential conflict in terms of emotional expression."

[0143] Thus, the original sentence is input into the sentence optimization model to be trained, and the predicted optimized sentence corresponding to the original sentence is output. Then, the model parameters of the sentence optimization model to be trained can be adjusted according to the difference between the predicted optimized sentence corresponding to the original sentence and the corresponding label optimized sentence, and iterative training can continue until the trained sentence optimization model meets the training termination condition, and training stops, thus obtaining the pre-trained sentence optimization model.

[0144] The technical solution of this embodiment involves inputting the expanded dialogue state topic statement into a pre-trained statement optimization model and outputting the statement to be analyzed. The pre-trained statement optimization model employs the following training method: obtaining a training sample set; the training sample set includes the original statement and the corresponding tag-optimized statement; the richness of dialogue state information in the tag-optimized statement is higher than that in the original statement; inputting the original statement into the statement optimization model to be trained and outputting the predicted optimized statement corresponding to the original statement; adjusting the model parameters of the statement optimization model to be trained based on the difference between the predicted optimized statement corresponding to the original statement and the corresponding tag-optimized statement and continuing iterative training until the trained statement optimization model meets the training termination condition, at which point training stops, thus obtaining the pre-trained statement optimization model.

[0145] Thus, by inputting the expanded dialogue state topic statement into a pre-trained statement optimization model, the output statement to be analyzed is obtained. The pre-trained statement optimization model is trained as follows: a training sample set is obtained; the training sample set includes the original statement and the corresponding label-optimized statement; the richness of dialogue state information in the label-optimized statement is higher than that in the original statement; the original statement is input into the statement optimization model to be trained, and the output statement is the predicted optimized statement corresponding to the original statement; based on the difference between the predicted optimized statement corresponding to the original statement and the corresponding label-optimized statement, the model parameters of the statement optimization model to be trained are adjusted and iterative training continues until the trained statement optimization model meets the training termination condition, at which point training stops, and the pre-trained statement optimization model is obtained.

[0146] Thus, the sentence optimization model to be trained is iteratively trained using the original sentence and the corresponding label-optimized sentence. The richness of dialogue state information in the label-optimized sentence is higher than that in the original sentence, which allows the pre-trained sentence optimization model to optimize the dialogue state information of the input dialogue sentence, effectively improving the richness of the dialogue state information of the input dialogue sentence. By testing the dialogue state awareness ability of the pre-trained large language model with such dialogue sentences that have richer dialogue state information, the difficulty of the test can be effectively increased.

[0147] In some other embodiments, such as Figure 3 As shown, this is a flowchart illustrating a method for constructing an evaluation dataset for a large language model. Figure 3 As shown, the process is divided into two parts: constructing the statements to be analyzed (question construction) and designing the dialogue state analysis result set (answer design). The specific process is as follows: Figure 3 As shown:

[0148] For the raw data (dialogue data) in the multi-turn dialogue dataset, dialogue data containing dialogue state information and topic scene information are selected to form a dialogue statement set. The dialogue statement set is then filtered according to preset dialogue quality conditions, and dialogue statements that meet the preset dialogue quality conditions are selected as initial data (target dialogue statements). Based on the definition of each dialogue state information subclass in the emotion recognition and understanding process, data annotators design specific emotion statements (dialogue state topic statements) related to the subclass based on the initial data. The emotion statements are then expanded and optimized to obtain emotion problem statements (statements to be analyzed). Finally, a corresponding dialogue state analysis result set is designed for the emotion problem statements.

[0149] In this way, a complete dataset for each emotion recognition and understanding test question is ultimately constructed. This process ensures the diversity and complexity of the questions, so that they not only test the model's emotion recognition ability but also assess its ability to understand the reasons for emotion. Each question and answer undergoes multiple optimizations and reviews, aiming to provide a scientific, rigorous, and challenging benchmark set that can realistically reflect the model's performance in various emotional scenarios. This multi-dimensional and multi-layered method of constructing emotion intelligence evaluation datasets, through a refined emotion classification system and a rigorous data construction process, comprehensively evaluates the recognition ability of large language models in complex emotional scenarios and their ability to understand the reasons for emotion, providing a scientific benchmark for improving the emotional intelligence of dialogue systems.

[0150] In another embodiment, such as Figure 4 The diagram illustrates a method for constructing an evaluation dataset for a large language model, including the following steps:

[0151] Step S402: Obtain the training sample set, input the original statement into the statement optimization model to be trained, and output the predicted optimized statement corresponding to the original statement.

[0152] Step S404: Based on the difference between the predicted optimized statement and the corresponding label optimized statement corresponding to the original statement, adjust the model parameters of the statement optimization model to be trained and continue iterative training until the trained statement optimization model meets the training termination condition, and then stop training to obtain the pre-trained statement optimization model.

[0153] Step S406: Obtain the dialogue statement set from the preset multi-turn dialogue dataset.

[0154] Step S408: Filter the dialogue statement set to obtain the target dialogue statement set.

[0155] Step S410: Obtain the expanded target dialogue statement and use the expanded target dialogue statement as the dialogue state topic statement.

[0156] Step S412: Expand the dialog state topic statement to obtain the expanded dialog state topic statement.

[0157] Step S414: Input the expanded dialogue state topic statement into the pre-trained statement optimization model and output the statement to be analyzed.

[0158] Step S416: Obtain the first correct sub-result and the corresponding second correct sub-result for the statement to be analyzed.

[0159] Step S418: Obtain at least one first sub-result that is different from the first correct sub-result, and obtain at least one second sub-result that is different from the second correct sub-result.

[0160] Step S420: Use the statement to be analyzed and the corresponding dialogue state analysis result set as the evaluation dataset.

[0161] It should be noted that the specific limitations of the above steps can be found in the specific limitations of the construction method of an evaluation dataset for a large language model described above.

[0162] 2. Emotional Management

[0163] like Figure 5 The diagram illustrates a flowchart of a dialogue state response mechanism. Dialogue emotion management refers to the process in human-computer dialogue scenarios where, when the model identifies a user's emotional state (including emotion category and reason for emotion), the robot makes a corresponding judgment based on the current topic and specific emotional state, thereby determining the appropriate response mode to advance the next step of the dialogue. Emotion management closely integrates two main aspects: "emotion recognition" and "emotion response." Its core lies in the model's ability to judge and predict suitable response strategies based on the user's current emotional state, ensuring that the user receives emotional support and enhancing the interactive experience.

[0164] Emotion management tasks provide emotional dialogue robots with a reference for automatic emotion expression, helping the robot make reasonable and evidence-based responses in emotional decision-making, thereby improving the overall user experience and user stickiness. In dialogue, common emotion management patterns can be divided into the following five types:

[0165] (1) Soothing: Soothing refers to the process by which a robot takes a series of measures to stabilize a user's emotions when the user experiences emotional fluctuations. As an emotion management strategy, soothing is characterized by reducing the intensity of emotions and restoring normal emotional levels. At the same time, through good communication, the robot can understand the reasons for the user's emotional fluctuations and provide targeted emotional guidance.

[0166] (2) Companionship: Companionship means that the robot spends a period of time with the user as a "companion" and provides emotional support. In emotional management, the role of companionship includes providing users with emotional support, enhancing their sense of security and confidence in facing difficulties, and releasing emotions through emotional communication.

[0167] (3) Small talk: Small talk refers to informal conversations conducted in a relaxed and pleasant manner. In emotional management, small talk helps users release stress and regulate their emotions. Through small talk, users can find common topics, increase their interest in the conversation, and enhance their enthusiasm for interaction.

[0168] (4) Discussion: Discussion refers to the exchange of views between the robot and the user on a certain issue during a dialogue, helping the user to broaden their thinking and find suitable solutions. Discussion not only allows users to fully express their own views and reduce emotional stress, but also helps them seek solutions to problems and improve their problem-solving abilities.

[0169] (5) Psychological Status Monitoring: Psychological status monitoring refers to the robot's ability to continuously assess the emotions and psychological state of users with extremely unstable emotions, promptly identifying potential psychological problems and providing a basis for subsequent intervention. Its functions include assessing psychological status, developing intervention plans, and conducting targeted psychological interventions based on monitoring results.

[0170] In the era of large language models, emotion management capability has become a crucial foundation for enhancing the model's emotional dialogue ability. Enabling large language models to track and manage emotions helps them better consider coping strategies and emotional response patterns. Therefore, emotion management capability is an indispensable ability for large language models in emotional dialogue. To evaluate the coping and judgment capabilities of different large language models in emotional dialogue, an emotion management task evaluation benchmark was designed. This benchmark includes the following evaluation dimensions:

[0171] (1) Personal Scenarios: including Emotional Support, which involves personal emotional expression and psychological needs; Self-Development, which focuses on personal growth and self-improvement; and Health and Wellness, which involves events related to physical and mental health.

[0172] (2) Social Scenarios: including Interpersonal Communication, which covers social activities and interpersonal relationship issues; Academic and Professional, which are events related to learning and work; and Sociocultural Dynamics, which involve social events, cultural trends, or public issues.

[0173] This evaluation benchmark simulates the thinking and judgment abilities of a large language model in responding to information received from a conversation partner at both personal and social levels, thereby assessing its emotional management level. The personal scenario test mainly evaluates the emotional support strategies that the model can propose when the conversation partner has doubts or problems; while the social scenario focuses on testing the optimal response that the model can make after receiving socially relevant chat content from the user.

[0174] In practical applications, the evaluation dataset in the emotion management phase can include the statements to be analyzed and the corresponding dialogue state analysis results sets corresponding to multiple dialogue topics such as emotional support and self-development in the personal scenario dimension, as well as the statements to be analyzed and the corresponding dialogue state analysis results sets corresponding to multiple dialogue topics such as interpersonal communication and academic / career in the social scenario dimension. For example, in the emotion management phase, the evaluation dataset corresponding to the personal scenario dimension and the evaluation dataset corresponding to the social scenario dimension can each account for 50%. In practical applications, the proportion of evaluation datasets under different evaluation dimensions in the same evaluation phase can be set according to actual needs and is not specifically limited here.

[0175] To accurately evaluate the performance of emotion management, similar to the benchmarks for emotion recognition and understanding, this solution designs an evaluation scheme for emotion management. The evaluation constructs emotion questions (statements to be analyzed) and designs answer options (a set of dialogue state analysis results), ultimately forming a standard benchmark for emotion management. The specific process is similar to the construction method in the previous section, but in the answer design, emotion experts combine or derive from five emotion management models (soothing, companionship, small talk, discussion, and psychological monitoring) as the final correct analysis result. In addition, other analysis results are semi-automatically generated by data labelers, also based on these five emotion management models, with supplementary designs involving combinations or derivative behaviors. In practical application, each statement to be analyzed has four analysis results, including one correct analysis result and three different analysis results, for the large language model of the evaluation to choose from.

[0176] For example, in the emotion management stage, the process of evaluating a pre-trained large language model using a specific statement to be analyzed within the social scenario dimension and the corresponding dialogue state analysis results can be shown in Table 1 below:

[0177] Table 1

[0178]

[0179] As shown in Table 1, a statement to be analyzed can include dialogue scenario information and dialogue question information. The dialogue scenario information is: "A music creator begins to realize that audiences want to hear music with emotion, not so-called perfection." The dialogue question information is: "In this situation, what is the most effective action for the music creator?" The dialogue state analysis result set corresponding to this statement includes: "A: [Encourage them to stick to their own style], B: [Discuss the expression of music], C: [Analyze audience preferences], D: [Provide professional music production advice]", where "B: [Discuss the expression of music]" is the correct analysis result. The statement to be analyzed and the corresponding dialogue state analysis result set are input into a pre-trained large language model, which outputs the target prediction result. For example, if the pre-trained large language model outputs the target prediction result as: "C: [Analyze audience preferences]", it is determined that the pre-trained large language model did not select the correct analysis result, and the computer device can further analyze the reason for the pre-trained large language model's incorrect selection.

[0180] For another example, consider a statement to be analyzed within the personal scenario dimension, specifically the topic of "Health and Wellness" and its corresponding dialogue state analysis results. The statement to be analyzed might be, "Robert recently got a satisfying job, but he also feels a lot of work pressure. He often works overtime and his diet is irregular." The corresponding dialogue state analysis result set could include: "Sharing work tips," "Providing health advice," "Listening to his stress," and "Planning rest time together." The correct analysis result is "Providing health advice."

[0181] 3. General emotional responses

[0182] In robotics and artificial intelligence systems, the ability to express emotions is a crucial indicator of advanced emotional functions. Possessing emotionally rich and appropriate responses not only makes dialogue systems more human-like but also reflects the long-term goal of AI technology in promoting machine emotional intelligence. The core of the emotional response task lies in generating dialogue or non-dialogue responses that incorporate emotional elements, enabling robots to exhibit human-like emotional expression when interacting with humans. In the emotional response task, the robot first perceives and analyzes the user's emotional state, and then, based on the current dialogue context, provides the user with optimal emotional feedback.

[0183] Emotional response is a typical generative task. Specifically, it involves recognizing and understanding the emotions in user utterances, combining this with emotion management strategies to select the most appropriate response state, and ultimately generate a targeted and meaningful response. Emotional response generation methods in related technologies often rely on emotion tags to produce fixed-pattern responses. For example, when a user inputs "I lost my keys on the street today," the system detects the emotion tag as "sad" and generates a comforting response such as "Don't be sad, things will be alright." The limitation of this approach is that the generated emotional response types are relatively limited and cannot effectively address diverse emotional needs. Another improvement is to combine human templates with deep learning models to optimize response diversity and avoid stereotypical responses. However, traditional emotion generation models often neglect to deeply integrate the user's emotional information into the response, resulting in responses lacking emotional depth and appearing rather bland.

[0184] Large language models demonstrate significant advantages in emotional response. Due to their powerful dialogue understanding capabilities, large language models can fully capture the information contained in the input text, generating more diverse and user-relevant emotional responses. Therefore, to more accurately evaluate the performance of different large language models in handling emotional subtasks, this application designs a benchmark system specifically for evaluating emotional responses. This benchmark assesses the emotional response capabilities of large language models from two evaluation dimensions:

[0185] (1) Response Scenarios: These include personal and social aspects. At the personal level, emotional responses involve addressing emotional issues, life difficulties, personal development, and other related topics (specific aspects may include but are not limited to the examples mentioned above). At the social level, emotional responses focus on emotional needs related to interpersonal communication, career and academic challenges, and socio-cultural dynamics.

[0186] (2) Action Scenarios: These also include personal and social aspects. Personal action scenarios cover the emotional needs that users may face in areas such as emotional problems, life difficulties, and developmental obstacles. Social action scenarios include emotional responses to topics such as interpersonal relationships, workplace issues, and socio-cultural dynamics.

[0187] This benchmark system comprehensively covers users' potential emotional needs by constructing various real-life problem scenarios. During the design process, a third-person perspective from a large language model was introduced, enabling it to provide effective solutions to specific emotional issues based on the emotional and contextual information provided by the user. These solutions can be emotionally resonant verbal responses or specific action suggestions.

[0188] In the baseline design of the emotional response section, the data generation and answer construction process continues the approach from the previous section. The difference lies in that, in this section, emotional experts design scenarios for the core characters in each emotional theme, ensuring that these scenarios reflect the specific emotional needs of the characters. Based on these needs, experts generate the most appropriate emotional responses as the correct analysis results. These responses can be in verbal form or action suggestions. The design requirements for correct emotional response analysis results include: possessing characteristics such as humor, attractiveness, interactivity, gentleness, and politeness. When constructing analysis results that differ from the correct analysis results, data annotators use semi-automated techniques to design three logically sound but not entirely identical emotional needs based on the correct analysis results, ensuring that the dialogue state analysis result set corresponding to each analyzed statement has sufficient discriminative power.

[0189] 4. Emotional responses in specific scenarios

[0190] To better apply the sentiment evaluation of large language models to specific business scenarios, this application designs a set of sentiment response evaluation benchmarks specifically for large language models in handling user feedback in specific scenarios. The construction of this benchmark aims to evaluate the model's performance in handling emotional issues in special scenarios, and it is divided into two evaluation dimensions: emotional counseling and psychological guidance.

[0191] The emotional counseling section covers four main categories of issues: romantic relationships, friendship, family relationships, and other emotional issues. Within romantic relationships, it is further subdivided into five specific scenarios: early stages of a relationship, middle stages of a relationship, late stages of a relationship, relationship crises, and breakups and reconciliations. Friendship counseling focuses on three typical scenarios: making new friends, maintaining friendships, and friendship crises. The family counseling section includes counseling on parent-child relationships, sibling relationships, and other kinship relationships. Other emotional scenarios involve broader social relationships, such as virtual relationships, cross-cultural relationships, professional relationships, and emotional counseling scenarios related to hobbies and interests.

[0192] The emotional and psychological counseling section focuses on three common sources of stress: economic stress, life stress, and work or study stress. Economic stress includes situations such as insufficient income, financial planning, mortgage and car loan payments, education investment, and retirement pressure. Life stress encompasses family problems, relationships with friends, social anxiety, health issues, and other life challenges. Work or study stress involves workplace competition, entrepreneurial pressure, academic pressure, career planning, and work-life balance issues.

[0193] In these sub-scenarios, the questions for emotional counseling and psychological support are designed to assess the performance of the large language model in complex emotional situations. Emotional counseling covers most everyday emotional issues, while psychological support focuses on the most common sources of stress in real life. These scenarios are designed to be more challenging and closely resemble real-life emotional dialogue needs, making the evaluation results more practically meaningful.

[0194] The data construction process is similar to the previous stages, but each sub-category scenario in this section is collaboratively designed by multiple psychological experts to ensure that the scenario descriptions are more realistic, logical, and detailed. These scenarios include the background and detailed information of the dialogue, presenting a smooth dialogue logic. Psychological experts identify the protagonists of the dialogue and design reasonable and specific dialogue plots. The protagonists can ask questions or engage in casual conversation, with an overall conversational style that meets the needs of real-life dialogue scenarios. When designing the analysis results, emotional experts and data annotators use semi-automated technology to generate four different analysis results. All analysis results must be reasonable, emotionally intelligent, clear, and conversational, easy to communicate, and able to encourage continued interaction from the participants. Although these analysis results all have their own rationality, only one analysis result is unanimously agreed upon by experts and annotators as the most suitable and correct one that maximizes emotional satisfaction. The entire dialogue scenario, plot, and options must remain natural, coherent, and consistent with realistic logic.

[0195] There are some subtle differences in the design of the emotional counseling and emotional psychological guidance sections. For emotional counseling, each subcategory includes a set of three multiple-choice questions, designed for different age groups: juvenile (15-22 years old), youth (23-30 years old), and middle-aged (30 years and older). This design is closer to real-world needs. The emotional psychological guidance section, on the other hand, designs two multiple-choice questions for each emotional scenario. Each question is customized for a specific age group and identity, ensuring logical consistency and appropriate challenge. Through these designs, the benchmark is more targeted and challenging, thus better assessing the emotional responsiveness of the large language model.

[0196] Thus, this application proposes a complete evaluation framework covering emotion recognition, understanding, management, and response, encompassing the entire process of emotion perception ability evaluation to ensure a systematic assessment of model capabilities. Furthermore, a semi-automated benchmark construction method, collaboratively developed by multiple psychology or sociology experts, ensures the diversity and logical rigor of the evaluation data. Simultaneously, it includes multi-dimensional scenarios such as personal, social, and special events, closely reflecting real-world dialogue needs and enhancing the practical application value of the evaluation results. The evaluation system is constructed from four dimensions: emotion source, expression mode, context, and dynamic changes, supporting testing in complex emotional scenarios. It can accurately assess whether the model possesses deeper reasoning abilities and the capacity to flexibly respond to diverse emotional scenarios. In addition, standardized labels are constructed based on emotion theory, while also supporting the dynamic expansion of new scenarios and emotion types to adapt to technological development needs.

[0197] In some embodiments, the evaluation dataset corresponding to the evaluation stage is input into a pre-trained large language model, and the target prediction result corresponding to the statement to be analyzed in the evaluation dataset is output. This includes: for any statement to be analyzed in the evaluation dataset, the statement to be analyzed and the corresponding dialogue state analysis result set are repeatedly input into the pre-trained large language model for a first preset number of times, and the prediction result corresponding to any statement to be analyzed is output for each input; the prediction result that appears most frequently is taken as the target prediction result corresponding to any statement to be analyzed.

[0198] The first preset number of times refers to the preset number of data inputs required to determine the target prediction result corresponding to any statement to be analyzed.

[0199] In specific implementation, when the computer device inputs the evaluation dataset corresponding to the evaluation stage into the pre-trained large language model and outputs the target prediction result corresponding to the statement to be analyzed in the evaluation dataset, it can repeatedly input any statement to be analyzed and the corresponding dialogue state analysis result set into the pre-trained large language model according to a first preset number of times, and output the prediction result corresponding to any statement to be analyzed under each input, so that the prediction result that appears most frequently is taken as the target prediction result corresponding to any statement to be analyzed.

[0200] It is understandable that, during the process of repeatedly inputting any statement to be analyzed and its corresponding dialogue state analysis result set into the pre-trained large language model according to the first preset number of times, the order of the results in the dialogue state analysis result set corresponding to any statement to be analyzed should be the same, and the order of the results is the same as the order of the analysis results. For example, the dialogue state analysis result set corresponding to any statement to be analyzed includes four analysis results: "encourage them to stick to themselves", "discuss the way music is expressed", "analyze the audience's preferences", and "provide suggestions for professional music production". During the process of inputting the dialogue state analysis result set corresponding to any statement to be analyzed into the pre-trained large language model according to the first preset number of times, the order of the results should be the same each time. For example, the order of the results could be: "A: [encourage them to stick to themselves], B: [discuss the way music is expressed], C: [analyze the audience's preferences], D: [provide suggestions for professional music production]".

[0201] Continuing with the previous example, assuming the first preset number of attempts is 5, if the prediction results of the 5 outputs include 3 outputs of "C: [Analyze the audience's preferences]", 1 output of "D: [Provide professional music production advice]", and 1 output of "A: [Encourage them to stick to their own beliefs]", since "C: [Analyze the audience's preferences]" is output the most times, "C: [Analyze the audience's preferences]" is used as the target prediction result output by the pre-trained large oracle model for any statement to be analyzed.

[0202] The technical solution of this embodiment involves repeatedly inputting any statement to be analyzed and its corresponding dialogue state analysis result set into a pre-trained large language model for any given statement in the evaluation dataset a first preset number of times, and outputting the prediction result corresponding to any given statement under each input; the prediction result that appears most frequently is taken as the target prediction result for any given statement. In this way, by repeatedly inputting any given statement and its corresponding dialogue state analysis result set into the pre-trained large language model, and taking the prediction result that appears most frequently as the target prediction result, the reliability and statistical stability of the results are ensured.

[0203] In one embodiment, the number of times the evaluation dataset corresponding to each evaluation stage is input to the pre-trained large language model satisfies a second preset number, and the order of the results of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset is different in each input; there are multiple target prediction results corresponding to the statement to be analyzed in the evaluation dataset, which respectively correspond to the analysis results selected by the pre-trained large language model in the dialogue state analysis result set of each input; the correct analysis results and the corresponding target prediction results corresponding to the statement to be analyzed in the evaluation dataset are compared to obtain the evaluation information of the pre-trained large language model under each evaluation stage, including: comparing the target prediction results output for each statement to be analyzed in the evaluation dataset with the corresponding correct analysis results to obtain the comparison results corresponding to each input of the evaluation dataset; and determining the evaluation information of the pre-trained large language model under the evaluation stage based on the comparison results corresponding to each input of the evaluation dataset.

[0204] In specific implementation, when the computer device compares the correct analysis results and the corresponding target prediction results of the statements to be analyzed in the evaluation dataset to obtain the evaluation information of the pre-trained large language model in each evaluation stage, since there are multiple target prediction results corresponding to the statements to be analyzed in the evaluation dataset, which correspond to the analysis results selected by the pre-trained large language model in the dialogue state analysis result set with different order of results for each input, the computer device can compare the target prediction results output by the statements to be analyzed in the evaluation dataset with the corresponding correct analysis results to obtain the comparison results of the evaluation dataset under each input. Based on the comparison results of the evaluation dataset under each input, the evaluation information of the pre-trained large language model under the evaluation stage is determined.

[0205] Furthermore, the second preset number of times refers to the preset number of statistical counts required to determine the evaluation information of the pre-trained large language model at each evaluation stage; wherein, the evaluation information can be represented by the average accuracy of the pre-trained large language model on the evaluation dataset corresponding to the evaluation stage, and the number of statistical counts refers to the number of accuracy statistics required to statistically measure the average accuracy of the pre-trained large language model on the evaluation dataset corresponding to the evaluation stage. The accuracy is calculated through comparison results.

[0206] For each evaluation stage, the evaluation dataset contains the statements to be analyzed and the corresponding dialogue state analysis results. These are then input into a pre-trained large language model to determine the target prediction result output by the pre-trained large language model for the statements to be analyzed. The method for determining the target prediction result for any statement to be analyzed has been described in the above embodiments and is not specifically limited here. After determining the target prediction result for each statement to be analyzed in the evaluation dataset for any evaluation stage, the target prediction result for each statement to be analyzed is compared with its corresponding correct analysis result. Based on the comparison results, the accuracy of the pre-trained large language model under the current result arrangement order for that evaluation stage is calculated. Then, the result arrangement order of the dialogue state analysis results in the evaluation dataset for that evaluation stage is reordered. The statements to be analyzed and the reordered dialogue state analysis results in the evaluation dataset for that evaluation stage are then input into the pre-trained large language model, and the accuracy of the pre-trained large language model under the current result arrangement order for that evaluation stage is calculated again. Thus, by repeating the second preset number of times, the average accuracy is calculated based on the accuracy obtained each time under different result arrangements, thereby obtaining the evaluation information of the pre-trained large language model under any evaluation stage.

[0207] For example, the accuracy of the pre-trained large language model under the current result arrangement order in any evaluation stage is statistically analyzed. Then, the result arrangement order of any evaluation stage is reordered and re-input into the pre-trained large language model to recalculate the accuracy. This process is repeated four times (the second preset number of times is four). The average accuracy obtained based on the accuracy obtained from the four statistical analyses is calculated as the evaluation information of the pre-trained large language model under any evaluation stage.

[0208] In this embodiment, the evaluation dataset for each evaluation stage is input to the pre-trained large language model multiple times, satisfying a second preset number of times. The order of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset is different in each input. There are multiple target prediction results for the statement to be analyzed in the evaluation dataset, each corresponding to the analysis result selected by the pre-trained large language model in the dialogue state analysis result set for each input. By comparing the target prediction result output for each statement to be analyzed in the evaluation dataset with the corresponding correct analysis result, the comparison result corresponding to each input of the evaluation dataset is obtained. Based on the comparison result corresponding to each input of the evaluation dataset, the evaluation information of the pre-trained large language model in the evaluation stage is determined. Thus, by inputting the evaluation dataset for each evaluation stage multiple times to the pre-trained large language model, with the order of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset being different in each input, and by comparing the target prediction result output for each statement to be analyzed in the evaluation dataset with the corresponding correct analysis result, the comparison result corresponding to each input of the evaluation dataset can be obtained. This eliminates the model's potential bias towards the order of results and improves the reliability of the evaluation information obtained based on the comparison result.

[0209] In one embodiment, the comparison result includes ratio information. The target prediction result for each output of the statement to be analyzed in the evaluation dataset is compared with the corresponding correct analysis result to obtain the comparison result for each input of the evaluation dataset. This includes: during the process of obtaining the ratio information for any input of the evaluation dataset, comparing the correct analysis result corresponding to the statement to be analyzed in the evaluation dataset with the target prediction result corresponding to any input to determine a first quantity; the first quantity is the number of statements to be analyzed whose target prediction result and corresponding correct analysis result are the same for any input; obtaining the ratio information between the first quantity and a second quantity as the ratio information for any input of the evaluation dataset; the second quantity is the total number of statements to be analyzed in the evaluation dataset.

[0210] In specific implementation, when the computer device compares the target prediction result of each output of the statement to be analyzed in the evaluation dataset with the corresponding correct analysis result to obtain the comparison result of the evaluation dataset under each input, it can compare the correct analysis result of the statement to be analyzed in the evaluation dataset with the target prediction result of the statement to be analyzed under any input to determine a first quantity. The first quantity is the number of statements to be analyzed whose target prediction result and corresponding correct analysis result are the same under any input. The ratio of the first quantity to the second quantity is obtained as the ratio information of the evaluation dataset under any input. The second quantity is the total number of statements to be analyzed in the evaluation dataset. This ratio information is the accuracy rate.

[0211] The technical solution of this embodiment includes a comparison result comprising ratio information. During the process of obtaining the ratio information corresponding to any input in the evaluation dataset, the correct analysis result corresponding to the statement to be analyzed in the evaluation dataset is compared with the target prediction result corresponding to any input to determine a first quantity. The first quantity is the number of statements to be analyzed whose target prediction result and corresponding correct analysis result are the same under any input. The ratio information between the first quantity and a second quantity is obtained as the ratio information corresponding to any input in the evaluation dataset. The second quantity is the total number of statements to be analyzed in the evaluation dataset. Thus, based on the ratio information corresponding to any input in the evaluation dataset, the accuracy of the pre-trained large language model under the order of results for that input can be determined.

[0212] It is understood that, in some embodiments, for any evaluation stage, the corresponding accuracy can also be calculated for different evaluation dimensions corresponding to that evaluation stage. Specifically, for each evaluation dimension, the number of statements to be analyzed whose target prediction result and corresponding correct analysis result are the same under that evaluation dimension is counted to obtain the third quantity for each evaluation dimension. Then, the total number of statements to be analyzed under each evaluation dimension is determined to obtain the fourth quantity for each evaluation dimension. The ratio information between the third quantity and the fourth quantity corresponding to each evaluation dimension is obtained as the accuracy of the pre-trained large language model under each evaluation dimension.

[0213] In one embodiment, the evaluation information of the pre-trained large language model in the evaluation stage is determined based on the comparison results corresponding to each input of the evaluation dataset, including: determining the average ratio information based on the ratio information corresponding to each input of the evaluation dataset; and using the average ratio information as the evaluation information of the pre-trained large language model in the evaluation stage.

[0214] The average ratio information is the average accuracy rate.

[0215] In practice, when the computer device determines the evaluation information of the pre-trained large language model in the evaluation stage based on the comparison results corresponding to each input of the evaluation dataset, it can determine the average ratio information based on the ratio information corresponding to each input of the evaluation dataset; and use the average ratio information as the evaluation information of the pre-trained large language model in the evaluation stage.

[0216] The technical solution of this embodiment determines the average ratio information based on the ratio information corresponding to each input in the evaluation dataset; this average ratio information is then used as the evaluation information for the pre-trained large language model during the evaluation process. Thus, since the order of the dialogue state analysis results set differs for each input, by statistically averaging the ratio information corresponding to each input, random fluctuations can be reduced, stability improved, and the reliability of the evaluation information is enhanced by using the average value as the evaluation information for the pre-trained large language model during the evaluation process.

[0217] In one embodiment, the evaluation dataset corresponding to the evaluation stage is input into a pre-trained large language model, and the target prediction result corresponding to the statement to be analyzed in the evaluation dataset is output. This includes: inputting the statement to be analyzed and the corresponding dialogue state analysis result set in the evaluation dataset into the pre-trained large language model, and outputting the target prediction result corresponding to the statement to be analyzed in the evaluation dataset under this input; reordering the result arrangement order of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset, using the reordered dialogue state analysis result set as the new dialogue state analysis result set, and returning to the step of inputting the statement to be analyzed and the corresponding dialogue state analysis result set in the evaluation dataset into the pre-trained large language model, until the second preset number of times is met.

[0218] In specific implementation, during the process of inputting the evaluation dataset corresponding to the evaluation stage into the pre-trained large language model and outputting the target prediction result corresponding to the statement to be analyzed in the evaluation dataset, the computer device can input the statement to be analyzed in the evaluation dataset and the corresponding dialogue state analysis result set into the pre-trained large language model, output the target prediction result corresponding to the statement to be analyzed in the evaluation dataset under this input, reorder the results of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset, and use the reordered dialogue state analysis result set as the new dialogue state analysis result set. The process returns to the step of inputting the statement to be analyzed in the evaluation dataset and the corresponding dialogue state analysis result set into the pre-trained large language model until the second preset number of times is met, ensuring that the order of the results of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset is different in each input, thereby obtaining the target prediction result corresponding to the statement to be analyzed in the evaluation dataset corresponding to the evaluation stage under each input.

[0219] The technical solution of this embodiment involves inputting the statement to be analyzed and its corresponding dialogue state analysis result set from the evaluation dataset into a pre-trained large language model, outputting the target prediction result corresponding to the statement to be analyzed in the evaluation dataset under this input; reordering the results of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset, using the reordered dialogue state analysis result set as the new dialogue state analysis result set, and returning to the step of inputting the statement to be analyzed and its corresponding dialogue state analysis result set from the evaluation dataset into the pre-trained large language model, until a second preset number of times is met. In this way, it can be ensured that the order of the results of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset is different in each input, obtaining the target prediction answer corresponding to the statement to be analyzed under the dialogue state analysis result set with different result ordering, thereby eliminating the model's potential bias towards the result ordering.

[0220] In some other embodiments, for any evaluation stage, the evaluation dimension under any evaluation stage can correspond to different evaluation weights. Thus, in the process of calculating the accuracy of the pre-trained large language model under any evaluation stage, the number of sentences to be analyzed that have the same target prediction result and the same correct analysis result can be counted according to different evaluation dimensions to obtain the third quantity corresponding to different evaluation dimensions. Then, based on the fourth quantity, the third quantity, and the evaluation weight corresponding to each of the different evaluation dimensions, the accuracy of the pre-trained large language model under different evaluation dimensions can be calculated.

[0221] In another embodiment, such as Figure 6 The diagram illustrates a flowchart of a method for outputting evaluation information from a large language model, comprising the following steps:

[0222] Step S602: Obtain the evaluation dataset corresponding to each evaluation stage of the preset dialogue state evaluation process.

[0223] Step S604: For any statement to be analyzed in the evaluation dataset, input the statement to be analyzed and the corresponding dialogue state analysis result set into the pre-trained large language model repeatedly according to the first preset number of times, and output the prediction result corresponding to the statement to be analyzed under each input.

[0224] Step S606: The prediction result that appears most frequently is taken as the target prediction result for any statement to be analyzed.

[0225] Step S608: Input the statement to be analyzed and the corresponding dialogue state analysis result set in the evaluation dataset into the pre-trained large language model, and output the target prediction result corresponding to the statement to be analyzed in the evaluation dataset under this input.

[0226] Step S610: Reorder the order of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset, take the reordered dialogue state analysis result set as the new dialogue state analysis result set, and return to the step of inputting the statement to be analyzed and the corresponding dialogue state analysis result set in the evaluation dataset into the pre-trained large language model until the second preset number of times is met.

[0227] Step S612: In the process of obtaining the ratio information corresponding to the evaluation dataset under any input, the correct analysis result corresponding to the statement to be analyzed in the evaluation dataset is compared with the target prediction result corresponding to any input to determine the first quantity.

[0228] Step S614: Obtain the ratio information of the first quantity to the second quantity, as the ratio information corresponding to the evaluation dataset under any input.

[0229] Step S616: Determine the average ratio information based on the ratio information corresponding to each input in the evaluation dataset.

[0230] Step S618: Use the average ratio information as the evaluation information for the pre-trained large language model in the evaluation phase.

[0231] It should be noted that the specific limitations of the above steps can be found in the specific limitations of the evaluation information output method for a large language model described above.

[0232] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0233] Based on the same inventive concept, this application also provides an apparatus for constructing an evaluation dataset for a large language model to implement the method for constructing the evaluation dataset of the large language model described above. The solution provided by this apparatus is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the apparatus for constructing an evaluation dataset for a large language model provided below can be found in the limitations of the method for constructing the evaluation dataset of a large language model described above, and will not be repeated here.

[0234] In one exemplary embodiment, such as Figure 7 As shown, an apparatus for constructing an evaluation dataset for a large language model is provided, comprising: a sentence acquisition module 710, an expansion module 720, an optimization module 730, and a construction module 740, wherein:

[0235] The statement acquisition module 710 is used to acquire dialogue state topic statements corresponding to different evaluation dimensions; the dialogue state topic statements contain dialogue state information and topic scene information; the evaluation dimensions are dimensions used to evaluate the dialogue state perception ability of the pre-trained large language model.

[0236] The extension module 720 is used to extend the dialogue state topic statement to obtain an extended dialogue state topic statement; the extended dialogue state topic statement includes preset key elements for analyzing the dialogue state.

[0237] The optimization module 730 is used to optimize the expanded dialogue state topic statement to obtain the statement to be analyzed; the richness of the dialogue state information of the statement to be analyzed is higher than the richness of the dialogue state information of the expanded dialogue state topic statement.

[0238] The construction module 740 is used to obtain the dialogue state analysis result set corresponding to the statement to be analyzed, and to use the statement to be analyzed and the corresponding dialogue state analysis result set as an evaluation dataset; the dialogue state analysis result set includes correct analysis results and at least one analysis result that is different from the correct analysis results; the evaluation dataset is used to evaluate the dialogue state perception capability of the pre-trained large language model.

[0239] In one embodiment, the correct analysis result includes a first correct sub-result and a second correct sub-result. Each analysis result different from the correct analysis result includes a first sub-result and a second sub-result. The construction module 740 is specifically used to obtain the first correct sub-result and the corresponding second correct sub-result corresponding to the statement to be analyzed. The first correct sub-result represents the correct dialogue state category corresponding to the statement to be analyzed. The second correct sub-result represents the reason for the generation of the correct dialogue state category. At least one first sub-result different from the first correct sub-result is obtained. The first sub-result represents a dialogue state category different from the correct dialogue state category. At least one second sub-result different from the second correct sub-result is obtained. The second sub-result represents the reason for the generation of the dialogue state category different from the correct dialogue state category.

[0240] In one embodiment, the apparatus further includes: a result acquisition module, configured to acquire a first analysis result set corresponding to the statement to be analyzed; the first analysis result set includes a first category label and at least one second category label different from the first category label; the first category label represents the correct dialogue state category corresponding to the statement to be analyzed; the second category label represents a dialogue state category different from the correct dialogue state category; the first category label is used as the first correct sub-result, and the second category label is used as the first sub-result.

[0241] In one embodiment, the acquisition module is specifically used to acquire a second analysis result set corresponding to the statement to be analyzed; the second analysis result set includes first reason information and at least one second reason information different from the first reason information; the first reason information represents the reason for the generation of the correct dialogue state category; the second reason information represents the reason for the generation of a dialogue state category different from the correct dialogue state category; the first reason information is used as the second correct sub-result, and the second reason information is used as the second sub-result.

[0242] In one embodiment, the statement acquisition module 710 is specifically used to acquire a set of dialogue statements from a preset multi-turn dialogue dataset; the dialogue statements in the set of dialogue statements contain the dialogue state information; the set of dialogue statements is filtered to obtain a target set of dialogue statements; the target dialogue statements in the target set of dialogue statements meet preset dialogue quality conditions; an expanded target dialogue statement is acquired, and the expanded target dialogue statement is used as the dialogue state topic statement; the expanded target dialogue statement is obtained by expanding the topic scenario of the target dialogue statement according to the definition of the evaluation dimension.

[0243] In one embodiment, the optimization module 730 is specifically used to input the expanded dialogue state topic statement into a pre-trained statement optimization model and output the statement to be analyzed. The pre-trained statement optimization model adopts the following training method: obtaining a training sample set; the training sample set includes the original statement and the corresponding tag-optimized statement; the richness of dialogue state information of the tag-optimized statement is higher than the richness of dialogue state information of the original statement; inputting the original statement into the statement optimization model to be trained and outputting the predicted optimized statement corresponding to the original statement; adjusting the model parameters of the statement optimization model to be trained based on the difference between the predicted optimized statement corresponding to the original statement and the corresponding tag-optimized statement and continuing iterative training until the trained statement optimization model meets the training termination condition and training stops, thus obtaining the pre-trained statement optimization model.

[0244] Each module in the apparatus for constructing the evaluation dataset of the aforementioned large language model can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can invoke and execute the operations corresponding to each module.

[0245] Based on the same inventive concept, this application also provides an apparatus for outputting evaluation information of a large language model to implement the aforementioned method for outputting evaluation information of a large language model. The solution provided by this apparatus is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the large language model evaluation information output apparatus provided below can be found in the limitations of the large language model evaluation information output method described above, and will not be repeated here.

[0246] In one exemplary embodiment, such as Figure 8 As shown, a device for outputting evaluation information of a large language model is provided, including: a data acquisition module 810, a prediction module 820, and a comparison module 830, wherein:

[0247] The data acquisition module 810 is used to acquire the evaluation dataset corresponding to each evaluation stage of the preset dialogue state evaluation process; the dialogue state evaluation process is a process for evaluating the dialogue state perception ability of a pre-trained large language model; the evaluation dataset corresponding to each evaluation stage includes the statement to be analyzed under different evaluation dimensions and the corresponding dialogue state analysis result set; the statement to be analyzed contains dialogue state information; the dialogue state analysis result set corresponding to each statement to be analyzed includes the correct analysis result and at least one analysis result that is different from the correct analysis result.

[0248] The prediction module 820 is used to input the evaluation dataset corresponding to each evaluation stage into the pre-trained large language model in each evaluation stage, and output the target prediction result corresponding to the statement to be analyzed in the evaluation dataset; the target prediction result is the analysis result selected by the pre-trained large language model in the dialogue state analysis result set corresponding to the input statement to be analyzed.

[0249] The comparison module 830 is used to compare the correct analysis results and the corresponding target prediction results of the sentences to be analyzed in the evaluation dataset to obtain the evaluation information of the pre-trained large language model in each evaluation stage; the evaluation information is used to characterize the dialogue state perception ability of the pre-trained large language model.

[0250] In one embodiment, the prediction module 820 is specifically used to repeatedly input any statement to be analyzed and its corresponding dialogue state analysis result set into the pre-trained large language model for any statement to be analyzed in the evaluation dataset according to a first preset number of times, and output the prediction result corresponding to any statement to be analyzed under each input; and take the prediction result that appears most frequently as the target prediction result corresponding to any statement to be analyzed.

[0251] In one embodiment, the number of times the evaluation dataset corresponding to each evaluation stage is input to the pre-trained large language model satisfies a second preset number, and the order of the results of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset is different in each input; there are multiple target prediction results corresponding to the statement to be analyzed in the evaluation dataset, which respectively correspond to the analysis results selected by the pre-trained large language model in the dialogue state analysis result set of each input; the comparison module 830 is specifically used to compare the target prediction result output for the statement to be analyzed in the evaluation dataset with the corresponding correct analysis result each time, to obtain the comparison result corresponding to the evaluation dataset under each input; and to determine the evaluation information of the pre-trained large language model under the evaluation stage based on the comparison result corresponding to the evaluation dataset under each input.

[0252] In one embodiment, the comparison result includes ratio information. The comparison module 830 is specifically used to compare the correct analysis result corresponding to the statement to be analyzed in the evaluation dataset with the target prediction result corresponding to the statement under any input during the process of obtaining the ratio information corresponding to the evaluation dataset under any input, and determine a first quantity; the first quantity is the number of statements to be analyzed whose target prediction result and corresponding correct analysis result are the same under any input; obtain the ratio information between the first quantity and the second quantity as the ratio information corresponding to the evaluation dataset under any input; the second quantity is the total number of statements to be analyzed in the evaluation dataset.

[0253] In one embodiment, the comparison module 830 is specifically used to determine the average ratio information based on the ratio information corresponding to each input of the evaluation dataset; and to use the average ratio information as the evaluation information of the pre-trained large language model in the evaluation process.

[0254] In one embodiment, the prediction module 820 is specifically used to input the statement to be analyzed and the corresponding dialogue state analysis result set in the evaluation dataset into a pre-trained large language model, output the target prediction result corresponding to the statement to be analyzed in the evaluation dataset under this input; reorder the result arrangement order of the dialogue state analysis result set corresponding to the statement to be analyzed in the evaluation dataset, take the reordered dialogue state analysis result set as the new dialogue state analysis result set, and return to the step of inputting the statement to be analyzed and the corresponding dialogue state analysis result set in the evaluation dataset into the pre-trained large language model, until the second preset number of times is met.

[0255] Each module in the evaluation information output device of the aforementioned large language model can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0256] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 9As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements the output of evaluation information for a large language model and / or a method for constructing an evaluation dataset for a large language model. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0257] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0258] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0259] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0260] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0261] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0262] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0263] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0264] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for constructing an evaluation dataset for a large language model, characterized in that, The method includes: Obtain dialogue state topic statements corresponding to different evaluation dimensions; the dialogue state topic statements have dialogue state information and topic scene information; the evaluation dimensions are dimensions used to evaluate the dialogue state perception ability of the pre-trained large language model. The dialogue state topic statement is expanded to obtain an expanded dialogue state topic statement; the expanded dialogue state topic statement includes preset key elements for analyzing the dialogue state. The expanded dialogue state topic statement is optimized to obtain the statement to be analyzed; the richness of the dialogue state information of the statement to be analyzed is higher than that of the expanded dialogue state topic statement. Obtain the dialogue state analysis result set corresponding to the statement to be analyzed, and use the statement to be analyzed and the corresponding dialogue state analysis result set as the evaluation dataset; the dialogue state analysis result set includes correct analysis results and at least one analysis result that is different from the correct analysis results; the evaluation dataset is used to evaluate the dialogue state awareness capability of the pre-trained large language model; The correct analysis result includes a first correct sub-result and a second correct sub-result, and each analysis result that differs from the correct analysis result includes a first sub-result and a second sub-result; obtaining the dialogue state analysis result set corresponding to the statement to be analyzed includes: Obtain the first correct sub-result and the corresponding second correct sub-result for the statement to be analyzed; the first correct sub-result represents the correct dialogue state category corresponding to the statement to be analyzed; the second correct sub-result represents the reason for the generation of the correct dialogue state category. Obtain at least one first sub-result that is different from the first correct sub-result; the first sub-result represents a dialogue state category that is different from the correct dialogue state category; Obtain at least one second sub-result that is different from the second correct sub-result; the second sub-result represents the reason for the occurrence of a dialogue state category that is different from the correct dialogue state category.

2. The method according to claim 1, characterized in that, The method further includes: Obtain a first analysis result set corresponding to the statement to be analyzed; the first analysis result set includes a first category label and at least one second category label that is different from the first category label; the first category label represents the correct dialogue state category corresponding to the statement to be analyzed; the second category label represents a dialogue state category that is different from the correct dialogue state category. The first category label is used as the first correct sub-result, and the second category label is used as the first sub-result.

3. The method according to claim 1, characterized in that, The method further includes: Obtain the second analysis result set corresponding to the statement to be analyzed; the second analysis result set includes first reason information and at least one second reason information that is different from the first reason information; the first reason information represents the reason for the generation of the correct dialogue state category; the second reason information represents the reason for the generation of a dialogue state category that is different from the correct dialogue state category; The first reason information is used as the second correct sub-result, and the second reason information is used as the second sub-result.

4. The method according to claim 1, characterized in that, The step of obtaining the dialogue state topic statements corresponding to different evaluation dimensions includes: Obtain a set of dialogue statements from a pre-defined multi-turn dialogue dataset; the dialogue statements in the set of dialogue statements contain the dialogue state information. The set of dialogue statements is filtered to obtain a target set of dialogue statements; the target dialogue statements in the target set of dialogue statements meet the preset dialogue quality conditions. Obtain the expanded target dialogue statement and use the expanded target dialogue statement as the dialogue state topic statement; the expanded target dialogue statement is obtained by expanding the topic scenario of the target dialogue statement according to the definition of the evaluation dimension.

5. The method according to claim 1, characterized in that, The optimization of the expanded dialogue state topic statement to obtain the statement to be analyzed includes: The expanded dialogue state topic statement is input into a pre-trained statement optimization model, which outputs the statement to be analyzed. The pre-trained statement optimization model uses the following training method: Obtain a training sample set; the training sample set includes the original statement and the corresponding tag-optimized statement; the richness of the dialogue state information of the tag-optimized statement is higher than that of the original statement; The original statement is input into the statement optimization model to be trained, and the predicted optimized statement corresponding to the original statement is output. Based on the difference between the predicted optimized statement and the corresponding labeled optimized statement corresponding to the original statement, the model parameters of the statement optimization model to be trained are adjusted and iterative training continues until the trained statement optimization model meets the training termination condition, at which point training stops, and the pre-trained statement optimization model is obtained.

6. A device for constructing an evaluation dataset for a large language model, characterized in that, The device includes: The statement acquisition module is used to acquire dialogue state topic statements corresponding to different evaluation dimensions; the dialogue state topic statements have dialogue state information and topic scene information; the evaluation dimensions are dimensions used to evaluate the dialogue state perception ability of the pre-trained large language model. An extension module is used to expand the dialogue state topic statement to obtain an expanded dialogue state topic statement; the expanded dialogue state topic statement includes preset key elements for analyzing the dialogue state. An optimization module is used to optimize the expanded dialogue state topic statement to obtain a statement to be analyzed; the richness of the dialogue state information of the statement to be analyzed is higher than the richness of the dialogue state information of the expanded dialogue state topic statement. A construction module is used to obtain the dialogue state analysis result set corresponding to the statement to be analyzed, and to use the statement to be analyzed and the corresponding dialogue state analysis result set as an evaluation dataset; the dialogue state analysis result set includes correct analysis results and at least one analysis result that is different from the correct analysis results; the evaluation dataset is used to evaluate the dialogue state awareness capability of the pre-trained large language model; the correct analysis results include a first correct sub-result and a second correct sub-result, and each analysis result that is different from the correct analysis result includes a first sub-result and a second sub-result; The construction module is specifically used to obtain the first correct sub-result and the corresponding second correct sub-result corresponding to the statement to be analyzed; the first correct sub-result represents the correct dialogue state category corresponding to the statement to be analyzed; the second correct sub-result represents the reason for the generation of the correct dialogue state category. Obtain at least one first sub-result that is different from the first correct sub-result; the first sub-result represents a dialogue state category that is different from the correct dialogue state category; Obtain at least one second sub-result that is different from the second correct sub-result; the second sub-result represents the reason for the occurrence of a dialogue state category that is different from the correct dialogue state category.

7. The apparatus according to claim 6, characterized in that, The device further includes: a result acquisition module, used to acquire a first analysis result set corresponding to the statement to be analyzed; the first analysis result set includes a first category label and at least one second category label different from the first category label; the first category label represents the correct dialogue state category corresponding to the statement to be analyzed; the second category label represents a dialogue state category different from the correct dialogue state category; the first category label is used as the first correct sub-result, and the second category label is used as the first sub-result.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Dialogue type question and answer implementation method oriented to intelligent data visualization

    CN113111158A

  • Dialogue processing method and device based on large model, electronic equipment and storage medium

    CN119149695A