Dialogue model evaluation and dialogue remodeling method, equipment and medium

By setting scene elements and multi-dimensional quantization and hierarchical analysis structure methods, the performance of large language models in human-computer dialogue is evaluated and reshape the problem of bias in evaluation results, and the evaluation accuracy and dialogue quality are improved.

CN120012946AInactive Publication Date: 2025-05-16CENT SOUTH UNIV
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510486691.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

There is a bias in the performance evaluation of large language models in human-computer dialogue, resulting in inaccuracy and difficulty of evaluation results.

Method used

By setting scene elements, building virtual dialogue roles and dialogue datasets, a multi-dimensional quantitative and hierarchical analysis structure method is adopted, including the analysis layer, the review layer and the evaluation layer, to evaluate the dialogue model and dialogue reshaping.

Benefits of technology

It improves the evaluation accuracy and logical consistency of the dialogue model, ensures that the generated dialogue is closely linked to the core topics and is goal-oriented, and improves the conversation quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012946A_ABST
    Figure CN120012946A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of model evaluation, in particular to a dialogue model evaluation and dialogue remodeling method and device and a medium, and the method comprises the following steps: constructing a dialogue data set based on scene elements; constructing input data based on the dialogue data set, constructing an analysis layer, and inputting the input data into the analysis layer to obtain a dimension analysis result; constructing an examination layer, and inputting the dimension analysis result into the examination layer to obtain a dimension examination result; constructing an evaluation layer, and comprehensively evaluating the input data to obtain an evaluation result; multi-dimensional deconstruction is realized through a thinking chain and an auto-reflection technology, and a multi-dimensional deconstruction initial reply is obtained; and performing rehearsal feedback to realize dialogue remodeling. According to the method, scene elements are defined, so that a dialogue data set is accurately constructed, the model is evaluated in a multi-dimensional quantization and hierarchical analysis structure mode, dialogue remodeling is carried out, and dialogue content generated by the dialogue model is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of model evaluation technology, and in particular to a dialogue model evaluation and dialogue reshaping method, device and medium. Background Art

[0002] Large Language Model (LLM) is essentially a generative model. It can generate corresponding dialogue responses based on input dialogue information, so it can be widely used in consulting, analysis, chat and other scenarios.

[0003] For large language models, how to evaluate the performance of the model in human-computer dialogue is one of the basic issues. The dialogue performance evaluation of large language models mainly evaluates the dialogue interaction ability of large language models. For example, based on the reaction and behavior characteristics of large language models in different situations, the understanding ability, generation ability, logical reasoning ability, emotional understanding ability and other aspects of the large language model are evaluated and analyzed.

[0004] In related technologies, the dialogue information output by a large language model can be used to evaluate the dialogue interaction ability of the model. However, due to the unpredictability of the output of a large language model, even if the same information is input, the dialogue information output by the model will have certain differences due to the different probabilities obtained each time, which will cause deviations in the evaluation results based on a single dialogue information, affecting the accuracy of the evaluation results and increasing the difficulty of model evaluation.

[0005] Therefore, it is necessary to design a dialogue model evaluation and dialogue reshaping method, equipment and medium to solve the above technical problems. Summary of the invention

[0006] The present invention aims to provide a method, device and medium for dialogue model evaluation and dialogue reshaping. The specific technical scheme is as follows: A method for evaluating a dialogue model and reshaping a dialogue includes the following steps: S1, setting scene elements, constructing two virtual dialogue roles based on the scene elements, calling the dialogue model to implement the two virtual dialogue roles, and then generating dialogue data. Different dialogue data are generated by setting different scene elements to form a dialogue data set; S2, constructing input data based on the conversation data set, constructing an analysis layer, wherein the analysis layer includes multiple analyzers of different dimensions, inputting the input data into the analyzer, and obtaining a dimensional analysis result; S3: constructing a review layer, the review layer including multiple reviewers of different dimensions, the reviewers corresponding to the analyzers one by one, and used to review the dimension analysis results and adjust the dimension analysis results to obtain the dimension review results; S4: constructing an evaluation layer, wherein the evaluation layer is used to integrate the review results of multiple dimensions, comprehensively evaluate the input data, obtain an evaluation result, and complete the evaluation of the dialogue model based on the evaluation result; S5, multidimensional deconstruction, which is achieved through thinking chain and self-reflection techniques; S6, dialogue reshaping, rehearsing feedback on the initial response to the multidimensional deconstruction, obtaining the reaction of the interlocutor, and guiding the initial response to the multidimensional deconstruction based on the evaluation results, forming feedback suggestions, and generating the final response through the reaction of the interlocutor and the feedback suggestions.

[0007] Optionally, in S1, the scene elements include the scene theme, the first speaker scene information, the second speaker scene information and the scene goal, and the construction expression of the virtual character is as follows: ; ; in, A virtual dialogue character representing the first speaker, A virtual dialogue character representing the second speaker, Indicates the role construction process of the first speaker, through the scene theme and the first speaker scene information , complete the construction of the virtual dialogue role of the first speaker; Indicates the role construction process of the second speaker, through the scene theme , Second speaker scene information And the scene goal , completing the construction of the virtual dialogue role of the second speaker.

[0008] Optionally, in S2, the input data includes conversation records, replies to be evaluated and other conversation data, the conversation records specifically refer to the content before the replies to be evaluated in the conversation data, and the other conversation data specifically refer to the content after the replies to be evaluated in the conversation data.

[0009] Optionally, in S2, the analyzer is constructed by combining a large language model with analysis prompt words, and the analysis prompt words include dimension information, conversation records, responses to be evaluated, scenario elements, analysis output format, and analysis tasks; The dimension information includes dimension name, dimension definition and scoring criteria; The output format includes dimension information, dimension performance analysis and dimension scoring; The analysis task is to evaluate and analyze the input data and output the dimensional analysis results in an analysis output format.

[0010] Optionally, in S3, the reviewer is constructed by combining a large language model with review prompt words, and the review prompt words include dimension information, conversation records, responses to be evaluated, scenario elements, dimension analysis results, output format, and review tasks; The review task is to review the dimension analysis results of the current dimension and adjust the dimension performance analysis and dimension scoring in the dimension analysis results, and output the dimension review results in an output format.

[0011] Optionally, in S3, the evaluation layer is constructed by combining a large language model with evaluation prompt words, and the evaluation prompt words include conversation records, responses to be evaluated, scenario data, dimension review results, evaluation criteria, evaluation output format, and evaluation tasks; The assessment criteria are assessment criteria set based on each communication ability level; The assessment output format includes communication ability level, performance and dimension scores of each dimension, total score, cause analysis and feedback suggestions; The assessment task is to conduct cause analysis and feedback suggestions based on the review results of each dimension, and output the assessment results in the assessment output format.

[0012] Optionally, in S5, a thought chain is constructed based on the dialogue record and scenario data, and a multi-dimensional deconstructed initial response is obtained through self-reflection technology. The expression of the thought chain is as follows: ; ; ; ; in, The result of the thought chain reasoning in the perception dimension indicates whether the other party has expressed needs and emotions based on the conversation records and scene data from the perception dimension. A chain of thought that represents the dimension of perception; Indicates a conversation record; The result of the thought chain reasoning in the expectation dimension, which means analyzing the information or goals that the other party expects to obtain from the conversation based on the conversation records and scene data from the expectation dimension; A chain of thoughts that represents the desired dimension; The result of the thought chain reasoning of the participation dimension indicates whether the participant actively participates in the conversation and whether the topic can be promoted to ensure the fluency of communication based on the conversation records and scene data. A chain of thoughts that represents the dimension of participation; The result of the thought chain reasoning in the information dimension indicates that from the information dimension, based on the conversation records and scene data, the original conversation clearly and completely conveyed the information needed by the other party, and whether there were any misunderstandings or ambiguities. A chain of thoughts that represents the dimension of participation; The self-reflection technique is used to guide the dialogue model to generate a multi-dimensional deconstructed initial response, which is expressed as follows: ; in, represents the initial response of multidimensional deconstruction, Indicates the generation of the initial response. Represents self-reflective technology.

[0013] Optionally, the dialogue model evaluation and dialogue reshaping method further includes dialogue reshaping, which is as follows: Optionally, in S6, a preview feedback is performed on the multi-dimensional deconstructed initial response, the response of the interlocutor is obtained, and feedback suggestions are generated based on the evaluation results. The final response is generated through the response of the interlocutor and the feedback suggestions. The calculation expression of the final response is as follows: ; ; ; in, Indicates the reaction of the interlocutor; Represents a simulation function for generating initial responses to multidimensional deconstruction the interlocutor's reaction; Express feedback and suggestions; Represents an analytical function for initial response to multidimensional deconstruction Generate feedback suggestions; Indicates a reply awaiting evaluation; Represents the output of the assessment layer; Indicates the final reply; Represents the output function, which is used to evaluate the dialogue model's comprehensive preview of the dialogue partner's reactions and feedback suggestions to generate the final response.

[0014] In addition, the present invention also includes a computer device, including a memory and a processor; The memory is used to store a computer program executable on the processor; The processor is used to implement the steps of the above-mentioned dialogue model evaluation and dialogue reshaping method when executing the computer program.

[0015] In addition, the present invention also includes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the dialogue model evaluation and dialogue reshaping method as described above are implemented.

[0016] The application of the technical solution of the present invention has the following beneficial effects: The present invention provides a method for evaluating a dialogue model and reshaping a dialogue. The method of the present invention accurately constructs a dialogue dataset by defining scene elements. Compared with existing methods, the advantage of the method of the present invention is that by clarifying the role background and goals, the generated dialogue is ensured to be closely related to the core topic and goal-oriented, thereby improving the logical consistency and practicality of the dialogue. In addition, the method of the present invention adopts a multi-dimensional quantification and hierarchical analysis structure to divide the evaluation process into an analysis layer, a review layer, and an assessment layer, which can ensure the gradual deepening and accuracy of the evaluation process.

[0017] In addition to the above-described purposes, features and advantages, the present invention has other purposes, features and advantages. The present invention will be further described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 It is a flowchart of the steps of the dialogue model evaluation and dialogue reshaping method in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0021] like Figure 1 As shown, this embodiment provides a method for evaluating a dialogue model and reshaping a dialogue, comprising the following steps: S1, setting scene elements, building two virtual dialogue roles based on the scene elements, calling the dialogue model to implement the two virtual dialogue roles, and then generating dialogue data. By setting different scene elements to generate different dialogue data, a dialogue data set is formed.

[0022] Specifically, the scene elements include the scene theme, the first speaker scene information, the second speaker scene information and the scene goal. The construction expression of the virtual character is as follows: ; ; in, A virtual dialogue character representing the first speaker, A virtual dialogue character representing the second speaker, Indicates the role construction process of the first speaker, through the scene theme and the first speaker scene information , complete the construction of the virtual dialogue role of the first speaker; Indicates the role construction process of the second speaker, through the scene theme , Second speaker scene information And the scene goal , completing the construction of the virtual dialogue role of the second speaker.

[0023] Furthermore, when constructing the dialogue data, the virtual dialogue character and Need to be in the process of dialogue, according to the dialogue record Generate a corresponding response In order to ensure the continuity of the conversation, the newly generated reply needs to be spliced ​​with the previous conversation record to achieve the conversation record The virtual characters have dialogues under each scene element to build the dialogue data under that scene. ,in, Indicates Virtual dialogue character in the round The generated reply has the following content format: "Speaker 1: reply content"; Indicates Virtual dialogue character in the round The generated reply has the following content format: "Speaker 2: reply content". The method in this embodiment constructs a dialogue dataset by conducting a dialogue between virtual dialogue characters under multiple scene elements. .

[0024] It should be noted that the method of this embodiment accurately constructs a dialogue data set by defining scene elements. Compared with existing methods, the advantage of the method of this embodiment is that by clarifying the role background and goals, it ensures that the generated dialogue is closely related to the core topic and is goal-oriented, thereby improving the logical consistency and practicality of the dialogue. At the same time, by updating the dialogue record in real time, the coherence and natural fluency of the dialogue are guaranteed, so that the virtual characters can interact more intelligently and naturally in complex scenes.

[0025] S2, constructing input data based on the conversation data set, constructing an analysis layer, wherein the analysis layer includes multiple analyzers of different dimensions, inputting the input data into the analyzer, and obtaining dimensional analysis results.

[0026] The input data includes conversation records, replies to be evaluated, and other conversation data. The conversation records are specifically the content before the replies to be evaluated in the conversation data, and the other conversation data are specifically the content after the replies to be evaluated in the conversation data.

[0027] The analyzer is constructed by combining a large language model with analysis prompt words, which include dimension information, conversation records, responses to be evaluated, scenario elements, analysis output format, and analysis tasks; The dimension information includes dimension name, dimension definition and scoring criteria; The output format includes dimension information, dimension performance analysis and dimension scoring; The analysis task is to evaluate and analyze the input data and output the dimensional analysis results in an analysis output format.

[0028] It should be noted that the dimensions in this embodiment include perception, expectation, participation and information, and the definitions of the dimensions are shown in Table 1.

[0029] Table 1 Communication ability assessment dimensions and their definitions

[0030] S3: Construct a review layer, which includes multiple reviewers of different dimensions. The reviewers correspond to the analyzers one by one and are used to review the dimensional analysis results and adjust the dimensional analysis results to obtain dimensional review results.

[0031] The reviewer is constructed by combining a large language model with review prompt words, which include dimension information, conversation records, responses to be evaluated, scenario elements, dimension analysis results, output format, and review tasks; The review task is to review the dimension analysis results of the current dimension and adjust the dimension performance analysis and dimension scoring in the dimension analysis results, and output the dimension review results in an output format.

[0032] S4: Construct an assessment layer, which is used to integrate the review results of multiple dimensions, conduct a comprehensive assessment on the input data, obtain an assessment result, and complete the evaluation of the dialogue model based on the assessment result.

[0033] The assessment layer is constructed by combining a large language model with assessment prompt words, and the assessment prompt words include conversation records, responses to be evaluated, scenario data, dimension review results, assessment criteria, assessment output format, and assessment tasks; The assessment criteria are assessment criteria set based on each communication ability level; The assessment output format includes communication ability level, performance and dimension scores of each dimension, total score, cause analysis and feedback suggestions; The assessment task is to conduct cause analysis and feedback suggestions based on the review results of each dimension, and output the assessment results in the assessment output format.

[0034] It should be noted that the large language model (LLM) used in this embodiment is an advanced artificial intelligence model trained based on large-scale text data, which can generate fluent and natural language text and has the ability to deeply analyze text semantics. In the field of natural language processing, LLM is widely used in various tasks, such as summary summarization, dialogue question and answer, and multilingual translation. In this embodiment, LLM can be used by calling the API interface or by local deployment. Specifically, when selecting a model, it is selected according to the application scenario and actual effect. The GPT series can be used, or models such as Qwen and DeepSeek can be used, which are not limited here.

[0035] The method of this embodiment adopts a multi-dimensional quantification and hierarchical analysis structure, and divides the evaluation process into an analysis layer, a review layer, and an assessment layer, which can ensure the gradual deepening and accuracy of the evaluation process. The analysis layer first performs a preliminary performance analysis of the input data, scores the response content according to predefined dimensions and scoring criteria, and provides basic data for subsequent layers. On this basis, the review layer conducts a detailed review of the results of each dimension, verifies the accuracy and comprehensiveness of the scoring of the analysis layer, and ensures that the evaluation results of each dimension meet actual needs. The assessment layer conducts a comprehensive assessment of the overall conversation quality based on the feedback from the review layer, and provides detailed cause analysis and improvement suggestions. The hierarchical progressive structure not only ensures the independence and focus of each level, but also makes the entire evaluation process more systematic, accurate and credible, and improves the reliability and integrity of the final evaluation results.

[0036] S5, multidimensional deconstruction: multidimensional deconstruction is achieved through thinking chain and self-reflection techniques.

[0037] In S5, the expression of thought chain technology is as follows: ; ; ; ; in, The result of the thought chain reasoning in the perception dimension indicates whether the other party has expressed needs and emotions based on the conversation records and scene data from the perception dimension. A chain of thought that represents the dimension of perception; Indicates a conversation record; The result of the thought chain reasoning in the expectation dimension, which means analyzing the information or goals that the other party expects to obtain from the conversation based on the conversation records and scene data from the expectation dimension; A chain of thoughts that represents the desired dimension; The result of the thought chain reasoning of the participation dimension indicates whether the participant actively participates in the conversation and whether the topic can be promoted to ensure the fluency of communication based on the conversation records and scene data. A chain of thoughts that represents the dimension of participation; The result of the thought chain reasoning in the information dimension indicates that from the information dimension, based on the conversation records and scene data, the original conversation clearly and completely conveyed the information needed by the other party, and whether there were any misunderstandings or ambiguities. A chain of thoughts that represents the dimension of participation; The self-reflection technique is used to guide the dialogue model to generate a multi-dimensional deconstructed initial response, which is expressed as follows: ; in, represents the initial response of multidimensional deconstruction, Indicates the generation of the initial response. Represents self-reflective technology.

[0038] S6, dialogue reshaping: conduct rehearsal feedback on the initial response to the multidimensional deconstruction, obtain the reaction of the interlocutor, and guide the initial response to the multidimensional deconstruction based on the evaluation results, form feedback suggestions, and generate the final response through the reaction of the interlocutor and the feedback suggestions.

[0039] Specifically, the expression of dialogue reshaping is as follows: ; ; ; in, Indicates the reaction of the interlocutor; Represents a simulation function for generating initial responses to multidimensional deconstruction the interlocutor's reaction; Express feedback and suggestions; Represents an analytical function for initial response to multidimensional deconstruction Generate feedback suggestions; Indicates a reply awaiting evaluation; Represents the output of the assessment layer; Indicates the final reply; Represents the output function, which is used to evaluate the dialogue model's comprehensive preview of the dialogue partner's reactions and feedback suggestions to generate the final response.

[0040] The method of this embodiment reshapes the conversation based on multi-dimensional deconstruction and preview feedback. Through gradual deconstruction, simulated feedback and optimized reshaping, it can better adapt to complex conversation scenarios. Compared with the traditional simple reply generation method, the method of this embodiment can understand the conversation record at a higher level, optimize the logic, emotional consistency and goal orientation of the conversation, and significantly improve the quality of the conversation. For fields such as intelligent customer service and virtual assistants that have strong demands for communication skills, the method of this embodiment can enhance the interactive experience with users and enhance flexibility and intelligence.

[0041] In addition, the present invention also includes a computer device, including a memory and a processor; The memory is used to store a computer program executable on the processor; The processor is used to implement the steps of the above-mentioned dialogue model evaluation and dialogue reshaping method when executing the computer program.

[0042] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, which are used to describe the execution process of the computer program in the computer device.

[0043] The computer device may be a computing device such as a mobile phone, a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may include, but is not limited to, a processor and a memory. For example, the computer device may also include an input / output device, a network access device, a bus, etc.

[0044] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device, and uses various interfaces and lines to connect various parts of the entire computer device.

[0045] The memory can be used to store the computer program and / or module, and the processor implements the computer program by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0046] Wherein, if the module / unit integrated in the computer device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0047] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned dialogue model evaluation and dialogue reshaping method are implemented.

[0048] This embodiment provides a method for evaluating a dialogue model and reshaping a dialogue. The method of this embodiment defines scene elements to accurately construct a dialogue data set. Compared with the existing methods, the advantage of the method of this embodiment is that by clarifying the role background and goals, the generated dialogue is ensured to be closely related to the core topic and goal-oriented, thereby improving the logical consistency and practicality of the dialogue. In addition, the method of this embodiment adopts a multi-dimensional quantification and hierarchical analysis structure to divide the evaluation process into an analysis layer, a review layer, and an assessment layer, which can ensure the gradual deepening and accuracy of the evaluation process. In addition, the method of this embodiment also includes dialogue reshaping. This embodiment generates a multi-dimensional deconstruction initial response through thinking chain technology and self-reflection technology, and performs preview feedback on the existing dialogue based on the multi-dimensional deconstruction initial response to obtain the dialogue party's response. According to the evaluation results and the dialogue party's response, the final response is calculated. Compared with the existing technology, the method of this embodiment can optimize the logic, emotional consistency and goal orientation of the dialogue, improve the quality of the dialogue generated by the dialogue model, and improve the user experience.

[0049] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0050] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for evaluating a dialogue model and reshaping a dialogue, characterized in that: The following steps are involved: S1, setting scene elements, constructing two virtual dialogue roles based on the scene elements, calling the dialogue model to implement the two virtual dialogue roles, and then generating dialogue data. Different dialogue data are generated by setting different scene elements to form a dialogue data set; S2, constructing input data based on the conversation data set, constructing an analysis layer, wherein the analysis layer includes multiple analyzers of different dimensions, inputting the input data into the analyzer, and obtaining a dimensional analysis result; S3: constructing a review layer, the review layer including multiple reviewers of different dimensions, the reviewers corresponding to the analyzers one by one, and used to review the dimension analysis results and adjust the dimension analysis results to obtain the dimension review results; S4: constructing an evaluation layer, wherein the evaluation layer is used to integrate the review results of multiple dimensions, comprehensively evaluate the input data, obtain an evaluation result, and complete the evaluation of the dialogue model based on the evaluation result; S5, multidimensional deconstruction, achieves multidimensional deconstruction through thought chain and self-reflection technology, and obtains the initial response of multidimensional deconstruction; S6, dialogue reshaping, rehearsing feedback on the initial response to the multidimensional deconstruction, obtaining the reaction of the interlocutor, and guiding the initial response to the multidimensional deconstruction based on the evaluation results, forming feedback suggestions, and generating the final response through the reaction of the interlocutor and the feedback suggestions.

2. The method for dialogue model evaluation and dialogue reconstruction according to claim 1, characterized in that: In S1, the scene elements include the scene theme, the first speaker scene information, the second speaker scene information and the scene goal. The construction expression of the virtual role is as follows: ; ; in, A virtual dialogue character representing the first speaker, A virtual dialogue character representing the second speaker, Indicates the role construction process of the first speaker, through the scene theme and the first speaker scene information , complete the construction of the virtual dialogue role of the first speaker; Indicates the role construction process of the second speaker, through the scene theme , Second speaker scene information And the scene goal , completing the construction of the virtual dialogue role of the second speaker.

3. The method for dialogue model evaluation and dialogue reconstruction according to claim 2, characterized in that: In S2, the input data includes conversation records, replies to be evaluated, and other conversation data. The conversation records are specifically the content before the replies to be evaluated in the conversation data, and the other conversation data are specifically the content after the replies to be evaluated in the conversation data.

4. The method for dialogue model evaluation and dialogue reconstruction according to claim 3, characterized in that: In S2, the analyzer is constructed by combining a large language model with analysis prompt words, and the analysis prompt words include dimension information, conversation records, responses to be evaluated, scene elements, analysis output format, and analysis tasks; The dimension information includes dimension name, dimension definition and scoring criteria; The output format includes dimension information, dimension performance analysis and dimension scoring; The analysis task is to evaluate and analyze the input data and output the dimensional analysis results in an analysis output format.

5. The method for dialogue model evaluation and dialogue reconstruction according to claim 4, characterized in that: In S3, the reviewer is constructed by combining a large language model with review prompt words, and the review prompt words include dimension information, conversation records, responses to be evaluated, scene elements, dimension analysis results, output format, and review tasks; The review task is to review the dimension analysis results of the current dimension and adjust the dimension performance analysis and dimension scoring in the dimension analysis results, and output the dimension review results according to the output format.

6. The method for dialogue model evaluation and dialogue reconstruction according to claim 5, characterized in that: In S3, the evaluation layer is constructed by combining a large language model with evaluation prompt words, and the evaluation prompt words include conversation records, responses to be evaluated, scenario data, dimension review results, evaluation criteria, evaluation output format, and evaluation tasks; The assessment criteria are assessment criteria set based on each communication ability level; The assessment output format includes communication ability level, performance and dimension scores of each dimension, total score, cause analysis and feedback suggestions; The assessment task is to conduct cause analysis and feedback suggestions based on the review results of each dimension, and output the assessment results in the assessment output format.

7. The method for dialogue model evaluation and dialogue reconstruction according to claim 6, characterized in that: In S5, a thought chain is constructed based on the dialogue records and scene data, and a multi-dimensional deconstructed initial response is obtained through self-reflection technology. The expression of the thought chain is as follows: ; ; ; ; in, The result of the thought chain reasoning in the perception dimension indicates whether the other party has expressed needs and emotions based on the conversation records and scene data from the perception dimension. A chain of thought that represents the dimension of perception; Indicates a conversation record; The result of the thought chain reasoning in the expectation dimension, which means analyzing the information or goals that the other party expects to obtain from the conversation based on the conversation records and scene data from the expectation dimension; A chain of thoughts that represents the desired dimension; The result of the thought chain reasoning of the participation dimension indicates whether the participant actively participates in the conversation and whether the topic can be promoted to ensure the fluency of communication based on the conversation records and scene data. A chain of thoughts that represents the dimension of participation; The result of the thought chain reasoning in the information dimension indicates that from the information dimension, based on the conversation records and scene data, the original conversation clearly and completely conveyed the information needed by the other party, and whether there were any misunderstandings or ambiguities. A chain of thoughts that represents the dimension of participation; The self-reflection technique is used to guide the dialogue model to generate a multi-dimensional deconstructed initial response, which is expressed as follows: ; in, represents the initial response of multidimensional deconstruction, Indicates the generation of the initial response. Represents self-reflective technology.

8. The method for dialogue model evaluation and dialogue reconstruction according to claim 7, characterized in that: In S6, the initial response of the multi-dimensional deconstruction is previewed and feedback is obtained to obtain the response of the interlocutor. Feedback suggestions are generated based on the evaluation results. The final response is generated through the response of the interlocutor and the feedback suggestions. The calculation expression of the final response is as follows: ; ; ; in, Indicates the reaction of the interlocutor; Represents a simulation function for generating initial responses to multidimensional deconstruction the interlocutor's reaction; Express feedback and suggestions; Represents an analytical function for initial response to multidimensional deconstruction Generate feedback suggestions; Indicates a reply awaiting evaluation; Represents the output of the assessment layer; Indicates the final reply; Represents the output function, which is used to evaluate the dialogue model's comprehensive preview of the dialogue partner's reactions and feedback suggestions to generate the final response.

9. A computer device, characterized in that: including memory and processor; The memory is used to store a computer program executable on the processor; The processor is used to implement the steps of the dialogue model evaluation and dialogue reshaping method as described in any one of claims 1 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the dialogue model evaluation and dialogue reshaping method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Assessment method and device of large language model and electronic equipment

    CN117112744A

  • Intelligent scoring method and system based on knowledge graph

    CN117312532A

  • Interaction method, device and product for increasing AI teaching-aided thinking conclusion ability

    CN117689028A

  • Question processing method and device, equipment and storage medium

    CN117743544A

  • Emotion support dialogue generation method, system and device based on thinking chain reasoning

    CN117932041A