Data processing method for human-machine interaction, server, and storage medium
By using the second human-computer interaction model to conduct multiple rounds of dialogue with the first human-computer interaction model, multiple rounds of dialogue information are solved, and the existing technology cannot accurately evaluate the multi-round dialogue capabilities of the human-computer interaction model is achieved, and high-quality model iteration and online launch are achieved.
Patent Information
- Application Number
- PCT/CN2024/125021
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-10-15
- Publication Date
- 2025-06-05
AI Technical Summary
The prior art cannot accurately evaluate the multi-round dialogue capabilities of human-computer interaction models, resulting in poor quality of model iteration and online launch.
By obtaining the pre-constructed conversation topic and the first human-computer interaction model to be evaluated, multiple rounds of dialogue are used to conduct multiple rounds of dialogue with the first human-computer interaction model, and multiple rounds of dialogue information are generated in order to accurately evaluate the model's multi-round dialogue capabilities.
It realizes accurate evaluation of the multi-round dialogue capabilities of human-computer interaction models, improves the quality of model iteration and online, and improves the quality of multi-round dialogues of human-computer interaction.
Smart Images

Figure CN2024125021_05062025_PF_FP_ABST
Abstract
Description
Human-computer interaction data processing method, server and storage medium
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on November 30, 2023, with application number 202311618092.5 and application name “Data processing method, server and storage medium for human-computer interaction”, the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0002] The present disclosure relates to computer technology, and in particular to a data processing method, a server, and a storage medium for human-computer interaction. Background Art
[0003] With the development of artificial intelligence, large models are widely used in human-computer interaction in the field of natural language processing. Currently, the evaluation of human-computer interaction models, such as large-scale language models, focuses on their single-round conversational capabilities. However, multi-round conversational capabilities are widely used in various scenarios, such as chatting, interactive games, and AI assistants. Many users do not initially formulate clear instructions and questions, necessitating multi-round conversations. This may also require the human-computer interaction model to provide support through Socratic-style questioning. Multi-round conversational capabilities are a critically valuable capability of current human-computer interaction models, and it is crucial to objectively evaluate these models' multi-round conversational capabilities.
[0004] The current evaluation method for multi-round dialogue capabilities uses pre-designed multi-round dialogue data. At least the questions input into the human-computer interaction model in each round are fixed. Regardless of whether the human-computer interaction model responds in the previous round and what the response is, the questions input into the human-computer interaction model in the next round will not change. This cannot simulate the real and natural interaction process of humans well, and cannot truly and accurately evaluate the multi-round dialogue capabilities of the human-computer interaction model. It is not conducive to selecting high-quality models with better multi-round dialogue capabilities in model iteration, and is not conducive to controlling the quality of multi-round dialogue of the online model, resulting in poor quality of human-computer interaction.
[0005] Summary of the Invention
[0006] The present disclosure provides a human-computer interaction data processing method, server, and storage medium, which are used to solve the problem that the existing technology cannot truly and accurately evaluate the multi-round dialogue capability of the human-computer interaction model, is not conducive to selecting high-quality models with better multi-round dialogue capabilities in model iteration, and is not conducive to controlling the quality of multi-round dialogue of the online model, resulting in poor human-computer interaction quality.
[0007] In a first aspect, the present disclosure provides a data processing method for human-computer interaction, comprising:
[0008] Obtain pre-built conversation topics and the first human-computer interaction model to be evaluated;
[0009] Conducting multiple rounds of dialogue with the first human-computer interaction model based on the conversation topic using a second human-computer interaction model to obtain multi-round dialogue information of the conversation topic, the dialogue information including: question information input to the first human-computer interaction model, and response information generated by the first human-computer interaction model to the question information;
[0010] Determine evaluation information of the multi-round dialogue capability of the first human-computer interaction model based on the multi-round dialogue information of the conversation topic.
[0011] In a second aspect, the present disclosure provides a data processing method for human-computer interaction, comprising:
[0012] A request for evaluating the multi-round conversation capability of the language model sent by a receiving device is received, and at least one pre-built conversation topic is obtained;
[0013] Conducting multiple rounds of dialogue with the language model based on the conversation topic using a second human-computer interaction model to obtain multiple rounds of dialogue information on the conversation topic, the dialogue information including: question information input to the language model and response information generated by the language model to the question information;
[0014] Determine evaluation information of the multi-round dialogue capability of the language model based on the multi-round dialogue information of the conversation topic.
[0015] In a third aspect, the present disclosure provides a server, comprising:
[0016] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to cause the server to execute the method described in the first aspect or the second aspect.
[0017] In a fourth aspect, the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in the first aspect or the second aspect is implemented.
[0018] In a fifth aspect, the present disclosure provides a computer program, which, when executed in a computer, causes the computer to execute the method described in the first aspect or the second aspect.
[0019] The data processing method, server, and storage medium for human-computer interaction provided by the present disclosure obtain a pre-built conversation topic and a first human-computer interaction model to be evaluated; use a second human-computer interaction model to conduct multiple rounds of dialogue with the first human-computer interaction model based on the conversation topic, thereby obtaining multiple rounds of dialogue information based on the conversation topic, the dialogue information including: question information input to the first human-computer interaction model and response information to the question information generated by the first human-computer interaction model; use the second human-computer interaction model to simulate the process of a human user conducting multiple rounds of dialogue with the first human-computer interaction model around the conversation topic, thereby generating multiple rounds of dialogue information based on the conversation topic. Furthermore, based on the multiple rounds of dialogue information on the conversation topic, evaluation information of the multiple rounds of dialogue capability of the first human-computer interaction model is determined, which can truly realize the evaluation of the multiple rounds of dialogue capability of the first human-computer interaction model and obtain accurate and high-quality evaluation information. The evaluation information is used to guide the online decision of the first human-computer interaction model or update the optimized version of the first human-computer interaction model. It can accurately select high-quality models in the iteration of the first human-computer interaction model, improve the quality of the multiple rounds of dialogue of the iteratively updated first human-computer interaction model, improve the quality of the multiple rounds of dialogue of the online model, and thus improve the quality of the multiple rounds of dialogue in human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0021] FIG1 is a schematic diagram of an example system architecture applicable to the present disclosure;
[0022] FIG2 is a flow chart of a data processing method for human-computer interaction provided by an exemplary embodiment of the present disclosure;
[0023] FIG3 is a flowchart of obtaining multi-round dialogue information based on a conversation topic according to an exemplary embodiment of the present disclosure;
[0024] FIG4 is a flowchart of evaluating context-independent multi-round dialogue capabilities according to an exemplary embodiment of the present disclosure;
[0025] FIG5 is a flowchart of evaluating context-independent multi-round dialogue capabilities provided by another exemplary embodiment of the present disclosure;
[0026] FIG6 is an interaction flow chart of a data processing method for human-computer interaction provided by an exemplary embodiment of the present disclosure;
[0027] FIG7 is a schematic diagram of the structure of a server provided in an embodiment of the present disclosure.
[0028] The above drawings illustrate specific embodiments of the present disclosure, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the present disclosure in any way, but rather to illustrate the concepts of the present disclosure to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0029] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0030] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and corresponding operation entrances must be provided for users to choose to authorize or refuse.
[0031] First, the terms involved in this disclosure are explained:
[0032] Conversation: A computer term that refers to the communication process between an end user and a human-computer interaction system. For example, a conversation occurs from the moment a user enters the system and begins using its functions to the moment they exit the system and end their interaction. During a conversation, a user enters a command or question, and the system responds to that command or question. This constitutes a round of dialogue, and a conversation can include one or more rounds of dialogue between the user and the system.
[0033] Visual question answering task: Given an input image and a question, determine the answer to the question from the visual information of the input image.
[0034] Image description task: Generate description text for the input image.
[0035] Visual entailment task: predict the semantic relevance of an input image and text, i.e., entailment, neutrality, or contradiction.
[0036] Referential expression and comprehension task: locate the image area corresponding to the input text in the input image based on the input text.
[0037] Image generation task: Generate an image based on the input description text.
[0038] Text-based sentiment classification task: predict the sentiment classification information of the input text.
[0039] Text summarization task: Generate summary information of the input text.
[0040] Multimodal tasks: refers to downstream tasks whose input and output data involve multiple modal data such as images and text, such as visual question answering tasks, image description tasks, visual implication tasks, referential expression and understanding tasks, image generation tasks, etc.
[0041] Multimodal pre-trained model: refers to a pre-trained model whose input and output data involve multiple modal data such as images and text. After fine-tuning and training, it can be applied to multimodal task processing.
[0042] Pre-trained language model: A pre-trained model obtained by pre-training a large-scale language model (LLM).
[0043] Large models refer to deep learning models with large-scale model parameters, typically containing hundreds of millions, tens of billions, or even hundreds of billions of model parameters. Large models, also known as foundation models (FMs), are pre-trained on large amounts of unlabeled corpora, producing pre-trained models with parameters exceeding 100 million. These models are adaptable to a wide range of downstream tasks and exhibit good generalization capabilities. Examples include large language models (LLMs) and multi-modal pre-training models.
[0044] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0045] When applied to human-computer interaction scenarios (such as intelligent robots), the large model generates responses based on user input commands, also known as the human-computer interaction model. During the iteration process of the human-computer interaction model, it is necessary to evaluate the advantages and disadvantages of different versions of the human-computer interaction model to achieve iterative updates of the human-computer interaction model. Before the human-computer interaction model is launched, it is necessary to evaluate whether the performance of the human-computer interaction model meets the launch requirements, and to launch human-computer interaction models with excellent performance and avoid launching human-computer interaction models with poor performance. Multi-round dialogue capability is an important and valuable capability of current human-computer interaction models. How to objectively evaluate the multi-round dialogue capability of these human-computer interaction models is very important.
[0046] Current methods for evaluating multi-turn conversational capabilities use pre-designed multi-turn conversational data. At the very least, the questions fed into the human-computer interaction model in each round are fixed. Regardless of whether the human-computer interaction model responded in the previous round or what the response was, the questions fed into the model in the next round remain unchanged. For example, the first round question might be, "In quantum physics, what is superposition, and how does it relate to quantum entanglement?" The second round question might be, "What assumptions did you make in your answer? Are they valid?" If the answer to the first round question is, "I'm sorry, I don't know," then the second round question would seem incomprehensible and fail to accurately simulate the natural interaction process of real people. Therefore, current methods for evaluating multi-turn conversational capabilities cannot truly and accurately assess the multi-turn conversational capabilities of human-computer interaction models. This hinders the selection of high-quality models with excellent multi-turn conversational capabilities during model iteration and the quality of multi-turn conversations performed by the production model, resulting in poor human-computer interaction quality.
[0047] The present disclosure provides a data processing method for human-computer interaction, which comprises obtaining a pre-constructed conversation topic and a first human-computer interaction model to be evaluated; using a second human-computer interaction model to conduct multiple rounds of dialogue with the first human-computer interaction model based on the conversation topic, thereby obtaining multi-round dialogue information on the conversation topic, wherein the dialogue information includes: question information input to the first human-computer interaction model and response information to the question information generated by the first human-computer interaction model; using the second human-computer interaction model to simulate the process of a human user conducting multiple rounds of dialogue with the first human-computer interaction model around the conversation topic, thereby generating multi-round dialogue information on the conversation topic. Furthermore, based on the multi-round dialogue information on the conversation topic, evaluation information of the multi-round dialogue capability of the first human-computer interaction model is determined, which can truly implement the evaluation of the multi-round dialogue capability of the first human-computer interaction model and obtain accurate and high-quality evaluation information. The evaluation information is used to guide the online decision of the first human-computer interaction model or update the optimized version of the first human-computer interaction model. It can accurately select a high-quality model in the iteration of the first human-computer interaction model, improve the multi-round dialogue quality of the iteratively updated first human-computer interaction model, improve the multi-round dialogue quality of the online model, and thus improve the quality of multi-round dialogue in human-computer interaction.
[0048] The first human-computer interaction model refers to the human-computer interaction model to be evaluated, and may be a large-scale pre-trained language model (LLM), a multimodal pre-trained model, or the like, and is not specifically limited here. The second human-computer interaction model simulates a human posing questions to the first human-computer interaction model to be evaluated, and may be any existing human-computer interaction model with relatively high performance, such as a relatively mature existing language model or a pre-trained model, and is not specifically limited here in this embodiment.
[0049] FIG1 is a schematic diagram of an example system architecture applicable to the present disclosure. As shown in FIG1 , the system architecture includes a first server responsible for evaluating the multi-round interaction capability of the first human-computer interaction model, a second server running the first human-computer interaction model, and an end-side device. A communication link is provided between the first server and the second server, enabling communication connection between the first server and the second server. A communication link is provided between the first server and the end-side device, enabling communication connection between the first server and the end-side device.
[0050] The second server can be a server cluster deployed in the cloud or a local device with computing power. The second server is responsible for running the human-computer interaction model to be evaluated (the first human-computer interaction model) and generating response information based on the given question information. One or more human-computer interaction models to be evaluated can be deployed on a second server. If multiple human-computer interaction models to be evaluated are to be deployed on one or more second servers.
[0051] A terminal device is an electronic device used by a user. Specifically, it may be a hardware device with network communication, computing, and information display functions, including but not limited to smartphones, tablet computers, desktop computers, and servers. The user sends a request for evaluating a first human-computer interaction model to the first server via the terminal device. The request includes information about one or more first human-computer interaction models to be evaluated.
[0052] The first server can be a server cluster deployed in the cloud, or a local device with computing power. The first server runs a second human-computer interaction model, which simulates a human user through the second human-computer interaction model. Based on the response information provided by the first human-computer interaction model, the first human-computer interaction model is asked the next round of questions, thereby simulating a multi-round dialogue between a human and the first human-computer interaction model, and generating multi-round dialogue information based on the conversation topic. The first server is also responsible for evaluating the multi-round dialogue capability of the first human-computer interaction model based on the multi-round dialogue information of the conversation topic, generating evaluation information of the multi-round dialogue capability of the first human-computer interaction model, and guiding the online decision of the first human-computer interaction model, updating the optimized version of the first human-computer interaction model, or selecting a high-quality first human-computer interaction model with strong multi-round dialogue capability.
[0053] In one example scenario, before a first human-computer interaction model for implementing human-computer interaction goes online, a user sends a request for evaluating the multi-round interaction capabilities of the first human-computer interaction model to be launched to a first server via a client-side device. The request includes relevant information about the first human-computer interaction model to be evaluated, such as an application program interface for calling the first human-computer interaction model and an access address for the first human-computer interaction model. In response to the evaluation request, the first server obtains a pre-built conversation topic, uses a second human-computer interaction model to conduct a multi-round conversation with the first human-computer interaction model based on the conversation topic, and obtains multi-round conversation information about the conversation topic. Based on the multi-round conversation information about the conversation topic, the first server determines evaluation information about the multi-round conversation capabilities of the first human-computer interaction model, thereby accurately evaluating the multi-round conversation capabilities of the human-computer interaction model.
[0054] Furthermore, the evaluation information of the multi-round dialogue capability of the first human-computer interaction model can be used to guide the online decision of the first human-computer interaction model. Optionally, the terminal-side device can also output the evaluation information of the multi-round dialogue capability of the first human-computer interaction model to guide relevant technical personnel to make an online decision on the first human-computer interaction model. Optionally, the first server determines whether the first human-computer interaction model meets the online conditions based on the evaluation information of the multi-round dialogue capability of the first human-computer interaction model, and outputs the online prompt information of the first human-computer interaction model, and the online prompt information indicates whether the first human-computer interaction model meets the online conditions. Optionally, the first server sends the evaluation information of the multi-round dialogue capability of the first human-computer interaction model to the terminal-side device. The terminal-side device outputs evaluation information of the multi-round dialogue capability of the first human-computer interaction model to guide the user to determine whether the first human-computer interaction model meets the online conditions; or, the terminal-side device determines whether the first human-computer interaction model meets the online conditions based on the evaluation information of the multi-round dialogue capability of the first human-computer interaction model, and outputs online prompt information of the first human-computer interaction model, and the online prompt information indicates whether the first human-computer interaction model meets the online conditions.
[0055] In another example scenario, during the iterative optimization process of the first human-computer interaction model, the obtained new version is evaluated. The user can send an evaluation request for the multi-round dialogue capability of the new version of the first human-computer interaction model to the first server through the terminal device. The evaluation request contains relevant information of the new version of the first human-computer interaction model, such as the application interface for calling the first human-computer interaction model, the access address of the first human-computer interaction model, etc. In response to the evaluation request, the first server obtains a pre-built conversation topic, uses the second human-computer interaction model to conduct a multi-round dialogue with the first human-computer interaction model based on the conversation topic, and obtains multi-round dialogue information based on the conversation topic; according to the multi-round dialogue information of the conversation topic, the evaluation information of the multi-round dialogue capability of the first human-computer interaction model is determined to achieve accurate evaluation of the multi-round dialogue capability of the new version of the first human-computer interaction model.
[0056] Furthermore, the evaluation information of the multi-round dialogue capability of the new version of the first human-computer interaction model can be used to guide the update of the optimized version of the first human-computer interaction model. Optionally, the first server compares the evaluation information of the multi-round dialogue capability of the new version and the previous version based on the evaluation information of the multi-round dialogue capability of the new version of the first human-computer interaction model and the evaluation information of the multi-round dialogue capability of the previous version to obtain a comparison result, and the comparison result is used to guide the update of the optimized version of the first human-computer interaction model. Specifically, the first server can send the comparison result to the end-side device. The end-side device outputs the comparison result of the evaluation information of the multi-round dialogue capability of different versions of the first human-computer interaction model to guide the user to select the optimized version with stronger multi-round dialogue capability for iterative update of the first human-computer interaction model.
[0057] In another example scenario, based on the evaluation information of the multi-round dialogue capabilities of multiple first human-computer interaction models to be selected, the user can select a first human-computer interaction model with stronger multi-round dialogue capabilities for human-computer interaction to improve the quality of human-computer interaction. The user can send a request for evaluating the multi-round dialogue capabilities of multiple first human-computer interaction models to the first server through the terminal device. The evaluation request contains relevant information of the multiple first human-computer interaction models, such as calling the application program interface of each first human-computer interaction model, the access address of each first human-computer interaction model, etc. In response to the evaluation request, the first server evaluates each first human-computer interaction model in the following manner: obtaining a pre-constructed conversation topic, using the second human-computer interaction model to conduct a multi-round dialogue with the first human-computer interaction model based on the conversation topic, and obtaining multi-round dialogue information based on the conversation topic; based on the multi-round dialogue information of the conversation topic, determining the evaluation information of the multi-round dialogue capability of the first human-computer interaction model, so as to fairly and accurately evaluate the multi-round dialogue capability of each first human-computer interaction model.
[0058] Furthermore, the first server compares the multi-round dialogue capability evaluation information of each first human-computer interaction model to obtain a comparison result of the multi-round dialogue capability evaluation information of each first human-computer interaction model. The first server sends the comparison result of the multi-round dialogue capability evaluation information of each first human-computer interaction model to the end-side device. The end-side device outputs the comparison result to guide the user to select the first human-computer interaction model with stronger multi-round dialogue capability as the first human-computer interaction model of their choice. Optionally, the end-side device can select the first human-computer interaction model with stronger multi-round dialogue capability based on the comparison result of the multi-round dialogue capability evaluation information of each first human-computer interaction model, and download and obtain the first human-computer interaction model based on the relevant information of the selected first human-computer interaction model, or use the first human-computer interaction model to implement human-computer interaction.
[0059] It should be noted that the first human-computer interaction model to be evaluated can be provided to the first server by a third party through an end-side device. The first server obtains the model to be evaluated provided by the end-side device. Exemplarily, the third party can upload the first human-computer interaction model to be evaluated to the first server through the end-side device, and the first server can deploy the first human-computer interaction model to be evaluated to the second server or the first server. Exemplarily, the third party can send a download link of the first human-computer interaction model to be evaluated to the first server through the end-side device, and the first server obtains the first human-computer interaction model through the download link, and deploys the first human-computer interaction model to the second server or the first server. Exemplarily, the third party can send the application programming interface (API) or access address of the first human-computer interaction model to be evaluated to the first server through the end-side device. The first server inputs information into the first human-computer interaction model to be evaluated through the application programming interface (API) or access address, and receives response information output by the first human-computer interaction model.
[0060] The following detailed description of the technical solution of the present disclosure and how the technical solution of the present disclosure solves the above-mentioned technical problems is provided with specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present disclosure will be described below in conjunction with the accompanying drawings.
[0061] FIG2 is a flow chart of a data processing method for human-computer interaction provided by an exemplary embodiment of the present disclosure. The execution subject of this embodiment is the first server in the aforementioned system architecture. As shown in FIG2, the specific steps of the method are as follows:
[0062] Step S201: Obtain a pre-built conversation topic and a first human-computer interaction model to be evaluated.
[0063] Among them, the conversation topic refers to the topic-related information of a conversation, including the following topic content-related information such as the keyword, topic description, etc. of at least one conversation, the opening words of the conversation (that is, the first question asked to the first human-computer interaction model to be evaluated), etc., the purpose and requirements of the conversation, and other conversation task information. This embodiment does not make specific limitations here.
[0064] In this embodiment, conversation topics of various conversation types and scenarios are pre-constructed to simulate various conversation types and scenarios and generate multi-round dialogue information of various conversation topics.
[0065] Different conversation types include, but are not limited to, small talk and task-based conversations. Small talk conversations involve users engaging in an unrestricted conversation without a specific purpose, such as chatting or playing games. Task-based conversations involve users completing a specific task, which ends upon completion. These include math problems, coding exercises, writing assignments, and knowledge quizzes. Different conversation scenarios include, but are not limited to, AI assistants, intelligent customer service, intelligent question-and-answer systems, chatbots, and intelligent educational robots.
[0066] The first human-computer interaction model to be evaluated can specifically be a model for implementing human-computer interaction, and can specifically be a large-scale pre-trained language model (LLM), a multimodal pre-trained model, etc., and is not specifically limited here. The first human-computer interaction model can be applied to various human-computer interaction scenarios, such as artificial intelligence assistants, intelligent customer service, intelligent question-and-answer systems, chatbots, intelligent educational robots, etc. The method of this embodiment can be applied to the evaluation of various human-computer interaction systems, and this embodiment does not specifically limit the first human-computer interaction model.
[0067] Step S202: Use the second human-computer interaction model to conduct multiple rounds of dialogue with the first human-computer interaction model based on the conversation topic to obtain multiple rounds of dialogue information on the conversation topic, where the dialogue information includes: question information input to the first human-computer interaction model, and response information to the question information generated by the first human-computer interaction model.
[0068] Among them, the second human-computer interaction model is a model that simulates humans asking questions to the first human-computer interaction model to be evaluated. It can be any existing human-computer interaction model with better performance, such as a more mature language model, a pre-trained model, etc. This embodiment does not make specific limitations here.
[0069] In this embodiment, a second human-computer interaction model is introduced and used to conduct multiple rounds of dialogue with the first human-computer interaction model based on the conversation topic. The second human-computer interaction model simulates a human being to dynamically ask questions to the first human-computer interaction model to be evaluated, thereby simulating multiple rounds of dialogue between humans and the first human-computer interaction model to be evaluated around the conversation topic, and obtaining multiple rounds of dialogue information on the conversation topic.
[0070] Among the multiple rounds of dialogue information for a conversation topic, one round of dialogue information includes: question information input to a first human-computer interaction model, and response information to the question information generated by the first human-computer interaction model. The question information input to the first human-computer interaction model is the opening statement in a given conversation topic, or question information generated by a second human-computer interaction model.
[0071] It should be noted that in this embodiment, for the same conversation topic, step S202 can be executed once to simulate the process of multiple rounds of dialogue in a conversation based on the conversation topic. Step S202 can be executed multiple times for the same conversation topic to simulate the process of multiple different conversations based on the same conversation topic, generating multiple rounds of dialogue information in each conversation. The multi-round dialogue information in different conversations based on the same conversation topic is not exactly the same. For example, for different response information of the first human-computer interaction model, the second human-computer interaction model generates different next-round question information. Even for the same response information of the first human-computer interaction model, the second human-computer interaction model will generate different question information, thereby simulating a real human-computer dialogue scenario.
[0072] Step S203: Determine evaluation information of the multi-round dialogue capability of the first human-computer interaction model based on the multi-round dialogue information of the conversation topic.
[0073] After obtaining multi-round dialogue information based on each conversation topic, the response quality of the first human-computer interaction model in the multi-round dialogue based on each conversation topic is evaluated according to the multi-round dialogue information of the conversation topic, so as to realize the evaluation of the multi-round dialogue capability of the first human-computer interaction model and obtain the evaluation information of the multi-round dialogue capability of the first human-computer interaction model.
[0074] In this embodiment, after obtaining the evaluation information of the multi-round dialogue capability of the first human-computer interaction model, the first server can output the evaluation information of the multi-round dialogue capability of the first human-computer interaction model. By visually outputting the evaluation information of the multi-round dialogue capability of the first human-computer interaction model and outputting the evaluation results of the multi-round dialogue capability of the first human-computer interaction model to the user, the evaluation information of the multi-round dialogue capability of the first human-computer interaction model can guide the user to make a judgment on whether the first human-computer interaction model is online; or, by comparing the evaluation information of the multi-round dialogue capability of multiple versions of the first human-computer interaction model, determine the high-quality version of the first human-computer interaction model and perform iterative optimization of the first human-computer interaction model; or, by comparing the evaluation information of the multi-round dialogue capability of multiple first human-computer interaction models, select the first human-computer interaction model with stronger multi-round dialogue capability as the target human-computer interaction model for implementing human-computer interaction.
[0075] For example, after obtaining the evaluation information of the multi-round dialogue capability of the first human-computer interaction model, the first server may determine whether the first human-computer interaction model meets the online conditions based on the evaluation information of the multi-round dialogue capability of the first human-computer interaction model, and output online prompt information of the first human-computer interaction model and / or the evaluation information of the multi-round dialogue capability of the first human-computer interaction model. The online prompt information indicates whether the first human-computer interaction model meets the online conditions.
[0076] For example, the first server may select one of the first human-computer interaction models as the target human-computer interaction model based on the evaluation information of the multi-round dialogue capabilities of multiple first human-computer interaction models, and output the information of the target human-computer interaction model to the client-side device. For example, the first server may select one of the multiple different versions of the first human-computer interaction model as the optimized version based on the evaluation information of the multi-round dialogue capabilities of the first human-computer interaction model, and update the optimized version of the first human-computer interaction model.
[0077] The launch condition includes a first threshold for the evaluation information of the multi-round dialogue capability of the first human-computer interaction model. If the evaluation information of the multi-round dialogue capability of the first human-computer interaction model is greater than or equal to the first threshold, the first human-computer interaction model meets the launch condition; otherwise, the first human-computer interaction model does not meet the launch condition. The first threshold in the launch condition can be customized by the user according to the needs of the specific application scenario.
[0078] The solution of this embodiment obtains a pre-built conversation topic and a first human-computer interaction model to be evaluated; uses a second human-computer interaction model to conduct multiple rounds of dialogue with the first human-computer interaction model based on the conversation topic, thereby obtaining multi-round dialogue information on the conversation topic, the dialogue information including: question information input to the first human-computer interaction model and response information to the question information generated by the first human-computer interaction model; uses the second human-computer interaction model to simulate the process of a human user conducting multiple rounds of dialogue with the first human-computer interaction model around the conversation topic, thereby generating multi-round dialogue information on the conversation topic. Furthermore, based on the multi-round dialogue information on the conversation topic, evaluation information of the multi-round dialogue capability of the first human-computer interaction model is determined, which can truly implement the evaluation of the multi-round dialogue capability of the first human-computer interaction model and obtain accurate and high-quality evaluation information. The evaluation information is used to guide the online decision of the first human-computer interaction model or update the optimized version of the first human-computer interaction model. It can accurately select high-quality models in the iteration of the first human-computer interaction model, improve the multi-round dialogue quality of the iteratively updated first human-computer interaction model, improve the multi-round dialogue quality of the online model, and thus improve the quality of multi-round dialogue in human-computer interaction.
[0079] In an optional embodiment, in the aforementioned step S202, the second human-computer interaction model is used to conduct multiple rounds of dialogue with the first human-computer interaction model based on the conversation topic to obtain multiple rounds of dialogue information on the conversation topic. Specifically, this can be implemented in the following manner: the conversation topic includes an opening statement, and the opening statement is used as the question information of the first round, which is input into the first human-computer interaction model, and response information of the question information is generated by the first human-computer interaction model to obtain the first round of dialogue information on the conversation topic; the second human-computer interaction model is used to generate the question information of the next round based on the dialogue information of each round of the conversation topic that has been obtained, and the question information of the next round is input into the first human-computer interaction model, and the response information of the next round is generated by the first human-computer interaction model to obtain the next round of dialogue information, and so on, until the conversation based on the conversation topic ends, and multiple rounds of dialogue information on the conversation topic are obtained.
[0080] FIG3 is a flowchart of obtaining multi-round dialogue information based on a conversation topic according to this embodiment. As shown in FIG3 , in step S202 , the second human-computer interaction model is used to conduct multi-round dialogues with the first human-computer interaction model based on the conversation topic to obtain multi-round dialogue information on the conversation topic. Specifically, the following steps can be used to implement the multi-round dialogue information:
[0081] Step S31: The conversation topic includes an opening statement. The opening statement is used as the first round of question information and input into a first human-computer interaction model. The first human-computer interaction model generates response information to the question information to obtain the first round of dialogue information of the conversation topic.
[0082] In this embodiment, the pre-constructed conversation topic includes an opening statement, and the opening statement serves as the first round of question information in this conversation, that is, the first round of question information in the conversation is pre-constructed.
[0083] In this step, the opening words in the pre-constructed conversation topic are used as the first round of question information and input into the first human-computer interaction model to simulate a human user initiating a conversation to the first human-computer interaction model, so that the first human-computer interaction model generates response information to the input question information. In this way, the first round of dialogue information of a conversation on the conversation topic can be obtained, including the first round of question information and response information.
[0084] In another optional embodiment, the pre-constructed conversation topic may not include an opening statement. The conversation topic includes information related to the topic content, such as the conversation's keyword and topic description, or includes conversation task information, such as the purpose and requirements of the conversation. In this case, the second human-computer interaction model is used to generate the opening statement based on the conversation topic, that is, the first round of question information is generated by the second human-computer interaction model. Furthermore, the first round of question information generated by the second human-computer interaction model is input into the first human-computer interaction model to simulate a human user initiating a conversation with the first human-computer interaction model, causing the first human-computer interaction model to generate response information to the question information, thereby obtaining the first round of dialogue information for the conversation topic.
[0085] Step S32: Use the second human-computer interaction model to generate the next round of question information based on the obtained dialogue information of each round of the conversation topic, input the next round of question information into the first human-computer interaction model, generate the next round of response information through the first human-computer interaction model, and obtain the next round of dialogue information.
[0086] After the first human-computer interaction model generates responses to the first round, the second human-computer interaction model generates questions for the next round based on the first round of dialogue. These questions are then fed into the first human-computer interaction model, simulating a human asking the first human-computer interaction model questions based on the first round of dialogue. The first human-computer interaction model then generates responses to the next round of questions, which are the second round of responses, resulting in the second round of dialogue.
[0087] Similar to the process of generating the second round of dialogue information, the second human-computer interaction model generates the next (third) round of questions based on the first two rounds of dialogue information. These next round of questions are then fed into the first human-computer interaction model, simulating the process of a human asking the first human-computer interaction model a third round of questions based on the first two rounds of dialogue information. The first human-computer interaction model then generates responses to the next round of questions, which are the responses to the third round, thus generating the third round of dialogue information.
[0088] By analogy, a third or even more rounds of dialogue information based on the conversation topic can be generated until the current conversation ends, and multiple rounds of dialogue information of a conversation based on the conversation topic are obtained.
[0089] In this embodiment, task prompts corresponding to each conversation topic are pre-built. The task prompts include the conversation topic. The task prompts prompt the second human-computer interaction model to simulate the interaction between a human and the first human-computer interaction model based on the conversation topic and the conversation task prompts, and to generate the next round of question information. Different conversation topics can use different task prompts.
[0090] For example, a task prompt message is: "I am not an AI assistant, but a user who is currently willing to chat. I am chatting with an AI assistant. My opening is ```{task}```. I will continue the conversation based on the chat history and the topic I started with, and will not deviate from the topic. When the AI assistant responds to me during a round of conversation and I can no longer continue the conversation, I will directly output<CHAT_TASK_DONE> , don't force a conversation, and don't treat yourself as an AI assistant. "The task prompt information corresponds to the conversation topic as the opening statement, and the {task} in the task prompt information refers to the opening statement."<CHAT_TASK_DONE> " is the pre-set conversation end information. When the next round of question information output by the second human-computer interaction model is<CHAT_TASK_DONE> When , it means that the second human-computer interaction model can no longer chat and ends the current session.
[0091] For example, a task prompt might read: "I'm not an AI assistant, but a user who wants to play games for fun, playing with an AI assistant. I begin by stating my purpose and requirements: `{task}```. I will carefully review the other party's response based on the purpose and requirements I've stated above. If the other party isn't playing the game or is breaking the rules, I will point out the problem and ask them to correct it. Otherwise, I will actively continue playing." The task prompt corresponds to the conversation topic as an opening statement, or it can include conversation task information such as the purpose and requirements of the conversation. In the task prompt, {task} refers to the conversation topic, and `\n` is a line break.
[0092] In this step, the task prompt information corresponding to the conversation topic and the responses from the previous round are input into the second human-computer interaction model. Based on the task prompt information and responses from the previous round, as well as the conversation context, the second human-computer interaction model generates the next round of questions and updates the conversation context. In this way, guided by the conversation topic and task prompt information, the second human-computer interaction model can engage in multiple rounds of dialogue with the first human-computer interaction model around the conversation topic, based on the requirements of the conversation topic and task prompt information, thereby improving the quality of the questions generated in each round.
[0093] In an optional embodiment, the task prompt information corresponding to the conversation topic and at least one previous round of dialogue information of the conversation topic can also be input into the second human-computer interaction model. The second human-computer interaction model generates the next round of question information based on the task prompt information of the conversation topic and at least one previous round of dialogue information, as well as the context information of the conversation, and updates the context information of the conversation.
[0094] In practical applications, during a multi-round conversation, both the first and second human-computer interaction models will locally store the conversation context information. This context information includes information about each round of conversation that has occurred in the current conversation. Specifically, it can include information about all rounds of conversation that have occurred, or only the most recent rounds. The number of rounds of conversation information retained in the context information can be set by setting a sliding window or other methods, and is not specifically limited here. The recording of conversation context information by each human-computer interaction model is a fundamental capability for implementing multi-round conversations and will not be further elaborated here.
[0095] In an optional embodiment, a topic guidance script for the conversation topic can be pre-built. The topic guidance script includes the main points of the conversation topic, user requirements, conversation objectives, and precautions, and is used to guide the second human-computer interaction model to raise new questions around the conversation topic. Different topic guidance scripts can be set for different conversation topics, and of course, no topic guidance script for the conversation topic is required.
[0096] In this step, the second human-computer interaction model generates the next round of questions based on the acquired conversation information for each round of the conversation topic and the topic guidance script for the conversation topic. By using the guidance of the topic guidance script, the second human-computer interaction model can engage in multiple rounds of dialogue with the first human-computer interaction model around the conversation topic, based on the key points of the conversation topic, user requirements, conversation objectives, and precautions in the topic guidance script. This can further improve the quality of the generated question information for each round.
[0097] Optionally, for task-based conversation topics, a topic guidance script corresponding to the conversation topic needs to be constructed to encourage the second human-computer interaction model, under the guidance of the topic guidance script, to generate the next round of question information centered around the purpose of the conversation task without deviating from the conversation topic. For chat-type conversation topics, a topic guidance script corresponding to the conversation topic may not be constructed. Of course, for some chat-type conversation topics, a topic guidance script can also be constructed to encourage the second human-computer interaction model, under the guidance of the topic guidance script, to generate the next round of question information centered around the topic guidance script, to avoid the conversation content from diverging too much and deviating from the conversation topic.
[0098] For example, Table 1 provides examples of topic guidance scripts for several chat-type conversation topics and task-based conversation topics:
[0099] Table 1
[0100] For a conversation topic for which a topic guiding script is constructed, the task prompt information of the conversation topic includes the conversation topic and the topic guiding script.
[0101] For example, an example of a task prompt message for a conversation topic with a topic-guided script is as follows: "I am an expert in reasoning. I am now going to evaluate the reasoning ability of an AI assistant by asking questions. The question I gave is: \n```{task}```. \nI have obtained the answer based on the prompts: ```{tips}```. I will use these prompts to determine whether the assistant has completed the task. If not, I will gradually give some prompts to guide the other party to complete the task, but will never give the answer. If the other party's latest round of responses has allowed me to achieve my goal, I will directly output<CHAT_TASK_DONE> ", where {task} refers to the conversation topic, which is the question that needs to be answered. {tips} refers to the topic guidance script of the conversation topic, which can be the answer to the question or the reasoning process to obtain the answer. "\n" is a line break character.<CHAT_TASK_DONE> " is the pre-set conversation end information. When the next round of question information output by the second human-computer interaction model is<CHAT_TASK_DONE> When , it means that the second human-computer interaction model can no longer chat and ends the current session.
[0102] For example, an example of a task prompt message for a conversation topic with a topic-guiding script is as follows: "I am an expert in science, technology, engineering, and mathematics. I am now going to evaluate the relevant capabilities of an AI assistant by asking questions. The topic I gave is: \n```{task}```. \nI have obtained the key points to pay attention to based on the prompts: ```{tips}```. I will use these prompts to determine whether the assistant has completed the task. If not, I will gradually give some prompts to guide the other party to complete the task, but I will never give an answer. If the other party's latest round of responses has allowed me to achieve my goal, I will directly output<CHAT_TASK_DONE> ", where {task} refers to the conversation topic, which is the question that needs to be answered. {tips} refers to the topic guide script of the conversation topic. "\n" is a line break character.<CHAT_TASK_DONE> " is the pre-set conversation end information. When the next round of question information output by the second human-computer interaction model is<CHAT_TASK_DONE> When , it means that the second human-computer interaction model can no longer chat and ends the current session.
[0103] For a conversation topic with a built-in topic guide script, this step inputs the topic guide script, task prompts, and responses from the previous round into the second human-computer interaction model. Based on the topic guide script, task prompts, responses from the previous round, and the context of the conversation, the second human-computer interaction model generates the next round of questions and updates the context of the conversation. In this way, guided by the conversation topic, topic guide script, and task prompts, the second human-computer interaction model can engage in multiple rounds of dialogue with the first human-computer interaction model around the conversation topic, based on the requirements of the conversation topic, topic guide script, and task prompts, improving the quality of the questions generated in each round.
[0104] Step S33: Determine whether the conversation based on the conversation topic is ended.
[0105] Specifically, when the next round of question information output by the second human-computer interaction model is session end information, the current session is ended; or when the number of interaction rounds in the current session is greater than or equal to a preset maximum number of rounds, the current session is ended.
[0106] The session end information is pre-configured specific information, such as<CHAT_TASK_DONE> , can be configured according to the needs of the actual application scenario, and is not specifically limited here.
[0107] The preset maximum number of rounds refers to the maximum number of dialogue rounds contained in a conversation based on a conversation topic, and can be configured according to the needs of the actual application scenario. For example, the preset maximum number of rounds can be set to 5, 10, etc., and is not specifically limited here.
[0108] In this step, if it is determined that the conversation based on the conversation topic has not ended, step S32 is continuously executed in a loop until the conversation ends.
[0109] Step S34: If the process ends, multiple rounds of conversation information on the conversation topic are obtained.
[0110] After a conversation based on the conversation topic ends, multiple rounds of dialogue information of the conversation topic can be obtained.
[0111] For example, taking the maximum number of dialogue rounds as 5, the second human-computer interaction model is used to simulate a human user and conduct 5 dialogue rounds with the first human-computer interaction model. The obtained 5-round dialogue information is shown in Table 2:
[0112] Table 2
[0113] The method of this embodiment pre-constructs a conversation opening statement as the conversation topic. The opening statement is used as the first round of question information and is input into a first human-computer interaction model. The first human-computer interaction model then generates response information to the question information, thereby obtaining the first round of dialogue information on the conversation topic. A second human-computer interaction model is then used to generate the next round of question information based on the previously obtained dialogue information on the conversation topic. The next round of question information is then input into the first human-computer interaction model. This simulates the process of a human user initiating a new round of questions to the first human-computer interaction model based on each round of historical dialogue information. This process continues until the conversation based on the conversation topic ends, obtaining multiple rounds of dialogue information on the conversation topic. Based on the constructed conversation topic, multiple rounds of dialogue information based on the conversation topic can be automatically generated, achieving a true multi-round dialogue. This simulates the process of a human user and the first human-computer interaction model engaging in multiple rounds of dialogue around the conversation topic. This generates true multi-round dialogue information, providing high-quality data for evaluating the multi-round dialogue capability of the first human-computer interaction model, thereby improving the evaluation quality of the multi-round dialogue capability of the first human-computer interaction model.
[0114] In practical applications, the multi-round dialogue capabilities of human-computer interaction models fall into at least the following categories: context-sensitive and context-independent. Context-sensitive means that information from previous rounds of a multi-round dialogue may be helpful (i.e., have a positive impact) in generating the model's response for the next round. Context-independent means that information from previous rounds of a multi-round dialogue may have a negative impact on the model's response for the next round.
[0115] For example, Table 3 provides examples of multi-turn dialogue information covering different types of multi-turn dialogue capabilities:
[0116] Table 3
[0117] Based on any of the aforementioned method embodiments, a second human-computer interaction model is used to conduct multiple rounds of dialogue with the first human-computer interaction model based on the conversation topic. In the multiple rounds of dialogue information obtained on the conversation topic, the question information of the next round of dialogue information is generated based on the conversation topic and the historical rounds of dialogue information, which affects the response information of the next round. Therefore, these multiple rounds of dialogues based on the conversation topic are context-related.
[0118] In the aforementioned step S203, based on the multi-round dialogue information of the conversation topic, the evaluation information of the multi-round dialogue capability of the first human-computer interaction model is determined to implement the evaluation of the context-related multi-round dialogue capability of the first human-computer interaction model.
[0119] Exemplarily, based on the multi-round dialogue information of the conversation topic, the context-related multi-round dialogue capability of the first human-computer interaction model can be evaluated from at least one of the three dimensions of micro-dialogue quality, macro-dialogue quality, and dialogue round quality to obtain evaluation information in multiple dimensions.
[0120] Among them, the micro-dialogue quality is the response quality of the response information given by the first human-computer interaction model in each round of dialogue information, and is evaluated from the "micro" perspective of a single round of dialogue.
[0121] Macro conversation quality refers to the overall response quality of the first human-computer interaction model in multiple rounds of conversation based on the conversation topic, and is evaluated from a "macro" perspective of a conversation based on the same conversation topic.
[0122] Conversation turn quality refers to the quality of the conversation length of the first human-computer interaction model within a conversation based on the conversation topic. Conversation types corresponding to conversation topics include, but are not limited to, small talk and task-based. For small talk conversation topics, a higher number of conversation turns is preferred. For task-based conversation topics, a lower number of conversation turns is preferred to ensure prompt task results. Small talk can be further subdivided into categories such as chatting and gaming, while task-based conversations can be further subdivided into categories such as math problem solving and writing. The division of conversation types corresponding to conversation topics can be configured and adjusted based on the needs of actual application scenarios and is not specifically limited here.
[0123] Specifically, from the perspective of micro-dialogue quality, the context-related multi-round dialogue capability of the first human-computer interaction model can be evaluated in the following ways:
[0124] Based on the multi-round dialogue information of the conversation topic, the response quality information of the first human-computer interaction model in each round of dialogue information is determined; the response quality information of the first human-computer interaction model in each round of dialogue information of each conversation topic is comprehensively analyzed to determine the response quality evaluation information of the first human-computer interaction model in the context-related multi-round dialogues.
[0125] For example, in the five rounds of dialogue information of the conversation topic given in Table 2, the response information given by the first human-computer interaction model is used as the five response information to be evaluated. The response information in each round is evaluated separately to obtain the response quality information of the first human-computer interaction model in each round of dialogue information of the conversation topic; the response quality information of the first human-computer interaction model in each round of dialogue information of each conversation topic is combined to calculate the response quality evaluation information of the first human-computer interaction model in the context-related multi-round dialogue.
[0126] Optionally, the average of the response quality information of the first human-computer interaction model in each round of dialogue information of each conversation topic may be calculated as the response quality evaluation information of the first human-computer interaction model in the context-related multi-round dialogues.
[0127] Optionally, the average of the response quality information of the first human-computer interaction model in each round of dialogue information for each conversation topic can be used as the average quality value corresponding to the conversation topic. The quality averages of each conversation topic are weighted averaged according to a first preset weight coefficient for each conversation topic to obtain response quality evaluation information of the first human-computer interaction model in multiple rounds of dialogue related to the context. The first preset weight coefficient for each conversation topic can be configured according to the needs of the actual application scenario and is not specifically limited here.
[0128] Specifically, from the perspective of macro-dialogue quality, the context-related multi-round dialogue capability of the first human-computer interaction model can be evaluated in the following ways:
[0129] Based on the multi-round dialogue information of the conversation topic, the overall response quality information of the first human-computer interaction model in the multi-round dialogue of the conversation topic is determined; the overall response quality information of the first human-computer interaction model in the multi-round dialogue of each conversation topic is comprehensively analyzed to determine the overall response quality evaluation information of the first human-computer interaction model in the context-related multi-round dialogue.
[0130] For example, in the five rounds of dialogue information of the conversation topic given in Table 2, the response information given by the first human-computer interaction model is used as the response information to be evaluated. From the perspective of the overall response quality in this conversation, the overall response quality of the first human-computer interaction model in this conversation is evaluated, and the overall response quality information of the first human-computer interaction model in the multiple rounds of dialogue on the conversation topic is obtained; the overall response quality information of the first human-computer interaction model in the multiple rounds of dialogue on each conversation topic is integrated to calculate the overall response quality evaluation information of the first human-computer interaction model in the context-related multiple rounds of dialogue.
[0131] Optionally, the average of the overall response quality information of the first human-computer interaction model in multiple rounds of dialogue for each conversation topic may be calculated as the overall response quality evaluation information of the first human-computer interaction model in the context-related multiple rounds of dialogue.
[0132] Optionally, based on a second preset weight coefficient for each conversation topic, the overall response quality information of the first human-computer interaction model in multiple rounds of dialogue for each conversation topic can be weighted averaged to obtain overall response quality evaluation information of the first human-computer interaction model in multiple rounds of dialogue related to the context. The second preset weight coefficient for each conversation topic can be configured according to the needs of the actual application scenario and is not specifically limited here.
[0133] Specifically, the context-sensitive multi-turn dialogue capability of the first human-computer interaction model is evaluated from the perspective of dialogue turn quality. This can be achieved in the following ways:
[0134] According to the number of conversation turns of the conversation topic and the conversation type corresponding to the conversation topic, conversation turn evaluation information of the first human-computer interaction model in the context-related multi-round conversation is determined.
[0135] The conversation types corresponding to the conversation topics include, but are not limited to, small talk and task-based. For small talk, a higher number of conversation turns is preferred. For task-based conversations, a lower number of conversation turns is preferred to ensure faster task results. Different conversation turn evaluation rules are configured for different conversation types. Based on the conversation turn evaluation rules corresponding to each conversation topic, the number of conversation turns for each topic is quantified into conversation turn evaluation information.
[0136] For example, for any conversation type, multiple different evaluation scores can be assigned to the conversation turn evaluation information, with different evaluation scores corresponding to different conversation turn intervals. Based on the conversation turn count and conversation type of the conversation topic, the conversation turn interval to which the conversation turn belongs is determined, and the evaluation score corresponding to the conversation turn interval is used as the conversation turn evaluation information for the conversation topic. Alternatively, other methods can be used to quantify the conversation turn count of a conversation topic into a conversation turn evaluation value, such as manual scoring, which are not specifically limited here.
[0137] Optionally, the context-sensitive multi-round dialogue capability of the first human-computer interaction model can be evaluated along the aforementioned multiple dimensions to obtain evaluation information for the multiple dimensions. Based on the weight coefficients corresponding to the various dimensions, the evaluation information for the multiple dimensions can be weighted averaged or summed to provide comprehensive evaluation information for the context-sensitive multi-round dialogue capability of the first human-computer interaction model. This comprehensive evaluation information can be used to guide the launch decision for the first human-computer interaction model, update an optimized version of the first human-computer interaction model, or select a high-quality first human-computer interaction model with strong multi-round dialogue capability.
[0138] The solution of this embodiment can comprehensively and accurately evaluate the context-related multi-round dialogue capability of the first human-computer interaction model by evaluating the context-related multi-round dialogue capability of the first human-computer interaction model from at least one of the three dimensions of micro-dialogue quality, macro-dialogue quality, and dialogue round quality, thereby improving the quality of the evaluation information of the first human-computer interaction model and thus improving the quality of human-computer interaction.
[0139] Based on any of the aforementioned embodiments, in an optional embodiment, the context-independent multi-turn dialogue capability of the first human-computer interaction model may also be evaluated. FIG4 is a flowchart of the context-independent multi-turn dialogue capability evaluation provided by this embodiment. As shown in FIG4 , the specific steps for evaluating the context-independent multi-turn dialogue capability of the first human-computer interaction model are as follows:
[0140] Step S41: Acquire multiple conversation data, where the conversation data includes multiple rounds of independent question information.
[0141] In this embodiment, multiple session data are pre-constructed, each session data includes multiple rounds of question information, wherein the question information of different rounds are independent of each other, and the question information and response information of the previous round of question information should not affect the response information of the next round of question information.
[0142] The multiple conversation data can be obtained by collecting questions generated in existing human-computer interaction scenarios, or randomly generated by a large language model, or manually constructed, which is not specifically limited in this embodiment. In order to ensure that the multiple rounds of question information in the conversation data are independent of each other, the conversation data can be annotated.
[0143] For example, Table 4 shows a conversation data containing 5 rounds of independent question information:
[0144] Table 4
[0145] Step S42: Input multiple rounds of question information in each conversation data into the first human-computer interaction model in sequence to conduct multiple rounds of dialogue, and generate first response information for each round of question information in the conversation data.
[0146] In this embodiment, the first round of questions in a session data set is input into a first human-computer interaction model to obtain responses to the first round of questions generated by the first human-computer interaction model. The second round of questions in the session data is then input into the first human-computer interaction model to obtain responses to the second round of questions generated by the first human-computer interaction model, and so on, until responses to the last round of questions are obtained. This allows for the generation of multi-round conversation information based on the session data. By sequentially inputting the multi-round questions in the session data into the first human-computer interaction model, a multi-round conversation process with the first human-computer interaction model is simulated, resulting in the generation of context-independent multi-round conversation information based on the session data.
[0147] It should be noted that when multiple rounds of question information from each conversation data are sequentially input into the first human-computer interaction model for multiple rounds of dialogue, the first human-computer interaction model will use the context information of the current conversation when generating responses to the questions in each round of conversation data. In a single round of dialogue, the first human-computer interaction model generates responses based solely on the current input and does not use context information.
[0148] In this embodiment, in order to distinguish the response information generated in the multi-round dialogue process simulated based on the conversation data and the response information generated in the single-round dialogue process simulated based on the conversation data, the response information generated in the multi-round dialogue process with the first human-computer interaction model simulated based on the conversation data is recorded as the first response information; and the response information generated in the single-round dialogue process with the first human-computer interaction model simulated based on the conversation data is recorded as the second response information.
[0149] Step S43: Determine first evaluation information of the context-independent multi-round dialogue capability of the first human-computer interaction model based on the response quality information of the first response information of each round of question information in each conversation data.
[0150] Optionally, in this step, response quality information of the first response information of each round of question information in each conversation data is obtained, and the response quality information of the first response information of each round of question information in each conversation data is integrated to determine the first evaluation information of the context-independent multi-round dialogue capability of the first human-computer interaction model.
[0151] Exemplarily, based on the conversation data shown in Table 4, this step simulates multiple rounds of dialogue with the first human-computer interaction model based on the conversation data, and the first response information and evaluation information (evaluation score) of each round of question information obtained are shown in Table 5:
[0152] Table 5
[0153] For example, the first responses to the questions in the five rounds of conversation data shown in Table 5 are used as the response information to be evaluated. The first responses in each round are evaluated to obtain the response quality information (evaluation score) of the first responses to the questions in each round of the conversation data. The response quality information of the first responses to the questions in each round of the conversation data is combined to calculate the first evaluation information of the context-independent multi-round dialogue capability of the first human-computer interaction model.
[0154] Optionally, the average of the response quality information (assessment score) of the first response information to each round of question information in the conversation data may be calculated as the first assessment information of the context-independent multi-round dialogue capability of the first human-computer interaction model.
[0155] Optionally, the average of the response quality information (evaluation score) of the first response information for each round of question information in each conversation data set can be used as the average quality score corresponding to the conversation data set. The quality scores corresponding to the respective conversation data sets are weighted averaged according to a third preset weight coefficient for each conversation data set to obtain first evaluation information of the context-independent multi-round dialogue capability of the first human-computer interaction model. The third preset weight coefficient for each conversation data set can be configured based on the needs of the actual application scenario and is not specifically limited here.
[0156] In an optional embodiment, as shown in FIG5 , after step S43, the method further includes:
[0157] Step S44: input the multiple rounds of question information in each conversation data into the first human-computer interaction model respectively to perform a single round of dialogue, and generate second response information for each round of question information in each conversation data.
[0158] In this step, any question from the conversation data is input into the first human-computer interaction model, which then generates a second response to the question. When generating the second response, the first human-computer interaction model ignores the questions and responses from other rounds and generates the second response based solely on the question input in the current round. During this process, the generated second response is unaffected by the questions and responses from other rounds.
[0159] Step S45: Determine second evaluation information of the context-independent multi-round dialogue capability of the first human-computer interaction model based on the response quality information of the first response information and the difference between the response quality information of the second response information of the same question information in each conversation data.
[0160] In this step, the response quality information of the second response information of each round of question information in each conversation data is obtained. For example, based on the conversation data shown in Table 4, this step simulates a single-round dialogue with the first human-computer interaction model based on the conversation data, and the second response information and evaluation information (evaluation score) of each round of question information are obtained as shown in Table 6:
[0161] Table 6
[0162] For example, the second response information of the five rounds of question information in the conversation data given in Table 5 is used as the response information to be evaluated. The second response information in each row is evaluated separately to obtain the response quality information (evaluation score) of the second response information of each round of question information in the conversation data.
[0163] Furthermore, second evaluation information of the context-independent multi-turn dialogue capability of the first human-computer interaction model is determined based on the difference between the response quality information of the first response information and the response quality information of the second response information for the same question information in each conversation data. The larger the difference, the more degraded the context-independent multi-turn dialogue capability of the first human-computer interaction model is compared to the single-turn dialogue capability.
[0164] For example, by comparing the second response information for the third round of questions in Table 6 with the first response information for the third round of questions in Table 5, it can be found that during the multi-round dialogue, due to the influence of the second round of dialogue information, the response quality evaluation score of the first response information for the third round of questions in Table 5 is lower (compared to the response quality evaluation score of the second response information for the third round of questions in Table 6). Similarly, by comparing the second response information for the fifth round of questions in Table 6 with the first response information for the fifth round of questions in Table 5, it can be found that during the multi-round dialogue, due to the influence of the fourth round of dialogue information, the response quality evaluation score of the first response information for the fifth round of questions in Table 5 is lower (compared to the response quality evaluation score of the second response information for the fifth round of questions in Table 6).
[0165] Optionally, the average of the absolute values of the difference between the response quality information of the first response information and the response quality information of the second response information for the same question information in each session data is calculated as the second evaluation information for the context-independent multi-round dialogue capability of the first human-computer interaction model. Alternatively, the average of the squares of the difference between the response quality information of the first response information and the response quality information of the second response information for the same question information in each session data can be calculated as the second evaluation information for the context-independent multi-round dialogue capability of the first human-computer interaction model. In addition, the second evaluation information for the context-independent multi-round dialogue capability of the first human-computer interaction model can also be determined based on other preset rules by calculating the difference between the response quality information of the first response information and the response quality information of the second response information for the same question information in each session data, which is not specifically limited here.
[0166] The method of this embodiment is based on multiple rounds of question information in the same conversation data, obtains the second response information of each round of question information by simulating a single-round dialogue process, and obtains the first response information of each round of question information by simulating a multi-round dialogue process, and uses the response quality information of the second response information obtained through the unit dialogue as a reference. According to the gap between the response quality information of the first response information and the response quality information of the second response information, the second evaluation information of the context-independent multi-round dialogue capability of the first human-computer interaction model is determined, which can better measure the context-independent multi-round dialogue capability of the first human-computer interaction model and prompt the evaluation quality of the first human-computer interaction model.
[0167] It should be noted that the above-mentioned embodiments can be combined with each other to obtain evaluation information of the context-related multi-round dialogue capability of the first human-computer interaction model, as well as evaluation information of the context-independent multi-round dialogue capability of the first human-computer interaction model. These are all evaluation information of the multi-round dialogue capability of the first human-computer interaction model.
[0168] FIG6 is an interactive flow chart of a data processing method for human-computer interaction provided by an exemplary embodiment of the present disclosure. In this embodiment, taking the language model as an example, the process of evaluating the multi-round dialogue capability of the language model is exemplified. As shown in FIG6, the interaction process between the first server and the end-side device is as follows:
[0169] Step S601: The client device sends a request to the first server for evaluating the multi-round dialogue capability of the language model.
[0170] Among them, the language model can be a pre-trained language model, which can be applied to natural language processing (NLP), computer vision and other fields. Specifically, it can be applied to tasks in the intersection of NLP and computer vision such as visual question answering (VQA), image description (IC), visual entailment (VE), referential expression and understanding (REC), as well as tasks in the field of natural language processing such as text-based sentiment classification tasks and text summarization tasks. It can be applied to various application scenarios such as digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0171] The evaluation request includes relevant information of the language model to be evaluated, such as the application program interface for calling the language model and the access address of the language model.
[0172] Step S602: The first server receives a request for evaluating the multi-round dialogue capability of the language model sent by the terminal-side device.
[0173] After receiving the evaluation request for the multi-round dialogue capability of the language model, the first server extracts relevant information of the language model to be evaluated from the evaluation request, such as the application program interface for calling the language model and the access address of the language model.
[0174] Step S603: The first server obtains at least one pre-built conversation topic.
[0175] The specific implementation method of this step is consistent with the implementation method of the aforementioned step S201. Please refer to the relevant content of the aforementioned embodiment for details, and no further details will be given here.
[0176] Step S604: The first server uses the second human-computer interaction model to conduct multiple rounds of dialogue based on the conversation topic and the language model to obtain multiple rounds of dialogue information of the conversation topic.
[0177] This step is similar to the specific implementation of the aforementioned step S202. In step S604, the first human-computer interaction model to be evaluated is a language model specified by the user. The specific implementation method is described in the aforementioned embodiment and will not be repeated here.
[0178] Step S605: The first server determines evaluation information of the multi-round dialogue capability of the language model based on the multi-round dialogue information of the conversation topic.
[0179] This step is similar to the specific implementation of the aforementioned step S203. For the specific implementation, please refer to the relevant content in the aforementioned embodiment and will not be repeated here.
[0180] In an optional embodiment, after the first server determines the evaluation information of the multi-round dialogue capability of the language model based on the multi-round dialogue information of the conversation topic, it can also output the evaluation information of the multi-round dialogue capability of the language model through steps S606-S608.
[0181] Step S606: The first server outputs evaluation information of the multi-round dialogue capability of the language model to the terminal device.
[0182] Step S607: The client-side device receives the evaluation information of the multi-round dialogue capability of the language model sent by the first server.
[0183] Step S608: The client-side device outputs evaluation information of the multi-round dialogue capability of the language model.
[0184] In this embodiment, a pre-constructed conversation topic is obtained, and a second human-computer interaction model is used to conduct multiple rounds of dialogue with a language model based on the conversation topic, thereby obtaining multi-round dialogue information based on the conversation topic. The dialogue information includes: question information input to the language model, and response information generated by the language model to the question information; the second human-computer interaction model is used to simulate the process of a human user and the language model conducting multiple rounds of dialogue around the conversation topic, thereby generating multi-round dialogue information based on the conversation topic. Furthermore, based on the multi-round dialogue information about the conversation topic, evaluation information of the language model's multi-round dialogue capability is determined, which can truly implement the evaluation of the language model's multi-round dialogue capability and obtain accurate and high-quality evaluation information. This evaluation information is used to guide the online decision of the language model or update the optimized version of the language model. It can accurately select high-quality models during language model iteration, improve the multi-round dialogue quality of the iteratively updated language model, and improve the multi-round dialogue quality of the online model, thereby improving the quality of multi-round dialogue in human-computer interaction.
[0185] FIG7 is a schematic diagram of the structure of a server provided in an embodiment of the present disclosure. As shown in FIG7 , the server includes a memory 701 and a processor 702. Memory 701 is used to store computer-executable instructions and can be configured to store various other data to support operations on the server. Processor 702 is communicatively connected to memory 701 and is used to execute the computer-executable instructions stored in memory 701 to implement the technical solutions provided in any of the above-mentioned method embodiments. The specific functions and technical effects achieved by these embodiments are similar and will not be further described here.
[0186] Optionally, as shown in Figure 7, the server further includes other components such as a firewall 703, a load balancer 704, a communication component 705, and a power supply component 706. Figure 7 only schematically illustrates some components, which does not mean that the server only includes the components shown in Figure 7.
[0187] The embodiments of the present disclosure also provide a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method of any of the aforementioned embodiments is implemented. The specific functions and technical effects that can be achieved are not repeated here.
[0188] The present disclosure also provides a computer program product, including a computer program. When executed by a processor, the computer program implements the method of any of the aforementioned embodiments. The computer program is stored in a readable storage medium. At least one processor of an electronic device can read the computer program from the readable storage medium. The at least one processor executes the computer program, causing the electronic device to implement the technical solution provided by any of the aforementioned method embodiments. The specific functions and technical effects achieved are not further described here.
[0189] The present disclosure provides a chip comprising: a processing module and a communication interface. The processing module is capable of executing the technical solutions of the electronic device described in the aforementioned method embodiments. Optionally, the chip further comprises a storage module (e.g., a memory) configured to store instructions, and the processing module configured to execute the instructions stored in the storage module. Execution of the instructions stored in the storage module causes the processing module to execute the technical solutions provided in any of the aforementioned method embodiments.
[0190] The integrated modules implemented in the form of software function modules can be stored in a computer-readable storage medium. The software function modules stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform some of the steps of the methods of various embodiments of the present disclosure.
[0191] It should be understood that the above-mentioned processor can be a processing unit (Central Processing Unit, referred to as CPU), or it can be other general-purpose processors, digital signal processors (Digital Signal Processor, referred to as DSP), application specific integrated circuits (Application Specific Integrated Circuit, referred to as ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The memory may include high-speed random access memory (Random Access Memory, referred to as RAM), and may also include non-volatile storage, such as at least one disk storage, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.
[0192] The above-mentioned memory may be an object storage service (OSS). The above-mentioned memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0193] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as a mobile hotspot (WiFi), a second-generation mobile communication system (2G), a third-generation mobile communication system (3G), a fourth-generation mobile communication system (4G) / Long Term Evolution (LTE), a fifth-generation mobile communication system (5G), or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth technology and other technologies. The above-mentioned power supply component provides power to various components of the device where the power supply component is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located. The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0194] An exemplary storage medium is coupled to a processor, such that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an application-specific integrated circuit. Of course, the processor and storage medium can also exist as discrete components in an electronic device or a host control device.
[0195] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0196] The order of the above-mentioned embodiments of the present disclosure is for description only and does not represent the advantages and disadvantages of the embodiments. In addition, in some of the processes described in the above-mentioned embodiments and the accompanying drawings, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or in parallel. They are only used to distinguish between different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to different types. The meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.
[0197] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present disclosure.
[0198] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein.
[0199] The above are only preferred embodiments of the present disclosure and are not intended to limit the patent scope of the present disclosure. Any equivalent structure or equivalent process transformation made using the contents of the present disclosure and the drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present disclosure.
Claims
1. A data processing method for human-computer interaction, wherein: include: Obtain pre-built conversation topics and the first human-computer interaction model to be evaluated; Using the second human-computer interaction model to conduct multiple rounds of dialogue with the first human-computer interaction model based on the conversation topic, to obtain multiple rounds of dialogue information of the conversation topic, wherein the dialogue information includes: question information input into the first human-computer interaction model, and response information of the question information generated by the first human-computer interaction model; Determine evaluation information of the multi-round dialogue capability of the first human-computer interaction model based on the multi-round dialogue information of the conversation topic.
2. The method according to claim 1, wherein: Using the second human-computer interaction model to conduct multiple rounds of dialogue with the first human-computer interaction model based on the conversation topic to obtain multiple rounds of dialogue information of the conversation topic, including: The conversation topic includes an opening statement, the opening statement is used as the first round of question information, input into the first human-computer interaction model, and response information of the question information is generated by the first human-computer interaction model to obtain the first round of dialogue information of the conversation topic; Using the second human-computer interaction model, based on the obtained dialogue information of each round of the conversation topic, the next round of question information is generated, the next round of question information is input into the first human-computer interaction model, the next round of response information is generated through the first human-computer interaction model, and the next round of dialogue information is obtained, and so on, until the conversation based on the conversation topic ends, and multiple rounds of dialogue information on the conversation topic are obtained.
3. The method according to claim 2, wherein: The using the second human-computer interaction model to generate the next round of question information based on the obtained dialogue information of each round of the conversation topic includes: The task prompt information corresponding to the conversation topic and the response information of the previous round are input into the second human-computer interaction model. Through the second human-computer interaction model, the question information of the next round is generated according to the task prompt information of the conversation topic and the response information of the previous round, as well as the context information of the conversation, and the context information of the conversation is updated.
4. The method according to claim 2, wherein: The using the second human-computer interaction model to generate the next round of question information based on the obtained dialogue information of each round of the conversation topic includes: The second human-computer interaction model is used to generate question information for the next round based on the obtained dialogue information of each round of the conversation topic and the topic guiding script of the conversation topic.
5. The method according to claim 2, wherein: When the next round of question information output by the second human-computer interaction model is session end information, end the session; or, When the number of interaction rounds of the session is greater than or equal to a preset maximum number of rounds, the session is terminated.
6. The method according to claim 3, wherein: Also includes: Construct task prompt information corresponding to the conversation topic, wherein the task prompt information is used to prompt the second human-computer interaction model to simulate the dialogue between humans and the first human-computer interaction model based on the conversation topic and the conversation task prompt, and generate the next round of question information.
7. The method according to claim 1, wherein: The determining, based on the multi-round dialogue information of the conversation topic, evaluation information of the multi-round dialogue capability of the first human-computer interaction model includes: Determining, according to multiple rounds of dialogue information of the conversation topic, response quality information of the first human-computer interaction model in each round of dialogue information; The response quality information of the first human-computer interaction model in each round of dialogue information of each conversation topic is integrated to determine the response quality evaluation information of the first human-computer interaction model in multiple rounds of dialogue related to the context.
8. The method according to claim 1, wherein: The determining, based on the multi-round dialogue information of the conversation topic, evaluation information of the multi-round dialogue capability of the first human-computer interaction model includes: Determining overall response quality information of the first human-computer interaction model in the multi-round dialogue based on the conversation topic according to the multi-round dialogue information of the conversation topic; The average value of the overall response quality information of the first human-computer interaction model in the multiple rounds of dialogue based on each of the conversation topics is used as the overall response quality evaluation information of the first human-computer interaction model in the context-related multiple rounds of dialogue.
9. The method according to claim 1, wherein: The determining, based on the multi-round dialogue information of the conversation topic, evaluation information of the multi-round dialogue capability of the first human-computer interaction model includes: According to the number of conversation turns of the conversation topic and the conversation type corresponding to the conversation topic, conversation turn evaluation information of the first human-computer interaction model in the context-related multi-round conversation is determined.
10. The method according to claim 1, wherein: Also includes: Acquire multiple conversation data, wherein the conversation data includes multiple rounds of independent question information; Inputting multiple rounds of question information in each of the conversation data into the first human-computer interaction model in sequence to conduct multiple rounds of dialogue, and generating first response information for each round of question information in the conversation data; First evaluation information of the context-independent multi-round dialogue capability of the first human-computer interaction model is determined according to response quality information of the first response information of each round of question information in each of the conversation data.
11. The method according to claim 10, wherein: Also includes: Inputting the multiple rounds of question information in each of the conversation data into the first human-computer interaction model respectively to perform a single round of dialogue, and generating second response information for each round of question information in each of the conversation data; According to the difference between the response quality information of the first response information and the response quality information of the second response information of the same question information in each of the conversation data, the second evaluation information of the context-independent multi-round dialogue capability of the first human-computer interaction model is determined.
12. The method according to any one of claims 1 to 11, wherein: Also includes: Determining whether the first human-computer interaction model meets an online condition according to the evaluation information of the first human-computer interaction model; Output online prompt information of the first human-computer interaction model and / or evaluation information of the first human-computer interaction model, wherein the online prompt information indicates whether the first human-computer interaction model meets online conditions.
13. The method according to any one of claims 1 to 11, wherein: Also includes: According to the evaluation information of the multiple first human-computer interaction models, select one of the first human-computer interaction models as a target human-computer interaction model, and output information of the target human-computer interaction model to the terminal side device; or, According to evaluation information of multiple different versions of the first human-computer interaction model, one of the versions is selected as an optimized version of the first human-computer interaction model; or, Output evaluation information of the multi-round dialogue capability of the first human-computer interaction model.
14. A data processing method for human-computer interaction, wherein: include: A request for evaluating the multi-round dialogue capability of the language model sent by a receiving device, and obtaining at least one pre-built conversation topic; Using the second human-computer interaction model to conduct multiple rounds of dialogue with the language model based on the conversation topic, obtaining multiple rounds of dialogue information of the conversation topic, wherein the dialogue information includes: question information input to the language model, and response information to the question information generated by the language model; Determine evaluation information of the multi-round dialogue capability of the language model based on the multi-round dialogue information of the conversation topic.
15. A server, wherein: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the server to execute the method described in any one of claims 1-14.
16. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the method according to any one of claims 1 to 14 is implemented.
17. A computer program, wherein: When the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Question and answer dialogue evaluation method, device and equipment and storage medium
CN112487140A
Man-machine interaction data processing method and server
CN116501592A
Dialogue model evaluation text acquisition method and device, electronic equipment and storage medium
CN116756284A
Man-machine interaction data processing method, server and storage medium
CN117332071A
Dialogue Model Training Method and Device Therefor
US20230080930A1