Large model and agent evaluation method and device based on multi-round session data set
By receiving and storing multi-round conversation data sets, determining dialogue group data, and displaying evaluation results based on data type, the problem in existing technologies of being difficult to effectively evaluate large models and intelligent agents in complex multi-round dialogue scenarios is solved, achieving more accurate evaluation and tuning.
Patent Information
- Application Number
- CN202511152279.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies make it difficult to effectively evaluate large models and agents in complex multi-round dialogue scenarios, resulting in poor evaluation results.
A large model and agent evaluation method based on a multi-round conversation dataset is provided. By receiving and storing the evaluation dataset, multiple conversation group data are determined, and the evaluation results are displayed according to the correspondence between the data type and the display method.
It enables comprehensive evaluation of large models and intelligent agents in complex multi-round dialogue scenarios. The evaluation results more accurately reflect their capabilities and support better tuning and improvement.
Smart Images

Figure CN120653950A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of large models, intelligent agents, and computer technology, and in particular, to a large model and intelligent agent evaluation method and device based on a multi-round conversation data set. Background Art
[0002] With the development of artificial intelligence (AI) technology, large models and intelligent agents are increasingly being used in various fields, making their evaluation increasingly important. Currently, common evaluation methods for large models and agents are often limited to simple grammatical and semantic judgments for single-round conversations. However, as the scenarios in which large models and agents are used become increasingly complex, especially when dealing with complex conversations that closely resemble real-world scenarios, current evaluation methods often suffer from poor results. Summary of the Invention
[0003] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, the present disclosure provides a large model and agent evaluation method based on a multi-round conversation dataset, the method comprising: Receiving and storing an evaluation data set, the evaluation data set comprising at least one data entry, each of the data entries comprising a dialogue group identifier, a dialogue turn, and at least one conversation data, the at least one conversation data comprising questions for interacting with an evaluation subject and reference answers to the questions, the evaluation subject comprising an intelligent agent or a large model; Determining a plurality of dialogue group data according to the evaluation data set, wherein one dialogue group data includes conversation data of a plurality of dialogue turns corresponding to the same dialogue group identifier; In response to receiving a first instruction for viewing target conversation group data among the plurality of conversation group data, each target conversation data is displayed separately according to a data type of each target conversation data in the target conversation group and a correspondence between a preset data type and a data display method.
[0005] In a second aspect, the present disclosure provides a large model and agent evaluation device based on a multi-round conversation dataset, the device comprising: a first receiving module, configured to receive and store an evaluation dataset, the evaluation dataset comprising at least one data entry, each of the data entries comprising a dialogue group identifier, a dialogue turn, and at least one conversation data, the at least one conversation data comprising questions for interacting with an evaluation subject and reference answers to the questions, the evaluation subject comprising an agent or a large model; A first determining module is configured to determine a plurality of dialogue group data according to the evaluation data set, wherein each dialogue group data includes conversation data of a plurality of dialogue turns corresponding to the same dialogue group identifier; The evaluation module is configured to evaluate the evaluation object using the plurality of dialogue group data in response to receiving an evaluation instruction for the evaluation object to obtain an evaluation result.
[0006] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect of the present disclosure.
[0007] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the method described in the first aspect of the present disclosure.
[0008] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method described in the first aspect of the present disclosure.
[0009] Through the above technical solution, an evaluation dataset is received and stored. The evaluation dataset includes at least one data entry, each data entry including a dialogue group identifier, a dialogue turn, and at least one conversation data item. The at least one conversation data item includes questions used to interact with the evaluation subject and reference answers to the questions. Based on the evaluation dataset, multiple dialogue group data items are determined, each dialogue group data item including conversation data from multiple dialogue turns corresponding to the same dialogue group identifier. Thus, by establishing a unified structure for the evaluation dataset's data items, the evaluation dataset's content can be quickly parsed, and multiple conversation turns belonging to the same dialogue group can be quickly integrated together. This allows the evaluation dataset to contain a variety of multi-turn conversations, allowing the evaluation dataset to more realistically simulate complex interaction scenarios and facilitate a more comprehensive evaluation of the evaluation subject. Thus, intelligent agents and large models can be evaluated based on dialogue group data with multi-turn conversations. The evaluation dataset used is more closely aligned with the complexity of actual conversation scenarios, enabling a more comprehensive and accurate assessment of the capabilities of the intelligent agents and large models, facilitating better tuning of the intelligent agents and large models.
[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a flowchart of a large model and agent evaluation method based on a multi-round conversation dataset provided according to an embodiment of the present disclosure; Figure 2 This is an exemplary schematic diagram of the first interface in the large model and agent evaluation method based on a multi-round conversation dataset provided by the present disclosure; Figure 3 This is an exemplary interface diagram for data uploading in the large model and agent evaluation method based on a multi-round conversation dataset provided by the present disclosure; Figure 4 This is an exemplary schematic diagram of a preset interface in the large model and agent evaluation method based on a multi-round conversation dataset provided by the present disclosure; Figure 5 This is an exemplary schematic diagram of the second interface in the large model and agent evaluation method based on a multi-round conversation dataset provided by the present disclosure; Figure 6 is a block diagram of a large model and an agent evaluation device based on a multi-round conversation dataset provided according to an embodiment of the present disclosure; Figure 7 A schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0012] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0013] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0014] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0015] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0016] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0017] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0018] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0019] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0020] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0021] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0022] At the same time, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0023] Figure 1 This is a flowchart of a large model and agent evaluation method based on a multi-round conversation dataset provided according to an embodiment of the present disclosure. Figure 1 As shown, the method provided by the present disclosure may include the following steps 11 to 13.
[0024] In step 11, an evaluation data set is received and stored.
[0025] The evaluation data set may include at least one data entry, each data entry may include a dialogue group identifier, a dialogue turn and at least one session data, and the at least one session data includes questions used to interact with the evaluation object (i.e., content used to input the evaluation object) and reference answers to the questions (i.e., content output by the evaluation object).
[0026] In the present disclosure, an evaluation subject can be an object evaluated based on an evaluation dataset, including but not limited to a large language model (LLM), an agent, or a prompt. Evaluation of the evaluation subject can be implemented based on a dialogue group. A dialogue group can correspond to a dialogue scenario and can include multiple rounds of dialogue for interacting with the evaluation subject. A dialogue round can include at least one question input to the evaluation subject during an interaction and its reference answer. The content of subsequent dialogue rounds should depend on the context of the content of previous dialogue rounds. By sequentially inputting questions from the dialogue group to the evaluation subject, collecting the actual answers output by the evaluation subject, and comparing the actual answers output by the evaluation subject with the reference answers, the evaluation subject's performance in the dialogue scenario of the dialogue group can be evaluated. In this way, based on multiple different dialogue groups, a comprehensive evaluation of the evaluation subject's performance in different scenarios can be achieved.
[0027] The evaluation dataset may include at least one data entry, each of which may include a conversation group identifier, a conversation turn, and at least one piece of conversation data. The at least one piece of conversation data may include questions used to interact with the evaluation subject and reference answers to the questions. The conversation group identifier uniquely identifies a conversation group and distinguishes it from other conversation groups.
[0028] Session data can be data involved in the interaction with the evaluation object, typically including at least two types of session data: questions used to interact with the evaluation object and reference answers to those questions. In addition, other session data can be configured based on actual needs. For example, for a weather forecast agent, the evaluation object can include four types of session data: input question, address, time, and reference answer. For example, the question could be "What's the weather like today?", the address could be location A, the event could be time B, and the reference answer could be "Hello, the weather at location A at time B is sunny."
[0029] A conversation turn may represent the conversation turn in which the conversation data under the data entry is located in a conversation group. A complete conversation turn may correspond to questions and answers in an interaction. For example, a conversation turn may be represented by a turn sequence number. For example, if the turn sequence number is 1, it indicates that the data entry is the data used in the first round of interaction with the evaluation object, and may include questions and answers in the first round of interaction.
[0030] The data type of each session data can be one of a number of predefined data types. In the present disclosure, predefined data types include, but are not limited to, text, rich text, images, audio, video, PDF (Portable Document Format), CSV (Comma-Separated Values, a plain text format that separates data fields by commas), and Word documents (commonly suffixed with .doc or .docx).
[0031] Based on this, conversation data from different conversation groups, conversation data from different rounds of the same conversation group, and even different conversation data from the same round of the same conversation group can all use different data types as needed, facilitating multimodal evaluation of the evaluation object. In other words, different conversation data from different rounds of different conversation groups can all be set to any usable data type as needed based on the preset data type. For example, if a conversation group contains two rounds of conversation, and each round of conversation contains two conversation data items: questions and answers, then each conversation data item in these two rounds can have a different data type. For example, the two conversation data items of the second round's questions and answers can use different data types, and the two conversation data items of the first and second rounds' questions can each use different data types. Other scenarios are not listed one by one.
[0032] In a possible implementation, step 11 may include the following steps: receiving a second instruction for instructing batch uploading of evaluation data sets; In response to receiving the second instruction, receiving a data compression package, the data compression package including a description table and a plurality of conversation files, the description table including a plurality of rows of data, each row of data including a conversation group identifier, a conversation turn, a file path, and a data type; For each row of data, a target session file is determined from multiple session files based on the file path included in the row of data, and the content in the target session file is stored as a data entry in association with the conversation group identifier and conversation turn included in the row of data as conversation data.
[0033] Optionally, the present disclosure may display a first button for batch uploading of evaluation data sets through a preset interface, and receive a second instruction generated by triggering the first button. Figure 4 The "Data Import" button shown triggers the generation of the second instruction. In a batch upload scenario, the data to be uploaded is usually uploaded in the form of a compressed package. The compressed package can be pre-set with a unified data format. The present disclosure can parse and further process the compressed package.
[0034] In response to receiving the second instruction, a data compression package is received. The data compression package may include a description table and multiple session files, wherein the description table may be a CSV table, and the multiple session files constitute a multimodal file. The description table may include multiple rows of data, each row of data may include a conversation group identifier, a conversation turn, a file path, and a data type. The file path is the file path of the session file associated with the row of data, and the data type is the data type of the session file associated with the row of data, i.e., one of the multiple preset data types described above.
[0035] Each row of data in the description table corresponds to the conversation data of a conversation turn in a conversation group. Therefore, for each row of data, the target conversation file can be determined from multiple conversation files based on the file path included in the row data, and the content in the target conversation file can be stored as a data entry as the conversation data in association with the conversation group identifier and conversation turn included in the row data.
[0036] Optionally, after receiving the compressed data package, it can be decompressed (for example, using a Python compression library) to read the description table. When parsing the description table, a verification process can be performed to verify the validity of the file path and the format of the session file. After verification, a metadata index (which may include information such as file name, data type, size, and upload time) can be generated for each session file and associated with the corresponding conversation group identifier and conversation turn, thereby enabling the associated storage of conversation data, conversation group identifier, and conversation turn. For example, verification of the session file format can be implemented based on the data type of the session file. For example, if the data type of the session file is an image, verification can be implemented using a regular expression or file header signature.
[0037] Optionally, when storing session files of different data types, they can be stored in a targeted manner according to pre-set rules. For example, binary-based session files such as images, audio, and video can be stored in a distributed file system, and a unique hash value and access link can be generated; session files such as plain text, rich text, PDF, and CSV can be stored in a relational database in a structured or semi-structured form; paragraph styles and attachment reference relationships of rich text can be stored in JSON format (the paragraph format of rich text can be recorded using Markdown syntax, and the storage path of attachment files can be mapped by associating fields to ensure the integrity of text styles and attachments). In this way, a combination of distributed file system and database storage methods can be implemented to achieve file storage, so that different types of files can be stored in appropriate locations as needed.
[0038] In this way, the evaluation data set can be quickly constructed, parsed, and stored by uploading data compression packages, which is highly efficient.
[0039] In another possible implementation, step 11 may include the following steps: receiving a third instruction for instructing to upload the evaluation data sets one by one; In response to receiving the third instruction, displaying a first interface, wherein the first interface provides configuration items for configuring data entries, the configuration items including a first configuration item for configuring a conversation group identifier, a second configuration item for configuring a conversation turn, and at least one third configuration item for configuring session data; In response to a configuration operation on a configuration item in the first interface, at least one data entry is obtained and stored.
[0040] Optionally, the present disclosure may display a second button for uploading the evaluation data set one by one through a preset interface, and receive a third instruction generated by triggering the second button. In the uploading scenario one by one, a separate configuration can usually be performed for a conversation group. For example, Figure 4 The "Add Conversation Group" button shown triggers the generation of this third instruction.
[0041] In response to receiving the third instruction, a first interface may be displayed. The first interface may include configuration items for configuring data entries. The configuration items may include a first configuration item for configuring a conversation group identifier, a second configuration item for configuring conversation turns, and at least one third configuration item for configuring conversation data. For example, if the conversation data only includes questions and reference answers, the configuration items may include a third configuration item for configuring the questions and a third configuration item for configuring the reference answers. Typically, the conversation group identifier can be automatically generated, so the first configuration item may not be provided.
[0042] like Figure 2 The figure shows an exemplary diagram of the first interface. In the first interface, the numbers 1, 2, and 3 on the left represent conversation rounds. More conversation round data can be added by triggering the add button B1 in the figure. The question column is used to enter questions for different rounds, and the reference answer column is used to enter reference answers for different rounds. The question and reference answer corresponding to round 1 are the two conversation data of round 1. In each round, an input box for each conversation data is provided. The editing mode of the input box can be switched by the mode switch button. Taking the question column as an example, if button B2 selects rich text mode, the input box corresponding to the first round will switch to a rich text edit box for users to edit rich text questions. If button B3 selects text mode, the input box corresponding to the second round will switch to text editing mode. If button B4 selects file mode, the input box of the third round will provide a file upload interface for receiving conversation files.
[0043] In this way, a flexible configuration method can be provided for users to configure conversation groups. In response to the configuration operation for the configuration item in the first interface, the conversation data configured for a conversation group can be obtained, where each row corresponds to a data entry, and at least one data entry obtained can be added to the evaluation data set.
[0044] In addition, whether uploading in batches or uploading one by one, you can choose the data upload method when uploading, such as Figure 3 The data uploading method may include a full coverage method, that is, the currently uploaded evaluation data set overwrites the existing evaluation data set, and the data uploading method may also include an appending method, that is, the currently uploaded evaluation data set is appended to the existing evaluation data set to form a new evaluation data set.
[0045] In step 12, a plurality of dialogue group data are determined based on the evaluation data set.
[0046] Among them, one conversation group data may include session data of multiple conversation turns corresponding to the same conversation group identifier.
[0047] That is, based on multiple data entries in the evaluation data set, conversation data corresponding to the same conversation group identifier can be aggregated to form a complete conversation group, which includes conversation data of multiple rounds.
[0048] In a possible implementation, step 12 may include the following steps: For each dialogue group identifier, the target data items corresponding to the dialogue group identifier are obtained, and the target data items are combined in ascending order of round numbers to obtain the dialogue group data corresponding to the dialogue group identifier.
[0049] For each conversation group ID, the data entry corresponding to the conversation group ID can be retrieved from the stored evaluation dataset as the target data entry. This allows the acquisition of conversation data for all conversation turns within a conversation group, where a conversation turn is the turn number. Furthermore, by combining the target data entries in ascending order of turn numbers, the complete multi-turn conversation data for the conversation group can be obtained.
[0050] In step 13, in response to receiving the evaluation instruction for the evaluation object, the evaluation object is evaluated using the plurality of dialogue group data to obtain an evaluation result.
[0051] Optionally, an evaluation task can be constructed for the evaluation object, and a dataset used for the evaluation task, namely, the evaluation dataset in this disclosure, can be set. Upon receiving an evaluation instruction for the evaluation object, a determination is made in response to the instruction that the evaluation dataset can be used to evaluate the evaluation object. Based on this, the multiple conversation group data corresponding to the evaluation dataset can be utilized. For each conversation group, conversation group data (i.e., all conversation data (with contextual information)) is obtained for each conversation group and used in the evaluation of the evaluation object. When obtaining conversation group data for a conversation group, the conversation data within the conversation group data is obtained one by one in ascending order of conversation turns, ensuring that the evaluation data used encompasses multiple conversation turns and contains highly relevant contextual information.
[0052] Based on this, the evaluation results of each dialogue group data corresponding to the evaluation data set can be obtained. Based on the evaluation results, the evaluation object can be updated and optimized in a targeted manner.
[0053] In one possible implementation, Figure 1 Based on the steps shown, the method provided by the present disclosure may further include the following steps: In response to receiving a first instruction for viewing target conversation group data in a plurality of conversation group data, each target conversation data is displayed separately according to the data type of each target conversation data in the target conversation group data and the correspondence between the preset data type and the data display mode.
[0054] Optionally, data presentation methods may include but are not limited to plain text, rich text, JSON, etc.
[0055] After the evaluation data set is stored, it can be preliminarily displayed in units of dialogue groups through the preset interface, such as Figure 4 The first instruction for viewing target conversation group data among multiple conversation group data can be triggered through a preset interface. For example, after a user clicks on a display area of a conversation group, the first instruction for the conversation group is generated.
[0056] After receiving the first instruction, the content of the target conversation group may be further displayed through the second interface according to the target conversation group indicated by the first instruction.
[0057] In a possible implementation, step 13 may include the following steps: In response to receiving the first instruction, for each target conversation data in the target dialogue group data, the target data display method corresponding to the data type of the target conversation data is determined according to the correspondence between the preset data type and the data display method, and the target conversation data is displayed according to the target data display method through the second interface.
[0058] Generally, the correspondence between preset data types and data display methods can be pre-set, and can generally be set as needed to ensure user readability. Specifically, text data display methods may include directly displaying text content, rich text data display methods may include parsing HTML styles and displaying them, and JSON data display methods may include formatting and displaying structured data.
[0059] In response to the first instruction, each session data in the target conversation group data can be used as the target session data and the above steps provided by the present disclosure can be performed. For each target session data, the target data display mode corresponding to the data type of the target session data can be determined based on the correspondence between the preset data type and the data display mode, and the target session data can be displayed according to the target data display mode through the second interface. During the display, the dynamic component mechanism can be used to render different views according to the data display mode, such as Figure 5 As shown, each conversation data of each dialogue round is rendered and displayed separately, the questions of the first round are displayed in JSON format, the questions of the second round are displayed in text format, the reference answers of the first round are displayed in rich text format, and the reference answers of the second round are displayed in text format.
[0060] In a possible embodiment, the method provided by the present disclosure may further include the following steps: Displaying a preset component in the display area of each target session data, the preset component is used to switch the target session data between multiple data display modes; In response to receiving a trigger instruction for a preset component in a target display area, determining a data display mode corresponding to the trigger instruction; The display mode of the target session data in the target display area is switched to the data display mode corresponding to the trigger instruction.
[0061] For example, the preset component can be as follows Figure 5 As shown in B5, it can switch the data display format between text, rich text, and JSON. If a trigger instruction for a preset component is received, it can determine the data display format selected by the trigger instruction and display the session data according to the selected data display format.
[0062] Through the above method, each target conversation data in the target conversation group is displayed separately based on its data type and the correspondence between the preset data type and data display method. This allows for parsing conversation data of different data types, supporting multimodal datasets. Multimodal data can also be displayed separately by type. Multiple rounds of conversation data in the same conversation group can be configured as different data types as needed and displayed separately. This facilitates displaying each round of conversation data in the most appropriate and readable format, improving visualization. Based on this, users can more intuitively view evaluation datasets and apply them to desired evaluation tasks as needed, improving the efficiency of task construction.
[0063] In a possible embodiment, the method provided by the present disclosure may further include the following steps: receiving a fourth instruction for adding session data; In response to receiving the fourth instruction, for each dialogue group data: based on the target number of dialogue turns of the dialogue group data, a target number of input areas are displayed, each input area is used to receive the dialogue data of one dialogue turn in the dialogue group data; in response to an input operation on the input area, the input content corresponding to the input operation is added to the dialogue group data as new dialogue data.
[0064] For example, you can Figure 4 The "Add Column" button shown triggers this fourth instruction.
[0065] Upon receiving the fourth instruction, a corresponding number of input areas can be generated and displayed for each conversation group data item based on the number of conversation rounds in the conversation group data item. Each input area is used for the user to enter conversation data for one conversation round. For example, if a conversation group item contains conversation data for two conversation rounds, when a column of conversation data is added, two input areas will be generated, one for entering new conversation data for the first round and one for entering new conversation data for the second round. For example, if a new column is added for adding conversation data item location, two input areas will be generated, one for the user to enter the location for the first round and one for the second round. This allows for flexible configuration of conversation groups.
[0066] Through the above technical solution, an evaluation dataset is received and stored. The evaluation dataset includes at least one data entry, each data entry including a dialogue group identifier, a dialogue turn, and at least one conversation data item. The at least one conversation data item includes questions used to interact with the evaluation subject and reference answers to the questions. Based on the evaluation dataset, multiple dialogue group data items are determined, each dialogue group data item including conversation data from multiple dialogue turns corresponding to the same dialogue group identifier. Thus, by establishing a unified structure for the evaluation dataset's data items, the evaluation dataset's content can be quickly parsed, and multiple conversation turns belonging to the same dialogue group can be quickly integrated together. This allows the evaluation dataset to contain a variety of multi-turn conversations, allowing the evaluation dataset to more realistically simulate complex interaction scenarios and facilitate a more comprehensive evaluation of the evaluation subject. Thus, intelligent agents and large models can be evaluated based on dialogue group data with multi-turn conversations. The evaluation dataset used is more closely aligned with the complexity of actual conversation scenarios, enabling a more comprehensive and accurate assessment of the capabilities of the intelligent agents and large models, facilitating better tuning of the intelligent agents and large models.
[0067] Figure 6 This is a block diagram of a large model and an agent evaluation device based on a multi-round conversation dataset provided according to an embodiment of the present disclosure. Figure 6 As shown, the device 60 includes: A first receiving module 61 is configured to receive and store an evaluation dataset, wherein the evaluation dataset includes at least one data entry, each of which includes a dialogue group identifier, a dialogue turn, and at least one conversation data, wherein the at least one conversation data includes questions used to interact with an evaluation subject and reference answers to the questions, wherein the evaluation subject includes an agent or a large model; A first determining module 62 is configured to determine a plurality of dialogue group data according to the evaluation data set, wherein each dialogue group data includes conversation data of a plurality of dialogue turns corresponding to the same dialogue group identifier; The evaluation module 63 is configured to evaluate the evaluation object using the plurality of dialogue group data in response to receiving an evaluation instruction for the evaluation object, and obtain an evaluation result.
[0068] Optionally, the data type of each session data is one of a plurality of preset data types; The device 60 further comprises: The first display module is configured to, in response to receiving a first instruction for viewing target conversation group data among the plurality of conversation group data, display each target conversation data according to a data type of each target conversation data in the target conversation group data and a correspondence between a preset data type and a data display method.
[0069] Optionally, the first receiving module 61 includes: A first receiving submodule, configured to receive a second instruction for instructing to batch upload the evaluation data set; a second receiving submodule, configured to receive, in response to receiving the second instruction, a data compression package, the data compression package including a description table and a plurality of session files, the description table including a plurality of rows of data, each row of data including a conversation group identifier, a conversation turn, a file path, and a data type; A storage submodule is used to determine, for each row of data, a target session file from among the multiple session files based on the file path included in the row of data, and store the content of the target session file as session data in association with the conversation group identifier and conversation turn included in the row of data as one of the data entries.
[0070] Optionally, the first receiving module 61 includes: A third receiving submodule, configured to receive a third instruction for instructing to upload the evaluation data sets one by one; a display submodule, configured to display a first interface in response to receiving the third instruction, wherein the first interface provides configuration items for configuring the data entry, the configuration items including a first configuration item for configuring a conversation group identifier, a second configuration item for configuring a conversation turn, and at least one third configuration item for configuring session data; The configuration submodule is configured to obtain and store at least one data entry in response to a configuration operation on the configuration item in the first interface.
[0071] Optionally, the dialogue turn is a turn sequence number; The first determining module 62 is configured to obtain target data entries corresponding to each conversation group identifier, and combine the target data entries in ascending order of round numbers to obtain conversation group data corresponding to the conversation group identifier.
[0072] Optionally, the device 60 further includes: a second receiving module, configured to receive a fourth instruction for adding session data; a processing module configured to, in response to receiving the fourth instruction, for each dialogue group data: display a target number of input areas based on a target number of dialogue turns for the dialogue group data, each of the input areas being configured to receive dialogue data for one dialogue turn in the dialogue group data; and, in response to an input operation performed on the input area, add input content corresponding to the input operation as new dialogue data to the dialogue group data.
[0073] Optionally, the first display module is used to, in response to receiving the first instruction, determine, for each target conversation data in the target dialogue group data, a target data display method corresponding to the data type of the target conversation data according to a correspondence between a preset data type and a data display method, and display the target conversation data according to the target data display method through a second interface.
[0074] Optionally, the device 60 further includes: A second display module is used to display a preset component in a display area of each target session data, wherein the preset component is used to switch the target session data between multiple data display modes; a second determining module, configured to, in response to receiving a trigger instruction for a preset component in a target display area, determine a data display mode corresponding to the trigger instruction; The display switching module is used to switch the display mode of the target session data in the target display area to a data display mode corresponding to the trigger instruction.
[0075] Optionally, the preset data types include: text, rich text, picture, audio, video, PDF, CSV, and Word document.
[0076] Optionally, the data display mode includes plain text, rich text, and JSON.
[0077] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0078] Based on the same inventive concept, the present disclosure also provides a computer-readable medium on which a computer program is stored. When the computer program is executed by a processing device, it implements the steps of the large model and intelligent agent evaluation method based on multi-round conversation data sets provided in any embodiment of the present disclosure.
[0079] Based on the same inventive concept, the present disclosure further provides an electronic device, comprising: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the large model and agent evaluation method based on a multi-round conversation data set provided in any embodiment of the present disclosure.
[0080] Based on the same inventive concept, the present disclosure also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the large model and agent evaluation method based on multi-round conversation data sets provided in any embodiment of the present disclosure.
[0081] Reference below Figure 7 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing an embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0082] like Figure 7 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.
[0083] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Figure 7 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0084] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0085] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.
[0086] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0087] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0088] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: receives and stores an evaluation data set, wherein the evaluation data set includes at least one data entry, each of the data entry includes a dialogue group identifier, a dialogue turn, and at least one conversation data, wherein the at least one conversation data includes questions used to interact with an evaluation object and reference answers to the questions, and the evaluation object includes an intelligent agent or a large model; determines a plurality of dialogue group data based on the evaluation data set, wherein one dialogue group data includes conversation data of a plurality of dialogue turns corresponding to the same dialogue group identifier; and in response to receiving an evaluation instruction for the evaluation object, evaluates the evaluation object using the plurality of dialogue group data to obtain an evaluation result.
[0089] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0091] The modules described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself. For example, the first receiving module may also be described as a "module for receiving and storing an evaluation dataset."
[0092] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0093] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0094] According to one or more embodiments of the present disclosure, a large model and agent evaluation method based on a multi-round conversation dataset is provided, the method comprising: Receiving and storing an evaluation data set, the evaluation data set comprising at least one data entry, each of the data entries comprising a dialogue group identifier, a dialogue turn, and at least one conversation data, the at least one conversation data comprising questions for interacting with an evaluation subject and reference answers to the questions, the evaluation subject comprising an intelligent agent or a large model; Determining a plurality of dialogue group data according to the evaluation data set, wherein one dialogue group data includes conversation data of a plurality of dialogue turns corresponding to the same dialogue group identifier; In response to receiving an evaluation instruction for the evaluation object, the evaluation object is evaluated using the plurality of conversation group data to obtain an evaluation result.
[0095] According to one or more embodiments of the present disclosure, a large model and agent evaluation method based on a multi-round conversation data set is provided, wherein the data type of each of the conversation data is one of a plurality of preset data types; The method further comprises: In response to receiving a first instruction for viewing target conversation group data among the plurality of conversation group data, each target conversation data is displayed separately according to a data type of each target conversation data in the target conversation group data and a correspondence between a preset data type and a data display method.
[0096] According to one or more embodiments of the present disclosure, a large model and agent evaluation method based on a multi-round conversation dataset is provided, wherein receiving and storing the evaluation dataset includes: receiving a second instruction for instructing batch uploading of the evaluation data set; In response to receiving the second instruction, receiving a data compression package, the data compression package including a description table and a plurality of session files, the description table including a plurality of rows of data, each row of data including a conversation group identifier, a conversation turn, a file path, and a data type; For each row of data, a target session file is determined from the multiple session files based on the file path included in the row of data, and the content of the target session file is stored as session data in association with the conversation group identifier and conversation turn included in the row of data as one of the data entries.
[0097] According to one or more embodiments of the present disclosure, a large model and agent evaluation method based on a multi-round conversation dataset is provided, wherein receiving and storing the evaluation dataset includes: receiving a third instruction for instructing uploading the evaluation data sets one by one; In response to receiving the third instruction, displaying a first interface, wherein the first interface provides configuration items for configuring the data entry, the configuration items including a first configuration item for configuring a conversation group identifier, a second configuration item for configuring a conversation turn, and at least one third configuration item for configuring session data; In response to a configuration operation on the configuration item in the first interface, at least one data entry is obtained and stored.
[0098] According to one or more embodiments of the present disclosure, a large model and agent evaluation method based on a multi-turn conversation dataset is provided, wherein the conversation turns are turn sequence numbers; Determining a plurality of dialogue group data according to the evaluation data set includes: For each dialog group identifier, target data entries corresponding to the dialog group identifier are obtained, and the target data entries are combined in ascending order of round numbers to obtain dialog group data corresponding to the dialog group identifier.
[0099] According to one or more embodiments of the present disclosure, a large model and agent evaluation method based on a multi-round conversation dataset is provided, the method further comprising: receiving a fourth instruction for adding session data; In response to receiving the fourth instruction, for each dialogue group data: based on the target number of dialogue turns of the dialogue group data, a target number of input areas are displayed, each of the input areas is used to receive the dialogue data of one dialogue turn in the dialogue group data; in response to an input operation on the input area, the input content corresponding to the input operation is added to the dialogue group data as new dialogue data.
[0100] According to one or more embodiments of the present disclosure, a large model and agent evaluation method based on a multi-round conversation dataset is provided. In response to receiving a first instruction for viewing target conversation group data among the plurality of conversation group data, each target conversation data is displayed according to the data type of each target conversation data in the target conversation group data and the correspondence between a preset data type and a data display mode, including: In response to receiving the first instruction, for each target conversation data in the target dialogue group data, the target data display method corresponding to the data type of the target conversation data is determined according to the correspondence between the preset data type and the data display method, and the target conversation data is displayed according to the target data display method through the second interface.
[0101] According to one or more embodiments of the present disclosure, a large model and agent evaluation method based on a multi-round conversation dataset is provided, the method further comprising: Displaying a preset component in a display area of each target session data, wherein the preset component is used to switch the target session data between multiple data display modes; In response to receiving a trigger instruction for a preset component in a target display area, determining a data display mode corresponding to the trigger instruction; The display mode of the target session data in the target display area is switched to a data display mode corresponding to the trigger instruction.
[0102] According to one or more embodiments of the present disclosure, a large model and agent evaluation method based on a multi-round conversation data set is provided, and the preset data types include: text, rich text, pictures, audio, video, PDF, CSV, and Word documents.
[0103] According to one or more embodiments of the present disclosure, a large model and agent evaluation method based on a multi-round conversation dataset is provided, and the data display methods include plain text, rich text, and JSON.
[0104] According to one or more embodiments of the present disclosure, a large model and agent evaluation device based on a multi-round conversation dataset is provided, the device comprising: a first receiving module, configured to receive and store an evaluation dataset, the evaluation dataset comprising at least one data entry, each of the data entries comprising a dialogue group identifier, a dialogue turn, and at least one conversation data, the at least one conversation data comprising questions for interacting with an evaluation subject and reference answers to the questions, the evaluation subject comprising an agent or a large model; A first determining module is configured to determine a plurality of dialogue group data according to the evaluation data set, wherein each dialogue group data includes conversation data of a plurality of dialogue turns corresponding to the same dialogue group identifier; The evaluation module is configured to evaluate the evaluation object using the plurality of dialogue group data in response to receiving an evaluation instruction for the evaluation object to obtain an evaluation result.
[0105] According to one or more embodiments of the present disclosure, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processing device, the steps of the large model and agent evaluation method based on a multi-round conversation data set provided by any embodiment of the present disclosure are implemented.
[0106] According to one or more embodiments of the present disclosure, there is provided an electronic device, including: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the large model and agent evaluation method based on a multi-round conversation data set provided in any embodiment of the present disclosure.
[0107] According to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the large model and agent evaluation method based on a multi-round conversation data set provided by any embodiment of the present disclosure.
[0108] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the scope of the above disclosure. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0109] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0110] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.
Claims
1. A large model and agent evaluation method based on a multi-round conversation dataset, characterized by: The method comprises: Receiving and storing an evaluation data set, the evaluation data set comprising at least one data entry, each of the data entries comprising a dialogue group identifier, a dialogue turn, and at least one conversation data, the at least one conversation data comprising questions for interacting with an evaluation subject and reference answers to the questions, the evaluation subject comprising an intelligent agent or a large model; Determining a plurality of dialogue group data according to the evaluation data set, wherein one dialogue group data includes conversation data of a plurality of dialogue turns corresponding to the same dialogue group identifier; In response to receiving an evaluation instruction for the evaluation object, the evaluation object is evaluated using the plurality of conversation group data to obtain an evaluation result.
2. The method according to claim 1, characterized in that The data type of each session data is one of a plurality of preset data types; The method further comprises: In response to receiving a first instruction for viewing target conversation group data among the plurality of conversation group data, each target conversation data is displayed separately according to a data type of each target conversation data in the target conversation group data and a correspondence between a preset data type and a data display method.
3. The method according to claim 1, characterized in that The receiving and storing of the evaluation data set includes: receiving a second instruction for instructing batch uploading of the evaluation data set; In response to receiving the second instruction, receiving a data compression package, the data compression package including a description table and a plurality of session files, the description table including a plurality of rows of data, each row of data including a conversation group identifier, a conversation turn, a file path, and a data type; For each row of data, a target session file is determined from the multiple session files based on the file path included in the row of data, and the content of the target session file is stored as session data in association with the conversation group identifier and conversation turn included in the row of data as one of the data entries.
4. The method according to claim 1, wherein The receiving and storing of the evaluation data set includes: receiving a third instruction for instructing uploading the evaluation data sets one by one; In response to receiving the third instruction, displaying a first interface, wherein the first interface provides configuration items for configuring the data entry, the configuration items including a first configuration item for configuring a conversation group identifier, a second configuration item for configuring a conversation turn, and at least one third configuration item for configuring session data; In response to a configuration operation on the configuration item in the first interface, at least one data entry is obtained and stored.
5. The method according to claim 1, wherein The dialogue turn is a turn sequence number; Determining a plurality of dialogue group data according to the evaluation data set includes: For each dialog group identifier, target data entries corresponding to the dialog group identifier are obtained, and the target data entries are combined in ascending order of round numbers to obtain dialog group data corresponding to the dialog group identifier.
6. The method according to claim 2, characterized in that The method further comprises: receiving a fourth instruction for adding session data; In response to receiving the fourth instruction, for each dialogue group data: based on the target number of dialogue turns of the dialogue group data, a target number of input areas are displayed, each of the input areas is used to receive the dialogue data of one dialogue turn in the dialogue group data; in response to an input operation on the input area, the input content corresponding to the input operation is added to the dialogue group data as new dialogue data.
7. The method according to claim 2, characterized in that In response to receiving the first instruction for viewing target conversation group data among the plurality of conversation group data, displaying each target conversation data according to the data type of each target conversation data in the target conversation group data and the correspondence between the preset data type and the data display mode includes: In response to receiving the first instruction, for each target conversation data in the target dialogue group data, the target data display method corresponding to the data type of the target conversation data is determined according to the correspondence between the preset data type and the data display method, and the target conversation data is displayed according to the target data display method through the second interface.
8. The method according to claim 7, characterized in that The method further comprises: Displaying a preset component in a display area of each target session data, wherein the preset component is used to switch the target session data between multiple data display modes; In response to receiving a trigger instruction for a preset component in a target display area, determining a data display mode corresponding to the trigger instruction; The display mode of the target session data in the target display area is switched to a data display mode corresponding to the trigger instruction.
9. The method according to claim 2 or 7, characterized in that The preset data types include: text, rich text, picture, audio, video, PDF, CSV, and Word document.
10. The method according to any one of claims 2, 7 and 8, characterized in that The data display methods include plain text, rich text, and JSON.
11. A large model and agent evaluation device based on a multi-round conversation dataset, characterized in that: The device comprises: a first receiving module, configured to receive and store an evaluation dataset, the evaluation dataset comprising at least one data entry, each of the data entries comprising a dialogue group identifier, a dialogue turn, and at least one conversation data, the at least one conversation data comprising questions for interacting with an evaluation subject and reference answers to the questions, the evaluation subject comprising an agent or a large model; A first determining module is configured to determine a plurality of dialogue group data according to the evaluation data set, wherein each dialogue group data includes conversation data of a plurality of dialogue turns corresponding to the same dialogue group identifier; The evaluation module is configured to evaluate the evaluation object using the plurality of dialogue group data in response to receiving an evaluation instruction for the evaluation object to obtain an evaluation result.
12. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 10 are implemented.
13. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 10.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Dialogue evaluation method and device
CN119621536A
Intelligent agent evaluation method and device based on multiple rounds of dialogues, electronic equipment and medium
CN119829964A
Large model evaluation method, device, equipment, system and program product
CN120106210A
Goal segmentation in speech dialogs
US10236017B1
Cited By
Evaluation task processing method, device and equipment based on agent platform
CN122240277A
Agent platform-based evaluation task processing method and device, and equipment
CN122240277B