Intelligent agent interaction and evaluation method, system and related device

By acquiring multi-round agent dialogue content and evaluating query content using behavioral tags and historical round dialogue content, the problems of agent interaction realism and evaluation accuracy are solved, achieving agent interaction and evaluation with higher realism and accuracy.

CN122021874APending Publication Date: 2026-05-12ANHUI IFLYHEALTH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI IFLYHEALTH CO LTD
Filing Date
2025-12-22
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, it is difficult to collect highly realistic data that matches the actual scenario during the interaction process of intelligent agents, and the evaluation accuracy is low, resulting in inaccurate evaluation of the interaction effect.

Method used

By acquiring dialogue content from multiple rounds of intelligent agents, controlling the dialogue process using behavioral tags, and evaluating the inquiry content based on the dialogue content and behavioral intent of historical rounds, a deep integration of automated control and evaluation processes for the dialogue flow is achieved.

Benefits of technology

It improves the realism of intelligent agent interaction and the accuracy of interaction effect evaluation, ensures that the evaluation process matches the actual scenario, and enhances the reliability of the evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122021874A_ABST
    Figure CN122021874A_ABST
Patent Text Reader

Abstract

The invention discloses an agent interaction and evaluation method and system and a related device, and the method comprises the steps: obtaining the inquiry content outputted by a target agent and the reply content outputted by a reference agent according to a behavior label in a plurality of rounds, and obtaining a dialogue set; wherein each behavior tag corresponds to a respective dialogue mode, and the behavior tag of a single round is determined based on dialogue content of at least a part of historical rounds or is set as a preset behavior tag of a target question included in inquiry content of a current round; traversing each round in the dialogue set, obtaining a behavior intention corresponding to the behavior label of the previous round, and evaluating the inquiry content of the current round based on the dialogue content and the behavior intention of at least part of historical rounds to obtain a single round score of the inquiry content of the current round; wherein the single-round scores of all rounds are used for evaluating the target agent. According to the scheme, the authenticity of agent interaction and the accuracy of interaction effect evaluation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, system, and related apparatus for intelligent agent interaction and evaluation. Background Technology

[0002] With the development of artificial intelligence, intelligent agents have been more widely used. The training process of these agents requires a large number of interactive dialogue samples. However, collecting real-world interactive dialogues is difficult. Typically, two agents are used to play different roles in the dialogue, asking and answering questions to collect data. However, the interaction between agents is difficult to control effectively, resulting in the inability to collect data that closely matches real-world scenarios. Furthermore, the evaluation of agent performance often relies on customized rules, which are disconnected from the actual interaction data, leading to low evaluation accuracy. Therefore, improving the realism of agent interactions and the accuracy of interaction effect evaluation has become an urgent problem to be solved. Summary of the Invention

[0003] The main technical problem addressed by this application is to provide a method, system, and related apparatus for intelligent agent interaction and evaluation, which can improve the realism of intelligent agent interaction and the accuracy of interaction effect evaluation.

[0004] To address the aforementioned technical problems, a first aspect of this application provides a method for intelligent agent interaction and evaluation, comprising: acquiring query content output by a target intelligent agent and response content output by a reference intelligent agent according to behavior tags in multiple rounds to obtain a dialogue set; wherein, each behavior tag corresponds to its respective dialogue mode, and the behavior tag for a single round is determined based on the dialogue content of at least some historical rounds, or set as a preset behavior tag for the target question included in the query content of the current round; traversing each round in the dialogue set, acquiring the behavior intent corresponding to the behavior tag of the previous round, and evaluating the query content of the current round based on the dialogue content of at least some historical rounds and the behavior intent to obtain a single-round score for the query content of the current round; wherein, the single-round scores of all rounds are used to evaluate the target intelligent agent.

[0005] To address the aforementioned technical problems, a second aspect of this application provides an intelligent agent interaction and evaluation system, comprising: an acquisition module, configured to acquire query content output by a target intelligent agent and response content output by a reference intelligent agent according to behavior tags in multiple rounds, thereby obtaining a dialogue set; wherein each behavior tag corresponds to its respective dialogue mode, and the behavior tag for a single round is determined based on the dialogue content of at least some historical rounds, or set as a preset behavior tag for the target question included in the query content of the current round; and an execution module, configured to traverse each round in the dialogue set, acquire the behavior intent corresponding to the behavior tag of the previous round, and evaluate the query content of the current round based on the dialogue content of at least some historical rounds and the behavior intent, thereby obtaining a single-round score for the query content of the current round; wherein the single-round scores of all rounds are used to evaluate the target intelligent agent.

[0006] To address the aforementioned technical problems, a third aspect of this application provides an electronic device comprising: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor invokes the program data to execute the method described in the first aspect.

[0007] To address the aforementioned technical problems, a fourth aspect of this application provides a computer-readable storage medium storing program data thereon, wherein the program data, when executed by a processor, implements the method described in the first aspect.

[0008] The beneficial effects of this application are as follows: Unlike existing technologies, this application obtains a dialogue set during multiple rounds of dialogue between a target agent and a reference agent. The dialogue set includes the query content output by the target agent in each round and the response content output by the reference agent according to behavior labels. Each behavior label corresponds to a specific dialogue mode. The behavior label for a single round is determined based on the dialogue content of at least some historical rounds. Therefore, the generated dialogue content is used to generate behavior labels that control the response content output by the reference agent. The generated dialogue content is used to automatically control the dialogue mode in at least some rounds of the dialogue process. Alternatively, the behavior label for a single round is set based on the pre-set behavior labels for the target questions included in the current round's query content. By pre-setting behavior labels for the target questions included in the query content, the dialogue mode in the corresponding round for the target question can be directly controlled, thereby effectively controlling the dialogue flow, ensuring the matching degree between the dialogue flow and the actual scenario, and thus improving the realism of agent interaction. Iterate through each round in the dialogue set, obtain the behavioral intent corresponding to the behavioral label of the previous round, evaluate the query content of the current round based on the dialogue content of at least some historical rounds and the behavioral intent of the previous round, and obtain a single-round score for the query content of the current round. This ensures that the behavioral intent of the response content of the previous round can be understood when scoring each round, so that the evaluation process is deeply integrated with the entire dialogue process. In this way, the target agent is evaluated through single-round scores of all rounds, thereby improving the accuracy of interaction effect evaluation. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating one implementation method of the intelligent agent interaction and evaluation method of this application; Figure 2 This is a flowchart illustrating another implementation of the intelligent agent interaction and evaluation method of this application; Figure 3 This is a schematic diagram of one embodiment of the intelligent agent interaction and evaluation system of this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the electronic device of this application; Figure 5 This is a schematic diagram of one embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0010] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments, and different implementation methods can be adaptively combined. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0011] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper means two or more.

[0012] The intelligent agent interaction and evaluation method provided in this application is used to collect data during intelligent agent interaction and evaluate the interaction effect of the intelligent agent. The corresponding execution subject is a processing unit capable of data processing.

[0013] Please see Figure 1 , Figure 1 This is a flowchart illustrating one embodiment of the intelligent agent interaction and evaluation method of this application. The method includes: S101: Obtain the query content output by the target agent and the response content output by the reference agent according to the behavior label in multiple rounds to obtain a dialogue set; wherein, each behavior label corresponds to its own dialogue mode, and the behavior label of a single round is determined based on the dialogue content of at least some historical rounds, or is set as the behavior label preset by the target question included in the query content of the current round.

[0014] Specifically, the dialogue set between the target agent and the reference agent is obtained in multiple rounds. The dialogue set includes the query content output by the target agent in each round and the response content output by the reference agent according to the behavior label.

[0015] It should be noted that each behavior label corresponds to its own dialogue mode. The behavior label for a single round is determined based on the dialogue content of at least some historical rounds. Thus, the behavior label is generated through the generated dialogue content to control the response content output by the reference agent. The generated dialogue content is used to automatically control the dialogue mode of at least some rounds in the dialogue process. Alternatively, the behavior label for a single round is set according to the behavior label preset for the target question included in the current round's query content. By pre-setting behavior labels for the target question included in the query content, the dialogue mode in the corresponding round of the target question can be directly controlled, thereby effectively controlling the dialogue flow, ensuring the matching degree between the dialogue flow and the actual scenario, and thus improving the realism of the agent's interaction.

[0016] In one implementation, the target agent outputs query content based on the dialogue content of all historical rounds, and the reference agent estimates the dialogue mode of the current round based on the dialogue content of all historical rounds, generates a behavior label that matches the dialogue mode, outputs a response based on the behavior label and the dialogue content of all historical rounds, and obtains a dialogue set of the target agent and the reference agent when they have a dialogue in multiple rounds.

[0017] In one embodiment, the target agent outputs query content based on the dialogue content of the previous round and historical rounds with a correlation degree exceeding a correlation degree threshold. The query content of each round includes a target question selected from multiple candidate questions. Each candidate question is pre-set with a behavior label. The reference agent outputs response content based on the behavior label and at least some of the dialogue content of historical rounds, thereby obtaining a dialogue set when the target agent and the reference agent have a dialogue in multiple rounds.

[0018] In one embodiment, the target agent outputs query content based on dialogue content from at least some historical rounds. Each round's query content includes a target question selected from multiple candidate questions. Some candidate questions are pre-set with behavior tags. When the target question included in the query content matches a behavior tag, the reference agent outputs response content based on the behavior tag and dialogue content from at least the previous round. When the target question included in the query content does not match a behavior tag, the reference agent estimates the dialogue mode of the current round based on dialogue content from historical rounds whose relevance to the current round exceeds a relevance threshold, generates a behavior tag matching the dialogue mode, and outputs response content based on the behavior tag and dialogue content from at least the previous round, thereby obtaining a dialogue set of the target agent and the reference agent during dialogues across multiple rounds.

[0019] Optionally, the dialogue content of the reference agent and the target agent in the historical rounds applied at least includes the dialogue content of the previous round. In different implementation scenarios, it may be the dialogue content of the previous round, or the dialogue content of all historical rounds, or the dialogue content of the previous round and the historical rounds whose correlation with the previous round exceeds the correlation threshold, or the dialogue content of the previous round and up to a preset number of historical rounds before it.

[0020] It is understandable that by setting behavior labels during the dialogue process for the reference agent or by pre-setting behavior labels for some questions, the progress of the dialogue can be effectively controlled. Behavior labels in multiple rounds together form the dialogue behavior path of the dialogue process, enabling the target agent and the reference agent to conduct dialogue according to a dialogue behavior path that is more adapted to the actual scenario, resulting in a higher quality and more valuable dialogue set.

[0021] It should be noted that the dialogue process is usually triggered and terminated by the target agent. The first round directly outputs the query content. In some implementation scenarios, it can also be triggered by the reference agent and terminated by the target agent, or triggered and terminated by the reference agent.

[0022] In some implementation scenarios, the target intelligent agent includes a doctor intelligent agent, and the reference intelligent agent includes a patient intelligent agent. The doctor intelligent agent is used for medical diagnosis or follow-up visits. The dialogue process is triggered and terminated by the doctor intelligent agent. The dialogue set of the doctor intelligent agent and the patient intelligent agent during the multi-round interaction is obtained.

[0023] In some implementation scenarios, the target intelligent agent includes a teacher intelligent agent, and the reference intelligent agent includes a student intelligent agent. The teacher intelligent agent is used to evaluate teaching effectiveness or guide students to answer questions. The dialogue process is triggered and terminated by the student intelligent agent. The dialogue set of the teacher intelligent agent and the student intelligent agent during the multi-round interaction is obtained.

[0024] Optionally, the behavioral labels include one or more of the following: positive expression, indirect expression, rhetorical question, irrelevant answer, and abnormal expression. Different combinations correspond to different dialogue styles.

[0025] For ease of explanation, let's take the doctor's agent as the target agent and the patient's agent as the reference agent in a patient follow-up scenario as an example. Positive expression: The patient's agent directly answers the question posed by the doctor's agent. For example, "Doctor: Can you take care of yourself in daily life? Patient: Yes." That is, the patient directly answers the question. Indirect expression: The patient's agent does not directly answer the question posed by the doctor's agent, but instead answers it indirectly by replying with other content. For example, "Doctor: Can you take care of yourself in daily life? Patient: I was just discharged from the hospital, how can I take care of myself?" That is, the patient indirectly indicates that they were just discharged and cannot take care of themselves yet by asking a question in return. Irrelevant answer: The patient's agent replies with content unrelated to the question posed by the doctor's agent. For example, "Doctor: Can you take care of yourself in daily life? Patient: I'm taking my medication on time." That is, the patient replies with the irrelevant information of "taking medication on time." Asking a question in return: The patient does not answer the question posed by the doctor's agent, but instead asks a question that requires an answer from the doctor's agent. For example, "Doctor: Can you take care of yourself in daily life?" Patient: What does it mean to be able to take care of oneself? That is, the patient is unclear about the term "self-care" and asks a question to require further clarification from the doctor's AI agent. Abnormal Expression: The patient's intention is ambiguous or their expression is unclear in response to the question posed by the doctor's AI agent. For example, "Doctor: Can you take care of yourself in daily life? Patient: Seventeen." That is, the patient did not answer the question directly but mentioned "seventeen" without knowing what it refers to, indicating ambiguous intention. Furthermore, the above different behaviors can be superimposed to generate new behavioral labels, such as positive expression + questioning (i.e., the patient asks a question after making a positive statement), or indirect expression + abnormal expression (i.e., the patient expresses ambiguous or unclear content after making an indirect statement). Different combinations can be customized in specific scenarios, which will not be elaborated upon in this application.

[0026] S102: Traverse each round in the dialogue set, obtain the behavioral intent corresponding to the behavioral label of the previous round, evaluate the query content of the current round based on the dialogue content and behavioral intent of at least some historical rounds, and obtain the single-round score of the query content of the current round; wherein, the single-round scores of all rounds are used to evaluate the target agent.

[0027] Specifically, each round in the dialogue set is traversed to obtain the behavioral intent corresponding to the behavioral label of the previous round. Based on the dialogue content of at least some historical rounds and the behavioral intent of the previous round, the query content of the current round is evaluated to obtain a single-round score for the query content of the current round. This ensures that the behavioral intent of the response content of the previous round can be understood when scoring each round, so that the evaluation process is deeply integrated with the entire dialogue process. In this way, the target agent is evaluated through single-round scores of all rounds, thereby improving the accuracy of interaction effect evaluation.

[0028] In one embodiment, the behavioral intent corresponding to the behavioral tag of the previous round is obtained. Based on the dialogue content of at least some historical rounds and the behavioral intent of the previous round, the inquiry content of the current round is evaluated according to a preset scoring rule to obtain a single-round score of the inquiry content of the current round.

[0029] In some implementation scenarios, the scoring rules are related to the relevance of the current round's inquiry content to the behavioral intent of the previous round, as well as the redundancy and relevance of the current round's inquiry content to the dialogue content of historical rounds.

[0030] In one implementation, the behavioral intent corresponding to the behavioral label of the previous round is obtained. At least a portion of the dialogue content from previous rounds, the behavioral intent of the previous round, and the response content of the current round are input into the evaluation model to obtain a single-round score of the current round's query content output by the evaluation model. The evaluation model is trained using multi-round dialogues labeled with the behavioral intent of each round.

[0031] In some implementation scenarios, the evaluation model is a large language model, which is obtained by fine-tuning multiple rounds of dialogue labeled with the behavioral intentions of each round.

[0032] It is understandable that the single-round scores of all rounds can be used to evaluate the interaction effect of the target intelligent agent, and the single-round scores are highly consistent with the behavioral intentions and content in the actual dialogue process. Therefore, based on the single-round scores of all rounds, the overall interaction effect of the target intelligent agent can be effectively evaluated. Moreover, the automated evaluation scheme can be adapted to the needs of different application scenarios, realize cross-scenario application migration, and ensure the accuracy and scenario adaptability of the evaluation results.

[0033] Optionally, the target intelligent agent includes a doctor intelligent agent, and the reference intelligent agent includes a patient intelligent agent. This allows the patient's dialogue behavior path to be constructed according to behavioral labels in multiple rounds. Based on the dialogue behavior path, the patient intelligent agent can accurately reproduce real patient behavior in different scenarios (such as consultation response and symptom description), possessing the core capability to simulate diverse and highly realistic patient behaviors. Furthermore, after evaluation, the doctor intelligent agent can accurately determine its effectiveness in real-world scenarios, facilitating decisions on whether to deploy it or continue optimization.

[0034] The above scheme acquires a dialogue set of the target agent and reference agent during multiple rounds of dialogue. The dialogue set includes the query content output by the target agent in each round and the response content output by the reference agent according to behavior tags. Each behavior tag corresponds to a specific dialogue mode. The behavior tag for a single round is determined based on the dialogue content of at least some historical rounds. Therefore, the generated dialogue content is used to generate behavior tags to control the response content output by the reference agent. This automatically controls the dialogue mode in at least some rounds of the dialogue process using the generated dialogue content. Alternatively, the behavior tag for a single round is set based on the pre-set behavior tags for the target questions included in the current round's query content. By pre-setting behavior tags for the target questions included in the query content, the dialogue mode in the corresponding round for the target question is directly controlled, thereby effectively controlling the dialogue flow, ensuring the matching degree between the dialogue flow and the actual scenario, and improving the realism of the agent interaction. Iterate through each round in the dialogue set, obtain the behavioral intent corresponding to the behavioral label of the previous round, evaluate the query content of the current round based on the dialogue content of at least some historical rounds and the behavioral intent of the previous round, and obtain a single-round score for the query content of the current round. This ensures that the behavioral intent of the response content of the previous round can be understood when scoring each round, so that the evaluation process is deeply integrated with the entire dialogue process. In this way, the target agent is evaluated through single-round scores of all rounds, thereby improving the accuracy of interaction effect evaluation.

[0035] Please see Figure 2 , Figure 2 This is a flowchart illustrating another embodiment of the intelligent agent interaction and evaluation method of this application, the method comprising: S201: Obtain the query content output by the target agent and the response content output by the reference agent according to the behavior label in multiple rounds to obtain a dialogue set; wherein, each behavior label corresponds to its own dialogue mode, and the behavior label of a single round is determined based on the dialogue content of at least some historical rounds, or set as the behavior label preset by the target question included in the query content of the current round.

[0036] Specifically, a dialogue set is obtained when the target agent and the reference agent engage in dialogue across multiple rounds. The dialogue set includes the query content output by the target agent in each round and the response content output by the reference agent according to behavior labels. Each behavior label corresponds to a specific dialogue mode. The behavior label for a single round can be determined by the reference agent based on the dialogue content of at least some historical rounds, or set as a pre-defined behavior label for the target question included in the query content of the current round.

[0037] It should be noted that the reference agent is pre-trained and can predict the dialogue mode of the current round using at least some of the dialogue content from previous rounds, filter out the behavior labels for a single round, and output the response content for the current round based on the behavior labels and at least some of the dialogue content from previous rounds.

[0038] Specifically, the generation of response content by the reference agent is divided into two stages. In the first stage, the dialogue mode of the current round is estimated by using at least some of the dialogue content from previous rounds, and the behavior label of a single round is selected from multiple behavior labels to control the subsequent output response content. In the second stage, the response content of the current round is output according to the instructions of the behavior label and in combination with at least some of the dialogue content from previous rounds, ensuring the relevance of the response content to the generated dialogue content and the matching degree with the behavior label of the current round.

[0039] Optionally, behavioral tags include one or more of the following: positive expression, indirect expression, rhetorical question, irrelevant answer, and abnormal expression. Different combinations correspond to different dialogue styles. Behavioral tags may also include direct questioning, which can be set according to the actual scenario.

[0040] It should be noted that the reference agent is pre-trained and can estimate the probability of each dialogue method based on the dialogue content of at least some historical rounds, obtain the behavior label matching the dialogue method with the highest probability, and generate the response content for the current round to respond to the query content according to the instructions of the behavior label and in combination with the dialogue content of at least some historical rounds.

[0041] Furthermore, the target agent can output the query content for the next round based on the dialogue content of at least some historical rounds, and the query content includes the target question selected from multiple candidate questions, or, based on the dialogue content of at least some historical rounds, choose to terminate the dialogue; wherein, at least some candidate questions are pre-labeled with behavioral tags.

[0042] Specifically, the target agent can output further dialogue queries based on at least a portion of the dialogue content from previous rounds. These queries include a target question selected from multiple candidate questions. In other words, the target agent can make further queries based on the already conducted dialogue, ensuring the queries are context-sensitive. Furthermore, the target agent can determine whether to terminate the dialogue based on the collected necessary content from at least a portion of the dialogue content from previous rounds, thereby effectively controlling the number of dialogue rounds.

[0043] It should be noted that at least some candidate questions have pre-set behavioral labels, so by setting behavioral labels for candidate questions, the behavioral labels of at least some rounds can be directly intervened to achieve direct control of the dialogue process.

[0044] It is understandable that the dialogue content of the reference agent and the target agent in the historical rounds should include at least the dialogue content of the previous round. In different implementation scenarios, it may be the dialogue content of the previous round, or the dialogue content of all historical rounds, or the dialogue content of the previous round and the historical rounds whose correlation with the previous round exceeds the correlation threshold, or the dialogue content of the previous round and at most a preset number of historical rounds before it.

[0045] It should be noted that at least some behavior labels include corresponding sub-labels, which are used to adjust the output ratio of responses matching the candidate question in a single round. These sub-labels are selected by the reference agent based on the dialogue content of at least some historical rounds, or, at least some candidate questions are pre-configured with sub-labels based on their pre-defined behavior labels.

[0046] Specifically, the candidate question matching has complete response content, and at least some behavior labels include multiple corresponding sub-labels. The sub-labels are used to adjust the output ratio of the response content that matches the candidate question in a single round, so as to generate richer dialogue behavior paths by using the sub-labels, so as to simulate different levels of difficulty in obtaining effective content and evaluate the interaction effect of the target intelligent agent in different scenarios.

[0047] Furthermore, the sub-labels are selected by the reference agent based on the dialogue content of at least some historical rounds. The reference agent is pre-trained to select a sub-label when the behavior label is configured with sub-labels, or at least some candidate questions are pre-configured with sub-labels for the behavior labels. Thus, by selecting the sub-labels configured with the behavior labels in advance by the reference agent, the output ratio of the response content can be effectively controlled.

[0048] In some implementation scenarios, when a behavior label includes either positive or indirect expression, it may contain multiple sub-labels, with different sub-labels controlling the proportion of the response content output during the expression. Furthermore, in different implementation scenarios, corresponding sub-labels can be set for other behavior labels; this application does not impose specific restrictions on this.

[0049] For ease of understanding, taking a medical diagnosis scenario as an example, the sub-tags include information hiding and detailed response. Under the information hiding sub-tag, the output ratio of symptom description information is set to 30%-50%, while under the detailed response sub-tag, this output ratio is increased to over 80%. Ultimately, this achieves dynamic and controllable patient response generation and generates more detailed dialogue behavior paths. The specific sub-tags and their corresponding output ratios can be customized in specific scenarios, and this application does not impose specific restrictions on them.

[0050] It should be noted that the candidate questions with behavioral labels, and the pre-set behavioral labels for these candidate questions, are adjusted by the target entity during the evaluation process. In other words, the target entity can adjust which candidate questions have behavioral labels, and the specific behavioral labels set for those candidate questions, during the evaluation process. This allows for continuous optimization and intervention in the dialogue behavior path, ensuring the controllability of the response content generated by the reference agent and the comprehensiveness of the target agent's evaluation.

[0051] S202: Traverse each round in the dialogue set, obtain the behavioral intent corresponding to the behavioral label of the previous round, evaluate the query content of the current round based on the dialogue content and behavioral intent of at least some historical rounds, and obtain the single-round score of the query content of the current round; wherein, the single-round scores of all rounds are used to evaluate the target agent.

[0052] Specifically, each round in the dialogue set is traversed to obtain the behavioral intent corresponding to the behavioral label of the previous round. Based on the dialogue content of at least some historical rounds and the behavioral intent of the previous round, the query content of the current round is evaluated to obtain a single-round score of the query content of the current round.

[0053] In one embodiment, the behavioral intent corresponding to the behavioral tag of the previous round is obtained, and the inquiry content of the current round is evaluated based on the dialogue content and behavioral intent of at least some historical rounds to obtain a single-round score for the inquiry content of the current round. This includes: obtaining the behavioral intent corresponding to the behavioral tag of the previous round; evaluating the inquiry content of the current round from multiple scoring dimensions based on the dialogue content and behavioral intent of at least some historical rounds to obtain a single-dimensional score for each scoring dimension; wherein each scoring dimension corresponds to a preset scoring point; and obtaining a single-round score for the inquiry content of the current round based on the single-dimensional scores of all scoring dimensions.

[0054] Specifically, the behavioral intent corresponding to the behavioral tag of the previous round is obtained. Based on the dialogue content of at least some historical rounds and the behavioral intent of the previous round, the inquiry content of the current round is evaluated from multiple pre-set scoring dimensions to determine the matching degree between the inquiry content and the scoring points of each scoring dimension, thereby obtaining the single-dimensional score of each scoring dimension and ensuring the accuracy of the single-dimensional score.

[0055] Furthermore, by merging the single-dimensional scores of all scoring dimensions, a single-round score for the current round of inquiries is obtained, thereby improving the accuracy of single-round scoring.

[0056] Optionally, multiple scoring dimensions include inquiry reasonableness, fluency of expression, and content redundancy rate. Inquiry reasonableness corresponds to the relevance of the current round's inquiry content to the behavioral intent of the previous round; fluency of expression corresponds to the relevance of the current round's inquiry content to the dialogue content of previous rounds; and content redundancy rate corresponds to the redundancy of the current round's inquiry content to the dialogue content of previous rounds. In different implementation scenarios, different scoring dimensions and their corresponding scoring points can be dynamically set via commands; this application does not impose specific limitations on this.

[0057] S203: Based on the multi-turn query content in multiple dialogue sets and the recall rate of the multi-turn response content relative to the content to be recalled, determine the overall reward of each word element in the query content within the dialogue set, and based on the single-turn score of the query content in each turn in the dialogue set, determine the single-turn reward of each word element in the query content of the corresponding turn.

[0058] Specifically, each dialogue set is matched with its own dialogue behavior path. Within the same dialogue behavior path, more than a preset number or a preset proportion of rounds share the same behavior label. Each dialogue behavior path includes multiple dialogue sets, and the reference agent is configured with content to be recalled. Specifically, each dialogue set is matched with its own dialogue behavior path. When more than a preset number or a preset proportion of rounds within a dialogue behavior path share the same behavior label, they are determined to be the same dialogue behavior path. Each dialogue behavior path includes multiple dialogue sets, and the reference agent matches content to be recalled. The content to be recalled is related to at least some of the responses to the target question.

[0059] Understandably, based on multi-turn query content from multiple dialogue sets and the recall rate of multi-turn responses relative to the content to be recalled, a unified overall reward is determined for each token in the query content across the entire dialogue set. This overall reward provides feedback on the overall performance of the target agent during the dialogue process. Based on the single-turn score of the query content in each round of the dialogue set, a corresponding single-turn reward is assigned to each token in the query content of that round. This single-turn reward provides feedback on the effectiveness of the tokens in the query content during a single round of interaction. Here, a token is also known as a Token.

[0060] It should be noted that the overall reward is matched with the entire dialogue set, and the same overall reward is set for each word in the query content within the dialogue set.

[0061] In one embodiment, each dialogue set corresponds to a dialogue reward evaluation rule. The dialogue reward evaluation rule is related to the content correlation and content redundancy rate between the multi-turn query content in the dialogue set, as well as the recall rate of the multi-turn response content relative to the content to be recalled. The dialogue reward evaluation rule is used to evaluate each dialogue set and obtain the dialogue reward for each dialogue set. The average reward corresponding to the dialogue reward of all dialogue sets under the dialogue behavior path is obtained. Based on the difference between the dialogue reward of each dialogue set and the average reward, the overall reward of each word in the query content within each dialogue set is determined.

[0062] It should be noted that the dialogue reward obtained using the dialogue reward evaluation rule is positively correlated with the content relevance between multi-turn queries, negatively correlated with the content redundancy rate between multi-turn queries, and positively correlated with the recall rate of multi-turn responses compared to the content to be recalled. The overall reward is obtained based on Generalized Regularized Policy Optimization (GRPO). The average reward of all dialogue sets under the dialogue behavior path can be used as the dialogue comparison reward. The overall reward of each word in the query content within each dialogue set is the relative result obtained by comparing the dialogue reward with the dialogue comparison reward.

[0063] In one implementation, the dialogue set matching has an objective evaluation dimension and a subjective evaluation dimension. The reward for the objective evaluation dimension is determined based on the total number of dialogue rounds in the dialogue set and the recall rate of multi-round response content compared with the content to be recalled. The reward for the subjective evaluation dimension is determined based on the interaction fluency between multi-round query content in the dialogue set. The rewards of the two dimensions are combined to obtain the overall reward for each word element in the query content within each dialogue set.

[0064] Understandably, recall corresponds to the ratio of the number of hits of the content to be recalled included in the multi-turn responses to the total number of content to be recalled. Recall corresponds to a base reward, while the total number of dialogue turns corresponds to additional rewards and penalties. The reward for the objective evaluation dimension is the sum of the base reward and the additional reward / penalty value. When the total number of turns exceeds the total number of content to be recalled, an additional penalty is set; when the total number of turns is less than or equal to the total number of content to be recalled, an additional reward is set, thereby aiming to obtain more effective information with fewer dialogue turns. Furthermore, the subjective evaluation dimension can be evaluated using a pre-trained evaluation model.

[0065] It should be noted that the single-round reward is matched with the word units in the single-round query content in the dialogue set, and the single-round score of each round is converted into the single-round reward for each word unit in the corresponding query content.

[0066] In some implementation scenarios, the overall reward for each word in the query content within a dialogue set is determined based on the multi-turn query content in multiple dialogue sets and the recall rate of the multi-turn response content relative to the content to be recalled. This includes: determining the dialogue reward for each dialogue set under the dialogue behavior path based on the total number of dialogue turns in each dialogue set under the dialogue behavior path, the content redundancy rate of the multi-turn query content, and the recall rate of the multi-turn response content relative to the content to be recalled; determining the dialogue comparison reward for all dialogue sets under the dialogue behavior path based on the dialogue reward for each dialogue set under the dialogue behavior path; and determining the overall reward for each word in the query content within a dialogue set based on the dialogue reward and the dialogue comparison reward of the dialogue set.

[0067] Specifically, for each dialogue set under the dialogue behavior path, based on the total number of dialogue rounds in the dialogue set, the content redundancy rate of multi-round query content, and the recall rate of multi-round response content compared to the content to be recalled, the total number of dialogue rounds and the recall rate can objectively evaluate the dialogue set, while the content redundancy rate of multi-round query content can subjectively evaluate the content of the dialogue set, thereby determining the high-precision dialogue reward for each dialogue set under the dialogue behavior path.

[0068] Furthermore, based on the dialogue reward of each dialogue set under the dialogue behavior path, the dialogue comparison reward corresponding to all dialogue sets under the dialogue behavior path is determined. The dialogue comparison reward is usually the average of the dialogue rewards of all dialogue sets under the dialogue behavior path, thus obtaining the dialogue comparison reward at a reduced cost. The dialogue reward of the dialogue set and the dialogue comparison reward are compared to determine the overall reward of each word in the query content within the dialogue set, so that the overall reward can accurately reflect whether each dialogue set as a whole has an advantage over other dialogue sets.

[0069] S204: Based on the overall reward and single-round reward for each word, determine the target reward for each word and adjust the target agent using the target reward; wherein, the adjusted target agent is used to communicate with the reference agent or with the target object.

[0070] Specifically, based on the overall reward and single-round reward of each word, the target reward of each word is determined from two dimensions. The target agent is adjusted using the target reward so that the adjusted target agent can adapt to different dialogue methods and guide the dialogue object to output more accurate content, thereby increasing the probability of obtaining the required information.

[0071] It should be noted that the adjusted target agent is used to engage in dialogue with the reference agent to collect a larger and more precise set of dialogue data for further agent adjustments, or to engage in dialogue with the target object to collect more accurate information.

[0072] In one implementation, the overall reward and single-round reward for each word are superimposed to obtain the target reward for each word, and the target agent is adjusted using the target reward.

[0073] In one implementation, the overall reward and single-round reward for each word are weighted and summed to obtain the target reward for each word, and the target agent is adjusted using the target reward.

[0074] Optionally, the target intelligent agent includes a doctor intelligent agent, and the reference intelligent agent includes a patient intelligent agent. This allows the patient's dialogue behavior path to be constructed according to behavior labels in multiple rounds. Based on the dialogue behavior path, the patient intelligent agent can generate dialogue sets in different scenarios and simulate different levels of difficulty in obtaining effective content. The doctor intelligent agent can then be evaluated and optimized to obtain a doctor intelligent agent that can ultimately have direct dialogue with the target object, thereby improving the accuracy of the doctor intelligent agent's interaction with the target object.

[0075] Please see Figure 3 , Figure 3 This is a schematic diagram of an embodiment of the intelligent agent interaction and evaluation system of this application. The intelligent agent interaction and evaluation system 30 includes an acquisition module 301 and an execution module 302. The acquisition module 301 is used to acquire the query content output by the target intelligent agent in multiple rounds and the response content output by the reference intelligent agent according to the behavior label to obtain a dialogue set. Each behavior label corresponds to its own dialogue mode. The behavior label of a single round is determined based on the dialogue content of at least some historical rounds, or set as the behavior label preset by the target question included in the query content of the current round. The execution module 302 is used to traverse each round in the dialogue set, acquire the behavior intention corresponding to the behavior label of the previous round, and evaluate the query content of the current round based on the dialogue content and behavior intention of at least some historical rounds to obtain a single-round score of the query content of the current round. The single-round scores of all rounds are used to evaluate the target intelligent agent.

[0076] In one embodiment, the reference agent is pre-trained to predict the dialogue mode of the current round using at least some of the dialogue content from previous rounds, filter out behavior labels for a single round, and output the response content for the current round based on the behavior labels and at least some of the dialogue content from previous rounds; the target agent is able to output the query content for the next round based on at least some of the dialogue content from previous rounds, and the query content includes a target question selected from multiple candidate questions, or, based on at least some of the dialogue content from previous rounds, choose to terminate the dialogue; wherein, at least some of the candidate questions are pre-labeled with behavior labels.

[0077] In one embodiment, at least some behavior labels include corresponding sub-labels, which are used to adjust the output ratio of responses matching the candidate question in a single round; wherein, the sub-labels are selected by the reference agent based on the dialogue content of at least some historical rounds, or, at least some candidate questions are configured with sub-labels in their preset behavior labels.

[0078] In one embodiment, candidate questions with behavioral labels and the preset behavioral labels of the candidate questions are adjusted by the target object during the evaluation process.

[0079] In one embodiment, the execution module 302 is further configured to obtain the behavioral intent corresponding to the behavioral tag of the previous round, and evaluate the query content of the current round from multiple scoring dimensions based on the dialogue content and behavioral intent of at least some historical rounds, so as to obtain a single-dimensional score for each scoring dimension; wherein, each scoring dimension corresponds to a preset scoring point; and a single-round score for the query content of the current round is obtained based on the single-dimensional scores of all scoring dimensions.

[0080] In one embodiment, each dialogue set is matched with its own dialogue behavior path. More than a preset number or a preset proportion of rounds within the same dialogue behavior path have the same behavior label. Each dialogue behavior path includes multiple dialogue sets, and the reference agent is configured with content to be recalled. The execution module 302 is further configured to determine the overall reward for each word element in the query content within the dialogue set based on the multi-round query content in the multiple dialogue sets and the recall rate of the multi-round response content relative to the content to be recalled; determine the single-round reward for each word element in the query content of each round based on the single-round score of the query content in each round of the dialogue set; determine the target reward for each word element based on the overall reward and single-round reward of each word element; and adjust the target agent using the target reward. The adjusted target agent is used to converse with the reference agent or with the target object.

[0081] In one embodiment, the execution module 302 is further configured to determine the dialogue reward for each dialogue set under the dialogue behavior path based on the total number of dialogue rounds in each dialogue set under the dialogue behavior path, the content redundancy rate of multi-round query content, and the recall rate of multi-round response content relative to the content to be recalled; determine the dialogue comparison reward corresponding to all dialogue sets under the dialogue behavior path based on the dialogue reward for each dialogue set under the dialogue behavior path; and determine the overall reward for each word in the query content within the dialogue set based on the dialogue reward and dialogue comparison reward of the dialogue set.

[0082] In one embodiment, the target agent includes a doctor agent, and the reference agent includes a patient agent.

[0083] Please see Figure 4 , Figure 4This is a schematic diagram of the structure of an embodiment of the electronic device of this application. The electronic device 40 includes a memory 401 and a processor 402 coupled to each other. The memory 401 stores program data (not shown in the figure). The processor 402 calls the program data to implement the method in any of the above embodiments. For the description of the relevant content, please refer to the detailed description of the above method embodiments, which will not be repeated here.

[0084] Please see Figure 5 , Figure 5 This is a schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The computer-readable storage medium 50 stores program data 500. When the program data 500 is executed by a processor, it implements the method in any of the above embodiments. For related descriptions, please refer to the detailed description of the above method embodiments, which will not be repeated here.

[0085] It should be noted that the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0086] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0087] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0088] The above description is merely an embodiment of this application and does not limit the scope of protection of this application. Any equivalent structural or procedural transformations made based on the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of protection of this application.

Claims

1. A method for intelligent agent interaction and evaluation, characterized in that, include: The query content output by the target agent and the response content output by the reference agent according to the behavior label are obtained in multiple rounds to obtain a dialogue set; wherein, each behavior label corresponds to its own dialogue mode, and the behavior label of a single round is determined based on the dialogue content of at least some historical rounds, or is set as the behavior label preset by the target question included in the query content of the current round. Each round in the dialogue set is traversed to obtain the behavioral intent corresponding to the behavioral label of the previous round. Based on the dialogue content of at least some historical rounds and the behavioral intent, the query content of the current round is evaluated to obtain a single-round score of the query content of the current round. The single-round scores of all rounds are used to evaluate the target agent.

2. The intelligent agent interaction and evaluation method according to claim 1, characterized in that, The reference agent is pre-trained and can predict the dialogue mode of the current round using at least some of the dialogue content of the previous rounds, filter out the behavior labels of a single round, and output the response content of the current round based on the behavior labels and at least some of the dialogue content of the previous rounds. The target agent can output the query content for the next round based on the dialogue content of at least some historical rounds, and the query content includes a target question selected from multiple candidate questions, or, based on the dialogue content of at least some historical rounds, choose to terminate the dialogue; wherein, at least some of the candidate questions are preset with behavioral labels.

3. The intelligent agent interaction and evaluation method according to claim 2, characterized in that, The at least part of the behavior tags include corresponding sub-tags, which are used to adjust the output ratio of response content that matches the candidate question in a single round; The sub-labels are selected by the reference agent based on the dialogue content of at least some historical rounds, or the sub-labels are configured with the pre-set behavior labels of at least some of the candidate questions.

4. The intelligent agent interaction and evaluation method according to claim 2, characterized in that, The candidate questions with the aforementioned behavioral labels, and the pre-defined behavioral labels for the candidate questions, are adjusted by the target object during the evaluation process.

5. The intelligent agent interaction and evaluation method according to claim 1, characterized in that, The step of obtaining the behavioral intent corresponding to the behavioral tag from the previous round, and evaluating the inquiry content of the current round based on at least some historical rounds of dialogue content and the behavioral intent, to obtain a single-round score for the inquiry content of the current round, includes: Obtain the behavioral intent corresponding to the behavioral tag in the previous round. Based on the dialogue content of at least some historical rounds and the behavioral intent, evaluate the query content of the current round from multiple scoring dimensions to obtain a single-dimensional score for each scoring dimension. Each scoring dimension has a preset scoring point. Based on the single-dimensional scores of all the aforementioned scoring dimensions, a single-round score for the query content in the current round is obtained.

6. The intelligent agent interaction and evaluation method according to claim 1, characterized in that, Each dialogue set is matched with its own dialogue behavior path. More than a preset number or a preset proportion of rounds in the same dialogue behavior path have the same behavior label. Each dialogue behavior path includes multiple dialogue sets. The reference agent is configured with content to be recalled. The process involves traversing each round of the dialogue set, obtaining the behavioral intent corresponding to the behavioral label of the previous round, evaluating the query content of the current round based on at least some historical rounds' dialogue content and the behavioral intent, and obtaining a single-round score for the query content of the current round. After using the single-round scores of all rounds to evaluate the target agent, the process further includes: Based on the multiple rounds of query content in the multiple dialogue sets, and the recall rate of the multiple rounds of response content compared with the content to be recalled, the overall reward of each word element in the query content within the dialogue set is determined, and based on the single-round score of the query content in each round in the dialogue set, the single-round reward of each word element in the query content of the corresponding round is determined. Based on the overall reward and the single-round reward for each word, a target reward for each word is determined, and the target agent is adjusted using the target reward; wherein, the adjusted target agent is used to communicate with the reference agent or with the target object.

7. The intelligent agent interaction and evaluation method according to claim 6, characterized in that, The determination of the overall reward for each term in the query content within the dialogue set, based on the multi-turn query content from multiple dialogue sets and the recall rate of the multi-turn response content relative to the content to be recalled, includes: Based on the total number of dialogue rounds in each dialogue set under the dialogue behavior path, the content redundancy rate of the multi-round query content, and the recall rate of the multi-round response content compared with the content to be recalled, the dialogue reward for each dialogue set under the dialogue behavior path is determined. Based on the dialogue reward of each dialogue set under the dialogue behavior path, determine the dialogue comparison reward corresponding to all dialogue sets under the dialogue behavior path; Based on the dialogue reward and the dialogue comparison reward of the dialogue set, the overall reward for each word in the query content within the dialogue set is determined.

8. The intelligent agent interaction and evaluation method according to any one of claims 1-7, characterized in that, The target intelligent agent includes a doctor intelligent agent, and the reference intelligent agent includes a patient intelligent agent.

9. A smart agent interaction and evaluation system, characterized in that, include: The acquisition module is used to acquire the query content output by the target agent and the response content output by the reference agent according to the behavior label in multiple rounds to obtain a dialogue set; wherein, each behavior label corresponds to its own dialogue mode, and the behavior label of a single round is determined based on the dialogue content of at least some historical rounds, or is set as the behavior label preset by the target question included in the query content of the current round. An execution module is used to traverse each round in the dialogue set, obtain the behavioral intent corresponding to the behavioral label of the previous round, evaluate the query content of the current round based on the dialogue content of at least some historical rounds and the behavioral intent, and obtain a single-round score of the query content of the current round; wherein, the single-round scores of all rounds are used to evaluate the target agent.

10. An electronic device, characterized in that, include: A memory and a processor are coupled to each other, wherein the memory stores program data, and the processor invokes the program data to perform the method as described in any one of claims 1-8.

11. A computer-readable storage medium storing program data thereon, characterized in that, When the program data is executed by the processor, it implements the method as described in any one of claims 1-8.