A multi-agent system, a method for generating dialogue data, an apparatus, and a storage medium.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-08-14
AI Technical Summary
目前,面向儿童的情感陪伴与情绪对话的相关方案存在对话数据缺乏多样性、质量差等技术缺陷因此,如何解决上述技术缺陷已成为本领域技术人员亟待解决的技术问题
[0040]本申请所提供的对话数据生成方法,包括:通过情绪场景生成智能体根据情绪与主题生成场景;通过用户画像生成智能体生成儿童用户画像;通过儿童对话智能体根据所述场景与所述儿童用户画像生成儿童对话;通过陪伴者对话智能体根据所述场景、所述儿童用户画像、所述儿童对话生成陪伴对话;通过对话审查智能体根据所述儿童对话与所述陪伴对话对目标维度进行评分,根据所述评分确定样本,根据所述样本优化所述儿童对话智能体与所述陪伴者对话智能体。
Smart Images

Figure CN122575363A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a multi-agent system, a method for generating dialogue data, an apparatus, and a storage medium. Background Technology
[0002] Emotional companionship and dialogue for children utilize language they can understand to guide them in recognizing various emotions such as happiness, hurt, sadness, and anger, and encourages them to express their feelings. Through interactive dialogue, it soothes children's anxieties and low moods, helping them learn to manage their emotions and express their needs appropriately. Currently, existing solutions for emotional companionship and dialogue for children suffer from technical deficiencies such as a lack of diversity and poor quality of dialogue data. Therefore, addressing these technical shortcomings has become a pressing technical problem for those skilled in the art. Summary of the Invention
[0003] The purpose of this application is to provide a multi-agent system, a method for generating dialogue data, an apparatus, and a storage medium that can improve the diversity and quality of dialogue data.
[0004] To address the aforementioned technical problems, this application provides a multi-agent system, comprising:
[0005] The system includes an intelligent agent for generating emotional scenarios, an intelligent agent for generating user profiles, an intelligent agent for child dialogue, an intelligent agent for companion dialogue, and an intelligent agent for dialogue review. The intelligent agent for generating emotional scenarios, the intelligent agent for generating user profiles, and the intelligent agent for dialogue review are all connected to the intelligent agent for child dialogue and the intelligent agent for companion dialogue, respectively. The intelligent agent for child dialogue is connected to the intelligent agent for companion dialogue.
[0006] The emotional scene generation agent is used to generate scenes based on emotions and themes, and output the scenes to the child dialogue agent and the companion dialogue agent.
[0007] The user profile generation agent is used to generate a child user profile and output the child user profile to the child dialogue agent and the companion dialogue agent.
[0008] The child dialogue agent is used to generate a child dialogue based on the scenario and the child user profile, and output the child dialogue to the companion dialogue agent and the dialogue review agent.
[0009] The companion dialogue agent is used to generate companion dialogue based on the scenario, the child user profile, and the child's dialogue, and output the companion dialogue to the dialogue review agent.
[0010] The dialogue review agent is used to score the target dimension based on the child's dialogue and the companion's dialogue, and to select samples based on the scores to optimize the child dialogue agent and the companion dialogue agent.
[0011] In some embodiments, scoring the target dimension based on the child's dialogue and the companion's dialogue includes:
[0012] The dialogue between the child and the companion dialogue are scored based on the dimensions of dialogue strategy, language naturalness, and logical depth.
[0013] In some embodiments, the step of selecting samples based on scores to optimize the child's dialogue agent and the companion's dialogue agent includes:
[0014] If any of the dimensions of dialogue strategy, language naturalness, or logic depth fails to meet the requirements, positive and negative sample pairs are selected to optimize the child dialogue agent and the companion dialogue agent.
[0015] In some embodiments, the step of selecting samples based on scores to optimize the child dialogue agent and the companion dialogue agent further includes:
[0016] If the scores for the dialogue strategy dimension, the language naturalness dimension, and the logical depth dimension are all satisfactory, then the child's dialogue and the companion dialogue are saved.
[0017] In some embodiments, the step of selecting samples based on scores to optimize the child's dialogue agent and the companion's dialogue agent includes:
[0018] Determine the direct preference optimization loss based on the sample;
[0019] The child dialogue agent and the companion dialogue agent are optimized based on the direct preference optimization loss.
[0020] In some embodiments, the step of outputting the child's dialogue to the companion dialogue agent and the dialogue review agent includes:
[0021] The companion dialogue agent selects a dialogue strategy based on the child's emotions and the scenario.
[0022] In some embodiments, the emotional scene generation agent, the user profile generation agent, the child dialogue agent, the companion dialogue agent, and the dialogue review agent are built on the Dify platform.
[0023] To address the aforementioned technical problems, this application also provides a dialogue data generation method, applied to the multi-agent system described above, comprising:
[0024] Intelligent agents are generated based on emotions and themes to create scenarios.
[0025] Generate child user profiles using intelligent agents generated from user profiles;
[0026] A child dialogue agent generates a child dialogue based on the scenario and the child user profile.
[0027] The companion dialogue agent generates a companion dialogue based on the scenario, the child user profile, and the child's dialogue.
[0028] The dialogue review agent scores the target dimensions based on the child's dialogue and the companion's dialogue, and selects samples based on the scores to optimize the child dialogue agent and the companion dialogue agent.
[0029] To address the aforementioned technical problems, this application also provides a computer device, comprising:
[0030] Memory, used to store computer programs;
[0031] A processor for executing the computer program to implement the steps of the dialogue data generation method as described above.
[0032] To address the aforementioned technical problems, this application also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the steps of the dialogue data generation method described above.
[0033] The multi-agent system provided in this application includes: an emotional scene generation agent, a user profile generation agent, a child dialogue agent, a companion dialogue agent, and a dialogue review agent; the emotional scene generation agent, the user profile generation agent, and the dialogue review agent are all connected to the child dialogue agent and the companion dialogue agent, respectively; the child dialogue agent is connected to the companion dialogue agent.
[0034] The emotional scene generation agent is used to generate scenes based on emotions and themes, and output the scenes to the child dialogue agent and the companion dialogue agent.
[0035] The user profile generation agent is used to generate a child user profile and output the child user profile to the child dialogue agent and the companion dialogue agent.
[0036] The child dialogue agent is used to generate a child dialogue based on the scenario and the child user profile, and output the child dialogue to the companion dialogue agent and the dialogue review agent.
[0037] The companion dialogue agent is used to generate companion dialogue based on the scenario, the child user profile, and the child's dialogue, and output the companion dialogue to the dialogue review agent.
[0038] The dialogue review agent is used to score the target dimension based on the child's dialogue and the companion's dialogue, and to select samples based on the scores to optimize the child dialogue agent and the companion dialogue agent.
[0039] It is evident that the complex task of providing emotional support to children can be broken down into independent sub-tasks through collaboration among emotional scene generation agents, user profile generation agents, child dialogue agents, companion dialogue agents, and dialogue review agents within a multi-agent system. This avoids the monotonous output of a single model and allows for more diverse and higher-quality dialogue data. Furthermore, the dialogue review agent scores the dialogue data and optimizes the child and companion dialogue agents, forming a closed loop that enables continuous improvement of both agents and ultimately enhances the quality of the dialogue data.
[0040] The dialogue data generation method provided in this application includes: generating a scene based on emotions and themes using an emotional scene generation agent; generating a child user profile using a user profile generation agent; generating a child dialogue using a child dialogue agent based on the scene and the child user profile; generating a companion dialogue using a companion dialogue agent based on the scene, the child user profile, and the child dialogue; scoring a target dimension using a dialogue review agent based on the child dialogue and the companion dialogue; determining samples based on the scores; and optimizing the child dialogue agent and the companion dialogue agent based on the samples.
[0041] As can be seen, the dialogue data generation method provided in this application, through the collaboration of an emotional scene generation agent, a user profile generation agent, a child dialogue agent, a companion dialogue agent, and a dialogue review agent, breaks down the complex task of providing emotional support to children into independent sub-tasks. This avoids the monotonous output of a single model and allows for more diverse and higher-quality generated dialogue data. Furthermore, the dialogue review agent scores the dialogue data, forming a closed loop for optimizing the child and companion dialogue agents. This enables continuous improvement of both agents, thereby enhancing the overall quality of the dialogue data.
[0042] The computer equipment and computer storage media provided in this application have the aforementioned technical effects. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the prior art and embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a schematic diagram of a multi-intelligent system provided in an embodiment of this application;
[0045] Figure 2 A flowchart illustrating a dialogue data generation method provided in an embodiment of this application;
[0046] Figure 3 This is a schematic diagram of a dialogue data generation process provided in an embodiment of this application;
[0047] Figure 4 This is a schematic diagram of a computer device provided in an embodiment of this application. Detailed Implementation
[0048] The core of this application is to provide a multi-agent system, a method for generating dialogue data, an apparatus, and a storage medium that can improve the diversity and quality of dialogue data.
[0049] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0050] Please refer to Figure 1 , Figure 1 A multi-agent system provided in this application embodiment, combined with Figure 1 As shown, the system includes:
[0051] The system comprises an emotional scene generation agent 10, a user profile generation agent 20, a child dialogue agent 30, a companion dialogue agent 40, and a dialogue review agent 50; the emotional scene generation agent 10, the user profile generation agent 20, and the dialogue review agent 50 are all connected to the child dialogue agent 30 and the companion dialogue agent 40, respectively; the child dialogue agent 30 is connected to the companion dialogue agent 40.
[0052] The emotional scene generation agent 10 is used to generate scenes based on emotions and themes, and output the scenes to the child dialogue agent 30 and the companion dialogue agent 40.
[0053] The user profile generation agent 20 is used to generate a child user profile and output the child user profile to the child dialogue agent 30 and the companion dialogue agent 40.
[0054] The child dialogue agent 30 is used to generate a child dialogue based on the scenario and the child user profile, and output the child dialogue to the companion dialogue agent 40 and the dialogue review agent 50.
[0055] The companion dialogue agent 40 is used to generate companion dialogue based on the scenario, the child user profile, and the child dialogue, and output the companion dialogue to the dialogue review agent 50.
[0056] The dialogue review agent 50 is used to score the target dimension based on the child's dialogue and the companion's dialogue, and to select samples based on the scores to optimize the child dialogue agent 30 and the companion dialogue agent 40.
[0057] The emotional scene generation agent 10 can generate a scene (event cause) based on the input emotion and theme. Emotions can include happiness, sadness, etc. Themes can include amusement parks, kindergartens, etc. Scenes can include playing in an amusement park, having a toy snatched from a kindergarten, etc. For example, if the emotion is happiness and the theme is amusement park, the emotional scene generation agent 10 will generate the scene of playing in an amusement park.
[0058] User profile generation agent 20 is used to generate diverse child user profiles. These profiles can include age, gender, personality traits, and behavioral habits. Furthermore, user profile generation agent 20 can connect with emotion / scene generation agent 10, which can output any one or more of the following: emotion, theme, or scene, to user profile generation agent 20 as a reference for generating diverse child user profiles.
[0059] The child dialogue agent 30 simulates the language characteristics and expression patterns of children in different emotional states, and generates children's dialogues based on the scenario and the child user profile.
[0060] The Companion Dialogue Agent 40 can output companion dialogues based on the scenario, the child user profile, and the child's conversations.
[0061] The dialogue review agent 50 scores the target dimensions of the child dialogue and the companion dialogue based on the child dialogue and the companion dialogue, and optimizes the child dialogue agent 30 and the companion dialogue agent 40 based on the scores, thereby achieving closed-loop optimization of the child dialogue agent 30 and the companion dialogue agent 40.
[0062] Based on the above embodiments, as a specific implementation method, the scoring of the target dimension based on the child's dialogue and the companion's dialogue includes:
[0063] The dialogue between the child and the companion dialogue are scored based on the dimensions of dialogue strategy, language naturalness, and logical depth.
[0064] The dialogue review agent 50 scores dialogues across three dimensions: dialogue strategy, naturalness of language, and depth of logic. By evaluating children's and companion's dialogues based on these three dimensions, the dialogue review agent 50 achieves automated quality assessment of the dual-agent interaction content, overcoming the subjective bias and inefficiency of manual scoring. Furthermore, it accurately identifies deficiencies such as inadequate dialogue strategy adaptation, stiff language, and disorganized logical connections, providing a reliable basis for optimizing the children's dialogue agent 30 and the companion's dialogue agent 40.
[0065] Based on the above embodiments, as a specific implementation method, the step of optimizing the child dialogue agent and the companion dialogue agent by selecting samples according to the scores includes:
[0066] If any of the dimensions of dialogue strategy, language naturalness, or logic depth fails to meet the requirements, positive and negative sample pairs are selected to optimize the child dialogue agent and the companion dialogue agent.
[0067] The dialogue review agent 50 scores the dialogue strategy, language naturalness, and logical depth across three dimensions. A score is considered acceptable if it reaches a preset threshold, and unacceptable if it falls below. If any of the three dimensions (dialogue strategy, language naturalness, or logical depth) fails to meet the threshold, a positive-negative sample pair is identified. This pair includes both positive and negative samples, which are used to optimize the model parameters of the child dialogue agent 30 and the companion dialogue agent 40.
[0068] In this embodiment, if any dimension score is unqualified, a positive and negative sample pair is determined. In this way, unqualified dialogues are transformed into positive and negative sample pairs that can be used for model training and optimization. This provides training samples for model iteration of dialogue strategies, language expression, and logical levels, continuously improving the strategy adaptability, language naturalness, and logical depth of children's emotional companionship dialogues, and continuously iterating and optimizing the interaction effect between children's dialogue agent 30 and companion dialogue agent 40.
[0069] Based on the above embodiments, as a specific implementation method, the step of selecting samples based on scores to optimize the child dialogue agent and the companion dialogue agent further includes:
[0070] If the scores for the dialogue strategy dimension, the language naturalness dimension, and the logical depth dimension are all satisfactory, then the child's dialogue and the companion dialogue are saved.
[0071] If the scores for the dialogue strategy dimension, the language naturalness dimension, and the logical depth dimension are all satisfactory, then the child's dialogue and the accompanying dialogue are saved.
[0072] In this embodiment, massive amounts of dialogue data are automatically generated by the child dialogue agent 30 and the companion dialogue agent 40, and the dialogue review agent 50 filters high-quality samples, thereby automatically expanding the training dataset and reducing the cost of manual annotation.
[0073] Based on the above embodiments, as a specific implementation method, the step of optimizing the child dialogue agent and the companion dialogue agent by selecting samples according to the scores includes:
[0074] Determine the direct preference optimization loss based on the sample;
[0075] The child dialogue agent and the companion dialogue agent are optimized based on the direct preference optimization loss.
[0076] In this embodiment, the DPO (Direct Preference Optimization) reinforcement learning method is used to continuously update the child dialogue agent 30 and the companion dialogue agent 40, thereby continuously improving the dialogue generation capabilities of the child dialogue agent 30 and the companion dialogue agent 40.
[0077] After completing multiple rounds of dialogue, the dialogue review agent 50 determines whether the dialogue data meets the standards in three dimensions: dialogue strategy, language naturalness, and logical depth. If the scores in each dimension are satisfactory, the dialogue data is saved. If the scores in any dimension are unsatisfactory, the dialogue that the dialogue review agent considers non-compliant is saved as a negative sample, and dialogues that the dialogue review agent considers compliant are generated as positive samples.
[0078] The DPO loss function is expressed as follows:
[0079] .
[0080] This indicates the current model to be optimized (child dialogue agent 30 or companion dialogue agent 40). This refers to the reference model (the initial base model, such as Qwen2.5-32B). This refers to the temperature parameter used to control the optimization intensity. This indicates a positive sample. This represents a negative sample. x represents the input (emotion, user profile, scenario, dialogue context).
[0081] Based on the direct preference optimization loss, the large dialogue model (e.g., Qwen2.5-32b) in the child dialogue agent 30 and the companion dialogue agent 40 is updated. Through the trained large dialogue model, the child dialogue agent 30 and the companion dialogue agent 40 can be more inclined to generate multi-turn dialogues that conform to the standards.
[0082] By determining the direct preference optimization loss through samples, and using the direct preference optimization loss to optimize the child dialogue agent 30 and the companion dialogue agent 40, it is possible to achieve closed-loop self-iterative optimization of the child dialogue agent 30 and the companion dialogue agent 40, so that the emotional companion dialogue ability of the two agents continuously converges to the optimal preference standard.
[0083] Based on the above embodiments, as a specific implementation method, after outputting the child's dialogue to the companion dialogue agent and the dialogue review agent, the following steps are included:
[0084] The companion dialogue agent selects a dialogue strategy based on the child's emotions and the scenario.
[0085] In this embodiment, the child dialogue agent 30 can also output the current emotion and determine the child's current state, which helps the companion dialogue agent 40 to select a dialogue strategy and output a companion dialogue that is more suitable for the current emotion.
[0086] By having the child's dialogue agent output emotions to the companion's dialogue agent, the companion's dialogue agent can obtain the child's current emotional state. Then, based on the child's emotional state, it can match an appropriate dialogue strategy and give personalized responses that fit the child's emotions, thereby improving the quality of the conversation.
[0087] Furthermore, the companion dialogue agent 40 can dynamically select dialogue strategies based on the child's emotional state and the scene, and output companion dialogue based on the scene, the child's user profile, and the child's dialogue.
[0088] Based on the above embodiments, as a specific implementation method, the emotional scene generation agent, the user profile generation agent, the child dialogue agent, the companion dialogue agent, and the dialogue review agent are built on the Dify platform.
[0089] The emotional scene generation agent 10, user profile generation agent 20, child dialogue agent 30, companion dialogue agent 40, and dialogue review agent 50 are built on the Dify platform. The entire process is based on the visual orchestration of the Dify platform, which supports modular adjustment, visual parameter control, and rapid expansion. This can greatly reduce data production costs and make the generation of emotional dialogue data more automated, controllable, and reusable.
[0090] In summary, the multi-agent system provided in this application can decompose complex child emotional companionship tasks into independent sub-tasks through collaboration among an emotional scene generation agent, a user profile generation agent, a child dialogue agent, a companion dialogue agent, and a dialogue review agent. This avoids the monotonous output of a single model and allows for more diverse and higher-quality dialogue data. Furthermore, the dialogue review agent scores the dialogue data and optimizes the child and companion dialogue agents, forming a closed loop that enables continuous improvement of both agents and enhances the overall quality of the dialogue data.
[0091] Please refer to Figure 2 , Figure 2 This is a flowchart illustrating a dialogue data generation method provided in an embodiment of this application. (Refer to...) Figure 2 As shown, the method includes:
[0092] S101: Generating intelligent agents based on emotions and themes;
[0093] S102: Generate a child user profile by generating an intelligent agent based on the user profile;
[0094] S103: Generate a child dialogue based on the scenario and the child user profile using a child dialogue agent;
[0095] S104: The companion dialogue agent generates a companion dialogue based on the scenario, the child user profile, and the child's dialogue.
[0096] S105: The dialogue review agent scores the target dimensions of the child dialogue and the companion dialogue based on the child dialogue and the companion dialogue, and selects samples based on the scores to optimize the child dialogue agent and the companion dialogue agent.
[0097] The dialogue data generation method provided in this application is applied to the multi-intelligent system described above, in which multiple intelligent agents work together to generate dialogue data, with each intelligent agent undertaking a specific function.
[0098] The emotional scene generation agent generates a scene (event cause) based on the input emotion and theme. Emotions can be happy, sad, etc. Themes can be amusement parks, kindergartens, etc., and scenes can be playing in an amusement park, having a toy snatched from a kindergarten, etc. For example, if the emotion is happy and the theme is an amusement park, the emotional scene generation agent will generate the scene of playing in an amusement park.
[0099] The intelligent agent generates diverse user profiles for children. These profiles can include age, gender, personality traits, and behavioral habits.
[0100] The child dialogue agent simulates the language characteristics and expression patterns of children in different emotional states. Based on the emotional scene and user profile generated by the agent, a child user profile is generated, and a child dialogue is generated.
[0101] In some embodiments, the method further includes: selecting a dialogue strategy based on the child's emotions and the scenario through the companion dialogue agent.
[0102] In this application, the child dialogue agent can also output the current emotion and judge the child's current state, which helps the companion dialogue agent to select dialogue strategies and output companion dialogue that is more suitable for the current emotion.
[0103] By having the child's dialogue agent output emotions to the companion's dialogue agent, the companion's dialogue agent can obtain the child's current emotional state. Then, based on the child's emotional state, it can match an appropriate dialogue strategy and give personalized responses that fit the child's emotions, thereby improving the quality of the conversation.
[0104] The companion dialogue agent can dynamically select dialogue strategies based on the child's emotional state and the scene, and output companion dialogue based on the scene, the child's user profile, and the child's conversation.
[0105] The dialogue review agent scores the target dimensions of the child dialogue and the companion dialogue based on the child dialogue and the companion dialogue. The sample is determined based on the score, and the child dialogue agent and the companion dialogue agent are optimized based on the sample, thereby realizing the closed-loop optimization of the child dialogue agent and the companion dialogue agent.
[0106] In some embodiments, the scoring of the target dimension by the dialogue review agent based on the child's dialogue and the companion's dialogue includes:
[0107] The dialogue review agent scores the dialogue strategy dimension, language naturalness dimension, and logical depth dimension based on the child's dialogue and the companion's dialogue.
[0108] The dialogue review agent scores the dialogue strategy, language naturalness, and logical depth. By scoring the child dialogue and the companion dialogue from the three dimensions of dialogue strategy, language naturalness, and logical depth, on the one hand, it can achieve the automated evaluation of the quality of the interaction content of the two agents, overcome the problems of subjective bias and inefficiency in manual scoring, and on the other hand, it can accurately identify defects such as insufficient adaptation of dialogue strategies, rigid language expressions, and chaotic logical connections, providing a reliable basis for optimizing the child dialogue agent and the companion dialogue agent.
[0109] In some embodiments, selecting samples according to the scores to optimize the child dialogue agent and the companion dialogue agent includes:
[0110] If the score of any dimension in the dialogue strategy dimension, language naturalness dimension, and logical depth dimension is unqualified, then select positive and negative sample pairs to optimize the child dialogue agent and the companion dialogue agent.
[0111] The dialogue review agent detects and scores the three dimensions of dialogue strategy, language naturalness, and logical depth. If the score reaches the preset threshold, the score is qualified; if the score does not reach the preset threshold, the score is unqualified. If the score of any dimension in the dialogue strategy dimension, language naturalness dimension, and logical depth dimension is unqualified, then determine the positive and negative sample pairs. The positive and negative sample pairs include positive samples and negative samples, so as to optimize the child dialogue agent and the companion dialogue agent by fine-tuning the model parameters of the child dialogue agent and the companion dialogue agent according to the positive and negative sample pairs.
[0112] In this embodiment, when the score of any dimension is unqualified, the positive and negative sample pairs are determined. In this way, the unqualified dialogues are converted into positive and negative sample pairs that can be used for model training and optimization, providing training samples for the model iteration of dialogue strategies, language expressions, and logical levels, continuously improving the strategy adaptability, language naturalness, and logical depth of children's emotional companion dialogues, and continuously iterating and optimizing the interaction effect of the child dialogue agent and the companion dialogue agent.
[0113] In some embodiments, the selecting samples according to the scores to optimize the child dialogue agent and the companion dialogue agent further includes:
[0114] If the scores of the dialogue strategy dimension, language naturalness dimension, and logical depth dimension are all qualified, then save the child dialogue and the companion dialogue.
[0115] If the scores of the dialogue strategy dimension, language naturalness dimension, and logical depth dimension are all qualified, then save the child dialogue and the companion dialogue.
[0116] In this embodiment, a massive amount of dialogue data is automatically generated by a child dialogue agent and a companion dialogue agent. A dialogue review agent filters high-quality samples, thereby automatically expanding the training dataset and reducing the cost of manual annotation.
[0117] In some embodiments, the step of selecting samples based on scores to optimize the child's dialogue agent and the companion's dialogue agent includes:
[0118] Determine the direct preference optimization loss based on the sample;
[0119] The child dialogue agent and the companion dialogue agent are optimized based on the direct preference optimization loss.
[0120] This embodiment employs the DPO (Direct Preference Optimization) reinforcement learning method to continuously update the child's dialogue agent and the companion's dialogue agent, thereby continuously improving their dialogue generation capabilities.
[0121] After completing multiple rounds of dialogue, the dialogue review agent determines whether the dialogue data meets the standards in three dimensions: dialogue strategy, language naturalness, and logical depth. If the scores in all dimensions are satisfactory, the dialogue data is saved. If the scores in any dimension are unsatisfactory, the dialogue data is saved as a negative sample because it is deemed unsatisfactory by the dialogue review agent, and a dialogue data that is deemed satisfactory by the dialogue review agent is generated as a positive sample.
[0122] The DPO loss function is expressed as follows:
[0123] .
[0124] This indicates the current model to be optimized (child dialogue agent or companion dialogue agent). This refers to the reference model (the initial base model, such as Qwen2.5-32B). This refers to the temperature parameter used to control the optimization intensity. This indicates a positive sample. This represents a negative sample. x represents the input (emotion, user profile, scenario, dialogue context).
[0125] Based on the direct preference optimization loss, update the large dialogue model (e.g., Qwen2.5-32b) in the child dialogue agent and the companion dialogue agent. Through the trained large dialogue model, the child dialogue agent and the companion dialogue agent can be more inclined to generate multi-turn dialogues that conform to the standards.
[0126] By determining the direct preference optimization loss through samples, and using the direct preference optimization loss to optimize the child dialogue agent and the companion dialogue agent, a closed-loop self-iterative optimization of the child dialogue agent and the companion dialogue agent can be achieved, so that the emotional companion dialogue ability of the two agents can continuously converge to the optimal preference standard.
[0127] In some embodiments, the emotional scene generation agent, the user profile generation agent, the child dialogue agent, the companion dialogue agent, and the dialogue review agent are built on the Dify platform.
[0128] The emotional scene generation agent, the user profile generation agent, the child dialogue agent, the companion dialogue agent, and the dialogue review agent are all built on the Dify platform. The entire process is based on the visual orchestration of the Dify platform, which supports modular adjustment, visual parameter control, and rapid expansion. This can greatly reduce data production costs and make the generation of emotional dialogue data more automated, controllable, and reusable.
[0129] refer to Figure 3 As shown, the following describes a specific embodiment:
[0130] The five core agents defined and running on the Dify platform are: the emotional scene generation agent, the user profile generation agent, the child dialogue agent, the companion dialogue agent, and the censorship agent.
[0131] Five intelligent agents interact in an orderly manner through workflow connections within the Dify platform, forming a closed-loop architecture.
[0132] The emotional scene generation agent generates scenes based on the input emotions and themes.
[0133] User profile generation: Intelligent agents generate diverse user profiles for children.
[0134] The child dialogue agent simulates the language characteristics and expression patterns of children (e.g., 4-6 year olds) under different emotional states, generates children's dialogues, outputs the current emotion, judges the child's current state, and helps the companion change the dialogue strategy.
[0135] The companion dialogue agent dynamically selects companionship strategies (e.g., soothing, encouraging, guiding, etc.) based on the child's emotional state and the scene, and outputs companionship dialogue;
[0136] The dialogue review agent scores the dialogue strategy, language naturalness, and logical depth based on the dialogue data between the child's dialogue agent and the companion's dialogue agent. If any dimension's score is unsatisfactory, positive and negative sample pairs that meet the reinforcement learning DPO (Discretionary Point of Interest) are generated. If all three dimensions' scores are satisfactory, the final dialogue data is saved. When a certain number of positive and negative samples are accumulated, such as 500, the child's dialogue agent and the companion's dialogue agent are fine-tuned based on the positive and negative samples.
[0137] This embodiment, in a child emotional support scenario, establishes a scalable multi-agent architecture encompassing emotional scene generation, user profile generation, a child dialogue agent, a companion dialogue agent, and a dialogue review agent, significantly enhancing scenario diversity and realism. Reinforcement learning DPO is introduced to determine positive and negative sample pairs based on the scoring results of the dialogue review agent, updating the parameters of the child and companion dialogue agents to form a learnable closed loop. The child and companion dialogue agents automatically generate massive amounts of dialogue data, while the dialogue review agent filters high-quality samples, automatically expanding the training dataset and reducing manual annotation costs. The Dify platform is used as the underlying execution platform, organizing the multi-agent architecture into a closed loop through a low-code workflow.
[0138] In summary, the dialogue data generation method provided in this application breaks down the complex task of providing emotional support to children into independent sub-tasks through the collaboration of an emotional scene generation agent, a user profile generation agent, a child dialogue agent, a companion dialogue agent, and a dialogue review agent. This avoids the monotonous output of a single model and allows for more diverse and higher-quality generated dialogue data. Furthermore, the dialogue review agent scores the dialogue data, creating a closed loop for optimizing the child and companion dialogue agents, thus continuously improving the quality of the dialogue data.
[0139] This application also provides a computer device, referenced... Figure 4 As shown, the device includes a memory 1 and a processor 2.
[0140] Memory 1 is used to store computer programs;
[0141] Processor 2 is used to execute computer programs to perform the following steps:
[0142] The system generates an intelligent agent based on emotions and themes to create scenarios; it generates a child user profile based on user profiles; it generates a child dialogue agent based on the scenarios and the child user profile; it generates a companion dialogue agent based on the scenarios, the child user profile, and the child dialogue; and it scores the target dimensions based on the child dialogue and the companion dialogue agent based on the scores, and selects samples to optimize the child dialogue agent and the companion dialogue agent.
[0143] For a description of the equipment provided in this application, please refer to the above method embodiments; further details will not be provided here.
[0144] This application also provides a computer storage medium storing a computer program, which, when executed by a processor, can perform the following steps:
[0145] The system generates an intelligent agent based on emotions and themes to create scenarios; it generates a child user profile based on user profiles; it generates a child dialogue agent based on the scenarios and the child user profile; it generates a companion dialogue agent based on the scenarios, the child user profile, and the child dialogue; and it scores the target dimensions based on the child dialogue and the companion dialogue agent based on the scores, and selects samples to optimize the child dialogue agent and the companion dialogue agent.
[0146] The computer storage medium may include: USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and other media that can store program code.
[0147] For a description of the computer storage medium provided in this application, please refer to the above method embodiments; further details will not be repeated here.
[0148] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. The systems, devices, and computer storage media disclosed in the embodiments are described simply because they correspond to the methods disclosed in the embodiments; relevant details can be found in the method section.
[0149] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0150] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0151] The multi-agent system, dialogue data generation method, device, and storage medium provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A multi-agent system, characterized in that, include: Intelligent agents for generating emotional scenarios, intelligent agents for generating user profiles, intelligent agents for children's dialogue, intelligent agents for companion's dialogue, and intelligent agents for dialogue review; The emotional scene generation agent, the user profile generation agent, and the dialogue review agent are all connected to the child dialogue agent and the companion dialogue agent, respectively. The child dialogue agent is connected to the companion dialogue agent; The emotional scene generation agent is used to generate scenes based on emotions and themes, and output the scenes to the child dialogue agent and the companion dialogue agent. The user profile generation agent is used to generate a child user profile and output the child user profile to the child dialogue agent and the companion dialogue agent. The child dialogue agent is used to generate a child dialogue based on the scenario and the child user profile, and output the child dialogue to the companion dialogue agent and the dialogue review agent. The companion dialogue agent is used to generate companion dialogue based on the scenario, the child user profile, and the child's dialogue, and output the companion dialogue to the dialogue review agent. The dialogue review agent is used to score the target dimension based on the child's dialogue and the companion's dialogue, and to select samples based on the scores to optimize the child dialogue agent and the companion dialogue agent.
2. The multi-agent system according to claim 1, characterized in that, The scoring of the target dimension based on the child's dialogue and the companion's dialogue includes: The dialogue between the child and the companion dialogue are scored based on the dimensions of dialogue strategy, language naturalness, and logical depth.
3. The multi-agent system according to claim 2, characterized in that, The step of selecting samples based on scores to optimize the child's dialogue agent and the companion's dialogue agent includes: If any of the dimensions of dialogue strategy, language naturalness, or logic depth fails to meet the requirements, positive and negative sample pairs are selected to optimize the child dialogue agent and the companion dialogue agent.
4. The multi-agent system according to claim 2, characterized in that, The step of selecting samples based on scores to optimize the child dialogue agent and the companion dialogue agent further includes: If the scores for the dialogue strategy dimension, the language naturalness dimension, and the logical depth dimension are all satisfactory, then the child's dialogue and the companion dialogue are saved.
5. The multi-agent system according to claim 1, characterized in that, The step of selecting samples based on scores to optimize the child's dialogue agent and the companion's dialogue agent includes: Determine the direct preference optimization loss based on the sample; The child dialogue agent and the companion dialogue agent are optimized based on the direct preference optimization loss.
6. The multi-agent system according to claim 1, characterized in that, After outputting the child's dialogue to the companion dialogue agent and the dialogue review agent, the process includes: The companion dialogue agent selects a dialogue strategy based on the child's emotions and the scenario.
7. The multi-agent system according to claim 1, characterized in that, The emotional scene generation agent, the user profile generation agent, the child dialogue agent, the companion dialogue agent, and the dialogue review agent are all built on the Dify platform.
8. A method for generating dialogue data, characterized in that, Applied to the multi-agent system as described in claim 1, comprising: Intelligent agents are generated based on emotions and themes to create scenarios. Generate child user profiles using intelligent agents generated from user profiles; A child dialogue agent generates a child dialogue based on the scenario and the child user profile. The companion dialogue agent generates a companion dialogue based on the scenario, the child user profile, and the child's dialogue. The dialogue review agent scores the target dimensions based on the child's dialogue and the companion's dialogue, and selects samples based on the scores to optimize the child dialogue agent and the companion dialogue agent.
9. A computer device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the dialogue data generation method as described in claim 8 when executing the computer program.
10. A computer storage medium, characterized in that, The computer storage medium stores a computer program, which, when executed by a processor, implements the steps of the dialogue data generation method as described in claim 8.