Multi-agent digital human collaborative interaction and scene deduction method and method
By introducing a fixed persona anchoring benchmark and dynamic strategy branches into a multi-agent digital human system, and combining global situational characteristics and multi-dimensional feedback, the problems of persona collapse and poor global coordination in long-term interaction of multi-agent digital humans are solved, and the adaptive expression of roles and continuous optimization of models are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AOYING TECHNOLOGY (GUANGDONG) CO LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-28
AI Technical Summary
Existing multi-agent digital humans suffer from problems such as easy collapse of character settings over long periods of time, poor global coordination, inability to autonomously evolve strategies and models, resulting in character settings deviation, dialogue deviation, content conflicts and high manual debugging costs in long-term interactions.
It adopts a two-layer architecture with a fixed character anchoring benchmark and dynamic strategy branches. Through global situation feature synchronization, multi-agent collaborative adaptation and temporal synchronization, combined with multi-dimensional feedback-driven dynamic strategy optimization and high-quality experience accumulation, it achieves adaptive character expression and model iteration.
It achieves consistency in character design, global collaboration, and model self-evolution in long-term interactions, eliminates dialogue repetition and content conflicts, reduces manual debugging costs, and improves collaborative effects.
Smart Images

Figure CN122472073A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence digital human technology, specifically to a method and approach for multi-agent digital human collaborative interaction and scene interpretation. Background Technology
[0002] With the rapid development of multimodal large-scale models and digital human rendering technology, multi-agent digital human collaborative interaction has been widely applied in scenarios such as live-streaming e-commerce, content creation, and simulation. Currently, the mainstream implementation methods are mainly divided into two categories: one is a pre-set script-triggered scheme, which controls the execution of digital humans through fixed lines, actions, and timing codes, resulting in poor flexibility and an inability to respond to real-time user input; the other is a simple single-agent splicing scheme, where each digital human uses an independent large-scale model and a fixed persona prompt to independently generate content, lacking a global collaborative mechanism.
[0003] The existing technology has several core flaws:
[0004] Firstly, the problem of character deterioration in long-term interactions: the solution of statically locking character attributes with a fixed Prompt can only maintain character consistency in short rounds of interaction. After multiple rounds of collaborative interaction in long-term interactions, it is very easy for character deviation, language style deviation, unclear character positioning or even complete collapse to occur. It cannot be adapted to scenarios such as long live broadcasts, long drama series, and long-term wargame simulations.
[0005] Secondly, the lack of global collaboration is a defect: the architecture of independent decision-making by single agents and simple splicing is without a collaborative scheduling mechanism based on global role division and scene goals. Each digital human only focuses on its own content generation, which easily leads to problems such as interruption, repetitive lines, content conflicts, logical fragmentation, and misaligned action sequences, resulting in a strong sense of disjointed collaborative performance.
[0006] Third, static and lack of evolutionary defects: the general algorithm has not been specifically optimized for digital human collaboration scenarios, lacks high-quality collaboration experience accumulation and self-iteration mechanism, and requires manual re-tuning of the Prompt and optimization of model parameters every time the scenario or role is changed, which is extremely costly and cannot improve the collaboration effect with the increase of usage. Summary of the Invention
[0007] This application primarily addresses the technical problems of existing multi-agent digital humans in collaborative interactions, such as the easy collapse of long-term character profiles, poor global coordination, and the inability to autonomously evolve strategies and models.
[0008] According to the first aspect, embodiments of this application provide a method for multi-agent digital human collaborative interaction and scene deduction, including:
[0009] Configure a fixed persona anchoring benchmark for each digital human agent participating in the current collaborative interaction, and a dynamic policy branch that runs within the constraint framework of the persona anchoring benchmark, to complete the initialization of the decision model for each digital human agent;
[0010] Real-time collection of multi-source interaction data in the current collaborative interaction scenario, extraction and generation of scenario situation features, and synchronous distribution of the scenario situation features to all digital human intelligent agents participating in the current collaborative interaction;
[0011] Each digital human agent generates candidate interaction actions based on its own persona anchoring benchmark and scene situation characteristics through corresponding dynamic strategy branches. After performing persona compliance verification on the candidate interaction actions, it outputs compliant candidate interaction actions.
[0012] For all compliant candidate interactive actions of the digital human intelligent agents participating in the current collaborative interaction, perform multi-agent collaborative adaptation and timing synchronization, generate synchronous execution instructions and send them to each digital human intelligent agent;
[0013] Each digital human intelligent agent completes multimodal collaborative interpretation according to synchronous execution instructions, and collects the interpretation execution results of the current collaborative interaction;
[0014] Based on the current collaborative interaction, multi-dimensional feedback data is obtained, and the model parameters of the dynamic strategy branches of each digital human agent are updated according to the feedback data. During the update process, the human setting anchor benchmark is not modified.
[0015] Select high-quality collaborative interaction experiences that meet preset conditions and accumulate them. Based on the accumulated high-quality experiences, iteratively update the decision-making foundation model of all digital human intelligent agents.
[0016] According to the second aspect, the embodiments of this application provide a multi-agent digital human collaborative interaction and scene interpretation system, including an initialization configuration module, a scene situation awareness module, a multi-agent decision-making module, a collaborative management and control module, a multimodal interpretation driving module, a feedback and strategy update module, and an experience iteration module;
[0017] The initialization configuration module is communicatively connected to the multi-agent decision-making module and the experience iteration module, respectively, and is used to configure a fixed human setting anchor benchmark and dynamic policy branches within the constraint framework for the digital human agents participating in the current collaborative interaction, thereby completing the initialization of the decision model.
[0018] The scene situation awareness module is communicatively connected to the multi-agent decision-making module and the collaborative control module, respectively. It is used to collect multi-source interaction data of the current collaborative interaction scene, extract and generate scene situation features, and then synchronously distribute them to all digital human agents participating in the current collaborative interaction.
[0019] The multi-agent decision-making module is communicatively connected to the collaborative management and control module, and has built-in decision-making models corresponding to multiple digital human agents in the system. It is used to generate candidate interactive actions based on the person setting anchor benchmark and scene situation characteristics, and output compliant candidate interactive actions after performing person setting compliance verification.
[0020] The collaborative management module is communicatively connected to the multimodal interpretation driving module, and is used to perform multi-agent collaborative adaptation and timing synchronization on all compliant candidate interactive actions of the digital human intelligent agents participating in the current collaborative interaction, and generate synchronous execution instructions.
[0021] The multimodal interpretation driving module is communicatively connected to the feedback and strategy update module, and is used to drive the digital human intelligent agent to complete multimodal collaborative interpretation, collect and feedback the interpretation execution results of the current collaborative interaction;
[0022] The feedback and strategy update module is connected to the multi-agent decision-making module and the experience iteration module respectively. It is used to collect multi-dimensional feedback data, calculate reward values and update the model parameters of the dynamic strategy branch. The update process does not modify the human anchoring benchmark.
[0023] The experience iteration module is connected to the initialization configuration module to accumulate high-quality collaborative interaction experience and iteratively update the decision-making foundation model of all digital human agents in the system.
[0024] The multi-agent digital human collaborative interaction and scene interpretation method and system described in the above embodiments structurally locks the core attributes of the role by permanently fixing the underlying persona anchoring benchmark. No strategy update will modify this benchmark, preventing persona deviation in long-term interactions. At the same time, dynamic strategy branches are retained within the persona constraint framework to achieve scene adaptation of the role's expression, perfectly solving the core contradiction of "persona rigidity and persona loss of control". Through globally unified scene situation feature distribution, all agents make decisions based on the same global information. Through global collaborative verification and time synchronization mechanism, problems such as interruption, repetition, conflict, and misalignment are eliminated at the execution level, achieving seamless collaborative interpretation of multiple digital humans. Through multi-dimensional feedback to drive real-time optimization of dynamic strategy branches, and through the accumulation of high-quality experience and knowledge distillation, the basic model is continuously iterated. The longer the system is used and the richer the data, the better the collaborative effect. Attached Figure Description
[0025] Figure 1 A flowchart of the multi-agent digital human collaborative interaction and scene interpretation method provided in the embodiments of this application;
[0026] Figure 2 This is an overall architecture diagram of the multi-agent digital human collaborative interaction and scene interpretation system provided in the embodiments of this application. Detailed Implementation
[0027] The present application will now be described in further detail with reference to the accompanying drawings and specific embodiments. Similar elements in different embodiments are referred to by related, similar element designations.
[0028] In the following embodiments, many details are described in order to enable a better understanding of this application. However, those skilled in the art will readily recognize that some of these features may be omitted in different circumstances, or may be replaced by other elements, materials, or methods.
[0029] In some cases, certain operations related to this application are not shown or described in the specification. This is to avoid the core parts of this application being overwhelmed by excessive description. For those skilled in the art, it is not necessary to describe these related operations in detail. They can fully understand the related operations based on the description in the specification and general technical knowledge in the field.
[0030] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.
[0031] The serial numbers assigned to components in this document, such as "first" and "second," are used only to distinguish the described objects and have no sequential or technical meaning. The terms "connection" and "linkage" used in this application, unless otherwise specified, include both direct and indirect connections (linkages).
[0032] Please refer to Figure 1 and Figure 2 To address the technical problems of existing multi-agent digital humans in collaborative interaction, such as easy collapse of long-term personas, poor global coordination, and inability to autonomously evolve strategies and models, this application provides a method for multi-agent digital human collaborative interaction and scene deduction, which includes the following steps:
[0033] Step S101: Configure a fixed persona anchoring benchmark for each digital human agent participating in the current collaborative interaction, and a dynamic policy branch that runs within the constraint framework of the persona anchoring benchmark, to complete the initialization of the decision model of each digital human agent.
[0034] In some embodiments, the number of digital human agents participating in the collaboration, their roles, core objectives, and collaboration rules are determined according to scenario requirements. A unique 128-dimensional dense basic persona anchor vector is generated for each digital human, encoding six core attributes: core identity, basic personality, language style, professional field, taboo rules, and role positioning, which are permanently embedded in the input layer of the policy network. An improved PPO policy network is initialized, a pre-trained basic model is loaded, and the reward function weight coefficients for the corresponding scenario are configured. The global collaborative scheduling center, situational awareness module, experience pool, and model distillation unit are initialized.
[0035] Step S102: Collect multi-source interaction data of the current collaborative interaction scenario in real time, extract and generate scenario situation features, and synchronously distribute the scenario situation features to all digital human intelligent agents participating in the current collaborative interaction.
[0036] In some embodiments, real-time data from the entire scenario is collected, including scenario-based rules, real-time user input / questions / comments, historical dialogue / actions / state data of all digital human agents, and real-time external access data (such as wargame situation data and live sales data). Through a hybrid coding network of **"CNN feature extraction + Transformer semantic encoding"**, feature extraction and fusion of multi-source data are performed to generate a standardized global situation feature vector, which is then synchronously distributed to all digital human agents and the global collaborative scheduling center.
[0037] Step S103: Each digital human agent generates candidate interactive actions based on its own persona anchoring benchmark and scene situation characteristics through the corresponding dynamic strategy branch. After performing persona compliance verification on the candidate interactive actions, it outputs compliant candidate interactive actions.
[0038] In some embodiments, each digital human agent inputs its own basic persona anchor vector, global situation feature vector, and its own historical interaction state features into an improved PPO policy network; within the constraint framework of the basic persona anchor, the policy network generates a sequence of candidate actions including dialogue text, tone of voice, body movements, and screen interactions; each agent performs preliminary persona compliance verification on the candidate actions, filters out illegal content that deviates from the basic persona, and uploads compliant candidate actions to the global collaborative scheduling center.
[0039] Step S104: For all compliant candidate interactive actions of the digital human agents participating in the current collaborative interaction, perform multi-agent collaborative adaptation and timing synchronization, generate synchronous execution instructions and send them to each digital human agent.
[0040] In some embodiments, the global collaborative scheduling center receives candidate actions from all agents and performs conflict verification (checking and filtering candidate actions with content conflicts, interruptions, and repetitions), logic verification (ensuring that the output of all agents conforms to the dialogue context logic), and timing alignment (assigning a unified timestamp to all compliant actions to ensure that dialogue, actions, and images are synchronized at the millisecond level). After verification, the final execution instruction is generated and synchronously sent to the multimodal driving unit of each digital human agent.
[0041] Step S105: Each digital human agent completes multimodal collaborative interpretation according to the synchronous execution instructions and collects the interpretation execution results of the current collaborative interaction.
[0042] In some embodiments, the multimodal driving unit receives execution instructions and synchronously drives all digital humans to complete dialogue speech generation, lip-sync, body movement performance, and screen content switching, thereby realizing seamless collaborative scene performance of multiple digital humans, and simultaneously synchronizing the execution results to the feedback acquisition module.
[0043] Step S106: Obtain multi-dimensional feedback data based on the current collaborative interaction execution results, and update the model parameters of the dynamic strategy branches of each digital human agent according to the feedback data, without modifying the human setting anchoring benchmark during the update process.
[0044] In some embodiments, four types of feedback data are collected in real time: persona matching degree, collaborative coherence, scene goal completion degree, and real-time user feedback. Based on the reconstructed five-dimensional weighted reward function, the immediate reward value and global collaborative reward value of each digital human agent are calculated and synchronously fed back to the corresponding PPO policy network. Based on the calculated reward values, policy gradient updates are performed through an improved PPO algorithm, strengthening high-reward, high-quality policies and suppressing low-reward / penalty violation policies. Throughout the policy update process, the underlying basic persona anchor vector is not modified; only the upper-level dynamic policy branches are optimized, ensuring persona consistency while completing policy auto-evolution.
[0045] Step S107: Select high-quality collaborative interaction experiences that meet the preset conditions and accumulate them. Based on the accumulated high-quality experiences, iteratively update the decision-making foundation model of all digital human intelligent agents.
[0046] In some embodiments, high-quality collaborative strategies, character adaptation experience, and scenario-based interpretation data with reward values exceeding a preset threshold are deposited into a dedicated experience pool and stored in categories according to scenario and role type; offline knowledge distillation is performed at fixed intervals to distill high-quality experience from the experience pool into the basic pre-trained model, update the initial network weights of all agents, and realize the continuous self-evolution of the model.
[0047] The persona anchoring benchmark is a standardized feature benchmark encoded with the core attributes of the digital human intelligent agent. The core attributes include the core identity of the role, basic personality, language style, professional field, taboo rules, and role positioning. The persona anchoring benchmark is permanently embedded in the input layer of the corresponding digital human intelligent agent decision-making model and is not modified with strategy updates throughout the entire process of collaborative interaction and strategy evolution.
[0048] The decision model adopts a hybrid architecture that connects a feature extraction unit, a semantic encoding unit, and a policy output unit. The feature extraction unit is used to extract structured features of scene situation features, the semantic encoding unit is used to encode the semantic context of the current collaborative interaction and the attribute features of the persona anchoring benchmark, and the policy output unit is used to output candidate interaction actions.
[0049] Furthermore, updating the model parameters of the dynamic policy branches of each digital human agent based on the feedback data includes:
[0050] A multi-objective fusion reward function is constructed based on four dimensions: persona matching degree, collaborative coherence, scene goal completion degree, and user interaction feedback, to calculate the real-time reward value for each digital human agent participating in the current collaborative interaction.
[0051] Based on the real-time reward value, the policy gradient optimization algorithm is used to strengthen the high-reward corresponding high-quality policies and suppress the low-reward corresponding illegal policies, thereby completing the parameter update of the dynamic policy branch.
[0052] Furthermore, the decision-making foundation model of all digital human agents is iteratively updated based on accumulated high-quality experience, including:
[0053] High-quality collaborative interaction experiences with reward values exceeding a preset threshold are categorized and stored in a dedicated experience pool according to application scenarios and role types.
[0054] At a preset fixed period, high-quality collaborative interaction experiences from the experience pool are iteratively updated into the decision-making foundation model of the digital human intelligent agent through knowledge distillation, thereby completing the model's self-evolution.
[0055] Furthermore, the multi-source interaction data of the current collaborative interaction scenario includes the basic rule data of the current collaborative interaction, real-time user input data, historical interaction data of all digital human intelligent agents participating in the current collaborative interaction, and real-time status data of the current collaborative interaction scenario;
[0056] By using a pre-defined feature encoding network, features are extracted and fused from the aforementioned multi-source interactive data to generate standardized scene situation features.
[0057] Furthermore, all compliant candidate interaction actions of the digital human agents participating in the current collaborative interaction perform multi-agent collaborative adaptation and timing synchronization, including:
[0058] Investigate and filter candidate interactive actions that have content conflicts, overlapping speaking times, or repetitive content;
[0059] Verify the logical coherence between the remaining candidate interaction actions and the current collaborative interaction context, and eliminate candidate interaction actions that have logical contradictions.
[0060] Assign a unified timestamp to the final compliant candidate interactive actions, complete the timing synchronization of lines, actions, and screen, and generate synchronous execution instructions.
[0061] The more specific implementation details of the method in this application are as follows:
[0062] In some embodiments, the present invention abandons the existing fixed Prompt static persona mode and constructs a two-layer persona control architecture for each digital human agent, consisting of "bottom-level basic persona anchoring + upper-level dynamic strategy evolution":
[0063] Bottom Layer: Basic Character Anchor Vector (Fixed): A dense 128-dimensional basic character anchor vector is generated for each digital human. The six core attributes of the character, namely core identity, basic personality, language style, professional field, taboo rules, and role positioning, are vectorized and permanently embedded into the input layer of the intelligent agent policy network. The core attributes of the character are locked throughout the process. No matter what strategy is updated, the basic character anchor vector will not be modified, thus preventing the problem of character collapse from the bottom layer.
[0064] Upper layer: Dynamic strategy evolution layer (real-time adaptive adjustment): Within the constraint framework anchored by the basic persona, dynamic strategy evolution branches are set for each agent. Based on the global scene situation, real-time user input, the interactive behavior of other agents, and the completion degree of scene goals, the improved PPO algorithm dynamically adjusts the character's expression strategy, tone of voice, interaction rhythm, division of labor weight, and emotional tendency to achieve strategy autonomous evolution that conforms to the character's persona and adapts to scene changes.
[0065] This invention is based on the mature CNN-PPO wargame multi-agent algorithm of enterprises, and makes three core improvements for multi-digital human collaboration scenarios: policy network structure reconstruction, multi-objective fusion reward function reconstruction, and dynamic pruning threshold optimization.
[0066] The policy network structure is reconstructed. Traditional PPO policy networks only input their own state features, while the policy network input layer of this invention permanently integrates three types of features: basic persona anchor vector, global situation features, and historical interaction features of other agents. This ensures that each agent's decisions conform to its own persona while also considering the global scenario and the behavior of other digital humans. The network structure adopts a hybrid architecture of "CNN feature extraction + Transformer semantic encoding + PPO policy head," where CNN is used to extract structured features of the global situation, Transformer is used to encode semantic context and persona attributes, and the PPO policy head is used to output the final dialogue and action decisions.
[0067] The multi-objective fusion reward function is reconstructed, abandoning the traditional PPO single-task completion reward and constructing a 5-dimensional weighted reward function for multi-digital human collaboration scenarios. The total reward formula is as follows: ;
[0068] Wherein, α, β, γ, ε, and δ are configurable weight coefficients that can be dynamically adjusted according to different scenarios (e.g., increasing the collaboration weight β in live streaming scenarios, and increasing the scenario adaptation weight γ in wargame scenarios). The definitions of each sub-reward are as follows:
[0069] Character consistency reward: The matching degree between the output content and the basic character anchor quantity is calculated by the semantic similarity model. The higher the matching degree, the higher the reward. The penalty is output that deviates from the character.
[0070] : Coherence reward, which rewards outputs that are logically coherent in dialogue with other intelligent agents and whose actions are synchronized in timing, and punishes behaviors such as interrupting, repetitive content, and logical breaks;
[0071] : Scenario-adaptive rewards, which reward outputs that align with the core objectives of the scenario (such as conversion guidance in live-streaming e-commerce or the accuracy of situational analysis in wargame simulations).
[0072] User feedback rewards are based on real-time user comments, questions, and interactions, rewarding outputs that meet user needs.
[0073] Conflict penalty items penalize content conflicts, character violations, and out-of-order outputs to ensure orderly collaboration.
[0074] Dynamic pruning threshold optimization involves dynamically adjusting the pruning threshold of the PPO algorithm to meet the differentiated real-time requirements of digital human scenarios: a low pruning threshold of 0.1-0.2 is used for real-time interactive scenarios (such as live streaming and real-time wargame command) to limit large fluctuations in the strategy and ensure stable output and low latency; a high pruning threshold of 0.2-0.3 is used for offline performance scenarios (such as situational drama production) to expand the strategy exploration space and achieve richer performance effects.
[0075] Please refer to Figure 2 This application also provides a multi-agent digital human collaborative interaction system, including an initialization configuration module, a scene situation perception module, a multi-agent decision-making module, a collaborative control module, a multimodal deduction driving module, a feedback and strategy update module, and an experience iteration module;
[0076] The initialization configuration module is communicatively connected to the multi-agent decision-making module and the experience iteration module, respectively, and is used to configure a fixed human setting anchor benchmark and dynamic policy branches within the constraint framework for the digital human agents participating in the current collaborative interaction, thereby completing the initialization of the decision model.
[0077] The scene situation awareness module is communicatively connected to the multi-agent decision-making module and the collaborative control module, respectively. It is used to collect multi-source interaction data of the current collaborative interaction scene, extract and generate scene situation features, and then synchronously distribute them to all digital human agents participating in the current collaborative interaction.
[0078] The multi-agent decision-making module is communicatively connected to the collaborative management and control module, and has built-in decision-making models corresponding to multiple digital human agents in the system. It is used to generate candidate interactive actions based on the person setting anchor benchmark and scene situation characteristics, and output compliant candidate interactive actions after performing person setting compliance verification.
[0079] The collaborative management module is communicatively connected to the multimodal interpretation driving module, and is used to perform multi-agent collaborative adaptation and timing synchronization on all compliant candidate interactive actions of the digital human intelligent agents participating in the current collaborative interaction, and generate synchronous execution instructions.
[0080] The multimodal interpretation driving module is communicatively connected to the feedback and strategy update module, and is used to drive the digital human intelligent agent to complete multimodal collaborative interpretation, collect and feedback the interpretation execution results of the current collaborative interaction;
[0081] The feedback and strategy update module is connected to the multi-agent decision-making module and the experience iteration module respectively. It is used to collect multi-dimensional feedback data, calculate reward values and update the model parameters of the dynamic strategy branch. The update process does not modify the human anchoring benchmark.
[0082] The experience iteration module is connected to the initialization configuration module to accumulate high-quality collaborative interaction experience and iteratively update the decision-making foundation model of all digital human agents in the system.
[0083] The decision model built into the multi-agent decision module adopts a hybrid architecture that connects a feature extraction unit, a semantic encoding unit, and a policy output unit. The input of the feature extraction unit is connected to the scene situation features, the input of the semantic encoding unit is connected to the persona anchoring benchmark and the semantic context data of the current collaborative interaction, and the output of the policy output unit outputs candidate interaction actions.
[0084] The feedback and strategy update module has a built-in multi-objective reward calculation unit. The multi-objective reward calculation unit constructs a multi-objective fusion reward function based on four dimensions: persona matching degree, collaborative coherence, scene goal completion degree, and user interaction feedback, and calculates the real-time reward value of each digital human agent participating in the current collaborative interaction.
[0085] In some embodiments, the multi-agent digital human collaborative interaction system of the present invention can be deployed on a cloud server, adopting a distributed architecture, and supporting simultaneous access of at least 100 digital human agents working collaboratively. The modules communicate using the MQTT protocol, with a communication latency of no more than 100ms, meeting the low-latency requirements of real-time interaction scenarios. The system supports multi-scene switching and can quickly adapt to different scenarios such as live streaming, situational dramas, and wargames by loading different basic models and character anchor vectors.
[0086] Those skilled in the art will understand that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or external hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.
[0087] The above examples illustrate this application only to aid understanding and are not intended to limit its scope. Those skilled in the art to which this application pertains can make various simple deductions, modifications, or substitutions based on the ideas presented.
Claims
1. A multi-agent digital human collaborative interaction method, characterized in that, include: Configure a fixed persona anchoring benchmark for each digital human agent participating in the current collaborative interaction, and a dynamic policy branch that runs within the constraint framework of the persona anchoring benchmark to complete the initialization of the decision model for each digital human agent; Real-time collection of multi-source interaction data in the current collaborative interaction scenario, extraction and generation of scenario situation features, and synchronous distribution of the scenario situation features to all digital human intelligent agents participating in the current collaborative interaction; Each digital human agent generates candidate interaction actions based on its own persona anchoring benchmark and scene situation characteristics through corresponding dynamic strategy branches. After performing persona compliance verification on the candidate interaction actions, it outputs compliant candidate interaction actions. For all compliant candidate interactive actions of the digital human intelligent agents participating in the current collaborative interaction, perform multi-agent collaborative adaptation and timing synchronization, generate synchronous execution instructions and send them to each digital human intelligent agent; Each digital human intelligent agent completes multimodal collaborative interpretation according to synchronous execution instructions, and collects the interpretation execution results of the current collaborative interaction; Based on the current collaborative interaction, multi-dimensional feedback data is obtained, and the model parameters of the dynamic strategy branches of each digital human agent are updated according to the feedback data. During the update process, the human setting anchor benchmark is not modified. Select high-quality collaborative interaction experiences that meet preset conditions and accumulate them. Based on the accumulated high-quality experiences, iteratively update the decision-making foundation model of all digital human intelligent agents.
2. The multi-agent digital human collaborative interaction method as described in claim 1, characterized in that, The persona anchoring benchmark is a standardized feature benchmark encoded with the core attributes of the digital human intelligent agent. The core attributes include the core identity of the role, basic personality, language style, professional field, taboo rules, and role positioning. The persona anchoring benchmark is permanently embedded in the input layer of the corresponding digital human intelligent agent decision-making model and is not modified with strategy updates throughout the entire process of collaborative interaction and strategy evolution.
3. The multi-agent digital human collaborative interaction method as described in claim 1, characterized in that, The decision model adopts a hybrid architecture that connects a feature extraction unit, a semantic encoding unit, and a policy output unit. The feature extraction unit is used to extract structured features of scene situation features, the semantic encoding unit is used to encode the semantic context of the current collaborative interaction and the attribute features of the persona anchoring benchmark, and the policy output unit is used to output candidate interaction actions.
4. The multi-agent digital human collaborative interaction method as described in claim 1, characterized in that, The step of updating the model parameters of the dynamic policy branches of each digital human agent based on the feedback data includes: A multi-objective fusion reward function is constructed based on four dimensions: persona matching degree, collaborative coherence, scene goal completion degree, and user interaction feedback, to calculate the real-time reward value for each digital human agent participating in the current collaborative interaction. Based on the real-time reward value, the policy gradient optimization algorithm is used to strengthen the high-reward corresponding high-quality policies and suppress the low-reward corresponding illegal policies, thereby completing the parameter update of the dynamic policy branch.
5. The multi-agent digital human collaborative interaction method as described in claim 4, characterized in that, The decision-making foundation model for all digital human agents is iteratively updated based on accumulated high-quality experience, including: High-quality collaborative interaction experiences with reward values exceeding a preset threshold are categorized and stored in a dedicated experience pool according to application scenarios and role types. At a preset fixed period, high-quality collaborative interaction experiences from the experience pool are iteratively updated into the decision-making foundation model of the digital human intelligent agent through knowledge distillation, thereby completing the model's self-evolution.
6. The multi-agent digital human collaborative interaction method as described in claim 1, characterized in that, The multi-source interaction data of the current collaborative interaction scenario includes the basic rule data of the current collaborative interaction, real-time user input data, historical interaction data of all digital human intelligent agents participating in the current collaborative interaction, and real-time status data of the current collaborative interaction scenario; By using a pre-defined feature encoding network, features are extracted and fused from the aforementioned multi-source interactive data to generate standardized scene situation features.
7. The multi-agent digital human collaborative interaction method as described in claim 1, characterized in that, All compliant candidate interaction actions of the digital human agents participating in the current collaborative interaction shall be subject to multi-agent collaborative adaptation and timing synchronization, including: Investigate and filter candidate interactive actions that have content conflicts, overlapping speaking times, or repetitive content; Verify the logical coherence between the remaining candidate interaction actions and the current collaborative interaction context, and eliminate candidate interaction actions that have logical contradictions. Assign a unified timestamp to the final compliant candidate interactive actions, complete the timing synchronization of lines, actions, and screen, and generate synchronous execution instructions.
8. A multi-agent digital human collaborative interaction system, characterized in that, It includes an initialization configuration module, a scene situation awareness module, a multi-agent decision-making module, a collaborative management and control module, a multimodal deduction and driving module, a feedback and policy update module, and an experience iteration module; The initialization configuration module is communicatively connected to the multi-agent decision-making module and the experience iteration module, respectively, and is used to configure a fixed human setting anchor benchmark and dynamic policy branches within the constraint framework for the digital human agents participating in the current collaborative interaction, thereby completing the initialization of the decision model. The scene situation awareness module is communicatively connected to the multi-agent decision-making module and the collaborative control module, respectively. It is used to collect multi-source interaction data of the current collaborative interaction scene, extract and generate scene situation features, and then synchronously distribute them to all digital human agents participating in the current collaborative interaction. The multi-agent decision-making module is communicatively connected to the collaborative management and control module, and has built-in decision-making models corresponding to multiple digital human agents in the system. It is used to generate candidate interactive actions based on the person setting anchor benchmark and scene situation characteristics, and output compliant candidate interactive actions after performing person setting compliance verification. The collaborative management module is communicatively connected to the multimodal deduction driving module, and is used to perform multi-agent collaborative adaptation and timing synchronization on all compliant candidate interactive actions of the digital human intelligent agents participating in the current collaborative interaction, and generate synchronous execution instructions. The multimodal interpretation driving module is communicatively connected to the feedback and strategy update module, and is used to drive the digital human intelligent agent to complete multimodal collaborative interpretation, collect and feedback the interpretation execution results of the current collaborative interaction; The feedback and strategy update module is connected to the multi-agent decision-making module and the experience iteration module respectively. It is used to collect multi-dimensional feedback data, calculate reward values and update the model parameters of the dynamic strategy branch. The update process does not modify the human anchoring benchmark. The experience iteration module is connected to the initialization configuration module to accumulate high-quality collaborative interaction experience and iteratively update the decision-making foundation model of all digital human agents in the system.
9. The multi-agent digital human collaborative interaction system as described in claim 8, characterized in that, The decision model built into the multi-agent decision module adopts a hybrid architecture that connects a feature extraction unit, a semantic encoding unit, and a policy output unit. The input of the feature extraction unit is connected to the scene situation features, the input of the semantic encoding unit is connected to the persona anchoring benchmark and the semantic context data of the current collaborative interaction, and the output of the policy output unit outputs candidate interaction actions.
10. The multi-agent digital human collaborative interaction system as described in claim 8, characterized in that, The feedback and strategy update module has a built-in multi-objective reward calculation unit. The multi-objective reward calculation unit constructs a multi-objective fusion reward function based on four dimensions: persona matching degree, collaborative coherence, scene goal completion degree, and user interaction feedback, and calculates the real-time reward value of each digital human agent participating in the current collaborative interaction.