A modal layered enhanced multi-agent collaborative control method and related device
Through multimodal data fusion and dynamic role migration network, a multi-level strategy system is constructed to solve the problems of system robustness and decision rationality in multi-agent dynamic confrontation environment, and realize efficient collaboration and rapid response in complex scenarios.
Patent Information
- Application Number
- CN202510984997.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-07-17
AI Technical Summary
In complex multi-agent dynamic confrontation environments, existing technologies have difficulty improving system robustness and decision-making rationality. This is especially true in scenarios such as RoboCup 3D simulated football, where static agent role allocation, low communication efficiency, and poor strategy adaptability lead to delayed responses and insufficient confrontation capabilities.
Through multimodal sensor fusion of visual, auditory, spatial, motion and communication modal data, a high-level environmental state representation is generated. Role assignment is performed in combination with a dynamic role migration network, a multi-level strategy system is constructed, and an event triggering mechanism and adversarial meta-learning mechanism are introduced to achieve low-latency strategy synchronization and continuous optimization.
It improves the real-time perception capability, role adaptation capability and confrontation strategy adaptability of the multi-agent system in dynamic confrontation environments, improves the robustness of the system and the rationality of decision-making, and ensures efficient collaboration and rapid response in highly dynamic and high-uncertainty scenarios.
Smart Images

Figure CN120508137B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of multi-agent dynamic confrontation and collaboration, and in particular to a modal layered enhanced multi-agent collaborative control method and related devices. Background Art
[0002] In recent years, with the advancement of artificial intelligence (AI), multi-agent systems, deep reinforcement learning, and game theory, multi-agent collaborative control systems have played a key role in dynamic confrontation scenarios at the intersection of AI and robotic control. Multi-agent collaborative decision-making and confrontational control have been widely researched and applied in various fields. Typical dynamic confrontation scenarios include RoboCup 3D simulated football matches, intelligent basketball tactical simulations, wargame confrontation simulation systems, and intelligent gaming platforms such as Go and chess. These scenarios all involve the complex interplay of high dynamics, high uncertainty, and complex confrontational strategies among multiple agents, making them important testing grounds for research on autonomous multi-agent collaborative and confrontational intelligent decision-making systems.
[0003] Therefore, how to enhance system robustness and improve decision-making rationality in a complex multi-agent dynamic confrontation environment has become a technical problem that needs to be solved urgently. Summary of the Invention
[0004] In order to enhance the robustness of the system and improve the rationality of decision-making in a complex multi-agent dynamic confrontation environment, the present application provides a modal hierarchical enhanced multi-agent collaborative control method and related devices.
[0005] In the first aspect, the present application provides a modal hierarchical enhanced multi-agent collaborative control method using the following technical solutions:
[0006] A modal hierarchical enhanced multi-agent collaborative control method, comprising:
[0007] Through multimodal sensors, real-time collection and fusion of visual, auditory, spatial, motion and communication modal data can generate high-level environmental state representation;
[0008] Based on the high-level environmental state representation, combined with historical data and real-time situation, dynamic roles are assigned to each agent through a dynamic role migration network to achieve adaptive role switching and conflict coordination;
[0009] Build a multi-level strategy system, decomposing decisions into a high-level strategy layer, a middle-level collaboration layer, and a low-level control layer, generating tactical intent, collaborative relationships, and physical control instructions respectively;
[0010] Based on the event trigger mechanism, intent-level communication information is compressed and transmitted to achieve low-latency policy synchronization under key events;
[0011] Introducing an adversarial meta-learning mechanism to quickly adapt to changes in opponent strategies through strategy prototype retrieval and online fine-tuning;
[0012] Execute underlying control instructions and provide real-time feedback on execution status to form a closed-loop optimization;
[0013] Dynamically adjust role allocation and strategy weights based on feedback data to achieve continuous evolution of the multi-agent collaborative system.
[0014] Optionally, the step of collecting and fusing visual, auditory, spatial, motion, and communication modality data in real time through multimodal sensors to generate a high-level environmental state representation includes:
[0015] Real-time collection and fusion of visual, auditory, spatial, motion and communication modality data through multimodal sensors;
[0016] Through visual modality analysis, panoramic and local visual information are converted into global coordinates and a field potential thermal map is generated.
[0017] Receive sound signals through the auditory modality and use Bayesian estimation and particle filtering to locate the sound source;
[0018] By integrating inertial navigation data and scene analysis results through spatial modality, a spatial coverage probability field is constructed;
[0019] Correct movement abnormalities by modeling inverse kinematics and joint timing relationships through motion modalities;
[0020] Construct the communication graph topology and generate node representations through communication modalities;
[0021] Multimodal time series data are aligned through a dynamic differentiable interpolation layer, and the contrastive loss is used to constrain the similarity between modalities to generate a unified high-order state representation.
[0022] Optionally, the dynamic role migration network includes:
[0023] The dynamic role migration network uses a dual-attention mechanism to calculate the role adaptation score based on the global strategic demand map and local perception information;
[0024] Periodically review role assignments through an event-triggered mechanism, and trigger role reconstruction when policy conflicts are detected.
[0025] Output the agent's probabilities of tendencies towards multiple roles, combine them with conflict adjustment factors to avoid resource competition, and achieve decentralized collaborative allocation.
[0026] Optionally, the high-level strategy layer solves the global tactical field weight distribution through Nash equilibrium to generate attack and defense tendency coefficients and risk tolerance instructions;
[0027] The middle collaboration layer generates collaborative relationship topology based on the graph attention mechanism and eliminates intention overlap through the conflict coordinator;
[0028] The underlying control layer uses a prediction error compensation model and hierarchical model predictive control to decode the collaborative intention into continuous action instructions under physical constraints.
[0029] Optionally, the intent-level communication information includes strategic-layer hotspot area boundaries, tactical-layer action ontology libraries, and execution-layer parameter encodings;
[0030] Compress semantic information through variational autoencoders and design channel utility functions to optimize communication resource allocation;
[0031] The receiving end verifies consistency, calculates coordination and dynamically reconstructs intent through a distributed intent fuser to ensure tactical spatiotemporal consistency.
[0032] Optionally, the step of introducing an adversarial meta-learning mechanism to quickly adapt to changes in the opponent's strategy through strategy prototype retrieval and online fine-tuning includes:
[0033] The adversarial meta-learning framework in the adversarial meta-learning mechanism is introduced to retrieve historical strategy meta-features through the strategy prototype library and generate adversarial perturbation samples in combination with the opponent's behavior pattern matrix;
[0034] A two-stream update rule is used to quickly fine-tune the policy network and value network, and the policy distillation loss is used to achieve continuous evolution of the meta-knowledge base.
[0035] Optionally, the step of executing the underlying control instructions and providing real-time feedback on the execution status to form a closed-loop optimization includes:
[0036] Collect action execution deviation, collision risk and tactical achievement data in real time, and compensate for execution errors through inverse kinematics models;
[0037] Dynamic role reconstruction and strategy weight adjustment are triggered based on feedback, while the control accuracy under communication delay is optimized through the Kalman gain matrix.
[0038] In a second aspect, the present application proposes a modality-layered enhanced multi-agent collaborative control device, comprising:
[0039] The information acquisition module is used to collect and fuse visual, auditory, spatial, motion and communication modal data in real time through multimodal sensors to generate a high-level environmental state representation;
[0040] A role assignment module is used to assign dynamic roles to each agent through a dynamic role migration network based on the high-level environmental state representation, combined with historical data and real-time situation, to achieve adaptive role switching and conflict coordination;
[0041] The system building module is used to build a multi-level strategy system, decomposing decisions into a high-level strategy layer, a middle-level collaboration layer, and a low-level control layer, generating tactical intent, collaborative relationships, and physical control instructions respectively;
[0042] The transmission module is used to compress and transmit intent-level communication information based on the event trigger mechanism to achieve low-latency policy synchronization under key events;
[0043] The fine-tuning module is used to introduce an adversarial meta-learning mechanism to quickly adapt to changes in the opponent's strategy through policy prototype retrieval and online fine-tuning;
[0044] The optimization module is used to execute the underlying control instructions and provide real-time feedback on the execution status, forming a closed-loop optimization;
[0045] The control module is used to dynamically adjust role allocation and strategy weights based on feedback data to achieve continuous evolution of the multi-agent collaborative system.
[0046] In a third aspect, the present application provides a computer device, comprising: a memory and a processor, wherein the processor executes the method described above when running computer instructions stored in the memory.
[0047] In a fourth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on a computer, enable the computer to execute the method described above.
[0048] In summary, this application uses multimodal sensors to collect and fuse visual, auditory, spatial, motion, and communication modal data in real time to generate a high-level representation of the environment state. Based on the high-level representation of the environment state, a dynamic role is assigned to each agent through a dynamic role migration network. A multi-level strategy system is constructed to decompose the decision into a high-level strategy layer, a middle-level collaboration layer, and a low-level control layer, generating tactical intent, collaborative relationships, and physical control instructions respectively. Low-latency strategy synchronization under key events is achieved based on an event trigger mechanism. An adversarial meta-learning mechanism is introduced to retrieve and fine-tune strategy prototypes online. The low-level control instructions are executed and the execution status is fed back in real time to form a closed-loop optimization. The role allocation and strategy weights are dynamically adjusted based on the feedback data to achieve the continuous evolution of the multi-agent collaborative system. This system realizes hierarchical strategic decision-making, and at the same time has the ability to communicate event-driven compressed intentions. It integrates an adversarial meta-learning evolver to achieve extremely fast adaptation to new opponents and continuous strategy self-evolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic diagram of the computer device structure of the hardware operating environment involved in the embodiment of the present application;
[0050] Figure 2 This is a flow chart of the first embodiment of the modal hierarchical enhanced multi-agent collaborative control method of the present application;
[0051] Figure 3 It is a structural block diagram of the first embodiment of the modal layered enhanced multi-agent collaborative control device of the present application. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below through the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0053] Reference Figure 1 , Figure 1 This is a schematic diagram of the computer device structure of the hardware operating environment involved in the embodiment of the present application.
[0054] like Figure 1 As shown, the computer device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display and an input unit, such as a keyboard. Optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a wireless fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk storage device. The memory 1005 may also be a storage device independent of the processor 1001.
[0055] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0056] like Figure 1 As shown, the memory 1005 as a storage medium may include an operating system, a network communication module, a user interface module, and a modality-layered enhanced multi-agent collaboration control program.
[0057] exist Figure 1In the computer device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in this application can be set in the computer device, and the computer device calls the modal layered enhanced multi-agent collaborative control program stored in the memory 1005 through the processor 1001, and executes the modal layered enhanced multi-agent collaborative control method provided in the embodiment of this application.
[0058] The present invention provides a method for multi-agent collaborative control with modal layering enhancement. Figure 2 , Figure 2 This is a flow chart of the first embodiment of the modal hierarchical enhanced multi-agent collaborative control method of this application.
[0059] In this embodiment, the modality layered enhanced multi-agent collaborative control method includes the following steps:
[0060] Step S10: Collect and fuse visual, auditory, spatial, motion and communication modal data in real time through multimodal sensors to generate a high-level environmental state representation.
[0061] In practice, RoboCup 3D soccer simulation is a highly complex test platform recognized internationally in the field of artificial intelligence. It requires multiple heterogeneous robotic agents to perform real-time perception, role allocation, tactical planning, coordinated control, and precise execution in a dynamic, high-speed, and randomly disturbed environment. Existing technologies for RoboCup 3D soccer simulation commonly suffer from the following technical pain points:
[0062] The traditional RoboCup 3D robot collaboration system uses static role allocation, which fixes the responsibilities of each agent. For example, it uses a pre-set role allocation system of forwards, midfielders, defenders, etc., which cannot be dynamically adjusted according to the competition format and is difficult to respond to unexpected situations during the game. It can easily put the team into a passive position and respond lags when facing sudden strategic changes or temporary offensive and defensive transformations from opponents. In terms of communication, existing methods mostly use full-frequency broadcast information communication, which leads to channel congestion and information redundancy, making it difficult to ensure low-latency and high-quality strategy sharing under high real-time conditions. In terms of strategy, traditional rule-based or static reinforcement learning models have difficulty adapting quickly to strategic changes of different types of opponents during the game. They are easily cracked or put into a passive position when facing advanced adversarial strategies. Traditional strategies have certain vulnerabilities and their adversarial capabilities are poor. During training, centralized reinforcement learning methods face the curse of dimensionality in the high-dimensional strategy space of multiple agents, with long training cycles and difficult to achieve stable convergence.
[0063] In addition to RoboCup 3D simulated football, dynamic confrontation scenarios are also widely present in the following fields: Basketball tactical deduction simulation: involving dynamic role switching (point guard, scorer, defender), on-the-spot tactical reconstruction and high-speed on-field communication, there is a high demand for real-time game solving and dynamic collaboration; Wargame confrontation simulation system: The scenario contains multiple arms and multiple equipment intelligent units, which require dynamic formation adjustment, on-the-spot battlefield situation reasoning and efficient collaborative communication under low-channel conditions; Chess (Go, chess) intelligent game platform: Although it is a complete information game on the surface, high-level AI models also have an urgent need for "strategy perturbation induction" and "rapid adaptation of diverse opponents", especially when confronting changing strategy AI or strategy confusion systems.
[0064] Common challenges in these complex, dynamic confrontation scenarios include: Agents must be able to dynamically reconfigure their roles in real time (flexibly switching from "leader" to "follower"); multimodal information (motion, spatial, and communication) must be integrated and processed to form a high-level perception of the environment; multiple agents must achieve intention-level compressed communication and policy consensus under limited bandwidth and high latency risks; and meta-learning capabilities are required to adapt to changing opponent strategies and achieve extremely fast adaptive policy transfer. This embodiment uses the RoboCup 3D simulated soccer game as an example to illustrate the multi-agent collaborative control scenario.
[0065] In specific implementations, multimodal sensors (underlying physics engines) interface with virtual simulation environments to collect and fuse visual, auditory, spatial, motion, and communication modal data in real time. The visual modality utilizes panoramic and local vision sensors to acquire dynamic information about the ball, opponents, teammates, poles, and boundaries in the environment, and uses sensors to obtain the current state of the world model. The auditory modality receives strategic information, instruction content, and directional cues from teammates, opponents, or system broadcasts through auditory sensors, enabling spatial positioning based on the direction of the sound source to assist in determining event occurrences. The spatial modality integrates inertial navigation (gyroscopes, accelerometers), scene analysis, and the world model to obtain the relative position of the intelligent agent and real-time position information in the global coordinate system. The motion modality provides individual motion state and joint execution feedback, combining joint angle feedback, gait engine status, and foot pressure sensors to perceive the robot's motion state, stability, and landing conditions in real time. The communication modality receives state and intention information from other intelligent agents in real time. All perception information is fused through a multimodal encoder to form a high-order state representation, providing rich and consistent environmental input for subsequent dynamic role migration and strategy generation, ensuring that the perception system is noise-resistant and real-time in high-interference and high-dynamic scenarios.
[0066] It should be noted that the step of collecting and fusing visual, auditory, spatial, motion and communication modal data in real time through multimodal sensors to generate a high-order environmental state representation includes: collecting and fusing visual, auditory, spatial, motion and communication modal data in real time through multimodal sensors, parsing panoramic and local visual information through the visual modality, converting polar coordinates into global coordinates and generating a field potential thermal map; receiving sound signals through the auditory modality, and locating the sound source using Bayesian estimation and particle filtering; integrating inertial navigation data and scene analysis results through the spatial modality to construct a spatial coverage probability field; modeling inverse kinematics and joint timing relationships through the motion modality to correct motion abnormalities; constructing a communication graph topology and generating node representations through the communication modality; aligning multimodal time series data through a dynamic differentiable interpolation layer, combining contrast loss to constrain the similarity between modalities, and generating a unified high-order state representation.
[0067] In practice, in the RoboCup 3D dynamic competition scenario, the intelligent agent must first integrate five heterogeneous modal data types in real time: vision, hearing, space, motion, and communication, to achieve comprehensive environmental perception and precise decision-making and control. Each intelligent agent (player) has a camera mounted at the center of its body, with a 360-degree field of view, enabling it to see all objects on the court.
[0068] For visual modality information, parse the polar coordinate observation values in the See message. or , extract features from visual information, and convert the acquired position information from polar coordinates to global coordinates. The formula is as follows:
[0069]
[0070] in, and Represent two parameters in polar coordinates, is the coordinate of the intelligent body (obtained through WorldModel::getMyPosition()). The detection result is converted into the field potential thermal position using the following formula:
[0071]
[0072] in, Indicates hot position The value at the court indicates the position The intensity of the "threat level" or "importance" of a location. Higher values indicate more important locations. Indicates the total number of target objects on the field (such as footballs, opponent players, teammate players, goals, etc.), Indicates the The location coordinates of the target object. Represents the standard deviation of the Gaussian function, which controls the width of the Gaussian kernel (i.e., the range of influence). is the category weighting coefficient (football > opponent > teammate > goal). This formula innovatively reflects the global threat level and opportunity distribution of the field through "continuous heat mapping", providing intuitive tactical input for the strategy allocation layer. Then, the target type embedding vector is constructed. :
[0073]
[0074] Generate an embedding vector for the target type, allowing the system to distinguish between a ball (which needs to be actively approached) and an opponent (which needs to be avoided). is the weight matrix that maps discrete types to continuous vector space.
[0075] For the auditory modality, the simulation platform microphone array interface is used to obtain a collection of sound signals from teammates, opponents, or system broadcasts: in, for the sound category (teammate instructions, opponent calls, referee signals); is the volume (distance back-calculation indicator); is the relative angle of the sound source. The sound source position is estimated using Bayesian estimation: By using particle filtering to perform multi-frame fusion, auditory perception can be used as an auxiliary signal to correct visual uncertainty areas (such as when there is occlusion), forming "auditory blindness compensation". , extract the direction angle and message content . And encode the direction, mapping the angle to the bearing vector:
[0076]
[0077] The message is compressed and event-driven encoding is performed to compress the 20-byte message into 4 bits. The auditory event vector is extracted through the event encoder: This vector is used as input for mid- and high-level strategies, especially as a non-visual supplementary criterion in strategy adjustment. For event encoders.
[0078] For spatial modal information, the body posture is first estimated through the gyroscope angular velocity and foot pressure, and the orientation is updated through the quaternary differential equation:
[0079]
[0080] in, Represents the current posture quaternion of the intelligent body, used to describe its orientation, is the angular velocity vector measured by the gyroscope, which represents the rotation rate of the agent around the three axes. is the time derivative of the quaternion, which represents the rate of change of the attitude.
[0081] According to the self-position and teammate position perception interface provided by the simulation platform:
[0082]
[0083] Representing an agent and agents The distance between them and the angle matrix are calculated at the same time:
[0084]
[0085] Represents the agent Pointing to the agent direction angle.
[0086] The above distance matrix and the angle matrix As the edge features of the graph, a time-varying spatial relationship graph is constructed, and an innovative event-varying spatial relationship graph is constructed for the input of the neural network. Combined with the ionized area network, the distribution of the agent's coverage ability is mapped into a spatial probability field:
[0087]
[0088] Indicates the location The strength of spatial coverage at a location, that is, the degree to which the location is "controlled" or "sensed" by the agent team. The standard deviation of the Gaussian distribution controls the "width" or "attenuation speed" of the coverage area. This coverage field is used to assess the threat level of the area and the formation of gaps in real time.
[0089] For motion modal data, obtain motion state quantities in real time:
[0090]
[0091] First, the inverse kinematics model is built, and the target pose of the end effector is given. , solve for the joint angle :
[0092]
[0093] in is the forward kinematics function, is the joint angle vector to be solved, is the default value of the joint angle (such as standing posture), is the regularization coefficient, which is used to penalize solutions that deviate too much from the default posture. Then, the joint timing relationship is constructed through LSTM:
[0094]
[0095] in, is the joint angle at the current moment, is the joint angular velocity at the current moment, , is the weight matrix and bias vector; is the hidden state at the previous moment, is the hidden state at the current moment, used to capture the temporal dependencies of joint motion. Motion anomalies are corrected through a self-supervised dynamic noise estimator. A motion feedback loss metric is designed:
[0096]
[0097] is the desired joint angle, The squared error is calculated for the actual joint angles, measuring the difference between the desired and actual execution. This error is then fed back to the underlying SAC algorithm controller in real time for adaptive fine-tuning.
[0098] For communication modality data (Say&Hear), first build a communication graph and define the graph ,node is the agent, edge weight is the communication frequency. Topological embedding generates node representations through graph convolution (GCN):
[0099]
[0100] in is the adjacency matrix, is the degree matrix, For the The node feature matrix of the layer, For the The weight matrix of the layer.
[0101] Finally, the features of each modality are fused. Considering the asynchronous and frequency-different characteristics of different modalities, a dynamic differentiable interpolation layer is designed to align the multimodal time series:
[0102]
[0103] in, Indicates time The fused low-dimensional feature vector (used for subsequent policy network input), Indicates the High-dimensional feature vectors of each modality (such as visual features, auditory direction vectors, motion states, etc.), For the The weight of the modal feature represents its effect on the current moment The degree of contribution, To control the smoothness of the weight distribution (smaller is sharper, larger is smoother), By lightweight MLP Dynamically generate and input high-dimensional features and timestamp , optimize the timing offset of alignment.
[0104] The integration of a dynamic time weighting mechanism enables real-time compensation of multimodal asynchronous data, forming a unified time-series input stream, significantly alleviating the negative impact of asynchronous sampling on policy update delays. A learnable time-series deviation estimation network is introduced to replace fixed time difference weight calculation.
[0105] Encode each modality, project the features of each modality into a shared representation space, use nonlinear projection to enhance the interaction between modalities, and combine contrast loss to constrain similarity:
[0106] Indicates the The original eigenvectors of the modes, is the first layer linear transformation parameter, is the linear transformation parameter of the second layer, and GELU is the activation function to enhance the nonlinear expression ability. For the The embedded representation of the modalities in the shared space.
[0107] Contrastive loss design:
[0108]
[0109] in, is the similarity function, is a semantically related modality, For irrelevant modes, is the temperature coefficient, which controls the smoothness of contrast loss.
[0110] Convert the multimodal raw inputs into the same vector space to ensure that query, key, and value operations can be performed on each other during subsequent attention fusion.
[0111] A hierarchical cross-attention mechanism is adopted to introduce intra-modal and inter-modal dual-stage attention to enhance multimodal interaction capabilities.
[0112] First, the first stage of self-attention within the modality:
[0113]
[0114] Indicates the Input features of each modality (such as vision, hearing, motion, etc.), They are the linear projection matrices of Query, Key, and Value respectively. is modal The output features after self-attention integrate information within each modality, extracting the structure and dependencies within the modality, and providing a richer and more structured representation for subsequent cross-modal fusion.
[0115] Second stage inter-modal cross attention:
[0116]
[0117] Represented as the output of the visual modality (as a Query), Outputs of other modalities (such as hearing, movement) (as Key and Value), is the projection matrix of the corresponding mode, Represented as visual modality and The output features after modal fusion, It is a scaling factor to prevent the gradient from disappearing due to the dot product being too large.
[0118] This structure allows each modality to first optimize its own representation and then aggregate cross-modal information.
[0119] Then the global state memory unit (GRU) is introduced to dynamically adjust the gating weights based on historical information:
[0120]
[0121] in, represents the global state vector at the previous moment, Represents the embedded features of the current visual modality, Represents the fusion features of other modalities (auditory, spatial, motion, communication) at the current moment. GRU is a gated recurrent unit used to update the global state. represents the global state vector at the current moment, represents the learnable weight matrix, Indicates the The gating weight of each modality indicates its contribution to the final fusion output. Indicates the The features of each modality after intra-modal attention, Indicates the The head features generated by the modalities in the cross attention, Represents the final multimodal fusion feature, which is used in the subsequent strategy network.
[0122] The temporal dependencies are captured by GRU memory units, and the gated weights fuse local and global information.
[0123] With the visual modality as the "core leading query" and the auditory, spatial, motion, and communication modalities as the "response modalities," this approach achieves dynamic master-slave information focusing in dynamic environments. A multi-head mechanism is introduced to ensure diversity and robustness in cross-modal feature capture.
[0124] Introducing gating weights , guided by the visual modality, the participation of each modality is regulated to achieve scene-sensitive dynamic fusion. For example, when the system detects auditory interference, it adaptively reduces the weight of the auditory modality and increases the proportion of the spatial and communication modalities. Dynamically adjust the contribution of each modality to the final output:
[0125]
[0126] Step S20: Based on the high-level environmental state representation, combined with historical data and real-time situation, dynamic roles are assigned to each agent through a dynamic role migration network to achieve adaptive role switching and conflict coordination.
[0127] It can be understood that the dynamic role migration network includes: the dynamic role migration network uses a dual attention mechanism to calculate the role fitness score based on the global strategic demand map and local perception information; periodically reviews the role allocation through an event trigger mechanism, and triggers role reconstruction when a strategy conflict is detected; outputs the agent's tendency probability for multiple roles, combines the conflict adjustment factor to avoid resource competition, and realizes decentralized collaborative allocation.
[0128] It should be noted that after completing multimodal information fusion, the system enters the dynamic role calculation and migration phase. Based on multimodal sensory input and historical data, a dynamic role migration network is constructed. Instead of fixed agent positions, agents dynamically assess their appropriate roles (offense, support, defense, or interference) based on the on-field situation, positional relationships, teammate coverage, and predicted opponent intentions. Based on the current scenario and historical trend information, combined with team collaboration goals, the system autonomously assesses the optimal role allocation for each agent. By constructing a temporal memory and relationship network, this module not only considers individual capabilities and local positions, but also perceives changes in opposing strategies in real time, dynamically adjusting the switching between offense, support, and defense roles. A game equilibrium allocation mechanism is also integrated to resolve role conflicts and overlaps in real time. When detecting a change of possession, an open space, or a sudden change in opponent strategy, the system triggers role reallocation, enabling millisecond-level adaptive migration and tactical reconstruction, ensuring rapid response and flexible coordination within the multi-agent swarm.
[0129] It is understandable that in dynamic confrontation scenarios, intelligent agents need to dynamically switch roles (forward / support / defense) according to the real-time situation (ball possession, opponent's strategy, teammate status). Multiple intelligent agents competing for the same role can easily lead to a decrease in overall efficiency, and multiple players attacking at the same time can easily lead to collisions and fouls and be sent off. Therefore, team collaboration and balance are particularly important in this.
[0130] The system obtains a complete description of the environmental status through the fusion of the previous five modes of vision, hearing, space, movement, and communication, including the position and movement trend of the ball, the distribution of enemy and friendly formations, the situation of open areas, the intention of teammates, and the limits of their own movement capabilities. Through multi-modal cross-sensing information, the system can capture "changes in events on the field" and "strategic opportunities" in real time, and then build a high-dimensional situation vector internally. .in is the number of perceptual dimensions, is the spatial grid resolution, For the length of the time window, this high-dimensional situation vector contains not only static information (such as distance and angle), but also dynamic information (such as the ball speed change trend, opponent's pressing speed, teammates' coordination plan), providing a basis for role allocation.
[0131] In actual games, task requirements will constantly change in different positions and situations. To this end, this system innovatively proposes the idea of "state-driven task priority generation": triggering events based on the perceived on-field situation and environmental situation: external triggering events include: ball possession changes, opponents breaking through the defense line, key gaps, balls out of bounds, free kicks, etc.; internal triggering events include: the agent detecting that its own target task completion rate deviates from the threshold, the communication channel receives a teammate's strategy adjustment signal, etc. When the conditions are met, When the role reconstruction is triggered, is the preset threshold, The system uses a sampling period. After a triggering event is sensed, a multimodal fusion module generates a "high-dimensional state description tensor" to calculate the most pressing tactical requirements at the current time (such as ball control and advancement, receiving support, defensive cover, and gap filling). High-priority tasks are included in a candidate set of roles, and each role is dynamically assigned an urgency and suitability score to form a candidate task pool. Each agent selects the most suitable role intention from the candidate pool based on its perception state, its own ability model, and the on-field strategic situation. A "conflict adjustment factor" is also introduced to avoid resource conflicts caused by multiple agents simultaneously preferring the same high-priority role.
[0132]
[0133] in Select the role for the current The number of agents, is the expected quantity at strategic equilibrium, To adjust the sensitivity parameter.
[0134] Dynamic character construction is not just an independent decision-making process for a single agent, but rather a highly collaborative problem. This embodiment introduces the "global-local dual-layer fusion concept":
[0135] At the global level: Based on the overall formation and the distribution of field control areas, a team's strategic needs map is constructed in real time. Based on the global situation assessment results (from the global strategy planning module), the role allocation probability matrix of each agent is dynamically normalized and suppressed to avoid resource concentration or role vacancies. At the local level: Each agent dynamically adjusts its candidate role preferences based on local perception information, relative position to surrounding teammates, and role intentions to avoid role conflicts and duplication. Through information exchange among multiple agents, a decentralized, distributed, and coordinated dynamic role allocation strategy is formed.
[0136] The role suitability score is calculated based on the dual attention mechanism:
[0137]
[0138] in, For global level role assessment, For local level role assessment, is the global feature dimension, For intelligent agents The field collection, is the conflict adjustment factor, Representing a role intensity of competition.
[0139] Dynamic competitive environments are characterized by extreme uncertainty. A single assignment cannot adapt to continuous changes. Therefore, this paper designs an event-triggered dynamic adjustment mechanism: When a change occurs on the field, such as a change in ball possession, a key area breach, a sudden gap in the open area, or a task timeout, the system immediately restarts the role assignment process. In the absence of event triggers, the system periodically reviews and fine-tunes role assignment probabilities, achieving a dynamic and smooth transition and avoiding frequent role oscillation. Ultimately, the system does not directly assign fixed role labels, but instead outputs agent preference probabilities and compatibility scores for multiple roles. In complex scenarios, agents select the optimal role based on a comprehensive score. The system remains open, allowing for online adjustment of role definitions and priority rules based on tactical strategies. The core of this dynamic role construction module lies in "dynamically generating role and task assignments based on multimodal situational dynamics, combining global strategic requirements with local adaptation." This method enables adaptive adjustment and real-time transition of agent roles in highly dynamic and non-deterministic environments, improving the coordination, agility, and tactical execution of multi-agent systems in complex competitive scenarios.
[0140] Step S30: Build a multi-level strategy system, decompose the decision into a high-level strategy layer, a middle-level collaboration layer, and a bottom-level control layer, and generate tactical intentions, collaborative relationships, and physical control instructions respectively.
[0141] It should be noted that the high-level strategy layer solves the global tactical field weight distribution through Nash equilibrium to generate attack and defense tendency coefficients and risk tolerance instructions. The middle-level collaboration layer generates collaborative relationship topology based on the graph attention mechanism and eliminates overlapping intentions through a conflict coordinator. The bottom-level control layer uses a prediction error compensation model and hierarchical model predictive control to decode collaborative intentions into continuous action instructions under physical constraints. This organically decouples and links high-level strategic planning, middle-level multi-agent collaborative decision-making, and bottom-level motion control to achieve an optimal balance between real-time performance, stability, flexibility, and strategic planning for the multi-agent system.
[0142] Understandably, the system is designed with a multi-layered strategy architecture, decoupling global decision-making from local execution into a high-level strategy layer, a mid-level collaboration layer, and a low-level physical control layer. The high-level strategy layer generates tactical-level intent instructions, such as zone defense, quick counterattacks, or high-pressure, based on global scenario input and dynamic role configuration. The mid-level collaboration layer, centered around a graph attention mechanism, calculates collaborative relationships between agents, pass probability distributions, coordinated attacks, and support paths. The low-level control layer, based on deep reinforcement learning algorithms, achieves high-precision motion control and physical constraint adjustment. Relying on a gait engine, inverse kinematics, and a SAC algorithm controller, it achieves precise motion output under physical constraints, such as dynamic acceleration, offensive and defensive switching, and precise arcing shots. This layered architecture effectively avoids the problem of strategy dimension explosion and, through a curriculum-based training mechanism, enables steady progression from simple rule-based confrontation to complex intelligent confrontation, improving the model's generalization and robustness.
[0143] In practice, the high-level strategy layer, at the top of the decision-making pyramid, generates a global "strategy-driven map" in real time based on global perception information and the tendency matrix output by dynamic role allocation. This is achieved by combining: field zoning potential field modeling, key event prediction (such as the probability of ball possession changes and the probability of offensive channel formation), and tactical priority planning to construct a multi-objective optimization problem:
[0144]
[0145] in, represents the positional advantage reward (such as ball possession area dominance), Indicates the success rate of pressure (interference rate on enemy transmission and control), Indicates the completeness of the gap coverage. Indicates risk penalty items (including offside risk and collision risk), represents the expectation operator, is the discount factor, are the weight coefficients of each sub-goal, which are used to adjust the importance of different goals.
[0146] Based on the global situation tensor ( is the pitch rasterization resolution, Generate the team's offensive and defensive strategy primitives (number of feature channels), define tactical phases (ball control, quick counterattack, zone defense, etc.), and output multi-objective optimization weights and tactical phase probability distribution, and outputs the strategic instruction vector ,in is the attack and defense tendency coefficient, For risk tolerance.
[0147]
[0148] Define the state-action value function (Q function) , used to evaluate the expected value of long-term strategic returns that can be obtained by adopting a certain strategy (or action) under the current state s, Indicates that the status The real-time strategic return under the opponent's strategy Solve the Nash equilibrium of , maximize our minimum benefit guarantee. When the opponent's defense line is detected ( ) generates a "Quick Side Transfer" instruction, activating the forward insertion coefficient of the wing agent If the ball is lost and , triggering the "high-pressure" mode and reconstructing the defensive role tendency probability matrix.
[0149] We further introduce the deep reinforcement learning (DRL) framework to build a two-layer policy optimization network:
[0150]
[0151] in, (weighted sum of target returns), is the risk penalty item, is the discount factor. Then the adaptive potential energy field is modeled:
[0152]
[0153] potential energy field Defined by the court zones, dynamic potential field Real-time updates through LSTM capture the opponent's movement trends. They are the weight coefficients of the static potential energy field and the dynamic potential energy field, respectively, which dynamically adapt to the opponent's strategy and improve the robustness and real-time performance of the global strategy.
[0154] The high-level strategy layer outputs the dynamic “tactical field” weight distribution and sends it down to the middle-level collaboration layer to form dynamic guidance of tactical templates.
[0155] The middle-level collaboration layer serves as the "cooperation hub" among multiple agents, responsible for information mapping and coordination between the global and individual components. It dynamically combines high-level tactical field distribution, agent role allocation probability, and current local spatial distribution information to generate a distributed collaborative intention matrix:
[0156]
[0157] Agent At the moment Preference for each task, Local spatial tensor (reflecting the spatial opportunity and risk density of the area where the agent is located), Represents the global tactic tensor.
[0158] High-level strategic directives Decomposed into a multi-agent writing task topology graph , where the nodes Represents tactical actions (cross running, triangle passing, pressure double team, etc.), side Representing spatiotemporal dependencies between actions.
[0159] This embodiment designs a "conflict coordinator" mechanism. When multiple agents intend to gather the same resource (such as the same receiving position), the conflict coordinator reallocates priorities based on the game balance principle.
[0160] First, define the conflict resolution function:
[0161]
[0162] in, Represented as an agent strategy, is the individual preference strategy, is the communication neighbor set, is the collaborative consistency weight.
[0163]
[0164] Defining adjusted collaboration intent , which is used to dynamically avoid strategy overlap or conflict when multiple intelligent agents (such as robots, football players, etc.) make collaborative decisions. Indicates the original collaborative intention, represents the conflict cost, is the conflict penalty coefficient. This mechanism can dynamically avoid strategy overlap among multiple agents, improving overall coordination efficiency and system stability.
[0165] The bottom-level execution layer is responsible for decoding the collaborative intent tensor and specific task assignment results from the middle-level collaboration layer into continuous control instructions, executing the agent's micro-operation behaviors in complex adversarial scenarios. This layer innovatively introduces the "Predictive Error Compensation Model" (PECM), which achieves a closed-loop correction of the "model-perception" by reversely correcting the deviation between historical state-action pairs and real-time perception:
[0166]
[0167] in, It is the action sequence issued by the middle layer. is the current perception state, For the early forecast status, In order to compensate the gain matrix, in addition, the underlying execution layer designs a multi-task multi-constraint optimizer:
[0168]
[0169] Effectively ensure smooth movement, collision avoidance, and posture stability. is the control instruction to be optimized, is the energy consumption cost function, is the collision risk cost function, is the attitude stability cost function, is the weight coefficient of each constraint.
[0170] Collaboration layer task topology Convert to continuous motion control instructions , satisfying the dynamic constraints .
[0171] At the same time, hierarchical model predictive control (HMPC) is designed:
[0172]
[0173]
[0174] in is the reference state trajectory, is the actual state trajectory, To control the input sequence, is the state error weight matrix, is the control input weight matrix, Include opponent position Then, according to the intensity of the environmental disturbance Adaptively adjust the cost function weights:
[0175]
[0176] Represents the weight of the cost function after dynamic adjustment at the current moment, represents the initial cost function weight (baseline value without disturbance), is the current environmental disturbance intensity (such as opponent interference, noise, uncertainty, etc.), It serves as a reference for disturbance intensity (for normalization) to improve control accuracy and anti-interference capabilities in complex adversarial scenarios.
[0177] Finally, we implement cross-layer coordination mechanisms and ensure policy consistency by defining semantic consistency constraints on inter-layer interfaces.
[0178]
[0179] represents the inter-layer policy consistency loss, represents the square of the Frobenius norm, which indicates the degree of difference between matrices or tensors. Representing a tactical map Strategic Directives The gradient, Representing a tactical map Execute the action The gradient of ,ensures the consistency of gradient propagation from strategic instructions to action execution.
[0180] Dynamic reconfiguration trigger: When a mismatch between inter-layer strategies is detected (such as the deviation between the actual running position and the tactical map), ), triggering local re-planning of the middle collaborative layer to avoid global strategy shock.
[0181] The three layers of strategy don't flow in a one-way manner, but rather form a real-time feedback loop. The high-level strategy layer receives execution deviations and actual situation data from the middle and bottom layers, dynamically updating the tactical field potential energy distribution. The middle-level collaboration layer continuously adjusts the intention allocation tensor based on the bottom-level execution feedback, achieving self-adaptation. The bottom-level execution layer uses the PEC model to correct execution errors in real time, achieving precise execution at the agent level.
[0182] This hierarchical decision-making architecture uses multimodal perception as input and dynamic role allocation output as a guiding clue, mapping highly complex dynamic confrontation scenarios into a hierarchical and controllable problem system. This three-layer linkage not only ensures the achievement of overall strategic goals but also maximizes the autonomy and flexibility of intelligent agents.
[0183] Step S40: Based on the event triggering mechanism, compress and transmit intent-level communication information to achieve low-latency policy synchronization under key events.
[0184] It should be noted that the intent-level communication information includes the boundaries of hot spots at the strategic layer, the action ontology library at the tactical layer, and the parameter encoding at the execution layer; semantic information is compressed through the variational autoencoder, and the channel utility function is designed to optimize the allocation of communication resources; the receiving end verifies consistency, calculates coordination, and dynamically reconstructs intent through a distributed intent fuser to ensure tactical spatiotemporal consistency.
[0185] In practice, to address multi-agent communication congestion and latency in dynamic confrontation scenarios, the system innovatively designs an event-triggered, intent-level sparse communication protocol. When key events (such as ball possession changes, opponent breakthroughs, and key openings) are detected, the system uses a lightweight compression algorithm to send only packets containing strategic intent and key status summaries. During non-critical moments, local prediction and inference are used to maintain behavioral consistency, reducing redundant communication overhead. Furthermore, the communication module integrates with the higher-level strategy layer to support intent-level synchronization and team tactical consensus, ensuring rapid dissemination of decisions to each agent and significantly reducing communication resource consumption and latency risks in the simulation environment.
[0186] It's important to note that after completing the aforementioned multimodal perception, dynamic role construction, and hierarchical decision-making, the agents already possess clear global and local tactical orientations and mission planning. However, in adversarial scenarios, even a single agent with a comprehensive autonomous strategy cannot guarantee system-level integrity, flexibility, and robustness solely through its own perception and reasoning when faced with rapid change, uncertain events, and adversary strategy perturbations. Therefore, the innovative "intention-level communication" mechanism becomes the core link between agents to achieve high-level collaboration, avoid conflicts, and jointly complete tactical missions. Through a three-level communication mechanism (strategic intent broadcast, tactical semantic sharing, and action parameter synchronization), cross-agent cognitive alignment is achieved, from abstract strategies to concrete behaviors.
[0187] Convert the hierarchical decision information output in step 3 into a transmittable semantic communication packet ,in Including aging grade , Three-layer encoding is used:
[0188] Strategic intention:
[0189]
[0190] in Indicates the attack tendency coefficient (such as attack intensity), represents risk tolerance, are the coordinates of the bounding box of the hotspot area (such as the weak area of the opponent's defense line), is the encoding vector of strategic layer intention.
[0191] Tactical layer semantics: define the tactical action ontology library:
[0192]
[0193] Using graph attention encoding: ,in The collaboration topology generated in step 3, is the communication adjacency matrix.
[0194] Execution layer parameters:
[0195]
[0196] Adopting Variational Autoencoder (VAE) to achieve low-dimensional bandwidth transmission of high-level tactical semantics:
[0197]
[0198] Represents the input tactical semantic information, are the parameters of the encoder, represents the low-dimensional latent variables, are the mean and variance of the Gaussian distribution, output by the encoder network.
[0199] Dynamically optimize communication resource allocation according to battlefield situation and construct channel utility function:
[0200]
[0201] in For information priority, For intelligent agents The communication coverage area, is the broadband occupancy rate, Expressed as a tactical value function For intelligent agents Tactical Semantics The gradient, Represented as an agent The communication coverage area, Represented as an agent and The distance between Expressed as bandwidth cost weight.
[0202] In order to solve the problem of coordination failure caused by inconsistent intentions of multiple agents, a distributed consensus protocol is constructed:
[0203] Define the tactical intention of agent i The confidence update rule is:
[0204]
[0205] in, Represented as an agent exist Always Tactical The confidence level of It is the tactical graph similarity calculation function.
[0206] When conflicting intentions are detected (e.g. multiple agents competing for the same slot at the same time), find the Pareto optimal strategy:
[0207]
[0208]
[0209] in, For intelligent agents The profit function, For example, in the coordination of dynamic offside traps, the full-backs and center-backs synchronize their forward pressing timing through intention broadcasts, ensuring the temporal and spatial consistency of the offside trap tactics.
[0210] To optimize cross-layer communication, we adopt strategic-tactical cascade feedback, design an intention achievement evaluation function, and reversely adjust high-level strategies:
[0211]
[0212] in To actually attack the hot zone, Tactical tolerance is used to dynamically adjust the attack tendency coefficient in the high-level strategy according to the deviation between the actual attack hot zone and the target hot zone. .
[0213] Then, communication-control joint optimization is performed to build a motion control model under communication delay constraints:
[0214]
[0215] in is the network-induced delay, is the communication scheduling strategy. The communication-control joint optimization problem aims to solve the problem of and communication scheduling strategies Under the constraint of , making the system state as close to the desired trajectory as possible.
[0216] In this embodiment, “intention-level communication” means that the communication information is no longer a static state quantity, but a high-level action intention tensor formed by the agent based on its own perception, role decision-making and strategy planning. After the receiving agent obtains the intention tensor of the teammate, it does not passively receive it, but completes the following three steps through the distributed intention fusion: (1) Intention Figure 1Consistency check: detect whether there is a conflict or contradiction between the received intention and the self-perception; (2) Synergy matching calculation: calculate the synergy efficiency score between the self-role task and the intention of teammates; (3) Intention adjustment and reconstruction: if a high-risk conflict is found (for example, multiple agents intend to rush to the same support area at the same time), the candidate role tendency will be automatically corrected to form a global-local consistency scheduling.
[0217] Building on the previous “hierarchical decision-making” approach, this intent-level communication design dynamically maps high-level strategic planning and mid-level collaboration into a real-time interactive intent stream between agents, breaking the rigidity of traditional directive-based collaboration and endowing the multi-agent system with powerful flexibility, adaptability, and resilience.
[0218] Step S50: Introduce an adversarial meta-learning mechanism to quickly adapt to changes in the opponent's strategy through strategy prototype retrieval and online fine-tuning.
[0219] In the specific implementation, the steps of introducing the adversarial meta-learning mechanism and quickly adapting to the opponent's strategy changes through strategy prototype retrieval and online fine-tuning include: introducing the adversarial meta-learning framework in the adversarial meta-learning mechanism to retrieve historical strategy meta-features through the strategy prototype library, and generating adversarial perturbation samples in combination with the opponent's behavior pattern matrix; using the two-stream update rule to quickly fine-tune the strategy network and value network, and realizing the continuous evolution of the meta-knowledge base through strategy distillation loss.
[0220] It's important to note that to account for the uncertainty and mutability of adversarial strategies in adversarial scenarios, the system incorporates an adversarial meta-learning mechanism. By designing adversarial perturbations and changing scenarios during training, the model learns to rapidly adapt to varying adversary strategies. In practice, upon detecting a change in the adversary's strategy (via entropy monitoring or game feedback), the agent immediately invokes meta-parameters via the meta-learning optimizer and rapidly adjusts the policy generation network, completing policy migration in a fraction of the time without requiring retraining. This mechanism significantly improves the system's responsiveness and adaptability in unknown scenarios and complex games, enabling autonomous, intelligent behavior capable of "on-the-spot adjustment and self-evolution."
[0221] In the specific implementation, to address the problem of strategy lag caused by sudden changes in opponent strategies and dynamic evolution of the environment in football confrontations, this system proposes a dual-stream adversarial meta-learning (DSAML) framework, which realizes strategy adaptive evolution through a collaborative mechanism of offline meta-knowledge base construction and online strategy rapid fine-tuning.
[0222] Extract cross-scenario strategy invariance features from historical adversarial data and build a strategy prototype library ,in is the strategy parameter, is the element-feature vector, is the applicable scene descriptor.
[0223] Design adversarial course learning to adapt to different opponent strategies and design strategy difficulty evaluation function:
[0224]
[0225] in, It's a strategy The difficulty rating of the confrontation, is the opponent's strategy in state s The value function estimate of is the strategy of the team in state s The value function estimate of It's a strategy The entropy at state s represents the uncertainty of the policy and dynamically generates a sequence of training courses. , satisfy when , design a strategy decoupling encoder and use contrastive learning to extract strategy meta-features:
[0226]
[0227] in To map the opponent’s situation, the CLIP module aligns the policy gradient with the scene features. is a projection function that maps features to a uniform space, For the opponent's situation map, It is the meta-feature representation of the strategy.
[0228] Design a dual-stream adaptation mechanism to implement the perception of opponent strategy offset , through strategy prototype retrieval and meta-gradient update, rapid adjustment is achieved, and the opponent behavior pattern matrix is constructed based on the communication data in step 4:
[0229]
[0230] in is the opponent behavior pattern matrix, The opponent is at all times The behavioral feature vector of is the time decay factor, Represents a tensor product.
[0231] Then perform meta-strategy matching retrieval and define the strategy similarity metric:
[0232]
[0233] in, is the strategy similarity score, Is a candidate strategy The meta-features of is the weight matrix of the meta-feature space, is the weight coefficient of KL divergence, real-time retrieval of Top-K candidate strategies , then conduct adversarial fast fine-tuning and build a dual-stream update rule (strategy stream + value stream):
[0234]
[0235]
[0236] in Expressed as policy network parameters, is the inner learning rate, is the cross entropy loss between the current strategy and the optimal strategy, is the value network parameter, is the current value function, Estimate the adversary's value network.
[0237] The adapted policy knowledge is deposited into the meta-knowledge base to achieve continuous evolution, using the policy distillation loss function:
[0238]
[0239] Among them, JS is Jensen-Shannon divergence, is the hidden layer feature of the policy network, is the weight coefficient of feature matching loss. Then the adversarial strategy is enhanced to train the adversarial generator Sample synthetic criticality strategy:
[0240]
[0241] Where G is the generator, D is the discriminator, z is the random noise vector, Based on the original characteristics of the strategy Policy samples for conditional generation, generating data for enhanced meta-training.
[0242] At the same time, the cross-module collaborative interface is modified and the P in step 4 is encoded into a strategy condition vector:
[0243]
[0244] Use Transformer encoder to transform strategic intent and tactical semantics Encoded as a policy condition vector , used to guide the conditional generation process of meta-policies.
[0245] Define the consistency loss between policy layers:
[0246]
[0247] in, For macro strategy (high-level strategy), For micro-strategy (underlying strategy), It is the gradient of the policy with respect to the meta-features, ensuring that the high-level strategy is aligned with the meta-features of the low-level actions.
[0248] Step S60: Execute the underlying control instructions and provide real-time feedback on the execution status to form a closed-loop optimization.
[0249] In specific implementation, the steps of executing underlying control instructions and providing real-time feedback on the execution status to form a closed-loop optimization include: real-time collection of action execution deviation, collision risk and tactical achievement data, and compensation of execution errors through an inverse kinematics model; triggering dynamic role reconstruction and strategy weight adjustment based on feedback, and optimizing control accuracy under communication delay through the Kalman gain matrix.
[0250] In practice, the final strategy is executed in real time by the underlying physical control layer based on the strategic intent and action instructions generated by the high- and mid-level layers. The agent completes specific actions in the virtual environment, such as positioning, passing, shooting, or defending, while simultaneously transmitting execution status and scenario feedback in real time to the multimodal perception system and high-level strategy network. This feedback mechanism includes not only basic movement feedback but also information on confrontation results, tactical execution effectiveness, and energy consumption, providing a closed-loop adjustment basis for the dynamic role migration network and meta-learning optimizer. The system supports multi-cycle incremental updates, ensuring continuous adaptive optimization of team strategies and individual control strategies, enabling continuous evolution and multi-scenario transferable applications.
[0251] It is understandable that in order to address the problems of accumulated strategic execution deviations, unpredictable environmental disturbances, and the emergence of opponent counter-strategies in football confrontations, this system is designed with a "three-ring nested feedback" architecture, which uses a collaborative mechanism of real-time control at the execution layer, online learning at the tactical layer, and offline optimization at the strategic layer.
[0252] Multi-granularity feedback channel design: Build a multi-level feedback loop from micro-actions to macro-strategies to achieve end-to-end optimization from execution effects to strategy parameters.
[0253] Real-time tactical feedback loop:
[0254]
[0255] in, is the tactical feedback vector, is the number of successful oppressions, is the total number of attempts, is the actual speed, is the planned speed when a collision risk is detected When the emergency retracement strategy is triggered:
[0256]
[0257] in, For emergency control instructions, is the repulsive potential field on the position The gradient of , the policy residual compensation, and the execution deviation are calculated by the inverse kinematics model:
[0258]
[0259] in, is the action deviation, is the inverse kinematics model, is the actual state of the current and previous moments, is the control instruction, which is based on the TD error to update the value network meta-strategy for fast calibration:
[0260]
[0261]
[0262] in, is the TD error, Indicates immediate reward, is the discount factor, represents the value function, Is the learning rate. Used to quickly calibrate the value network , to adapt to new tactics or adversary behavior.
[0263] Finally, a closed-loop coordination mechanism was designed, execution-communication joint optimization was performed, and an experimental compensation observer was designed:
[0264]
[0265] in For communication delay, is the Kalman gain matrix, is the delayed observation value, is the observation matrix. It is used in the presence of communication delay In the case of .
[0266] Feedback-driven role refactoring when persistent execution deviations are detected , the dynamic role reconstruction of step 2 is triggered:
[0267]
[0268] in, is the trigger signal, For intelligent agents At the moment The rate of action or state change.
[0269] Step S70: Dynamically adjust role allocation and strategy weights based on feedback data to achieve continuous evolution of the multi-agent collaboration system.
[0270] It should be noted that this embodiment obtains comprehensive environmental information through multimodal perception fusion, providing a basis for dynamic role reconstruction and hierarchical decision-making. The dynamic role reconstruction mechanism adjusts the role of the intelligent agent according to the real-time situation, thereby improving the flexibility of the team. The hierarchical decision-making architecture decomposes complex decision-making problems into different levels, realizing collaborative optimization from global strategy to local execution. The intention-level communication mechanism realizes efficient information exchange and policy coordination between intelligent agents. The dual-stream adversarial meta-learning framework enhances the system's ability to cope with new opponents and environmental changes through rapid strategy adaptation and fine-tuning. The three-loop nested feedback architecture realizes multi-level optimization from micro-actions to macro-strategies, continuously improving system performance. These technical features work together to solve the problem of multi-agent collaborative control in complex dynamic confrontation scenarios. Multimodal perception and dynamic role reconstruction improve the system's environmental adaptability, hierarchical decision-making and intent-level communication enhance the efficiency of collaboration between intelligent agents, and adversarial meta-learning and nested feedback mechanisms ensure the system's continuous optimization and evolution capabilities.
[0271] In the specific implementation, the working process of this embodiment is described by taking the RoboCup 3D simulated football game as an example:
[0272] Multimodal perception fusion: The system collects and fuses multimodal data such as vision (ball, teammates, opponent positions), hearing (referee whistle, teammates' shouts), space (own position, orientation), motion (joint angles, gait status) and communication (teammates' intentions) in real time to form comprehensive environmental cognition.
[0273] Dynamic role reconfiguration: When a ball possession transition is detected, the system immediately reevaluates the role of each agent. For example, a forward who was originally in the frontcourt may be assigned to a defensive role and quickly return to defense.
[0274] Hierarchical decision-making: The high-level strategy layer generates tactical instructions for "quick counterattacks" based on the overall situation. The middle-level collaboration layer translates these instructions into specific running and passing routes. The bottom-level execution layer precisely controls the movement trajectory and action execution of each agent.
[0275] Intent-level communication: The forward broadcasts "I will sprint to receive the long pass" through intent communication, and the midfielder chooses the appropriate time and route to pass the ball accordingly.
[0276] Adversarial meta-learning: When the system detects that the opponent is using a high-pressure strategy, it quickly retrieves and fine-tunes the response strategy from the meta-knowledge base, such as increasing the frequency of short passes.
[0277] Feedback Optimization: The execution layer adjusts action parameters, such as pass strength, in real time. The tactical layer optimizes passing decisions through online learning based on pass success rates. The strategic layer analyzes match data offline to optimize the overall tactical system.
[0278] In practice, to address the issue of insufficient adaptability during agent role switching, this embodiment proposes an optimization solution based on gradual role transition. This solution introduces a role transition buffer mechanism, allowing the agent to experience a brief mixed state during the role switching process, gradually adapting to the behavior pattern of the new role.
[0279] Role Fusion Matrix: A dynamic role fusion matrix is designed for each agent to represent the weight distribution of the current role and the target role. For example, an agent transitioning from a forward to a defender might start at [0.8, 0.2] (80% forward, 20% defender) and then gradually transition to [0, 1].
[0280] Gradual Behavior Adjustment: Based on the role fusion matrix, the agent's decisions and behaviors will be a weighted combination of the two roles. For example, in the early stages of the transition, the agent may still retain some offensive awareness while also beginning to perform some basic defensive tasks.
[0281] Adaptive Learning Module: A fast adaptive learning module is introduced to enable the agent to quickly learn and adapt to the key skills and decision-making patterns required for the new role during role transitions. This module uses meta-learning techniques to extract general adaptation strategies from past role transition experiences.
[0282] Context-Aware Transition Controller: Design a context-aware transition controller that dynamically adjusts the speed and method of character transitions based on the current game situation (e.g., ball position, teammate distribution, opponent pressure, etc.). This allows for faster transitions in critical situations, while allowing for smoother transitions in less demanding situations.
[0283] Collaborative transition mechanism: When an agent begins a role transition, nearby teammate agents are notified and adjust their behavior accordingly to coordinate with the transitioning teammate. This collaborative mechanism ensures the coherence and efficiency of the entire team during the role transition process.
[0284] This gradual role transition scheme allows agents to gradually adapt to the requirements of their new roles while maintaining a certain degree of continuity. This not only reduces decision-making delays and behavioral inconsistencies associated with role switching, but also improves the team's overall adaptability and performance in dynamic competitive environments. By smoothing the role transition process, this scheme effectively addresses the technical issue of insufficient adaptability in agent role switching in specific application scenarios, further enhancing the competitive advantage of the multimodal hierarchical reinforcement multi-agent collaborative control system in the RoboCup 3D simulated football competition.
[0285] This embodiment uses multimodal sensors to collect and fuse visual, auditory, spatial, motion, and communication modal data in real time to generate a high-level representation of the environment state. Based on this high-level representation, a dynamic role transition network is used to assign dynamic roles to each agent. A multi-level policy system is constructed, decomposing decision-making into a high-level policy layer, a mid-level collaboration layer, and a low-level control layer, generating tactical intent, coordination relationships, and physical control instructions respectively. An event-triggered mechanism is used to achieve low-latency policy synchronization under critical events. An adversarial meta-learning mechanism is introduced to retrieve policy prototypes and fine-tune them online. Low-level control instructions are executed and the execution status is fed back in real time, forming a closed-loop optimization loop. Role assignments and policy weights are dynamically adjusted based on feedback data, enabling the continuous evolution of the multi-agent collaborative system. This system implements hierarchical strategic decision-making, while also possessing event-driven compressed intent communication capabilities. The integration of an adversarial meta-learning evolver enables rapid adaptation to new adversaries and continuous policy self-evolution.
[0286] In addition, an embodiment of the present application also proposes a computer-readable storage medium, on which a program for modal layered enhancement of multi-agent collaborative control is stored. When the program for modal layered enhancement of multi-agent collaborative control is executed by a processor, the steps of the method for modal layered enhancement of multi-agent collaborative control as described above are implemented.
[0287] Reference Figure 3 , Figure 3 This is a structural block diagram of the first embodiment of the modal layered enhanced multi-agent collaborative control device of this application.
[0288] like Figure 3 As shown, the modal layered enhanced multi-agent collaborative control device proposed in the embodiment of the present application includes:
[0289] The information acquisition module 10 is used to collect and fuse visual, auditory, spatial, motion and communication modal data in real time through multimodal sensors to generate a high-level environmental state representation;
[0290] A role assignment module 20 is configured to assign dynamic roles to each agent based on the high-level environmental state representation, combined with historical data and real-time situation, through a dynamic role migration network to achieve adaptive role switching and conflict coordination;
[0291] System construction module 30 is used to build a multi-level strategy system, decomposing the decision into a high-level strategy layer, a middle-level collaboration layer, and a bottom-level control layer, generating tactical intent, coordination relationships, and physical control instructions respectively;
[0292] The transmission module 40 is used to compress and transmit intent-level communication information based on the event trigger mechanism to achieve low-latency policy synchronization under key events;
[0293] Fine-tuning module 50, which is used to introduce adversarial meta-learning mechanism and quickly adapt to the opponent's strategy changes through strategy prototype retrieval and online fine-tuning;
[0294] Optimization module 60, used to execute the underlying control instructions and provide real-time feedback on the execution status, forming a closed-loop optimization;
[0295] The control module 70 is used to dynamically adjust role allocation and strategy weights according to feedback data to achieve continuous evolution of the multi-agent collaboration system.
[0296] It should be understood that the above is only an example and does not constitute any limitation to the technical solution of the present application. In specific applications, technicians in this field can make settings as needed, and the present application does not impose any restrictions on this.
[0297] This embodiment uses multimodal sensors to collect and fuse visual, auditory, spatial, motion, and communication modal data in real time to generate a high-level representation of the environment state. Based on this high-level representation, a dynamic role transition network is used to assign dynamic roles to each agent. A multi-level policy system is constructed, decomposing decision-making into a high-level policy layer, a mid-level collaboration layer, and a low-level control layer, generating tactical intent, coordination relationships, and physical control instructions respectively. An event-triggered mechanism is used to achieve low-latency policy synchronization under critical events. An adversarial meta-learning mechanism is introduced to retrieve policy prototypes and fine-tune them online. Low-level control instructions are executed and the execution status is fed back in real time, forming a closed-loop optimization loop. Role assignments and policy weights are dynamically adjusted based on feedback data, enabling the continuous evolution of the multi-agent collaborative system. This system implements hierarchical strategic decision-making, while also possessing event-driven compressed intent communication capabilities. The integration of an adversarial meta-learning evolver enables rapid adaptation to new adversaries and continuous policy self-evolution.
[0298] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this application. In actual applications, technicians in this field can select part or all of it according to actual needs to achieve the purpose of this embodiment scheme, and no restrictions are imposed here.
[0299] In addition, for technical details not fully described in this embodiment, please refer to the method of modal layered enhanced multi-agent collaborative control provided in any embodiment of this application, which will not be repeated here.
[0300] In addition, it should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0301] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0302] Through the above description of the embodiments, those skilled in the art will clearly understand that the above-mentioned embodiments and methods can be implemented using software plus the necessary general-purpose hardware platform. Of course, hardware can also be used, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / RAM, a magnetic disk, or an optical disk) and includes several instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of this application. The above are only preferred embodiments of this application and do not limit the scope of the patent application. Any equivalent structure or equivalent process transformation made using the contents of this application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the scope of patent protection of this application.
Claims
1. A modal hierarchical enhanced multi-agent collaborative control method, characterized in that: include: Through multimodal sensors, real-time collection and fusion of visual, auditory, spatial, motion and communication modal data can generate high-level environmental state representation; Based on the high-level environmental state representation, combined with historical data and real-time situation, dynamic roles are assigned to each agent through a dynamic role migration network to achieve adaptive role switching and conflict coordination; Build a multi-level strategy system, decomposing decisions into a high-level strategy layer, a middle-level collaboration layer, and a low-level control layer, generating tactical intent, collaborative relationships, and physical control instructions respectively; Based on the event trigger mechanism, intent-level communication information is compressed and transmitted to achieve low-latency policy synchronization under key events; Introducing an adversarial meta-learning mechanism to quickly adapt to changes in opponent strategies through strategy prototype retrieval and online fine-tuning; Execute underlying control instructions and provide real-time feedback on execution status to form a closed-loop optimization; Dynamically adjust role allocation and strategy weights based on feedback data to achieve continuous evolution of the multi-agent collaborative system.
2. The method according to claim 1, characterized in that The step of collecting and fusing visual, auditory, spatial, motion, and communication modal data in real time through multimodal sensors to generate a high-level environmental state representation includes: Real-time collection and fusion of visual, auditory, spatial, motion and communication modality data through multimodal sensors; Through visual modality analysis, panoramic and local visual information are converted into global coordinates and a field potential thermal map is generated. Receive sound signals through the auditory modality and use Bayesian estimation and particle filtering to locate the sound source; By integrating inertial navigation data and scene analysis results through spatial modality, a spatial coverage probability field is constructed; Correct movement abnormalities by modeling inverse kinematics and joint timing relationships through motion modalities; Construct the communication graph topology and generate node representations through communication modalities; Multimodal time series data are aligned through a dynamic differentiable interpolation layer, and the contrastive loss is used to constrain the similarity between modalities to generate a unified high-order state representation.
3. The method according to claim 1, characterized in that The dynamic role migration network includes: The dynamic role migration network uses a dual-attention mechanism to calculate the role adaptation score based on the global strategic demand map and local perception information; Periodically review role assignments through an event-triggered mechanism, and trigger role reconstruction when policy conflicts are detected. Output the agent's probabilities of tendencies towards multiple roles, combine them with conflict adjustment factors to avoid resource competition, and achieve decentralized collaborative allocation.
4. The method according to claim 1, wherein The high-level strategy layer solves the global tactical field weight distribution through Nash equilibrium to generate attack and defense tendency coefficients and risk tolerance instructions; The middle collaboration layer generates collaborative relationship topology based on the graph attention mechanism and eliminates intention overlap through the conflict coordinator; The underlying control layer uses a prediction error compensation model and hierarchical model predictive control to decode the collaborative intention into continuous action instructions under physical constraints.
5. The method according to claim 1, characterized in that The intention-level communication information includes the strategic layer hotspot area boundary, the tactical layer action ontology library and the execution layer parameter coding; Compress semantic information through variational autoencoders and design channel utility functions to optimize communication resource allocation; The receiving end verifies consistency, calculates coordination and dynamically reconstructs intent through a distributed intent fuser to ensure tactical spatiotemporal consistency.
6. The method according to claim 1, wherein The steps of introducing the adversarial meta-learning mechanism and rapidly adapting to changes in the opponent's strategy through strategy prototype retrieval and online fine-tuning include: The adversarial meta-learning framework in the adversarial meta-learning mechanism is introduced to retrieve historical strategy meta-features through the strategy prototype library and generate adversarial perturbation samples in combination with the opponent's behavior pattern matrix; A two-stream update rule is used to quickly fine-tune the policy network and value network, and the policy distillation loss is used to achieve continuous evolution of the meta-knowledge base.
7. The method according to claim 1, characterized in that The steps of executing the underlying control instructions and providing real-time feedback on the execution status to form a closed-loop optimization include: Collect action execution deviation, collision risk and tactical achievement data in real time, and compensate for execution errors through inverse kinematics models; Dynamic role reconstruction and strategy weight adjustment are triggered based on feedback, while the control accuracy under communication delay is optimized through the Kalman gain matrix.
8. A modal layered enhanced multi-agent collaborative control device, characterized in that: include: The information acquisition module is used to collect and fuse visual, auditory, spatial, motion and communication modal data in real time through multimodal sensors to generate a high-level environmental state representation; A role assignment module is used to assign dynamic roles to each agent through a dynamic role migration network based on the high-level environmental state representation, combined with historical data and real-time situation, to achieve adaptive role switching and conflict coordination; The system building module is used to build a multi-level strategy system, decomposing decisions into a high-level strategy layer, a middle-level collaboration layer, and a low-level control layer, generating tactical intent, collaborative relationships, and physical control instructions respectively; The transmission module is used to compress and transmit intent-level communication information based on the event trigger mechanism to achieve low-latency policy synchronization under key events; The fine-tuning module is used to introduce an adversarial meta-learning mechanism to quickly adapt to changes in the opponent's strategy through policy prototype retrieval and online fine-tuning; The optimization module is used to execute the underlying control instructions and provide real-time feedback on the execution status, forming a closed-loop optimization; The control module is used to dynamically adjust role allocation and strategy weights based on feedback data to achieve continuous evolution of the multi-agent collaborative system.
9. A computer device, characterized in that: The device comprises: a memory and a processor, wherein the processor executes the method according to any one of claims 1 to 7 when running computer instructions stored in the memory.
10. A computer-readable storage medium, characterized in that The method comprises instructions which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.