Heterogeneous multi-agent cooperative control test system and method and electronic equipment

By constructing a unified heterogeneous multi-agent cooperative control test system, using the MuJoCo engine and ROS2 framework, the collaborative operation of multiple types of agents in the same simulation environment was realized. This solved the problems of platform fragmentation and communication incompatibility, improved the real-time performance and stability of the system, simplified the control test process, and enhanced the collaborative efficiency of the multi-agent system.

CN120848465APending Publication Date: 2025-10-28SHENZHEN GUOCHUANG EMBODIED INTELLIGENT ROBOT CO LTD +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511095050.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In existing technologies, heterogeneous multi-agent systems suffer from platform fragmentation, incompatible communication protocols, and difficulty in data sharing during collaborative simulation testing, resulting in low development efficiency, poor real-time performance, and poor stability.

Method used

This paper presents a heterogeneous multi-agent cooperative control test system. It uses the MuJoCo engine to build a unified simulation platform, integrates multiple types of agent models, defines a unified communication protocol and data format through the ROS2 framework, performs modular modeling, integrates multiple control algorithms, and provides a visualization test module and a task control module to realize the cooperative operation and data interaction of different types of agents.

Benefits of technology

It improves the integration, adaptability and efficiency of simulation testing, solves the platform fragmentation problem, realizes efficient data sharing and state synchronization, improves the real-time performance and stability of the system, simplifies the control testing process, and enhances the collaborative efficiency of multi-agent systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120848465A_ABST
    Figure CN120848465A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous multi-agent cooperative control test system, and the system comprises a physical simulation engine module which is used for constructing a simulation platform through employing a MuJoCo engine, and the simulation platform integrates various types of agent models; the communication interaction module is used for establishing a preset standard interface with a uniform communication protocol and a data format among the intelligent agent models of various types; the modular modeling module is used for constructing various types of agent models in a modular mode; the visual test module is used for providing a graphical interface to drag and deploy various control algorithms, presenting state information and task execution progress of different types of agent models, and receiving a test debugging instruction triggered by a user; the control algorithm module is used for integrating a plurality of control algorithms and deploying the control algorithms into different types of agent models through a preset standard interface; and the task control module is used for controlling the plurality of different types of agent models to perform behavior scheduling, path planning and action distribution according to the received cooperative task request.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics, and more particularly to heterogeneous multi-agent cooperative control testing systems, methods, and electronic devices. Background Technology

[0002] With the development of robotics technology, heterogeneous multi-agent systems have been widely used in fields such as search and rescue, and environmental monitoring. These systems consist of different types of intelligent agents (such as humanoid robots, robotic dogs, and drones) that collaborate to complete complex tasks, and simulation platforms have become an important tool for verifying their control strategies.

[0003] In existing technologies, physics engines such as MuJoCo (Multi-Joint dynamics with Contact) are commonly used in conjunction with the Robot Operating System 2 (ROS2) framework for multi-agent simulation and control. This involves loading models, building scenarios, and deploying algorithms to achieve task simulation and testing. Some platforms support environment setup via graphical interfaces or scripts and integrate basic control algorithms for verification.

[0004] However, existing simulation testing technologies supporting collaborative work among multiple heterogeneous intelligent agents still face the following bottlenecks: First, mainstream simulation platforms often focus on a specific type of robot, requiring users to switch between multiple platforms, resulting in low development efficiency. Second, due to significant differences in sensors, communication protocols, and control strategies among various intelligent agents, collaboration between multiple agents is difficult, and data sharing between different agents is not easy, affecting the overall real-time performance and stability of the system. Summary of the Invention

[0005] Based on the above problems, this application provides a heterogeneous multi-agent cooperative control test system, method and electronic device, with the aim of providing a unified and scalable simulation platform to support the collaborative operation of multiple types of heterogeneous agents in the same physical simulation environment.

[0006] In a first aspect, embodiments of this application provide a heterogeneous multi-agent cooperative control test system, including:

[0007] The physics simulation engine module is used to build a simulation platform using the MuJoCo engine; wherein, the simulation platform integrates multiple types of intelligent agent models, and different types of intelligent agent models interact within the same simulation platform;

[0008] The communication and interaction module is used to establish a unified communication protocol and a preset standard interface for data format between various types of intelligent agent models, so as to realize data interaction between different types of intelligent agent models.

[0009] The modular modeling module is used to construct various types of intelligent agent models in a modular manner. The modules of the intelligent agent model include structural modules, motor modules, and sensor modules. The modules interact with each other on the simulation platform through a preset standard interface.

[0010] The visualization testing module provides a graphical interface for dragging and dropping to deploy various control algorithms, presenting the status information of different types of intelligent agent models, presenting the task execution progress of different types of intelligent agent models, and receiving test and debugging commands triggered by users.

[0011] The control algorithm module is used to integrate multiple control algorithms, including Proportional-Integral-Derivative (PID) algorithm, Model Predictive Control (MPC) algorithm, and Linear Quadratic Regulator (LQR) algorithm, and deploys them to different types of intelligent agent models through a preset standard interface;

[0012] The task control module is used to control multiple intelligent agent models of different types to perform behavior scheduling, path planning and action allocation based on the received collaborative task requests.

[0013] In one embodiment, the modular modeling module includes:

[0014] The sensor module is used for configuring the physical properties of the sensors and simulating data; the sensors include a camera, a lidar, and an inertial measurement unit (IMU).

[0015] The actuator module provides a drive model and feedback control PID interface for the actuator, which includes a DC motor, a servo motor, and a servo motor.

[0016] The structural module is used to define the structural components of the intelligent agent model and their physical properties; the physical properties include mass, friction coefficient, elastic coefficient, and stiffness parameter.

[0017] In one embodiment, the communication interaction module adopts a binary message format defined by the ROS2 framework, the message format including a timestamp, an agent state array, and a control command array;

[0018] The communication interaction module is also used to perform data compression processing on the output data of the sensor according to the sensor type and the current collaborative task; the data compression processing includes cropping high-frequency sensor data, or merging or interval sampling low-frequency sensor data.

[0019] In one embodiment, the visualization testing module includes:

[0020] A graphical programming workspace for dragging and dropping sensor modules, control algorithm modules, and actuator modules;

[0021] The real-time rendering unit is used to display the position, task progress, and environmental obstacles of different types of intelligent agent models based on Unity3D.

[0022] The parameter adjustment panel is used to adjust PID control parameters or reinforcement learning strategy hyperparameters.

[0023] In one embodiment, the visualization testing module further includes a scene construction module for generating virtual environments required for simulation testing of multiple types of intelligent agent models. The scene construction module includes the following various scene modeling methods:

[0024] A graphical interface is constructed, allowing users to drag and drop preset environmental components into the simulated virtual environment and set the type, parameters, and spatial position of each component. The environmental components include terrain, obstacles, and targets.

[0025] Scripted construction allows you to define the type, parameters, and spatial location of each environmental component in the scene by writing configuration scripts.

[0026] Natural language construction utilizes large language models to perform semantic parsing of the natural language input by the user and generates the type, parameters, and spatial location of each environmental component;

[0027] Reinforcement learning-assisted construction automatically generates the type, parameters, and spatial location of each environmental component based on the collaborative task of the selected agent model.

[0028] In one embodiment, the control algorithm module includes:

[0029] Basic control layer: PID controllers used to provide feedback control for various types of intelligent agent models, so as to control the joint position or speed;

[0030] Collaborative Control Layer: Integrates Model Predictive Control (MPC) algorithms for path trajectory planning in various agent models;

[0031] Autonomous Learning Layer: Integrates reinforcement learning policy interfaces for model training and policy inference based on Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) algorithms, enabling each agent model to make decisions in different virtual environments.

[0032] In one embodiment, the system further includes a task scheduling module, which specifically includes:

[0033] The task monitoring module is used to monitor the remaining battery power, communication quality, and priority changes of the currently executing tasks of each intelligent agent in real time; the communication quality includes communication latency and packet loss rate.

[0034] The mode switching module is used to select the appropriate collaboration mode for each agent model based on the information provided by the task monitoring module; the collaboration modes include cooperative mode, competitive mode and hybrid mode, and to issue mode switching instructions to each agent.

[0035] Secondly, embodiments of this application also provide a heterogeneous multi-agent cooperative control testing method, applied to a heterogeneous multi-agent cooperative control testing system, comprising:

[0036] A simulation platform is built using the MuJoCo engine; the simulation platform integrates multiple types of intelligent agent models, and different types of intelligent agent models interact within the same simulation platform;

[0037] Establish a unified communication protocol and a pre-defined standard interface for data format among various types of intelligent agent models to enable data interaction between different types of intelligent agent models;

[0038] Various types of intelligent agent models are constructed in a modular manner. The modules of the intelligent agent model include structural modules, motor modules, and sensor modules. The modules interact with each other on the simulation platform through a preset standard interface.

[0039] The graphical interface provided by the visualization testing module is used to receive various control algorithms deployed by the user and test debugging commands triggered by the user. The graphical interface provided by the visualization testing module is also used to present the state information of different types of intelligent agent models and the task execution progress of different types of intelligent agent models.

[0040] The control algorithms provided by the control algorithm module are deployed to different types of intelligent agent models through a preset standard interface. The control algorithms include feedback control PID algorithm, model predictive control MPC algorithm, and linear quadratic regulator LQR algorithm.

[0041] Based on the received collaborative task requests, control multiple intelligent agent models of different types to perform behavior scheduling, path planning, and action allocation.

[0042] In one embodiment, data compression processing is performed on the output data of the sensor based on the sensor type and the current collaborative task; the data compression processing includes cropping high-frequency sensor data or merging or interval sampling low-frequency sensor data.

[0043] Thirdly, embodiments of this application also provide an electronic device, including:

[0044] Central processing unit, memory, input / output interfaces;

[0045] The memory is either a short-term storage memory or a persistent storage memory;

[0046] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method described in any of the second aspects above.

[0047] Fourthly, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it performs the method described in any one of the second aspects above.

[0048] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0049] This application embodiment significantly improves the integration, adaptability, and efficiency of simulation testing by constructing a unified heterogeneous multi-agent cooperative control testing system. Firstly, the system builds a high-precision, high-performance unified simulation platform based on the MuJoCo physics engine, supporting the integration of multiple heterogeneous agent models into the same simulation environment for interaction, effectively solving the problems of fragmentation and difficulty in unified operation of existing simulation platforms.

[0050] Secondly, the standardized communication protocols and data formats established through the communication interaction module enable efficient data sharing and state synchronization among different types of intelligent agent models, overcoming communication barriers caused by differences in sensors and inconsistent control architectures, and improving the overall real-time performance and stability of the system. Simultaneously, the modular modeling module employs a combination of functional modules such as structures, motors, and sensors, giving various intelligent agent models good reusability and scalability, reducing modeling complexity, and improving simulation configuration efficiency.

[0051] Furthermore, this application provides a visualization testing module, enabling algorithm deployment, task debugging, and status monitoring through a graphical interface, further simplifying the control testing process and lowering the development threshold. The control algorithm module integrates multiple mainstream control strategies and supports flexible deployment to different types of intelligent agents, enhancing the simulation platform's adaptability to multi-task scenarios. Combined with the behavior scheduling and path planning implemented by the task control module, the system can flexibly handle complex and ever-changing collaborative tasks, improving the collaborative efficiency of multi-agent systems and the practical value of simulation. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0053] Figure 1 A schematic diagram of the architecture of a heterogeneous multi-agent cooperative control test system provided in this application embodiment;

[0054] Figure 2 This is a schematic diagram of an electronic device structure provided in an embodiment of this application. Detailed Implementation

[0055] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0056] This application provides a heterogeneous multi-agent cooperative control test system, such as... Figure 1 As shown, the system includes:

[0057] The physics simulation engine module is used to build a simulation platform using the MuJoCo engine; wherein, the simulation platform integrates multiple types of intelligent agent models, and different types of intelligent agent models interact within the same simulation platform;

[0058] The communication and interaction module is used to establish a unified communication protocol and a preset standard interface for data format between various types of intelligent agent models, so as to realize data interaction between different types of intelligent agent models.

[0059] The modular modeling module is used to construct various types of intelligent agent models in a modular manner. The modules of the intelligent agent model include structural modules, motor modules, and sensor modules. The modules interact with each other on the simulation platform through a preset standard interface.

[0060] The visualization testing module provides a graphical interface for dragging and dropping to deploy various control algorithms, presenting the status information of different types of intelligent agent models, presenting the task execution progress of different types of intelligent agent models, and receiving test and debugging commands triggered by users.

[0061] The control algorithm module is used to integrate multiple control algorithms, including feedback control PID algorithm, model predictive control MPC algorithm and linear quadratic regulator LQR algorithm, and to deploy them to different types of intelligent agent models through a preset standard interface;

[0062] The task control module is used to control multiple intelligent agent models of different types to perform behavior scheduling, path planning and action allocation based on the received collaborative task requests.

[0063] The heterogeneous multi-agent cooperative control test system provided in this application includes multiple functional modules, among which the physics simulation engine module is used to build a unified simulation platform. This module uses the MuJoCo physics engine as the underlying dynamics simulation tool to achieve high-precision simulation of the physical interaction processes of various types of intelligent agent models, such as rigid body motion, joint motion, collision detection, and sensor feedback. The simulation platform has real-time computing capabilities and can simultaneously handle the motion states and interactions of multiple heterogeneous intelligent agent models in a three-dimensional virtual environment, ensuring the modeling and state response of different types of intelligent agents in a unified physical environment. Through the implementation of this module, the system has the ability to support the collaborative operation of multiple types of intelligent agents in the same simulation scenario, providing a precise and stable physical foundation for the subsequent deployment and testing of control algorithms.

[0064] The communication and interaction module is used to establish data interaction methods between various types of intelligent agent models. This module unifies the communication protocol and data structure format among intelligent agent models through a pre-defined standard interface. Based on the ROS2 framework, this module defines a binary message format containing timestamp fields, agent state arrays, control command arrays, and sensor data fields to ensure consistent timing and structural uniformity in information transmission. The communication and interaction module enables different intelligent agents to perform state synchronization, task sharing, and sensor information exchange during simulation. It can also configure data publishing frequency and transmission methods through Quality of Service (QoS) policies to ensure the real-time performance and robustness of the overall system communication. Through this module, intelligent agents with different physical attributes, control models, and sensor configurations can achieve seamless communication, improving the system's collaborative efficiency and interactive performance.

[0065] The modular modeling module is used to construct various types of intelligent agent models. Its modeling approach is modular, primarily composed of a structural module, a motor module, and a sensor module. The structural module defines the physical components of the intelligent agent, including rigid bodies, joints, and connection structures. Their physical properties, such as mass, coefficient of friction, elasticity coefficient, and stiffness parameters, can be set in the simulation platform according to actual needs. The motor module controls the movement of various joints or drive components, allowing for the design of various drive models (such as DC motors and servo motors), and integrates a PID control interface for performing position or speed control tasks. The sensor module simulates sensing devices such as cameras, lidar, and inertial measurement units, generating corresponding sensor data by setting parameters such as viewing angle, resolution, and sampling frequency. The three modules are combined and connected in the simulation platform through a unified interface standard, and exchange parameters and update states through data channels, thereby achieving high coupling and configurability between model components.

[0066] The visualization testing module provides users with an intuitive interface and control entry point, facilitating the deployment, observation, and debugging of various types of intelligent agent models. This module presents the real-time position, attitude, and state parameters of the intelligent agent model within the simulation platform through a graphical interface, displaying state changes and task progress feedback during task execution. It can receive control algorithms selected and deployed by the user via drag-and-drop, configured task parameters, set sensors, and configured actuator attributes. It can also receive user control commands and debugging operations in real time, such as task pause, parameter adjustment, and state rollback. This module works in conjunction with the control algorithm module and task control module to achieve fully visualized operation from simulation environment configuration and control strategy deployment to execution status monitoring, significantly improving the system's interactivity and efficiency.

[0067] The control algorithm module provides control strategy support for the system, integrating various mainstream robot control algorithms and deploying these strategies to different types of intelligent agent models through pre-defined standard interfaces. This module includes Proportional-Integral-Derivative (PID) control from feedback control algorithms for precise adjustment of joint angles and velocities, suitable for low-level control tasks. It also integrates Model Predictive Control (MPC) to predict the state evolution of the multi-agent system within future time windows and generate optimal control sequences for path optimization and obstacle avoidance. Furthermore, it includes a Linear Quadratic Regulator (LQR) algorithm to optimize state feedback control under known system model conditions, improving the stability and response speed of the control system. The control algorithm module can select appropriate control strategies for scheduling and binding according to different task types and model characteristics, enabling rapid deployment, parallel operation, and multi-instance application of control algorithms in the simulation environment.

[0068] The task control module handles task scheduling and behavior planning for collaborative tasks involving multiple types of intelligent agent models. This module receives task request information from users or the upper-level scheduling module and, based on the state data, control strategies, and resource distribution of each intelligent agent model in the simulation platform, performs behavior scheduling, path planning, and action allocation operations. Behavior scheduling rationally allocates sub-tasks among multiple agents; path planning generates collision-free paths based on agent model states and environmental constraints; and action allocation binds path points with control commands and transmits them to the control algorithm module. This module supports multiple scheduling strategies, including priority scheduling, dynamic allocation, and policy coordination mechanisms, ensuring the coordination and stability of task execution.

[0069] Through the integration of the above modules, the heterogeneous multi-agent cooperative control test system of this application realizes the functions of unified modeling of heterogeneous models, interconnection of standard communication interfaces, efficient deployment of control algorithms, and graphical management of the simulation environment. The system has a high degree of modularity and scalability, and can support multi-level control verification from low-level joint control to high-level cooperative task scheduling. It effectively solves the problems of fragmented simulation environment, incompatible communication protocols, and low debugging efficiency of control algorithms in the prior art, and improves the simulation integration of multi-agent systems.

[0070] In one embodiment, the modular modeling module includes: a sensor module for configuring the physical properties of sensors and simulating data; the sensors include a camera, a lidar, and an inertial measurement unit (IMU); an actuator module for providing a drive model and feedback control PID interface for the actuators, the actuators including a DC motor, a servo motor, and a servo motor; and a structure module for defining the structural components of the intelligent agent model and their physical properties; the physical properties include mass, coefficient of friction, elastic coefficient, and stiffness parameter.

[0071] In this embodiment, the sensor module, actuator module, and structural module in the modular modeling module are used to implement physical modeling and behavioral control of heterogeneous intelligent agent models in the simulation platform, so as to meet the modeling requirements of collaborative operation of multiple types of intelligent agents. Each module is connected to a unified simulation platform through standardized interface specifications, realizing data interaction and behavioral coupling between modules, supporting the simulation platform in building and multi-task testing of various types of intelligent agent models such as humanoid robots, robot dogs, and drones.

[0072] The sensor module is used for configuring the physical properties and simulating data from sensors, including those such as cameras, lidar, and inertial measurement units (IMUs). This module simulates the operating characteristics of actual sensors in a simulation platform. By setting simulation attributes such as measurement range, sampling frequency, resolution, accuracy, and noise parameters, it achieves the function of acquiring perception data in a virtual environment. For example, lidar sensors support the generation of two-dimensional or three-dimensional point cloud data based on scanning angle and ranging range, while the IMU module simulates angular velocity and acceleration data and outputs it in a unified format for the intelligent agent to use for attitude estimation or motion inference. The sensor module and the structural module are bound to specific mounting points on the virtual intelligent agent entity to ensure the consistency between the perceived information and the actual physical location, thus building high-fidelity perception capabilities for different task scenarios.

[0073] The actuator module provides the drive model and feedback control PID interface for actuators, including DC motors, servo motors, and servo motors. This module can flexibly call different types of actuator models according to the control accuracy requirements and dynamic characteristics of the simulation task, and provides a control interface for setting control input parameters such as torque, voltage, and angular velocity. The module integrates a PID control strategy, which can calculate the adjustment amount based on the error signal to achieve stable control of joints or control targets, enabling the agent to perform complex operations such as movement, steering, and attitude adjustment. The dynamic response parameters of the motor model, such as maximum output torque, speed limit, and friction loss, can be configured according to platform requirements to ensure physical consistency of the action responses of different types of agents and the simulation accuracy of real-time control.

[0074] The structure module defines the structural components and physical properties of the agent model. This module constructs the agent's motion skeleton structure using various rigid body and joint types supported by the MuJoCo engine. Rigid bodies can include solid components such as the torso, limbs, and wheels, while joint types include rotary joints, sliding joints, and ball joints, each defining the agent's degrees of freedom and structural connections. Each structure module requires configuration of physical properties such as mass, coefficient of friction, elasticity coefficient, and stiffness parameters to influence the dynamic behavior and contact response between rigid bodies. The structure module allows receiving user-defined shapes, sizes, and assembly methods for each structural component and interfaces with actuator and sensor modules through pre-defined standard interfaces to achieve complete agent structure construction and interactive behavior simulation.

[0075] Through the integration of the above modules, the modular modeling module enables users to build different types of intelligent agent models in a building-block manner, providing highly reusable and easily scalable modeling capabilities. Each module maintains good physical consistency and behavioral coupling when running in the physics simulation engine, enabling the intelligent agent model to possess motion and perception capabilities highly matched to the real environment in collaborative tasks. The modular modeling approach also provides standardized interfaces for subsequent algorithm deployment, control strategy verification, and system integration, which is beneficial for improving the modeling efficiency, scalability, and task adaptability of the simulation platform.

[0076] In summary, the embodiments of this application decouple and encapsulate the functions of sensors, actuators, and structures through modular modeling modules, significantly improving the flexibility and engineering practicality of the simulation platform in constructing heterogeneous intelligent agents. Users can quickly configure models according to collaborative task scenarios, reducing the cost of repetitive modeling. Modeled modules can be directly reused when needed, improving system simulation accuracy and development efficiency, and effectively solving the problems of non-universal models, difficulty in expansion, and cluttered control interfaces under existing modeling methods.

[0077] In one embodiment, the communication interaction module adopts a binary message format defined by the ROS2 framework. The message format includes a timestamp, an agent state array, and a control command array. The communication interaction module is also used to perform data compression processing on the output data of the sensor according to the sensor type and the current collaborative task. The data compression processing includes cropping high-frequency sensor data or merging or interval sampling low-frequency sensor data.

[0078] In this embodiment, the communication interaction module adopts the binary message format defined by the ROS2 framework to meet the real-time, efficient, and stable data transmission requirements of heterogeneous multi-agent systems in the simulation environment. The ROS2 message format has high scalability and customization capabilities, enabling consistent encapsulation and transmission of cross-platform, multi-type data through the definition of structured binary messages. In this embodiment, the message format includes a timestamp field, an agent state array field, and a control command array field. The timestamp is used to identify the time precision of data generation and supports synchronous processing; the agent state array encapsulates the current state information of multiple agents, such as position, velocity, and attitude; the control command array contains control commands issued by the system scheduling module to each agent, such as target position, velocity vector, and other control parameters.

[0079] Furthermore, considering that different agents (such as humanoid robots, robot dogs, drones, etc.) may have different sensors in multi-agent collaboration, such as LiDAR, cameras, IMU, Global Positioning System (GPS), etc., in order to further support the joint use and cross-agent sharing of multi-agent system sensor data, the communication interaction module defines a unified message encapsulation structure to encapsulate the sensor status, control commands and other auxiliary information of multiple agents in the same ROS2 message, realizing a one-time release and multi-agent readable efficient message transmission mechanism. The unified data format can adapt to these heterogeneous sensor data and ensure efficient data exchange.

[0080] For example, the communication interaction module can be designed with a general data structure to adapt to various heterogeneous sensors carried by different types of intelligent agents, including but not limited to LiDAR, IMU, and GPS. This general data structure is built based on the ROS2 Messages and Services mechanism. Each type of sensor data can be defined in a specific message format, while supporting on-demand expansion and ensuring that each message structure is concise, efficient, and easy to parse. For example, message types can be represented as follows:

[0081] sensor_msgs / LaserScan: Used to represent lidar scanning data;

[0082] sensor_msgs / Imu: Used to represent inertial measurement unit data;

[0083] sensor_msgs / NavSatFix: Used to represent GPS positioning data;

[0084] std_msgs / Float32MultiArray: Used to represent control commands or other multidimensional numerical arrays (such as speed, direction, attitude control, etc.).

[0085] Taking lidar data as an example, its message format can be defined as:

[0086] std_msgs / Header header: Used to carry timestamp and reference coordinate system information;

[0087] float32[] ranges: Represents distance measurement data in different directions;

[0088] float32[] intensities: Represents the laser reflection intensity data in the corresponding direction.

[0089] Secondly, to further improve data transmission efficiency, the communication module adopts a ROS2-based binary message format definition mechanism, using fixed-length data types such as float32 and int32 to replace traditional text formats, avoiding redundant characters and improving data compression ratio. In the custom ROS2 message types, the data structure is compacted through the sequential arrangement of structured fields, ensuring that the transmitted content maintains a high transmission rate and stability even in bandwidth-constrained environments, effectively reducing parsing latency and improving system responsiveness.

[0090] To address the data generation characteristics of different sensor types and the real-time requirements of collaborative tasks, the communication module introduces a scenario-based dynamically adjusted data compression mechanism. For high-frequency sensor data such as LiDAR, data can be cropped based on environmental changes or task importance, retaining only significantly changing measurement results to reduce the uploading of redundant information. For low-frequency sensor data such as GPS, time window merging or interval sampling can be used to reduce the data transmission frequency and alleviate the system communication burden. This mechanism can dynamically adapt to task scenarios and data loads, ensuring the timely delivery of critical sensor information and improving the utilization of communication resources.

[0091] Furthermore, in multi-agent collaborative tasks, to support the sharing and unified scheduling of multi-source data, the communication interaction module further designed a combined ROS2 message format. This format carries the state data and control commands of multiple agents within a single message, thereby achieving synchronous transmission and parsing. For example, the combined message structure can be represented as follows:

[0092] std_msgs / Header header: Records unified timestamp and coordinate reference information;

[0093] std_msgs / Float32MultiArray control_commands: Records the array of control commands corresponding to each agent;

[0094] robot_state_msgs / RobotState[] robot_states: Contains a list of current state data for multiple agents, with each element representing the complete state of an agent.

[0095] This application's embodiments, by employing a custom messaging mechanism and efficient middleware support provided by the ROS2 framework, ensure consistent data exchange between various types of sensor data and commands under the same communication protocol, avoiding format incompatibility issues and improving the overall system's data interaction efficiency. Simultaneously, a task-aware compression mechanism optimizes data transmission content and frequency, effectively improving communication bandwidth utilization and reducing system latency. In summary, through the above technical solutions, the communication interaction module significantly improves the uniformity and real-time performance of information transmission among heterogeneous multi-agent systems in the simulation platform, effectively alleviating data incompatibility and communication redundancy problems caused by sensor heterogeneity, and enhancing the system's adaptability and collaborative efficiency in complex task scenarios.

[0096] Furthermore, in another feasible implementation, to achieve efficient processing of sensor data and control commands in a multi-agent system, the communication interaction module further integrates a parsing algorithm for heterogeneous sensor data. The data parsing algorithm is designed based on the aforementioned unified data format, combined with mechanisms such as high-frequency sensor data processing, multi-threaded scheduling, and multi-source information fusion to ensure the system's data processing capabilities and real-time response performance in dynamic environments.

[0097] First, to improve data decoding and parsing efficiency, the communication module employs memory mapping technology to read received binary sensor data. By directly mapping data to memory space, the overhead of data copying in traditional input / output (I / O) operations is significantly reduced, improving access speed. Simultaneously, to address rate differences and data bursts in the sensor data stream, the system incorporates a high-performance buffer structure at the data receiving end, such as a circular buffer or a concurrent safe queue, to temporarily store and distribute incoming data. This ensures data integrity and timing consistency during the parsing phase, adapts to the parsing processing rates of different modules, and prevents data loss or backlog.

[0098] Secondly, in scenarios supporting concurrent communication among multiple agents, the communication interaction module relies on the multi-threaded scheduling mechanism of the ROS2 system, constructing an asynchronous execution environment through rclcpp::Executor. This multi-threaded scheduling mechanism allows multiple data topics to be processed in parallel, and can also prioritize tasks by setting thread priorities, ensuring that topics with high real-time requirements, such as control commands and position updates, are processed first. The ROS2 system's multi-threaded scheduling mechanism supports dynamic expansion and thread reuse, adapting to the concurrent data processing needs of various types of agents and improving the system's response speed and stability in complex task scenarios.

[0099] Furthermore, to improve the reliability and accuracy of multi-source sensor data in collaborative tasks, the communication module integrates various data fusion and filtering algorithms to reduce sensor measurement errors and redundant data interference. In practical applications, appropriate algorithms can be selected for combined processing based on the sensor combination and task requirements. For example, using a Kalman filter to fuse IMU and GPS data can effectively compensate for the noise or accuracy deficiencies of a single sensor, achieving accurate estimation of position and velocity. When facing non-Gaussian noise or highly dynamic environments, a particle filter can be used for position estimation in dynamic environments, especially when there is significant uncertainty in the environment, to improve positioning robustness and adaptability. The aforementioned fusion and filtering mechanisms can dynamically adjust parameters as needed to match the sensor sampling frequency, thereby improving the motion control accuracy and task execution stability of the agent in dynamic environments.

[0100] Through the integrated design of the above-mentioned analytical algorithms, the communication interaction module not only has the ability to process high-frequency, high-concurrency sensor data, but also can realize data decoding, scheduling and fusion processing in collaborative control tasks, providing stable, reliable and real-time data processing support for multi-agent systems, and significantly improving the system's perception accuracy, control effect and collaborative efficiency.

[0101] In one embodiment, the visualization testing module includes: a graphical programming workspace for dragging and dropping sensor modules, control algorithm modules, and actuator modules; a real-time rendering unit for displaying the position, task progress, and environmental obstacles of different types of intelligent agent models based on Unity3D; and a parameter debugging panel for adjusting PID control parameters or reinforcement learning strategy hyperparameters.

[0102] To improve the efficiency of visual control development and modular configuration capabilities of heterogeneous multi-agent systems, the system in this application provides a highly modular and scalable graphical programming interface. As part of the visual testing module, this interface supports users to drag and drop and combine functional units such as sensor modules, control algorithm modules, and actuator modules in a graphical manner, so as to realize the rapid construction, strategy deployment, and debugging and verification of the control system.

[0103] This graphical programming interface has a built-in modular component library, which includes multiple functional modules, specifically: a sensor module for simulating the sensing functions of devices such as cameras, lidar, IMUs, and force sensors, and configuring the physical properties of sensors (such as measurement range, accuracy, sampling frequency, etc.) and output format; an actuator module for describing the control models of various drive components such as motors and servos, and setting control parameters such as target voltage and angular velocity, as well as configuring feedback signals; a control algorithm module, which includes traditional control algorithms such as proportional-integral-derivative (PID) control and model predictive control (MPC), as well as reinforcement learning strategies, and can flexibly select the control method according to task requirements; a logic and flow control module for providing logic units such as conditional judgments, loop structures, and triggering mechanisms, supporting the construction of complex control flow logic; and a communication interaction module for completing data interaction, signal filtering, and format conversion between modules, ensuring that the functional units in the control system work together.

[0104] In addition to meeting general functional requirements, the system also supports the development and integration of user-defined modules. Users can define new modules using a graphical configuration interface or scripting languages ​​(such as Python and C++), including input / output port configuration, logical processing flow design, and parameter setting methods. Once defined, custom modules can be integrated with existing system modules through a unified interface to participate in the construction of graphical processes. Modules can be reused and shared across tasks. Furthermore, the system can be configured with module management functions, including module version control, permission management, and documentation maintenance, to ensure the maintainability of custom modules.

[0105] This graphical programming interface is deeply integrated with the control algorithm module, helping users select the desired control strategy module (such as PID, MPC, reinforcement learning, etc.) within the interface and configure specific algorithm parameters through the parameter panel to achieve behavioral control and task response for various types of intelligent agents. For reinforcement learning strategies, the system can visualize the training process, track key indicators, and export the model, facilitating the deployment of the trained strategy to simulation environments or actual control systems. Simultaneously, the interface provides real-time monitoring of the control logic, allowing users to dynamically view the execution path, module status, and control effects. Combined with debugging tools, users can fine-tune parameters and evaluate performance, thereby optimizing the agent's behavior.

[0106] To enhance the user experience, this graphical interface supports drag-and-drop programming. Users can directly drag required modules from the component library into the workspace and establish connections between modules via lines to build a complete control flow. Each module is equipped with an independent parameter configuration panel, allowing users to set specific values ​​and adjust them in real time according to control objectives. The system also provides process visualization, dynamically displaying the control data flow path and the running status of logic nodes, helping users understand the overall control logic. Furthermore, the interface features error prompts, automatically identifying connection errors, missing parameters, and other issues during programming. It supports breakpoint debugging and log output, improving system debugging efficiency and development reliability.

[0107] To ensure standardized and efficient communication between modules, the system also constructs an extensible interface system. Each functional module interacts with data through a unified interface, including but not limited to: a sensor data interface for transmitting real-time environmental perception information such as position, velocity, and acceleration; a motor control interface for sending control commands such as target angle and velocity output from the control algorithm module to the actuator module; and a simulation platform interface for connecting the modular modeling system with the underlying simulation engine (such as MuJoCo), supporting remote calls to various modules in the simulation environment via its C language application programming interface (API) or Python interface, enabling simulation-driven system control and status feedback.

[0108] In summary, the above-mentioned graphical programming interface design, by providing a modular component library, supporting custom development, realizing control strategy integration and debugging visualization, and ensuring the consistency and scalability of communication between modules through a unified interface system, not only significantly reduces the construction threshold of heterogeneous multi-agent control systems, but also greatly improves the efficiency and flexibility of control strategy development, verification and deployment.

[0109] The real-time rendering unit, implemented using the Unity3D graphics engine, presents the state information, task execution progress, and obstacle distribution in the virtual environment of different types of intelligent agent models in a real-time visual interface. This unit can dynamically display the position, posture, sensor range, and control feedback effects of the intelligent agent model, and synchronously display key events and behavioral trajectories during task execution, facilitating intuitive understanding and intervention by users of the system's operational status. The rendering results are synchronized with the state variables in the simulation engine, ensuring consistency between the displayed content and the system's internal logic.

[0110] The parameter tuning panel is used to configure and adjust key parameters of various control algorithm modules online. This includes setting the proportional, integral, and derivative parameters of the PID controller, as well as setting hyperparameters such as learning rate, discount factor, and exploration rate in reinforcement learning strategies. Users can adjust parameter values ​​in real time during algorithm execution, and the system automatically records and applies the changes, facilitating comparison of control effects under different parameter configurations and enabling rapid strategy optimization. Furthermore, the parameter tuning panel provides module-level performance indicator displays, enabling real-time monitoring of dimensions such as response time and error convergence curves, which helps guide control system debugging and performance analysis.

[0111] In summary, the visualization testing module enables modular construction of control flows through a graphical programming workspace, achieves dynamic visual display of simulation scenarios through a real-time rendering unit, and supports interactive configuration and optimization of control strategies through a parameter debugging panel. This effectively improves the development efficiency and task debugging flexibility of multi-agent systems and lowers the technical threshold for users in the simulation testing process.

[0112] In one embodiment, the system constructs a modular intelligent scene modeling mechanism to adapt to the simulation testing needs of heterogeneous multi-agent models in different task scenarios. The modular intelligent scene modeling platform pre-configures various types of environmental component modules, including but not limited to: terrain module, obstacle module, target object module, and environmental interaction module. The terrain module provides the basic terrain structure in the simulation environment, supporting configuration of flat land, slopes, rugged terrain, etc., and allowing setting of relevant parameters such as slope and undulation to meet the terrain adaptability testing needs of different types of agents during walking, jumping, or flying. The obstacle module includes static and dynamic obstacles; the former includes walls and fences, while the latter includes movable platforms or dynamic obstacle objects, which can be used to test the effectiveness of path planning and obstacle avoidance algorithms. The target object module defines the agent's task objectives, such as target areas, objects to be picked up, or destinations, and is used to evaluate the agent's navigation, localization, and interaction capabilities. The environmental interaction module constructs dynamic interactive elements in the simulation environment, such as buttons and switches, to simulate the task environment requiring perception and feedback control in a multi-agent system.

[0113] Based on the aforementioned various types of environment component modules, in one embodiment, the visualization testing module further includes a scene construction module for generating virtual environments required for simulation testing of multiple types of intelligent agent models. The scene construction module includes the following four scene modeling methods to improve modeling efficiency and flexibility:

[0114] The graphical user interface (GUI) construction method is based on visual interactive operation, allowing users to quickly add environmental components such as terrain, obstacles, and targets to the simulated virtual environment by dragging and dropping them. Users can also set the type, scale, behavior mode, and position and orientation information of each component in 3D space through interface operations. For example, users can drag and drop a slope module to set the slope angle, or set a dynamic obstacle module to specify its movement trajectory. This method is intuitive and easy to use, suitable for interactive development and rapid verification of scene configuration effects.

[0115] The script-based construction approach allows users to programmatically define environmental components by writing configuration scripts, including component type, physical properties (such as size, material, and coefficient of friction), initial position, motion parameters, and interaction logic with the intelligent agent. This approach is suitable for the automatic generation and batch management of large-scale or complex scenes, and can be more flexibly expanded, making it easier for users to meet highly customized scene modeling needs.

[0116] The natural language-based construction method leverages the semantic understanding capabilities of Large Language Models (LLMs), allowing users to describe their desired simulation environment requirements in natural language, such as "create a slope area and place two obstacles." The platform then automatically generates the corresponding environment component configurations through semantic parsing of this natural language input, thus automating the construction of the virtual environment. This approach significantly lowers the modeling threshold, making it particularly suitable for non-professional users or rapid iterative development scenarios.

[0117] The reinforcement learning-assisted construction method is designed for autonomous training tasks of intelligent agents. The system can automatically generate challenging and adaptive training environments based on the currently selected agent model and its collaborative task objectives. This method analyzes the agent's training state through reinforcement learning algorithms and dynamically adjusts the arrangement of components and parameter configurations in the environment to improve training efficiency and policy generalization ability. It is widely applicable to complex policy training and multi-agent interaction research.

[0118] Through the modular component design and diverse modeling methods described above, the system can flexibly and quickly construct diverse virtual environments that meet the testing needs of heterogeneous multi-agent systems, and supports efficient evaluation and optimization of agent behavior strategies under different simulation task scenarios. Furthermore, the scenario construction mechanism, combined with efficient communication formats, real-time scheduling mechanisms, and data fusion algorithms, provides the system with scalable, stable, and real-time environmental support, significantly improving the testing efficiency and development reliability of multi-agent cooperative control systems.

[0119] The control algorithm module provides multi-level, multi-strategy control capabilities for heterogeneous multi-agent models in simulation testing environments. In one embodiment, the control algorithm module is divided into three hierarchical structures based on control functions: a basic control layer, a collaborative control layer, and an autonomous learning layer. Each layer is integrated into the corresponding agent model through a platform-preset standard interface, collaboratively realizing the control process from low-level action control to high-level path planning and intelligent decision-making.

[0120] The basic control layer provides feedback control capabilities for the underlying joints or motor units of various types of intelligent agent models. This layer integrates a proportional-integral-derivative (PID) controller, which allows adjustment of joint position or velocity response by setting the proportional (P), integral (I), and derivative (D) coefficients, ensuring high stability and response accuracy of the actuators in simulation tasks. Users can configure or dynamically adjust the parameters of each controller in the platform's parameter panel to adapt to the structural characteristics and task requirements of different types of intelligent agent models, thereby achieving coordinated control of single-joint or multi-joint drive components.

[0121] The collaborative control layer is used to achieve path planning and dynamic trajectory generation among multiple types of intelligent agent models, enhancing the system's global collaborative capabilities. This layer integrates model predictive control algorithms, based on the principle of objective function optimization, to perform dynamic state modeling and control input solving within a finite prediction time window. It combines the current environmental state with predictions of future behavioral trends to achieve efficient planning and control for complex tasks such as multi-agent obstacle avoidance, path convergence, and collaborative navigation to target points. The MPC algorithm is deployed on each intelligent agent model through a platform-preset interface, allowing for flexible switching of control objectives based on task changes.

[0122] The autonomous learning layer provides autonomous policy optimization capabilities for each agent model in simulation environments with high uncertainty and where rules are difficult to model precisely. This layer integrates various reinforcement learning control policy interfaces, supporting training and policy inference using algorithms such as Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC). Users can generate optimal policy models through interactive training under specific simulation tasks based on the system's training framework, or import existing pre-trained models to perform policy transfer inference operations, thereby enabling the agent models to possess dynamic environmental adaptation and autonomous decision-making capabilities.

[0123] In one embodiment, to improve resource utilization efficiency and task completion stability in complex task scenarios, the system further includes a task scheduling module for dynamic task management and collaboration mode selection for multiple heterogeneous intelligent agent models in a simulation environment. The task scheduling module includes a task monitoring module and a mode switching module, which operate collaboratively through information interaction to ensure the real-time performance and accuracy of the scheduling strategy.

[0124] The task monitoring module is used to collect and evaluate the operational metrics of various types of intelligent agent models in real time during task execution, specifically including the agent's remaining battery power, communication quality, and task priority. Here, communication quality includes two main parameters: communication link latency and packet loss rate, used to measure the reliability and timeliness of data exchange between agents. The system periodically acquires these parameters through a built-in monitoring mechanism and communication protocol interface, and records and visualizes key data changes in real time, enabling the task control module and mode switching module to respond accurately. Simultaneously, the task monitoring module supports a multi-threaded data update mechanism, maintaining high reliability and low latency monitoring capabilities even under high load and multi-tasking conditions.

[0125] The mode switching module dynamically selects a collaboration mode that matches the agent's current resource status and task requirements based on real-time status information provided by the task monitoring module, and issues mode switching commands to the corresponding agent model. Collaboration modes include three categories: cooperative mode, competitive mode, and hybrid mode. Their adaptation rules are as follows: when agent communication quality is stable, remaining battery power is sufficient, and task priority is high, cooperative mode is prioritized to improve overall system efficiency; when communication quality deteriorates or battery power approaches a threshold, the system tends to use competitive mode to ensure priority execution of critical tasks; if some agents in the system are in a cooperative state while others need to work independently due to resource constraints, hybrid mode is used to balance overall system stability and local task completion efficiency. The mode switching module selects modes through a preset rule table or reinforcement learning strategy module and supports secondary dynamic adjustments during execution based on feedback from the task monitoring module.

[0126] This task scheduling module, through the linkage of task monitoring and mode switching, enables dynamic perception and response scheduling of resource status, communication status, and task status in heterogeneous multi-agent collaborative systems. It can effectively reduce the risk of task failure caused by uneven resources or communication obstacles, and improve the adaptability and stability of the simulation test system to complex scenarios.

[0127] Secondly, embodiments of this application also provide a heterogeneous multi-agent cooperative control testing method, applied to a heterogeneous multi-agent cooperative control testing system, comprising:

[0128] A simulation platform is built using the MuJoCo engine; wherein, the simulation platform integrates multiple types of intelligent agent models, and the different types of intelligent agent models interact within the same simulation platform;

[0129] Establish a unified communication protocol and a pre-defined standard interface for data format among various types of intelligent agent models to enable data interaction between different types of intelligent agent models;

[0130] Various types of intelligent agent models are constructed in a modular manner. The modules of the intelligent agent model include structural modules, motor modules, and sensor modules. The modules interact with each other on the simulation platform through a preset standard interface.

[0131] The graphical interface provided by the visualization testing module is used to receive various control algorithms deployed by the user and test debugging commands triggered by the user. The graphical interface provided by the visualization testing module is also used to present the state information of different types of intelligent agent models and the task execution progress of different types of intelligent agent models.

[0132] The control algorithms provided by the control algorithm module are deployed to different types of intelligent agent models through a preset standard interface. The control algorithms include feedback control PID algorithm, model predictive control MPC algorithm, and linear quadratic regulator LQR algorithm.

[0133] Based on the received collaborative task requests, control multiple intelligent agent models of different types to perform behavior scheduling, path planning, and action allocation.

[0134] In one embodiment, the method of this application further includes: performing data compression processing on the output data of the sensor according to the sensor type and the current collaborative task; the data compression processing includes cropping high-frequency sensor data, or performing merging or interval sampling on low-frequency sensor data.

[0135] The following describes an embodiment of this application with a specific task execution example: In a virtual simulation test scenario of a post-disaster search and rescue mission, the simulation platform is built based on the MuJoCo engine to evaluate the collaborative performance of a search and rescue team composed of heterogeneous intelligent agents in a complex environment. This search and rescue team includes three types of intelligent agent models: a quadruped robot (simulating a robot dog), an aircraft (simulating a drone), and a wheeled ground robot (simulating a vehicle). These three types of intelligent agent models each possess different structural modules, motor modules, and sensor modules, such as lidar, cameras, and inertial measurement units (IMUs), and interact and collaborate within the simulation platform.

[0136] In the three types of intelligent agent models mentioned above, each type of intelligent agent model is first constructed in a modular manner. The models are defined separately according to their structure, sensor, and motor functions, and data interaction between modules is achieved through a unified standard interface. Simultaneously, the system configures a unified communication protocol and binary format message interface for all heterogeneous intelligent agents. The messages contain timestamps, state arrays, and control command arrays, enabling multi-type intelligent agents to achieve task collaboration under a unified communication interface.

[0137] Users can use the system's visualization testing module in a graphical interface to drag and drop PID controllers to control the leg joints of a quadruped robot, deploy Model Predictive Control (MPC) algorithms for path tracking and dynamic obstacle avoidance on a drone, and LQR control algorithms for wheeled robots to optimize steering paths. Furthermore, users can view the task progress, position information, and control parameter adjustment effects of various agents in real time through the interface.

[0138] In response to the collaborative objectives set for this mission (such as "searching for and marking the locations of trapped personnel in earthquake rubble"), the system can adaptively compress sensor data based on the current mission attributes and the types of sensors: it performs cropping on high-frequency updated LiDAR data, retaining only the edge change data of key obstacles; it performs interval sampling on low-frequency GPS positioning data, updating only at key nodes where the agent's path changes, thereby reducing the system's communication load and ensuring the real-time collaborative performance among multiple agents.

[0139] Finally, based on the user's collaborative task request, the task scheduling module assigned search and rescue paths and action targets to the three types of intelligent agents: drones were responsible for high-altitude reconnaissance and target identification, wheeled robots were responsible for rapid mid-range traversal and establishing relay communication, and quadruped robots entered complex ruin areas to perform marking and identification tasks. Various control algorithms were deployed to the corresponding intelligent agent models through standard interfaces, and the system completed the full-process simulation and performance evaluation of the search and rescue mission.

[0140] To implement the heterogeneous multi-agent cooperative control testing method provided in this application embodiment, this application embodiment also provides an electronic device, such as... Figure 2 As shown, the electronic device 200 includes:

[0141] Central processing unit 201, memory 202, and input / output interface 203;

[0142] The memory 202 is a short-term storage memory or a persistent storage memory;

[0143] The central processing unit 201 is configured to communicate with the memory 202 and execute the instructions in the memory 202 to perform any of the above heterogeneous multi-agent cooperative control test methods.

[0144] Of course, in practical applications, the various components in the electronic device 200 are coupled together through the bus system 204. It is understood that the bus system 204 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 204 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 2 The general labeled all buses as Bus System 204.

[0145] The memory 202 in this embodiment is used to store various types of data to support the operation of the electronic device 200. Examples of such data include any computer program used to operate on the electronic device 200.

[0146] It is understood that when the processor in the above-described electronic device executes the computer program, it can also realize the functions of each unit in the corresponding device embodiments described above, which will not be repeated here. Exemplarily, the computer program can be divided into one or more modules / units, one or more modules / units are stored in memory and executed by the processor to complete the various embodiments of this application. One or more modules / units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device. For example, the computer program can be divided into units in the above-described electronic device, and each unit can realize the specific functions described in the corresponding electronic devices above.

[0147] Electronic devices can be computing devices such as desktop computers, laptops, handheld computers, and cloud servers. Electronic devices may include, but are not limited to, processors and memory. Those skilled in the art will understand that processors and memory are merely examples of electronic devices and do not constitute a limitation on electronic devices. They may include more or fewer components, combinations of certain components, or different components. For example, electronic devices may also include input / output devices, network access devices, buses, etc.

[0148] A processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of an electronic device, connecting all parts of the device through various interfaces and lines.

[0149] Memory can be used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. Furthermore, memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0150] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, performs the heterogeneous multi-agent cooperative control test method described in any of the above embodiments.

[0151] This application also provides a computer program product storing a computer program / instruction, which, when executed by a processor, is used to implement the heterogeneous multi-agent cooperative control test method described in the second aspect or any specific implementation of the second aspect of this application.

[0152] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0153] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0154] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0155] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0156] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A heterogeneous multi-agent cooperative control test system, characterized in that, include: The physics simulation engine module is used to build a simulation platform using the MuJoCo engine; wherein, the simulation platform integrates multiple types of intelligent agent models, and different types of intelligent agent models interact within the same simulation platform; The communication and interaction module is used to establish a unified communication protocol and a preset standard interface for data format between various types of intelligent agent models, so as to realize data interaction between different types of intelligent agent models. The modular modeling module is used to construct various types of intelligent agent models in a modular manner. The modules of the intelligent agent model include structural modules, motor modules, and sensor modules. The modules interact with each other on the simulation platform through a preset standard interface. The visualization testing module provides a graphical interface for dragging and dropping to deploy various control algorithms, presenting the status information of different types of intelligent agent models, presenting the task execution progress of different types of intelligent agent models, and receiving test and debugging commands triggered by users. The control algorithm module is used to integrate multiple control algorithms, including feedback control PID algorithm, model predictive control MPC algorithm and linear quadratic regulator LQR algorithm, and to deploy them to different types of intelligent agent models through a preset standard interface; The task control module is used to control multiple intelligent agent models of different types to perform behavior scheduling, path planning and action allocation based on the received collaborative task requests.

2. The system according to claim 1, characterized in that, The modular modeling module includes: A sensor module is used for configuring the physical properties of sensors and simulating data; the sensors include a camera, a lidar, and an inertial measurement unit (IMU). The actuator module provides a drive model and feedback control PID interface for the actuator, which includes a DC motor, a servo motor, and a servo motor. The structural module is used to define the structural components of the intelligent agent model and their physical properties; the physical properties include mass, friction coefficient, elastic coefficient, and stiffness parameter.

3. The system according to claim 1, characterized in that, The communication interaction module adopts a binary message format defined by the ROS2 framework. The message format includes a timestamp, an agent state array, and a control command array. The communication interaction module is also used to perform data compression processing on the output data of the sensor according to the sensor type and the current collaborative task; the data compression processing includes cropping high-frequency sensor data, or merging or interval sampling of low-frequency sensor data.

4. The system according to claim 1, characterized in that, The visualization testing module includes: A graphical programming workspace for dragging and dropping sensor modules, control algorithm modules, and actuator modules; The real-time rendering unit is used to display the position, task progress, and environmental obstacles of different types of intelligent agent models based on Unity3D. The parameter adjustment panel is used to adjust PID control parameters or reinforcement learning strategy hyperparameters.

5. The system according to claim 1, characterized in that, The visualization testing module also includes a scene construction module, which is used to generate virtual environments required for simulation testing of various types of intelligent agent models. The scene construction module includes the following various scene modeling methods: A graphical interface is constructed, allowing users to drag and drop preset environmental components into the simulated virtual environment and set the type, parameters, and spatial position of each component. The environmental components include terrain, obstacles, and targets. Scripted construction allows you to define the type, parameters, and spatial location of each environmental component in the scene by writing configuration scripts. Natural language construction utilizes large language models to perform semantic parsing of the natural language input by the user and generates the type, parameters, and spatial location of each environmental component; Reinforcement learning-assisted construction automatically generates the type, parameters, and spatial location of each environmental component based on the collaborative task of the selected agent model.

6. The system according to claim 1, characterized in that, The control algorithm module includes: Basic control layer: PID controllers used to provide feedback control for various types of intelligent agent models, so as to control the joint position or speed; Collaborative Control Layer: Integrates Model Predictive Control (MPC) algorithms for path trajectory planning in various agent models; Autonomous Learning Layer: Integrates reinforcement learning policy interfaces for model training and policy inference based on probabilistic policy optimization (PPO) and soft behavior policy optimization (SAC) algorithms, enabling each agent model to make decisions in different virtual environments.

7. The system according to claim 1, characterized in that, The system also includes a task scheduling module, which specifically includes: The task monitoring module is used to monitor the remaining battery power, communication quality, and priority changes of the currently executing tasks of each intelligent agent in real time; the communication quality includes communication latency and packet loss rate. The mode switching module is used to select the appropriate collaboration mode for each agent model based on the information provided by the task monitoring module; the collaboration modes include cooperative mode, competitive mode and hybrid mode, and to issue mode switching instructions to each agent.

8. A heterogeneous multi-agent cooperative control testing method, characterized in that, Applications include heterogeneous multi-agent cooperative control test systems, including: A simulation platform is built using the MuJoCo engine; the simulation platform integrates multiple types of intelligent agent models, and different types of intelligent agent models interact within the same simulation platform; Establish a unified communication protocol and a pre-defined standard interface for data format among various types of intelligent agent models to enable data interaction between different types of intelligent agent models; Various types of intelligent agent models are constructed in a modular manner. The modules of the intelligent agent model include structural modules, motor modules, and sensor modules. The modules interact with each other on the simulation platform through a preset standard interface. The graphical interface provided by the visualization testing module is used to receive various control algorithms deployed by the user and test debugging commands triggered by the user. The graphical interface provided by the visualization testing module is also used to present the state information of different types of intelligent agent models and the task execution progress of different types of intelligent agent models. The control algorithms provided by the control algorithm module are deployed to different types of intelligent agent models through a preset standard interface. The control algorithms include feedback control PID algorithm, model predictive control MPC algorithm, and linear quadratic regulator LQR algorithm. Based on the received collaborative task requests, control multiple intelligent agent models of different types to perform behavior scheduling, path planning, and action allocation.

9. The method according to claim 8, characterized in that, The method further includes: Based on the sensor type and the current collaborative task, data compression processing is performed on the output data of the sensor; the data compression processing includes cropping high-frequency sensor data, or merging or interval sampling of low-frequency sensor data.

10. An electronic device, characterized in that, include: Central processing unit, memory, and input / output interfaces; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method of any one of claims 8 to 9.

Citation Information

Cited By

  • Multi-joint robot cooperative control method and system based on reinforcement learning

    CN121680096A

  • Multi-agent collaborative particle accelerator debugging method, device, equipment and medium

    CN122194962A