A UAV swarm digital twin system and method based on biological swarm intelligence learning

Through the digital twin system of drone swarm intelligence learning, combined with Bayesian reasoning and swarm intelligence learning optimization algorithm, the problems of local decision-making of individual behavior and group behavior evaluation in drone swarms are solved, and efficient intelligent task execution and operation optimization of drone swarms are realized.

CN119538700BActive Publication Date: 2025-09-23BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411359218.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-09-23
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

The existing drone swarm digital twin technology has not yet effectively solved the problems of local decision-making of individual behavior and centralized evaluation of group behavior in multi-agent scenarios, resulting in insufficient intelligence level and operational efficiency of drone swarm mission execution.

Method used

A drone swarm digital twin system based on biological swarm intelligence learning is adopted, including resource layer, simulation layer, communication layer and application layer. It combines Bayesian reasoning and swarm intelligence learning-inspired optimization algorithms, updates model parameters through chaos mapping and marine predator algorithm, and uses multi-agent proximal strategy optimization algorithm for deep reinforcement learning training.

Benefits of technology

It has significantly improved the intelligence level and operational efficiency of drone swarm mission execution, realized comprehensive monitoring, real-time prediction and efficient learning and training of drone swarms, and improved the efficiency and accuracy of model parameter updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119538700B_ABST
    Figure CN119538700B_ABST
Patent Text Reader

Abstract

The present invention discloses a digital twin system of a drone cluster based on biological swarm intelligence learning and a method thereof. The system architecture includes a resource layer, a simulation layer, a communication layer and an application layer. The resource layer is responsible for managing software platform resources. The simulation layer is the core of the platform and supports the simulation process of virtual models. The communication layer establishes communication links between various entities and models connected to the platform software. The application layer manages application services based on digital twins and provides a user interface. The method includes building a drone twin model, installing a communication interface, updating model parameters using Bayesian reasoning and optimization algorithms, and using twin models and simulations to support drone cluster training tasks based on swarm intelligence learning and deep reinforcement learning. The present invention can fully meet the functional requirements of drone cluster monitoring, prediction, updating and learning and training, improve the efficiency and accuracy of model parameter updates, and significantly improve the intelligence level and operational efficiency of drone cluster task execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a drone cluster digital twin system and method based on biological swarm intelligence learning, belonging to the field of drone digital twins. Background Art

[0002] Digital twin technology, a cutting-edge integrated technology, focuses on creating virtual representations of physical entities, achieving a seamless connection between the physical and digital worlds. This technology uses real-time data feedback to simulate, analyze, predict, and optimize the state of physical entities, playing a vital role in multiple fields. Digital twin technology has broad applications in aerospace, industrial manufacturing, healthcare, transportation, and other fields. In these fields, digital twins not only improve product quality and production efficiency, but also ensure safety and conserve resources. As a key technical support for Industry 4.0, digital twin technology, through its highly accurate simulation and predictive capabilities, provides a powerful tool for analyzing and making decisions about complex systems.

[0003] The characteristics of digital twins can be summarized from five dimensions: model, data, connectivity, services / functions, and physical. The model dimension requires digital twin models to possess high fidelity, high reliability, and high accuracy. The data dimension emphasizes data comprehensiveness and the ability to be dynamically updated in real time. The connectivity dimension highlights the seamless connection between the physical and virtual worlds. The service / function dimension encompasses a range of application services, such as product design, operation monitoring, and energy optimization. The physical dimension reflects the inseparable relationship between digital twins and physical entities. Based on these characteristics, a widely accepted digital twin system model architecture is the five-dimensional digital twin model, which provides a comprehensive framework for digital twin technology. This architecture emphasizes the importance of physical entities, virtual models, data connectivity, service functions, and cyber-physical interactions. As the foundation of digital twins, physical entities collect data through sensors and measurement technology, providing raw information for virtual models. Virtual models, in turn, utilize advanced simulation technologies to create accurate digital representations of physical entities. Data, as the link between the physical and virtual worlds, ensures that virtual models can reflect the status and changes of physical entities in real time.

[0004] At present, research on digital twin technology for drone swarms, especially the local decision-making of individual behavior and centralized evaluation of group behavior in multi-agent scenarios, and how to improve the intelligence level and operational efficiency of drone swarm mission execution, need to be urgently addressed. Summary of the Invention

[0005] 1. Purpose of the invention:

[0006] This paper provides a digital twin system and method for unmanned aerial vehicle (UAV) swarm intelligence learning. This platform aims to create a digital twin software platform for unmanned swarms. Through precise digital twin technology, it enables comprehensive monitoring, intelligent prediction, real-time updates, and efficient learning and training for both single and swarm drones performing typical tasks. This invention provides a comprehensive solution to enhance the intelligence level of unmanned swarms, optimize mission execution strategies, and improve overall operational efficiency and reliability.

[0007] 2. Technical solution:

[0008] Aiming at the typical mission execution scenarios of unmanned swarms, this paper takes unmanned single equipment and swarms as research objects and establishes a UAV swarm digital twin system based on biological swarm intelligence learning. The system architecture is as follows Figure 1 As shown:

[0009] A drone swarm digital twin system based on biological swarm intelligence learning is divided into four levels, namely resource layer, simulation layer, communication layer, and application layer.

[0010] (1) The resource layer is the foundation of the unmanned cluster digital twin platform and is mainly responsible for managing static software platform resources. It includes modules such as the virtual model library, twin database, application algorithm library, and communication protocol library.

[0011] Virtual Model Library: This library contains digital twin models corresponding to physical entities. These virtual models can include physical appearance, dynamics, mission logic, and simulation interface models. These models are typically stored in a modular format, including configurable model parameters. Each virtual model is mapped to a unique physical entity.

[0012] Twin database: The twin database is responsible for storing all necessary data generated during the platform's operation. This includes all data generated within the digital twin platform, such as direct simulation data of physical entities and virtual models, virtual-real communication data between modules, and log data.

[0013] Application Algorithm Library: This library includes packaged specific algorithm models, such as optimization algorithms and machine learning algorithms. Some applications in the digital twin platform software involve calling different algorithms depending on the scenario. For example, when using the twin environment for deep reinforcement learning training, algorithms with packaged input and output interfaces can be called as needed to achieve more efficient learning for different scenarios and objects.

[0014] Communication protocol library: The communication protocol library is responsible for storing various communication protocol formats. Different communication protocols and interface definitions are required between physical entities, virtual entities, and virtual and real entities in the digital twin platform. For example, wireless data transmission networking based on serial port protocols or data exchange between virtual models based on TCP / UDP are required.

[0015] The virtual model is used by users to perform basic operations on the virtual model, including:

[0016] Model import: The model import function can import platform virtual models in a specific format. These models are composed of multiple modular component models (such as controllers, motors or batteries, etc.) and then serve as simulation instances in a twin world.

[0017] Platform model creation: Platform model creation is used to organize modular component models into a platform instance model. In this link, the key parameters in the module will be determined, and the platform model will become a mapping of a specific physical object.

[0018] Model definition: Digital twin models can include models from multiple fields, such as mechanical models, electrical models, thermal models, etc., and can also include multi-dimensional models, such as physical models, three-dimensional models, behavioral models, etc. [2].

[0019] Model Storage: Physical components of the same type correspond to a fixed virtual model template. Different individuals modify key parameters based on this template. Therefore, model storage can be simplified to the storage of types and key parameters. In the digital twin platform software, model storage uses XML format to store the model template address and key parameter list.

[0020] The last part of the resource layer involves the processing of all key data involved in the operation of the digital twin platform software. It includes the following modules:

[0021] Historical data import: The data import module enables the digital twin platform software to import various types of data generated by users before, such as simulation data, actual installation data, log data, etc., and is compatible with common mainstream data formats such as CSV, XML or TXT.

[0022] Data filtering: Data filtering allows users to search and view data based on its fields, and repackage and save the data of interest to facilitate subsequent processing and support top-level complex applications such as feature extraction or machine learning.

[0023] Data processing: Data processing mainly pre-processes specific types of data such as sensor data, including outlier processing, noise smoothing, and duplicate data removal.

[0024] Data Reinjection: Supports output data generated from simulation models and physical implementations to be used as input for further simulations, allowing top-level applications to reanalyze past simulation routines.

[0025] (2) The simulation layer is the core of the digital twin platform, focusing on supporting the simulation process of virtual models. It consists of two parts: simulation configuration management and simulation engine. Simulation configuration management covers simulation instance configuration, algorithm parameter configuration, scenario configuration, and communication configuration, allowing users to set the simulation environment and parameters according to specific needs.

[0026] Simulation instance configuration: Simulation instance configuration is used to clarify the initialization requirements of the platform model instance involved in the simulation scenario of the twin world according to prior requirements, including setting its initial position, load configuration, performance envelope, etc.

[0027] Algorithm parameter configuration: Algorithm parameter configuration is used to configure the parameters of the algorithm executed in the top-level application according to user expectations and application scenarios, such as the search weight parameters and boundary values ​​of the optimization algorithm, the learning rate, discount factor, exploration rate of the reinforcement learning algorithm, etc.

[0028] Scenario configuration: Scenario configuration involves configuring environmental factors of the simulation execution scenario, such as the number and size of obstacles, the configuration of virtual targets, and the materials and general models of the rendering design. It is configured according to user needs and randomly deployed within a certain range to support interaction in learning.

[0029] Communication configuration: The digital twin platform involves multiple connections, and the data content of each communication party varies. Therefore, configuration is required to clearly define the field composition between the two communicating parties. This typically includes a fixed packet header and body to distinguish different data types and specific content.

[0030] The simulation engine is based on Isaac SIM software, part of the Nvidia Omniverse platform. This advanced robotics simulation and development platform aims to provide researchers and developers with a powerful tool for designing, testing, and optimizing the behavior and performance of robots in complex environments. Leveraging Nvidia's powerful Graphical Processing Units (GPUs) and deep learning technologies, the platform delivers a highly realistic virtual environment, enabling users to conduct large-scale parallel simulations with unprecedented speed and efficiency. Advantages of Isaac SIM include its ease of use, flexibility, and scalability. It supports multiple Robot Operating Systems (ROS) and common robotic hardware interfaces, allowing existing ROS users to seamlessly migrate to the platform. Its modular design allows users to customize the simulation environment and experiments based on their needs. Furthermore, Isaac SIM offers a rich API and plugin system for integrating third-party tools and extending functionality, further enhancing its potential in robotics simulation. Therefore, ROS-based control programs installed on task robots in a ROS-based deployment platform can be easily migrated to the Isaac SIM environment, significantly facilitating real-time, parallel simulation of digital twin models.

[0031] (3) The communication layer plays a vital role in the digital twin technology framework. It establishes the communication links between various entities and models connected to the digital twin platform software. The communication layer not only includes direct serial interface communication with physical devices, but also covers wireless communication hardware such as XBEE radio modules or radio frequency data transmission, as well as network communication based on wired LAN. In addition, communication supports the transmission and storage of configuration files and log files. In order to handle the data interaction between remote servers and on-site edge computing resources, the communication layer adopts intranet penetration technology to build a communication bridge based on the public Internet, which enhances the applicability and flexibility of the platform. In order to support isomorphism with the actual model, the twin model supports the same MAVLink protocol as the actual flight control, and communicates with the ROS control program of the mission aircraft through the interface.

[0032] (4) The application layer is the top layer of the digital twin platform architecture, dedicated to managing digital twin-based applications and services and creating a user-friendly operation interface. Relying on the support of the resource layer, simulation layer, and communication layer, the application layer builds a comprehensive access point for application services.

[0033] White-box monitoring allows users to indirectly monitor the physical drone by monitoring the virtual model, saving communication resources and achieving effective monitoring;

[0034] Model update: Correct and update model parameters through data recording during mission execution to maintain consistency between the digital twin model and the physical drone;

[0035] Task prediction: Using simulation environments to create highly similar virtual environments, quickly predict potential task outcomes and provide users with decision-making references;

[0036] Training and learning assist machine learning algorithms in training, reducing conversion errors and improving algorithm performance and efficiency. The application layer enables users to easily interact with the digital twin platform software to achieve comprehensive management and optimization of unmanned cluster tasks.

[0037] Based on a drone swarm digital twin system based on biological swarm intelligence learning, a drone swarm digital twin method of the present invention is as follows:

[0038] Step 1: Build a digital twin model based on the drone's physical structure. The virtual drone model can be divided into several sub-modules based on autopilot, sensor, power, and communication. Each module is modeled based on the mechanisms of the actual physical entity.

[0039] First, an autopilot model is established based on the flight control law of the physical drone object. The autopilot is an important component of the virtual drone model and is the core component that determines the flight status of the drone. It is responsible for converting the desired control input into the control signal of the motor.

[0040] Next, we build a powertrain model based on the physical drone object. The powertrain is responsible for converting the controller's control signals into the forces that change the drone's motion. The status of each component within the system is also of great interest to users, such as the actual motor speed, the electronic speed controller current, and the battery voltage and remaining charge. At the end of the controller modeling in the previous section, we obtain the desired motor speed control output signal. In the powertrain modeling section, we use this desired speed output to calculate the drone's actual motor speed and subsequently build virtual models for the brushless motor, electronic speed controller, and battery.

[0041] Finally, sensor models are built based on the payload of the physical drone. These sensors collect information about the surrounding environment within the simulation system, enabling the drone to perceive and adapt to complex flight conditions. Sensor modeling is essential for physical drones. To more realistically simulate the drone's control input conditions, simulated sensors are designed based on the physical drone's configuration and provide input for key communication messages. To ensure realistic simulation, necessary noise and errors should be added to the modeling.

[0042] Step 2: Build the actual communication interface. In the actual drone's communication logic, only the mission aircraft uses the ROS system for internal control. Communication and control between the flight control and the mission aircraft uses the MAVLink communication protocol. Therefore, it is necessary to increase support for the initialization and processing of MAVLink protocol messages in the drone model. Therefore, a communication interface model for the drone model is built based on Pymavlink, enabling the drone to process and send MAVLink messages, thereby supporting external programs or software and hardware to control the virtual drone model via the MAVLink protocol. At the same time, considering that MAVROS is widely used in ROS to improve the efficiency of controlling drones via the MAVLink protocol, the communication interface model will also consider the usage requirements of MAVROS's common topics and services to support the corresponding MAVLink protocol.

[0043] Step 3: Update the digital twin model parameters. Using Bayesian reasoning and crowd-intelligence-inspired optimization algorithms, the model data parameters are updated using data samples collected during the operation of the actual drone. A digital twin is a virtual mirror of the physical entity, providing a valuable reference for understanding the drone's behavior and state throughout its lifecycle. This requires continuous updating of model parameters to adapt to changes caused by wear, environmental conditions, and other operational factors.

[0044] When updating parameters, multiple combinations may produce similar model outputs, and their performance parameters generally do not change suddenly. Therefore, in order to prevent large jumps in results, Bayesian inference methods are used to perform regular parameter updates. The maximum a posteriori estimate of the parameter θ is The calculation of is as follows:

[0045]

[0046] Here, L(D|θ) is the maximum likelihood obtained using a dataset D of length k given the drone model parameters θ. P(θ) represents the prior probability of θ, which is assumed to follow a Gaussian distribution. In addition, logarithmic operations are used to simplify the calculation. Using the model output x| under parameters θ, θ , and the actual state of the next state in the dataset D The negative sign is added because we want the error to be as small as possible, while the likelihood is expected to be as large as possible.

[0047] Furthermore, we introduce an optimization algorithm inspired by swarm intelligence learning to solve Bayesian inference. We first use a chaotic map to initialize the population position. Logistic mapping is a common method used in chaotic mapping to simulate biological population growth. The algorithm's initial population is randomly initialized within the range [0, 1], then undergoes several chaotic mapping cycles before finally transitioning to the search space. The mathematical expression for the logistic map is as follows:

[0048] x n+1 =ax n (1-x n ) (3)

[0049] Among them, x n represents the position vector of the individual in the range [0,1] before the nth chaotic mapping. In practice, to improve efficiency, the fitness of the original randomly initialized population is compared with the fitness of the initialized population processed by the chaotic mapping, and the better half is finally retained for algorithm iteration.

[0050] Furthermore, the Marine Predators Algorithms (MPA) algorithm is used as the basis for optimization iteration. The algorithm iteration is divided into three stages. The first stage uses the following formula:

[0051]

[0052] Among them, Prey and Elite represent the prey and predator matrices, and the predator in each round is transformed from the prey with the highest fitness. B is a normally distributed random vector representing Brownian motion, P and R are constants, usually taking values ​​of (0,1).

[0053] The second stage formula is as follows:

[0054]

[0055] where R L is a Lévy random vector, is an adaptive parameter that controls the compensation of prey motion.

[0056] The formula for the third stage is as follows:

[0057]

[0058] Furthermore, a Trap-Avoidance Operator (TAO) is introduced to improve the effectiveness of the algorithm in dealing with real-world problems, enabling individuals to achieve significant quality improvements. Specifically, at the end of each iteration, if the random number is less than a threshold, a TAO operation is performed on the individual. The mathematical expression of TAO update is as follows:

[0059]

[0060] Where λ1 and λ2 are random uniform numbers in the range of (-1, 1) and (-0.5, 0.5), respectively. μ1 and μ2 are values ​​based on a random number Δ in the range of 0 to 1. TAO uses these random conditions, using the optimal individual and the average position as a reference, to make the population more diverse, thereby escaping the local optimal solution.

[0061] Step 4: Conduct deep reinforcement learning training based on swarm intelligence. By introducing swarm intelligence, training tasks are carried out in a multi-agent scenario using distributed control and centralized learning. Leveraging established digital models and high-precision simulations, along with the computing resources of remote servers, this system can effectively support deep reinforcement learning training tasks for drone swarms.

[0062] Similar to the self-organizing behavior in swarm learning, in a multi-agent system, intelligence is distributed among the agents, each possessing only local information and computational capabilities. Furthermore, individuals typically perceive only their local environment and the partial states of other units within the group. Local interactions influence global behavior. Therefore, local observation information is established based on the mission scenario requirements and the drone's observation capabilities.

[0063] In a group, the decision-making process of each agent is usually based on simple rules or strategies, which are continuously learned and optimized through interaction with the environment. In deep reinforcement learning tasks, the behavior of individual drones is relatively simple, such as speed and position control.

[0064] One of the goals of swarm learning is to generate complex collective behaviors through collaboration or competition among individuals. This is reflected in deep reinforcement learning, where actions are generated by a distributed policy network, while an evaluation network receives information from all agents and updates their parameters. This embodies the characteristics of local decision-making for individual behavior and centralized evaluation of group behavior.

[0065] To achieve this training goal, Multi-agent proximal policy optimization (MAPPO) is used as a deep reinforcement learning training algorithm. In MAPPO, each drone is regarded as an agent, and the drone cluster formation process is regarded as a Markov decision process, which can be defined as Where S, A, P, and R represent the state space, action space, state transition probability, and reward function, respectively. The drone, acting as a distributed intelligent agent, generates an action based on local observations and a shared policy, interacts with the environment, and generates rewards. Its policy network uses only local state information, while the critic network utilizes global state information. Specifically, each agent receives a local observation and outputs an action probability, and all agents use a single policy network. Each agent's critic network receives observations from all agents and outputs a value, which is used to update the policy network.

[0066] 3. Advantages and effects:

[0067] The present invention proposes a digital twin system of a drone cluster based on biological swarm intelligence learning and a method thereof: 1. The present invention provides an innovative digital twin system framework for a drone cluster, which integrates the resource layer, simulation layer, communication layer and application layer to form an efficient and reasonable workflow that can fully meet the functional requirements of drone cluster monitoring, prediction, updating and learning and training; 2. The present invention adopts an optimization algorithm based on swarm intelligence inspiration to update model parameters in real time, and uses Bayesian reasoning as the basic framework to solve the Bayesian reasoning problem by initializing the population position through chaos mapping and the improved marine predator algorithm for trap avoidance operation, thereby improving the efficiency and accuracy of model parameter updating; 3. The present invention constructs a set of drone deep reinforcement learning methods that combine swarm intelligence learning and digital twin technology, and realizes distributed control and centralized learning through the multi-agent proximal policy optimization (MAPPO) algorithm, which solves the problems of local decision-making of individual behavior and centralized evaluation of group behavior in multi-agent scenarios, and significantly improves the intelligence level and operation efficiency of drone cluster task execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 This is a framework diagram of the drone swarm digital twin system and its method.

[0069] Figure 2 It is the trajectory diagram of speed control mode.

[0070] Figure 3 It is the motor simulation curve.

[0071] Figure 4 This is the ESC simulation curve.

[0072] Figure 5 It is a battery simulation curve chart.

[0073] Figure 6 This is a battery life prediction chart.

[0074] Figure 7 It is the comparison curve between the GPS sensor output position and the true value.

[0075] Figure 8 It is the comparison curve between the angular velocity output by the IMU sensor and the true value.

[0076] Figure 9 It is a comparison chart of the fitness curve of model parameter update.

[0077] Figure 10 It is the comparison between the simulated trajectory and the real trajectory before and after the model parameters are updated.

[0078] Figure 11 This is a comparison chart of the reward curve during the policy network training process.

[0079] Figure 12 This is a comparison chart of the average reward curve per round for a single drone. DETAILED DESCRIPTION

[0080] The effectiveness of the system and method proposed in this invention is verified through a specific drone swarm digital twin example. The specific steps of the system and method are as follows:

[0081] Step 1: Create a digital twin model based on the physical structure of the drone.

[0082] (1) Autopilot

[0083] Taking the actual flight control system equipped with Ardupilot firmware as an example, the design of the autopilot follows its actual logical architecture.

[0084] In terms of horizontal control, the control law of the UAV virtual controller model can be described by the following formula:

[0085]

[0086] in, The desired horizontal speed control quantity is firstly controlled by the position P controller using the position Pos of the UAV on the horizontal plane. x,y With expected position Then, the speed PID controller uses the horizontal speed deviation to convert it into the desired roll angle φ and pitch angle θ. Then, in the attitude control loop, a PD controller is used to calculate the speed change component required to bring the aircraft to the desired attitude angle. It should be noted that the expected yaw angle The calculation of the desired roll angle and the desired pitch angle are different. The latter is obtained according to formula (12), while the former is divided into two cases: when the control mode includes yaw angle control, the desired yaw angle is directly specified by the control command of the upper task machine, otherwise

[0087]

[0088] For vertical control, the control law of the UAV virtual controller model can be described by the following formula:

[0089]

[0090]

[0091] As shown in equations (15) and (16), the position P controller is first used to set the position Pos on the Z axis to z and expectations The deviation is converted into the expected vertical speed That is, the climb rate, which is then distinguished from the level control, and the speed PID controller converts the speed deviation into the required desired longitudinal acceleration Furthermore, considering the motor speed ω in the hovering state h satisfy

[0092] 4k F ω h 2 =mg (17)

[0093] where K F is the thrust coefficient of the motor propeller, m is the mass of the drone, and g is the acceleration due to gravity.

[0094] Combining the above horizontal and vertical control laws and considering the X-frame configuration of the drone, the vertical speed variable, the speed variable causing attitude change, and the speed of the four motors satisfy the following relationship:

[0095]

[0096] Through this hierarchical control strategy, users can input the desired control amount from different positions of the control loop according to the scene in which the drone is located, thereby achieving accurate and stable control in various flight modes, further providing more flexible options for the execution of top-level tasks, and improving flight efficiency and safety. Taking speed control as an example, the simulation curve is as follows Figure 2 shown.

[0097] (2) Power system

[0098] The final output of the motor modeling part is the average input voltage of the motor Um and operating current I m The expression is:

[0099]

[0100]

[0101] Among them, and R m The no-load speed of the motor is proportional to the no-load input voltage U of the motor. m0 , this ratio is called the rated no-load motor constant K V0 , M is the torque generated by a single propeller.

[0102] For the ESC, considering the internal resistance of the battery, the input voltage U e and current I e It can be expressed as:

[0103] Ue=Ub-IbRb (21)

[0104] I e =σI m ,I e ≤I eMax (twenty two)

[0105] Among them I eMax The maximum operating current of the ESC is usually provided by the ESC manufacturer. b and R b Represent the battery current and internal resistance respectively. r For a multi-rotor drone with 2 motors, each motor is equipped with an ESC, and the power supply is used to power all the ESCs and onboard devices. The battery current can be expressed as:

[0106] Ib=nrIe+IFC+Icom (23)

[0107] Among them, I FC and I com The working current of the flight control device and the mission machine carried by the drone respectively. For the quadcopter drone hardware platform involved in this article, the power supply specifications of the flight control device and the Raspberry Pi mission machine are both 5V1A. Therefore, I FC and I com Usually 1A is acceptable.

[0108] When it comes to batteries, users are generally most concerned about the current remaining battery charge and the estimated battery life. The remaining battery charge is generally expressed using the State of Charge (SoC). Experiments have shown that for a lithium polymer (Li-Po) battery consisting of four cells, the relationship between its open circuit voltage and SoC can be expressed using the following polynomial:

[0109]

[0110] Among them, the open circuit voltage V oc is the sum of the battery's operating voltage and the voltage drop across the battery's internal resistance, so V oc ≈U b The Soc of a battery or battery pack at a given time is the ratio of the available charge at that moment to the total charge available when fully charged. Expressed as a percentage, a full charge is 100% and an empty charge is 0%. The most common method for calculating Soc is the ampere-hour method:

[0111]

[0112] Among them, t0 is the initial time, C T Indicates the total battery capacity in milliampere hours (mAh). Average battery discharge current and the current battery capacity to estimate the battery life T b .

[0113]

[0114] The simulations were conducted using three configurations of the powertrain, such as Figures 3 to 6 shown.

[0115] (3) Sensor system

[0116] For the research of physically installed drones, the most basic and critical sensors include the Global Positioning System (GPS) and the Inertial Measurement Unit (IMU).

[0117] Taking into account the measurement error, the GPS measurement value and the real value of the simulation environment satisfy the following relationship:

[0118]

[0119] Where p is the real position in the local ENU coordinate system obtained directly from the simulation system. is the measurement value given by the GPS simulation module, and Δt represents the simulation step size. prepresents Gaussian white noise with zero mean, simulated using a normally distributed random variable. p The random walk diffusion process represents a slow-changing process. It is a noise process with "memory", that is, the current noise value is affected by the previous noise value. It is worth noting that for a continuous-time random process, the variance grows linearly with time. Therefore, when simulating a continuous-time random process in a discrete-time system, it is necessary to ensure that the variance grows consistently, that is, the standard deviation is equal to the square root of the step size. Directly proportional.

[0120] IMU consists of a gyroscope and an accelerometer, which are used to measure the angular velocity and acceleration in the drone's body coordinate system. Similar to GPS, the angular velocity measurement value output by the IMU simulation module is and acceleration measurements With Gaussian white noise and random walk process, the expression is as follows:

[0121]

[0122] Where ω and a are the true angular velocity and acceleration in the body coordinate system obtained directly from the simulation system. g and η a The table indicates that the mean is zero and the variance is σ ng and σ na Gaussian white noise. g and b a A random walk diffusion process representing slowly changing angular velocity and acceleration measurements.

[0123] The sensor was simulated and verified, and the output curves and true value comparison curves of GPS and IMU were obtained as follows: Figures 7 and 8 shown.

[0124] Step 2: Build the actual communication interface

[0125] In the communication logic of the actual drone, only the mission aircraft uses the ROS system for internal control. Communication and control between the flight control and the mission aircraft uses the MAVLink communication protocol. Therefore, it is necessary to increase support for the initialization and processing of MAVLink protocol messages in the drone model. Therefore, a communication interface model for the drone model is constructed based on Pymavlink, enabling the drone to process and send MAVLink messages, thereby supporting external programs or software and hardware to control the virtual drone model through the MAVLink protocol. At the same time, considering that MAVROS is widely used in ROS to improve the efficiency of drone control via the MAVLink protocol, the communication interface model will also consider the usage requirements of MAVROS's common topics and services to support the corresponding MAVLink protocol.

[0126] Step 3: Update digital twin model parameters

[0127] Before updating the model, we first need to collect flight data from the physical drone. The data samples include the position and velocity in the North East Earth (NED) coordinate system, the attitude angle and angular rate expressed in Euler angles, and the desired velocity calculated by the mission computer as control input, and are collected at a frequency of 10Hz. This paper selects a drone that meets the above configuration to perform basic parameter identification. Assume that the baseline values ​​of the drone identification parameters are, I xx =0.0153, I yy =0.0153, I zz =0.0257, k f =1.23×10 -5 , k M =1.68×10 -7 .

[0128] A variety of classic and latest optimization algorithms are used for comparison, including: the original marine predator algorithm (MPA), pigeon optimization algorithm (PIO), Harris Hawk optimization algorithm (HHO), gray wolf optimization algorithm (GWO), golden sine algorithm (GSA), rime optimization algorithm (ROA), and the fitness curve changes are shown in Figure 2. Figure 9 The comparison of the trajectory diagram of the simulation model in the same task before and after parameter update and the real data is shown in Figure 10 shown.

[0129] Step 4: Conduct deep reinforcement learning training based on crowd learning

[0130] Consider a drone that can obtain the state of surrounding units through sensors. Given its target position, the drone's state vector s is defined as:

[0131] s=[g,n,o] (29)

[0132] Among them, g represents the target state information of the UAV in the surrounding direction, g i (i=0,1,…,m) represents the vector formed by the inverse of the target distance in m directions. If no target is detected in this direction, the corresponding g i = 0. Similarly, n and o represent the status information of neighboring drones and obstacles, satisfying

[0133]

[0134] Among them, d g Indicates the distance from the current position of the drone to the target position, d n and d o They represent the distance to the neighboring drone and the physical boundary of the obstacle respectively.

[0135] Considering the actual control law of UAVs, the action space of any individual UAV is defined as follows:

[0136]

[0137] Among them, h and v are the horizontal direction and speed of the UAV respectively. represents the action strategy set of UAV i, the value range of h is m directions evenly divided on the horizontal plane, and v is a discrete speed value.

[0138] For the reward function, define the formation reward function r f for

[0139]

[0140] tr[·] represents the trace of the matrix, and represent the normalized Laplace matrix of the current cluster configuration and the normalized Laplace matrix of the ideal configuration, respectively.

[0141] Scale reward function r s , defined as follows:

[0142]

[0143] Among them, Φ is the neighboring drone of drone i, D des is the expected distance between UAVs.

[0144] Target reward function r a , defined as follows:

[0145] r a =D t-1 -D t (34)

[0146] Among them, D t Indicates the distance between the drone and the target at that moment, including the ideal distance and noise. t-1 Indicates the distance between the drone and the target at the last moment.

[0147] Consider three drones flying in a triangular formation with an expected separation of 5 meters. They take off from a certain area and fly in formation over a target. The normalized Laplace matrix of the expected formation is:

[0148]

[0149] Figure 11The weighted total reward curve for each round of a single drone is shown. The policy network is also deployed on a real drone for flight testing. Since the control law logic is taken into account during training, the policy network trained in the twin environment can be easily transferred to the real drone. Figure 12 It can be seen that the test flight trajectory is basically consistent with the twin decision trajectory. It can be seen that the deep reinforcement learning decision model trained through the twin environment is also reliable in the real world.

Claims

1. A drone swarm digital twin method based on biological swarm intelligence learning, characterized by: The method includes: Step 1: Modeling based on the physical structure of the drone The virtual drone model is divided into several sub-modules based on the autopilot, sensor, power, and communication categories. Each module is modeled according to the mechanisms of actual physical entities. Specifically, the autopilot model is first established based on the flight control laws of the physical drone object; then the power system model is established based on the physical drone object; and finally, the sensor model is established based on the payload of the physical drone object. Step 2: Build the actual communication interface In terms of communication logic of the actual drone, the mission aircraft uses the ROS system for internal control, and the communication and control between the flight control and the mission aircraft uses the MAVLink communication protocol; A communication interface model for the drone model is built based on Pymavlink, enabling the drone to process and send MAVLink messages, thereby supporting external programs or software and hardware to control the virtual drone model through the MAVLink protocol; at the same time, the communication interface model also supports the corresponding MAVLink protocol using the common topics and service requirements of MAVROS; Step 3: Update digital twin model parameters Using Bayesian reasoning and an optimization algorithm inspired by swarm intelligence learning, the model data parameters are updated using data samples collected during the operation of the actual drone. Step 4: Conduct deep reinforcement learning training based on crowd learning Introducing swarm learning, training tasks are carried out in a multi-agent scenario using distributed control and centralized learning. Multi-agent proximal strategy optimization is used as a training algorithm for deep reinforcement learning. For UAV swarms, each UAV is regarded as an intelligent agent in MAPPO, and the UAV swarm formation process is regarded as a Markov decision process, which is defined as Among them, S, A, P and R are state space, action space, state transition probability and reward function respectively; the drone, as a distributed intelligent agent, generates an action based on local observations and a shared strategy, interacts with the environment and generates rewards; its policy network only uses local state information, while the evaluation network uses global state information; specifically, each agent receives a local observation and outputs an action probability, and all agents use a policy network; the evaluation network of each agent receives the observations of all agents and outputs a value, which is used to update the policy network.

2. The method according to claim 1, wherein: In the step three, the Bayesian inference method is used to perform regular parameter updates; the model output x| under the parameter θ is used. θ , and the actual state of the next state in the dataset D The likelihood is calculated by the mean squared error between ; the logarithmic operation is used to simplify the calculation, and the negative sign is added because we hope that the error is as small as possible, while the likelihood is expected to be as large as possible.

3. The method according to claim 2, wherein: An optimization algorithm inspired by swarm intelligence learning is further introduced to solve the Bayesian inference problem. First, the population position is initialized using chaotic mapping. The initial population of the algorithm is randomly initialized in the range of [0,1]. Then, after several chaotic mappings, it is finally converted to the search space range.

4. The method according to claim 3, wherein: The ocean predator algorithm is further used as the basis for optimization iteration.

5. The method according to claim 4, characterized in that: A trap avoidance operation is further introduced: at the end of each round of iteration, if the random number is less than the threshold DF, a TAO operation is performed on the individual.

6. A drone swarm digital twin system based on biological swarm intelligence learning, used to implement the method described in any one of claims 1 to 5, characterized in that: The system includes resource layer, simulation layer, communication layer and application layer; The resource layer is responsible for managing static software system resources, which includes virtual model library, twin database, application algorithm library and communication protocol library modules; The simulation layer is the core of the digital twin platform, supporting the simulation process of virtual models. It consists of two parts: simulation configuration management and simulation engine. Simulation configuration management covers simulation instance configuration, algorithm parameter configuration, scenario configuration, and communication configuration, allowing users to set the simulation environment and parameters according to specific needs. The communication layer establishes communication links between various entities and models connected to the digital twin system software. The communication layer not only includes direct serial interface communication with physical devices, but also covers wireless communication hardware and network communication based on wired LAN. In addition, the communication layer integrates the FTP protocol to support the transmission and storage of configuration files and log files. The communication layer uses intranet penetration technology to build a communication bridge based on the public Internet. The application layer is the top layer in the digital twin platform architecture, managing digital twin-based applications and services and creating a user-friendly operation interface.

7. The system according to claim 6, characterized in that: The virtual model library includes digital twin models corresponding to physical entities. The virtual models include physical appearance, dynamics model, task logic model, and simulation interface model. These models are stored in a modular form and include configurable model parameters. Each virtual model has a unique physical entity as a mapping. The twin database includes various data generated in the digital twin platform, such as direct simulation data of physical entities and virtual models, virtual-real communication data between modules, and log data; The application algorithm library includes packaged specific algorithm models. The communication protocol library: the communication protocol is responsible for storing various communication protocol formats; Different communication protocols and interface definitions are used between physical entities, virtual entities, and virtual and real entities in the digital twin platform; The virtual model is used by the user to perform basic operations on the virtual model, including: Model import: used to import a platform virtual model in a specific format. The model is composed of multiple modular component models and then used as a simulation instance in the twin world; Platform model creation: used to build modular component models into a platform instance model. In this step, the key parameters of the modules will be determined, and the platform model will become a mapping of a specific physical object; Model definition; Model storage: Model storage is simplified to the storage of types and key parameters. In the digital twin platform software, model storage uses XML format to store the address and key parameter list of the model template. The final part of the resource layer involves processing all key data involved in the operation of the digital twin platform software. It includes the following modules: Historical data import: The data import module enables the digital twin platform software to import various types of data generated by the user before; Data filtering: Data filtering allows users to search and view data based on its fields, and repackage and save the data of interest; Data processing: Data processing is performed on specific types of data, including outlier processing, noise smoothing, and duplicate data removal; Data Reinjection: Supports output data generated from simulation models and physical implementations to be used as input for further simulations, allowing top-level applications to reanalyze past simulation routines.

8. The system according to claim 6, characterized in that: The simulation configuration specifically includes: Simulation instance configuration: Simulation instance configuration is used to clarify the initialization requirements of the platform model instance involved in the simulation scenario of the twin world according to the pre-requirements; Algorithm parameter configuration: Algorithm parameter configuration is used to configure the parameters of the algorithm executed in the top-level application according to user expectations and application scenarios; Scenario configuration: Scenario configuration involves configuring the environmental factors of the simulation execution scenario, which is configured and randomized according to user needs to support interaction in learning; Communication configuration: The digital twin platform involves multiple connections. The data content of the communication between the parties is different. It is necessary to clarify the field composition between the two communicating parties through configuration; including fixed packet headers and packet bodies to distinguish the types and specific contents of different data.

9. The system according to claim 6, characterized in that: The application layer manages digital twin-based applications and services and creates a user-friendly operation interface, specifically including: White-box monitoring allows users to indirectly monitor the physical drone by monitoring the virtual model, saving communication resources and achieving effective monitoring; Model update: Correct and update model parameters through data recording during mission execution to maintain consistency between the digital twin model and the physical drone; Task prediction: Using simulation environments to create highly similar virtual environments, quickly predict potential task outcomes and provide users with decision-making references; Training learning, assists machine learning algorithms in training and reduces conversion errors.

Citation Information

Patent Citations

  • Real-time intelligent leakage monitoring and positioning method for hydrogen delivery pipe network without abnormal samples

    CN117072891A

  • Cluster collaborative electronic interference method based on digital twinning and deep reinforcement learning

    CN117890860A