Multi-unmanned aerial vehicle cooperative autonomous adjustment method under sudden failure
By using a deep deterministic policy gradient network and an improved generalized linear model, the load swing and stability issues of a multi-UAV cooperative sling system under sudden failures were solved, thereby improving the system's robustness and control accuracy and ensuring formation stability.
Patent Information
- Application Number
- CN202411596594.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-11-11
AI Technical Summary
In the event of a sudden failure, the load swing of a multi-UAV collaborative hoisting system is difficult to control, leading to system instability and making coordinated control challenging. Existing fault-tolerant algorithms have insufficient response speed and are unable to cope with dynamic changes in the hoisting system.
A deep deterministic policy gradient network and an improved generalized linear model are established. By acquiring information about UAVs and their payloads, a virtual model is built for iterative optimization and real-time fault detection. The position, payload allocation, and flight trajectory of the UAVs are adjusted to achieve formation reorganization and minimize payload oscillation.
It improves the robustness, stability, and control precision of the multi-UAV collaborative hoisting system, ensuring that the system maintains stable operation in the event of a failure.
Smart Images

Figure CN119472720B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of unmanned aerial vehicles, and particularly relates to a method for autonomous adjustment of multiple unmanned aerial vehicles in cooperation with load under sudden failure. BACKGROUND
[0002] In recent years, unmanned aerial vehicles have made certain progress in the field of load transportation, and especially a four-rotor hanging system has been proved to be able to complete the task of load transportation, including basic tasks such as vertical take-off and landing, hovering, and the like, and has certain maneuverability and safety. However, due to the low rated load capacity of a single four-rotor unmanned aerial vehicle, the operation capacity thereof is subject to the limited load capacity of a single unmanned aerial vehicle, and a multiple unmanned aerial vehicle cooperative hanging system can well solve the limitation of load capacity.
[0003] The stability of the formation control of the multiple unmanned aerial vehicle cooperative hanging system has a significant influence on the swing of the load, and the instability factors of a single unmanned aerial vehicle can easily lead to the collapse of the entire hanging system. At present, the formation control algorithm of the multiple unmanned aerial vehicle cooperative hanging system includes the leader-follower method, the behavior-based method, the virtual structure method, the graph theory method and the consensus-based method, among which the leader-follower method is more widely used. In the leader-follower method, the leader is usually directly controlled by a person or a computer, and is responsible for determining the direction and speed of the entire formation, while the follower autonomously adjusts its own position according to the relative position and direction information of the leader to maintain the formation of the formation. However, this method relies on one or more leaders to guide the entire formation, and if the leader fails or is disturbed, the entire formation can lose direction, leading to the collapse of the formation structure and the crash of the load; the behavior-based method can provide distributed control and better adaptability, but the complexity of behavior fusion can lead to conflicts in control instructions, making it difficult to maintain consistency in the formation, and when facing dynamic changes in the environment or task requirements, behavior conflicts can reduce the coordination between formation members, increasing the risk of collision; the virtual structure method requires unmanned aerial vehicles to maintain fixed relative positions, which can not be practical when passing through narrow spaces or performing complex maneuvers, and lacks flexibility and adaptability in obstacle avoidance, and strict formation constraints can lead to frequent position adjustments when encountering obstacles or other unmanned aerial vehicles, increasing energy consumption and possibly causing actuator saturation, thereby affecting the stability of the formation; the graph theory method can be difficult to quickly adapt to structural changes in large-scale or dynamic changes in the formation, leading to control delay and decision-making errors; the consensus method relies on the consistency of information between formation members, but in actual application, communication delay and information loss can damage the consistency, making it difficult to maintain synchronization between formation members, thereby affecting the overall performance and stability of the formation.
[0004] The unmanned aerial vehicle formation fault-tolerant control method refers to a technology and method for taking corresponding control strategies to maintain stable flight of the formation in the face of possible individual unmanned aerial vehicle failures during unmanned aerial vehicle formation flight. The existing unmanned aerial vehicle formation fault-tolerant control methods mainly include fault-tolerant algorithms based on distributed control, sliding mode control and model predictive control. The fault-tolerant algorithm based on distributed control only requires each unmanned aerial vehicle to communicate and coordinate with adjacent unmanned aerial vehicles without the participation of a global control center. When a certain unmanned aerial vehicle fails, other unmanned aerial vehicles can adjust the formation through local information to continue to perform the task. However, when this method is applied to a suspended system, if some unmanned aerial vehicles fail, the load swing may increase, resulting in an increase in control complexity; using the sliding mode control algorithm, the control strategy is quickly switched when the unmanned aerial vehicle fails to maintain the stability of the system, and the fault is tolerated by online estimation of the unmanned aerial vehicle and timely adjustment of the control input. However, this method is prone to chattering when implemented, affecting control accuracy, and the load control of the suspended system is weak, making it difficult to cope with complex load dynamics and multi-unmanned aerial vehicle coordination; the fault-tolerant algorithm based on model predictive control uses the dynamic model of the system for prediction, and makes optimal control decisions according to the current system state and future output trajectory. When the unmanned aerial vehicle fails, the task is redistributed by adjusting the model prediction. This method has strong dependence on the system model, and when the system model error is large, the control effect will decrease significantly. Especially in the case of dramatic changes in the dynamics of the suspended load, this method has large computational load, poor real-time performance, and is prone to large delays, affecting system stability.
[0005] The main problem of the multi-unmanned aerial vehicle cooperative suspended system is the control of the load swing and swing angle. Once some unmanned aerial vehicles fail, the load swing of the suspended system may become difficult to control, leading to instability of the overall system. At the same time, the coordination control of the multi-unmanned aerial vehicle cooperative suspended system is more difficult, and the failure of a single unmanned aerial vehicle will break the original formation balance, requiring other unmanned aerial vehicles to quickly redistribute tasks. However, in the suspended system, the dynamic response of the load is more complex, increasing the difficulty of task redistribution. In addition, the suspended system has high real-time requirements for control, and any delay or mistake may lead to instability of the load. The response speed of many fault-tolerant algorithms after the failure of the unmanned aerial vehicle may not be sufficient to cope with the dynamic changes in the suspended system. It can be seen that although the multi-unmanned aerial vehicle cooperative suspended system can improve the upper limit of the load weight, its safety also brings greater challenges, and the failure of a single unmanned aerial vehicle is easy to lead to the collapse of the entire suspended system.
[0006] Therefore, there is an urgent need for a new multi-unmanned aerial vehicle cooperative load autonomous adjustment method under sudden failure to solve the above technical problems. SUMMARY
[0007] The application provides a multi-unmanned aerial vehicle cooperative load autonomous adjustment method under a sudden failure, aiming at the robustness, stability and control accuracy of the multi-unmanned aerial vehicle cooperative load.
[0008] The application provides a multi-unmanned aerial vehicle cooperative load autonomous adjustment method under a sudden failure, comprising the following steps:
[0009] S1, acquiring unmanned aerial vehicle information and load information of a multi-unmanned aerial vehicle cooperative suspension system, and establishing an unmanned aerial vehicle and load physical model according to the unmanned aerial vehicle information and the load information;
[0010] S2, establishing a multi-unmanned aerial vehicle cooperative suspension system virtual model according to the unmanned aerial vehicle and load physical model;
[0011] S3, establishing a deep deterministic policy gradient network, and iteratively optimizing the multi-unmanned aerial vehicle cooperative suspension system virtual model through the deep deterministic policy gradient network to obtain an optimized multi-unmanned aerial vehicle cooperative suspension system virtual model;
[0012] S4, dispatching the multi-unmanned aerial vehicle cooperative suspension system to perform a task, and controlling the multi-unmanned aerial vehicle cooperative suspension system based on the optimized multi-unmanned aerial vehicle cooperative suspension system virtual model;
[0013] S5, establishing a fault detection model based on an improved generalized linear model, and performing real-time fault detection on the multi-unmanned aerial vehicle cooperative suspension system through the fault detection model to determine whether an unmanned aerial vehicle in the multi-unmanned aerial vehicle cooperative suspension system has a fault, and if so, adjusting the position of the unmanned aerial vehicle, load distribution and flight trajectory in the multi-unmanned aerial vehicle cooperative suspension system according to the fault detection result.
[0014] Preferably, in step S1, the unmanned aerial vehicle information comprises unmanned aerial vehicle position and attitude change information, output thrust of a motor of the unmanned aerial vehicle and attitude adjustment torque information, and the load information comprises center of mass motion information, swing angle information and angular velocity information of the load.
[0015] Preferably, step S2 comprises the following substeps:
[0016] S21, setting a state space of the unmanned aerial vehicle according to the unmanned aerial vehicle and load physical model;
[0017] S22, setting an action space of the unmanned aerial vehicle according to the unmanned aerial vehicle and load physical model to obtain the multi-unmanned aerial vehicle cooperative suspension system virtual model.
[0018] Preferably, in step S21, the state space comprises position information, attitude information, load state, environment information and fault state information of the unmanned aerial vehicle.
[0019] Preferably, in step S22, the action space includes thrust vector adjustment information, attitude adjustment information, and speed adjustment information of the UAV. The thrust vector adjustment information is used to adjust the thrust to change the position and attitude of the UAV. The attitude adjustment information is used to adjust the pitch angle, roll angle, and yaw angle of the UAV. The speed adjustment information is used to adjust the speed of the UAV.
[0020] Preferably, step S3 includes the following sub-steps:
[0021] S31. Set up a reward mechanism for the drone in the deep deterministic policy gradient network;
[0022] S32. Establish a policy network and a value network for each UAV in the virtual model of the multi-UAV cooperative sling system through the deep deterministic policy network. The policy network is used to predict the optimal policy, and the value network is used to evaluate the action value of the UAV and initialize the experience replay pool.
[0023] S33. The virtual model of the multi-UAV collaborative hoisting system is trained and iterated according to the reward mechanism, and the data information of the UAVs in the virtual model of the multi-UAV collaborative hoisting system after training and iteration is stored in the experience replay pool; wherein, the data information includes the current state, action, reward and next state of the UAV;
[0024] S34. Randomly select the data information from the experience replay pool, and update the parameters of the policy network and the parameters of the value network according to the data information;
[0025] S35. Determine whether the current iteration number is the maximum iteration number. If yes, output the current virtual model of the multi-UAV collaborative hoisting system as the optimized virtual model of the multi-UAV collaborative hoisting system. If no, return to step S33.
[0026] Preferably, the reward mechanism is set based on the relative distance between drones, the flight trajectory of drones, the obstacle avoidance capability of drones, and the swing angle of the payload carried by drones.
[0027] Compared with existing technologies, this invention acquires drone and load information of a multi-UAV cooperative sling system, establishes a physical model of the drones and load based on this information, and then establishes a virtual model of the multi-UAV cooperative sling system based on the physical model. A deep deterministic policy gradient network is established, and the virtual model of the multi-UAV cooperative sling system is iteratively optimized using this network to obtain an optimized virtual model. The multi-UAV cooperative sling system is then dispatched to perform tasks, and control is applied based on the optimized virtual model. A fault detection model based on an improved generalized linear model is established, and real-time fault detection is performed on the multi-UAV cooperative sling system to determine if any drones in the system are faulty. If so, the position, load distribution, and flight trajectory of the drones in the system are adjusted according to the fault detection results. Thus, by training the drones in the multi-UAV cooperative sling system, this invention enables each drone to achieve optimal action using local information, and by detecting faults in real time through the fault detection model, it achieves drone formation reorganization and minimizes load sway, effectively increasing the robustness, stability, and control accuracy of the multi-UAV cooperative sling system. Attached Figure Description
[0028] The present invention will now be described in detail with reference to the accompanying drawings. The above and other aspects of the present invention will become clearer and more readily understood through the detailed description following the accompanying drawings. In the drawings:
[0029] Figure 1 This is a flowchart of a multi-UAV collaborative cargo autonomous adjustment method under sudden failure conditions provided in an embodiment of the present invention. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0031] Please refer to Figure 1 This invention provides a method for autonomous adjustment of cargo carrying capacity by multiple unmanned aerial vehicles (UAVs) in the event of a sudden failure, comprising the following steps:
[0032] S1. Obtain drone information and load information of the multi-drone collaborative hoisting system, and establish a physical model of drones and loads based on the drone information and load information.
[0033] In this embodiment of the invention, in step S1, the UAV information includes UAV position and attitude change information, UAV motor output thrust and attitude adjustment torque information, and the load information includes load center of mass motion information, swing angle information and angular velocity information.
[0034] Specifically, the physical model of the UAV and payload includes kinematic and dynamic models of both. The UAV's dynamic model needs to include changes in its position and attitude, as well as the thrust and attitude adjustment torque generated by its motors. The payload's dynamic model needs to consider the payload's oscillation in three-dimensional space, including the motion of its center of mass, oscillation angle, and angular velocity. Additionally, it needs to consider the tension transmitted by the suspension cable and its reaction force on the UAV. The suspension cable can be modeled using a spring-damped model to accurately describe the payload oscillation caused by tension changes. Based on this, obstacle modeling uses a rigid body model and introduces a stochastic distribution model of wind disturbances in the environment, simulating the impact of wind on the system through fluid dynamics disturbance equations.
[0035] The multi-drone collaborative suspension system consists of four quadcopter drones and a carrier. The carrier is a rigid body, and the quadcopters are connected to the carrier via flexible cables. The flexible cables are suspended from the center point of the quadcopter's bottom, with the connection points between the flexible cables and the load symmetrically distributed. The system's communication architecture is distributed, with direct point-to-point communication between the quadcopters. Each quadcopter makes autonomous decisions based on its own status. Each quadcopter is equipped with force sensors, positioning sensors, inertial measurement units (IMUs), lidar, and vision sensors. These sensors are used to measure the tension of the flexible cables on the load, the relative positions between the drones, the drone's attitude, acceleration, and angular velocity, detect surrounding obstacles and their relative distances to other drones, and identify environmental features and the status of other drones through the vision system. An IMU is also installed on the carrier, primarily for measuring the carrier's attitude and acceleration.
[0036] S2. Establish a virtual model of a multi-UAV collaborative hoisting system based on the physical model of the UAV and the load.
[0037] In this embodiment of the invention, the physical model of the UAV is linked to the physical model of the load. The UAV and the load are coupled through a suspension line, and the force transmission on the suspension line needs to be modeled. The virtual model of the multi-UAV collaborative suspension system needs to control the tension applied by each UAV to ensure the stability of the load and prevent excessive swinging or loss of control.
[0038] The core of a multi-UAV collaborative sling-mounting system is the collaborative work of multiple agents; therefore, it is also necessary to model the cooperation among multiple UAVs. Communication between UAVs employs a distributed control strategy, with each agent sharing local information to achieve global goals. To address communication delays or data loss, the virtual model of the multi-UAV collaborative sling-mounting system adopts a fault-tolerant control mechanism, ensuring that the failure of a single UAV does not affect the stability of the entire formation. The collaborative control algorithm is based on a distributed consensus algorithm, implementing UAV pull distribution and load stability adjustment to ensure that the virtual model of the multi-UAV collaborative sling-mounting system can maintain consistent cooperation when facing complex tasks.
[0039] In this embodiment of the invention, step S2 includes the following sub-steps:
[0040] S21. Set the state space of the UAV according to the physical model of the UAV and the payload;
[0041] S22. Set the motion space of the UAV according to the physical model of the UAV and the load to obtain the virtual model of the multi-UAV collaborative hoisting system.
[0042] In this embodiment of the invention, in step S21, the state space includes the UAV's position information, attitude information, load status, environmental information, and fault status information. Specifically, the state space includes local information of each UAV and its relationship with other UAVs, mainly including UAV position information, attitude information, load status, environmental information, and fault status. UAV position information includes its own spatial coordinates, relative positions between UAVs, and the distance between the formation center point and the target point. Attitude information includes UAV speed, attitude angles (pitch angle, yaw angle, roll angle), and motor output thrust status. Load status includes the load's center of mass position, swing angle, and tension of the sling. Environmental information includes distances to obstacles and wind disturbances. Fault status includes marking faulty UAVs and the details of their faults.
[0043] In this embodiment of the invention, in step S22, the action space determines the control method of each UAV. The action space includes the thrust vector adjustment information, attitude adjustment information and speed adjustment information of the UAV. The thrust vector adjustment information is to adjust the thrust to change the position and attitude of the UAV. The attitude adjustment information is to adjust the pitch angle, roll angle and yaw angle of the UAV. The speed adjustment information is to adjust the speed of the UAV.
[0044] S3. Establish a deep deterministic policy gradient network, and iteratively optimize the virtual model of the multi-UAV collaborative hoisting system through the deep deterministic policy gradient network to obtain an optimized virtual model of the multi-UAV collaborative hoisting system.
[0045] In this embodiment of the invention, step S3 includes the following sub-steps:
[0046] S31. Set up a reward mechanism for the drone in the deep deterministic policy gradient network;
[0047] In this embodiment of the invention, the reward mechanism is set based on the relative distance between drones, the drones' flight trajectories, their obstacle avoidance capabilities, and the swing angle of the payload carried by the drones. Specifically, it includes minimizing payload swing by quantifying the swing amplitude based on the payload's acceleration and the rope's swing amplitude; balancing tension to ensure uniform tension distribution across each drone; ensuring reasonable formation positioning to avoid collisions or loss of formation control between drones and to prevent entanglement between different ropes; tracking effectiveness to verify whether the system flies along the set flight trajectory; and obstacle avoidance by adjusting the flight trajectory in a timely manner to avoid collisions when obstacles are detected by lidar and visual sensors. Distributed control is achieved through drone autonomous learning, with each drone learning the optimal cooperation strategy based on local information and interactions with other drones. Even in the event of communication delays or missing local information, smoother and more flexible cooperation between drones can be achieved through autonomous learning. Its advantage lies in its ability to dynamically adjust formation and payload allocation, thereby adapting to various complex task requirements.
[0048] S32. Establish a policy network and a value network for each UAV in the virtual model of the multi-UAV cooperative sling system through the deep deterministic policy network. The policy network is used to predict the optimal policy, and the value network is used to evaluate the action value of the UAV and initialize the experience replay pool.
[0049] S33. The virtual model of the multi-UAV collaborative hoisting system is trained and iterated according to the reward mechanism, and the data information of the UAVs in the virtual model of the multi-UAV collaborative hoisting system after training and iteration is stored in the experience replay pool; wherein, the data information includes the current state, action, reward and next state of the UAV;
[0050] S34. Randomly select the data information from the experience replay pool, and update the parameters of the policy network and the parameters of the value network according to the data information;
[0051] In this embodiment of the invention, TD error is used simultaneously to minimize the gap between the predicted and actual rewards, and the network is updated accordingly. Additionally, the policy network uses gradient ascent optimization to find a policy that maximizes long-term cumulative rewards. An entropy regularization term is added during policy updates to encourage policy exploration and avoid getting trapped in local optima. Data in the experience replay pool is selected through a priority replay mechanism to ensure that important state transitions are updated first.
[0052] Simultaneously, the policy network and value network are updated every few iterations. The training process terminates when the set number of training steps is reached, or when the load swing angle and drone formation error reach a certain threshold. All drones learn in parallel during training, sharing experience and rewards. Using a shared experience replay pool improves learning efficiency and promotes cooperative behavior among all drones.
[0053] S35. Determine whether the current iteration number is the maximum iteration number. If yes, output the current virtual model of the multi-UAV collaborative hoisting system as the optimized virtual model of the multi-UAV collaborative hoisting system. If no, return to step S33.
[0054] Specifically, the optimal formation position is found through the learning strategy of the deep deterministic policy gradient network, which can effectively deal with complex nonlinear control problems and fault tolerance problems of UAV formation, and can perform policy optimization in a high-dimensional continuous action space. The deep deterministic policy gradient network has the following three characteristics: (1) The optimal policy learned can give the optimal action using only local information when applied. (2) It does not need to know the dynamic model of the environment or special communication requirements. (3) The network can be used not only in cooperative environments but also in competitive environments. The framework of the network is centralized training and decentralized execution. Each UAV has an independent policy network and value network, which are used to predict the optimal policy and evaluate the value of the action. The basic framework of reinforcement learning training consists of actor, state, and action. During the training process, the actor first selects an action based on the current state. Then, the critic calculates a Q value based on this state-action, which is an evaluation feedback of the action selected by the actor. The evaluator compares the estimated Q-value with the actual Q-value for self-training, while the executor adjusts and updates its strategy based on the evaluator's feedback. During training, additional information can be incorporated into the evaluator phase to obtain a more accurate Q-value, such as the status and actions of other drones. This means each drone evaluates the value of its current action not only based on its own situation but also on the behavior of other drones. Decentralized execution refers to the situation where, after each drone has been sufficiently trained, each executor can independently take appropriate actions based on its current state, without needing information on the status or actions of other drones.
[0055] S4. Dispatch the multi-UAV collaborative hoisting system to perform the task, and control the multi-UAV collaborative hoisting system based on the optimized virtual model of the multi-UAV collaborative hoisting system;
[0056] S5. Establish a fault detection model based on an improved generalized linear model, and use the fault detection model to perform real-time fault detection on the multi-UAV collaborative hoisting system to determine whether there is a fault in the UAVs in the multi-UAV collaborative hoisting system. If so, adjust the position, load distribution and flight trajectory of the UAVs in the multi-UAV collaborative hoisting system according to the fault detection results.
[0057] In this embodiment of the invention, when the fault detection model detects a fault, it marks the faulty drone in the multi-drone collaborative hoisting system, expands the state space, adds a load pull distribution identifier, adjusts the weight of the reward mechanism, and performs processing such as drone formation reconstruction and pull redistribution.
[0058] Specifically, when the fault detection model detects a malfunction in a drone, the state space of the multi-drone cooperative sling system automatically expands to include fault identification and load distribution change information. Based on the fault state transition reward function, rewards are emphasized for minimizing load sway, reorganizing formation, and ensuring drone safety. After a fault occurs, other drones automatically adjust their formation based on the faulty drone's state, employing new strategies to maintain load balance. The multi-drone cooperative sling system re-plans the target position, attitude, and thrust distribution of each drone. The remaining drones adjust their thrust and position to share the load of the faulty drone, ensuring the load does not drastically sway due to the drone's malfunction. Fault detection based on the fault detection model can promptly identify faults in individual drones within the multi-drone sling system, and a deep deterministic policy gradient network enables policy transitions under fault states. Through state space expansion and reward function adjustment, the system can quickly respond to fault situations, ensuring drone formation stability, reducing load sway, and achieving fault-tolerant control of the system.
[0059] When a multi-drone sling system experiences a non-fatal fault, such as sensor malfunction, reduced power supply, or minor structural damage, resulting in limited drone mobility and decreased load-bearing capacity, the state space needs to be expanded and adjusted to reflect the impact of the fault on the system. An identifier variable should be added to the state space to distinguish the faulty drone and its affected area. For example, a specific drone could be identified as faulty, and its control variables could be marked as unavailable. Furthermore, due to the faulty drone's removal, other drones need to bear additional loads. The state space should include an identifier for the current load carried by each drone. Since load balance may be disrupted due to drone malfunctions, the state space should also include information on the degree of load imbalance, such as the rate of change of load swing angle and load offset.
[0060] When a multi-drone sling system enters a fault state, the reward function needs to be redesigned, and different weights should be assigned to the multi-objective reward function. After a drone failure, the reward function should primarily encourage the remaining drones to quickly adjust their formation to minimize load sway, so the reward weight for load should be increased. The reward function should also encourage the remaining drones to take on more load and complete formation reorganization in the shortest possible time. At this point, the pull and load of each drone need to be included in the reward function to encourage even distribution of pull. In addition, preventing other drones from failing due to overload or lack of coordination is crucial in the failure situation; the reward function should penalize multiple drone failures or overloaded drones. Furthermore, the reward function should maintain system stability as much as possible, and its reward function should also include failure response time, minimization of load sway, stability of the new formation structure, and avoidance of rope entanglement.
[0061] In multi-UAV collaborative sling systems, fault diagnosis is crucial. To ensure stable and safe system operation, fault diagnosis methods based on the improved generalized linear model (BLS) can effectively identify and handle faults in individual UAVs or the entire system. BLS is an efficient neural network model capable of rapid online learning and fault diagnosis. Compared to traditional deep learning models, BLS avoids complex hierarchical structures by directly constructing generalized feature maps and augmenting nodes, significantly improving learning efficiency. Based on the improved BLS model, faults in multi-UAV systems can be quickly learned and diagnosed. Its online learning mechanism adapts to environmental changes, making it suitable for fault diagnosis in complex and dynamic multi-UAV systems.
[0062] Compared with existing technologies, this invention acquires drone and load information of a multi-UAV cooperative sling system, establishes a physical model of the drones and load based on this information, and then establishes a virtual model of the multi-UAV cooperative sling system based on the physical model. A deep deterministic policy gradient network is established, and the virtual model of the multi-UAV cooperative sling system is iteratively optimized using this network to obtain an optimized virtual model. The multi-UAV cooperative sling system is then dispatched to perform tasks, and control is applied based on the optimized virtual model. A fault detection model based on an improved generalized linear model is established, and real-time fault detection is performed on the multi-UAV cooperative sling system to determine if any drones in the system are faulty. If so, the position, load distribution, and flight trajectory of the drones in the system are adjusted according to the fault detection results. Thus, by training the drones in the multi-UAV cooperative sling system, this invention enables each drone to achieve optimal action using local information, and by detecting faults in real time through the fault detection model, it achieves drone formation reorganization and minimizes load sway, effectively increasing the robustness, stability, and control accuracy of the multi-UAV cooperative sling system.
[0063] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0064] The embodiments of the present invention have been described above with reference to the accompanying drawings. The disclosed embodiments are merely preferred embodiments of the present invention. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many equivalent changes in form without departing from the spirit and scope of the claims of the present invention, and all such changes are within the protection scope of the present invention.
Claims
1. A method for autonomous adjustment of cargo-carrying capabilities among multiple unmanned aerial vehicles (UAVs) under sudden failure conditions, characterized in that: Includes the following steps: S1. Obtain drone information and load information of the multi-drone collaborative hoisting system, and establish a physical model of drones and loads based on the drone information and load information; S2. Establish a virtual model of a multi-UAV collaborative hoisting system based on the physical model of the UAV and the load; S3. Establish a deep deterministic policy gradient network, and iteratively optimize the virtual model of the multi-UAV collaborative hoisting system through the deep deterministic policy gradient network to obtain an optimized virtual model of the multi-UAV collaborative hoisting system. S4. Dispatch the multi-UAV collaborative hoisting system to perform the task, and control the multi-UAV collaborative hoisting system based on the optimized virtual model of the multi-UAV collaborative hoisting system; S5. Establish a fault detection model based on an improved generalized linear model, and use the fault detection model to perform real-time fault detection on the multi-UAV collaborative hoisting system to determine whether there is a fault in the UAVs in the multi-UAV collaborative hoisting system. If so, adjust the position, load distribution and flight trajectory of the UAVs in the multi-UAV collaborative hoisting system according to the fault detection results. When the fault detection model detects a fault, it marks the faulty drone in the multi-drone collaborative hoisting system, expands the state space, adds a load-pull allocation identifier, adjusts the weight of the reward mechanism, and reconstructs and redistributes the drone formation.
2. The multi-UAV collaborative cargo autonomous adjustment method under sudden failure as described in claim 1, characterized in that, In step S1, the UAV information includes information on changes in the UAV's position and attitude, information on the output thrust and attitude adjustment torque of the UAV's motors, and information on the load, including information on the load's center of mass motion, swing angle, and angular velocity.
3. The multi-UAV collaborative cargo autonomous adjustment method under sudden failure as described in claim 1, characterized in that, Step S2 includes the following sub-steps: S21. Set the state space of the UAV according to the physical model of the UAV and the payload; S22. Set the motion space of the UAV according to the physical model of the UAV and the load to obtain the virtual model of the multi-UAV collaborative hoisting system.
4. The multi-UAV collaborative cargo autonomous adjustment method under sudden failure as described in claim 3, characterized in that, In step S21, the state space includes the UAV's position information, attitude information, load status, environmental information, and fault status information.
5. The multi-UAV cooperative cargo autonomous adjustment method under sudden failure as described in claim 3, characterized in that, In step S22, the action space includes thrust vector adjustment information, attitude adjustment information, and speed adjustment information of the UAV. The thrust vector adjustment information is used to adjust the thrust to change the position and attitude of the UAV. The attitude adjustment information is used to adjust the pitch angle, roll angle, and yaw angle of the UAV. The speed adjustment information is used to adjust the speed of the UAV.
6. The multi-UAV cooperative cargo autonomous adjustment method under sudden failure as described in claim 1, characterized in that, Step S3 includes the following sub-steps: S31. Set up a reward mechanism for the drone in the deep deterministic policy gradient network; S32. Establish a policy network and a value network for each UAV in the virtual model of the multi-UAV cooperative sling system through the deep deterministic policy gradient network. The policy network is used to predict the optimal policy, and the value network is used to evaluate the action value of the UAV and initialize the experience replay pool. S33. The virtual model of the multi-UAV collaborative hoisting system is trained and iterated according to the reward mechanism, and the data information of the UAVs in the virtual model of the multi-UAV collaborative hoisting system after training and iteration is stored in the experience replay pool; wherein, the data information includes the current state, action, reward and next state of the UAV; S34. Randomly select the data information from the experience replay pool, and update the parameters of the policy network and the parameters of the value network according to the data information; S35. Determine whether the current iteration number is the maximum iteration number. If yes, output the current virtual model of the multi-UAV collaborative hoisting system as the optimized virtual model of the multi-UAV collaborative hoisting system. If no, return to step S33.
7. The multi-UAV cooperative cargo autonomous adjustment method under sudden failure as described in claim 6, characterized in that, The reward mechanism is set based on the relative distance between drones, the drone's flight trajectory, the drone's obstacle avoidance capability, and the swing angle of the drone's payload.
Citation Information
Patent Citations
Fault-tolerant control method for unmanned aerial vehicle sensor fault based on reinforcement learning
CN113467248A
Unmanned aerial vehicle group fault isolation and formation reconstruction method and system based on bird flock algorithm
CN116560399A