A multi-unmanned aerial vehicle cooperative inspection method based on MASAC and MPC hierarchical control

CN122526282APending Publication Date: 2026-08-07NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANTONG UNIV
Filing Date
2026-03-25
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]本发明公开了一种基于MASAC与MPC分层控制的多无人机协同巡检方法,旨在解决现有端到端强化学习方法在多无人机协同轨迹规划中存在的策略收敛速度慢、终端控制精度不足以及控制输出缺乏物理可行性与安全性保障的技术问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122526282A_ABST
    Figure CN122526282A_ABST
Patent Text Reader

Abstract

The application provides a multi-unmanned aerial vehicle cooperative inspection method based on MASAC and MPC hierarchical control, and belongs to the technical field of unmanned aerial vehicle control and artificial intelligence. The technical scheme is as follows: S1, acquiring local observation state information of the multi-unmanned aerial vehicle in a complex inspection environment; S2, constructing a multi-unmanned aerial vehicle MASAC-MPC hierarchical cooperative control model; S3, completing model learning according to a full-state composite guidance interface and a multi-target reward mechanism; and S4, deploying and applying the multi-unmanned aerial vehicle cooperative inspection control method. The application can effectively decouple global cooperative decision of multi-agent reinforcement learning and bottom-layer constraint processing of model predictive control, explicitly deliver speed and attitude intentions through a full-state composite guidance mechanism to suppress control oscillation, and ensure optimization solving feasibility under physical limits through a complete soft constraint mechanism of the lower MPC.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of unmanned aerial vehicle (UAV) control and artificial intelligence technology, specifically to a multi-UAV collaborative inspection method based on MASAC and MPC hierarchical control. Background Technology

[0002] Intelligent drone technology is rapidly evolving towards a networked collaborative model, demonstrating significant value in critical infrastructure maintenance tasks such as power line inspection, search and rescue, and environmental monitoring. It can significantly improve efficiency and replace manual labor in hazardous environments. In these high-risk tasks, multiple drone systems need to navigate and maintain coordinated formation in unstructured environments, placing stringent demands on trajectory planning and control algorithms. Currently, collaborative trajectory planning is mainly divided into model-driven and data-driven methods. Traditional model-driven methods, such as model predictive control (MPC), while possessing explicit constraint handling capabilities, heavily rely on high-precision dynamic models and face enormous computational challenges in high-dimensional dynamic scenarios, struggling to cope with interference from unknown environments.

[0003] To overcome model dependence, data-driven methods, represented by deep reinforcement learning, have emerged and performed well in adapting to unknown environments. However, existing end-to-end multi-agent reinforcement learning (MARL) methods face significant bottlenecks in practical applications: on the one hand, relying solely on neural networks to directly output control quantities often leads to slow policy convergence, low terminal control accuracy, and difficulty in ensuring trajectory smoothness; on the other hand, end-to-end methods struggle to strictly meet the hard safety constraints of physical systems, and the output actions may exceed dynamic limits, resulting in serious flight safety hazards. Although some research has attempted to combine reinforcement learning with traditional control in hierarchical control architectures, there are still shortcomings in the inter-layer interaction mechanism. Dynamic inconsistencies often occur between upper-layer decisions and lower-layer execution, leading to control oscillations, and under extreme exploration strategies, the lower-layer controller is prone to failure due to constraint conflicts. Summary of the Invention

[0004] This invention discloses a multi-UAV cooperative inspection method based on MASAC and MPC hierarchical control, which aims to solve the technical problems of slow policy convergence speed, insufficient terminal control accuracy, and lack of physical feasibility and security guarantee of control output in the existing end-to-end reinforcement learning method for multi-UAV cooperative trajectory planning.

[0005] The inventive concept of this invention is as follows: This invention provides a multi-UAV collaborative inspection method based on MASAC and MPC hierarchical control. The method first constructs a three-dimensional inspection environment and kinematic model for multiple UAVs, including static facilities and random obstacles. Based on this, a MASAC-MPC hierarchical control architecture integrating decision-making and execution is established. In the upper-level decision-making stage, a multi-agent flexible Actor-Critic (MASAC) network is designed as a policy planner to dynamically plan robust short-term navigation waypoints based on the local observation states of the UAVs. To eliminate dynamic inconsistencies between layers, this invention designs a full-state composite guidance interface, enabling the upper-level network to not only output the desired position offset but also explicitly output the desired velocity vector and gimbal attitude intention, thus providing a complete state reference for the lower layer. In the lower-level execution stage, a fully soft-constrained Model Predictive Controller (MPC) is designed as the actuator to receive composite guidance commands from the upper layer. By introducing adaptive relaxation variables into the optimization problem, this controller can transform the upper-level reference commands into acceleration and angular velocity control sequences that satisfy dynamic constraints, ensuring the feasibility of the solver even when the upper-level commands exceed the physically feasible domain. The entire system adopts a centralized training and distributed execution paradigm, and combines a multi-objective reward function that includes progress, safety and stability to iteratively train the upper-layer network.

[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: a multi-UAV collaborative inspection method based on MASAC and MPC hierarchical control, comprising the following steps:

[0007] S1. Construct a multi-UAV 3D inspection environment that includes static line facilities, random obstacles, and electromagnetic interference areas of transmission lines, and establish a UAV kinematic model;

[0008] S2. At each decision moment, acquire local observation information of each UAV, the local observation information including at least its own motion state, relative target deviation, relative state of neighboring UAVs, and environmental obstacle perception information.

[0009] S3. Input the local observation information into the deep reinforcement learning policy network and output a full-state composite guidance command containing the desired position offset, desired velocity and desired gimbal attitude.

[0010] S4. Construct a model predictive controller that introduces slack variables, use the full-state composite guidance command as a tracking reference, solve for the optimal control sequence that satisfies the dynamic constraints, and apply the control sequence to the UAV;

[0011] S5. Based on a centralized training and distributed execution mechanism, the deep reinforcement learning policy network is iteratively updated using a multi-objective reward function to achieve multi-UAV collaborative inspection control.

[0012] Furthermore, step one includes the following steps:

[0013] S11. Construct a basic environment for multi-UAV 3D inspection: Establish a static facility model in 3D space that includes power transmission towers, power transmission cables and ground wires; at the same time, introduce a random obstacle generation mechanism, and at the beginning of each task, randomly generate cylindrical obstacles with uniformly distributed positions and sizes in the scene to simulate the uncertainty of unstructured environment.

[0014] S12. Establish a three-dimensional kinematic model of the UAV: ​​Abstract the UAV into a point mass motion system equipped with an independent gimbal, define its state vector including three-dimensional position, three-dimensional velocity and gimbal angle (yaw and pitch), and the control input vector including three-dimensional acceleration and gimbal angular velocity; the system state evolution follows the discrete-time first-order Euler integral rule and is subject to physical constraints of maximum velocity, maximum acceleration and gimbal rotation range.

[0015] S13. Establish a graded risk zone model for the high-voltage electromagnetic environment of transmission lines. First, define a strong interference no-fly zone (inner layer): This zone is centered on the cable axis and has a preset inner radius (denoted as ). The cylindrical space is defined as an insurmountable rigid geometric boundary. Any intrusion into this area is considered a serious incident and will cause the current mission to fail, in order to ensure physical safety.

[0016] S14. Define the weak interference risk zone (outer region): This region is a region with a radius of [missing information]. A ring-shaped columnar transition space is constructed within this area; a risk potential field inversely proportional to the distance from the cable axis is constructed within this area; this potential field is introduced as a soft penalty term into the environmental feedback of reinforcement learning, guiding the UAV to learn autonomously during the exploration process and to exhibit defensive flight behavior that actively maintains a safe distance.

[0017] Furthermore, step two includes the following steps:

[0018] S21. Define the local observation vector structure of the UAV, which specifically includes four types of key state information: its own motion state, which includes normalized absolute position coordinates, normalized linear velocity vector, and normalized gimbal yaw and pitch angles, used to describe the current dynamics and observation attitude of the UAV; relative target deviation, which includes the relative position vector relative to the final hovering target point, used to indicate the global navigation direction of the mission; neighbor cooperation state, which includes the set of relative position vectors of all neighboring UAVs within the communication range relative to itself, used to support obstacle avoidance and cooperation among multiple UAVs; and environmental perception features, which include relative position information relative to the center of the power transmission tower, the nearest power transmission cable, and obstacles within the detection range.

[0019] S22. Feature extraction and encoding of environmental perception information: For randomly distributed and variable-numbered obstacles in unstructured environments, a Top-N sparse coding mechanism is used to transform variable-length environmental information into fixed-dimensional feature inputs. First, the Euclidean distance between all perceived obstacles within the detection range and the current UAV is calculated. Second, all obstacles are sorted from nearest to farthest, and only the relative position vectors of the nearest N obstacles are selected for encoding. Finally, if the actual number of obstacles within the current detection range is less than N, the remaining input channels are filled with preset virtual far-point coordinates to ensure that the dimension of the observation vector input to the neural network remains consistent under different environmental complexities.

[0020] Furthermore, step three includes the following steps:

[0021] S31. Define the action vector structure: The upper-layer MASAC policy network outputs a full-state composite guidance instruction based on the current local observation. Its mathematical definition is:

[0022]

[0023] in, This is the desired position offset relative to the current position. The desired three-dimensional linear velocity when reaching this offset position. The yaw and pitch angles of the gimbal are the angles expected to be used when reaching this position.

[0024] S32. Establish a composite guidance mechanism: Use the position offset to generate a short-term navigation guidance point for the current moment. and will and As a dynamic attribute associated with this guiding point;

[0025] S33. Eliminate dynamic inconsistencies: The full-state composite guidance command is used as the reference input of the lower-level model predictive controller. By explicitly transmitting the velocity and attitude intentions, the zero velocity constraint implicit in traditional position guidance is eliminated, enabling the UAV to maintain momentum when passing through the guidance point, thereby effectively suppressing the sawtooth oscillation of the trajectory.

[0026] Furthermore, step four includes the following steps:

[0027] S41. Constructing a soft constraint mechanism: Abandoning the hard equality constraint in traditional MPC that requires the predicted state to be strictly equal to the reference target, adaptive slack variables are introduced for the tracking target's position, velocity, and gimbal angle. and The following soft constraint equations are established:

[0028]

[0029]

[0030]

[0031] Among them, the slack variable represents the minimum necessary deviation between the predicted state of the system and the upper reference target;

[0032] S42. Design a robust objective function: Construct the objective function for the quadratic programming problem. A joint penalty is applied to the control input, the smoothing term, and the slack variables:

[0033]

[0034] in, and These are the acceleration and gimbal angular velocity control inputs, respectively. and These are the corresponding weight matrices, and the configurations are as follows: The value of the matrix is ​​much greater than matrix;

[0035] S43. Ensuring solution feasibility: Through the soft constraint mechanism, it is mathematically ensured that the feasible domain of the optimization problem is always non-empty. Even when the upper-level reinforcement learning exploratory instructions exceed the boundary of the lower-level dynamics, the solver can still calculate the feasible control sequence with the least degree of violation, ensuring the continuous and stable operation of the system.

[0036] Furthermore, step five includes the following steps:

[0037] S51. Configure a multi-objective reward function: The multi-objective reward function is used to guide the system to smoothly transition from rapid maneuvering to precise hovering. The multi-objective reward function consists of progress reward, terminal performance reward, safety obstacle avoidance reward, decision continuity reward, and task completion reward, and its weighted sum is expressed as:

[0038]

[0039] in, , , , and These represent progress rewards, terminal performance rewards, safety and obstacle avoidance rewards, decision continuity rewards, and task completion rewards, respectively. , , and These are the corresponding weighting coefficients.

[0040] S52, the aforementioned progress reward This includes entity progress rewards based on the actual physical location of the drone and virtual progress rewards based on the location of upper-layer guide points, used to calibrate the consistency between the upper-layer strategy and the global mission direction. Its definition is:

[0041] in, and These represent the reduction in distance from the actual location of the drone and the guide point to the target point, respectively. and It is the weighting coefficient.

[0042] S53, the aforementioned terminal performance reward A distance-based exponential decay gating mechanism is employed, activating a constraint reward on the velocity modulus and a reward on the alignment of the camera's optical axis with the ideal line-of-sight only when the UAV approaches the final target point. These rewards are defined as follows:

[0043]

[0044] in, Let be the unit vector of the camera's orientation. The first term is the ideal unit vector pointing from the drone to the observed target; the second term rewards camera attitude alignment through vector dot product. These are the weight and the attenuation coefficient, respectively.

[0045] S54, the aforementioned safety obstacle avoidance penalty This is used to characterize the safety status between the drone and obstacles or neighboring drones; no penalty is imposed when the drone is outside the safe distance; a soft negative reward related to the degree of intrusion is imposed when the drone enters the safe warning distance; a huge negative reward for mission failure is imposed when a physical collision occurs or the drone intrudes into a strongly interfered no-fly zone, defined as:

[0046]

[0047] S55, the decision continuity penalty The norm squared of the difference between the composite guidance instructions of the full state at adjacent time points is used to suppress drastic changes in the output actions of the upper-layer policy network, thereby improving the smoothness of the lower-layer control. It is defined as follows:

[0048]

[0049] S56, Rewards for completing the task For sparse rewards, a pre-defined task completion reward is given if and only if all agents simultaneously satisfy the terminal constraints of position, velocity, and attitude. Its definition is:

[0050]

[0051] S57. Update the Critic network: Utilize the dual Q-network mechanism to mitigate overestimation bias by minimizing the following loss function. Update Critic network parameters:

[0052]

[0053] Where the target value It combines immediate rewards with soft value estimation in the next moment;

[0054] S58. Updating the Actor Network and Temperature Coefficient: The Actor network makes decisions based solely on local observations, updating parameters by minimizing the KL divergence. That is, to maximize the Q-value while increasing the policy entropy:

[0055]

[0056] At the same time, by minimizing the loss function Adaptive adjustment of temperature coefficient This allows the average entropy of the strategy to dynamically approximate the preset target entropy. .

[0057] Meanwhile, this invention proposes a multi-UAV collaborative inspection system based on MASAC and MPC hierarchical control. The system, applying the method described in this invention, includes the following steps:

[0058] The UAV kinematic model building module is configured to perform the following process: build a multi-UAV 3D inspection environment that includes static line facilities, random obstacles, and electromagnetic interference areas of power transmission lines, and establish a UAV kinematic model;

[0059] The UAV local observation information module is configured to perform the following process: at each decision moment, acquire local observation information of each UAV, the local observation information including at least its own motion state, relative target deviation, relative state of neighboring UAVs, and environmental obstacle perception information;

[0060] The output module is configured to perform the following process: inputting the local observation information into a deep reinforcement learning policy network and outputting a full-state composite guidance command containing the desired position offset, desired velocity, and desired gimbal attitude;

[0061] The model predictive controller module is configured to perform the following process: constructing a model predictive controller that introduces slack variables, using the full-state composite guidance command as a tracking reference, solving for the optimal control sequence that satisfies the dynamic constraints, and applying the control sequence to the UAV;

[0062] The multi-UAV collaborative inspection control module is configured to execute the following process: based on a centralized training and distributed execution mechanism, the deep reinforcement learning policy network is iteratively updated using a multi-objective reward function to achieve multi-UAV collaborative inspection control.

[0063] Meanwhile, the present invention proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed, it implements the steps of the method described in the present invention.

[0064] Furthermore, the present invention proposes a computer-readable storage medium having a computer program stored thereon, the computer program being configured to implement the steps of the method described in the present invention when invoked by a processor.

[0065] Finally, the present invention provides a computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method described in the present invention.

[0066] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0067] (1) Significantly improves the efficiency of collaborative decision-making and the success rate of missions involving multiple UAVs in unstructured environments. This invention decouples the high-dimensional and complex end-to-end control problem into low-dimensional macro-level decision-making and high-frequency micro-level execution through a MASAC-MPC hierarchical architecture. Experimental results show that in complex inspection scenarios containing randomly distributed obstacles, the success rate of the method in this invention reaches 93.30%, which is a significant improvement compared to the traditional end-to-end MASAC algorithm (71.23%) and MADDPG algorithm (25.73%).

[0068] (2) It effectively avoids the risk of lower-level solution failure caused by aggressive upper-level instructions in traditional hierarchical control. This invention designs a fully soft-constrained MPC controller at the lower level. By introducing adaptive relaxation variables and configuring high-priority penalty weights, the feasible region of the quadratic programming problem is guaranteed to be non-empty in most cases from the perspective of mathematical model construction. This enables the lower-level controller to calculate the suboptimal feasible control sequence with the least degree of violation even when the upper-level reinforcement learning is in the early exploration stage and outputs random guide points that deviate from the physical limits. This ensures the continuous and stable operation of the system under complex working conditions and greatly enhances the robustness of the control system to uncertain instructions. Attached Figure Description

[0070] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used to explain the invention in conjunction with Embodiment 1 and do not constitute a limitation thereof.

[0071] Figure 1This is a schematic diagram of the scenario and environment modeling for a multi-UAV power line inspection task in an embodiment of the present invention.

[0072] Figure 2 This is a schematic diagram of short-term navigation guidance points and full-state composite instructions generated by the upper-layer policy network in an embodiment of the present invention.

[0073] Figure 3 This is a schematic diagram of the MASAC-MPC hierarchical control framework in an embodiment of the present invention.

[0074] Figure 4 This is a graph showing the average reward convergence curves of the embodiments of the present invention and the comparison algorithm in barrier-free scenarios and random obstacle scenarios.

[0075] Figure 5 This is a flight trajectory diagram of the method of this invention (MASAC-MPC) and two baseline algorithms in a typical obstacle scenario.

[0076] Figure 6 This is a probability density distribution diagram of terminal control errors (position, speed, gimbal angle) in an embodiment of the present invention. Detailed Implementation

[0077] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0078] Example 1: This example provides a multi-UAV collaborative inspection method based on MASAC and MPC hierarchical control, including the following steps:

[0079] Step 1) Environment and task modeling: Construct a multi-UAV 3D inspection environment that includes static line facilities and random obstacles, establish a UAV 3D kinematic model, and establish a graded risk zone model for the electromagnetic environment of power transmission lines;

[0080] Step 2) State awareness and sparse coding: At each decision moment, acquire the local observation vector of each UAV and perform Top-N sparse coding on an indefinite number of obstacles;

[0081] Step 3) Upper-level full-state composite guidance: Input the local observation vector into the MASAC policy network and output a full-state composite guidance command that includes position offset, desired velocity and gimbal attitude;

[0082] Step 4) Execution of the lower-level fully soft-constrained trajectory: Construct a fully soft-constrained MPC controller, receive composite guidance instructions as a reference, solve the quadratic programming problem with introduced slack variables, and execute the first control term;

[0083] Step 5) Collaborative policy update: Based on the CTDE paradigm, the upper-level policy network is iteratively updated using a multi-objective reward function.

[0084] Step 1) includes the following steps:

[0085] 1-1) Constructing a multi-UAV 3D inspection scenario: This embodiment builds a multi-UAV collaborative simulation verification platform based on the Python programming language and the PyTorch deep learning framework. In the 3D space of this platform... The following elements are constructed in the middle:

[0086] Static facilities: Includes two cylindrical transmission towers (radius) ,high The transmission cables and overhead ground wires connecting the two towers are abstracted as three-dimensional line segments connecting the top of the towers.

[0087] Random obstacles: To improve the robustness of the algorithm, a domain randomization mechanism is introduced. At the beginning of each round, a random number of obstacles are generated. Cylindrical obstacles. The state of each obstacle is determined by its bottom center position. ,radius and height The description states that these parameters are all sampled uniformly within a preset range.

[0088] 1-2) Establishing a 3D kinematic model of the UAV: ​​Considering the computational efficiency of multi-UAV collaboration, the UAV is abstracted as a point mass motion system equipped with an independent gimbal. Define the UAV. exist State vector at time step ,in For location, For linear velocity, Let yaw and pitch be the gimbal angles. Define the control input vector. ,in This is a three-axis acceleration command. This refers to the gimbal angular velocity command. In the simulation platform's physics engine, the system state update follows the discrete-time first-order Euler integral rule, with a time step of... The specific state update formula is as follows:

[0089]

[0090] At the same time, the system is subject to physical constraints: , And mechanical limits on the gimbal angle and angular velocity.

[0091] 1-3) Establish a graded risk zone model: To address the nonlinear interference of strong electromagnetic fields around high-voltage transmission lines on UAV communication links and electronic equipment, instead of simply treating it as a geometric obstacle, a coaxial nested graded risk model is established, such as... Figure 1 As shown.

[0092] 1-3-1) Strong Interference No-Fly Zone (Inner Zone): Defined with the power transmission cable axis as the center and a radius of... The cylindrical space is a high-interference no-fly zone. Any intrusion into this area will be considered a serious incident, directly triggering mission failure and incurring a huge negative reward, in order to ensure physical safety.

[0093] 1-3-2) Weak Interference Risk Zone (Outer Region): Defined radius range as The annular cylindrical space is a low-interference risk zone. Within this zone, although the drone can maintain physical flight, it faces the risk of signal loss. This potential field is introduced as a soft penalty term into the reward function of reinforcement learning, guiding the agent to autonomously develop safe driving strategies away from cables during the exploration process.

[0094] Step 2) includes the following steps:

[0095] 2-1) Define the composition structure of the UAV local observation vector: at each decision time... drones Local observation vectors are obtained through airborne sensors and communication links. The components are then normalized to fit the neural network input; the specific mathematical expression of this vector is:

[0096]

[0097] In the formula, and drones Normalized self-position and linear velocity; For the gimbal's yaw and pitch angles; Relative to the final hovering target point The relative position vector is used to indicate the global navigation direction; For other neighboring drones within communication range Relative to itself The set of relative position vectors It is used to sense cluster formation to avoid inter-machine collisions; It is an environmental perception feature, containing key environmental information relative to transmission towers, transmission cables, and random obstacles.

[0098] 2-2) Environmental perception information Feature extraction and sparse coding are performed: For randomly distributed and variable-numbered obstacles in unstructured environments, a Top-N sparse coding mechanism is used to transform variable-length environmental information into fixed-dimensional feature inputs. The specific process is as follows: First, for static facilities, the relative position vector of the UAV's current position with respect to the center of the transmission tower, and the net distance with respect to the safety boundary of the transmission cable are calculated; second, for dynamic obstacles, the relative positions of all perceived obstacles within the detection range and the current UAV are calculated. The Euclidean distances are calculated and sorted in ascending order of distance; then, only the preset number of closest distances are selected. The relative position vectors of each obstacle are encoded as input, ignoring secondary threats at a distance; finally, if the actual number of obstacles within the current detection range is insufficient... If there are multiple virtual far points, then the preset virtual far point coordinates (e.g., ...) will be used. The remaining input channels are filled to ensure that the dimension of the observation vector input to the neural network remains consistent under different environmental complexities.

[0099] Step 3) includes the following steps:

[0100] 3-1) The local observation vector obtained in step 2) Input into the trained MASAC policy network Output full-state composite boot instructions .like Figure 2 As shown, this instruction includes not only spatial location information but also dynamic vector information upon reaching that location, specifically expressed mathematically as follows:

[0101]

[0102] The physical meanings of each component in the formula are as follows: position offset :like Figure 2 The dashed arrow indicates the spatial displacement vector from the current location of the UAV to the virtual guide point at the next moment, which is used to generate navigation waypoints. Expected speed :like Figure 2 The solid arrow at the virtual guide point indicates the three-dimensional instantaneous velocity vector that the UAV should possess upon reaching this virtual guide point. The direction of this vector indicates the direction of momentum during flight; desired attitude. As shown in the field-of-view cone below the guide point in the figure, it represents the target yaw and pitch angles of the gimbal camera when reaching that point, which are used to ensure that the field of view covers the observed target.

[0103] 3-2) The upper-layer policy network uses the full-state composite guidance instructions. This explicitly transmits the dynamic characteristics upon reaching the guide point to the lower-level MPC controller. Unlike traditional methods that only output position coordinates, this embodiment introduces... Figure 2 The non-zero velocity vector shown This forces the lower-level controller to track a non-zero momentum reference during optimization.

[0104] Step 4) includes the following steps:

[0105] 4-1) Constructing a linear prediction model and soft constraint equations:

[0106] The lower-level controller receives the full-state composite boot command output in step 3). As a reference trajectory. Figure 3 The system architecture diagram shown indicates that the MPC module is first based on a discrete-time dynamics model. State prediction is performed. In this embodiment, a prediction time domain is defined. To address the issue that traditional hard constraints often fail to provide solutions when dealing with aggressive instructions from higher levels, this embodiment abandons... Instead of strong constraints, slack variables are introduced. The following fully soft-constrained state equations are established:

[0107]

[0108] In the formula, The table represents the "minimum necessary deviation" between the system's predicted state and the upper-level reference target. This mathematical construction ensures that the feasible region of the optimization problem is non-empty under any operating condition.

[0109] 4-2) Constructing and solving the fully soft-constrained quadratic programming problem: Constructing the quadratic programming (QP) optimization objective function. This aims to minimize control costs and relaxation bias. The specific formula is as follows:

[0110]

[0111] This embodiment uses specific parameter configurations to establish a hierarchy of "priority tracking, suboptimal feasibility":

[0112] Controlling smoothing weights: Setting the acceleration weight matrix Angular velocity weight matrix of gimbal To suppress high-frequency jitter.

[0113] Relaxation penalty weights: Set relaxation penalty weights for position, velocity, and attitude. , , Both are 1.0. Because... The value of the matrix is ​​significantly greater than The solver will prioritize slack variables within the limits of physical capabilities. Optimize to near zero. Finally, the solver output contains the optimal acceleration sequence. and gimbal angular velocity sequence The control vector, and the first term of the sequence. It is applied to the motor control system of the drone.

[0114] Step 5) includes the following steps:

[0115] 5-1) Configure a multi-objective reward function model:

[0116] To guide the smooth transition of drones from rapid maneuvering to precise hovering, the following four core reward components are constructed. Its mathematical expression is:

[0117]

[0118] The specific definitions and parameter configurations of each item in this embodiment are as follows: As a reward for task progress, As a reward for terminal performance, Penalty for obstacle avoidance in safety For decision continuity penalties, Sparse rewards for task completion. These are the weighting coefficients for each item;

[0119] Progress Rewards This includes entity progress rewards based on the actual physical location of the drone and virtual progress rewards based on the location of upper-layer guide points, used to calibrate the consistency between the upper-layer strategy and the global mission direction. Its definition is:

[0120] in, and These represent the reduction in distance from the actual location of the drone and the guide point to the target point, respectively. and It is the weighting coefficient.

[0121] Terminal performance bonus A distance-based exponential decay gating mechanism is employed, activating a constraint reward on the velocity modulus and a reward on the alignment of the camera's optical axis with the ideal line-of-sight only when the UAV approaches the final target point. These rewards are defined as follows:

[0122]

[0123] in, Let be the unit vector of the camera's orientation. The first term is the ideal unit vector pointing from the drone to the observed target; the second term rewards camera attitude alignment through vector dot product. These are the weight and the attenuation coefficient, respectively.

[0124] Safety Obstacle Avoidance Penalty This is used to characterize the safety status between the drone and obstacles or neighboring drones; no penalty is imposed when the drone is outside the safe distance; a soft negative reward related to the degree of intrusion is imposed when the drone enters the safe warning distance; a huge negative reward for mission failure is imposed when a physical collision occurs or the drone intrudes into a strongly interfered no-fly zone, defined as:

[0125]

[0126] Decision continuity penalty The norm squared of the difference between the composite guidance instructions of the full state at adjacent time points is used to suppress drastic changes in the output actions of the upper-layer policy network, thereby improving the smoothness of the lower-layer control. It is defined as follows:

[0127]

[0128] Task completion reward For sparse rewards, a pre-defined task completion reward is given if and only if all agents simultaneously satisfy the terminal constraints of position, velocity, and attitude. Its definition is:

[0129]

[0130] 5-2) Execution of collaborative iterative training process: A centralized training and distributed execution paradigm is adopted. The system will use state transition tuples... Storage capacity is The experience replay pool. Critic network update: Randomly sample batch size from the replay pool. The data is used to calculate the loss function using a dual Q-network:

[0131]

[0132] Where the target value Discount factor The learning rate is set to .

[0133] Actor Network Update: Maximizing Expected Return and Entropy Regularization Based on Local Observations

[0134]

[0135] The learning rate is also set to Automatic entropy adjustment: dynamically adjusts the temperature coefficient through gradient descent. The target entropy is set to (i.e., negative action dimension).

[0136] Training results: such as Figure 4 The "average reward convergence curve for random obstacle scenarios" shown in this paper demonstrates that after approximately 10,000 training rounds, the average reward value of the method in this invention steadily converges to the high score range, proving the effectiveness of this collaborative training mechanism in complex dynamic environments.

[0137] Example 2: Based on Example 1, we conducted 3000 rounds of testing in both accessible and random obstacle scenarios, counting the number of successful multi-UAV collaborative inspection tasks. In the 3000-round test in the accessible scenario, the success rate of the proposed method (MASAC-MPC) and the MASAC algorithm was 99.97%, while the MADDPG algorithm achieved 93.97%. In the more challenging random obstacle scenario, the advantages of the proposed method were significantly enhanced: in the 3000-round test, the proposed method achieved 2799 successes, maintaining a task success rate of 93.30%; while the success rate of the comparative algorithms MASAC was 71.23%, and MADDPG's success rate was only 25.73%. The corresponding average reward convergence curve is shown below. Figure 4 As shown in the figure. Experimental data indicates that the method of this invention improves the task success rate by 22.07% compared to the traditional end-to-end MASAC algorithm and by 67.57% compared to the traditional MADDPG algorithm. Combined with... Figure 5 As can be seen from the flight trajectory diagrams in typical obstacle scenarios, traditional algorithms often fail to learn effective obstacle avoidance strategies, ultimately resulting in collisions near obstacles. In contrast, the method of this invention successfully selects a reasonable detour and smoothly reaches the target area. This reflects the synergistic effect of high-level guidance and low-level control within a layered framework: the upper-level strategy continuously updates guidance information based on environmental conditions, while the lower-level MPC performs rolling optimization and stable tracking while satisfying system constraints, significantly reducing the collision risk of multi-machine collaboration in complex unstructured environments.

[0138] Example 3: Based on Example 1, we conducted a 3000-round test to verify the control accuracy and trajectory stability. First, we statistically analyzed the control errors of each algorithm during the terminal hovering phase. The probability density distribution of the terminal control errors (position, velocity, gimbal angle) is as follows: Figure 6 As shown, in both unobstructed and random obstacle scenarios, the position, velocity, and gimbal angle error distributions of the proposed method (MASAC-MPC) all exhibit obvious zero-mean peak characteristics, demonstrating extremely high steady-state convergence accuracy. In contrast, the error distributions of the comparative algorithms MASAC and MADDPG are wider and significantly deviate from zero, making it difficult to achieve high-precision hovering.

[0139] Example 4: This example proposes a multi-UAV collaborative inspection system based on MASAC and MPC hierarchical control, applying the steps of the method described in this invention. The system includes:

[0140] The UAV kinematic model building module is configured to perform the following process: build a multi-UAV 3D inspection environment that includes static line facilities, random obstacles, and electromagnetic interference areas of power transmission lines, and establish a UAV kinematic model;

[0141] The UAV local observation information module is configured to perform the following process: at each decision moment, acquire local observation information of each UAV, the local observation information including at least its own motion state, relative target deviation, relative state of neighboring UAVs, and environmental obstacle perception information;

[0142] The output module is configured to perform the following process: inputting the local observation information into a deep reinforcement learning policy network and outputting a full-state composite guidance command containing the desired position offset, desired velocity, and desired gimbal attitude;

[0143] The model predictive controller module is configured to perform the following process: constructing a model predictive controller that introduces slack variables, using the full-state composite guidance command as a tracking reference, solving for the optimal control sequence that satisfies the dynamic constraints, and applying the control sequence to the UAV;

[0144] The multi-UAV collaborative inspection control module is configured to execute the following process: based on a centralized training and distributed execution mechanism, the deep reinforcement learning policy network is iteratively updated using a multi-objective reward function to achieve multi-UAV collaborative inspection control.

[0145] Example 5: This example proposes an electronic system, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method steps of the present invention.

[0146] Example 6: This example proposes a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the steps of the method described in this invention, which will not be repeated here.

[0147] Example 7: This example proposes a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps of the method described in this invention, which will not be repeated here.

[0148] It should be noted that the processing flow of embodiments 4-7 corresponds to the specific steps of the method provided in embodiment 1 of the present invention, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in this embodiment can be found in the method provided in embodiment 1 of the present invention.

[0149] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0150] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-UAV collaborative inspection method based on MASAC and MPC hierarchical control, characterized in that, Includes the following steps: S1. Construct a multi-UAV 3D inspection environment that includes static line facilities, random obstacles, and electromagnetic interference areas of transmission lines, and establish a UAV kinematic model; S2. At each decision moment, acquire local observation information of each UAV, the local observation information including at least its own motion state, relative target deviation, relative state of neighboring UAVs, and environmental obstacle perception information. S3. Input the local observation information into the deep reinforcement learning policy network and output a full-state composite guidance command containing the desired position offset, desired velocity and desired gimbal attitude. S4. Construct a model predictive controller that introduces slack variables, use the full-state composite guidance command as a tracking reference, solve for the optimal control sequence that satisfies the dynamic constraints, and apply the control sequence to the UAV; S5. Based on a centralized training and distributed execution mechanism, the deep reinforcement learning policy network is iteratively updated using a multi-objective reward function to achieve multi-UAV collaborative inspection control.

2. The multi-UAV collaborative inspection method based on MASAC and MPC hierarchical control according to claim 1, characterized in that, S1 includes: S21. Establish a hierarchical spatial risk model for the space surrounding power transmission cables; S22. Define the internal space area adjacent to the cable as a strong interference no-fly zone and set it as a hard geometric constraint boundary with repulsion properties; S23. Define the outer space area surrounding the strong interference no-fly zone as the weak interference risk zone. Construct a risk potential field within this area and introduce it into the environmental feedback as a soft penalty to guide the drone to autonomously maintain a safe distance.

3. The multi-UAV collaborative inspection method based on MASAC and MPC hierarchical control according to claim 1, characterized in that, In step S3, the specific construction method of the full-state composite boot instruction is as follows: S31, The action vector output by the upper-layer policy network is defined as follows: ; S32, among which Includes the desired position offset relative to the current position. Used to generate short-term navigation guide points; contains the three-dimensional velocity vector expected to reach the guide point. ; and the expected gimbal yaw and pitch angles when reaching the guide point. ; S33. The full-state composite guidance command explicitly transmits the desired velocity vector and gimbal angular velocity intention when reaching the guidance point, as a feedforward reference for the lower-level controller, eliminating the implicit constraint of zero velocity caused by the discrete position jump of the upper layer, thereby suppressing the sawtooth oscillation of the trajectory and ensuring the dynamic consistency of inter-layer decisions.

4. The multi-UAV collaborative inspection method based on MASAC and MPC hierarchical control according to claim 1, characterized in that, In step S4, constructing the fully soft-constrained model predictive controller specifically includes: S41. Abandon the hard constraints of forced state matching and establish soft constraint equations that introduce slack variables for tracking targets with position, velocity and gimbal angle. S42, The soft constraint is: ; ; ; in , , The vectors of relaxed variables for position, velocity, and attitude represent the minimum necessary deviation between the system's predicted state and the upper-level reference target. S43. This mechanism ensures that when the guidance commands output by the upper layer exceed the physical feasible domain of the UAV, the lower layer controller can still find the feasible solution with the least degree of violation, thus avoiding system crash due to the unsolvable optimization problem.

5. A multi-UAV collaborative inspection method based on MASAC and MPC hierarchical control according to claim 4, characterized in that, In step S4, the objective function of the model prediction controller is constructed as follows: S51. Constructing the objective function ; in, and These are acceleration control input and gimbal angular velocity control input, respectively. and The smoothing weight matrix is ​​used to control the input; , and Here is the penalty weight matrix for the slack variables; S52. The weight matrix W of the slack variables and the control input weight matrix R both have hierarchical priority, where the eigenvalues ​​of W are greater than the eigenvalues ​​of R, in order to establish that error-free state tracking is prioritized within the limits of physical constraints. In cases of conflicting or infeasible physical constraints, the feasibility of solving the control problem and the stability of the system are prioritized by tolerating the minimized state deviation.

6. The multi-UAV collaborative inspection method based on MASAC and MPC hierarchical control according to claim 1, characterized in that, In step S5, the multi-objective reward function guides the system to smoothly transition from rapid maneuvering to precise hovering. The multi-objective reward function consists of progress rewards, terminal performance rewards, safety obstacle avoidance rewards, decision continuity rewards, and task completion rewards, and its weighted sum is expressed as: ; in, , , , and These represent progress rewards, terminal performance rewards, safety and obstacle avoidance rewards, decision continuity rewards, and task completion rewards, respectively. , , and These are the corresponding weighting coefficients; S61, the aforementioned progress reward This includes entity progress rewards based on the actual physical location of the drone and virtual progress rewards based on the location of upper-layer guide points, used to calibrate the consistency between the upper-layer strategy and the global mission direction. Its definition is: ; in, and These represent the reduction in distance from the actual location of the drone and the guide point to the target point, respectively. and These are weighting coefficients; S62, the terminal performance reward A distance-based exponential decay gating mechanism is employed, activating a constraint reward on the velocity modulus and a reward on the alignment of the camera's optical axis with the ideal line-of-sight only when the UAV approaches the final target point. These rewards are defined as follows: ; in, Let be the unit vector of the camera's orientation. The first term is the ideal unit vector pointing from the drone to the observed target; the second term rewards camera attitude alignment through vector dot product. These are the weights and attenuation coefficients, respectively. S63, the aforementioned safety obstacle avoidance penalty This is used to characterize the safety status between the drone and obstacles or neighboring drones; no penalty is imposed when the drone is outside the safe distance; a soft negative reward related to the degree of intrusion is imposed when the drone enters the safe warning distance; a huge negative reward for mission failure is imposed when a physical collision occurs or the drone intrudes into a strongly interfered no-fly zone, defined as: ; S64, the decision continuity penalty The norm squared of the difference between the composite guidance instructions of the full state at adjacent time points is used to suppress drastic changes in the output actions of the upper-layer policy network, thereby improving the smoothness of the lower-layer control. It is defined as follows: ; S65, Rewards for Completing the Task For sparse rewards, a pre-defined task completion reward is given if and only if all agents simultaneously satisfy the terminal constraints of position, velocity, and attitude. Its definition is: 。 7. A multi-UAV collaborative inspection system based on MASAC and MPC hierarchical control, characterized in that, The system comprising the steps of applying the method according to any one of claims 1 to 6, wherein the system includes: The UAV kinematic model building module is configured to perform the following process: build a multi-UAV 3D inspection environment that includes static line facilities, random obstacles, and electromagnetic interference areas of power transmission lines, and establish a UAV kinematic model; The UAV local observation information module is configured to perform the following process: at each decision moment, acquire local observation information of each UAV, the local observation information including at least its own motion state, relative target deviation, relative state of neighboring UAVs, and environmental obstacle perception information; The output module is configured to perform the following process: inputting the local observation information into a deep reinforcement learning policy network and outputting a full-state composite guidance command containing the desired position offset, desired velocity, and desired gimbal attitude; The model predictive controller module is configured to perform the following process: constructing a model predictive controller that introduces slack variables, using the full-state composite guidance command as a tracking reference, solving for the optimal control sequence that satisfies the dynamic constraints, and applying the control sequence to the UAV; The multi-UAV collaborative inspection control module is configured to execute the following process: based on a centralized training and distributed execution mechanism, the deep reinforcement learning policy network is iteratively updated using a multi-objective reward function to achieve multi-UAV collaborative inspection control.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed, it implements the steps of the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is configured to implement the steps of the method according to any one of claims 1 to 6 when invoked by a processor.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.