Unmanned aerial vehicle group cooperative hoisting and carrying method based on AI vision algorithm
By using dual-source image acquisition based on AI vision algorithms and collaborative decision-making neural network optimization, combined with distributed model predictive control, the problem of coupled iterative optimization of path and force distribution in UAV swarm collaborative hoisting was solved, realizing spatiotemporal synchronous control of the UAV swarm hoisting process and avoiding trajectory fluctuations and load distribution imbalance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI HUADIAN ENGINEERING CONSULTATING & DESIGN CO LTD
- Filing Date
- 2026-04-23
- Publication Date
- 2026-07-31
AI Technical Summary
In existing UAV swarm collaborative hoisting technology, single-source visual acquisition cannot integrate environmental spatial features and target object morphological features, traditional path planning algorithms cannot achieve coupled iterative optimization of motion path and force distribution, centralized control is difficult to adapt to the dynamic changing scenarios of the swarm, and single-sequence model predictive control cannot simultaneously cover UAV status and load force parameters, which leads to trajectory fluctuations and load distribution imbalances during the hoisting process.
A method based on AI vision algorithms is adopted to construct a multi-dimensional perception data structure through dual-source image acquisition. The path and force distribution are optimized by using a collaborative decision neural network. The multi-moment state trajectory and load force sequence of the UAV are derived by combining a distributed model predictive control algorithm. Dynamic adjustments are made through trajectory smoothness and load balance scores, and finally, spatiotemporally synchronized cluster collaborative control commands are generated.
It achieves complete coverage of environmental spatial features and target object morphological features, coupled iterative calculation of path and force distribution, eliminates parameter disconnection problem, dynamically adjusts and adapts to changes in real-time sensing data, ensures accurate matching of UAV attitude and tension control quantity, and avoids hoisting trajectory fluctuations and load distribution imbalance.
Smart Images

Figure CN122086098B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent hoisting control technology for unmanned aerial vehicles (UAVs), specifically a collaborative hoisting and transportation method for UAV swarms based on AI vision algorithms. Background Technology
[0002] Conventional UAV swarm collaborative hoisting and transportation mostly use single-source environmental visual acquisition or discrete point cloud data to complete environmental perception, rely on traditional path planning algorithms to formulate hoisting motion schemes, complete UAV swarm scheduling through centralized control logic, the model predictive control link only derives a single motion state sequence, the planning scheme adjustment is completed based on a single evaluation index, and the force distribution and path planning adopt an independent solution mode.
[0003] Single-source visual acquisition cannot integrate environmental spatial features with target object morphological features, and cannot form a multi-dimensional unified perception data structure. Traditional planning algorithms cannot achieve coupled iterative optimization of motion path and force distribution, and centralized control is difficult to adapt to dynamically changing cluster scenarios. Single-sequence model predictive control cannot simultaneously cover UAV status and load force parameters, single index adjustment cannot take into account trajectory operation characteristics and load distribution characteristics, and cluster control commands are difficult to achieve spatiotemporal synchronization matching of pose and tension. The hoisting process is prone to trajectory fluctuations and load distribution imbalances.
[0004] It is necessary to build a multi-dimensional perception data structure based on dual-source original images, optimize the path and force distribution through neural network coupling and generate a global planning map, synchronously deduce the UAV's multi-moment state trajectory and load force sequence, and complete the dynamic adjustment of the planning scheme through dual scoring dimensions, so as to finally form a spatiotemporal synchronous cluster control command adapted to the hoisting scenario. Summary of the Invention
[0005] This invention aims to solve at least one of the technical problems existing in the prior art;
[0006] Therefore, this invention proposes a collaborative lifting and transportation method for UAV swarms based on AI vision algorithms, comprising:
[0007] Based on the original environmental images and the original morphological images of the target object collected by the hoisting drone, a dynamic environmental perception tensor is generated.
[0008] The dynamic environment perception tensor is input into the trained collaborative decision neural network, which generates a global collaborative hoisting planning diagram by iteratively optimizing the motion path and force distribution scheme of each UAV.
[0009] Based on the global collaborative hoisting planning diagram, a distributed model predictive control algorithm is used to derive the predicted state trajectory and predicted load force sequence of each individual in the hoisting drone swarm at multiple future moments;
[0010] For each predicted state trajectory and the predicted load force sequence, calculate its trajectory smoothness score and load balance score, and prioritize and dynamically adjust the planning schemes based on the trajectory smoothness score and the load balance score;
[0011] The predicted state trajectory and the predicted load force sequence after dynamic adjustment are integrated to form a spatiotemporally synchronized cluster collaborative control instruction set. Each instruction in the instruction set corresponds to the pose and tension control amount of a hoisting UAV at a specific moment.
[0012] Furthermore, the generation of a dynamic environment perception tensor based on the original environmental images and the original morphological images of the transported target object acquired by the hoisting drone includes:
[0013] The initial visual data stream is formed by using a hoisting drone equipped with a visual sensing device to collect original environmental images of the collaborative operation area and original morphological images of the transported target.
[0014] Perform multi-scale spatiotemporal fusion and feature extraction operations on the initial visual data stream to separate the background environment feature stream, the carrier target feature stream, and the UAV body feature stream;
[0015] Based on the separated background environment feature stream, the target object feature stream, and the UAV body feature stream, a dynamic environment perception tensor for collaborative hoisting is constructed.
[0016] Furthermore, multi-scale spatiotemporal fusion and feature extraction operations are performed on the initial visual data stream to separate the background environment feature stream, the target object feature stream, and the UAV body feature stream, including:
[0017] The original environmental image is subjected to pyramid downsampling to construct a multi-scale environmental image sequence. Static obstacle contours, dynamic obstacle motion vectors, and wind disturbance features are identified and extracted from the multi-scale environmental image sequence and fused into the background environmental feature stream.
[0018] The original morphological image is reconstructed into a three-dimensional point cloud and its skeleton is extracted. The centroid position, attitude angle, mass distribution and moment of inertia of the target object are calculated to form the feature flow of the target object.
[0019] By using visual odometry and feature point matching technology, the pose, velocity, acceleration and sling angle of each hoisting drone are calculated from continuous image frames to form the feature flow of the drone body;
[0020] A spatiotemporal alignment mechanism is established to synchronize and calibrate the background environment feature stream, the carrier target feature stream, and the UAV body feature stream based on a unified world coordinate system and high-precision timestamps.
[0021] An attention-guided feature separation network is used to decouple the synchronously calibrated hybrid feature streams to obtain spatially and semantically independent background environment feature streams, carrier target object feature streams, and UAV body feature streams.
[0022] Furthermore, the dynamic environment perception tensor is input into a trained collaborative decision-making neural network. This network iteratively optimizes the motion paths and force distribution schemes of each UAV to generate a global collaborative hoisting plan, including:
[0023] Initialize a fully connected interactive graph with each hoisting drone as a node, and encode the dynamic environment perception tensor as the initial features of each node in the graph;
[0024] In the collaborative decision-making neural network, the environmental and state information of each node and its neighboring nodes are aggregated through a graph attention layer to update node features;
[0025] The updated node features are input into the policy network, which outputs a series of waypoint sequences and expected lifting force sequences for each UAV in the planning time domain.
[0026] The long-term value of the combined action consisting of the waypoint sequence and the desired lifting pull sequence is evaluated through a value network, and constraints for avoiding collisions, maintaining formation, and minimizing energy consumption are introduced.
[0027] A gradient-based strategy optimization method is adopted to iteratively update the parameters of the strategy network and the value network until the joint action satisfies all constraints and the long-term value converges. Finally, the optimized set of node actions is mapped to the global collaborative hoisting planning diagram that includes spatial path, time arrangement and force allocation.
[0028] Furthermore, based on the aforementioned global collaborative hoisting planning diagram, a distributed model predictive control algorithm is used to derive the predicted state trajectory and predicted load force sequence of each individual in the hoisting drone swarm at multiple future moments, including:
[0029] The waypoint sequence and the expected lifting tension sequence assigned to a specific lifting UAV in the global collaborative lifting planning map are used as the reference trajectory and reference force sequence for the local rolling time domain optimization of the specific lifting UAV.
[0030] Based on the nonlinear dynamics model of the specific hoisting UAV, a local optimization problem is constructed with the goal of tracking the reference trajectory and the reference force sequence and constrained by the capabilities of its own actuators.
[0031] In each control cycle, the local optimization problem is solved to obtain the predicted state trajectory and the predicted load force sequence of the specific hoisting UAV in the future finite time domain;
[0032] During the solution process, the predicted state information is exchanged with other hoisting drones in the communication neighborhood. The predicted trajectory of the neighboring drones is used as a constraint on its own optimization problem in a coupled manner to ensure the coordination and collision-free nature of the cluster prediction.
[0033] Further, for each of the predicted state trajectories and the predicted load force sequence, a trajectory smoothness score and a load balance score are calculated, including:
[0034] For a predicted state trajectory, calculate its rate of change on each derivative, including the rate of change of position, the rate of change of velocity, and the rate of change of acceleration;
[0035] By combining the weighted sum of the rates of change of each derivative and evaluating their continuity, the trajectory smoothness score, which reflects the smoothness of the motion, is obtained.
[0036] For the predicted load force sequence, calculate the deviation between the predicted load force and the expected load force of each hoisting drone at the same time in the predicted state trajectory, as well as the variance of the predicted load force among each hoisting drone.
[0037] By combining the weighted evaluation of the deviation and the square value, the load balance score, which reflects the rationality of the stress on the cluster, is obtained.
[0038] Furthermore, the integration of the dynamically adjusted predicted state trajectory and the predicted load force sequence to form a spatiotemporally synchronized cluster collaborative control instruction set includes:
[0039] Based on the trajectory smoothness score and the load balance score, the predicted state trajectory and the predicted load force sequence of all hoisting drones are adjusted for consistency alignment.
[0040] In the consistency alignment adjustment, it is ensured that the spatial positions corresponding to all trajectories at the same timestamp meet the preset hoisting formation geometric constraints, and all load force sequences meet the preset system force closure constraints.
[0041] The adjusted, time- and space-aligned states and force sequences of each UAV are discretized into high-frequency control commands. Each command contains the position, attitude, and speed that the UAV should achieve at a specific moment, as well as the output cable tension vector.
[0042] All discrete control commands of the hoisting drones are globally sorted and packaged in chronological order to form the spatiotemporally synchronized cluster collaborative control command set.
[0043] Furthermore, it also includes:
[0044] The anti-interference robustness verification of the spatiotemporally synchronized cluster cooperative control instruction set is performed, and the complex instruction set is decomposed into multiple independently executable and fault-tolerant control sub-task clusters.
[0045] For each of the control subtask clusters, a state inversion calculation based on visual feedback is performed to quantify the tracking error and compensation amount of each control command to the actual hoisting state.
[0046] Based on the tracking error and compensation amount of each control command, and combined with the preset hoisting stability criterion, a real-time updated UAV swarm collaborative hoisting and transportation control signal is generated.
[0047] The process of performing anti-interference robustness verification on the spatiotemporal synchronized cluster collaborative control instruction set involves decomposing the complex instruction set into multiple independently executable and fault-tolerant control subtask clusters, including:
[0048] A simulated environmental disturbance model is injected into the cluster collaborative control command set. The environmental disturbance model includes wind field changes, sensor noise, and actuator errors.
[0049] Perform forward simulation to verify whether the cluster system can still stably track instructions and complete the hoisting task after the disturbance is injected, and identify the instruction segments that fail or conflict under the disturbance.
[0050] Based on the spatiotemporal dependencies and functional correlations between instructions, the cluster collaborative control instruction set is divided into multiple relatively independent instruction subsets. Each instruction subset and its corresponding UAV subset constitute a control subtask cluster.
[0051] A backup control law and intra-cluster communication protocol are designed for each of the control sub-task clusters, so that when some UAVs in a certain control sub-task cluster fail, the control sub-task cluster can autonomously adjust and maintain basic hoisting functions.
[0052] Furthermore, for each of the control subtask clusters, a state inversion calculation based on visual feedback is performed to quantify the tracking error and compensation amount of each control command to the actual hoisting state, including:
[0053] During the execution of one of the control subtask clusters, real-time image streams collected by the UAV visual sensing devices within the cluster are continuously acquired;
[0054] The actual pose and velocity of the target object, as well as the actual pose and direction of the sling tension of each UAV, are calculated in real time from the real-time image stream.
[0055] The calculated actual state is compared with the expected state of the corresponding instruction in the control subtask cluster to calculate the position tracking error, attitude tracking error and force tracking error.
[0056] The tracking error is input to the state inversion observer, which estimates the system state variables that are not directly measured based on the system dynamics model. It then combines the direct error with the estimated state to calculate the feedforward compensation and feedback compensation for environmental disturbances and model uncertainties, which together constitute the compensation amount.
[0057] Furthermore, based on the tracking error and compensation amount of each control command, and combined with a preset hoisting stability criterion, a real-time updated UAV swarm collaborative hoisting and transportation control signal is generated, including:
[0058] The tracking error is fused with the compensation amount to generate a preliminary hoisting control correction amount;
[0059] The preliminary hoisting control correction amount is superimposed on the control instructions in the original control subtask cluster to form the corrected desired control amount.
[0060] The modified desired control quantity is input into the hoisting stability criterion for verification. The hoisting stability criterion includes the target moment balance condition, the non-negativity condition of the sling tension, and the UAV thrust limiting condition.
[0061] If the modified expected control quantity passes the hoisting stability criterion test, it is converted into a low-level control signal that can directly drive the UAV flight control and actuator, namely the real-time updated UAV swarm collaborative hoisting and transportation control signal.
[0062] If the test fails, the corrected expected control quantity will be limited or adjusted based on the criterion feedback until the test passes before generating the real-time updated UAV swarm collaborative hoisting and transportation control signal.
[0063] Compared with the prior art, the beneficial effects of the present invention are:
[0064] A dynamic environmental perception tensor is generated by collecting original environmental images and original morphological images of the target object from hoisting drones. This tensor is then input into a trained collaborative decision neural network to iteratively optimize the motion paths and force distribution schemes of each drone and generate a global collaborative hoisting planning map. Dual-source image acquisition can fully cover environmental spatial features and target object morphological features. The dynamic environmental perception tensor can construct a multi-dimensional unified perception data carrier. The collaborative decision neural network can realize coupled iterative calculation of path and force distribution. The global collaborative hoisting planning map can integrate the planning parameters of the entire hoisting domain of the cluster, eliminating the parameter disconnect problem caused by independent calculation of path and force distribution.
[0065] Based on the global collaborative hoisting planning diagram, a distributed model predictive control algorithm is used to derive the predicted state trajectory and predicted load force sequence of individual hoisting drones at multiple future moments. The trajectory smoothness score and load balance score are calculated for the two types of sequences. Based on the dual scores, the planning scheme is prioritized and dynamically adjusted. The adjusted parameters are integrated to form a spatiotemporally synchronized cluster collaborative control command set. The command corresponds to the drone's pose and tension control quantity at a specific moment. The distributed model predictive control can synchronously solve the dual sequence parameters, covering the drone's state and force changes at multiple future moments. The dual scores can quantify the trajectory operation and load distribution characteristics separately. The dynamic adjustment can adapt to changes in real-time sensing data. The command set can achieve accurate matching of the pose and tension control quantity of a single drone, avoiding the phenomenon of cluster hoisting trajectory fluctuations and load distribution imbalances. Attached Figure Description
[0066] Figure 1 This is a flowchart illustrating the steps of the UAV swarm collaborative hoisting and transportation method based on AI vision algorithms described in this invention.
[0067] Figure 2 A flowchart for multi-scale spatiotemporal fusion and feature extraction;
[0068] Figure 3 A comparison chart of the normalized change rate of trajectory state variables for collaborative hoisting by UAV swarms;
[0069] Figure 4 A single-dimensional position trajectory diagram of a drone swarm collaborative hoisting system;
[0070] Figure 5 A diagram illustrating the visual feedback state inversion calculation and analysis for collaborative hoisting by a drone swarm. Detailed Implementation
[0071] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0072] See Figure 1This invention provides a collaborative lifting and transportation method for a swarm of unmanned aerial vehicles (UAVs) based on AI vision algorithms. The overall implementation scheme is as follows: Using visual sensing devices mounted on the lifting UAVs, original environmental images and original morphological images of the target object are acquired. A series of processing steps generate a dynamic environmental perception tensor, which serves as the basis for subsequent collaborative decision-making. The generated dynamic environmental perception tensor is input into a pre-trained collaborative decision-making neural network. This network, through internal iterative optimization, plans the motion path and force distribution scheme for each UAV in the swarm, and comprehensively generates a global collaborative lifting planning map. Based on this global planning map, a distributed model predictive control algorithm is used to independently derive the predicted state trajectory and predicted load force sequence for each individual UAV in the swarm at multiple future time points, while considering neighborhood coupling constraints. Subsequently, a trajectory smoothness score and a load balance score are calculated for each derived predicted state trajectory and predicted load force sequence. Based on these scores, the planning schemes are prioritized and dynamically adjusted as necessary. Finally, all dynamically adjusted predicted state trajectories and predicted load force sequences are integrated and spatiotemporally synchronized to form a set of cluster collaborative control instructions that can be directly issued. Each instruction in this set clearly corresponds to the pose that a hoisting drone needs to achieve at a specific moment and the amount of tension control it needs to output.
[0073] In one embodiment of the present invention, a visual sensing device is mounted on a hoisting drone to collect original environmental images of the collaborative operation area and original morphological images of the transported target object, forming an initial visual data stream. See also... Figure 2 The initial visual data stream undergoes multi-scale spatiotemporal fusion and feature extraction operations. Specifically, this includes: pyramid downsampling of the original environmental images to construct a multi-scale environmental image sequence; identifying and extracting static obstacle contours, dynamic obstacle motion vectors, and wind disturbance features from this sequence; and fusing this information into a background environmental feature stream. Three-dimensional point cloud reconstruction and skeleton extraction are performed on the original morphological images to calculate the centroid position, attitude angle, mass distribution, and moment of inertia of the target object, forming a target object feature stream. Visual odometry and feature point matching techniques are used to calculate the pose, velocity, acceleration, and sling angle of each hoisting UAV from consecutive image frames, forming a UAV body feature stream. A spatiotemporal alignment mechanism is established to synchronously calibrate the background environmental feature stream, target object feature stream, and UAV body feature stream based on a unified world coordinate system and high-precision timestamps. An attention-guided feature separation network is used to decouple the synchronously calibrated mixed feature streams, thereby spatially and semantically separating the background environmental feature stream, target object feature stream, and UAV body feature stream. Based on these three types of feature flows, a dynamic environment perception tensor is constructed for subsequent collaborative hoisting decisions.
[0074] In practice, the visual sensing equipment of the hoisting drone acquires original environmental images including terrain, buildings, and trees, as well as original morphological images of the container or structural component to be hoisted. These images form the initial visual data stream at a rate of thirty frames per second. The processing system receives the initial visual data stream, performs pyramid downsampling on the original environmental images, and generates a multi-scale environmental image sequence containing four scales with a fixed scaling factor. From the multi-scale environmental image sequence, the contour boundaries of static obstacles are identified in the coarse-scale images, the pixel displacement vectors of dynamic obstacles such as birds are calculated in the fine-scale images, and the distortion of image texture is analyzed between consecutive frames to estimate wind disturbance features. The extracted contour, vector, and disturbance features are fused and encoded into a background environmental feature stream. Three-dimensional point cloud reconstruction is performed on the original morphological image. A mesh model is generated through surface triangulation. The skeleton representing the main structure is extracted from the mesh model. Based on the mesh model and skeleton, the three-dimensional position of the centroid of the target object in the camera coordinate system, the attitude angle of the target object relative to the direction of gravity, the mass distribution obtained by volume integral and preset density, and the moment of inertia calculated by inertial tensor are calculated. These physical quantities constitute the feature flow of the target object. Visual odometry technology is used to match and triangulate natural or artificial feature points in continuous image frames to solve the position and attitude of each hoisting UAV in the world coordinate system. The velocity and acceleration are obtained through inter-frame difference, and the sling angle is calculated by analyzing the image coordinates of the sling connection point and the UAV body. The pose, velocity, acceleration, and sling angle data constitute the feature flow of the UAV body.
[0075] In some embodiments, the system establishes a spatiotemporal alignment mechanism to uniformly transform the background environment feature stream, the target object feature stream, and the UAV body feature stream from different UAVs and processing threads into a geodetic coordinate system with the center of the operating area as the origin. Each set of feature data is then stamped with a microsecond-level precision timestamp from a globally synchronized clock, achieving synchronous calibration of the data in both time and space dimensions. A feature separation network guided by a multi-head attention mechanism is employed. This network takes the synchronously calibrated mixed feature stream as input and learns to focus on background, target, and body-related feature channels through multiple parallel attention heads. The outputs of the attention heads are then decoupled into spatially and semantically independent background environment feature stream, target object feature stream, and UAV body feature stream through gating and feature projection.
[0076] Optionally, the process of constructing the dynamic environment perception tensor involves tensor concatenation and dimensionality upscaling of the separated background environment feature stream, target object feature stream, and UAV body feature stream. The background environment feature stream provides obstacle distribution and wind field information, the target object feature stream provides the mass and inertial properties of the payload, and the UAV body feature stream provides the real-time state of the actuator. The concatenated high-dimensional tensor is the dynamic environment perception tensor. One dimension of the dynamic environment perception tensor corresponds to the time series, one dimension corresponds to the feature channel, and the remaining dimensions correspond to spatial or UAV identification information. The construction process of the dynamic environment perception tensor can be represented by the following formula:
[0077]
[0078] in: Represents the dynamic environment-aware tensor. This represents the feature concatenation and fusion function. Represents the background environment feature flow. Indicates the characteristic flow of the transport target object, Represents the feature flow of the drone itself. This represents the set of learnable parameters for the fusion function.
[0079] It is understandable that during the feature flow synchronization calibration step, if the visual odometry of an individual UAV experiences a brief failure, the system uses predicted values based on historical motion models for interpolation and cross-validates and corrects the data from neighboring UAVs observing the same landmark to maintain the robustness of the spatiotemporal alignment of the feature flow. The feature separation network is trained using a large amount of labeled simulation and real hoisting scenario data. The training objective is to minimize the differences between the reconstructed feature flows and the real labeled features, and to maximize the lower bound of mutual information between different categories of feature flows, thereby ensuring the sufficiency of feature separation. A fully trained feature separation network can stably separate the background environment feature flow, the target object feature flow, and the UAV body feature flow from mixed inputs.
[0080] In one embodiment of the invention, a fully connected interactive graph is initialized, with each hoisting drone in the cluster as a node, and the dynamic environment-aware tensor is encoded as the initial features of each node in the graph. In the collaborative decision neural network, the environment and state information of each node and its neighboring nodes are aggregated through a graph attention layer to update the feature representation of each node. The updated node features are input to a policy network, which outputs a series of waypoint sequences and expected hoisting pull sequences for each drone in the planning time domain. The long-term value of the joint action consisting of all drone waypoint sequences and expected hoisting pull sequences is evaluated through a value network, and constraints such as collision avoidance, maintaining a preset formation, and minimizing energy consumption are introduced during the evaluation process. A gradient-based policy optimization method is used to iteratively update the parameters of the policy network and the value network until the generated joint action satisfies all constraints and its long-term value evaluation results converge. Finally, the optimized set of node actions is mapped into a global collaborative hoisting planning graph containing details of spatial paths, time arrangements, and force allocation.
[0081] In implementation, the system initializes a fully connected interactive graph, where each node corresponds to a hoisting drone, and the total number of nodes equals the drone swarm size. The environmental and state features associated with each drone in the dynamic environment perception tensor are encoded, and the encoded feature vectors serve as the initial features for the corresponding graph nodes. In the graph attention layer of the collaborative decision neural network, each node aggregates information by calculating the attention weights between its own features and those of its neighboring nodes. These attention weights reflect the importance of information from different neighboring nodes. The aggregation operation updates the feature representation of each node, ensuring that the node features contain local swarm environment and state information. The updated node features are input to the policy network, which is a multilayer perceptron. The policy network outputs a sequence of three-dimensional waypoint coordinates for the corresponding drone within the planning time domain and a time-varying sequence of the expected hoisting force scalar. The value network receives all drone node features and the joint action sequence output by the policy network. The value network evaluates the long-term cumulative reward value of this joint action from the current state, introducing constraints during the evaluation process. These constraints include a collision avoidance term that penalizes drones for being too close together, a formation-maintaining term that penalizes formation deviation from the reference geometry, and an energy-minimizing term that penalizes total energy consumption.
[0082] In some embodiments, the gradient-based policy optimization method employs a proximal policy optimization algorithm. This algorithm samples multiple sets of joint actions generated by a policy network and executes them in a simulation environment to evaluate their long-term value and constraint violations. The proximal policy optimization algorithm calculates the loss function of the policy network and the loss function of the value network. The loss function includes the long-term value maximization objective and the constraint satisfaction conditions. The loss function of the policy network encourages the generation of actions with high long-term value while limiting the difference between new and old policies. The loss function of the value network is the mean squared error between the predicted value and the actual reward. The weight parameters of the policy network and the value network are iteratively updated through backpropagation and gradient descent. The iterative update process continues until the generated joint action sequence causes the constraints of collision avoidance, maintaining formation, and minimizing energy consumption to all fall below a set threshold, and the long-term value evaluation fluctuates less than the convergence tolerance in consecutive iterations. At this point, the optimization process is considered converged.
[0083] Optionally, the optimized set of node actions is mapped to a global collaborative lifting planning graph. The mapping process connects the waypoint sequences output by each node in three-dimensional space to form path curves, and temporally correlates the desired lifting tension sequence with points on the path curves. The global collaborative lifting planning graph includes a set of path curves, a set of time-tension relationships, and spatiotemporal relationships between UAVs. The global collaborative lifting planning graph uses a hierarchical graph structure for storage and representation. The top-level graph represents the task allocation and collaborative relationships of UAVs as nodes. At the bottom level, each node is associated with a subgraph, which details the UAV's pathpoint sequence and tension sequence. The pathpoint sequence includes position, velocity, and acceleration information, while the tension sequence includes magnitude and direction information. Pathpoints and tension are strictly aligned on the time axis. The generation process of the global collaborative lifting planning graph can be represented by the following value function evaluation formula:
[0084]
[0085] in: Representation Strategy The long-term value of the underlying assets Indicates based on strategy Generate trajectory Expectations Indicates the length of the planning time domain. Indicates the discount factor. Indicates time Status-action instant reward, , , Representing trajectories Collision avoidance constraint violation, formation maintenance constraint violation, and energy consumption constraint violation. , , These are the corresponding constraint weight coefficients.
[0086] It is understandable that the attention weight calculation of the graph attention layer adopts an additive attention mechanism. The query vector comes from the node's own features, the key vector comes from the features of neighboring nodes, and the attention score is calculated through a single-layer feedforward network, thereby dynamically determining the contribution of each neighbor during information aggregation. The training data for the policy network and value network comes from offline simulation and online interaction. The simulation environment simulates various obstacle layouts, wind conditions, and parameters of the transport target, while the online interaction is conducted in a controlled real test scenario. By continuously interacting with the environment, state, action, reward, and constraint violation data are collected to expand the training dataset and improve the generalization ability of the collaborative decision-making neural network. During the iterative update process, if the long-term value fails to converge after multiple iterations or constraint violations persist, the weight coefficients of the constraint terms are adjusted or the network parameters are reinitialized, and the optimization process restarts until a globally collaborative hoisting planning diagram that meets the requirements is generated.
[0087] In one embodiment of the present invention, the waypoint sequence and the desired lifting force sequence assigned to a specific lifting UAV in the global collaborative lifting planning map are used as the reference trajectory and reference force sequence for the specific lifting UAV to perform local rolling time-domain optimization. Based on the nonlinear dynamic model of the specific lifting UAV, a local optimization problem is constructed with the primary objective of tracking the aforementioned reference trajectory and reference force sequence and the capability limitations of the UAV's own actuators as constraints. In each control cycle, this local optimization problem is solved to obtain the predicted state trajectory and predicted load force sequence of the specific lifting UAV in the future finite time domain. During the solution process, the UAV exchanges predicted state information with other lifting UAVs in its communication neighborhood and incorporates the predicted trajectories of neighboring UAVs into its own optimization problem in the form of coupling constraints, thereby ensuring the predictive behavior of the entire cluster is cooperative and avoiding collisions. A trajectory smoothness score is calculated for each predicted state trajectory by calculating the rate of change of the derivatives of position, velocity, and acceleration of the trajectory, and combining the weighted sum of these rates of change with a continuity evaluation. To calculate the load balance score of the predicted load force sequence, the method is to calculate the deviation between the predicted load force and the expected load force of each hoisting drone at the same time of the predicted state trajectory, as well as the variance of the predicted load force among the drones, and then perform a weighted evaluation of the deviation and variance.
[0088] In practical implementation, the global collaborative lifting planning map includes the waypoint sequence and expected lifting force sequence assigned to each lifting drone. For a specific lifting drone in the cluster, the planning system extracts the waypoint sequence and expected lifting force sequence assigned to that specific lifting drone from the global collaborative lifting planning map, and uses these two sequences as the reference trajectory and reference force sequence for that specific lifting drone when performing local rolling time-domain optimization. Based on the nonlinear dynamic model of the specific lifting drone, which involves mass, inertia, aerodynamic parameters, and actuator dynamics, a local optimization problem is constructed with tracking the reference trajectory and reference force sequence as the main optimization objective. The constraints of the local optimization problem include, but are not limited to, the maximum thrust limit, maximum tilt angle limit, velocity and acceleration boundaries of the specific lifting drone. In each control cycle, this local optimization problem is solved, and the state and control input sequence of the specific lifting drone in the future finite time domain are solved through numerical optimization algorithms to obtain the predicted state trajectory and predicted load force sequence of the specific lifting drone. In the process of solving local optimization problems, specific hoisting drones exchange their predicted state information, especially predicted position information, with other neighboring hoisting drones through communication links. When constructing their own optimization problems, specific hoisting drones incorporate the predicted trajectories of neighboring drones as obstacle avoidance constraints in the form of mathematical inequality constraints, thereby ensuring the coordination and collision-free nature of cluster predictions. This coupling constraint makes the optimization problems of each drone no longer independent.
[0089] In some embodiments, a trajectory smoothness score is calculated for each predicted state trajectory. The calculation process involves first performing discrete-time sampling on the predicted state trajectory to obtain a series of position, velocity, and acceleration state variables at various timestamps. Then, the rate of change of the magnitude of the position difference between adjacent sampling points is calculated as the position change rate, the rate of change of the velocity vector difference is calculated as the velocity change rate, and the rate of change of the acceleration vector difference is calculated as the acceleration change rate. The trajectory smoothness score is calculated by combining the rate of change of each order derivative. One calculation method is to normalize the rate of change of position, velocity, and acceleration by dividing each by a corresponding feature quantity to make them dimensionless. Then, weights are assigned to the normalized rate of change and summed. The continuity of the rate of change of the entire trajectory sequence is evaluated, which involves calculating the standard deviation of the normalized rate of change sequence. To calculate the load balance score for the predicted load force sequence, the calculation process involves comparing the predicted load force of all hoisting drones in the cluster with the expected load force obtained from the global plan at each identical moment in the predicted state trajectory. The absolute value of the force deviation for each drone at each moment is calculated, and the force deviations of all drones at all moments are summed to obtain the overall deviation metric. Simultaneously, the variance among the predicted load forces of all drones at each moment is calculated, and the variances at all moments are summed to obtain the overall variance metric. The load balance score is a weighted sum of the normalized overall deviation metric and the normalized overall variance metric. The normalization process involves dividing by a characteristic force and the square of that characteristic force, respectively.
[0090] Optionally, referring to Table 1, the trajectory smoothness score and load balance score can be calculated based on the data structure shown in Table 1. The score calculation module reads the predicted state trajectory and the predicted load force sequence and executes the above calculation process.
[0091] Table 1: Calculation Parameters for Trajectory Smoothness and Load Balancing Scores
[0092] Evaluation object Calculation indicators Detailed calculation description Predicting state trajectory Normalized rate of change of position The magnitude difference of the position vector difference between discrete points of the trajectory divided by the characteristic length . Predicting state trajectory Normalized rate of change The difference of velocity vectors between discrete points of the trajectory divided by the characteristic velocity . Predicting state trajectory Normalized rate of change of acceleration The difference of acceleration vectors between discrete points of the trajectory divided by the characteristic acceleration . Predicting state trajectory Smoothness rating Predicted load force sequence Normalized force tracking deviation The time series sum of the absolute values of the difference between the predictive and expected capabilities of each UAV divided by the characteristic force. . Predicted load force sequence Normalized force distribution variance The time-series sum of the variances of the predictive power among the various UAVs divided by the square of the eigenforce. . Predicted load force sequence Balance score
[0093] In the above table, This represents the number of discrete points in the prediction time domain. Let these represent the normalized position, velocity, and rate of change of acceleration at the k-th point, respectively. These are the corresponding dimensionless weighting coefficients. Normalized rate of change sequence standard deviation Dimensionless weighting coefficients for continuous evaluation. and These are the normalized force tracking deviation metric and the force distribution variance metric, respectively, both of which are dimensionless. These are the corresponding dimensionless weighting coefficients. Trajectory smoothness score. and load balancing score The calculation results are all dimensionless scalar scores, and the smaller the value, the smoother the trajectory and the more balanced the force on the cluster.
[0094] It is understandable that the distributed model predictive control algorithm can be solved using numerical optimization methods such as sequential quadratic programming or interior-point methods. A fixed-duration optimization problem is solved in each control cycle, and the first control variable of the optimized solution is applied to the UAV at the current moment. The specific form of the coupling constraint can be that the distance between the predicted position of each UAV and the predicted positions of all neighboring UAVs at each moment in the optimization time domain must be greater than a safety threshold. This constraint is added as an inequality constraint to the local optimization problem. The calculation of trajectory smoothness score and load balance score is a parallel process. The score results are used for subsequent priority ranking and dynamic adjustment of planning schemes. The score calculation module can be deployed in the central computing unit or the local computing unit of each UAV. If the smoothness score of the predicted state trajectory is too high, it indicates that the trajectory may have undergone drastic changes and needs adjustment; if the balance score of the predicted load force sequence is too high, it indicates that the force distribution may be uneven or seriously deviate from the plan, requiring adjustment. Feature length Characteristic velocity Characteristic acceleration and characteristic forces The settings are based on the typical motion and force dimensions of the specific hoisting task scenario.
[0095] See Figure 3 This is a comparison chart of the normalized change rates of the state variables in a coordinated lifting operation by a drone swarm, primarily showcasing the normalized change rates of three types of state variables. The acceleration change rate exhibits the most dramatic fluctuations, ranging from approximately 0.4 to 1.8, representing the largest amplitude among the three types. Acceleration directly corresponds to the drone's thrust / torque output; to track waypoints or avoid obstacles, frequent adjustments to control variables are required, hence the most significant fluctuations. The position change rate is the most stable overall, ranging from approximately 0.4 to 1.2, with the smallest amplitude, indicating good quality spatial position trajectory planning for the drone and the absence of drastic displacement changes. By comparing the fluctuations of the three curves, it is possible to accurately determine whether the problem lies in position planning, velocity tracking, or acceleration control. If the acceleration change rate at a certain time step exceeds a threshold, a system warning can be triggered, dynamically adjusting and optimizing targets or constraints to prevent load swaying or drone instability due to sudden trajectory changes during lifting.
[0096] In one embodiment of the present invention, the predicted state trajectories and predicted load force sequences of all hoisting drones are aligned and adjusted for consistency based on the calculated trajectory smoothness score and load balance score. During this alignment adjustment process, it is ensured that the spatial positions corresponding to the trajectories of all drones at the same timestamp satisfy the preset hoisting formation geometric constraints, and simultaneously that the load force sequences of all drones satisfy the preset system force closure constraints. The adjusted state sequences and force sequences of each drone, which are strictly aligned in both time and space, are discretized into high-frequency low-level control commands. Each command contains the position, attitude, and velocity information that the drone should reach at a specific time, as well as the output sling tension vector. The discrete control commands of all hoisting drones in the cluster are globally sorted and packaged in chronological order, ultimately forming a spatiotemporally synchronized cluster cooperative control command set.
[0097] In practical implementation, based on the calculation results of trajectory smoothness score and load balance score, a consistency alignment adjustment is performed on the predicted state trajectory and predicted load force sequence of all individuals in the hoisting drone swarm. The adjustment operation ensures that the spatial relative positions of all drones at the same time satisfy the pre-set formation geometry constraints of the hoisting task. For example, when four drones are hoisting a large structural component, they need to maintain a rectangular vertex distribution. At the same time, the adjustment operation ensures that the resultant force and resultant moment of the tension vector output by all drones at any time satisfy the static equilibrium condition of the transported target, i.e., the system force closed constraint. The consistency alignment adjustment process involves fine-tuning the time axis, correcting the translation or rotation of the spatial path, and redistributing the tension magnitude to ensure spatiotemporal synchronization and mechanical feasibility. For example, when the predicted trajectory smoothness score of one drone shows that its acceleration changes drastically at a certain time period, while the load balance score of another drone shows that its load force is consistently high, the adjustment algorithm will smooth the trajectory of the former and partially transfer the load to other drones. After dynamic adjustment, the predicted state trajectory and predicted load force sequence of each UAV are completely aligned in timestamps, satisfying formation constraints in space and closure constraints in mechanics. These adjusted sequences form the basis for the generation of subsequent control commands. Refer to Table 2. The changes in key parameters before and after adjustment can be explained by the comparison data shown in Table 2.
[0098] Table 2: Comparison of Key Parameters Before and After Trajectory and Force Sequence Adjustment
[0099] drone number Maximum positional deviation (meters) before adjustment Maximum positional deviation after adjustment (meters) Adjust the range of changes in the preload force (Newtons). Adjusted load force variation range (Newtons) Drone 1 0.15 0.05 [85,120] [92,110] Drone 2 0.22 0.04 [78,130] [95,108] Drone 3 0.18 0.06 [88,125] [94,112] Drone 4 0.20 0.05 [82,128] [90,105]
[0100] In some embodiments, the consistency alignment adjustment ensures that the spatial positions of all trajectories at the same timestamp satisfy a preset lifting formation geometric constraint. This geometric constraint is defined by a set of mathematical inequalities concerning the relative distances and angles between UAVs. The adjustment process is achieved by solving a constrained optimization problem, where the optimization variables are the time offset and spatial coordinate fine-tuning of each UAV trajectory point, and the objective function is to minimize the total adjustment magnitude. The consistency alignment adjustment also ensures that all load force sequences satisfy a preset system force closure constraint. This constraint requires that the sum of the pull vectors of all UAVs equals the weight of the carried target, and that the resultant torque of these pull vectors about the target's center of mass is zero. The adjustment process satisfies this constraint by correcting the pull direction and magnitude of each UAV, while simultaneously minimizing the load balance score. The adjusted, time- and space-aligned state and force sequences of each UAV are discretized into high-frequency control commands. The discretization process samples the continuous time-state-force curves at a fixed control cycle. Each sampling point generates a control command. Each control command includes the three-dimensional position coordinates that the UAV should achieve at a specific moment, the attitude quaternion that it should maintain, the velocity vector that it should have, and the component of the cable tension vector that it should output in the body coordinate system.
[0101] In some embodiments, the generation of control instructions can be expressed as the following mapping relationship:
[0102]
[0103] in: Indicates the first The drone in Control commands for each control cycle This indicates an instruction encapsulation function. Indicates at time No. The three-dimensional position coordinates of the drone. Indicates at time No. The attitude quaternion of the drone, Indicates at time No. The velocity vector of the drone Indicates at time No. The drone should output the sling tension vector. The command encapsulation function encodes the position, attitude, velocity, and tension information into a data packet format that the flight controller can directly parse.
[0104] Optionally, strict timestamp alignment is achieved through a globally synchronized clock. All UAV control command sequences generate timestamps based on the same high-precision clock source, with timestamp synchronization accuracy reaching the microsecond level to ensure timing consistency during distributed execution. Command sorting and packaging are handled by a central command scheduling module or a distributed consensus protocol. This globally sorts the discrete control commands of all hoisting UAVs according to their timestamp order, combining commands that need to be sent to different UAVs at the same or adjacent times into a single data packet. This ultimately forms a spatiotemporally synchronized cluster collaborative control command set. The command set's data structure contains an ordered list of commands, each item of which includes the target UAV identifier, the command's effective timestamp, and specific control quantity data.
[0105] It is understandable that the mathematical expression of a force-closed constraint is: and ,in It is the number of drones. It is the first The pull vector of the drone It refers to the mass of the target object being transported. It is the gravitational acceleration vector. From the center of gravity of the transport target material to the first The vector alignment adjustment of each suspension point must ensure that the adjusted force sequence satisfies these two equations at all times. The discretization frequency of the control command, i.e. the operating frequency of the control system, is usually between 100 Hz and 500 Hz. High-frequency discretization can track continuous trajectories more accurately, but it will increase the communication and computing load. The choice of frequency needs to strike a balance between accuracy and load.
[0106] See Figure 4 This is a single-dimensional position trajectory diagram of a drone swarm collaborative hoisting system, visually displaying the position changes of four drones within 0-10 seconds. All four drones start synchronously from 0 meters and increase their position linearly over time without significant overlap or misalignment, indicating that the globally planned waypoint sequence is effectively tracked under distributed MPC control, and the swarm formation remains stable. Drone 3 is the fastest, reaching 2.5 meters at 10 seconds, and is responsible for the main traction or attitude adjustment tasks; Drone 2 is the slowest, reaching 1.5 meters at 10 seconds, and may be responsible for load stabilization or auxiliary support; Drones 1 and 4 are at intermediate speeds, forming an intermediate echelon to ensure the geometric constraints of the hoisting formation. All curves show smooth linear growth without sharp fluctuations or inflection points, indicating a low trajectory smoothness score. The position difference between the four drones remains within a controllable range, meeting obstacle avoidance constraints and formation geometry requirements.
[0107] In one embodiment of the present invention, the robustness of the spatiotemporally synchronized cluster cooperative control instruction set is verified by decomposing the complex instruction set into multiple independently executable and fault-tolerant control sub-task clusters. This verification and decomposition process includes: injecting simulated environmental disturbance models into the cluster cooperative control instruction set, including wind field changes, sensor noise, and actuator errors; performing forward simulation to verify whether the cluster system can still stably track instructions and complete the hoisting task after the disturbance is injected, and identifying instruction segments that may fail or conflict under the disturbance; dividing the original cluster cooperative control instruction set into multiple relatively independent instruction subsets based on the spatiotemporal dependencies and functional correlations between instructions, with each instruction subset and its corresponding UAV subset constituting a control sub-task cluster; designing a backup control law and intra-cluster communication protocol for each control sub-task cluster, so that when some UAVs within a control sub-task cluster fail, the cluster can maintain basic hoisting functions through autonomous adjustment; and continuously acquiring real-time image streams collected by the visual sensing devices of the UAVs within the cluster during the execution of a control sub-task cluster. The actual pose and velocity of the target object, as well as the actual pose and sling tension direction of each UAV, are calculated in real time from the real-time image stream. The calculated actual state is compared with the expected state of the corresponding commands in the control sub-task cluster to calculate the position tracking error, attitude tracking error, and force tracking error. These tracking errors are input into a state inversion observer, which estimates the system state variables that are not directly measured based on the system dynamics model. It then combines the direct errors and the estimated state to calculate the feedforward compensation and feedback compensation for environmental disturbances and model uncertainties. These two factors together constitute the compensation amount. The tracking errors and compensation amounts are fused to generate a preliminary lifting control correction amount. This preliminary correction amount is superimposed on the control commands in the original control sub-task cluster to form the corrected expected control amount. The corrected expected control amount is input to a preset lifting stability criterion for verification. This criterion includes the target object torque balance condition, the sling tension non-negativity condition, and the UAV thrust limiting condition. If the revised expected control quantity passes the hoisting stability criterion test, it is converted into a low-level control signal that can directly drive the UAV flight control and actuators, i.e., a real-time updated UAV swarm coordinated hoisting and transportation control signal. If it fails the test, the revised expected control quantity is limited or adjusted based on the feedback of the criterion until it passes the test, and then a real-time updated control signal is generated.
[0108] In practical implementation, the anti-interference robustness verification of the spatiotemporally synchronized cluster cooperative control command set is performed. The verification process involves injecting simulated environmental disturbance models into the cluster cooperative control command set. These models include a random wind field variation model, a sensor Gaussian white noise model, and an actuator response delay and deviation model. The command set after disturbance injection is then subjected to forward simulation in a digital twin simulation environment that includes UAV dynamics, sling models, and target object dynamics. The forward simulation uses numerical integration to advance the system state. During the simulation, the actual trajectory of each UAV, sling tension, and the motion state of the target object are recorded. The verification verifies whether the cluster system can still stably track commands and complete the lifting task after disturbance injection. Stability criteria include whether the deviation between the UAV trajectory and the command trajectory exceeds the tolerance, whether the sling tension is always positive and does not exceed the upper limit, and whether the target object's attitude angle remains stable. Analysis of the simulation data identifies command segments that fail or conflict under disturbance. Failure manifests as the UAV failing to reach the commanded position, while conflict manifests as the distance between UAVs falling below a safety threshold or abnormal oscillations in the sling tension. Based on the spatiotemporal dependencies and functional correlations between instructions, the cluster collaborative control instruction set is divided into multiple relatively independent instruction subsets. Spatiotemporal dependencies refer to the temporal order of instructions and their coupling relationship along spatial paths. Functional correlations refer to the degree of cooperation between UAVs when completing specific sub-tasks such as lifting, translation, and rotation. Each instruction subset and its corresponding UAV subset constitute a control subtask cluster. The instructions within a control subtask cluster exhibit high spatiotemporal locality and functional cohesion. A backup control law and intra-cluster communication protocol are designed for each control subtask cluster. The backup control law is typically a simplified PID controller or model predictive controller, used to take over control when the main control command fails due to disturbances. The intra-cluster communication protocol specifies the frequency, content, and format of status information exchanged between UAVs within the cluster. This allows the control subtask cluster to detect the failure of some UAVs according to the communication protocol and autonomously adjust the control commands of the remaining UAVs through the backup control law, thereby maintaining basic lifting functions such as hovering or slow landing.
[0109] In some embodiments, a state inversion calculation based on visual feedback is performed for each control subtask cluster. During the execution of the control subtask cluster, each UAV within the cluster continuously acquires a real-time image stream collected by its own visual sensing device. The real-time image stream contains visual information about the environment, other UAVs, and the target object. The actual pose and actual velocity of the target object are calculated in real time from the real-time image stream. The calculation is achieved through feature point matching, 3D reconstruction, and optical flow. Simultaneously, the actual pose of each UAV relative to global features or the target object is calculated from the image, as well as the actual sling tension direction estimated by combining the projection direction of the sling in the image with depth information. The calculated actual state is compared with the expected state of the corresponding command in the control subtask cluster. The expected state is parsed from the command. The comparison calculates the position tracking error, attitude tracking error, and force direction tracking error. The position tracking error is the vector difference between the actual position and the expected position; the attitude tracking error is the difference between the actual attitude quaternion and the expected attitude quaternion; and the force direction tracking error is the angle between the actual sling tension direction vector and the expected direction. The tracking error is input to the state inversion observer, which is based on a system dynamics model of the UAV and the target vehicle coupled together. The model includes mass, inertia, kinematics, and dynamic equations. The observer uses the measurable tracking error as input and estimates the system state variables that are not directly measured by designing an internal state estimator, such as the center of mass acceleration of the target vehicle and the unknown aerodynamic disturbance force on the UAV. It then combines the directly measured tracking error with the estimated system state to calculate the feedforward compensation and feedback compensation for environmental disturbances and model uncertainties. The feedforward compensation is used to offset known or estimated disturbance models, and the feedback compensation is used to correct the control command in the closed loop to eliminate the tracking error. The feedforward compensation and feedback compensation together constitute the compensation.
[0110] Optionally, the robustness index of the control subtask cluster can be calculated and evaluated using the following formula:
[0111]
[0112] in: This represents the robustness index of the control subtask cluster; a higher value indicates better robustness. This represents the ratio of the number of instruction failures occurring within this control subtask cluster in multiple Monte Carlo perturbation simulations. This represents the average percentage of task completion for the control subtask cluster in a simulation that did not completely fail. and The weighted coefficients are and satisfy the following conditions: This formula is used to quantitatively evaluate the performance of each control subtask cluster under disturbances after decomposition, and to guide the design of backup control laws.
[0113] It is understandable that the real-time updated UAV swarm collaborative lifting and transport control signal is generated based on the tracking error and compensation amount of each control command. The generation process first fuses the tracking error and compensation amount, which is usually achieved by weighted summation or through a compensation controller such as a sliding mode controller, to generate an initial lifting control correction amount. The initial lifting control correction amount is superimposed with the control commands in the original control sub-task cluster to form the corrected expected control amount, which includes adjustments made to resist disturbances and eliminate errors. The corrected expected control amount is then input to a preset lifting stability criterion for verification. The lifting stability criterion includes the target moment balance condition, the non-negative cable tension condition, and the UAV thrust limiting condition. The moment balance condition requires that the resultant torque of all cable tensions about the target's center of gravity is balanced with the moment of inertia caused by the target's angular acceleration. The non-negative cable tension condition requires that the tension of each cable is greater than zero. The UAV thrust limiting condition requires that the total thrust output of each UAV is within its actuator capability range. If the corrected expected control quantity passes the hoisting stability criterion test, it is transformed into a low-level control signal that can directly drive the UAV flight control actuators through a control allocation algorithm. This is the real-time updated UAV swarm collaborative hoisting and transport control signal, which is typically a motor speed command or a servo deflection angle command. If the corrected expected control quantity fails the hoisting stability criterion test, it is limited or adjusted based on the criterion feedback. Limiting involves cutting off forces or torques exceeding the threshold to within the allowable range, while adjustment involves resolving for the closest feasible control quantity that satisfies the stability criterion until the test is passed, at which point the real-time updated UAV swarm collaborative hoisting and transport control signal is generated again.
[0114] See Figure 5 This is a visual feedback state inversion calculation and analysis diagram of UAV swarm collaborative hoisting. The force tracking error initially experiences the largest disturbance, converges rapidly, and reaches a maximum error of approximately 25N at around 2.5 seconds, which is the maximum deviation caused by the simulated environmental disturbance injected into the system. After 5 seconds, the error decreases rapidly and approaches 0 at 20 seconds, indicating that the Distributed Model Predictive Control (DMPC) algorithm has extremely strong anti-interference capabilities and can quickly correct thrust control commands. The attitude tracking error fluctuates progressively and gradually stabilizes, reaching a local peak of approximately 12 degrees between 2.5 and 5 seconds, and then slowly decreases. Attitude adjustment lags behind force adjustment, which is a typical physical characteristic of UAV sling systems. However, it eventually converges to an extremely low error range, verifying that the non-negativity condition of sling tension and the UAV thrust limiting condition are satisfied. The position tracking error remains at an extremely low level throughout, with the entire curve smooth and close to the zero axis, and the maximum error not exceeding 0.2 meters. This proves that the global collaborative hoisting plan is perfectly tracked, and the system can strictly maintain geometric formation constraints even under strong disturbances, satisfying the "position stability criterion".
[0115] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A method for collaborative lifting and transportation of unmanned aerial vehicles (UAVs) based on AI vision algorithms, characterized in that, include: Based on the original environmental images and the original morphological images of the target object collected by the hoisting drone, a dynamic environmental perception tensor is generated. The dynamic environment perception tensor is input into the trained collaborative decision neural network, which generates a global collaborative hoisting planning diagram by iteratively optimizing the motion path and force distribution scheme of each UAV. Based on the global collaborative hoisting planning diagram, a distributed model predictive control algorithm is used to derive the predicted state trajectory and predicted load force sequence of each individual in the hoisting drone swarm at multiple future moments; For each predicted state trajectory and predicted load force sequence, calculate its trajectory smoothness score and load balance score, and prioritize and dynamically adjust the planning schemes based on the trajectory smoothness score and the load balance score; The predicted state trajectory and the predicted load force sequence after dynamic adjustment are integrated to form a spatiotemporally synchronized cluster collaborative control instruction set. Each instruction in the instruction set corresponds to the pose and tension control amount of a hoisting UAV at a specific moment.
2. The method for collaborative lifting and transportation of unmanned aerial vehicles (UAVs) based on AI vision algorithms as described in claim 1, characterized in that, The process of generating a dynamic environment perception tensor based on the original environmental images and the original morphological images of the transported target object acquired by the hoisting drone includes: The initial visual data stream is formed by using a hoisting drone equipped with a visual sensing device to collect original environmental images of the collaborative operation area and original morphological images of the transported target. Perform multi-scale spatiotemporal fusion and feature extraction operations on the initial visual data stream to separate the background environment feature stream, the carrier target feature stream, and the UAV body feature stream; Based on the separated background environment feature stream, the target object feature stream, and the UAV body feature stream, a dynamic environment perception tensor for collaborative hoisting is constructed.
3. The unmanned aerial vehicle group cooperative hoisting and carrying method based on an AI vision algorithm according to claim 2, characterized in that, Perform multi-scale spatiotemporal fusion and feature extraction operations on the initial visual data stream to separate the background environment feature stream, the carrier target feature stream, and the UAV body feature stream, including: The original environmental image is subjected to pyramid downsampling to construct a multi-scale environmental image sequence. Static obstacle contours, dynamic obstacle motion vectors, and wind disturbance features are identified and extracted from the multi-scale environmental image sequence and fused into the background environmental feature stream. The original morphological image is reconstructed into a three-dimensional point cloud and its skeleton is extracted. The centroid position, attitude angle, mass distribution and moment of inertia of the target object are calculated to form the feature flow of the target object. By using visual odometry and feature point matching technology, the pose, velocity, acceleration and sling angle of each hoisting drone are calculated from continuous image frames to form the feature flow of the drone body; A spatiotemporal alignment mechanism is established to synchronize and calibrate the background environment feature stream, the carrier target feature stream, and the UAV body feature stream based on a unified world coordinate system and high-precision timestamps. An attention-guided feature separation network is used to decouple the synchronously calibrated hybrid feature streams to obtain spatially and semantically independent background environment feature streams, carrier target object feature streams, and UAV body feature streams.
4. The unmanned aerial vehicle group cooperative hoisting and carrying method based on an AI vision algorithm according to claim 3, characterized in that, The dynamic environment perception tensor is input into a trained collaborative decision-making neural network. This network iteratively optimizes the motion paths and force distribution schemes of each UAV to generate a global collaborative hoisting plan, including: Initialize a fully connected interactive graph with each hoisting drone as a node, and encode the dynamic environment perception tensor as the initial features of each node in the graph; In the collaborative decision-making neural network, the environmental and state information of each node and its neighboring nodes are aggregated through a graph attention layer to update node features; The updated node features are input into the policy network, which outputs a series of waypoint sequences and expected lifting force sequences for each UAV in the planning time domain. The long-term value of the combined action consisting of the waypoint sequence and the desired lifting pull sequence is evaluated through a value network, and constraints for avoiding collisions, maintaining formation, and minimizing energy consumption are introduced. A gradient-based strategy optimization method is adopted to iteratively update the parameters of the strategy network and the value network until the joint action satisfies all constraints and the long-term value converges. Finally, the optimized set of node actions is mapped to the global collaborative hoisting planning diagram that includes spatial path, time arrangement and force allocation.
5. The unmanned aerial vehicle group cooperative hoisting and carrying method based on an AI vision algorithm according to claim 4, characterized in that, Based on the aforementioned global collaborative hoisting planning diagram, a distributed model predictive control algorithm is used to derive the predicted state trajectory and predicted load force sequence of each individual in the hoisting drone swarm at multiple future moments, including: The waypoint sequence and the expected lifting tension sequence assigned to a specific lifting UAV in the global collaborative lifting planning map are used as the reference trajectory and reference force sequence for the local rolling time domain optimization of the specific lifting UAV. Based on the nonlinear dynamics model of the specific hoisting UAV, a local optimization problem is constructed with the goal of tracking the reference trajectory and the reference force sequence and constrained by the capabilities of its own actuators. In each control cycle, the local optimization problem is solved to obtain the predicted state trajectory and the predicted load force sequence of the specific hoisting UAV in the future finite time domain; During the solution process, the predicted state information is exchanged with other hoisting drones in the communication neighborhood. The predicted trajectory of the neighboring drones is used as a constraint on its own optimization problem in a coupled manner to ensure the coordination and collision-free nature of the cluster prediction.
6. The method for collaborative lifting and transportation of unmanned aerial vehicles (UAVs) based on AI vision algorithms as described in claim 5, characterized in that, For each of the predicted state trajectories and the predicted load force sequence, calculate its trajectory smoothness score and load balance score, including: For a predicted state trajectory, calculate its rate of change on each derivative, including the rate of change of position, the rate of change of velocity, and the rate of change of acceleration; By combining the weighted sum of the rates of change of each derivative and evaluating their continuity, the trajectory smoothness score, which reflects the smoothness of the motion, is obtained. For the predicted load force sequence, calculate the deviation between the predicted load force and the expected load force of each hoisting drone at the same time in the predicted state trajectory, as well as the variance of the predicted load force among each hoisting drone. By combining the weighted evaluation of the deviation and the variance, a load balance score reflecting the rationality of the stress on the cluster is obtained.
7. The unmanned aerial vehicle group cooperative hoisting and carrying method based on an AI vision algorithm according to claim 6, characterized in that, The integration of the dynamically adjusted predicted state trajectory and the predicted load force sequence forms a spatiotemporally synchronized cluster collaborative control instruction set, including: Based on the trajectory smoothness score and the load balance score, the predicted state trajectory and the predicted load force sequence of all hoisting drones are adjusted for consistency alignment. In the consistency alignment adjustment, it is ensured that the spatial positions corresponding to all trajectories at the same timestamp meet the preset hoisting formation geometric constraints, and all load force sequences meet the preset system force closure constraints. The adjusted, time- and space-aligned states and force sequences of each UAV are discretized into high-frequency control commands. Each command contains the position, attitude, speed that the UAV should achieve at a specific moment, as well as the output cable tension vector. All discrete control commands of the hoisting drones are globally sorted and packaged in chronological order to form the spatiotemporally synchronized cluster collaborative control command set.
8. The unmanned aerial vehicle group cooperative hoisting and carrying method based on an AI vision algorithm according to claim 7, characterized in that, Also includes: The anti-interference robustness verification of the spatiotemporally synchronized cluster cooperative control instruction set is performed, and the complex instruction set is decomposed into multiple independently executable and fault-tolerant control sub-task clusters. For each of the control subtask clusters, a state inversion calculation based on visual feedback is performed to quantify the tracking error and compensation amount of each control command to the actual hoisting state. Based on the tracking error and compensation amount of each control command, and combined with the preset hoisting stability criterion, a real-time updated UAV swarm collaborative hoisting and transportation control signal is generated. The process of performing anti-interference robustness verification on the spatiotemporal synchronized cluster cooperative control instruction set involves decomposing the complex instruction set into multiple independently executable and fault-tolerant control subtask clusters, including: A simulated environmental disturbance model is injected into the cluster collaborative control command set. The environmental disturbance model includes wind field changes, sensor noise, and actuator errors. Perform forward simulation to verify whether the cluster system can still stably track instructions and complete the hoisting task after the disturbance is injected, and identify the instruction segments that fail or conflict under the disturbance. Based on the spatiotemporal dependencies and functional correlations between instructions, the cluster collaborative control instruction set is divided into multiple relatively independent instruction subsets. Each instruction subset and its corresponding UAV subset constitute a control subtask cluster. A backup control law and intra-cluster communication protocol are designed for each of the control sub-task clusters, so that when some UAVs in a certain control sub-task cluster fail, the control sub-task cluster can autonomously adjust and maintain basic hoisting functions. 9.The AI vision algorithm-based unmanned aerial vehicle group cooperative hoisting and carrying method according to claim 8, wherein For each of the control subtask clusters, a state inversion calculation based on visual feedback is performed to quantify the tracking error and compensation amount of each control command to the actual hoisting state, including: During the execution of one of the control subtask clusters, real-time image streams collected by the UAV visual sensing devices within the cluster are continuously acquired; The actual pose and velocity of the target object, as well as the actual pose and direction of the sling tension of each UAV, are calculated in real time from the real-time image stream. The calculated actual state is compared with the expected state of the corresponding instruction in the control subtask cluster to calculate the position tracking error, attitude tracking error and force tracking error. The tracking error is input to the state inversion observer, which estimates the system state variables that are not directly measured based on the system dynamics model. It then combines the direct error with the estimated state to calculate the feedforward compensation and feedback compensation for environmental disturbances and model uncertainties, which together constitute the compensation amount.
10. The method for collaborative lifting and transportation of unmanned aerial vehicles (UAVs) based on AI vision algorithms as described in claim 9, characterized in that, Based on the tracking error and compensation amount of each control command, and combined with a preset hoisting stability criterion, a real-time updated UAV swarm collaborative hoisting and transportation control signal is generated, including: The tracking error is fused with the compensation amount to generate a preliminary hoisting control correction amount; The preliminary hoisting control correction amount is superimposed on the control instructions in the original control subtask cluster to form the corrected desired control amount. The modified desired control quantity is input into the hoisting stability criterion for verification. The hoisting stability criterion includes the target moment balance condition, the non-negativity condition of the sling tension, and the UAV thrust limiting condition. If the modified expected control quantity passes the hoisting stability criterion test, it is converted into a low-level control signal that can directly drive the UAV flight control and actuator, namely the real-time updated UAV swarm collaborative hoisting and transportation control signal. If the test fails, the corrected expected control quantity will be limited or adjusted based on the criterion feedback until the test passes before generating the real-time updated UAV swarm collaborative hoisting and transportation control signal.