Helicopter cluster level intelligent safety planning method based on data correlation sampling
By optimizing safe formation spacing through data correlation sampling and distributed formation coordination, combined with reinforcement learning methods based on experience replay, the problem of safety planning and autonomous control of helicopter formation systems in complex environments was solved. This achieved highly robust fault-tolerant tracking control, improving the stability and autonomy of the formation system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-21
AI Technical Summary
In complex environments, helicopter formation systems face challenges in safety planning, autonomous formation coordination, and robust fault-tolerant tracking control. In particular, existing methods struggle to achieve high-precision control and stability when actuator failures and dynamic obstacles are present.
A helicopter cluster hierarchical intelligent safety planning method based on data correlation sampling is adopted. The safe formation spacing is optimized through a greedy algorithm. Combined with distributed formation coordination and experience-based reinforcement learning, the control strategy is optimized to ensure stable flight and trajectory tracking of the formation system under fault conditions.
It significantly improves the stability and reliability of helicopter formation systems under fault conditions, enhances the efficiency of autonomous obstacle avoidance and formation recovery, strengthens the robustness and tracking accuracy of formation control, and solves the problems of low formation control accuracy and weak fault tolerance in existing technologies.
Smart Images

Figure CN122431407A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of safe path planning and fault-tolerant tracking control technology for helicopter formation systems, and in particular to a hierarchical intelligent safety planning method for helicopter clusters based on data correlation sampling. Background Technology
[0002] Due to the inherent nonlinear and underactuated dynamic characteristics of helicopter systems, their flight is often subject to strict kinematic constraints and is highly sensitive to external environmental disturbances. In low-altitude airspace, the complex electromagnetic environment can also lead to communication disruptions between the formation system and the ground control station, making it difficult for the formation to obtain timely and reliable control commands. Simultaneously, the low-altitude environment contains various dynamic obstacles, such as runaway drones, aerial debris, and balloons, which may collide with the helicopter formation, causing actuator failures or flight performance degradation. Therefore, achieving safe planning and stable control of helicopter formation systems under complex conditions involving dynamic obstacles and actuator failures remains a significant technical challenge.
[0003] In helicopter formation missions, to achieve safe flight and collaborative operations, it is necessary to rationally plan the formation structure and ensure that each helicopter maintains a predetermined relative position while tracking the desired trajectory. Existing formation planning and coordination methods are mostly based on traditional control theory, achieving coordinated movement among multiple aircraft through the design of different control strategies. These methods typically rely on a leader node in the formation, using reference information provided by the leader to guide the other members' formation. However, in complex environments, due to unstable communication links or leader node failure, formation methods relying on leader information struggle to achieve stable and reliable formation reconfiguration and autonomous adjustment. Furthermore, in environments with dynamic obstacles, the formation needs to adjust its structure in real time according to environmental changes to ensure safe obstacle avoidance and formation stability, placing higher demands on the real-time performance and autonomy of formation planning methods.
[0004] On the other hand, during long-term missions, the aerodynamic efficiency of the helicopter's main rotor may degrade due to factors such as mechanical wear and changes in aerodynamic conditions, leading to significant changes in the generated aerodynamic forces and consequently causing actuator performance degradation or even failure. Such failures can significantly impact the helicopter's trajectory tracking performance, and in severe cases, may even cause instability in the formation system. Therefore, designing a fault-tolerant tracking control method in helicopter formation systems is crucial for ensuring system stability and mission reliability. Existing fault-tolerant control methods are typically based on classical control theories such as adaptive control, robust control, or sliding mode control, which can ensure that the system tracking error remains bounded given a known disturbance range or fault model. However, these methods often rely on relatively accurate system models and prior fault information. In actual helicopter formation systems, due to the complexity of the environment and the diversity of fault types, relevant information is often difficult to obtain in advance, limiting the adaptability and scalability of traditional methods in complex environments.
[0005] In recent years, reinforcement learning methods have attracted widespread attention in the field of complex system control due to their ability to learn control strategies through data-driven approaches without relying on precise system models. By introducing reinforcement learning into fault-tolerant control frameworks, control systems can achieve adaptive adjustment in the face of actuator performance degradation or dynamic system changes, thereby improving the system's robustness and resilience. However, reinforcement learning methods still suffer from problems such as insufficient online learning stability and low data utilization efficiency in practical applications. Especially in helicopter fault scenarios, the fault data that the system can acquire is usually limited and difficult to repeat, which affects the training effect and convergence speed of reinforcement learning algorithms. Therefore, how to improve the data utilization efficiency and stability of reinforcement learning control methods under limited data conditions remains an important technical problem that needs to be solved.
[0006] In summary, achieving safe planning, autonomous formation coordination, and robust fault-tolerant tracking control for helicopter formation systems remains a significant technical challenge in complex environments with dynamic obstacles and actuator failures. Summary of the Invention
[0007] Purpose of the invention: The purpose of this invention is to provide a helicopter formation hierarchical intelligent safety planning method based on data correlation sampling, which solves the technical defects of existing technologies such as low formation control accuracy and weak fault tolerance in actuator failure scenarios, and improves the autonomy, safety and reliability of helicopter formation systems in complex environments.
[0008] Technical solution: A hierarchical intelligent safety planning method for helicopter clusters based on data correlation sampling, including the following steps:
[0009] S1, let it be... A helicopter formation system consisting of several helicopters is established, and a helicopter position loop model is created. Considering the failure of reduced aerodynamic efficiency of the rotor main wing, a helicopter failure model under this failure is created.
[0010] S2, combining speed and acceleration constraints, constructs a safe formation spacing planning problem; a greedy algorithm is used to optimize and adjust the safe formation spacing to ensure that the helicopter formation system can safely avoid obstacles and fly stably in the event of a fault.
[0011] S3. A distributed formation coordination algorithm is adopted to optimize the target of the distributed formation coordinator, so that each helicopter can independently adjust its own state and share the system matrix information, ensuring that all UAVs can adjust their formation according to the safe formation spacing obtained in step S2 in a dynamic environment and generate the desired reference trajectory.
[0012] S4 employs an experience-based reinforcement learning method to evaluate the state of the helicopter formation system and optimize the control strategy through a critic neural network; it updates the experience pool samples using cosine similarity; and it adaptively optimizes the control law in actuator failure scenarios to ensure the helicopter tracks the reference trajectory.
[0013] Furthermore, the helicopter position ring model is as follows:
[0014] ,
[0015] in, , Representing inertial frames of reference respectively Next The position and velocity vectors of the helicopter. Indicates the first The mass of a helicopter; Indicates the first The expected forces acting on a helicopter, This represents the lift of the main rotor, and T represents the matrix transpose. Represents the gravitational acceleration vector. Indicates the first External disturbances to the helicopter; Represents the coordinate system of the machine body to inertial coordinate system The transformation matrix is specifically represented as:
[0016] ,
[0017] in, , , These are the helicopter's pitch angle, roll angle, and yaw angle, respectively. , , ;
[0018] The helicopter fault model is represented as follows:
[0019] ,
[0020] in, , These represent the helicopter's position and speed under main rotor failure. This indicates the change in force caused by a main rotor malfunction. This indicates the efficiency of the helicopter rotor.
[0021] Furthermore, the safe formation spacing planning problem is described as a constrained optimization problem, where the decision variables are... Set as the formation spacing under null projection rate of change, Represents a vector under constraints A definite projection on the null space. This represents the expected formation spacing vector for all helicopters. Let the desired formation spacing vector of the i-th helicopter be denoted by ; then the optimization problem is designed as:
[0022] ,
[0023] ,
[0024] in, This represents the optimization index function. Indicates a finite forecast time window; This indicates the magnitude of change in the current control input compared to the previous step. This represents the long-term formation deviation within the future time window t. This indicates the deviation between the current formation and the initial desired formation. , and These are the weighting coefficients, Indicates the time integral term; This represents the basis for mapping null space vectors back to the original space; Indicates the location of the obstacle. Indicates the safe separation distance between the helicopter and the obstacle. Indicates the collision-free distance between helicopters. , These represent the velocity and acceleration of the change in formation spacing, respectively. This indicates the maximum normal helicopter speed. This represents the maximum acceleration of a normal helicopter. This indicates the maximum speed of the malfunctioning helicopter. This indicates the maximum acceleration of the malfunctioning helicopter.
[0025] Furthermore, in step S3, the target of the distributed formation coordinator is set as follows:
[0026] ,
[0027] in, , They represent the helicopter at time t. and Current desired position; Helicopter at time t and The expected relative distance between them Indicates helicopter The relative position with respect to a fixed reference point; for each helicopter, the proposed distributed formation coordinator can be represented as:
[0028] ,
[0029] in, This represents the system matrix, used to adjust the state of the helicopter reference trajectory; Indicates the topology connection strength; for abbreviation; , These are the system matrix adjustment gain and the system state adjustment gain, respectively, used to adjust the system's convergence performance. This represents the coordination interval that can be directly invoked by the coordinator, where The distance transformation matrix is used to define the conversion relationship between the planned distance and the cooperative distance; the formation spacing error is defined as... .
[0030] Furthermore, in step S4, data samples are selectively retained based on the cosine similarity between the new data and existing samples in the experience pool; newly generated data samples... With existing samples in the experience pool Cosine similarity between Represented as:
[0031] ,
[0032] Where T represents the matrix transpose;
[0033] If the cosine similarity between new data and any existing sample exceeds the historical maximum, the new data will be discarded; otherwise, the historical sample most similar to the new data will be deleted.
[0034] Furthermore, in step S4, the definition is... To account for positional error, a sliding mode term is introduced. The desired position of the i-th helicopter is achieved by constructing an optimization objective function. Trajectory tracking; with the goal of minimizing tracking error and controlling energy consumption, an optimization objective function is constructed. for:
[0035] ,
[0036] in, Represents the cost function, , These are the state weight matrix and the control weight matrix, respectively. Represents the sliding mode gain matrix. ;
[0037] sliding surface Represented as:
[0038] ,
[0039] in, This represents an unknown state matrix, which contains faults in the system. and disturbance ; Given the control gain matrix, It represents the control input and serves as an optimization control term in reinforcement learning algorithms;
[0040] The control law is designed as follows:
[0041] ,
[0042] The network of design critics is:
[0043] ,
[0044] Where T represents the matrix transpose; , It includes Activation functions of hidden neurons; Represents the estimated weight matrix; It is the optimal weight matrix. It is the approximation error of the neural network;
[0045] The weight update law for designing a commentator neural network is as follows: :
[0046] ,
[0047] ,
[0048] ,
[0049] ,
[0050] ,
[0051] in, It is the weighting adjustment factor;
[0052] For helicopter models that include faults and environmental disturbances, the aforementioned control law and weight update law ensure that the position tracking error is maintained. It is uniform and eventually converges to a bounded state.
[0053] Compared with the prior art, the significant advantages of this invention are as follows:
[0054] 1. This invention designs a hierarchical helicopter formation planning and control scheme that considers actuator failures and dynamic obstacles. By mapping the failure rate of the control layer to the velocity and acceleration constraints of the planning layer, the desired trajectory generated by the planning layer can accurately adapt to the failure state requirements of the control layer, thus realizing a hierarchical collaborative fault-tolerant mechanism. This scheme effectively mitigates the impact of main rotor failures on helicopter formation performance, solves the technical defects of low formation control accuracy and weak fault tolerance in existing technologies under actuator failure scenarios, and significantly improves the stability and reliability of the helicopter formation system under fault conditions.
[0055] 2. To address the obstacle avoidance trajectory planning requirements in leaderless scenarios, this invention proposes a distributed formation consensus strategy. By using a finite-time greedy algorithm to determine the formation configuration, it can simultaneously achieve obstacle avoidance and formation recovery functions. In complex environments with dynamic obstacles, it has stronger environmental adaptability and formation coordination capabilities, effectively solving the problem that existing formation planning methods cannot balance obstacle avoidance performance and formation stability, and improving the efficiency of autonomous obstacle avoidance and formation recovery of helicopter formations.
[0056] 3. This invention proposes a reinforcement learning fault-tolerant control framework based on an experience replay strategy. By calculating the cosine similarity between data samples, it implements an experience buffer update method based on data correlation. This method can efficiently extract fault-related information from the collected data, effectively eliminate temporal correlations between data, and ensure the persistence of reinforcement learning incentives. This experience buffer update method not only improves data utilization efficiency but also significantly enhances the fault tolerance capability of helicopter formation systems under fault scenarios. It solves the problems of low data utilization and poor fault adaptability in existing reinforcement learning fault-tolerant control methods, further improving the robustness and tracking accuracy of formation control. Attached Figure Description
[0057] Figure 1 This is a flowchart of the present invention;
[0058] Figure 2 This is a topology diagram of a helicopter formation system;
[0059] Figure 3 It is a three-dimensional relative position diagram of the helicopter and the obstacles;
[0060] Figure 4 This is a diagram showing the formation spacing error.
[0061] Figure 5 This is a trajectory tracking error diagram of a healthy helicopter;
[0062] Figure 6 This is a trajectory tracking error diagram of a malfunctioning helicopter. Detailed Implementation
[0063] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0064] like Figure 1 The present invention provides a helicopter cluster hierarchical intelligent safety planning method based on data correlation sampling, specifically including the following steps:
[0065] Step 1: Establish a dynamic model of the unmanned helicopter, considering the impact of position, velocity, main rotor force, attitude angle, and main rotor failure on system performance; and use an appropriate fault model to describe the impact of actuator failure and environmental disturbance.
[0066] Consider by A helicopter formation system consisting of several helicopters is established using a helicopter position loop model as follows:
[0067] (1)
[0068] (2)
[0069] in, , Representing inertial frames of reference respectively Next The position and velocity vectors of the helicopter. Indicates the first The mass of a helicopter; Indicates the first The expected forces acting on a helicopter, This represents the lift of the main rotor, and T represents the matrix transpose. Represents the gravitational acceleration vector. Indicates the first External disturbance to the helicopter. Represents the coordinate system of the machine body to inertial coordinate system The transformation matrix is specifically represented as:
[0070] (3)
[0071] in, , , These are the helicopter's pitch angle, roll angle, and yaw angle, respectively. , , .
[0072] During helicopter flight, the aerodynamic efficiency of the main rotor may decrease due to factors such as metal fatigue fracture or limited collective pitch control. A helicopter fault model for this type of failure can be represented as follows:
[0073] (4)
[0074] (5)
[0075] in, , These represent the helicopter's position and speed under main rotor failure. This indicates the change in force caused by a main rotor malfunction. This indicates the efficiency of the helicopter rotor.
[0076] Step 2: Using a finite-time planning framework, the formation spacing is adjusted using a greedy algorithm, while considering speed and acceleration constraints, to ensure that the system can avoid obstacles and maintain stable flight in the event of actuator failure.
[0077] To quantitatively describe the local formation planning process, the safe formation spacing planning problem is described as a constrained optimization problem. Decision variables Set as the formation spacing under null projection The rate of change. Represents a vector under constraints A definite projection on the null space. This represents the expected formation spacing vector for all helicopters. Let represent the desired formation spacing vector for the i-th helicopter. To enhance overall formation stability and prevent the cluster from deviating excessively from the desired formation while avoiding obstacles, the optimization problem is designed as follows:
[0078] (6)
[0079]
[0080] in, This represents the optimization index function. Indicates a finite forecast time window; This indicates the magnitude of change in the current control input compared to the previous step; The long-term formation deviation within the future time window t represents the smoothness and stability of the control decision. This indicates the deviation between the current formation and the initial desired formation; , and These are the weighting coefficients for the three evaluation indicators mentioned above. This represents the time integral term. This represents the basis for mapping null space vectors back to the original space. Indicates the location of the obstacle. Indicates the safe separation distance between the helicopter and the obstacle. Indicates the collision-free distance between helicopters. and These represent the velocity and acceleration of the change in formation spacing, respectively. This indicates the maximum normal helicopter speed. This represents the maximum acceleration of a normal helicopter. This indicates the maximum speed of the malfunctioning helicopter. This indicates the maximum acceleration of the malfunctioning helicopter.
[0081] To solve the proposed finite-time-window-constrained optimization model in real time and efficiently, an iterative greedy search strategy was adopted. The algorithm process is shown in Algorithm 1 in Table 1.
[0082] Table 1. Finite-range greedy formation planning algorithm
[0083]
[0084] Step 3 proposes a distributed formation coordination algorithm, which enables each helicopter to independently adjust its own state and share system matrix information, thereby ensuring that all UAVs can adjust their formation according to the safe formation spacing obtained in Step 2 in a dynamic environment, and generate the desired reference trajectory to ensure that the UAVs can effectively avoid obstacles and restore formation.
[0085] To ensure that all helicopters in the formation can reach a consensus and form the required spacing, the objective of setting up a distributed formation coordinator is:
[0086] (7)
[0087] in, , They represent the helicopter at time t. and Current expected position Helicopter at time t and The expected relative distance between them Indicates helicopter The relative position with respect to a fixed reference point. For each helicopter, the proposed distributed formation coordinator can be represented as:
[0088] (8)
[0089] (9)
[0090] in, This represents the system matrix, primarily used to adjust the state of the helicopter's reference trajectory. Indicates topological connectivity strength; for ease of description, use express The abbreviation of . , These are the system matrix adjustment gain and the system state adjustment gain, respectively, used to adjust the system's convergence performance. ,in This represents the coordination interval that can be directly invoked by the coordinator. This represents the distance transformation matrix, used to define the transformation relationship between the planned distance and the coordinated distance (the coordinated spacing of helicopters refers to the inter-aircraft spacing mapped to the coordinator; due to the mapping relationship between the coordinator and the actual spacing, it needs to be transformed through...). This transformation ensures that the coordinator can convert the desired spacing to a cooperative spacing that the coordinator can directly call upon. For ease of subsequent explanation, the formation spacing error is defined as... .
[0091] For any positive number System matrix It will converge at an exponential rate. To achieve consistency, where Laplace matrix The smallest non-zero eigenvalue, and Define this consensus as , Let represent the system matrix at time t. The spatial centroid of the initial formation position is chosen as the origin of the coordinate system. For any initial state... In accordance with the communication topology diagram Includes a spanning tree and an initial system matrix for each helicopter. It is a symmetric matrix and satisfies the Hurwitz stability criterion and conditions. In the case of the generated reference trajectory It will converge to the desired formation spacing distribution.
[0092] The following provides a proof that the designed distributed formation coordinator can achieve consistent convergence.
[0093] To simplify the symbols, let , as well as The system can then be written in the following compact form.
[0094] (10)
[0095] in, Represents the Kronecker product. For dimension The identity matrix, express The dimension. Let ,because The rate of change is much smaller than The system can be transformed into:
[0096] (11)
[0097] Further definition ,in , Then we can get:
[0098] (12)
[0099] For the nonhomogeneous differential equation of formula (12), its solution can be expressed as:
[0100] (13)
[0101] in For homogeneous solutions, This is a particular solution. A homogeneous solution can be obtained that satisfies... And under the conditions Down, ,in ,and Converging to a bounded vector To obtain a particular solution, the method of variation of constants is used. Let... Substituting into formula (12), we get:
[0102] (14)
[0103] eliminate Then, after further simplification, we get:
[0104] (15)
[0105] Therefore, we can conclude that:
[0106] (16)
[0107] special solution It can be represented as:
[0108] (17)
[0109] (18)
[0110] because You can get It is a symmetric negative definite matrix. Meanwhile... ,therefore It is also a symmetric negative definite matrix. Furthermore, due to the Laplace matrix... It is a symmetric positive semi-definite matrix, therefore and All are Hurwitz matrices. Because... For Hurwitz matrices, .because Also a Hurwitz matrix, we can obtain:
[0111] (19)
[0112] Therefore, combining and ,when In order to satisfy Under the given conditions, then:
[0113] (20)
[0114] because Therefore, we have:
[0115] (twenty one)
[0116] The final result is:
[0117] (twenty two)
[0118] in Distance transformation matrix A compact form of expression For dimension The identity matrix, For dimension The identity matrix. When it satisfies hour, lie in Therefore, a corresponding matrix must exist in the column space. , making This further illustrates:
[0119] (twenty three)
[0120] Therefore, the desired trajectory of the helicopter formation can converge to a safe and consistent position.
[0121] Step 4 employs an experience-replay-based reinforcement learning method to evaluate the state of the helicopter formation system and optimize the control strategy through a critic neural network. Cosine similarity calculation is introduced to update the experience pool, ensuring diversity and effectively improving fault tolerance. Reinforcement learning is used to learn and adapt to actuator failure scenarios, maintaining effective tracking of the reference trajectory.
[0122] To ensure that the formation can maintain stable trajectory tracking in complex environments and under fault conditions, a reinforcement learning method based on experience replay is proposed:
[0123] To achieve efficient experience replay, a filtering mechanism is introduced to selectively retain data samples based on the similarity between new data and existing experience. First, the cosine similarity between the newly added data and existing samples in the experience pool is calculated. Only when the similarity between the new data and existing data is sufficiently low is the data sample added to the experience pool, thus ensuring that the experience replay pool maintains a rich and diverse collection of experience samples. Newly generated data samples... With existing samples in the experience pool The cosine similarity between them is expressed as:
[0124] (twenty four)
[0125] If the cosine similarity between new data and any existing sample exceeds the historical maximum, the new data will be discarded; otherwise, in order to maintain the diversity of the experience pool, the historical sample most similar to the new data will be deleted.
[0126] To achieve effective trajectory tracking, we first define... This represents the positional error. To improve the response performance of the control system, a sliding mode term is introduced. This term incorporates both position and velocity error information. Therefore, to ensure that the reinforcement learning-based tracking controller can adaptively adjust its parameters to achieve the desired position of the i-th helicopter under varying degrees of rotor efficiency degradation, To achieve trajectory tracking and ensure that the tracking error converges to the desired range, while minimizing the tracking error and controlling energy consumption, an optimization objective function is constructed. for:
[0127] (25)
[0128] in, Represents the cost function, , These are the state weight matrix and the control weight matrix, respectively. Represents the sliding mode gain matrix. Used to simplify symbol representation.
[0129] To implement the control law based on reinforcement learning, the sliding surface Represented as
[0130] (26)
[0131] in, This represents an unknown state matrix, which contains uncertainties in the system, such as faults. and disturbance ; Given the control gain matrix, It represents the control input and is used as an optimization control term in reinforcement learning algorithms.
[0132] By defining the value function, the relevant Hamiltonian function can be derived as follows:
[0133] (27)
[0134] in, Indicators of performance Regarding sliding mode error variables The gradient of . Therefore, the optimal performance index function can be expressed as:
[0135] (28)
[0136] It satisfies the following Hamilton-Jacobi-Bellman (HJB) equations:
[0137] (29)
[0138] Based on the optimal control conditions, we can conclude that:
[0139] (30)
[0140] Based on this equation, the optimal control law can be derived as follows:
[0141] (31)
[0142] To evaluate the control performance, the critic neural network is designed as follows:
[0143] (32)
[0144] in It is the optimal weight matrix. It includes The activation function of a hidden neuron. This is the approximation error of the neural network. Next, Regarding the sliding mode error term The partial derivatives can be expressed as:
[0145] (33)
[0146] in , .
[0147] Due to the optimal weight matrix and approximation error gradient Since the weight matrix is not available, the estimated weight matrix is used. Approaching Thus, the estimated performance index gradient and the corresponding control law are derived:
[0148] (34)
[0149] (35)
[0150] Based on the principle of dynamic programming, the optimal Hamiltonian function can be derived to satisfy... However, due to the limitations of the commentator neural network in fitting the optimal performance metric... However, due to network approximation errors, the actually calculated Hamiltonian function may not satisfy the ideal zero-value condition. Correspondingly, the Hamiltonian function approximation error can be expressed as:
[0151] (36)
[0152] (37)
[0153] To enrich data diversity and reduce computational requirements, a cosine similarity-based data filtering strategy is employed for real-time training of the commentator neural network. The filtered data in the experience pool is defined as follows:
[0154] (38)
[0155] in, This indicates the number of experience points retained in the experience pool. One valid sample, This represents the maximum capacity of the experience pool. The corresponding Hamiltonian error can be expressed as:
[0156] (39)
[0157] (40)
[0158] To analyze the convergence of the Hamiltonian error, an expression for the error with empirical replay was established:
[0159] (41)
[0160] To ensure To ensure convergence, the weight update strategy for the commentator neural network is designed as follows:
[0161] (42)
[0162] in This is the weight adjustment coefficient. Then, the weight estimation error is defined as... and order , Based on the above equation, we can deduce that:
[0163] (43)
[0164] in , This is the residual.
[0165] For helicopter models that include faults and environmental disturbances, the aforementioned control law and weight update law can be designed to ensure position tracking error accuracy. It is uniform and eventually converges to a bounded state.
[0166] The following provides a proof that the reinforcement learning method based on experience replay designed above can achieve stable trajectory tracking.
[0167] The candidate form for Lyapunov functions is:
[0168] (44)
[0169] in , .
[0170] First, by analyzing the first part of the Lyapunov function, we can obtain:
[0171] (45)
[0172] Based on the relevant equations, we can further deduce:
[0173] (46)
[0174] definition We can obtain:
[0175] (47)
[0176] It can be further transformed into:
[0177] (48)
[0178] We can obtain the following by scaling Young's inequality:
[0179] (49)
[0180] Next, we will analyze the second term of the Lyapunov function:
[0181] (50)
[0182] According to Young's inequality, we can deduce that:
[0183] (51)
[0184] (52)
[0185] Therefore, we can obtain:
[0186] (53)
[0187] Combining formulas (51) and (52), we can deduce that:
[0188] (54)
[0189] Therefore, under the conditions and The following can be derived. ,in and They are defined as follows:
[0190] (55)
[0191] (56)
[0192] in, , , This represents the upper bound of the corresponding variable. Due to the sliding surface... It is UUB, therefore it corresponds to the sliding surface. Tracking error It's also from UUB.
[0193] In this embodiment, a simulation case is provided to verify the effectiveness of the proposed distributed formation coordination and commentator tracking control method in a helicopter formation system. The helicopter formation system consists of five helicopters, and its communication topology is as follows: Figure 2As shown, the formation lost connection with the ground station or leader. The fault occurred in... It was injected into helicopter 3 at that time, and the main rotor efficiency was when the failure occurred. All helicopters are subject to external environmental interference, with the interference level being [missing information]. The trajectory of dynamic obstacles is set as follows: .
[0194] For distributed formation coordination, the relevant parameter settings are as follows: Optimization metrics The weights are set as follows: , , Collision avoidance constraints are set to Obstacle avoidance constraints are set to The maximum speed and acceleration under normal and fault conditions are respectively , , , Initial system matrix The initial reference position of the helicopter is set to a randomized Hurwitz matrix. It coincides with its actual position, and the initial positions are set as follows: , , , , .
[0195] For the tracking controller, the relevant parameters are set as follows: State weight matrix Control weight matrix Sliding mode gain matrix Experience pool size The critic network consists of 21 neurons, with the activation function set to the radial basis function. The learning rate of the neural network is... .
[0196] Figure 3 The demonstration showed the dynamic adjustment process of the formation and the relative positions of the helicopters and obstacles. When an obstacle approaches, the helicopter formation can autonomously adjust its formation structure to avoid it, and after the obstacle is removed, the formation is reconfigured back to its original state. Figure 4 Formation spacing error exist The formation converged internally and then fluctuated only slightly as it evolved, ensuring real-time adjustments to the obstacle avoidance formation. Figure 5 and Figure 6 The trajectory tracking performance of five helicopters relative to a desired trajectory is demonstrated. It can be observed that the tracking error consistently remains within a certain range. Within this range, this verifies that the proposed controller can quickly update different flight modes through real-time learning, thereby ensuring adaptability in various flight scenarios. Even in scenarios such as... Figure 6 shown Even when the main rotor efficiency degrades, the proposed commentator network based on experience replay can still enhance the main rotor lift through autonomous learning, effectively mitigating the degradation of tracking performance and avoiding the risk of formation collapse due to malfunctions.
[0197] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A hierarchical intelligent safety planning method for helicopter clusters based on data correlation sampling, characterized in that, Including the following steps: S1, let it be... A helicopter formation system consisting of several helicopters is established, and a helicopter position loop model is created. Considering the failure of reduced aerodynamic efficiency of the rotor main wing, a helicopter failure model under this failure is created. S2, combining speed and acceleration constraints, constructs a safe formation spacing planning problem; a greedy algorithm is used to optimize and adjust the safe formation spacing to ensure that the helicopter formation system can safely avoid obstacles and fly stably in the event of a failure; S3. A distributed formation coordination algorithm is adopted to optimize the target of the distributed formation coordinator, so that each helicopter can independently adjust its own state and share the system matrix information, ensuring that all UAVs can adjust their formation according to the safe formation spacing obtained in step S2 in a dynamic environment and generate the desired reference trajectory. S4 employs an experience-based reinforcement learning method to evaluate the state of the helicopter formation system and optimize the control strategy through a critic neural network; it updates the experience pool samples using cosine similarity; and it adaptively optimizes the control law in actuator failure scenarios to ensure the helicopter tracks the reference trajectory.
2. The helicopter cluster hierarchical intelligent safety planning method based on data correlation sampling according to claim 1, characterized in that, The helicopter position ring model is as follows: , in, , Representing inertial frames of reference respectively Next The position and velocity vectors of the helicopter. Indicates the first The mass of a helicopter; Indicates the first The expected forces acting on a helicopter, This represents the lift of the main rotor, and T represents the matrix transpose. Represents the gravitational acceleration vector. Indicates the first External disturbances to the helicopter; Represents the coordinate system of the machine body to inertial coordinate system The transformation matrix is specifically represented as: , in, , , These are the helicopter's pitch angle, roll angle, and yaw angle, respectively. , , ; The helicopter fault model is represented as follows: , in, , These represent the helicopter's position and speed under main rotor failure. This indicates the change in force caused by a main rotor malfunction. This indicates the efficiency of the helicopter rotor.
3. The helicopter cluster hierarchical intelligent safety planning method based on data correlation sampling according to claim 1, characterized in that, The safe formation spacing planning problem is described as a constrained optimization problem, where the decision variables are... Set as the formation spacing under null projection rate of change, Represents a vector under constraints A definite projection on the null space. This represents the expected formation spacing vector for all helicopters. Let the desired formation spacing vector of the i-th helicopter be denoted by ; then the optimization problem is designed as follows: , , in, This represents the optimization index function. Indicates a finite forecast time window; This indicates the magnitude of change in the current control input compared to the previous step. This represents the long-term formation deviation within the future time window t. This indicates the deviation between the current formation and the initial desired formation. , and These are the weighting coefficients, Indicates the time integral term; This represents the basis for mapping null space vectors back to the original space; Indicates the location of the obstacle. Indicates the safe separation distance between the helicopter and the obstacle. Indicates the collision-free distance between helicopters. , These represent the velocity and acceleration of the change in formation spacing, respectively. This indicates the maximum normal helicopter speed. This represents the maximum acceleration of a normal helicopter. This indicates the maximum speed of the malfunctioning helicopter. This indicates the maximum acceleration of the malfunctioning helicopter.
4. The helicopter cluster hierarchical intelligent safety planning method based on data correlation sampling according to claim 1, characterized in that, In step S3, the target of the distributed formation coordinator is set as follows: , in, , They represent the helicopter at time t. and Current desired position; Helicopter at time t and The expected relative distance between them Indicates helicopter The relative position with respect to a fixed reference point; for each helicopter, the proposed distributed formation coordinator can be represented as: , in, This represents the system matrix, used to adjust the state of the helicopter reference trajectory; Indicates the topology connection strength; for abbreviation; , These are the system matrix adjustment gain and the system state adjustment gain, respectively, used to adjust the system's convergence performance; This represents the coordination interval that can be directly invoked by the coordinator, where The distance transformation matrix is used to define the conversion relationship between the planned distance and the cooperative distance; the formation spacing error is defined as... .
5. The helicopter cluster hierarchical intelligent safety planning method based on data correlation sampling according to claim 1, characterized in that, In step S4, data samples are selectively retained based on the cosine similarity between the new data and existing samples in the experience pool; newly generated data samples... With existing samples in the experience pool Cosine similarity between Represented as: , Where T represents the matrix transpose; If the cosine similarity between new data and any existing sample exceeds the historical maximum, the new data will be discarded; otherwise, the historical sample most similar to the new data will be deleted.
6. The helicopter cluster hierarchical intelligent safety planning method based on data correlation sampling according to claim 2, characterized in that, In step S4, define To account for positional error, a sliding mode term is introduced. The desired position of the i-th helicopter is achieved by constructing an optimization objective function. Trajectory tracking; with the goal of minimizing tracking error and controlling energy consumption, an optimization objective function is constructed. for: , in, Represents the cost function, , These are the state weight matrix and the control weight matrix, respectively. Represents the sliding mode gain matrix. ; sliding surface Represented as: , in, This represents an unknown state matrix, which contains faults in the system. and disturbance ; Given the control gain matrix, It represents the control input and serves as an optimization control term in reinforcement learning algorithms; The control law is designed as follows: , The network of design critics is: , Where T represents the matrix transpose; , It includes Activation functions of hidden neurons; Represents the estimated weight matrix; It is the optimal weight matrix. It is the approximation error of the neural network; The weight update law for designing a commentator neural network is as follows: : , , , , , in, It is the weighting adjustment factor; For helicopter models that include faults and environmental disturbances, the aforementioned control law and weight update law ensure that the position tracking error is maintained. It is uniform and eventually converges to a bounded state.