Autonomous obstacle avoidance method and system for low-altitude intelligent dynamic monitoring aircraft
Through the layered reinforcement learning framework and Bezier curve optimization technology, the problems of obstacle avoidance and cruise planning of low-altitude aircraft clusters in complex environments are solved, and the global optimization control and stable operation of the cluster are achieved.
Patent Information
- Application Number
- CN202510308509.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for low-altitude aircraft clusters to achieve autonomous obstacle avoidance and cruise area planning in complex environments, especially in unknown time-varying wind-hacking environments, traditional path planning and obstacle avoidance methods are difficult to ensure flight safety and mission effectiveness.
The hierarchical reinforcement learning framework is adopted to decompose the low-altitude aircraft cluster control problem into two levels: high-level task allocation and low-level trajectory execution. The optimal detection route is generated through the Bezier curve and genetic algorithm optimization, and the patrol airspace is determined based on ellipse fitting, and a distributed control strategy based on timing differential learning and policy gradient optimization is designed.
The global optimization control of low-altitude aircraft cluster is realized, ensuring stable operation in unknown wind-shocking environments, reducing system complexity, improving algorithm scalability and computing efficiency, improving the rationality of airspace division and the real-time and robustness of the system.
Smart Images

Figure CN120085671A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of low-altitude aircraft obstacle avoidance, and particularly to an autonomous obstacle avoidance method and system for low-altitude intelligent dynamic monitoring aircraft. Background Art
[0002] The low-altitude aircraft cluster technology has been widely applied in fields such as patrol. However, the autonomous obstacle avoidance and patrol airspace planning of low-altitude aircraft clusters in complex environments still face many challenges. Especially under the influence of unknown time-varying wind disturbances, traditional path planning and obstacle avoidance methods are difficult to ensure flight safety and mission effectiveness.
[0003] Currently, the control methods for low-altitude aircraft clusters mainly adopt a centralized architecture. This method has problems such as high communication pressure and poor real-time performance when facing large-scale cluster cooperative tasks. At the same time, existing obstacle avoidance algorithms are mostly based on the assumption of a deterministic environment, difficult to cope with the dynamically changing wind disturbance environment, and lack overall consideration of the multi-aircraft cooperation performance, which easily leads to local optimal solutions. Summary of the Invention
[0004] This application provides an autonomous obstacle avoidance method and system for low-altitude intelligent dynamic monitoring aircraft, thereby realizing the global optimal control of low-altitude aircraft clusters and ensuring the stable operation of low-altitude aircraft clusters in an unknown wind disturbance environment.
[0005] In the first aspect of this application, an autonomous obstacle avoidance method for low-altitude intelligent dynamic monitoring aircraft is provided. The autonomous obstacle avoidance method for low-altitude intelligent dynamic monitoring aircraft includes: Construct a state vector for the position coordinates, flight speed, and heading angle of the low-altitude aircraft cluster, and perform kinematic modeling in combination with obstacle position data and wind speed components to obtain the state space of the low-altitude aircraft cluster; Perform Bezier curve calculation and genetic optimization on the control point sequence in the state space of the low-altitude aircraft cluster to obtain the optimal detection route parameters; Perform geometric distance calculation and two-point removal processing according to the waypoint sequence of the optimal detection route to obtain the patrol airspace equation; Input the state space of the low-altitude aircraft cluster, the optimal detection route parameters, and the patrol airspace equation into a hierarchical reinforcement learning framework for multi-level reward function calculation to obtain high and low-level decision-making networks; Based on the high and low-level decision-making networks, perform temporal difference learning and policy gradient optimization to obtain a distributed control strategy.
[0006] In the second aspect of this application, an autonomous obstacle avoidance system for low-altitude intelligent dynamic monitoring aircraft is provided. The autonomous obstacle avoidance system for low-altitude intelligent dynamic monitoring aircraft includes: A modeling module, which is used to construct a state vector of the position coordinates, flight speed, and heading angle of a low-altitude aircraft cluster, and perform kinematic modeling by combining obstacle position data and wind speed components to obtain the state space of the low-altitude aircraft cluster; A calculation module, which is used to perform Bezier curve calculation and genetic optimization on the control point sequence in the state space of the low-altitude aircraft cluster to obtain the optimal detection route parameters; A processing module, which is used to perform geometric distance calculation and double-point removal processing according to the waypoint sequence of the optimal detection route to obtain a patrol airspace equation; A reinforcement learning module, which is used to input the state space of the low-altitude aircraft cluster, the optimal detection route parameters, and the patrol airspace equation into a hierarchical reinforcement learning framework for multi-level reward function calculation to obtain a high-level and low-level decision-making network; An optimization module, which is used to perform temporal difference learning and policy gradient optimization based on the high-level and low-level decision-making network to obtain a distributed control strategy.
[0007] Compared with the prior art, the present application has the following beneficial effects: By constructing a hierarchical reinforcement learning framework, the low-altitude aircraft cluster control problem is decomposed into two levels: high-level task allocation and low-level trajectory execution, reducing the system complexity and improving the scalability and computational efficiency of the algorithm. The optimal detection route is generated by combining Bezier curves with genetic algorithm optimization, and the patrol airspace is determined by the ellipse fitting method, making the planning result more in line with the actual flight requirements. A distributed control strategy based on temporal difference learning and policy gradient optimization is designed to achieve unified optimization of multi-aircraft cooperation performance and single-aircraft control performance. By introducing an adaptive compensation controller and a Kalman filter, the path following problem under unknown time-varying wind disturbances is solved, ensuring the stability of the system. The double-point removal algorithm based on geometric distance is adopted to effectively avoid the influence of abnormal waypoints on the patrol airspace planning and improve the rationality of airspace division. Through the design of a multi-level reward function, multiple objectives such as collision avoidance, task completion degree, and energy efficiency are considered uniformly, realizing the global optimal control of the low-altitude aircraft cluster. A wind disturbance compensation controller is designed based on the Lyapunov stability theory, and the parameters are adjusted in combination with the global cooperation index to ensure the stable operation of the cluster system in an unknown wind disturbance environment. The control strategy is implemented by a distributed deployment method, reducing the communication burden and improving the real-time performance and robustness of the system. Description of the Drawings
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0009] The structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those skilled in this technology to understand and read, and are not used to limit the conditions under which the present invention can be implemented. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0010] Figure 1 It is a schematic flowchart of the autonomous obstacle avoidance method of the low-altitude intelligent dynamic monitoring aircraft provided by an embodiment of the present invention; Figure 2 It is a schematic block diagram of the structure of the autonomous obstacle avoidance system of the low-altitude intelligent dynamic monitoring aircraft provided by an embodiment of the present invention. Detailed implementation manners
[0011] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0012] The flowchart shown in the drawings is only an example illustration, and does not necessarily include all the content and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may change according to the actual situation.
[0013] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0014] It should be further understood that the term " / and" as used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Please refer to Figure 1 , an embodiment of the autonomous obstacle avoidance method of the low-altitude intelligent dynamic monitoring aircraft in the embodiments of this application includes: Step 100: Construct a state vector for the position coordinates, flight speed, and heading angle of the low-altitude aircraft cluster, and perform kinematic modeling by combining the obstacle position data and wind speed components to obtain the state space of the low-altitude aircraft cluster; It can be understood that the execution entity of this application can be the autonomous obstacle avoidance system of a low-altitude intelligent dynamic monitoring aircraft, or it can also be a terminal or a server. Specifically, it is not limited here. In this embodiment of the application, the server is taken as an example of the execution entity for illustration.
[0015] Specifically, the two-dimensional plane coordinate values, flight speed values, and heading angle values of a single low-altitude aircraft in the cluster are combined to form a single-aircraft state vector. The single-aircraft state vectors of each low-altitude aircraft are arranged sequentially in the form of a matrix, and these state vectors are spliced to generate a cluster state matrix that comprehensively describes the dynamic state of the low-altitude aircraft cluster, reflecting the spatial distribution, motion characteristics, and heading information of all low-altitude aircraft in the cluster. The obstacle information in the obstacle avoidance environment of the low-altitude aircraft is processed to construct a static obstacle set. The spatial position coordinates and range parameters of each obstacle are extracted, and these data are used for set analysis of the obstacles to construct a static obstacle set containing m obstacles, describing the spatial distribution and size of the obstacles in the mission area. At the same time, the horizontal wind speed component and vertical wind speed component at the current moment are extracted and combined into a two-dimensional vector to construct a time-varying wind disturbance vector. In order to ensure that the mission execution boundary of the low-altitude aircraft is clear, the minimum coordinate values and maximum coordinate values in the horizontal and vertical directions of the specified mission area are extracted to generate a mission airspace boundary vector. The cluster state matrix and the time-varying wind disturbance vector are substituted into the differential equation for numerical solution to establish the kinematic state transition equation of the low-altitude aircraft cluster. This differential equation describes the dynamic changes of the low-altitude aircraft under the influence of the external environment and provides the ability to predict the state over time series. Through the numerical solution method, the evolution law of the state of the low-altitude aircraft cluster over time is obtained. The Euclidean distance is calculated for the position data of any two low-altitude aircraft in the cluster state matrix and compared with the preset risk threshold to obtain a collision risk metric function, which quantifies the collision risk between low-altitude aircraft in real time and provides a direct basis for dynamically adjusting the flight strategy. At the same time, in order to ensure that the low-altitude aircraft can safely avoid obstacles, for each low-altitude aircraft in the cluster state matrix and each obstacle in the static obstacle set, the minimum distance is calculated respectively and evaluated in combination with the safety margin to obtain an obstacle risk metric function. Based on the above kinematic state transition equation, collision risk metric function, and obstacle risk metric function, a state space of the low-altitude aircraft cluster is created.
[0016] Step 200: Perform Bezier curve calculation and genetic optimization on the control point sequence in the state space of the low-altitude aircraft cluster to obtain the optimal detection route parameters; Specifically, numerical integration operations are performed on the kinematic state transition equation in the state space of the low-altitude aircraft cluster. By discretizing the state transition equation and using numerical integration methods (such as the trapezoidal method or the Runge-Kutta method) for solution, the motion trajectory information of the low-altitude aircraft during flight is obtained, and the initial control point set is extracted. The control point set forms an initial path in space. The initial control point set is screened and adjusted according to the risk assessment mechanism. By combining the collision risk metric function and the obstacle risk metric function, the safety of each control point in the actual path is evaluated. For example, for each control point, the distances between it and neighboring low-altitude aircraft and environmental obstacles are evaluated. If it is found that some control points have potential collision risks or the possibility of deviating from the safety margin, they are marked and removed to ensure that the remaining control point set forms a candidate control point sequence. Based on the candidate control point sequence, the parametric equation of the Bezier curve is constructed. For each control point, the binomial coefficients associated with it are calculated. By matrix processing these basis functions, a basis function matrix of an nth-order Bezier curve is constructed. The basis function matrix is multiplied by the control point sequence matrix to generate the corresponding parametric equation of the Bezier curve, describing the motion characteristics of the low-altitude aircraft on the path. The advantage of the Bezier curve lies in its natural smoothness and flexibility. By adjusting the control points, the shape of the path can be changed, thereby realizing the dynamic optimization of the flight route. The Bezier curve is evaluated and optimized. By calculating the weighted sum of the route length, curve smoothness, and detection efficiency, an objective optimization function is constructed, which comprehensively considers multiple important indicators of path planning. For example, the shorter the route length, the higher the flight efficiency; the smoother the curve, the more stable the motion control of the low-altitude aircraft; the level of detection efficiency directly affects the time and effect of task completion. By adjusting the weighting coefficients, different priorities are assigned to different indicators according to the actual task requirements. The constructed control point sequence is chromosomally encoded, where each chromosome represents the coordinate value of a control point, and these chromosomes are randomly initialized to generate an initial population. The initial population serves as the starting point of the algorithm, providing a diverse solution space for the entire optimization process. During the optimization process, classical operations in genetic algorithms are adopted, including roulette wheel selection, uniform crossover, and Gaussian mutation operations. Roulette wheel selection makes a probabilistic selection based on the fitness values of individuals, ensuring that individuals with high fitness have a greater chance of survival and reproduction; uniform crossover generates new individuals by randomly exchanging gene segments of chromosomes, increasing the diversity of the population; Gaussian mutation explores local regions in the solution space by making small random adjustments to the genes of individuals, thus avoiding falling into local optimal solutions. Through multiple iterative optimizations, the individuals in the population gradually converge to the global optimal solution, generating an optimized control point sequence. The optimized control point sequence is substituted into the parametric equation of the Bezier curve for path reconstruction and generation, obtaining the final optimal detection route parameters.
[0017] Step 300: Calculate the geometric distances and perform two-point removal processing based on the waypoint sequence of the optimal detection route to obtain the patrol airspace equation; It should be noted that through parameter discretization processing and coordinate calculation of the optimal detection route parameters, an initial waypoint set is obtained. These waypoints are generated by uniformly discretizing the route parameters, and the coordinates of each point precisely correspond to the position of the detection route in the discretization interval. On this basis, calculate the curvature of adjacent waypoints in the initial waypoint set to evaluate the geometric change characteristics of the path. The curvature value can quantify the degree of bending of the route, and then perform a threshold judgment on these points to screen out the waypoint sequence that meets the task requirements. Calculate the mean value of the coordinate data of the waypoints and perform deviation matrix operations to extract the statistical characteristic parameters of the waypoints, including the central position and distribution range, reflecting the overall distribution trend of the waypoints in space. At the same time, in order to quantify the geometric relationship between waypoints, input the waypoint sequence into the Euclidean distance calculation module to calculate the distance between each two points and generate a distance matrix between waypoints. The distance matrix contains the relative position relationships between all waypoints. Perform statistical analysis on the distance matrix to judge the distribution of abnormal points. Based on the principle of k times the standard deviation, dynamically calculate the distance threshold between waypoints to obtain the abnormal point discrimination criterion. This criterion is set according to the normal distribution characteristics of waypoints. When the distance between two waypoints deviates from the normal distribution range (exceeds k times the standard deviation), it is determined as an abnormal point pair. Based on this discrimination criterion, iteratively screen and recursively remove the abnormal points in the waypoint sequence. For the point pairs that do not meet the criterion, gradually eliminate the abnormal points among them, so that the optimized waypoint sequence can be smoother and avoid the interference of redundant or error points. After the optimization is completed, model the geometric characteristics of the waypoints. By applying principal component analysis to the optimized waypoint sequence, extract its dominant characteristics in space. Principal component analysis determines the main axis direction and the ratio of the major and minor axes of the waypoint distribution through eigenvalue decomposition. This result describes the basic shape and direction of the distribution ellipse of the waypoints. Combine the basic parameters of the ellipse and the statistical characteristic parameters of the waypoints, substitute them into the standard equation of the quadratic curve, and construct the patrol airspace equation. The quadratic curve equation represents the distribution characteristics of these points in a mathematical form by transforming the waypoint coordinates and calculating the coefficients, and provides a clear patrol airspace range for the flight mission of the low-altitude aircraft. Obtain the patrol airspace equation.
[0018] Step 400: Input the low-altitude aircraft cluster state space, the optimal detection route parameters, and the patrol airspace equation into the hierarchical reinforcement learning framework for multi-level reward function calculation to obtain the high-level and low-level decision-making networks; Specifically, numerical integration operations are carried out using the kinematic state transition equation of the low-altitude aircraft, and combined with the Bessel curve parameters, the flight path of the low-altitude aircraft is predicted to provide the position information of the low-altitude aircraft at future moments, forming the position component in the high-level state space. At the same time, the collision risk metric function and the obstacle risk metric function are analyzed and matched with the patrol airspace equation to determine the task component of the low-altitude aircraft within a specific airspace. The high-level state space contains two parts of information: one is the predicted position of the low-altitude aircraft, and the other is the spatial constraint information related to the task. Based on the position component and the task component of the high-level state space, a task assignment decision matrix is constructed to describe the relevance and priority between the low-altitude aircraft and the task objectives. At the same time, the route planning parameters are discretized, and the continuous route planning problem is transformed into a computable high-level action space matrix. This matrix defines the set of actions that the low-altitude aircraft may perform in high-level decision-making, such as path adjustment and task priority assignment. The completion degree of the cooperative detection task, the route following deviation, and the energy consumption are weighted and calculated to generate the high-level reward function value. By reasonably designing the weights, the multi-objective requirements are balanced, ensuring both task efficiency and optimizing flight energy consumption and path accuracy. While constructing the high-level part, a low-level state space is constructed for the local control requirements of a single low-altitude aircraft. The state data of a single low-altitude aircraft is spatially mapped with the local obstacle information around it, and combined with the time-varying wind disturbance vector, the coordinates are transformed to generate the low-level state space vector, reflecting the state of the low-altitude aircraft in the local environment and external influences. Based on the low-level state space, a speed command set and a heading command set are constructed, and these commands define the basic actions that the low-altitude aircraft can perform in the local environment. By discretizing these actions, a low-level action space matrix is generated to describe the control strategy of the low-altitude aircraft on a short time scale. To evaluate the effectiveness of the low-level decision-making, multi-index weighted calculations are carried out on the flight path following accuracy, the obstacle avoidance safety distance, and the maneuver command amplitude to form the low-level reward function value. The low-level reward function aims to optimize the safety and flexibility of the low-altitude aircraft in the local environment while ensuring its response accuracy to high-level decisions. The high-level state space, the low-level state space, their corresponding action spaces, and the reward functions are input into a deep neural network for training. During the training process, a two-layer architecture of a policy network and a value network is adopted, where the policy network is used to learn the decision-making strategy, that is, to select the optimal action according to the state, and the value network is used to estimate the long-term benefits of each state. By combining high-level and low-level decisions, a high-low level decision network is formed, where the high level is responsible for global task planning and path optimization, while the low level focuses on local control and real-time obstacle avoidance. The hierarchical structure can effectively meet the multi-objective task requirements in complex environments and achieve global and local collaborative optimization at the same time. Through the hierarchical design of the deep reinforcement learning framework, the high-low level decision network dynamically adapts to environmental changes and provides a complete solution for the low-altitude aircraft from global task planning to local fine control.
[0019] Step 500: Perform temporal difference learning and policy gradient optimization based on the high-level and low-level decision-making networks to obtain a distributed control strategy.
[0020] Specifically, the current low-altitude aircraft cluster state is input into the Actor network in the high-low level decision-making network to calculate the high-level task allocation strategy. By analyzing the overall state of the low-altitude aircraft cluster, including position, speed, task requirements, and environmental information, etc., high-level task allocation instructions are generated to clarify the specific tasks that each low-altitude aircraft in the low-altitude aircraft cluster needs to execute, decoupling complex global tasks into subtasks suitable for a single low-altitude aircraft to execute. According to the high-level task allocation instructions, the tasks of the low-altitude aircraft cluster are decomposed to generate a single-aircraft execution instruction sequence for each low-altitude aircraft. The single-aircraft execution instruction sequence is input into the Actor network in the high-low level decision-making network to calculate the specific action strategy of each low-altitude aircraft in the local environment. The calculation of the low-level control instructions needs to comprehensively consider information such as local obstacle distribution, wind disturbance influence, and task target position to generate an action plan for the low-altitude aircraft on a short time scale. These instructions include specific operations such as adjusting flight speed, changing the heading angle, and avoiding obstacles. Temporal difference calculation is performed on the low-level control instructions to evaluate the performance difference of the instructions at consecutive moments. The temporal difference method gradually adjusts the policy parameters by comparing the difference between the predicted value function and the actual reward. Subsequently, the calculation results are input into the Critic network in the high-low level decision-making network for value evaluation to obtain the value of the current state and action, measuring the effectiveness of the low-level control instructions in the local environment, including the accuracy of track following, the safety of obstacle avoidance, and the execution efficiency of actions, etc. Based on the value evaluation results, policy gradient optimization is performed on the low-level control instructions. Through optimization, the parameters of the low-level control strategy are adjusted to improve it in the direction of maximizing the value, generating a locally optimal control strategy. The locally optimal control strategy ensures that a single low-altitude aircraft can efficiently complete the task objective within its local range while avoiding conflicts with obstacles or other low-altitude aircraft. The locally optimal control strategy is input into the Critic network for global value evaluation. Global value evaluation not only considers the local performance of a single low-altitude aircraft but also integrates the overall synergy effect of the low-altitude aircraft cluster. The evaluation indicators include task completion rate, the efficiency of collaboration between clusters, and the balance of global task execution, etc. According to the global synergy indicators, policy gradient correction is performed on the high-level task allocation instructions to optimize the task allocation among low-altitude aircraft, maximizing the task completion efficiency while ensuring the overall coordination of the low-altitude aircraft cluster. The corrected optimal task allocation strategy is combined with the locally optimal control strategy and deployed in a distributed manner to obtain the distributed control strategy of the low-altitude aircraft cluster. Through distributed deployment, each low-altitude aircraft in the low-altitude aircraft cluster independently executes its optimized control strategy while maintaining collaborative consistency under the guidance of the global task. Decompose and calculate the speed control command and the heading control command in the local optimal control strategy. By analyzing the control strategy of the low-altitude aircraft, refine the global command of the high-level planning into the nominal control input of each low-altitude aircraft. The nominal control input is a representation of the ideal motion behavior of the low-altitude aircraft, including speed components and heading angle information, providing a benchmark for the subsequent compensation control design. Combine the nominal control input with the current kinematic state of the low-altitude aircraft as the basic variables for dynamic modeling. Substitute the nominal control input and the kinematic state transition equation of the low-altitude aircraft into the Lyapunov function to evaluate the stability of the system. The Lyapunov function is the core tool for the stability analysis of the control system. Its form is selected as a quadratic function of the state variables. By calculating the derivative of this function, the stability conditions of the system are obtained. These stability conditions reflect the dynamic behavior of the system under external disturbances and can provide a theoretical basis for the design of wind disturbance compensation. On this basis, design an adaptive compensation term for the time-varying wind disturbance vector. The adaptive compensation generates an initial compensation control law by adjusting the parameters of the compensation law in real time, aiming to offset the influence of the wind disturbance on the motion of the low-altitude aircraft and ensure that the system meets the Lyapunov stability conditions. Construct a collaborative compensation term according to the optimal task allocation strategy. The collaborative compensation term is oriented towards the overall task objective of the low-altitude aircraft cluster. Combining the collaborative relationship and global constraints among multiple low-altitude aircraft, introduce a task consistency adjustment mechanism for the compensation control. Combine the collaborative compensation term with the initial compensation control law to generate an adaptive compensation controller. The adaptive compensation controller significantly improves the task execution ability and coordination of the cluster in a complex environment by dynamically adjusting the control parameters of each low-altitude aircraft. After generating the adaptive compensation controller, design a sliding mode surface for the control gain. Sliding mode control is a control method with robustness and anti-interference ability. By dynamically adjusting the controller gain, it realizes a fast response and high-precision control for external disturbances. Through the design of the sliding mode surface, combine the stability and adaptive ability of the controller to improve the performance of the system. At the same time, optimize and adjust the compensation control parameters in combination with the global collaboration index to balance the relationship between local stability and global collaboration and ensure the optimal overall performance of the low-altitude aircraft cluster. On the basis of optimizing the controller parameters, substitute the GPS position data of a single low-altitude aircraft and the nominal control input into the Kalman filter to estimate the wind disturbance state quantity. The Kalman filter is a classic state estimation tool. By weighted fusion of the system's observation data and prediction model, it filters out noise in real time and accurately estimates the wind disturbance suffered by the low-altitude aircraft. The wind disturbance state quantity is the key input variable for the compensation control, directly affecting the design and execution of the compensation control law. Calculate the adaptive law for the wind disturbance state quantity and the compensation control parameters to generate a real-time compensation control quantity. The core of the adaptive law calculation is to adjust the compensation parameters in real time according to the environmental changes, enabling the compensation controller to dynamically adapt to complex environmental changes.The nominal control input and the real-time compensation control amount are weighted and combined to generate the wind disturbance compensation control amount for each low-altitude aircraft. The wind disturbance compensation control amount is a dynamic correction to the nominal control input, which ensures that the low-altitude aircraft can perform tasks according to the ideal trajectory by offsetting the influence of external disturbances.
[0021] In the embodiments of this application, by constructing a hierarchical reinforcement learning framework, the low-altitude aircraft cluster control problem is decomposed into two levels: high-level task allocation and low-level trajectory execution, reducing the system complexity and improving the scalability and computational efficiency of the algorithm. The optimal detection route is generated by combining Bezier curves with genetic algorithm optimization, and the patrol airspace is determined by the ellipse fitting method, making the planning results more in line with the actual flight requirements. A distributed control strategy based on temporal difference learning and policy gradient optimization is designed to achieve the unified optimization of multi-aircraft cooperation performance and single-aircraft control performance. By introducing an adaptive compensation controller and a Kalman filter, the path following problem under unknown time-varying wind disturbances is solved, ensuring the stability of the system. The double-point removal algorithm based on geometric distance is adopted to effectively avoid the influence of abnormal waypoints on the patrol airspace planning and improve the rationality of airspace division. Through the design of a multi-level reward function, multiple objectives such as collision avoidance, task completion rate, and energy efficiency are considered together to achieve the global optimal control of the low-altitude aircraft cluster. The wind disturbance compensation controller is designed based on the Lyapunov stability theory, and the parameters are adjusted in combination with the global cooperation index to ensure the stable operation of the cluster system in an unknown wind disturbance environment. The control strategy is implemented in a distributed deployment manner, reducing the communication burden and improving the real-time performance and robustness of the system.
[0022] In a specific embodiment, the process of executing step 100 may specifically include the following steps: The two-dimensional plane coordinate values, flight speed values, and heading angle values of a single low-altitude aircraft in the low-altitude aircraft cluster are combined to obtain a single-aircraft state vector, and the single-aircraft state vectors are sorted in sequence and matrix-concatenated to obtain a cluster state matrix; The spatial position coordinates and range parameters of each obstacle are extracted and analyzed to obtain a static obstacle set composed of m obstacles; A two-dimensional vector of the horizontal wind speed component and the vertical wind speed component at the current moment is constructed to obtain a time-varying wind disturbance vector, and the minimum and maximum coordinate values in the horizontal and vertical directions of the specified task area are extracted to obtain a task airspace boundary vector; The cluster state matrix and the time-varying wind disturbance vector are substituted into the differential equation for numerical solution to obtain a kinematic state transition equation; Calculate the Euclidean distance and compare it with the risk threshold for the position data of any two low-altitude aircraft in the cluster state matrix to obtain the collision risk metric function. Also, calculate the minimum distance and evaluate the safety margin for each low-altitude aircraft in the cluster state matrix and each obstacle in the static obstacle set to obtain the obstacle risk metric function; Create the state space of the low-altitude aircraft cluster based on the kinematic state transition equation, the collision risk metric function, and the obstacle risk metric function.
[0023] Specifically, characterize the state of a single low-altitude aircraft in the low-altitude aircraft cluster. The state of each low-altitude aircraft is represented by two-dimensional plane coordinates indicating its position in the mission area, by velocity representing its current flight speed, and by the heading angle indicating its deflection angle relative to the due north direction. Combine this information into a single-aircraft state vector , where and are the horizontal and vertical coordinates of the -th low-altitude aircraft, is its velocity, is its heading angle. Assume there are low-altitude aircraft in the cluster. By arranging and concatenating each single-aircraft state vector in order of number, construct a cluster state matrix : : ; This matrix completely describes the spatial distribution, motion characteristics, and direction information of the low-altitude aircraft cluster. At the same time, extract and analyze the obstacle information in the mission environment. The position of each obstacle is represented by three-dimensional coordinates , where is the height range of the obstacle in the vertical direction. The geometric range of the obstacle is represented by the radius , applicable to spherical or cylindrical obstacles. By organizing all the obstacle information into a set , where , the static obstacle set can be constructed. This set provides a basis for calculating the relative relationship between low-altitude aircraft and obstacles. To consider the influence of the dynamic environment, model the wind disturbance information. The change of wind speed in the mission airspace is decomposed into the horizontal component and the vertical component . These components form a time-varying wind disturbance vector . At the same time, define the boundary of the mission area, and determine the boundary vector of the mission airspace using the minimum coordinate values and the maximum coordinate values in the horizontal and vertical directions Substitute the cluster state matrix and the time-varying wind disturbance vector into the kinematic model to establish a differential equation to describe the dynamic behavior of low-altitude aircraft. The kinematic state transition equation of the low-altitude aircraft is expressed as: ; where and are the velocity components of the low-altitude aircraft in the and directions respectively, and are the velocity components generated by the low-altitude aircraft under the drive of its own velocity and heading angle, and are the velocity components introduced by the wind disturbance. Solve the above equation by numerical integration methods (such as Euler's method or fourth-order Runge-Kutta method) to obtain the evolution law of the position of the low-altitude aircraft over time. Calculate the collision risk between low-altitude aircraft. For the states of any two low-altitude aircraft in the cluster state matrix , take their positions and , and calculate the Euclidean distance between them: ; Compare this distance with the preset risk threshold . If , it is considered that there is a collision risk, and then define the collision risk metric function as: ; Similarly, for each low-altitude aircraft, calculate the minimum distance between its position and each obstacle position in the obstacle set , and obtain the obstacle risk metric function by comparing the difference between it and the obstacle radius : ; Combine the kinematic state transition equation, the collision risk metric function and the obstacle risk metric function to create the complete state space of the low-altitude aircraft cluster.
[0024] In a specific embodiment, the process of executing step 200 may specifically include the following steps: Perform numerical integration operations on the kinematic state transition equation in the state space of the low-altitude aircraft cluster to obtain an initial control point set; Perform safety evaluation operations on the initial control point set according to the collision risk metric function and the obstacle risk metric function to obtain a candidate control point sequence; Calculate the binomial coefficients for each control point in the candidate control point sequence to obtain the basis function matrix of the nth-order Bezier curve, and perform matrix multiplication on the basis function matrix and the control point sequence to obtain the parametric equation of the Bezier curve; Perform a weighted summation calculation on the flight path length, curve smoothness, and detection efficiency of the parametric equation of the Bezier curve to obtain the objective optimization function; Construct the chromosome encoding of the genetic algorithm based on the objective optimization function, and randomly initialize the control point coordinates to obtain the initial population; Perform roulette wheel selection, uniform crossover, and Gaussian mutation operations on the initial population to obtain the iteratively optimized control point sequence, and substitute the iteratively optimized control point sequence into the parametric equation of the Bezier curve for curve reconstruction to obtain the optimal detection flight path parameters.
[0025] Specifically, perform numerical integration on the kinematic state transition equation in the state space of the low-altitude aircraft cluster to obtain the initial control point set. The kinematic state transition equation of the low-altitude aircraft is described as: ; Among them, and are the planar coordinates of the low-altitude aircraft, is the speed, is the heading angle, and are the components of the wind disturbance in the and directions respectively. Through numerical integration methods (such as the trapezoidal method or the Runge-Kutta method), according to the initial state and the time step calculate the motion trajectory of the low-altitude aircraft over a period of time to generate a series of discrete track points. These points form the initial control point set , where is the number of time steps. Perform a safety assessment operation on the initial control point set according to the collision risk metric function and the obstacle risk metric function, so as to screen out the candidate control point sequence. The collision risk metric function is used to evaluate the relative safety between low-altitude aircraft and is expressed as: ; Among them, is the number of low-altitude aircraft, is the th and th low-altitude aircraft, is the Euclidean distance between the is the minimum distance risk between the low-altitude aircraft and the obstacle: ; Among them, is the number of obstacles, is the position of the obstacle, is its radius. Perform risk assessment on each point in the initial control point set, and after removing high-risk points, obtain the candidate control point sequence Perform Bezier curve modeling on the candidate control point sequence. The Bezier curve is a smooth curve generated by weighting control points, and its order basis function matrix is expressed as: ; Among them is the combination number, is the curve parameter. Represent the candidate control point sequence in matrix form , through matrix multiplication , which describes the spatial trajectory of the path. In order to optimize the Bezier curve, comprehensively evaluate its route length, curve smoothness, and detection efficiency. The target optimization function is defined as: ; Among them is the curve length, is the curve smoothness, is the detection efficiency, is the weighting coefficient. The curve length is calculated by integration, the smoothness is represented by the sum of squares of curvature changes, while the detection efficiency is related to the route coverage area. Based on the target optimization function , construct the chromosome encoding of the genetic algorithm. Take the coordinates of the control points as the genes of the chromosome, and randomly initialize the control point positions to generate the initial population. Perform genetic operations on the initial population, including roulette wheel selection, uniform crossover, and Gaussian mutation. Roulette wheel selection selects high-quality individuals according to the fitness value, uniform crossover generates new individuals by randomly exchanging genes, and Gaussian mutation perturbs the gene values slightly to increase the exploration ability of the solution space. After multiple generations of iteration, the population gradually converges, and an optimized control point sequence is generated. Substitute the iteratively optimized control point sequence into the Bezier curve parameter equation for curve reconstruction to obtain the optimal detection route parameters.
[0026] In a specific embodiment, the process of executing step 300 may specifically include the following steps: Based on the optimal detection route parameters, perform parameter discretization processing and coordinate calculation to obtain the initial waypoint set, and calculate and perform threshold judgment on the curvature of adjacent waypoints in the initial waypoint set to obtain the waypoint sequence; Calculate the mean value of the coordinate data in the waypoint sequence and perform deviation matrix operations to obtain the statistical characteristic parameters of the waypoints. Then input the waypoint sequence into the Euclidean distance calculation module to calculate the pairwise distances and obtain the distance matrix between waypoints; Conduct statistical analysis on the distance matrix and calculate the dynamic distance threshold based on the k-fold standard deviation principle to obtain the abnormal point discrimination criterion; Iteratively screen the waypoint sequence according to the abnormal point discrimination criterion and recursively remove the point pairs that do not meet the criterion to obtain the optimized waypoint sequence; Perform principal component analysis on the optimized waypoint sequence, calculate the major axis direction and the ratio of the major and minor axes of the ellipse through eigenvalue decomposition to obtain the basic ellipse parameters, and substitute the basic ellipse parameters and the statistical characteristic parameters into the standard equation of the conic section for coordinate transformation and coefficient calculation to obtain the patrol airspace equation.
[0027] Specifically, perform parameter discretization processing and coordinate calculation based on the optimal detection route parameters to obtain the initial waypoint set. Assume that the optimal detection route is represented by the Bézier curve parametric equation as follows: ; where is the coordinate of the -th control point, is the order of the Bézier curve, is the curve parameter. By discretizing the parameter into , calculate the corresponding waypoint coordinates to form the initial waypoint set . After generating the initial waypoint set, calculate the curvature of adjacent waypoints to evaluate the smoothness of the path. The formula for the curvature is: ; where and represent the first derivative and the second derivative of the curve respectively. Approximate the derivative by numerical methods and calculate the curvature value for each segment of the curve. Compare the curvature with the preset threshold . If it exceeds the threshold, it is considered that this segment of the curve is too steep and needs to be adjusted or screened. After processing, a smoother waypoint sequence is obtained. Calculate the mean value of the coordinate data in the waypoint sequence and perform deviation matrix operations to extract the statistical characteristic parameters. Assume that the waypoint sequence is , and its mean value is: ; The deviation matrix element represents the difference between the th waypoint and the th waypoint, defined as: ; Input the waypoint sequence into the Euclidean distance calculation module to generate a complete distance matrix . Conduct statistical analysis on the distance matrix, and calculate the dynamic distance threshold based on the -fold standard deviation principle, and then define the abnormal point discrimination criterion. The mean value of each element in the distance matrix and the standard deviation are respectively: ; The dynamic distance threshold is . When the distance between a pair of waypoints exceeds this threshold, it is regarded as an abnormal point pair. According to the abnormal point discrimination criterion, iterate and screen the waypoint sequence, and recursively remove the point pairs that do not meet the criterion. This process can gradually eliminate the noise and abnormal points in the waypoints, making the optimized waypoint sequence more in line with the actual task requirements. Conduct principal component analysis on the optimized waypoint sequence to extract the geometric features of the path. By performing eigenvalue decomposition on the covariance matrix of the waypoints, determine the main axis direction and the ratio of the major and minor axes of the distribution ellipse. Assume the covariance matrix is , and its eigenvalues and eigenvectors are and respectively. Then the main axis direction is determined by the eigenvector , and the ratio of the major and minor axes is Substitute the basic parameters of the ellipse (such as the main axis direction, the ratio of the major and minor axes) and the statistical characteristic parameters (such as the mean value and the deviation matrix) into the standard equation of the conic curve to complete the mathematical modeling of the patrol airspace. The standard equation of the conic curve is: ; where is the center of the ellipse, and is the coefficient calculated from the main axis direction and the ratio of the major and minor axes. Through coordinate transformation and parameter optimization, finally obtain the equation describing the patrol airspace of the low-altitude aircraft.
[0028] In a specific embodiment, the process of executing step 400 may specifically include the following steps: Perform numerical integration on the kinematic state transition equation, and combine with the Bessel curve parameters for trajectory prediction to obtain the position component in the high-level state space; Conduct airspace constraint analysis on the collision risk metric function and the obstacle risk metric function, and perform matching operations with the patrol airspace equation to obtain the task component in the high-level state space; Construct a task assignment decision matrix based on the position component and task component of the high-level state space, and discretize the route planning parameters to obtain the high-level action space matrix; Perform multi-objective weight calculation on the completion degree, route following deviation, and energy consumption of the collaborative detection task to obtain the high-level reward function value; Perform spatial mapping on the state data of a single low-altitude aircraft and local obstacle information, and combine with the time-varying wind disturbance vector to perform coordinate transformation to obtain the low-level state space vector; Construct a speed instruction set and a heading instruction set based on the low-level state space vector, and perform discretization processing on the action amount to obtain the low-level action space matrix. Perform multi-index weighted calculation on the track following accuracy, obstacle avoidance safety distance, and maneuver instruction amplitude to obtain the low-level reward function value; Input the high-level state space, low-level state space, their corresponding action spaces, and reward functions into the deep neural network. After double-layer training of the policy network and value network, obtain the high-low level decision network.
[0029] Specifically, by performing numerical integration operations on the kinematic state transition equation and combining with the Bezier curve parameters, predict the future track of the low-altitude aircraft cluster to obtain the position component of the high-level state space. The kinematic state transition equation describes the motion characteristics of the low-altitude aircraft, and the formula is: ; Among them, represents the two-dimensional plane position of the low-altitude aircraft, is the speed, is the heading angle, and are the components of the wind disturbance in the and directions. By discretizing time , and using numerical integration methods (such as Euler method or Runge-Kutta method), calculate the position trajectory of the low-altitude aircraft at multiple time steps. The Bezier curve parameter equation describes the ideal track, where are the control points, is the curve order. Combine the numerical integration results with the Bezier curve prediction to correct the actual track deviation and obtain a more accurate position component. Analyze the collision risk metric function and the obstacle risk metric function, and combine with the patrol airspace equation constraint to obtain the task component of the high-level state space. The collision risk metric function is defined as: ; Among them is the distance between the th and the th low-altitude aircraft, is the safety distance threshold. The obstacle risk metric function is as follows: ; where is the number of obstacles, and are the position and extent of the obstacle, respectively. Combining with the patrol airspace equation of the low-altitude aircraft , the position of the low-altitude aircraft is matched with the mission area, and the mission relevance information is extracted to form the mission component. On this basis, a mission assignment decision matrix is constructed based on the position component and mission component in the high-level state space. The mission assignment decision matrix represents the priority of the th low-altitude aircraft to execute the th mission. The route planning parameters are discretized. For example, the mission space is divided into equally spaced grids, and each grid point is defined as a possible mission target, thereby generating a high-level action space matrix. To optimize the mission assignment, a high-level reward function value is defined, comprehensively considering the completion degree of the collaborative detection mission, the route following deviation, and the energy consumption. The high-level reward function is expressed as: ; where is the mission completion degree, is the track deviation, is the energy consumption, is the weighting coefficient. At the same time, for the local control requirements of a single low-altitude aircraft, its state data is spatially mapped with the local obstacle information, and coordinate transformation is combined with the time-varying wind disturbance vector to obtain the low-level state space vector. The low-level state space includes the current position, speed, heading angle of the low-altitude aircraft, and the relative position of the obstacle in the local reference frame. Based on the low-level state space, a speed command set and a heading command set are constructed. The speed command set represents different speed selections, and the heading command set represents different heading adjustments. Through the discretization of the action quantity, a low-level action space matrix is generated. The low-level reward function is obtained by weighted calculation of the track following accuracy, the obstacle avoidance safety distance, and the amplitude of the maneuver command: ; where represents the track following accuracy, represents the obstacle avoidance distance, represents the amplitude of the maneuver command. The high-level state space, low-level state space, their corresponding action spaces, and reward functions are input into the deep neural network to construct the policy network and value network. The policy network is used to learn the optimal action policy, and the formula is: ; where Is the state And the action Of the value function. The value network is used to estimate the long-term reward , and its goal is to minimize the Bellman error: ; Where Is the immediate reward, Is the discount factor. After double-layer training of the policy network and the value network, a high-low layer decision network is finally formed for real-time control and task allocation of the low-altitude aircraft cluster.
[0030] In a specific embodiment, the process of executing step 500 may specifically include the following steps: Input the current state into the Actor network in the high-low layer decision network to calculate the task allocation strategy, obtain the high-level task allocation instruction, and decouple the task of the low-altitude aircraft cluster according to the high-level task allocation instruction to obtain the single-aircraft execution instruction sequence; Input the single-aircraft execution instruction sequence into the Actor network in the high-low layer decision network to calculate the action strategy and obtain the low-level control instruction; Perform temporal difference calculation on the low-level control instruction and input the calculation result into the Critic network of the high-low layer decision network for value evaluation to obtain the state-action value; Optimize the low-level control instruction based on the state-action value by policy gradient to obtain the local optimal control strategy; Input the local optimal control strategy into the Critic network of the high-low layer decision network for global value evaluation to obtain the global cooperation index, and correct the high-level task allocation instruction by policy gradient according to the global cooperation index to obtain the optimal task allocation strategy; Perform distributed deployment on the optimal task allocation strategy and the local optimal control strategy to obtain the distributed control strategy of the low-altitude aircraft cluster.
[0031] Specifically, input the current state into the Actor network in the high-low layer decision network, calculate the task allocation strategy, and obtain the high-level task allocation instruction. The state of the low-altitude aircraft cluster is represented as , where Is the -th state vector of the low-altitude aircraft, including its two-dimensional position , speed And heading angle . Through the high-level Actor network , calculate the task allocation strategy , where Is the task allocation instruction, They are the parameters of the high-level Actor network. This instruction assigns mission objectives to each low-altitude aircraft, such as search areas or patrol routes. According to the high-level mission assignment instruction, the global mission is decoupled into a sequence of single-aircraft execution instructions , each corresponding to the specific mission of a low-altitude aircraft. The single-aircraft execution instruction sequence is input into the Actor network in the high-low level decision-making network to calculate the action strategy and obtain the low-level control instruction. For each low-altitude aircraft, its low-level state vector includes local position information, speed, heading angle, and local obstacle information , which is generated by the low-level Actor network to generate low-level action instructions , where represents the adjusted speed and heading angle. In this way, the low-level network generates executable action instructions for each low-altitude aircraft, ensuring local control accuracy and real-time obstacle avoidance capabilities. Temporal difference calculation is performed on the low-level control instructions to evaluate their dynamic performance. Assume that the state at the current moment is , the action is , the state at the next moment is , and the immediate reward is . Based on the temporal difference formula: ; where is the state-action value, is the discount factor, is the value function of the state at the next moment. The calculation result is input into the Critic network of the high-low level decision-making network to evaluate the long-term value of the state-action. Based on the state-action value, policy gradient optimization is performed on the low-level control instructions to update the parameters of the low-level Actor network. The goal of the policy gradient is to maximize the expected return , and its gradient calculation is: ; By optimizing the gradient, a locally optimal control strategy is generated to ensure the efficient control of low-altitude aircraft within the local range. The locally optimal control strategy is input into the Critic network of the high-low level decision-making network for global value evaluation and calculation of the global cooperation index. The global value function comprehensively considers factors such as mission completion, cooperation efficiency of low-altitude aircraft, and energy consumption, and is defined as: ; where is the mission completion degree, is the cooperation cost, is the energy consumption, is the weight coefficient. According to the global cooperation index, the high-level task allocation instruction is corrected through policy gradient, the parameters of the high-level Actor network are updated, and the optimal task allocation strategy is generated. The optimal task allocation strategy and the local optimal control strategy are distributedly deployed to the low-altitude aircraft cluster. Each low-altitude aircraft adjusts its flight path and task execution mode in real time according to the allocated high-level task instruction and the local control strategy , and forms the distributed control strategy of the low-altitude aircraft cluster.
[0032] In a specific embodiment, the autonomous obstacle avoidance method of the low-altitude intelligent dynamic monitoring aircraft further includes the following steps: Perform decomposition operations on the speed control instruction and the heading control instruction in the local optimal control strategy to obtain the nominal control input of each low-altitude aircraft; Substitute the nominal control input and the kinematic state transition equation into the Lyapunov function, calculate the derivative of the function, obtain the stability condition, design an adaptive compensation term for the time-varying wind disturbance vector based on the stability condition, and obtain the initial compensation control law; Construct a cooperative compensation term according to the optimal task allocation strategy, and combine the cooperative compensation term with the initial compensation control law to obtain an adaptive compensation controller; Design a sliding mode surface for the control gain of the adaptive compensation controller, and adjust the parameters through the global cooperation index to obtain the compensation control parameters. Substitute the GPS position data and the nominal control input of a single low-altitude aircraft into the Kalman filter to obtain the wind disturbance state quantity; Perform adaptive law calculation on the wind disturbance state quantity and the compensation control parameters to obtain the real-time compensation control quantity, and perform weighted combination of the nominal control input and the real-time compensation control quantity to obtain the wind disturbance compensation control quantity of each low-altitude aircraft.
[0033] Specifically, perform decomposition operations on the speed control instruction and the heading control instruction in the local optimal control strategy to obtain the nominal control input of each low-altitude aircraft. Assume that the speed control instruction provided by the local optimal control strategy is , and the heading control instruction is . By decomposition, the speed components of the low-altitude aircraft in the and directions are obtained: ; where and respectively represent the components of the nominal control input in the and directions. For a low-altitude aircraft, its nominal control input is expressed in vector form: ; Substitute the nominal control input and the kinematic state transition equation of the low-altitude aircraft into the Lyapunov function, and calculate its derivative to determine the stability condition. The kinematic state transition equation of the low-altitude aircraft is: ; where and are the velocities of the low-altitude aircraft in the and directions, and are the velocity components of the wind disturbance. To analyze the stability of the system, introduce the Lyapunov function: ; Its physical meaning is half of the kinetic energy of the low-altitude aircraft. Take the derivative of to get: ; Substitute the state transition equation, and we can get: ; To ensure the stability of the system, design the adaptive compensation terms and such that , that is, the system gradually tends to be stable. The initial compensation control law is: ; where is the proportional coefficient, which is used to dynamically weaken the influence of the wind disturbance. Construct the cooperative compensation term according to the optimal task allocation strategy to coordinate the task execution of the low-altitude aircraft cluster. Assume the optimal task allocation strategy is , and each represents the target task assigned to the th low-altitude aircraft. The goal of the cooperative compensation term is to reduce the conflicts between low-altitude aircraft and optimize the overall cooperation efficiency, and its form is: ; where is the cooperation coefficient, is the cooperation weight between the low-altitude aircraft and . Combine the initial compensation control law with the cooperative compensation term to obtain the adaptive compensation controller: ; Design the sliding mode surface for the control gain of the adaptive compensation controller to enhance its robustness. The sliding mode surface is defined as: ; where is the sliding mode control gain. Take the derivative of Derive the derivative, combine it with the output of the adaptive compensation controller, and design the sliding mode control law: ; where is the sliding mode control gain, which is used to dynamically adjust the compensation intensity. Optimize the compensation control parameters through the global cooperation index, substitute the GPS position data and the nominal control input of a single low-altitude aircraft into the Kalman filter to estimate the wind disturbance state quantity. The prediction equation of the Kalman filter is: ; where is the state estimate value at the th step, and are the system dynamic matrices. The update equation is: ; where is the measured value, is the Kalman gain. Calculate the adaptive law for the wind disturbance state quantity and the compensation control parameters to obtain the real-time compensation control quantity: ; Apply to each low-altitude aircraft to effectively offset the influence of wind disturbance and ensure the stability of the system.
[0034] The above describes the autonomous obstacle avoidance method of the low-altitude intelligent dynamic monitoring aircraft in the embodiment of the present application. Next, the autonomous obstacle avoidance system 10 of the low-altitude intelligent dynamic monitoring aircraft in the embodiment of the present application will be described. Please refer to Figure 2 , an embodiment of the autonomous obstacle avoidance system 10 of the low-altitude intelligent dynamic monitoring aircraft in the embodiment of the present application includes: A modeling module 11, which is used to construct a state vector for the position coordinates, flight speed, and heading angle of the low-altitude aircraft cluster, and perform kinematic modeling in combination with the obstacle position data and wind speed components to obtain the state space of the low-altitude aircraft cluster; A calculation module 12, which is used to perform Bezier curve calculation and genetic optimization on the control point sequence in the state space of the low-altitude aircraft cluster to obtain the optimal detection route parameters; A processing module 13, which is used to perform geometric distance calculation and two-point removal processing according to the waypoint sequence of the optimal detection route to obtain the patrol airspace equation; A reinforcement learning module 14, which is used to input the state space of the low-altitude aircraft cluster, the optimal detection route parameters, and the patrol airspace equation into a hierarchical reinforcement learning framework to perform multi-level reward function calculation to obtain a high-level and low-level decision-making network; An optimization module 15, which is used to perform temporal difference learning and policy gradient optimization based on the high-level and low-level decision-making network to obtain a distributed control strategy.
[0035] Through the collaborative cooperation of the above-mentioned various components, by constructing a hierarchical reinforcement learning framework, the low-altitude aircraft swarm control problem is decomposed into two levels: high-level task allocation and low-level trajectory execution, reducing the system complexity and improving the scalability and computational efficiency of the algorithm. The optimal detection route is generated by combining Bezier curves with genetic algorithm optimization, and the patrol airspace is determined by the ellipse fitting method, making the planning result more in line with the actual flight requirements. A distributed control strategy based on temporal difference learning and policy gradient optimization is designed to achieve the unified optimization of multi-aircraft collaborative performance and single-aircraft control performance. By introducing an adaptive compensation controller and a Kalman filter, the path following problem under unknown time-varying wind disturbances is solved, ensuring the stability of the system. The double-point removal algorithm based on geometric distance is adopted to effectively avoid the influence of abnormal waypoints on the patrol airspace planning and improve the rationality of airspace division. Through the design of a multi-level reward function, multiple objectives such as collision avoidance, task completion degree, and energy efficiency are considered together, realizing the global optimal control of the low-altitude aircraft swarm. A wind disturbance compensation controller is designed based on the Lyapunov stability theory, and the parameters are adjusted in combination with the global cooperation index to ensure the stable operation of the swarm system in an unknown wind disturbance environment. The control strategy is implemented by a distributed deployment method, reducing the communication burden and improving the real-time performance and robustness of the system.
[0036] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0037] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution of this application, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0038] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. An autonomous obstacle avoidance method for a low-altitude intelligent dynamic monitoring aircraft, characterized in that: The method comprises: The state vector of the position coordinates, flight speed, and heading angle of the low-altitude aircraft cluster is constructed, and kinematic modeling is performed in combination with obstacle position data and wind speed components to obtain the state space of the low-altitude aircraft cluster; Performing Bezier curve calculation and genetic optimization on the control point sequence in the state space of the low-altitude aircraft cluster to obtain optimal detection route parameters; Perform geometric distance calculation and double point removal processing according to the waypoint sequence of the optimal detection route to obtain a patrol airspace equation; Inputting the low-altitude aircraft cluster state space, the optimal detection route parameters and the patrol airspace equation into a hierarchical reinforcement learning framework to perform multi-level reward function calculations to obtain high- and low-level decision networks; Based on the high- and low-level decision networks, temporal difference learning and policy gradient optimization are performed to obtain a distributed control strategy.
2. The autonomous obstacle avoidance method for a low-altitude intelligent dynamic monitoring aircraft according to claim 1 is characterized in that: The state vector is constructed for the position coordinates, flight speed, and heading angle of the low-altitude aircraft cluster, and kinematic modeling is performed in combination with obstacle position data and wind speed components to obtain the state space of the low-altitude aircraft cluster, including: The two-dimensional plane coordinate values, flight speed values, and heading angle values of a single low-altitude aircraft in the low-altitude aircraft cluster are combined to obtain a single-aircraft state vector, and the single-aircraft state vector is sequentially arranged and matrix-joined to obtain a cluster state matrix; Perform data extraction and set analysis on the spatial position coordinates and range parameters of each obstacle to obtain a static obstacle set consisting of m obstacles; The horizontal and vertical wind speed components at the current moment are constructed into two-dimensional vectors to obtain the time-varying wind disturbance vector, and the minimum and maximum coordinate values in the horizontal and vertical directions of the specified task area are extracted to obtain the task airspace boundary vector. Substituting the cluster state matrix and the time-varying wind disturbance vector into the differential equation for numerical solution to obtain a kinematic state transfer equation; Performing Euclidean distance calculation and risk threshold comparison operations on the position data of any two low-altitude aircraft in the cluster state matrix to obtain a collision risk measurement function, and performing minimum distance calculation and safety margin evaluation operations on each low-altitude aircraft in the cluster state matrix and each obstacle in the static obstacle set to obtain an obstacle risk measurement function; A low-altitude aircraft cluster state space is created based on the kinematic state transfer equation, the collision risk measurement function and the obstacle risk measurement function.
3. The autonomous obstacle avoidance method for a low-altitude intelligent dynamic monitoring aircraft according to claim 2 is characterized in that: The step of performing Bezier curve calculation and genetic optimization on the control point sequence in the state space of the low-altitude aircraft cluster to obtain the optimal detection route parameters includes: Performing numerical integration operations on the kinematic state transfer equations in the state space of the low-altitude aircraft cluster to obtain an initial control point set; Performing a safety assessment operation on the initial control point set according to the collision risk measurement function and the obstacle risk measurement function to obtain a candidate control point sequence; Calculating binomial coefficients for each control point in the candidate control point sequence to obtain a basis function matrix of an n-order Bezier curve, and performing matrix multiplication operation on the basis function matrix and the control point sequence to obtain a Bezier curve parameter equation; Performing weighted sum calculation on the route length, curve smoothness and detection efficiency of the Bezier curve parameter equation to obtain a target optimization function; Constructing a chromosome code of a genetic algorithm based on the target optimization function, and randomly initializing the coordinates of the control points to obtain an initial population; Roulette wheel selection, uniform crossover and Gaussian mutation operations are performed on the initial population to obtain an iteratively optimized control point sequence, and the iteratively optimized control point sequence is substituted into the Bezier curve parameter equation for curve reconstruction to obtain the optimal detection route parameters.
4. The autonomous obstacle avoidance method for a low-altitude intelligent dynamic monitoring aircraft according to claim 3 is characterized in that: The geometric distance calculation and double point removal processing are performed according to the waypoint sequence of the optimal detection route to obtain the patrol airspace equation, including: Based on the optimal detection route parameters, parameter discretization processing and coordinate calculation are performed to obtain an initial waypoint set, and curvatures of adjacent waypoints in the initial waypoint set are calculated and threshold judgment is performed to obtain a waypoint sequence; Performing mean calculation and deviation matrix operation on the coordinate data in the waypoint sequence to obtain statistical characteristic parameters of the waypoints, and inputting the waypoint sequence into a Euclidean distance calculation module to perform pairwise distance calculation to obtain a distance matrix between the waypoints; Performing statistical analysis on the distance matrix, and calculating a dynamic distance threshold based on the k-times standard deviation principle to obtain an outlier discrimination criterion; Iteratively screening the waypoint sequence according to the abnormal point discrimination criterion, and recursively removing the point pairs that do not meet the criterion to obtain an optimized waypoint sequence; The optimized waypoint sequence is subjected to principal component analysis, and the direction of the major axis and the ratio of the major and minor axes of the ellipse are calculated by eigenvalue decomposition to obtain the basic parameters of the ellipse. The basic parameters of the ellipse and the statistical characteristic parameters are substituted into the standard equation of the quadratic curve for coordinate transformation and coefficient calculation to obtain the patrol airspace equation.
5. The autonomous obstacle avoidance method for a low-altitude intelligent dynamic monitoring aircraft according to claim 4 is characterized in that: The inputting of the low-altitude aircraft cluster state space, the optimal detection route parameters and the patrol airspace equation into a hierarchical reinforcement learning framework to perform multi-level reward function calculations to obtain high- and low-level decision networks includes: Performing numerical integration operation on the kinematic state transfer equation, and performing track prediction in combination with the Bezier curve parameters to obtain the position component of the high-level state space; Performing airspace constraint analysis on the collision risk measurement function and the obstacle risk measurement function, and performing matching operations with the patrol airspace equation to obtain task components in a high-level state space; A task allocation decision matrix is constructed based on the position component and the task component of the high-level state space, and the route planning parameters are discretized to obtain a high-level action space matrix; Perform multi-objective weight calculation on the completion degree, route following deviation and energy consumption of the collaborative detection task to obtain the high-level reward function value; The state data of a single low-altitude aircraft and the local obstacle information are spatially mapped, and the coordinate transformation is performed in combination with the time-varying wind disturbance vector to obtain a low-level state space vector; Based on the low-level state space vector, a speed instruction set and a heading instruction set are constructed, and the action amount is discretized to obtain a low-level action space matrix, and a multi-index weighted calculation is performed on the track following accuracy, obstacle avoidance safety distance and maneuver instruction amplitude to obtain a low-level reward function value; The high-level state space, the low-level state space and their corresponding action space and reward function are input into a deep neural network, and after double-layer training of a policy network and a value network, a high-level and low-level decision network is obtained.
6. The autonomous obstacle avoidance method for a low-altitude intelligent dynamic monitoring aircraft according to claim 5 is characterized in that: The distributed control strategy is obtained by performing temporal difference learning and policy gradient optimization based on the high- and low-level decision networks, including: Input the current state into the Actor network in the high-low layer decision network to calculate the task allocation strategy, obtain the high-level task allocation instruction, and perform task decoupling on the low-altitude aircraft cluster according to the high-level task allocation instruction to obtain a single-machine execution instruction sequence; Inputting the single-machine execution instruction sequence into the Actor network in the high- and low-level decision networks to calculate the action strategy and obtain the low-level control instructions; Performing time difference calculation on the low-level control instruction, and inputting the calculation result into the Critic network of the high- and low-level decision networks for value evaluation to obtain the state action value; Performing policy gradient optimization on the low-level control instructions based on the state action value to obtain a local optimal control strategy; The local optimal control strategy is input into the Critic network of the high- and low-level decision networks for global value evaluation to obtain a global coordination index, and the high-level task allocation instruction is corrected by a policy gradient according to the global coordination index to obtain an optimal task allocation strategy; The optimal task allocation strategy and the local optimal control strategy are deployed in a distributed manner to obtain a distributed control strategy for a low-altitude aircraft cluster.
7. The autonomous obstacle avoidance method for a low-altitude intelligent dynamic monitoring aircraft according to claim 6 is characterized in that: The autonomous obstacle avoidance method of the low-altitude intelligent dynamic monitoring aircraft also includes: Decomposing and calculating the speed control command and the heading control command in the local optimal control strategy to obtain the nominal control input of each low-altitude aircraft; Substituting the nominal control input and the kinematic state transfer equation into the Lyapunov function, and calculating the function derivative to obtain a stability condition, and designing an adaptive compensation term for the time-varying wind disturbance vector based on the stability condition to obtain an initial compensation control law; Constructing a coordinated compensation term according to the optimal task allocation strategy, and combining the coordinated compensation term with the initial compensation control law to obtain an adaptive compensation controller; The control gain of the adaptive compensation controller is designed with a sliding surface, and the parameters are adjusted through the global coordination index to obtain compensation control parameters, and the GPS position data of a single low-altitude aircraft and the nominal control input are substituted into the Kalman filter to obtain the wind disturbance state quantity; The wind disturbance state quantity and the compensation control parameter are adaptively calculated to obtain a real-time compensation control quantity, and the nominal control input and the real-time compensation control quantity are weighted combined to obtain a wind disturbance compensation control quantity for each low-altitude aircraft.
8. An autonomous obstacle avoidance system for low-altitude intelligent dynamic monitoring aircraft, characterized in that: A method for executing an autonomous obstacle avoidance method for a low-altitude intelligent dynamic monitoring aircraft as claimed in any one of claims 1 to 7, wherein the autonomous obstacle avoidance system for the low-altitude intelligent dynamic monitoring aircraft comprises: The modeling module is used to construct the state vector of the position coordinates, flight speed, and heading angle of the low-altitude aircraft cluster, and to perform kinematic modeling in combination with the obstacle position data and wind speed components to obtain the state space of the low-altitude aircraft cluster; A calculation module, used for performing Bezier curve calculation and genetic optimization on the control point sequence in the state space of the low-altitude aircraft cluster to obtain the optimal detection route parameters; A processing module, used for performing geometric distance calculation and double point removal processing according to the waypoint sequence of the optimal detection route to obtain a patrol airspace equation; A reinforcement learning module, used for inputting the low-altitude aircraft cluster state space, the optimal detection route parameters and the patrol airspace equation into a hierarchical reinforcement learning framework to perform multi-level reward function calculations to obtain high- and low-level decision networks; The optimization module is used to perform temporal difference learning and policy gradient optimization based on the high- and low-level decision networks to obtain a distributed control strategy.
Citation Information
Cited By
Trajectory planning method, device and equipment for variant near space aircraft and medium
CN121008487A
Remote intelligent lifting hook control method based on Internet of Things
CN121180862A
A hook remote intelligent control method based on an internet of things
CN121180862B
Layered adaptive safety reinforcement learning system and method applied to patrol robot
CN121300346A
Urban low-altitude distribution unmanned aerial vehicle self-adaptive navigation method and system
CN122261182A