A planning and control method and system for a rotary-wing unmanned aerial vehicle

By combining the proximal strategy optimization algorithm and B-spline curve with Lyapunov stability theory, a lift input and attitude torque controller is designed to solve the problems of low precision and collision in the planning and control of rotorcraft UAVs, and achieve efficient and safe path tracking.

CN118605592BActive Publication Date: 2025-09-23UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410507211.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-09-23
Estimated Expiration
2044-04-25

AI Technical Summary

Technical Problem

The existing planning and control methods of rotary-wing UAVs are divorced from the path planning task design, and the control algorithm does not conform to the actual situation, resulting in low control accuracy. Inaccurate tracking can easily cause the rotary-wing UAV to collide with obstacles and fail the mission.

Method used

A proximal strategy optimization algorithm is used to generate feasible waypoints, which are converted into a trackable continuous path using B-spline curves. Combined with the posture dynamics equations and Lyapunov stability theory, lift input and attitude torque controllers are designed, and precise path tracking is achieved through a closed-loop controller.

Benefits of technology

The control accuracy and flight safety of the rotorcraft UAV are improved, ensuring that the desired path can be tracked quickly and stably in complex environments and avoiding collisions with obstacles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118605592B_ABST
    Figure CN118605592B_ABST
Patent Text Reader

Abstract

The present invention provides a planning and control method and system for a rotary-wing unmanned aerial vehicle (UAV), which relates to the technical field of planning and control of rotary-wing UAVs. The method comprises the following steps: generating feasible waypoints of the rotary-wing UAV in a flight environment; converting the feasible waypoints into a trackable continuous path; establishing a posture dynamics equation of the rotary-wing UAV in an inertial coordinate system; establishing a path tracking error dynamics equation of the rotary-wing UAV in the path reference coordinate system in combination with a posture dynamics model and a conversion relationship between a path reference coordinate system and an inertial coordinate system; determining a path parameter update law of the trackable continuous path, an expected heading angle and an expected speed of an expected posture; designing a lift input controller and an attitude torque controller of the posture dynamics equation with the goal of eliminating the tracking error of the path tracking error dynamics equation; and controlling the rotary-wing UAV through the lift input controller and the attitude torque controller in combination with the path parameter update law, thereby improving the control accuracy of the rotary-wing UAV.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of rotary-wing UAV planning and control technology, and in particular to a rotary-wing UAV planning and control method and system. Background Art

[0002] Path planning and tracking control for rotary-wing UAVs are key technologies for numerous practical engineering tasks, including aerial search and rescue, logistics distribution, power inspections, and agricultural plant protection. Rotary-wing UAVs typically rely on path planning algorithms to generate a flyable path and path tracking control algorithms to accurately track the path, ensuring efficient mission completion.

[0003] The path planning algorithms for rotary-wing drones are relatively mature, including typical algorithms such as the A* algorithm, rapidly expanding random trees, and probabilistic roadmaps. Many improved versions have been proposed to address the shortcomings of the original algorithms, improving the algorithm's performance in terms of search time, space, and convergence speed. However, these algorithms have certain limitations, and the algorithms themselves lack the ability to analyze, extract information, and make autonomous decisions. Compared to traditional path planning algorithms, neural network-based learning algorithms, represented by deep reinforcement learning, have irreplaceable advantages in terms of iteration time, generalization ability, and intelligence level. By acquiring environmental information through various types of sensors carried by quadrotors and inputting this information into neural networks for feature extraction and learning, autonomous decision-making is achieved, effectively solving the path planning problem.

[0004] Research on path tracking control for rotary-wing UAVs is characterized by diversity and specificity. In recent years, extensive research has been conducted on the theory and application of path tracking control. For example, robust and adaptive model predictive control (MPC) has been proposed to combat model uncertainty and external disturbances in rotary-wing UAVs, improving flight performance and stability in complex environments. Nonlinear state feedback control methods have been used to achieve global convergence of path tracking errors while ensuring bounded control inputs. Velocity vector following controllers have been designed based on the concept of differential flatness, acquiring velocity, acceleration, and first- and second-order derivatives of acceleration. An inner-loop controller has been designed to enable the quadrotor to follow a specified velocity vector field. Path tracking control eliminates time dependence and uses scalar variables to parameterize the path. Compared to trajectory tracking control, it offers several advantages, such as generally smoother convergence to the desired path, smaller transient errors, smaller control amplitudes, and smoother responses. Furthermore, path planning tasks generally do not require tracking complex path curves, making path tracking control more advantageous.

[0005] In summary, the existing planning and control methods of rotorcraft UAVs are divorced from the path planning task design, and the control algorithm does not conform to the actual situation, resulting in low control accuracy. When the safety margin of the planned path is insufficient, inaccurate tracking can easily cause the rotorcraft UAV to collide with obstacles, thereby causing mission failure. Summary of the Invention

[0006] In order to solve the technical problems in the prior art that the existing planning and control methods of rotary-wing UAVs are divorced from the path planning task design, the control algorithm does not conform to the actual situation, resulting in low control accuracy, and when the safety margin of the planned path is insufficient, inaccurate tracking easily causes the rotary-wing UAV to collide with obstacles, thereby causing mission failure, the present invention provides a planning and control method and system for rotary-wing UAVs.

[0007] The technical solutions provided by the embodiments of the present invention are as follows:

[0008] First aspect

[0009] An embodiment of the present invention provides a planning and control method for a rotary-wing UAV, comprising:

[0010] S1: Generate feasible waypoints for the rotorcraft in the flight environment based on the proximal strategy optimization algorithm;

[0011] S2: Convert feasible waypoints into a trackable continuous path using B-spline curves;

[0012] S3: Establish the posture dynamics equations of the rotorcraft in the inertial coordinate system;

[0013] S4: Combined with the pose dynamics model and the conversion error between the path reference coordinate system and the inertial coordinate system, the path tracking error dynamics equation of the rotorcraft in the path reference coordinate system is established, where the origin of the path reference coordinate system is the desired path reference point for tracking the continuous path;

[0014] S5: Combined with Lyapunov stability theory, determine the path parameter update law that can track the continuous path, the desired heading angle and the desired velocity of the desired attitude;

[0015] S6: To eliminate the tracking error of the path tracking error dynamic equation, design the lift input controller of the attitude dynamic equation according to the desired velocity, and design the attitude torque controller according to the desired heading angle;

[0016] S7: Combined with the path parameter update law, the rotor UAV is controlled through the lift input controller and attitude torque controller.

[0017] Second aspect

[0018] An embodiment of the present invention provides a planning and control system for a rotary-wing UAV, comprising:

[0019] processor;

[0020] A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the planning and control method for the rotary-wing UAV as described in the first aspect is implemented.

[0021] The third aspect

[0022] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the planning and control method for a rotary-wing UAV as described in the first aspect is implemented.

[0023] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0024] In this paper, a path planning algorithm based on a proximal strategy optimization algorithm is designed. This algorithm models the flight environment and generates a series of feasible waypoints. These waypoints are then converted into a parameterized, trackable continuous path using B-spline curves. This allows for local modification of the path planning problem, and its order does not increase with the addition of control points. Even with the addition of more control points, the complexity of the curve does not increase, thus avoiding the Runge phenomenon in Bezier curves. Based on this trackable continuous path, a closed-loop controller is designed, combining Lyapunov stability theory with the goal of eliminating the tracking error in the path tracking error dynamics equation. This controller is used to control a rotary-wing UAV. While improving the control accuracy of the rotary-wing UAV, it also ensures rapid convergence, stability, and feasibility, enabling the UAV to autonomously plan and control itself according to the flight environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0026] Figure 1 A schematic flow chart of a planning and control method for a rotary-wing UAV provided in an embodiment of the present invention;

[0027] Figure 2 This is a schematic diagram of a simulation environment for a path planning task provided by an embodiment of the present invention;

[0028] Figure 3 This is a round reward curve in the model training process of a path planning task based on a proximal policy optimization algorithm provided by an embodiment of the present invention;

[0029] Figure 4 It is a feasible path map generated by a model in a training environment provided by an embodiment of the present invention;

[0030] Figure 5 It is a feasible path map generated by a model provided by an embodiment of the present invention in a test environment;

[0031] Figure 6 It is a feasible path map generated by a fine-tuned model in a test environment provided by an embodiment of the present invention;

[0032] Figure 7 This is a feasible path map generated based on the A* algorithm provided by an embodiment of the present invention;

[0033] Figure 8 It is a feasible path generated by a proximal strategy optimization algorithm and an A* algorithm provided by an embodiment of the present invention;

[0034] Figure 9 This is a comparison chart of the total lengths of feasible paths generated by a proximal strategy optimization algorithm and an A* algorithm provided by an embodiment of the present invention;

[0035] Figure 10 A roadmap provided by an embodiment of the present invention is obtained by smoothing a feasible path generated by a PPO algorithm based on a B-spline curve;

[0036] Figure 11 This is a geometric diagram of an inertial coordinate system, a body coordinate system, and a path reference coordinate system framework in a plane provided by an embodiment of the present invention;

[0037] Figure 12 This is a comparison diagram of the actual motion path and the expected path of a rotary-wing UAV provided by an embodiment of the present invention;

[0038] Figure 13 This is a speed and expected speed comparison diagram provided by an embodiment of the present invention;

[0039] Figure 14 This is a comparison diagram of a posture and an expected posture provided by an embodiment of the present invention;

[0040] Figure 15 is a graph showing the change of attitude error over time provided by an embodiment of the present invention;

[0041] Figure 16 This is a comparison diagram of the actual motion path and the smoothed path of a rotary-wing UAV provided by an embodiment of the present invention;

[0042] Figure 17 A schematic structural diagram of a planning and control system for a rotary-wing UAV provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0044] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.

[0045] In the embodiments of the present invention, the terms "image" and "picture" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same. The terms "of," "corresponding," and "corresponding" may be used interchangeably. It should be noted that, when the distinction between them is not emphasized, their intended meanings are the same.

[0046] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0047] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0048] Reference Manual Figure 1 , shows a flow chart of a rotor UAV planning and control method provided by an embodiment of the present invention.

[0049] An embodiment of the present invention provides a method for planning and controlling a rotary-wing UAV. The method can be implemented by a rotary-wing UAV planning and control device, which can be a terminal or a server. The processing flow of the method for planning and controlling a rotary-wing UAV can include the following steps:

[0050] S1: Generate feasible waypoints for the rotorcraft in the flight environment based on the proximal strategy optimization algorithm.

[0051] Optionally, the rotary wing UAV may be a rotary wing UAV. The flight environment may be an indoor environment or an outdoor environment.

[0052] Among them, the Proximal Policy Optimization (PPO) algorithm is currently considered to be the most effective deep reinforcement learning algorithm. It is an online learning algorithm based on policy gradients and an Actor-Critic architecture. First, the objective function of the policy gradient is generally defined as

[0053] L PG (θ)=E t [logπθ (a t |s t )A t ] (1)

[0054] where π θ is a policy network about the optimization of parameter vector θ, where subscript t represents the time step, E t Expressing expectations, A t is the estimated value of the advantage function of the model at time t, and the mathematical expression is

[0055] A t =-V(s t )+r t +γr t+1 +…+γ T-t+1 r T-1 +γ T-t V(s T ) (2)

[0056] Where V(s t ) is the state value function with time index t∈[0,T], T is the length of the trajectory segment, r t is the reward value at time t, and γ represents the discount factor. The general definition of the advantage function is

[0057] A t =δ t +(γλ)δ t+1 +…+(γλ) T-t+1 δ T-1 (3)

[0058] where δ t =r t +γV(s t+1 )-V(s t ).

[0059] Derivative the objective function (1) of the policy gradient to obtain the gradient estimator g = E t [▽logπ θ (a t |s t )A t ], and finally use the gradient ascent algorithm θ new ←θ old +αg updates the network parameters θ.

[0060] It should be noted that by generating feasible waypoints in the rotorcraft UAV flight environment based on the proximal strategy optimization algorithm, such a path planning method can take into account the complex factors of the flight environment, generate a safer and more efficient flight path, and by converting it into a trackable continuous path, ensure that the UAV maintains smooth movement during flight, thereby improving the safety, stability and efficiency of the flight.

[0061] In a possible implementation, S1 specifically includes:

[0062] S101: Based on the Markov decision process, a tuple model from the initial point to the target point in the flight environment is established. The tuple model includes the state space, action space, probability density function and reward function:

[0063] <S,A,P,R>

[0064]

[0065] A=[a k ×45°]

[0066] a k ={0,1,2,3,4,5,6,7}

[0067] R=r total =r step +r reach +r collision

[0068]

[0069]

[0070] r step =ρ(d last -d current ),ρ>0

[0071] Among them, S represents the state space, A represents the action space, P represents the probability density function of state transition, R represents the reward function, (g x ,g y ) represents the target point coordinates, (a x ,a y ) represents the current position coordinates of the rotary wing UAV, and They represent the coordinates of the first obstacle and the second obstacle closest to the rotorcraft, respectively. k represents the set of action decision variables, r total represents the total reward value of the flight environment feedback after the rotorcraft moves one step, r collision represents the reward function responsible for collision avoidance, d obs Indicates the distance between the rotorcraft and the obstacle, d wall Indicates the distance between the rotorcraft and the wall, r reach represents the reward function responsible for traveling to the goal point, d tar Indicates the distance between the rotorcraft and the target point, r step represents the single-step reward function, dlast Indicates the distance between the rotorcraft and the target point at the last moment, d current It represents the distance between the rotorcraft and the target point at the current moment, and ρ represents the weight coefficient.

[0072] S102: Determine the direction of motion of the rotorcraft:

[0073]

[0074] in, represents the direction of movement, a ki ∈a k .

[0075] S103: Based on the tuple model, the policy gradient algorithm is improved in combination with the importance sampling strategy to obtain the improved objective function:

[0076]

[0077] A t =-V(s t )+r t +γr t+1 +…+γ T-t+1 r T-1 +γ T-t V(s T )

[0078]

[0079] Among them, L CPI (θ) represents the improved objective function, E t represents expectation, π θ Represents the policy network model with respect to the parameter vector θ, a t represents the action taken at the current time t, s t represents the state at the current time t, π θ (a t |s t ) represents the action probability under the current strategy, π θold (a t |s t ) represents the action probability under the previous strategy, the subscript t represents the time step, A t represents the estimated value of the advantage function at time t, r t (θ) represents the importance weight, V(s t ) represents the state value function of the time index t∈[0,T], T represents the trajectory segment length, r t represents the total reward value at time t, and γ represents the discount factor.

[0080] S104: Add a truncation constraint to the improved objective function to avoid policy mutation, and obtain the objective function of the proximal policy optimization algorithm:

[0081] L CLIP (θ)=E t [min(r t (θ)A t ,clip(r t (θ),1-ε,1+ε)A t )]

[0082] Among them, L CLIP (θ) represents the objective function, min represents the minimum value, ε represents the truncation constant, and clip represents the truncation function.

[0083] S105: Derivative the objective function, calculate the gradient estimator, and use the gradient ascent algorithm to update the motion vector through the gradient estimator:

[0084]

[0085] θ new ←θ old +αg

[0086] Here, g represents the gradient estimator and α represents the learning rate.

[0087] S106: Determine the direction of motion by combining the objective function and the gradient estimator, and then generate multiple feasible waypoints for the rotorcraft in the flight environment:

[0088]

[0089]

[0090] in, represents the position coordinates of the rotorcraft UAV at time t, It represents the position coordinate of the UAV at time t+1, which is a feasible waypoint, and l represents the moving step length.

[0091] Specifically, the path planning problem is converted into a Markov decision process and expressed as a tuple model:<S,A,P,R> , where S is the state space of the system, A is the action space of the system, and P is the probability density function of the system state transition P(s′|s,a)=P(S t+1 =s′|S t =s,A t =a), which means the probability of the system transferring from the current state s to the next state s′ when action a is taken at time t. R:S×A→R is the reward function, which is used to evaluate the expected reward of the agent taking action a in state s.

[0092] The agent obtains the local state information s of its current location by interacting with the environment, and outputs the action a taken based on the current state through the policy network. After the agent executes the action, it obtains a new state s′. At the same time, the environment feedbacks reward information based on the agent's new state information. The agent obtains multiple tuple data by continuously exploring the environment and stores them in the memory bank. The agent updates the network parameters based on the data in the memory bank, and finally learns the strategy of avoiding obstacles and reaching the target point, thereby solving the path planning problem in indoor environments.

[0093] Policy gradient is an online learning algorithm. After an update, the policy network changes and cannot use the original data. It must interact with the environment again to collect new data. This results in low data utilization efficiency and a long learning time for the algorithm. Therefore, to overcome the above problems, the PPO algorithm adopts an importance sampling strategy to achieve data reuse and ensure that the expected values ​​of two sets of data with different distributions are equal:

[0094]

[0095] And by calculating the variance, it can be found that when performing importance sampling, if the probability density functions p(x) and q(x) of the two groups of distributions are slightly different, the variance of the two groups of data is also relatively small.

[0096] Var x~p [f(x)]=E x~p [f(x) 2 ]-(E x~p [f(x)]) 2 (5)

[0097]

[0098] Use the action probability π under the current policy θ (a|s) and the action probability π under the previous strategy θold The importance weight r is expressed as the ratio of (a|s) t (θ)=π θ (a t |s t ) / π θold (a t |s t ), r(θ old )=1. If the probability is r t (θ)>1, indicating that the probability of the current action under this strategy is higher than that of the previous strategy. t (θ)<1, the probability is lower than the previous strategy. Therefore, the improved objective function is designed as

[0099]

[0100] If there is no constraint L CPI Maximizing will lead to excessive policy updates, so in order to avoid policy mutations during parameter updates, the objective function must be constrained. The PPO algorithm uses two constraints, namely, adding a KL divergence regularization term to the objective function or adding a truncation operation to the gradient function to limit the policy update to a small range to improve the stability of the policy network. In practical applications, the truncation method is more effective and easier to implement. Therefore, the objective function of the PPO algorithm is designed to be

[0101] L CLIP (θ)=E t [min(r t (θ)A t ,clip(r t (θ),1-ε,1+ε)A t )] (8)

[0102] Where ε is the truncation constant used to assist in the range of policy updates, usually set to 0.1 or 0.2. The clip function is a truncation function that compares the old and new policies with the parameter r. t The value of (θ) is limited to the interval [1-ε,1+ε]. The objective function uses the min function to represent the smaller value between the probability ratio of the new and old strategies and the truncation function. When the advantage function A t When it is greater than 0, it means that the current action has a positive impact on the optimization target, and the probability of occurrence should be increased, but the update range should be limited to below 1+ε. t When it is less than 0, it means that the current action should be prevented and its probability is reduced to 1-ε. The core content of the PPO algorithm is to avoid using large policy updates to solve the problems of difficult to determine step size and low data utilization efficiency in the policy gradient algorithm.

[0103] State space and action space are important components of deep reinforcement learning algorithms. Combined with the needs of actual path planning tasks, the current position coordinates of the quadrotor (a x ,a y ), target position coordinates (g x ,g y ) and the position coordinates of n obstacles in the environment All of these are necessary information for planning and decision-making. In general, the state space needs to contain as much information as possible that is valuable for policy learning, but at the same time, problems such as high network complexity, non-convergence of training results, and poor decision-making timeliness caused by excessively high state space dimensions should be avoided.

[0104] The state space consists of two parts: the horizontal and vertical coordinate difference between the target position and the current position of the quadrotor (g x -ax ,g y -a y ), and the coordinate difference between the position of the two obstacles closest to the quadrotor and the current position So the state space is defined as

[0105]

[0106] Considering the requirements of the actual task, the horizontal movement direction of the rotorcraft is divided into 8 equal parts, with a heading angle interval of 45 degrees. At each decision point in the training process, the quadcopter can take a Move in a specific step length l in the direction, where the action decision variable a k ={0,1,2,3,4,5,6,7}, the action space is represented as A=[a k ×45°], the corresponding position coordinates at time t+1 are

[0107]

[0108] The intensive reward method is used to design the reward function for the UAV path planning task. According to the task requirements, the design of the reward function is divided into three parts:

[0109] a. To avoid collisions with obstacles or walls, design a reward function responsible for collision avoidance:

[0110]

[0111] When the distance d between the quadrotor and the obstacle obs Less than 0.5m or wall distance d wall If the distance is less than 1.5m, a collision is considered imminent and a negative reward is fed back to ensure that the quadrotor stays away from obstacles and walls during movement;

[0112] b. When the distance between the quadrotor and the target point is d tar When the distance is less than 0.5m, positive rewards are fed back to encourage the quadrotor to move towards the target point during the movement. The reward function is designed as

[0113]

[0114] c. Design a single-step reward function to avoid the problem of sparse rewards, that is, the environment needs to feedback rewards for each step of the quadrotor:

[0115] r step =ρ(d last -d current ) (13)

[0116] Where ρ is a weight coefficient greater than 0, d lastis the distance between the quadrotor and the target point at the last moment, d current Is the distance between the current moment and the target point, if d last >d current , r step is positive, otherwise it is negative;

[0117] In summary, after the quadrotor moves one step, the total reward of the environment feedback is

[0118] r total =r step +r reach +r collision (14)

[0119] Reference Manual Figure 2 , shows a schematic diagram of a simulation environment for a path planning task provided by an embodiment of the present invention. Figure 3 , shows the round reward curve in the model training process of a path planning task based on a proximal policy optimization algorithm provided by an embodiment of the present invention. Figure 4 , which shows a feasible path map generated by a model provided by an embodiment of the present invention in a training environment. Figure 5 , which shows a feasible path map generated by a model provided by an embodiment of the present invention under a test environment. Figure 6 , which shows a feasible path map generated by a fine-tuned model in a test environment provided by an embodiment of the present invention. Figure 7 , shows a feasible path map generated based on the A* algorithm provided by an embodiment of the present invention. Figure 8 , which shows a feasible path generated by a proximal strategy optimization algorithm and an A* algorithm provided by an embodiment of the present invention. Figure 9 , which shows a comparison diagram of the total length of feasible paths generated by a proximal strategy optimization algorithm provided by an embodiment of the present invention and the A* algorithm.

[0120] In the actual application process, computer simulation is used to verify the effectiveness of the path planning solution based on the proximal strategy optimization algorithm provided in this embodiment. The simulation environment is built based on the Coppeliasim platform under the win11x64 operating system. Figure 2 As shown, the algorithm and model training program are built based on the Python programming language and the Pytorch deep learning framework. Figure 2The cylinder in the figure is an obstacle, surrounded by walls. The PPO algorithm has a state space dimension of 6 and an action space dimension of 8. The actor network's input is the environment state, and its output is the next action to be executed. The hidden layer size is 64, and the Adam adaptive optimizer is used with a learning rate of 0.0003. The critic network's input is the environment state, and its output is the predicted value. The hidden layer size is 64, and the Adam adaptive optimizer is used with a learning rate of 0.0003.

[0121] The simulation results of the algorithm are as follows Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 、 Figure 7 、 Figure 8 and Figure 9 As shown, Figure 3 As shown in the figure, after 2200 rounds of training, the total reward curve of each round is obtained. After 400 rounds, the total reward obtained in each round basically reaches about 300, and the total reward curve rises very quickly, and can reach convergence results in fewer rounds of training time. Figure 4 、 Figure 5 and Figure 6 The following table shows the feasible paths generated by the model in the training environment after PPO algorithm training, with a 100% mission success rate. The feasible paths generated by the model in the test environment had a mission success rate that dropped to 86.7%. The feasible paths generated by the model after retraining and fine-tuning in the test environment also had a 100% mission success rate. However, these feasible paths were close to obstacles. Figure 7 and Figure 8 A visual comparison of the feasible paths generated by the A* algorithm and the PPO algorithm at the same start and end positions. Figure 9 The figure shows the comparison of the total length of feasible paths generated by the A* algorithm and the PPO algorithm under 5 different start and end positions. The difference between the two is acceptable. It can be seen that the path generated by the PPO algorithm meets the suboptimal requirements.

[0122] Reference Manual Figure 10 , shows a roadmap provided by an embodiment of the present invention after smoothing a feasible path generated by a PPO algorithm based on a B-spline curve.

[0123] S2: Convert feasible waypoints into a trackable continuous path using B-spline curves.

[0124] Among them, the B-spline curve is a linear combination of B-spline basis functions. Based on the B-spline curve, feasible waypoints are converted into trackable continuous paths, which makes up for the Runge phenomenon caused by the inability to modify the Bezier curve locally and the increase in curve order as the number of control points increases.

[0125] In actual application, the B-spline curve fitting method for multiple control points is as follows:

[0126] Define P0, P1, P2, ..., P n There are n+1 control points in total. The shape, direction and range of the curve can be defined by the control points. Then the expression of the k-order B-spline curve with n+1 control points is:

[0127]

[0128] Among them B i,k (u) is the i-th k-order B-spline basis function, and the control point P i Correspondingly, k≥1, u is the independent variable.

[0129] The basis functions have the Delbeux-Cox recursion:

[0130]

[0131] where [u0,u1,…,u k ,u k+1 ,…,u n ,u n+1 ,…,u n+k ] is a set of non-decreasing sequences called node vectors, the first and last values ​​of which are usually defined as 0 and 1, and it is agreed that 0 / 0=0.

[0132] The k-order B-spline curve is a k-1-order curve about the independent variable u, that is, the basis function B i,k (u) is the k-1 function of u, B i,k (u) involving u i ,u i+1 ,…,u i+k There are k+1 nodes and k intervals, so from B 0,k (u) to B n,k (u) There are n+k+1 nodes in total.

[0133] The B-spline curve is continuous. To find its derivative with respect to u, first we need to calculate the basis function B. i,k (u) Derivative:

[0134]

[0135] Substituting the derivative of the basis function into the B-spline curve expression, we can get:

[0136]

[0137] in,

[0138] In a possible implementation, the traceable continuous path is specifically:

[0139]

[0140]

[0141] Among them, P(u) represents a traceable continuous path, B i,k (u) represents the control point P i The corresponding i-th k-order B-spline basis function, u represents the independent variable, P′(u) represents the derivative of the traceable continuous path, B i+1,k-1 represents the i+1th k-1th order B-spline basis function that conforms to the Delbeux-Cox recursion, u i+k+1 and u i+1 They represent the i+k+1th node and the i+1th node respectively, Q i Represents the corresponding intermediate variable, P i+1 and P i denote the i+1th and ith feasible waypoints respectively.

[0142] In the actual application process, computer simulation is used to optimize the trajectory of the generated feasible path, and B-spline curves are used to convert feasible waypoints into trackable continuous parameter paths. The simulation is based on Matlab software under the win11x64 operating system. The simulation results are as follows Figure 10 As shown in Figure 3, the feasible waypoints are converted into third-order B-spline curves to achieve path smoothing.

[0143] Reference Manual Figure 11 , shows a geometric schematic diagram of an inertial coordinate system, a body coordinate system and a path reference coordinate system framework under a plane provided by an embodiment of the present invention.

[0144] S3: Establish the posture dynamics equations of the rotorcraft in the inertial coordinate system.

[0145] It's important to note that the established position and attitude dynamics equations for rotary-wing UAVs in an inertial coordinate system provide an important foundation for accurately simulating and controlling UAV flight. This mathematical model comprehensively describes the changes in the UAV's position, velocity, and attitude, facilitating the design of effective control algorithms, flight simulation verification, and a deeper understanding of the UAV's flight dynamics, thereby improving flight stability, efficiency, and safety.

[0146] In a possible implementation, the posture dynamics equation is established based on the Newton-Euler method, and the posture dynamics equation is specifically:

[0147]

[0148] p=[x,y,z] T

[0149]

[0150]

[0151]

[0152] ω=[p,q,r] T

[0153]

[0154]

[0155]

[0156] in, represents the derivative of the position coordinate p of the rotor UAV in the inertial coordinate system, v represents the speed of the rotor UAV in the inertial coordinate system, T represents the transpose, ξ represents the Euler angle of the rotor UAV, Derivatives of the Euler angles of the rotorcraft, φ,θ, They represent the rotation angles of the rotorcraft around the X-axis, Y-axis, and Z-axis in the inertial coordinate system, ω represents the angular velocity of the rotorcraft in the body coordinate system, p, q, r represent the rotation rates of the rotorcraft around the X-axis, Y-axis, and Z-axis of the body coordinate system whose origin is located at the center of mass of the rotorcraft, and W represents the kinematic Jacobian matrix, that is, the conversion relationship between the Euler angular velocity and the attitude angular velocity. represents the second-order derivative of the position coordinate p of the rotorcraft in the inertial coordinate system, that is, the acceleration, m represents the mass of the rotorcraft, u1 represents the lift of the rotorcraft, R represents the rotation matrix from the body coordinate system to the inertial coordinate system, g represents the acceleration of gravity, J and τ represent the inertia matrix and control torque of the rotorcraft, respectively.

[0157] Specifically, the dynamic model is derived based on the Newton-Euler method, such as Figure 11 As shown, three coordinate systems and vectors are defined, where is the inertial coordinate system, It is the body coordinate system with its origin at the center of mass of the rotorcraft. is a path reference coordinate system whose origin is at the desired path reference point.

[0158] Rotary wing UAV in inertial coordinate system The motion under the condition is divided into the translation of the center of mass and the rotation around the center of mass, where the position kinematics is

[0159]

[0160] where p = [x, y, z] T and In the coordinate system The position and velocity in .

[0161] Posture kinematics is

[0162]

[0163] in are Euler angles, representing the quadrotor rotating around the coordinate system The rotation angles of the X-axis, Y-axis, and Z-axis are ω=[p,q,r] T is the angular velocity, representing the quadrotor rotating around the coordinate system The rotation rates of the X-axis, Y-axis, and Z-axis are represented by the kinematic Jacobian matrix, which represents the conversion relationship between the Euler angular velocity and the attitude angular velocity.

[0164] According to Newton's second law, the coordinate system The position dynamics expressed in is

[0165]

[0166] Where m is the mass of the rotorcraft. u1 is the lift generated by the rotation of the four propellers. R is the coordinate system To the inertial system The rotation matrix of . is the acceleration.

[0167] According to the Euler dynamics equation of rigid body, the dynamics of the rotor UAV rotating around the center of mass is:

[0168]

[0169] Where J and τ are the inertia matrix and control torque of the quadrotor respectively.

[0170] S4: Combined with the pose dynamics model and the transformation relationship between the path reference coordinate system and the inertial coordinate system, the path tracking error dynamics equation of the rotorcraft in the path reference coordinate system is established.

[0171] The origin of the path reference coordinate system is the desired path reference point for tracing a continuous path.

[0172] It should be noted that by considering the path conversion error and establishing the path tracking error dynamic equation of the rotorcraft in the path reference coordinate system, the deviation between the actual path and the expected path of the UAV can be described more accurately, which provides an important basis for designing a more effective path tracking controller.

[0173] In one possible implementation, the path tracking error dynamics equation is specifically:

[0174]

[0175]

[0176] β=arctan(u,v)

[0177]

[0178]

[0179]

[0180] Among them, (x, y) represents the real-time position of the rotary-wing UAV, (x d ,y d ) represents the desired path reference point, It represents the offset angle of the speed direction of the rotor UAV in the X-axis direction of the inertial coordinate system, and β represents the offset angle of the speed direction of the rotor UAV in the X-axis direction of the body coordinate system. Indicates the heading angle of the rotorcraft, and Represents the rotation angle and rotation matrix from the inertial coordinate system to the path reference coordinate system, express The derivative of V t Indicates the plane flight speed of the rotorcraft, x e and y e They represent the X-axis position error and Y-axis position error of the rotorcraft in the path reference coordinates.

[0181] Specifically, assuming that the flying height of the rotor UAV remains unchanged, only its movement in the XY plane is considered. When the quadrotor moves to the coordinate system at a certain moment When the position {x, y} is below, its plane position motion is simplified to

[0182]

[0183] in is the plane flight speed of the quadrotor, is the velocity direction and coordinate system The X-axis direction of the offset angle, β = arctan (u, v) is the velocity direction and the coordinate system The offset angle in the X-axis direction, is the heading angle.

[0184] At this time, the corresponding expected path reference point P is {x d ,y d}, defined in the coordinate system The position error is

[0185]

[0186] in and From the coordinate system To coordinate system The rotation matrix and rotation angle.

[0187] like Figure 11 As shown, the quadrotor satisfies the geometric relationship when moving in a plane:

[0188]

[0189] in

[0190] Find the time derivative of equation (25):

[0191]

[0192] Where S(ω F ) is an antisymmetric matrix, and the parameters It is expressed in the coordinate system Next Relative to The angular velocity vector, Satisfies the following relationship:

[0193]

[0194] in is the reference point P relative to the coordinate system speed, is the curvature at the reference point P:

[0195]

[0196] Multiply both sides of the equation (26) by the matrix have

[0197]

[0198] in

[0199] After expanding Equation (29), the path tracking error dynamics equation expressed in the path reference coordinate system is obtained:

[0200]

[0201] S5: Combined with Lyapunov stability theory, determine the path parameter update law that can track the continuous path, the desired heading angle and the desired velocity of the desired attitude.

[0202] Lyapunov stability theory is a mathematical theory used to study the stability of dynamic systems. Here, it is used to analyze the stability of a drone tracking a desired path, ensuring that the drone remains stable and controllable during flight. The discount factor is a key parameter used to adjust the desired attitude and velocity in path-following control. By incorporating Lyapunov stability theory, a suitable path parameter update law can be determined, enabling the drone to quickly and smoothly track the desired path while maintaining system stability. The desired attitude refers to the desired posture the drone should assume during flight, including the desired heading angle and velocity. By incorporating Lyapunov stability theory, the desired heading angle and velocity for the desired attitude can be determined, ensuring good stability and controllability when the drone tracks the desired path. By incorporating Lyapunov stability theory, key parameters in path-following control can be determined, ensuring that the drone can stably and effectively track the desired path, improving flight safety and performance.

[0203] In a possible implementation, S5 specifically includes:

[0204] S501: Combined with the position error determined by the path tracking error dynamics equation, select the Lyapunov function:

[0205]

[0206] S502: Determine the path parameter update law based on the Lyapunov function Desired heading angle for desired attitude and the desired velocity v d :

[0207]

[0208]

[0209]

[0210] Among them, k x ,k y Indicates an adjustable parameter greater than 0, V d Indicates the expected flight speed.

[0211] Specifically, the control objective of the path tracking control problem is determined as

[0212] Therefore, a candidate Lyapunov function is designed:

[0213]

[0214] Find the time derivative of equation (31):

[0215]

[0216] in

[0217] because So yes Taking the absolute value satisfies the inequality:

[0218]

[0219] when Substitute equation (33) into Available for σ∈(0,1) has

[0220]

[0221] Therefore, the update law of the path parameters and the expected heading angle are selected as

[0222]

[0223]

[0224] Define desired speed

[0225] S6: With the goal of eliminating the tracking error of the path tracking error dynamic equation, the lift input controller of the attitude dynamic equation is designed according to the desired velocity, and the attitude torque controller is designed according to the desired heading angle.

[0226] It's important to note that the goal of designing the lift input controller and attitude torque controller for the pose dynamics equations is to eliminate path tracking errors and enable the drone to accurately track the desired path. The lift input controller adjusts lift based on the desired velocity, while the attitude torque controller adjusts attitude based on the desired heading angle. These two controllers work together to achieve stable and accurate path tracking, improving flight accuracy and reliability.

[0227] In a possible implementation, the lift input controller is specifically:

[0228]

[0229] Among them, R 13 ,R 23 ,R 33 Represents the corresponding element in the rotation matrix, K3=diag{k 11 ,k 22 ,k 33} indicates an adjustable parameter greater than 0.

[0230] Specifically, define the velocity tracking error ve =vv d , and then design the candidate Lyapunov function:

[0231]

[0232] Find the time derivative of equation (37):

[0233]

[0234] Designing a closed-loop controller Substituting into formula (38) we can get:

[0235]

[0236] Among them, K3∈R 3×3 is a diagonal matrix whose elements are all greater than 0.

[0237] That means the speed tracking error v e converges asymptotically to zero.

[0238] Therefore, the control input, namely the lift input controller, is designed:

[0239]

[0240] Reference Manual Figure 12 , which shows a comparison diagram of the actual motion path and the expected path of a rotary-wing UAV provided by an embodiment of the present invention.

[0241] Reference Manual Figure 13 , showing a comparison diagram of a speed and an expected speed provided by an embodiment of the present invention.

[0242] Reference Manual Figure 14 , showing a comparison diagram of a posture and an expected posture provided by an embodiment of the present invention.

[0243] Reference Manual Figure 15 , shows a graph showing the change of posture error over time provided by an embodiment of the present invention.

[0244] Reference Manual Figure 16 , which shows a comparison diagram of the actual motion path and the smoothed path of a rotary-wing UAV provided by an embodiment of the present invention.

[0245] In a possible implementation, the attitude torque controller is designed based on the backstepping method.

[0246] In one possible implementation, designing an attitude torque controller according to a desired heading angle specifically includes:

[0247] Determine the desired attitude of the rotary-wing UAV d :

[0248]

[0249] φ d =0

[0250] θ d =0

[0251]

[0252] Among them, φ d ,θ d and They represent the desired rotation angles of the rotorcraft around the X-axis, Y-axis, and Z-axis in the inertial coordinate system, respectively.

[0253] A command filter is introduced into the attitude loop of the rotorcraft UAV. The command filter is specifically:

[0254]

[0255] Where x1 and x2 represent the first state vector and the second state vector respectively, and They represent the derivatives of the first state vector and the second state vector respectively. The initial conditions of the first state vector and the second state vector are x1=ξ. d (0),x2=[0,0,0] T , η and ω n They represent the damping ratio and frequency respectively, and represent the command filter.

[0256] The first-order derivative and second-order derivative corresponding to the output of the command filter are used as the first-order derivative and second-order derivative of the desired posture:

[0257]

[0258] Among them, ξ c Indicates the command filter output, i.e., the command posture, represents the first-order derivative of the command filter output, Represents the second-order derivative of the command filter output.

[0259] Based on the output of the command filter and the corresponding first-order and second-order derivatives, the attitude torque controller is calculated, where the attitude torque controller includes the desired angular velocity and control torque:

[0260]

[0261]

[0262] Among them, ω d represents the desired angular velocity, represents the derivative of the desired angular velocity, τ represents the control torque, ξ e represents the posture tracking error, ω e represents the angular velocity tracking error, S(ω) represents the antisymmetric matrix, K1,K2∈R 3×3 Both represent diagonal matrices whose elements are all greater than 0.

[0263] It should be noted that the attitude stabilization controller is designed based on the backstepping method to achieve the desired attitude tracking and eliminate the path tracking error.

[0264] Specifically, assuming that the roll angle and pitch angle in planar motion are 0, the desired attitude is

[0265]

[0266] where φ d =0,θ d =0, heading angle

[0267] The first and second order derivatives of the desired attitude are relatively complex to solve by formula. The command filter is introduced into the attitude loop of the rotor UAV to solve the need to calculate the analytical derivatives. The existing technology has rigorously analyzed the impact of the command filter on the closed loop stability and performance. The structure of the command filter is

[0268]

[0269] The state vectors are x1 and x2, and the input is ξ d , η and ω n are the damping ratio and frequency, respectively.

[0270] The initial condition of the state vector of the command filter is x1 = ξ d (0),x2=[0,0,0] T The purpose of designing the filter is to use the output of the filter as the approximate first-order derivative and second-order derivative of the desired posture, and to ensure that it is infinitely close to the true value. Therefore, the output of the filter is defined as

[0271]

[0272] The command pose ξ generated by the command filter c To approximate the desired posture ξ d ,ξ c and ξ d The error between n Because the error between the two can be guaranteed to be small enough, in the design of the inner loop attitude torque controller of the rotor UAV, we choose to indirectly track ξ c, rather than directly tracking d Using the known ξ c 、 as well as Derived command angular velocity ω c and command angular acceleration satisfy

[0273]

[0274] Where W c =I 3×3 .

[0275] The attitude tracking error is ξ e =ξ-ξ c , design candidate Lyapunov functions:

[0276]

[0277] Find the time derivative of equation (45):

[0278]

[0279] Define the desired angular velocity Angular velocity tracking error ω e =ω-ω d , substituting into formula (46) we get:

[0280]

[0281] Design candidate Lyapunov functions:

[0282]

[0283] Find the time derivative of equation (48):

[0284]

[0285] Design control torque Substituting into formula (49) we can get:

[0286]

[0287] Among them, K1,K2∈R 3×3 is a diagonal matrix whose elements are all greater than 0.

[0288] That means the attitude tracking error ξ e and angular velocity tracking error ω e Asymptotically converges to zero, especially the heading angle error is globally uniformly asymptotically stable, and So the path tracking error x e ,ye converges asymptotically to zero.

[0289] For example, computer numerical simulation is used to verify the effectiveness of the path tracking control algorithm of the rotor UAV provided in this embodiment, and the simulation platform is based on Matlab software under the win11x64 operating system.

[0290] Consider the following physical parameters of the rotorcraft:

[0291] m=1.023(kg)

[0292] J=diag{0.0095, 0.0095, 0.0186} (kgm 2 )

[0293] The expression of the expected path is:

[0294]

[0295] The initial state of the rotorcraft is set to:

[0296] p(0)=[25,-10,10](m), v(0)=[0,0,0](m / s)

[0297] ξ(0)=[0.1,0.1,0.1](rad),ω(0)=[0,0,0](rad / s)

[0298] γ=0,V d =1m / s

[0299] The parameters of the designed controller are set as follows:

[0300] k x =0.85,k y =1

[0301] K1=diag{0.4, 0.4, 0.4}, K2=diag{0.4, 0.4, 0.4}, K3=diag{1.1, 1.1, 1.1}

[0302] η=0.707,ω n =5

[0303] The simulation results are as follows Figure 12-15 As shown, Figure 12-15 The dotted lines represent the corresponding expected values, and the solid lines represent the corresponding true values. Figure 12 The results showing the actual motion path and the expected path of the rotary-wing UAV indicate that the control algorithm provided by the embodiment of the present invention can control the rotary-wing UAV to track the expected trajectory. Figure 13It means that the speed of the rotor UAV in plane motion can converge to the desired speed quickly. Figure 14 and Figure 15 It means that the roll angle and pitch angle of the rotorcraft can converge to zero, the heading angle can track the expected heading angle, and the tracking error eventually converges to zero.

[0304] Next, the continuous parameter path smoothed by the B-spline curve is tracked to verify the effectiveness of the path planning and tracking control method proposed in this embodiment. The simulation platform is based on Matlab software under the win11x64 operating system.

[0305] The starting positions of the rotary wing drone are:

[0306] p start =[-10,-13.5],[-5,-13.5],[0,-13.5],[5,-13.5],[10,-13.5]

[0307] The simulation results are as follows Figure 16 As shown, Figure 16 The dashed line and the solid line represent the expected value and the actual value, respectively. Based on the control method proposed in this embodiment, the rotorcraft can well track the continuous parameter path generated by the path planning algorithm starting from different starting positions.

[0308] S7: Combined with the path parameter update law, the rotor UAV is controlled through the lift input controller and attitude torque controller.

[0309] As you can see, the rotorcraft is controlled through a lift input controller and an attitude torque controller, combined with the path parameter update law. This means the control system dynamically adjusts lift and attitude based on the real-time flight state and desired path to minimize path tracking error, ensuring the drone accurately follows the desired path and successfully completes its mission.

[0310] Specifically, this solution studies the path planning and tracking control of a rotorcraft UAV in constant-altitude flight. By simplifying and modeling the indoor flight environment, it proposes a solution to the path planning problem based on a proximal policy optimization algorithm. Transfer learning is introduced to enhance the generalization capability of the model, generating a series of feasible waypoints and converting them into trackable parameter paths using B-spline curves. The present invention designs a path tracking controller to enable the rotorcraft UAV to accurately track the parameter path generated by the planner. The path tracking error on the XY plane is uniformly represented in the path reference coordinate system. Based on the Lyapunov stability theory, the path parameter update law and the desired heading angle are derived. A closed-loop controller is designed for the speed tracking control problem, achieving accurate tracking of the desired speed. For the problem of tracking the desired attitude in the inner loop, an attitude tracking controller is designed based on the backstepping method. Furthermore, within the Lyapunov framework, it is rigorously proven that the proposed control scheme can ensure that the path tracking error converges to zero.

[0311] It should be noted that the designed path planning algorithm based on deep reinforcement learning addresses the slow convergence and weak generalization capabilities of traditional algorithms. It generates feasible waypoints in indoor environments and converts them into a trackable continuous path. Under the designed controller, the path tracking error of the quadrotor can converge asymptotically to zero, achieving the combined path planning and tracking control of the rotorcraft drone, meeting the operational requirements of fully autonomous planning and control in mission scenarios.

[0312] The beneficial effects brought about by the technical solutions provided by the embodiments of the invention include at least:

[0313] In this paper, a path planning algorithm based on a proximal strategy optimization algorithm is designed. This algorithm models the flight environment and generates a series of feasible waypoints. These waypoints are then converted into a parameterized, trackable continuous path using B-spline curves. This allows for local modification of the path planning problem, and its order does not increase with the addition of control points. Even with the addition of more control points, the complexity of the curve does not increase, thus avoiding the Runge phenomenon in Bezier curves. Based on this trackable continuous path, a closed-loop controller is designed, combining Lyapunov stability theory with the goal of eliminating the tracking error in the path tracking error dynamics equation. This controller is used to control a rotary-wing UAV. While improving the control accuracy of the rotary-wing UAV, it also ensures rapid convergence, stability, and feasibility, enabling the UAV to autonomously plan and control itself according to the flight environment.

[0314] Reference Manual Figure 17 , shows a structural schematic diagram of a rotary-wing UAV planning and control system provided by the present invention.

[0315] The present invention further provides a rotary-wing UAV planning and control system 20, which is applied to the above-mentioned rotary-wing UAV planning and control method, comprising:

[0316] Processor 201.

[0317] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201 , the planning and control method of the rotary-wing UAV in the method embodiment is implemented.

[0318] The rotary-wing UAV planning and control system 20 provided by the present invention can execute the above-mentioned rotary-wing UAV planning and control method and achieve the same or similar technical effects. To avoid repetition, the present invention will not go into details.

[0319] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0320] In this paper, a path planning algorithm based on a proximal strategy optimization algorithm is designed. This algorithm models the flight environment and generates a series of feasible waypoints. These waypoints are then converted into a parameterized, trackable continuous path using B-spline curves. This allows for local modification of the path planning problem, and its order does not increase with the addition of control points. Even with the addition of more control points, the complexity of the curve does not increase, thus avoiding the Runge phenomenon in Bezier curves. Based on this trackable continuous path, a closed-loop controller is designed, combining Lyapunov stability theory with the goal of eliminating the tracking error in the path tracking error dynamics equation. This controller is used to control a rotary-wing UAV. While improving the control accuracy of the rotary-wing UAV, it also ensures rapid convergence, stability, and feasibility, enabling the UAV to autonomously plan and control itself according to the flight environment.

[0321] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0322] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0323] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0324] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0325] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0326] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0327] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0328] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0329] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0330] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0331] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0332] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0333] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the planning and control method for a rotary-wing UAV as described in the method embodiment.

[0334] The computer-readable storage medium provided by the present invention can implement the steps and effects of the rotorcraft UAV planning and control method of the above-mentioned method embodiment. To avoid repetition, the present invention will not go into details.

[0335] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0336] In this paper, a path planning algorithm based on a proximal strategy optimization algorithm is designed. This algorithm models the flight environment and generates a series of feasible waypoints. These waypoints are then converted into a parameterized, trackable continuous path using B-spline curves. This allows for local modification of the path planning problem, and its order does not increase with the addition of control points. Even with the addition of more control points, the complexity of the curve does not increase, thus avoiding the Runge phenomenon in Bezier curves. Based on this trackable continuous path, a closed-loop controller is designed, combining Lyapunov stability theory with the goal of eliminating the tracking error in the path tracking error dynamics equation. This controller is used to control a rotary-wing UAV. While improving the control accuracy of the rotary-wing UAV, it also ensures rapid convergence, stability, and feasibility, enabling the UAV to autonomously plan and control itself according to the flight environment.

[0337] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

[0338] There are a few points to note:

[0339] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention. Other structures may refer to conventional designs.

[0340] (2) For the sake of clarity, the thickness of layers or regions in the drawings used to describe the embodiments of the present invention are exaggerated or reduced, that is, these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region, or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or intervening elements may be present.

[0341] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to form new embodiments.

[0342] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A planning and control method for a rotary-wing UAV, characterized in that: include: S1: generating feasible waypoints of the rotary-wing UAV in the flight environment based on a proximal strategy optimization algorithm; S2: Converting the feasible waypoints into a trackable continuous path using a B-spline curve; S3: Establishing the posture dynamic equation of the rotary wing UAV in the inertial coordinate system; S4: combining the pose dynamics equation and the conversion relationship between the path reference coordinate system and the inertial coordinate system, establishing a path tracking error dynamics equation of the rotorcraft in the path reference coordinate system, wherein the origin of the path reference coordinate system is the desired path reference point of the trackable continuous path; S5: Determine a path parameter update law of the trackable continuous path, a desired heading angle, and a desired velocity of the desired attitude in combination with Lyapunov stability theory; S6: with the goal of eliminating the tracking error of the path tracking error dynamic equation, designing a lift input controller of the posture dynamic equation according to the desired speed, and designing an attitude torque controller according to the desired heading angle; S7: In combination with the path parameter update law, the rotor UAV is controlled through the lift input controller and the attitude torque controller.

2. The planning and control method for a rotary-wing UAV according to claim 1, characterized in that: Said S1 specifically includes: S101: Based on the Markov decision process, a tuple model from the initial point to the target point in the flight environment is established, wherein the tuple model includes a state space, an action space, a probability density function, and a reward function: <S,A,P,R> A=[a k ×45°] a k ={0,1,2,3,4,5,6,7} R=r total =r step +r reach +r collision r step =ρ(d last -d current ),ρ>0 Where S represents the state space, A represents the action space, P represents the probability density function of state transition, R represents the reward function, (g x ,g y ) represents the target point coordinates, (a x ,a y ) represents the current position coordinates of the rotary-wing UAV, and Respectively represent the coordinates of the first obstacle and the second obstacle closest to the rotorcraft, a k represents the set of action decision variables, r total represents the total reward value of the flight environment feedback after the rotorcraft moves one step, r collision represents the reward function responsible for collision avoidance, d obs Indicates the distance between the rotorcraft and the obstacle, d wall Represents the distance between the rotor UAV and the wall, r reach represents the reward function responsible for traveling to the target point, d tar Represents the distance between the rotor UAV and the target point, r step represents the single-step reward function, d last Indicates the distance between the rotorcraft and the target point at the last moment, d current represents the distance between the rotary wing UAV and the target point at the current moment, and ρ represents the weight coefficient; S102: Determine the movement direction of the rotary-wing UAV: in, represents the direction of movement, a ki ∈a k ; S103: According to the tuple model, the policy gradient algorithm is improved in combination with the importance sampling strategy to obtain the improved objective function: A t =-V(s t )+r t +γr t+1 +…+γ T-t+1 r T-1 +γ T-t V(s T ) Among them, L CPI (θ) represents the improved objective function, E t represents expectation, π θ represents the policy network model with respect to the parameter vector θ, a t represents the action taken at the current time t, s t represents the state at the current time t, π θ (a t |s t ) represents the action probability under the current strategy, π θold (a t |s t ) represents the action probability under the previous strategy, the subscript t represents the time step, A t represents the estimated value of the advantage function at time t, r t (θ) represents the importance weight, V(s t ) represents the state value function of the time index t∈[0,T], T represents the trajectory segment length, r t represents the total reward value at time t, and γ represents the discount factor; S104: Add truncation constraints to the improved objective function to avoid policy mutations, and obtain the objective function of the proximal policy optimization algorithm: L CLIP (θ)=E t [min(r t (i)A t ,clip(r t (θ),1-ε,1+ε)A t )] Among them, L CLIP (θ) represents the objective function, min represents the minimum value, ε represents the truncation constant, and clip represents the truncation function. S105: Derivative the objective function, calculate the gradient estimator, and update the motion vector using the gradient estimator using a gradient ascent algorithm: i new ←θ old +ag Wherein, g represents the gradient estimator, and α represents the learning rate; S106: Determine the movement direction by combining the objective function and the gradient estimator, and then generate a plurality of feasible waypoints for the rotary-wing UAV in the flight environment: in, represents the position coordinates of the rotary wing UAV at time t, represents the position coordinate of the rotary wing UAV at time t+1, i.e., a feasible waypoint, and l represents the moving step length.

3. The planning and control method for a rotary-wing UAV according to claim 1, characterized in that: The traceable continuous path is specifically: Wherein, P(u) represents the traceable continuous path, B i,k (u) represents the control point P i The corresponding i-th k-order B-spline basis function, u represents the independent variable, P′(u) represents the derivative of the traceable continuous path, B i+1,k-1 represents the i+1th k-1th order B-spline basis function that conforms to the Delbeux-Cox recursion, u i+k+1 and u i+1 They represent the i+k+1th node and the i+1th node respectively, Q i Represents the corresponding intermediate variable, P i+1 and P i denote the i+1th and ith feasible waypoints respectively.

4. The planning and control method for a rotary-wing UAV according to claim 1, characterized in that: The posture dynamics equation is established based on the Newton-Euler method, and the posture dynamics equation is specifically: p=[x,y,z] T ω=[p,q,r] T e3=[0,0,1] T in, represents the derivative of the position coordinate p of the rotary-wing UAV in the inertial coordinate system, v represents the velocity of the rotary-wing UAV in the inertial coordinate system, T represents the transpose, ξ represents the Euler angle of the rotary-wing UAV, Derivatives of the Euler angles of the rotorcraft, φ, θ, They represent the rotation angles of the rotorcraft around the X-axis, Y-axis, and Z-axis in the inertial coordinate system, ω represents the angular velocity of the rotorcraft in the body coordinate system, p, q, and r represent the rotation rates of the rotorcraft around the X-axis, Y-axis, and Z-axis of the body coordinate system whose origin is located at the center of mass of the rotorcraft, respectively. W represents the kinematic Jacobian matrix, i.e., the conversion relationship between the Euler angular velocity and the attitude angular velocity. represents the second-order derivative of the position coordinate p of the rotorcraft in the inertial coordinate system, i.e., acceleration, m represents the mass of the rotorcraft, u1 represents the lift of the rotorcraft, R represents the rotation matrix from the body coordinate system to the inertial coordinate system, g represents the acceleration of gravity, J and τ represent the inertia matrix and control torque of the rotorcraft, respectively.

5. The planning and control method for a rotary-wing UAV according to claim 1, characterized in that: The path tracking error dynamic equation is specifically: β=arctan(u,v) Wherein, (x, y) represents the real-time position of the rotary wing UAV, (x d ,y d ) represents the desired path reference point, represents the offset angle of the speed direction of the rotary-wing UAV in the X-axis direction of the inertial coordinate system, β represents the offset angle of the speed direction of the rotary-wing UAV in the X-axis direction of the body coordinate system, Indicates the heading angle of the rotorcraft, and Represents the rotation angle and rotation matrix from the inertial coordinate system to the path reference coordinate system, express The derivative of V t represents the plane flight speed of the rotor UAV, x e and y e They respectively represent the X-axis position error and the Y-axis position error of the rotorcraft in the path reference coordinates.

6. The planning and control method for a rotary-wing UAV according to claim 1, characterized in that: The S5 specifically includes: S501: Based on the position error determined by the path tracking error dynamics equation, select the Lyapunov function: S502: Determine the path parameter update law according to the Lyapunov function The desired heading angle of the desired attitude and the desired speed v d : Among them, k x ,k y Indicates an adjustable parameter greater than 0, V d Indicates the expected flight speed.

7. The planning and control method for a rotary-wing UAV according to claim 1, characterized in that: The lift input controller is specifically: Among them, R 13 ,R 23 ,R 33 Represents the corresponding element in the rotation matrix, K3=diag{k 11 ,k 22 ,k 33 } indicates an adjustable parameter greater than 0.

8. The planning and control method for a rotary-wing UAV according to claim 1, characterized in that: The attitude torque controller is designed based on the backstepping method.

9. The planning and control method for a rotary-wing UAV according to claim 8, characterized in that: The designing of the attitude torque controller according to the desired heading angle specifically includes: Determine the desired attitude of the rotary wing UAV d : f d =0 i d =0 Among them, φ d ,θ d and represent the desired rotation angles of the rotary-wing UAV around the X-axis, Y-axis, and Z-axis in the inertial coordinate system, respectively; A command filter is introduced into the attitude loop of the rotary wing UAV. The command filter is specifically: Where x1 and x2 represent the first state vector and the second state vector respectively, and They represent the derivatives of the first state vector and the second state vector respectively, and the initial conditions of the first state vector and the second state vector are x1=ξ d (0),x2=[0,0,0] T , η and ω n denote the damping ratio and frequency respectively, indicating the command filter; The first-order derivative and the second-order derivative corresponding to the output of the instruction filter are used as the first-order derivative and the second-order derivative of the desired posture: Among them, ξ c Indicates the command filter output, i.e., the command posture, represents the first-order derivative of the command filter output, represents the second-order derivative of the command filter output; Based on the output of the command filter and the corresponding first-order derivatives and second-order derivatives, the attitude torque controller is calculated, wherein the attitude torque controller includes the desired angular velocity and the control torque: Among them, ω d represents the desired angular velocity, represents the derivative of the desired angular velocity, τ represents the control torque, ξ e represents the posture tracking error, ω e represents the angular velocity tracking error, S(ω) represents the antisymmetric matrix, K1,K2∈R 3×3 Both represent diagonal matrices whose elements are all greater than 0.

10. A planning and control system for a rotary wing UAV, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the planning and control method for a rotary-wing UAV according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Unmanned airship error limited moving path tracking control method

    CN117055604A

  • Method in which small fixed-wing unmanned aerial vehicle follows path and LGVF path-following controller using same

    US20210311503A1