A hybrid control strategy method and system for UAVs in highly dynamic and complex environments
By combining hybrid control strategies with model prediction control and reinforcement learning, the problems of low control accuracy, insufficient real-time performance and poor robustness in highly dynamic and complex environments are solved, and efficient and accurate real-time control of the drone is achieved.
Patent Information
- Application Number
- CN202510877609.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The existing UAV control technology has problems such as low control accuracy, insufficient real-time performance and poor robustness in highly dynamic and complex environments.
A hybrid control strategy combining model prediction control and reinforcement learning is adopted, and a hybrid control strategy framework is built by collecting drone status and environmental information in real time. The model prediction control module is used to calculate the short-term optimal control action sequence. The reinforcement learning module generates long-term optimization target state trajectory, and dynamically adjusts the weights through the strategy fusion module to generate mixed control instructions.
It realizes efficient and accurate real-time control of drones in complex environments, improving control accuracy and robustness.
Smart Images

Figure CN120370723B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) control, and in particular relates to a hybrid control strategy method and system for UAVs in highly dynamic and complex environments. Background Art
[0002] In recent years, drone technology has been widely used in industrial inspections, logistics and transportation, disaster relief, military reconnaissance, and other fields. In practical applications, drones are often required to complete high-precision, real-time flight missions in complex, dynamic, and uncertain environments, which places high demands on the drone's autonomous control strategies. Traditional drone control methods, such as PID control or simple trajectory tracking control, often struggle to maintain good control performance in highly complex and dynamic environments. This is particularly true when faced with dynamic obstacles, wind disturbances, sensor noise, or sudden changes in the flight environment, where control performance can deteriorate.
[0003] Model Predictive Control (MPC), an advanced real-time optimization control method, utilizes system dynamics models and optimization algorithms to calculate the optimal sequence of control actions in the short term within a prediction time window. This effectively addresses constraints and achieves high control accuracy. However, MPC generally relies on an accurate system model and has a short prediction window, making it difficult to optimize long-term flight performance. Furthermore, in highly uncertain environments, MPC is highly sensitive to model accuracy, and the accumulation of prediction errors can lead to a decline in control performance.
[0004] Reinforcement learning (RL), as a data-driven control strategy, can continuously learn and optimize policies through interaction with the environment, making it suitable for handling complex environments and long-term performance optimization problems. However, RL methods typically have high training costs and poor real-time performance, making it difficult to guarantee the real-time and accuracy of control action calculations in isolation.
[0005] In summary, existing UAV control technology still has problems such as low control accuracy, insufficient real-time performance, and poor robustness when dealing with highly dynamic and complex environments. It is urgent to propose a real-time hybrid control method and system that integrates the advantages of model predictive control and reinforcement learning strategies to improve the flight control performance of UAVs in complex environments. Summary of the Invention
[0006] The purpose of the present invention is to provide a hybrid control strategy method for UAVs in highly dynamic and complex environments, aiming to solve the problems of low control accuracy, insufficient real-time performance and poor robustness in existing UAV control technologies when dealing with highly dynamic and complex environments.
[0007] The present invention is implemented as follows: a hybrid control strategy method for unmanned aerial vehicles in a highly dynamic and complex environment, the method comprising:
[0008] S1, collects the current flight status information and environmental information of the UAV in real time, and performs preprocessing to generate standardized input data;
[0009] S2. Build a hybrid control strategy framework for UAVs, including a model predictive control module, a reinforcement learning strategy module, and a strategy fusion module;
[0010] S3, the model predictive control module predicts short-term state changes within the prediction time window based on the current input data according to the UAV dynamics model, and uses sequential quadratic programming, interior point method or fast gradient descent algorithm to calculate and output the short-term optimal control action sequence in real time;
[0011] S4, the reinforcement learning strategy module uses one of the pre-trained deep deterministic policy gradient algorithm, proximal policy optimization algorithm, or soft actor-critic algorithm to calculate the long-term optimal target state trajectory of the drone in real time based on the current input data;
[0012] S5. The strategy fusion module calculates the environmental complexity index in real time, dynamically adjusts the output weights of the reinforcement learning strategy module and the model predictive control module based on the environmental complexity index, uses the long-term optimization target state trajectory as a constraint or reference information, guides the calculation of short-term control actions, and generates hybrid control instructions;
[0013] S6. Send the hybrid control instructions to the UAV flight control actuator in real time to control the attitude and trajectory of the UAV in real time.
[0014] Preferably, the step of calculating the short-term optimal control action by the model predictive control module includes:
[0015] Using the dynamic model, the current state The prediction is started to generate multiple candidate state trajectories, which are expressed as:
[0016] ;
[0017] in, Indicates time When , the state in the next N steps is predicted based on the current state, and the evolution of the predicted state satisfies the dynamic model constraints:
[0018] ;
[0019] in, is the dynamic state transfer function of the UAV, is the control input to be optimized;
[0020] Define the cost function based on the candidate state trajectory:
[0021] ;
[0022] in, Indicates at time The desired reference state; , is the weight matrix of state deviation and control energy consumption; is the weight matrix of the terminal state deviation; is the length of the prediction time window; represents the penalty term of the control input; Indicates that at the current moment Expected in the future The reference state that the system should reach at any moment; Indicates time The moment prediction The state of the moment;
[0023] Use optimization algorithms to optimize the problem and obtain the optimal control action sequence within the prediction time window :
[0024] ;
[0025] in, Represents the control input sequence and outputs the first control action of the optimal control action sequence in the short term in real time , and use it as the execution control action of the model predictive control module in the current control cycle.
[0026] Preferably, the step of calculating the long-term optimization target trajectory by the reinforcement learning strategy module includes:
[0027] Build and pre-train reinforcement learning policy network , the network is based on the current state of the drone and environmental information As input, the optimal target state trajectory is output, and the policy network is expressed as:
[0028] ;
[0029] Among them, the network parameters Obtained through pre-training of deep reinforcement learning algorithms;
[0030] The reinforcement learning policy network is trained using deep deterministic policy gradient algorithms, proximal policy optimization, or soft actor-critic algorithms.
[0031] In the online control phase, based on the real-time state input, the trained reinforcement learning strategy network directly outputs the long-term optimal target state trajectory of the UAV:
[0032] ;
[0033] in: Represents the current policy network The parameter vector of Indicates that the parameter is The policy network function.
[0034] Preferably, a deep deterministic policy gradient algorithm is used to train the reinforcement learning policy network, and the steps include:
[0035] Step a: Initialize the reinforcement learning policy network Actor parameters , value network critic parameter , and set the target network parameters:
[0036] ;
[0037] in, Represents the parameters of the current Actor network, Represents the parameters of the target Actor network, is the parameter of the current Critic network, Represents the parameters of the target Critic network;
[0038] Step b: The drone interacts in the environment and selects actions based on the policy network:
[0039] ;
[0040] in, To explore noise;
[0041] Step c: Execute actions in the environment , observe the next state and rewards , and store the experience tuple in the experience replay buffer :
[0042] ;
[0043] Step d: Randomly extract small batches of data from the experience replay buffer for training, and update the critic network parameters by minimizing the critic network's loss function:
[0044] ;
[0045] in, Indicates the current state, represents the action, generated by the current policy network; Indicates the current reward; Indicates the state at the next moment; For the experience replay pool; Represents the current Critic network's valuation of the state-action pair value; Target Value, which is defined as:
[0046] ;
[0047] in, is the discount factor; is the current reward value; Represents the target Critic network's estimate of the action value of the next state; Indicates that the target Actor network is in state Next, select the action; Indicates the next state;
[0048] Step e: Update the reinforcement learning policy network Actor network parameters , optimize the expected return of the policy network by gradient ascent:
[0049] ;
[0050] Step f: Periodically soft-update target network parameters:
[0051] ;
[0052] Where, is the soft update coefficient, .
[0053] Preferably, step S5 includes:
[0054] Real-time computing environment complexity index , expressed as:
[0055] ;
[0056] in, The current speed of the drone; is the obstacle density in the current environment; is the UAV state prediction error;
[0057] , , is a predetermined weight coefficient used to adjust the contribution ratio of different indicators to complexity;
[0058] According to the complexity index of real-time calculation , dynamically adjust the output weight of the reinforcement learning strategy module and model predictive control module output weights , the weights satisfy the constraints:
[0059] ;
[0060] The weight adjustment rules are:
[0061] ;
[0062] in, is the Sigmoid function; Preset thresholds for environmental complexity; Adjust the sensitivity coefficient for the weight;
[0063] Fusion reinforcement learning module output long-term goal state trajectory The target state trajectory corresponding to the short-term optimal control action sequence output by the model predictive control module , obtain the reference state trajectory after fusion , specifically integrated as follows:
[0064] ;
[0065] The fused trajectory As the model predictive control module calculates the reference trajectory of the short-term control action sequence and outputs the final hybrid control command:
[0066] ;
[0067] in, is the current drone status; It represents the process of executing the model predictive control algorithm to solve the control action using the fused trajectory as the reference input.
[0068] Preferably, the step of calculating the environmental complexity index includes calculating the drone state prediction error trend, including:
[0069] Calculate the current prediction error in real time , defined as the deviation between the actual observed state of the drone and the predicted state:
[0070] ;
[0071] in, The current actual measured drone status; The current state of the drone predicted for the previous control cycle; represents the Euclidean norm of the state vector, Indicates the current control cycle number;
[0072] Real-time calculation of prediction error trends based on sliding window method , defined as the most recent consecutive The linear regression slope of the forecast error series at each moment:
[0073] Assume that the forecast error sequence is , then the linear regression calculation of the error trend is:
[0074] ;
[0075] in, is the sliding window size, which indicates the number of time series error points used to calculate the trend; When , the error tends to increase and the system uncertainty increases; when When , it indicates that the error tends to decrease and the system uncertainty decreases;
[0076] Forecast Error Trend Indicator Incorporating environmental complexity indicators The extended definition of the computational complexity index is:
[0077] ;
[0078] in, is the weight coefficient of the forecast error trend;
[0079] According to the updated complexity index Adjust the fusion weight of the reinforcement learning strategy module and the model predictive control module in real time. The fusion weight adjustment adopts:
[0080] .
[0081] Preferably, the UAV dynamics model is a data-driven neural network dynamics model or an explicit dynamics model based on physical parameters. When the data-driven neural network dynamics model is adopted:
[0082] Using historical flight datasets Training a neural network model , so that the state prediction error is minimized, and the optimization objective function is defined as:
[0083] ;
[0084] in, For drones at all times Status; For the moment Control input; For the moment the actual status of is the neural network to be trained, is the network parameter; is the Euclidean norm of the state vector prediction error;
[0085] Real-time prediction of the drone's status at the next moment, using a data-driven neural network model:
[0086] ;
[0087] In the case of a state transfer function constructed using a neural network prediction model, Indicates time The predicted The state of the moment.
[0088] Preferably, when an explicit physical parameter dynamic model is used:
[0089] According to the rigid body dynamics model of the UAV, the kinematic equation is defined as follows:
[0090] Kinematic equations of position and velocity:
[0091] ;
[0092] in, is the first derivative of position, indicating that the change of position with respect to time is velocity; is the first-order derivative of velocity, indicating that the linear acceleration of the drone is composed of the linear acceleration converted from gravity acceleration and thrust;
[0093] Attitude kinematic equation:
[0094] ;
[0095] in: Represents the first derivative of angular velocity, describing the change of angular velocity of a rigid body over time, including torque input and gyroscopic force; is the derivative of the attitude angle;
[0096] 、 Respectively represent the position and speed status of the drone; is the gravitational acceleration vector; For drone quality; is the attitude rotation matrix, where , , Represent the pitch angle, roll angle and yaw angle respectively; is the total thrust of the UAV; is the angular velocity of the drone, is the Euler angle posture; is the inertia matrix; is the total moment acting on the UAV; is the conversion matrix from angular velocity to attitude angular velocity; symbol represents vector cross product;
[0097] Real-time prediction of the drone state at the next moment, based on the dynamic model of physical parameters, using numerical integration method:
[0098] ;
[0099] in, is the prediction time step; is the state derivative based on the explicit physical model;
[0100] Using explicit physics modeling in state derivatives based on rigid body dynamics models When the prediction is made by numerical integration, Indicates time The predicted The state of the moment.
[0101] Another object of the present invention is to provide a hybrid control strategy system for unmanned aerial vehicles in a highly dynamic and complex environment, the system comprising:
[0102] Status information acquisition module, used to collect and pre-process the UAV flight status and environment information in real time;
[0103] Model predictive control module, which is used to generate short-term optimal control action sequences in real time based on state information and dynamic models;
[0104] Reinforcement learning strategy module, used to generate long-term optimization target state trajectory in real time based on state information;
[0105] The strategy fusion module is used to use the output of the reinforcement learning strategy module as constraints or reference information for the model predictive control module to generate hybrid control instructions;
[0106] The flight control execution module is used to control the UAV's flight attitude and trajectory in real time based on hybrid control instructions.
[0107] Preferably, the state information acquisition module includes at least one of an inertial measurement unit, a visual synchronous positioning and mapping module, a laser radar and an ultrasonic sensor.
[0108] The present invention provides a hybrid control strategy method for unmanned aerial vehicles in highly dynamic and complex environments, which adopts a data-driven neural network model or an explicit physical parameter dynamics model. The data-driven model is trained through historical flight data sets with the goal of minimizing state prediction errors; while the explicit dynamics model adopts rigid body dynamics equations to predict the next moment state of the unmanned aerial vehicle through numerical integration. Through the unmanned aerial vehicle hybrid control strategy system on the edge computing platform, including the above-mentioned various functional modules, efficient and precise real-time control of the unmanned aerial vehicle in complex environments is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0109] Figure 1 This is a flow chart of a hybrid control strategy method for a UAV in a highly dynamic and complex environment according to an embodiment of the present invention;
[0110] Figure 2 This is an architectural diagram of a hybrid control strategy system for drones in highly dynamic and complex environments according to an embodiment of the present invention. DETAILED DESCRIPTION
[0111] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0112] like Figure 1 FIG. 1 is a flow chart of a hybrid control strategy method for a UAV in a highly dynamic and complex environment according to an embodiment of the present invention. The method includes:
[0113] S1. Collect the current flight status information and environmental information of the UAV in real time and perform preprocessing to generate standardized input data.
[0114] In this step, the drone's current flight status and environmental information are collected in real time. Specifically, the drone's onboard status information collection module uses at least one device, such as an inertial measurement unit (IMU), a simultaneous localization and mapping (SLAM) module, a lidar (LiDAR), or an ultrasonic sensor, to obtain real-time status information such as the drone's speed, position, and attitude, as well as information such as the location and density of obstacles in the environment. This acquired data is preprocessed to form standardized input data, including denoising, filtering, and data normalization, to ensure the real-time and reliability of the input data.
[0115] S2. Build a hybrid control strategy framework for UAVs, including a model predictive control module, a reinforcement learning strategy module, and a strategy fusion module.
[0116] In this step, we build a hybrid control strategy framework for the drone. This framework consists of a model predictive control (MPC) module, a reinforcement learning (RL) strategy module, and a strategy fusion module. The MPC module is responsible for developing short-term, real-time control strategies, while the RL module generates long-term, optimized target trajectories. The fusion module integrates these two strategies in real time.
[0117] S3, the model predictive control module predicts short-term state changes within the prediction time window based on the current input data according to the UAV dynamics model, and uses sequential quadratic programming, interior point method or fast gradient descent algorithm to calculate and output the short-term optimal control action sequence in real time.
[0118] In this step, the dynamics model specifically includes a data-driven neural network dynamics model or an explicit dynamics model based on physical parameters. The specific algorithm steps are:
[0119] (1) When using a data-driven neural network dynamics model:
[0120] Using historical flight datasets Training a neural network model , so that the state prediction error is minimized, and the optimization objective function is defined as:
[0121] ;
[0122] in, For drones at all times Status; For the moment Control input; For the moment the actual status of is the neural network to be trained, is the network parameter; is the Euclidean norm of the state vector prediction error;
[0123] (2) When using an explicit physical parameter dynamic model:
[0124] According to the rigid body dynamics model of the UAV, the kinematic equation is defined as follows:
[0125] Kinematic equations of position and velocity:
[0126] ;
[0127] Attitude kinematic equation:
[0128] ;
[0129] in:
[0130] 、 Respectively represent the position and speed status of the drone; is the gravitational acceleration vector; For drone quality; is the attitude rotation matrix, where , , Represent the pitch angle, roll angle and yaw angle respectively;
[0131] is the total thrust of the UAV; is the angular velocity of the drone, is the Euler angle posture; is the inertia matrix; is the total moment acting on the UAV; is the conversion matrix from angular velocity to attitude angular velocity; symbol represents vector cross product;
[0132] (3) Real-time prediction of the drone’s status at the next moment:
[0133] Data-driven neural network model prediction:
[0134] ;
[0135] In the case of a state transfer function constructed using a neural network prediction model, Indicates time The predicted The state of the moment.
[0136] Kinetic models based on physical parameters, using numerical integration methods (e.g. Euler method) to predict:
[0137] ;
[0138] in, is the prediction time step; is the state derivative based on the explicit physical model;
[0139] Using explicit physics modeling in state derivatives based on rigid body dynamics models When the prediction is made by numerical integration, Indicates time The predicted The state of the moment.
[0140] The model predictive control module generates short-term control action sequences in real time based on the UAV's dynamic model. Specifically, the dynamic model is used to predict the current state and generate multiple candidate state trajectories. The candidate state evolution is calculated as follows:
[0141] ;
[0142] Define the cost function To evaluate the deviation between the predicted trajectory and the desired trajectory and the energy consumption of the control action:
[0143] ;
[0144] Use sequential quadratic programming (SQP), interior point method or fast gradient descent algorithm to solve the optimization problem, obtain the optimal control action sequence, and output the first control action in real time.
[0145] The short-term optimal control action calculation of the model predictive control module includes:
[0146] (1) Using the dynamic model, the current state The prediction is started to generate multiple candidate state trajectories, which are expressed as:
[0147] ;
[0148] Among them, the evolution of the predicted state satisfies the dynamic model constraints:
[0149] ;
[0150] in, is the dynamic state transfer function of the UAV, is the control input to be optimized;
[0151] (2) Define a cost function based on the candidate state trajectory to measure the deviation between the predicted state sequence and the desired state sequence and the energy consumption or cost of the control action:
[0152] ;
[0153] in, Indicates at time The desired reference state; , is the weight matrix of state deviation and control energy consumption; is the weight matrix of the terminal state deviation; is the length of the prediction time window; represents the penalty term of the control input; Indicates that at the current moment Expected in the future The reference state that the system should reach at any moment; Indicates time The moment prediction The state of the moment;
[0154] (3) Using an optimization algorithm, such as sequential quadratic programming (SQP), interior point method, or fast gradient descent algorithm, solve the following optimization problem to obtain the optimal control action sequence within the prediction time window: :
[0155] ;
[0156] (4) Real-time output of the first control action of the optimal control action sequence in the short term , and use it as the execution control action of the model predictive control module in the current control cycle.
[0157] S4, the reinforcement learning strategy module uses one of the pre-trained deep deterministic policy gradient algorithm, proximal policy optimization algorithm, or soft actor-critic algorithm to calculate the long-term optimal target state trajectory of the UAV in real time based on the current input data.
[0158] In this step, the reinforcement learning policy module generates the long-term optimized target state trajectory in real time. A reinforcement learning policy network is pre-built and trained. This network uses one of the Deep Deterministic Policy Gradient (DDPG), Proximal Policy Optimization (PPO), or Soft Actor Critic (SAC) algorithms. It takes the current state of the drone and the environment as input and outputs the long-term optimized target state trajectory.
[0159] For example, taking the DDPG algorithm as an example, the steps for training a reinforcement learning network include: initializing the parameters of the Actor network and the Critic network, collecting experience tuples through environmental interaction, storing them in the experience replay buffer, randomly extracting small batches of data from them to update the parameters of the Critic network and the Actor network, and regularly performing soft updates on the target network to ultimately achieve effective network training.
[0160] The long-term optimization target trajectory calculation of the reinforcement learning policy module includes:
[0161] (1) Build and pre-train reinforcement learning strategy network , the network is based on the current state of the drone and environmental information As input, the optimal target state trajectory is output, and the policy network is expressed as:
[0162] ;
[0163] Among them, the network parameters Obtained through pre-training of deep reinforcement learning algorithms;
[0164] (2) The training process of the reinforcement learning policy network adopts one of the following reinforcement learning optimization methods, including but not limited to deep deterministic policy gradient (DDPG), proximal policy optimization (PPO) or soft actor critic (SAC) algorithm; specifically, taking the DDPG algorithm as an example, the training steps are:
[0165] Step a: Initialize the policy network (Actor) parameters , Value Network (Critic) Parameters , and set the target network parameters:
[0166] ;
[0167] Step b: The drone interacts in the environment and selects actions based on the policy network:
[0168] ;
[0169] in, To explore noise;
[0170] Step c: Execute actions in the environment , observe the next state and rewards , and store the experience tuple in the experience replay buffer :
[0171] ;
[0172] Step d: Randomly extract small batches of data from the experience replay buffer for training, and update the critic network parameters by minimizing the critic network's loss function:
[0173] ;
[0174] Among them, the target value Defined as:
[0175] ;
[0176] in, is the discount factor;
[0177] Step e: Update the reinforcement learning policy network Actor network parameters , optimize the expected return of the policy network by gradient ascent:
[0178] ;
[0179] Step f: Periodically soft-update target network parameters:
[0180] ;
[0181] Where, is the soft update coefficient, ;
[0182] (3) In the online control phase, based on the real-time state input, the trained strategy network directly outputs the long-term optimal target state trajectory of the UAV:
[0183] ;
[0184] in: Represents the current policy network The parameter vector of Indicates that the parameter is The policy network function.
[0185] S5. The strategy fusion module calculates the environmental complexity index in real time, dynamically adjusts the output weights of the reinforcement learning strategy module and the model predictive control module based on the environmental complexity index, uses the long-term optimization target state trajectory as a constraint or reference information, guides the calculation of short-term control actions, and generates hybrid control instructions.
[0186] In this step, the strategy fusion module dynamically adjusts the output weights of reinforcement learning and model predictive control based on real-time environmental complexity indicators. Specific environmental complexity indicators include drone speed, obstacle density, state prediction error, and prediction error trend analysis.
[0187] The target state trajectories output by the reinforcement learning module and the model predictive control module are fused into a new reference trajectory to further guide the real-time calculation of short-term control actions.
[0188] The specific implementation of the strategy fusion module is as follows:
[0189] (1) Real-time computing environment complexity index , specifically expressed as:
[0190] ;
[0191] in, The current speed of the drone; is the obstacle density in the current environment; is the UAV state prediction error;
[0192] , , is a predetermined weight coefficient used to adjust the contribution ratio of different indicators to complexity;
[0193] (2) Based on the complexity index of real-time calculation , dynamically adjust the output weight of the reinforcement learning strategy module and model predictive control module output weights , the weights satisfy the constraints:
[0194] ;
[0195] The weight adjustment rules are:
[0196] ;
[0197] in, is the Sigmoid function; Preset thresholds for environmental complexity; Adjust the sensitivity coefficient for the weight;
[0198] (3) Long-term target state trajectory output by the fusion reinforcement learning module The target state trajectory corresponding to the short-term optimal control action sequence output by the model predictive control module , obtain the reference state trajectory after fusion , specifically integrated as follows:
[0199] ;
[0200] (4) The fused trajectory As the model predictive control module calculates the reference trajectory of the short-term control action sequence and outputs the final hybrid control command:
[0201] ;
[0202] in, is the current drone status; It represents the process of executing the model predictive control algorithm to solve the control action using the fused trajectory as the reference input.
[0203] The strategy fusion module calculates environmental complexity indicators in real time, including but not limited to the UAV's movement speed, obstacle density in the environment, and state prediction error. It dynamically adjusts the output weights of the reinforcement learning strategy module and the model predictive control module based on the complexity indicators calculated in real time, and uses the long-term optimized target state trajectory as a constraint or reference information to guide the calculation of short-term control actions, thereby generating hybrid control instructions.
[0204] Real-time calculation and analysis of the UAV state prediction error trend. The specific algorithm steps are as follows:
[0205] (1) Real-time calculation of the current prediction error , defined as the deviation between the actual observed state of the drone and the predicted state:
[0206] ;
[0207] in, The current actual measured drone status; The current state of the drone predicted for the previous control cycle; represents the Euclidean norm of the state vector; Indicates the current control cycle number;
[0208] (2) Real-time calculation of prediction error trends based on the sliding window method , defined as the most recent consecutive The linear regression slope of the forecast error series at each moment:
[0209] Assume that the forecast error sequence is , then the linear regression calculation of the error trend is:
[0210] ;
[0211] in, is the sliding window size, which indicates the number of time series error points used to calculate the trend; When , the error tends to increase and the system uncertainty increases; when When , it indicates that the error tends to decrease and the system uncertainty decreases;
[0212] (3) The forecast error trend indicator Incorporating environmental complexity indicators The extended definition of the computational complexity index is:
[0213] ;
[0214] in, is the weight coefficient of the forecast error trend; other symbols are defined as above;
[0215] (4) According to the updated complexity index The fusion weights of the reinforcement learning strategy module and the model predictive control module are adjusted in real time to improve the robustness of the system in complex and uncertain environments. The fusion weight adjustment still adopts:
[0216] .
[0217] S6. Send the hybrid control instructions to the UAV flight control actuator in real time to control the attitude and trajectory of the UAV in real time.
[0218] The present invention also provides a hybrid control strategy system for unmanned aerial vehicles in a highly dynamic and complex environment, the system comprising:
[0219] Status information acquisition module, used to collect and pre-process the UAV flight status and environment information in real time;
[0220] Model predictive control module, which is used to generate short-term optimal control action sequences in real time based on state information and dynamic models;
[0221] Reinforcement learning strategy module, used to generate long-term optimization target state trajectory in real time based on state information;
[0222] The strategy fusion module is used to use the output of the reinforcement learning strategy module as constraints or reference information for the model predictive control module to generate hybrid control instructions;
[0223] The flight control execution module is used to control the UAV's flight attitude and trajectory in real time based on hybrid control instructions.
[0224] The status information acquisition module includes at least one of an inertial measurement unit (IMU), a visual simultaneous localization and mapping (SLAM) module, a lidar, and an ultrasonic sensor to provide high-precision real-time status information.
[0225] The system also includes an edge computing platform that deploys a model predictive control module, a reinforcement learning strategy module, and a strategy fusion module in real time to achieve real-time calculation of the control strategy.
[0226] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A hybrid control strategy method for UAVs in highly dynamic and complex environments, characterized by: The method comprises: S1, collects the current flight status information and environmental information of the UAV in real time, and performs preprocessing to generate standardized input data; S2. Build a hybrid control strategy framework for UAVs, including a model predictive control module, a reinforcement learning strategy module, and a strategy fusion module; S3, the model predictive control module predicts short-term state changes within the prediction time window based on the current input data according to the UAV dynamics model, and uses sequential quadratic programming, interior point method or fast gradient descent algorithm to calculate and output the short-term optimal control action sequence in real time; S4, the reinforcement learning strategy module uses one of the pre-trained deep deterministic policy gradient algorithm, proximal policy optimization algorithm, or soft actor-critic algorithm to calculate the long-term optimal target state trajectory of the drone in real time based on the current input data; S5. The strategy fusion module calculates the environmental complexity index in real time, dynamically adjusts the output weights of the reinforcement learning strategy module and the model predictive control module based on the environmental complexity index, uses the long-term optimization target state trajectory as a constraint or reference information, guides the calculation of short-term control actions, and generates hybrid control instructions; S6. Send the hybrid control instructions to the UAV flight control actuator in real time to control the attitude and trajectory of the UAV in real time.
2. The hybrid control strategy method for UAV in a highly dynamic and complex environment according to claim 1 is characterized in that: The steps for calculating the short-term optimal control action using the model predictive control module include: Using the dynamic model, the current state The prediction is started to generate multiple candidate state trajectories, which are expressed as: ; in, It means that at time t, the state in the next N steps is predicted based on the current state, and the evolution of the predicted state satisfies the dynamic model constraints: ; in, is the dynamic state transfer function of the UAV, is the control input to be optimized; in the case of the state transfer function constructed using the dynamic model, Represents the state predicted at time t+1; Define the cost function based on the candidate state trajectory: ; in, Indicates at time The desired reference state; , is the weight matrix of state deviation and control energy consumption; is the weight matrix of the terminal state deviation; is the length of the prediction time window; represents the penalty term of the control input; It represents the reference state that the system should reach at the future time t+N, which is expected at the current time t. Represents the state at the kth moment predicted at time t; Use optimization algorithms to optimize the problem and obtain the optimal control action sequence within the prediction time window ; ; in, Represents the control input sequence and outputs the first control action of the optimal control action sequence in the short term in real time , and use it as the execution control action of the model predictive control module in the current control cycle.
3. The hybrid control strategy method for UAV in a highly dynamic and complex environment according to claim 1 is characterized in that: The steps for calculating the long-term optimal target trajectory through the reinforcement learning policy module include: Build and pre-train reinforcement learning policy network , the network is based on the current state of the drone and environmental information As input, the optimal target state trajectory is output, and the policy network is expressed as: ; Among them, the network parameters Obtained through pre-training of deep reinforcement learning algorithms; The reinforcement learning policy network is trained using deep deterministic policy gradient algorithms, proximal policy optimization, or soft actor-critic algorithms. In the online control phase, based on the real-time state input, the trained reinforcement learning strategy network directly outputs the long-term optimal target state trajectory of the UAV: ; in: Represents the current policy network The parameter vector of Indicates that the parameter is The policy network function.
4. The hybrid control strategy method for unmanned aerial vehicles in highly dynamic and complex environments according to claim 3 is characterized in that: The reinforcement learning policy network is trained using a deep deterministic policy gradient algorithm. The steps include: Step a: Initialize the reinforcement learning policy network Actor parameters , value network critic parameter , and set the target network parameters: ; in, Represents the parameters of the current Actor network, Represents the parameters of the target Actor network, is the parameter of the current Critic network, Represents the parameters of the target Critic network; Step b: The drone interacts in the environment and selects actions based on the policy network: ; in, To explore noise; Step c: Execute actions in the environment , observe the next state and rewards , and store the experience tuple in the experience replay buffer : ; Step d: Randomly extract small batches of data from the experience replay buffer for training, and update the critic network parameters by minimizing the critic network's loss function: ; in, Indicates the current state, represents the action, generated by the current policy network; r represents the current reward; Indicates the state at the next moment; For the experience replay pool; Represents the current Critic network's estimated Q value for the state-action pair; is the target Q value, which is defined as: ; in, is the discount factor; r is the current reward value; Represents the target Critic network's estimate of the action value of the next state; Indicates that the target Actor network is in state Next, select the action; Indicates the next state; Step e: Update the reinforcement learning policy network Actor network parameters , optimize the expected return of the policy network by gradient ascent: ; Step f: Periodically soft-update target network parameters: ; ; Where, is the soft update coefficient, .
5. The hybrid control strategy method for unmanned aerial vehicles in a highly dynamic and complex environment according to claim 1 is characterized in that: Step S5 includes: Real-time computing environment complexity index , expressed as: ; in, The current speed of the drone; is the obstacle density in the current environment; is the UAV state prediction error; , , is a predetermined weight coefficient used to adjust the contribution ratio of different indicators to complexity; According to the complexity index of real-time calculation , dynamically adjust the output weight of the reinforcement learning strategy module and model predictive control module output weights , the weights satisfy the constraints: ; The weight adjustment rules are: ; ; in, is the Sigmoid function; Preset thresholds for environmental complexity; Adjust the sensitivity coefficient for the weight; Fusion reinforcement learning module output long-term goal state trajectory The target state trajectory corresponding to the short-term optimal control action sequence output by the model predictive control module , obtain the reference state trajectory after fusion , specifically integrated as follows: ; The fused trajectory As the model predictive control module calculates the reference trajectory of the short-term control action sequence and outputs the final hybrid control command: ; in, is the current drone status; It represents the process of executing the model predictive control algorithm to solve the control action using the fused trajectory as the reference input.
6. The hybrid control strategy method for UAV in a highly dynamic and complex environment according to claim 5 is characterized in that: The steps to calculate the environmental complexity index, including calculating the trend of the drone state prediction error, include: Calculate the current prediction error in real time , defined as the deviation between the actual observed state of the drone and the predicted state: ; in, The current actual measured drone status; The current state of the drone predicted for the previous control cycle; represents the Euclidean norm of the state vector, Indicates the current control cycle number; Real-time calculation of prediction error trends based on sliding window method , defined as the most recent consecutive The linear regression slope of the forecast error series at each moment: Assume that the forecast error sequence is , then the linear regression calculation of the error trend is: ; in, is the sliding window size, which indicates the number of time series error points used to calculate the trend; When , the error tends to increase and the system uncertainty increases; when When , it indicates that the error tends to decrease and the system uncertainty decreases; Forecast Error Trend Indicator Incorporating environmental complexity indicators The extended definition of the computational complexity index is: ; in, is the weight coefficient of the forecast error trend; According to the updated complexity index Adjust the fusion weight of the reinforcement learning strategy module and the model predictive control module in real time. The fusion weight adjustment adopts: ; 。 7. The hybrid control strategy method for unmanned aerial vehicles in a highly dynamic and complex environment according to claim 1 is characterized in that: The UAV dynamics model is a data-driven neural network dynamics model or an explicit dynamics model based on physical parameters. When the data-driven neural network dynamics model is used: Using historical flight datasets Training a neural network model , so that the state prediction error is minimized, and the optimization objective function is defined as: ; in, For drones at all times Status; For the moment Control input; For the moment the actual status of is the neural network to be trained, is the network parameter; is the Euclidean norm of the state vector prediction error; Real-time prediction of the drone's status at the next moment, using a data-driven neural network model: ; In the case of a state transfer function constructed using a neural network prediction model, It represents the state at time t+1 predicted at time t.
8. The hybrid control strategy method for UAV in a highly dynamic and complex environment according to claim 7 is characterized in that: When using an explicit physical parameter dynamics model: According to the rigid body dynamics model of the UAV, the kinematic equation is defined as follows: Kinematic equations of position and velocity: ; ; in, is the first derivative of position, indicating that the change of position with respect to time is velocity; is the first-order derivative of velocity, indicating that the linear acceleration of the drone is composed of the linear acceleration converted from gravity acceleration and thrust; Attitude kinematic equation: ; ; in: Represents the first derivative of angular velocity, describing the change of angular velocity of a rigid body over time, including torque input and gyroscopic force; is the derivative of the attitude angle; 、 Respectively represent the position and speed status of the drone; is the gravitational acceleration vector; For drone quality; is the attitude rotation matrix, where , ; Represent the pitch angle, roll angle and yaw angle respectively; is the total thrust of the UAV; is the angular velocity of the drone, is the Euler angle posture; is the inertia matrix; is the total moment acting on the UAV; is the conversion matrix from angular velocity to attitude angular velocity; symbol represents vector cross product; Real-time prediction of the drone state at the next moment, based on the dynamic model of physical parameters, using numerical integration method: ; in, is the prediction time step; is the state derivative based on the explicit physical model; In the case of state derivatives based on rigid body dynamics models, explicit physical modeling + numerical integration is used for prediction. It represents the state at time t+1 predicted at time t.
9. A hybrid control strategy system for unmanned aerial vehicles in a highly dynamic and complex environment, the system being used to implement the hybrid control strategy method for unmanned aerial vehicles in a highly dynamic and complex environment as described in any one of claims 1 to 8, characterized in that: The system comprises: Status information acquisition module, used to collect and pre-process the UAV flight status and environment information in real time; Model predictive control module, which is used to generate short-term optimal control action sequences in real time based on state information and dynamic models; Reinforcement learning strategy module, used to generate long-term optimization target state trajectory in real time based on state information; The strategy fusion module is used to use the output of the reinforcement learning strategy module as constraints or reference information for the model predictive control module to generate hybrid control instructions; The flight control execution module is used to control the UAV's flight attitude and trajectory in real time based on hybrid control instructions.
10. The hybrid control strategy system for unmanned aerial vehicles in highly dynamic and complex environments according to claim 9 is characterized in that: The state information acquisition module includes at least one of an inertial measurement unit, a visual synchronous positioning and mapping module, a laser radar and an ultrasonic sensor.
Citation Information
Patent Citations
Construction method, control method and system of flight control network
CN119668277A
Ship motion trail optimization control method based on deep learning
CN119987374A