Unmanned aerial vehicle adaptive optimization method based on reinforcement learning training
By applying an adaptive optimization method based on enhanced learning in drone control, state space and action space are built, action learning model with a dual neural network architecture is built, and decision-making optimization is used to solve the problem of insufficient adaptive ability of drones to perform tasks in complex environments, and efficient and safe flight performance and task execution capabilities are achieved.
Patent Information
- Application Number
- CN202510330314.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing UAV control technology is difficult to achieve efficient and safe task execution in complex and changing environments, especially when facing dynamic obstacles and environmental changes, and lacks adaptability and decision-making accuracy.
Adaptive optimization method of drone based on enhanced learning training is adopted, by building state space and action space, an action learning model with a dual neural network architecture including policy networks and value networks is built, a multi-factor dynamic reward model is used for flight decision optimization, and model training is carried out through the empirical replay buffer.
It realizes the adaptive optimization of drones in complex scenarios, improves flight performance and mission execution capabilities, can sense environmental changes in real time, make optimal flight decisions, ensure flight safety and improve mission execution efficiency.
Smart Images

Figure CN120178680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) flight paths, and particularly to an adaptive optimization method for UAVs based on reinforcement learning training. Background Art
[0002] With the wide application of UAVs in many fields, such as logistics distribution, agricultural plant protection, surveying and exploration, etc., the requirements for the flight performance and adaptability of UAVs are increasing day by day. In complex and changeable task scenarios, UAVs need to have the ability of rapid response and intelligent decision-making to efficiently complete tasks and ensure their own safety. However, the existing UAV control technologies are difficult to meet the diverse needs in these complex scenarios. Therefore, it is of great practical significance to study an adaptive optimization method for UAVs based on reinforcement learning training.
[0003] Traditional UAV control technologies are mostly based on pre-set rules and algorithms. For example, in flight path planning, path planning methods based on geometric algorithms are often used. This method can plan feasible paths in simple scenarios, but has many limitations in complex environments. Its advantages are that the algorithm is relatively simple, the amount of calculation is small, and it is easy to implement. However, the disadvantages are also obvious. It lacks the adaptive ability to dynamic environmental changes. When encountering sudden obstacles or environmental interferences, it cannot adjust the path in time, resulting in low task execution efficiency and even putting the UAV in a dangerous situation. Moreover, this way of pre-setting rules is difficult to adapt to the diversity changes of different task requirements and complex scenarios.
[0004] The existing technologies have introduced intelligent algorithms to a certain extent to improve the performance of UAVs. For example, some studies use machine learning-based methods for UAV path planning and decision-making. These methods have made certain progress compared with traditional technologies and can learn environmental characteristics and task patterns from historical data to make relatively more reasonable decisions. However, the existing technologies still have deficiencies. On the one hand, the existing machine learning algorithms need to improve the learning efficiency and decision-making accuracy when dealing with complex environments with high dynamics and strong uncertainty. For example, when facing fast-moving obstacles and signal interferences, the UAV cannot timely and accurately perceive environmental changes and make correct responses. On the other hand, the decision-making models of UAVs in the existing technologies have poor generality, and the models need to be redesigned and trained for different tasks and scenarios, which greatly increases the development cost and application difficulty.
[0005] Therefore, neither traditional technology nor existing technology can well meet the requirements for drones to efficiently and safely perform tasks in complex and changing environments. To overcome the deficiencies of these technologies and achieve adaptive optimization flight of drones in complex scenarios, this application proposes a drone adaptive optimization method based on reinforcement learning training, enabling drones to perceive environmental changes in real time, autonomously learn, and make optimal flight decisions, thereby improving the flight performance and task execution ability of drones. Summary of the Invention
[0006] Based on the above, the drone adaptive optimization method based on reinforcement learning training in this application includes the following steps:
[0007] S1. Construct a drone state space, discretize the flight decision actions of the drone in the state space, and construct an action space;
[0008] S2. Build an action learning model through the proximal policy optimization algorithm and initialize the parameters of the action learning model;
[0009] S3. Obtain the dynamic data of the drone performing tasks in the state space, where the dynamic data includes the position, speed, and size information of dynamic obstacle data and the state information of the position and speed of the drone itself;
[0010] S4. Through the obtained state information, use the learning model to select a flight decision action from the action space. After executing the selected action, obtain a new state and the corresponding reward value;
[0011] S5. Store the current state, the executed action, the obtained reward value, and the new state in the experience replay buffer;
[0012] S6. When the number of data in the experience replay buffer reaches the set threshold, randomly sample a batch of data from the buffer, train the flight learning model, optimize the model parameters, and make flight decisions until the flight task is completed.
[0013] Preferably, in S1, constructing the drone state space specifically includes: obtaining the pose data of the attitude angle and three-dimensional position coordinates of the drone, and forming a state space by fusing the pose data in the time dimension and space dimension of the drone; in the time dimension, sampling pose information at fixed intervals and adding timestamps to form a pose sequence; in the space dimension, dividing the task area into grids, mapping pose vectors and combining grid coordinates; forming a state space by integrating the pose data encoded in the time dimension and space dimension.
[0014] Preferably, the attitude and position data of the drone are obtained through an inertial measurement unit and carrier phase differential technology; the inertial measurement unit adopts an inertial measurement unit combined with a three-axis accelerometer and gyroscope of a microelectromechanical system to obtain the attitude angle data of the drone; the carrier phase differential technology determines the three-dimensional position coordinates of the drone by obtaining the carrier phase observation values and position information of the drone in real time; the attitude angle data and three-dimensional position coordinate data of the drone are obtained to form the attitude and position data of the drone.
[0015] Preferably, in S1, the flight decision-making actions of the drone in the state space are discretized to obtain continuous variables of the drone flight, including flight speed, direction, and altitude. According to the flight scenario and mission requirements, the continuous value range of the variables is divided into multiple fuzzy intervals, and each interval corresponds to a basic discrete action. The basic discrete actions are initially divided, and the rationality of the action combination and the effectiveness of covering the state space are used as the fitness function. The formula is: where p j represents the probability of the jth action combination appearing in the state space, n is the total number of action combinations, is the information entropy of the action combination, R is the rationality score of the action combination, C is the coverage degree score of the action combination for the state space, ω1, ω2, and ω3 are the corresponding weight coefficients respectively. By iteratively adjusting the interval boundaries, discrete intervals are obtained, and the discrete actions of different variables are logically associated to form an action space.
[0016] Preferably, in S2, an action learning model is built through the proximal policy optimization algorithm, which specifically includes: constructing a dual neural network architecture of a policy network and a value network; the policy network outputs the action probability distribution in the action space according to the state space where the drone is located; the value network takes the state information as input and evaluates the value of the current state action; an action learning model is constructed by alternately updating the parameters of the policy network and the value network.
[0017] Preferably, the policy network adopts a multi-layer perceptron architecture. After the input layer receives the state vector S of the state space, it is passed to the hidden layer for feature extraction and transformation. After the feature transformation of multiple hidden layers, the output layer outputs the action feature vector f. Through the formula: the action probability distribution P(a i |s) is obtained, where H(P(·|s)) is the action probability distribution entropy, a i is the action output, f i is the ith action feature vector, f l is the lth action feature vector, m is the total number of action feature vectors, and λ is the entropy regularization coefficient; the policy network is used to convert the state vector into an action probability distribution.
[0018] Preferably, the value network adopts a double-layer architecture of a convolutional neural network and a long short-term memory network. The convolutional neural network extracts the features of the state information, and the long short-term memory network captures the temporal dependence relationship of the state information features. The feature vectors processed by the double-layer architecture of the convolutional neural network and the long short-term memory network are concatenated into a value feature vector f k , through the formula: Output the state-action value V(s), where w T is the weight vector, f z-k is the k-th state information of the value feature vector, α k is the attention weight, r is the total number of state information feature vectors, and b v is the bias term.
[0019] Preferably, the acquisition of the dynamic obstacle data in the state space in S3 specifically includes: using a drone to emit a light beam, generating a three-dimensional point cloud by calculating the time difference of the reflected light, and identifying the spatial position and size of the obstacle through the three-dimensional point cloud; acquiring the dynamic obstacle image data, identifying the obstacle type, constructing a dynamic obstacle information database, analyzing the motion law through a spatio-temporal sequence algorithm, and predicting the future position and speed change of the obstacle in combination with the real-time image data to obtain the position, speed and size information of the dynamic obstacle.
[0020] Preferably, after performing the selected action in S4, a multi-factor dynamic reward model is constructed to obtain a reward value; the multi-factor dynamic reward model is scored by task completion, flight safety and action efficiency, and outputs the reward value R z , R z = w1R t + w2R s + w3R e , where R t is the task completion reward value, R s is the flight safety reward value, R e is the action efficiency reward value, and w1, w2, and w3 are the corresponding weight coefficients.
[0021] Preferably, after the data in the experience replay buffer reaches the set threshold number N, a batch of data B is randomly sampled from the buffer to train the flight learning model. By optimizing the model parameters, a flight decision is made, which specifically includes:
[0022] Randomly sample a batch of data of size B from the buffer where s i is the state information, a i is the action output, R i is the reward value, s i+1 is the next state, and it is input into the flight learning model for training; the flight learning model includes a policy network Pθ (a|s) and value network V φ (s), where θ and φ are the model parameters of the policy network and the value network respectively. In each training iteration, the policy network parameter θ and the value network parameter φ are updated respectively through the stochastic gradient descent algorithm until the model training reaches the termination condition, and the flight decision is output; the termination conditions include the cumulative reward threshold condition and the training iteration number condition.
[0023] Compared with the prior art, the technical solution of the present application has the following technical effects:
[0024] By constructing a state space that includes the attitude angle and three-dimensional position coordinates of the unmanned aerial vehicle (UAV) and integrates the pose data in the time dimension and the space dimension, and constructing an action space by discretizing the flight decision actions, the present invention solves the technical problems in the prior art that it is difficult for the UAV to accurately perceive complex environmental information and the decision actions are single, and obtains the technical effect that the UAV can comprehensively obtain its own state and environmental information, has rich and reasonable decision action options, and thus can flexibly plan the flight path in complex scenarios.
[0025] By using the proximal policy optimization algorithm to build a dual neural network architecture action learning model that includes a policy network and a value network and initializing the model parameters, the present invention solves the technical problems that the existing machine learning algorithms have insufficient learning efficiency and decision accuracy when dealing with complex environments with high dynamics and strong uncertainty, and obtains the technical effect that the UAV can quickly learn according to the environmental state and output a reasonable action probability distribution, accurately evaluate the value of the current state action, and thus improve the decision accuracy and learning efficiency.
[0026] By adopting the technical solution of obtaining the dynamic data of the UAV performing tasks in the state space, including dynamic obstacle data and its own state information, and obtaining the reward value according to the multi-factor dynamic reward model, the present invention solves the technical problems that the UAV cannot effectively cope with dynamic obstacles during flight and it is difficult to balance task completion, flight safety and efficiency, and realizes the technical effect that the UAV can real-time perceive the dynamic environmental changes, comprehensively optimize the flight decision according to the task completion degree, flight safety and action efficiency, ensure flight safety and improve the task execution efficiency.
[0027] By means of storing the data in the experience replay buffer, randomly sampling and training the flight learning model after the data reaches the set threshold number, and updating the model parameters through the stochastic gradient descent algorithm until the termination condition is reached, the present invention solves the technical problems that the UAV decision model in the prior art has poor generality and high training cost, and achieves the technical effect of improving the generality and generalization ability of the flight learning model, reducing the cost of retraining the model for different tasks and scenarios, and enabling the UAV to efficiently execute tasks in multiple scenarios.
[0028] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, so as to be implemented in accordance with the content of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following will be described in detail with reference to the preferred embodiments of the present application and the accompanying drawings.
[0029] Those skilled in the art will understand the above and other purposes, advantages and features of the present application more clearly according to the following detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings. Description of the Drawings
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to actual scale.
[0031] Figure 1 Flowchart of the UAV adaptive optimization method based on reinforcement learning training of the present invention;
[0032] Figure 2 Flowchart of obtaining dynamic obstacle data of the UAV adaptive optimization method based on reinforcement learning training of the present invention. Detailed Embodiments
[0033] To make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. In the following description, specific details such as specific configurations and components are provided only to help a comprehensive understanding of the embodiments of the present application. Therefore, those skilled in the art should clearly understand that various changes and modifications can be made to the embodiments described here without departing from the scope and spirit of the present application. In addition, descriptions of known functions and structures are omitted for clarity and conciseness.
[0034] It should be understood that the "one embodiment" or "the present embodiment" mentioned throughout the specification means that a specific feature, structure or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the "one embodiment" or "the present embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in any suitable manner in one or more embodiments.
[0035] In addition, this application may repeat reference numerals and / or letters in different instances. This repetition is for the purpose of simplicity and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0036] The term “and / or” in this document is merely a description of the associated relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, B exists alone, and both A and B exist simultaneously. The term “ / and” in this document describes another associated object relationship, indicating that there can be two relationships. For example, A / and B can represent: A exists alone, and both A and B exist. Additionally, the character “ / ” in this document generally indicates that the associated objects before and after are in an “or” relationship.
[0037] The term “at least one” in this document is merely a description of the associated relationship of the associated objects, indicating that there can be three relationships. For example, at least one of A and B can represent: A exists alone, both A and B exist simultaneously, and B exists alone.
[0038] It should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms “comprising,” “including,” or any other variant thereof are intended to cover non-exclusive inclusion.
[0039] Embodiment 1
[0040] This embodiment mainly describes an adaptive optimization method for drones based on reinforcement learning training, as Figure 1 shown, including:
[0041] S1. Construct a drone state space, discretize the flight decision actions of the drone in the state space, and construct an action space;
[0042] S2. Build an action learning model through the proximal policy optimization algorithm and initialize the parameters of the action learning model;
[0043] S3. Obtain the dynamic data of the drone performing tasks in the state space, where the dynamic data includes the position, speed, and size information of dynamic obstacle data and the state information of the drone's own position and speed;
[0044] S4. Through the obtained state information, use the learning model to select a flight decision action from the action space. After executing the selected action, obtain the new state and the corresponding reward value;
[0045] S5. Store the current state, the executed action, the obtained reward value, and the new state in the experience replay buffer;
[0046] S6. When the data in the experience replay buffer reaches the set threshold number, randomly sample a batch of data from the buffer, train the flight learning model, and make flight decisions by optimizing the model parameters until the flight task is completed.
[0047] Furthermore, in S1, the construction of the UAV state space specifically includes: obtaining the pose data of the UAV's attitude angle and three-dimensional position coordinates, and forming the state space by fusing the pose data in the time dimension and space dimension of the UAV; in the time dimension, sample the pose information at fixed intervals and add timestamps to form a pose sequence; in the space dimension, divide the task area into grids, map the pose vectors and combine them with the grid coordinates; form the state space by integrating the pose data encoded in the time dimension and space dimension.
[0048] Furthermore, the acquisition of the UAV pose data is obtained through the inertial measurement unit and the carrier phase differential technology; the inertial measurement unit uses a microelectromechanical system's three-axis accelerometer and gyroscope combined inertial measurement unit to obtain the UAV's attitude angle data; the carrier phase differential technology determines the UAV's three-dimensional position coordinates by obtaining the carrier phase observation values and position information of the UAV in real time; the UAV pose data is formed by obtaining the UAV's attitude angle data and three-dimensional position coordinate data.
[0049] Furthermore, in S1, the flight decision-making actions of the UAV in the state space are discretized to obtain the continuous variables of the UAV flight, including flight speed, direction, and altitude. According to the flight scenario and task requirements, divide the continuous value range of the variables into multiple fuzzy intervals, and each interval corresponds to a basic discrete action. Initially divide the basic discrete actions, and use the rationality of the action combination and the effectiveness of covering the state space as the fitness function. The formula is: where p j represents the probability of the jth action combination appearing in the state space, n is the total number of action combinations, is the information entropy of the action combination, R is the rationality score of the action combination, C is the coverage degree score of the action combination for the state space, ω1, ω2, ω3 are the corresponding weight coefficients respectively. By iteratively adjusting the interval boundaries, obtain the discrete intervals, and logically associate the discrete actions of different variables to form the action space.
[0050] Furthermore, S2 constructs an action learning model through the proximal policy optimization algorithm, which specifically includes: constructing a dual neural network architecture of a policy network and a value network; the policy network outputs the action probability distribution in the action space according to the state space where the UAV is located; the value network takes the state information as input and evaluates the value of the current state action; by alternately updating the parameters of the policy network and the value network, an action learning model is constructed.
[0051] Furthermore, the policy network adopts a multi-layer perceptron architecture. After the input layer receives the state vector S of the state space, it is passed to the hidden layer for feature extraction and transformation. After the feature transformation of multiple hidden layers, the output layer outputs the action feature vector f. Through the formula: the action probability distribution P(a i |s) is obtained, where H(P(·|s)) is the action probability distribution entropy, a i is the action output, f i is the i-th action feature vector, f l is the l-th action feature vector, m is the total number of action feature vectors, and λ is the entropy regularization coefficient; the policy network is used to convert the state vector into the action probability distribution.
[0052] Furthermore, the value network adopts a double-layer architecture of a convolutional neural network and a long short-term memory network. The convolutional neural network extracts the features of the state information, and the long short-term memory network captures the time-dependent relationship of the state information features. The feature vectors processed by the double-layer architecture of the convolutional neural network and the long short-term memory network are concatenated into the value feature vector f k , and through the formula: the state-action value V(s) is output, where w T is the weight vector, f z-k is the k-th state information of the value feature vector, α k is the attention weight, r is the total number of state information feature vectors, and b v is the bias term.
[0053] Furthermore, as Figure 2 shown, the acquisition of dynamic obstacle data in the state space in S3 specifically includes: using the UAV to emit a light beam, generating a three-dimensional point cloud by calculating the time difference of the reflected light, and identifying the spatial position and size of the obstacle through the three-dimensional point cloud; acquiring dynamic obstacle image data, identifying the obstacle type, constructing a dynamic obstacle information database, analyzing the motion law through a spatio-temporal sequence algorithm, and predicting the future position and speed change of the obstacle in combination with the real-time image data to obtain the position, speed and size information of the dynamic obstacle.
[0054] Further, after performing the selected action in S4, a multi-factor dynamic reward model is constructed to obtain a reward value; the multi-factor dynamic reward model is scored based on task completion, flight safety, and action efficiency, and outputs a reward value R z , R z = w1R t + w2R s + w3R e , where R t is the task completion reward value, R s is the flight safety reward value, R e is the action efficiency reward value, and w1, w2, and w3 are the corresponding weight coefficients respectively.
[0055] Further, after the data in the experience replay buffer in S6 reaches the set threshold number N, a batch of data B is randomly sampled from the buffer to train the flight learning model. By optimizing the model parameters, flight decisions are made, specifically including:
[0056] Randomly sample a batch of data of size B from the buffer where s i is the state information, a i is the action output, R i is the reward value, s i+1 is the next state, and is input into the flight learning model for training; the flight learning model includes a policy network P θ (a|s) and a value network V φ (s), where θ and φ are the model parameters of the policy network and the value network respectively. In each training iteration, the policy network parameter θ and the value network parameter φ are updated respectively through the stochastic gradient descent algorithm until the model training reaches the termination condition, and the flight decision is output; the termination condition includes the cumulative reward threshold condition and the training iteration number condition.
[0057] This embodiment details the UAV adaptive optimization method based on reinforcement learning training. By constructing a more complete state space and action space, combining the proximal policy optimization algorithm and the multi-factor dynamic reward model, the UAV can perceive environmental changes in real time, autonomously learn and make optimal flight decisions, thereby improving the flight performance and task execution ability of the UAV.
[0058] Based on Embodiment 1, this embodiment describes the UAV state space in S1, and task target information can be added. At the same time, for the task target information, by combining the target position and task priority of the task target into the state space, specifically:
[0059] In the construction of the UAV state space, by considering the target position and task priority of the mission objective, the flight decision-making of the UAV is enhanced; the target position and task priority of the mission objective are innovatively integrated into the state space to endow the UAV with mission perception ability.
[0060] The target position includes the flight path planning and flight time estimation information required for the UAV to fly to this position; the task priority reflects the importance and urgency of the task at the target position in the entire mission system; high-priority tasks need to be completed by the UAV first, even if it requires more flight costs; low-priority tasks can be processed when resources permit. By combining these two into the state space, the UAV can comprehensively consider the urgency and difficulty of task implementation during the decision-making process, thus formulating a more reasonable and efficient flight strategy.
[0061] For example, in the scenario of multi-task execution, the UAV no longer simply sorts tasks according to the distance of the target position, but will consider the level of task priority. If a task with a far distance but extremely high priority appears, the UAV will abandon the currently executing low-priority task and fly to the target position of the high-priority task instead. This decision-making method based on comprehensive task information greatly improves the adaptability and execution efficiency of the UAV in complex task environments.
[0062] Effectively integrating the target position and task priority of the mission objective into the state space specifically includes: for the target position information, performing multi-dimensional encoding on it. In addition to the basic longitude and latitude coordinates, relative position information is introduced, such as the distance and direction angle relative to the current position of the UAV; at the same time, combining Geographic Information System (GIS) data to extract the characteristics of the environment where the target position is located, so that the multi-dimensional information can more comprehensively describe the characteristics of the target position and provide richer basis for the decision-making of the UAV.
[0063] For task priority, a numerical representation method is adopted, and different numerical values are assigned to different task priorities. For example, the priority is set from 1 to 10, and the larger the number, the higher the priority. In the state space, the task priority is stored as an independent dimension. To achieve the fusion of the target position and task priority, a feature fusion model based on neural network is constructed; the model takes the multi-dimensional encoded information of the target position and the task priority value as inputs, and through the non-linear transformation of the multi-layer neural network, deeply fuses the features of the two and outputs a comprehensive task feature vector.
[0064] When training the feature fusion model, a large amount of historical task data is used for supervised learning. By setting a reasonable loss function, the model can learn the internal relationship between the target location and the task priority. For example, the loss function takes into account the execution efficiency and success rate of the drone under different task combinations. By optimizing the model parameters, the fused task feature vector can more accurately reflect the actual requirements of the task. In practical applications, when the drone receives a new task each time, it inputs the target location and task priority information into the trained feature fusion model to obtain a comprehensive task feature vector, and adds it to the state space. The state space can then contain information about both the target location and the task priority, providing more comprehensive support for subsequent flight decisions.
[0065] When the state space fuses the target location and task priority information of the task objective, it will play an important role in the drone's flight decision-making; in terms of path planning, the fused state space enables the path planning algorithm to comprehensively consider the importance and urgency of the task; for example, when the drone faces multiple task objectives, the path planning algorithm can sort the objectives according to the task priority and prioritize planning the path to the high-priority task objective. At the same time, during the path selection process, the environmental characteristics of the target location will be fully considered to select a safer and more efficient flight path.
[0066] In terms of task scheduling, the fused state space can provide a more intelligent task scheduling strategy for the drone. The drone can dynamically adjust the task execution order according to the task information in the current state space. If a high-priority task appears near the target location of a low-priority task, the drone can interrupt the current low-priority task in a timely manner and switch to executing the high-priority task. This dynamic task scheduling ability based on the state space enables the drone to better adapt to complex and changing task environments.
[0067] In addition, the state space after fusing the task objective information can also support the drone's fault handling and emergency response. When the drone encounters a fault or an emergency, it can prioritize the completion of high-priority tasks according to the task information in the state space. For example, when the drone's energy is about to run out, it can choose to complete a high-priority task nearby according to the task priority and target location information, or return to the base for charging as soon as possible, while marking the uncompleted low-priority tasks for subsequent processing when conditions permit. This fault handling and emergency response mechanism based on comprehensive task information greatly improves the reliability of the drone and the success rate of task execution. By continuously expanding and optimizing the application of the state space, the drone can play a greater role in various complex task scenarios.
[0068] This embodiment details the integration of task target locations and priorities into the UAV state space, which can greatly enhance its decision-making and execution capabilities. In complex task environments, the UAV can reasonably plan flight paths and dynamically schedule tasks according to the importance and urgency of tasks. In the face of failures or emergencies, it can also prioritize high-priority tasks, improving reliability and success rates, and comprehensively enhancing the UAV's adaptability and task completion efficiency in various scenarios.
[0069] Based on Embodiment 1, this embodiment describes the experience replay buffer of S5, which specifically includes:
[0070] The experience replay buffer lies in the storage and utilization of data. During the flight of the UAV, the state information of each flight of the UAV, such as its position, speed, attitude, etc., the actions performed, and the corresponding reward values, such as obtaining a positive reward for successfully avoiding obstacles, an increasing reward for approaching the target location, and the new state entered, are recorded completely and orderly. As the number of flights increases, these data accumulate continuously, building a large dataset of flight experiences. This buffer breaks the sequential dependence of data in traditional reinforcement learning. Continuously collected data in the past often has strong correlations, making learning algorithms easily misled by local data features and falling into local optimal solutions. The experience replay buffer randomly samples experience segments from a large amount of historical data at different stages and in different flight scenarios. For example, during multiple flights, the coping strategies adopted by the UAV in different meteorological conditions such as sunny and rainy days, or in different geographical environments such as open fields and narrow lanes are stored in the buffer. During training, these diverse data are randomly sampled for learning, enabling the algorithm to be exposed to a wider range of flight scenarios, effectively avoiding ignoring other potential effective strategies due to the limitations of continuous data, thus enhancing the stability and comprehensiveness of the learning process and enabling the algorithm to better adapt to complex and changing flight environments.
[0071] The experience replay buffer can intelligently classify and manage the stored data according to task characteristics. Considering the diversity of tasks performed by the UAV, the buffer adds specific tags to each stored experience data according to the priority of the task and the characteristics of the target location. In complex scenarios with multiple tasks running in parallel, the flight experiences generated by high-priority tasks are prominently marked in the buffer. During subsequent training, the probability of sampling these key experience data is greatly increased, prompting the algorithm to preferentially learn efficient flight strategies for important tasks, achieving precise investment and efficient utilization of resources, and ensuring that the UAV can prioritize the completion of core tasks when performing tasks.
[0072] A dynamic update mechanism is introduced into the buffer. According to the dynamic change characteristics of the UAV flight environment and tasks, the buffer can flexibly adjust the data storage and sampling strategies based on real-time feedback. For example, when the UAV enters a new and complex flight area, the buffer will quickly respond, temporarily increase the storage volume of relevant data in this area, and increase the weight of such data during the sampling process, guiding the algorithm to quickly focus on optimizing the flight strategy in the new environment. At the same time, by using the automatic screening function, it can identify and eliminate stale data that has little effect on the current task optimization according to the training effect of the algorithm, always maintaining the timeliness and effectiveness of the buffer data, providing training data that highly fits the actual flight requirements for the flight learning model, and greatly enhancing the adaptability and optimization ability of the technical solution of this application in dynamic and complex flight scenarios.
[0073] This embodiment details that the experience replay buffer breaks the traditional sequential dependence in storing flight data, and random sampling improves the stability and comprehensiveness of learning. It can also intelligently classify and manage data according to task characteristics, highlight high-priority task data, and increase the sampling probability. Its dynamic update mechanism can flexibly adjust the storage and sampling strategies, eliminate stale data, provide efficient and practical training data for the flight learning model, and help the UAV adapt to complex flight scenarios.
[0074] Based on Embodiment 1, this embodiment describes the multi-factor dynamic reward model of S4, specifically including:
[0075] After performing the selected action in S4, by constructing a multi-factor dynamic reward model, a reward value is obtained; the multi-factor dynamic reward model is scored through task completion, flight safety, and action efficiency, and outputs the reward value R z , R z = w1R t + w2R s + w3R e , where R t is the task completion reward value, which is determined by calculating the Euclidean distance d between the current position of the UAV and the target position. The formula is: The closer the distance, the higher the reward; R s is the flight safety reward value. If the minimum distance D from an obstacle is less than the safety threshold D th , the reward value is R s , and the formula is: where k1 is the safety coefficient, otherwise R s = 1, ensuring that the UAV stays away from danger; R e is the action efficiency reward value, and the formula is: Compare the deviation Δe between the expected effect and the actual effect of the action to ensure the accuracy and efficiency of the action. w1, w2, and w3 are the corresponding weight coefficients.
[0076] After the actions are executed in detail in this embodiment, new states are obtained through the fusion of multiple types of sensors, comprehensively and accurately. The reward value is calculated by integrating multiple factors, such as task completion, flight safety, energy consumption, and action efficiency. The weights are dynamically adjusted according to the task stage and environmental adaptation evaluation to ensure that the rewards reasonably reflect the quality of the actions. This enables the drone to weigh the pros and cons in flight decisions, optimize the flight strategy, and significantly improve the efficiency, safety, and resource utilization rate of flight task execution.
[0077] The above are only the preferred embodiments of the present invention, and it does not limit the protection scope of the present invention accordingly. For those skilled in the art, the present invention can have various changes and modifications; within the spirit and principle of the present invention, any changes, modifications, substitutions, integrations, and parameter changes to these embodiments through conventional substitutions or the ability to achieve the same functions without departing from the principle and spirit of the present invention fall within the protection scope of the present invention.
Claims
1. The UAV adaptive optimization method based on reinforcement learning training is characterized by: include: S1. Construct the state space of the UAV, discretize the flight decision actions of the UAV in the state space, and construct the action space; S2. Build an action learning model through the proximal strategy optimization algorithm and initialize the action learning model parameters; S3, obtaining dynamic data of the UAV performing the task in the state space, wherein the dynamic data includes the position, speed and size information of the dynamic obstacle data and the state information of the position and speed of the UAV itself; S4, using the acquired state information, selecting a flight decision action from the action space using the learning model, and after executing the selected action, obtaining a new state and a corresponding reward value; S5, storing the current state, the executed action, the obtained reward value and the new state into the experience replay buffer; S6. When the data in the experience playback buffer reaches the set threshold, a batch of data is randomly sampled from the buffer to train the flight learning model. By optimizing the model parameters, flight decisions are made until the flight mission is completed.
2. The UAV adaptive optimization method based on reinforcement learning training according to claim 1 is characterized in that: The UAV state space is constructed in S1, specifically including: obtaining the posture data of the attitude angle and three-dimensional position coordinates of the UAV, and forming the state space by fusing the posture data of the time dimension and the space dimension of the UAV; in the time dimension, sampling the posture information at fixed intervals, adding timestamps to form a posture sequence; in the space dimension, dividing the task area into grids, mapping the posture vectors and combining the grid coordinates; and forming the state space by integrating the posture data encoded in the time dimension and the space dimension.
3. The UAV adaptive optimization method based on reinforcement learning training according to claim 2 is characterized in that: The acquisition of the drone's posture data is carried out through an inertial measurement unit and carrier phase differential technology; the inertial measurement unit uses an inertial measurement unit composed of a three-axis accelerometer and a gyroscope of a micro-electromechanical system to obtain the drone's attitude angle data; The carrier phase difference technology determines the three-dimensional position coordinates of the UAV by acquiring the carrier phase observation value and position information of the UAV in real time; and forms the UAV posture data by acquiring the attitude angle data and three-dimensional position coordinate data of the UAV.
4. The UAV adaptive optimization method based on reinforcement learning training according to claim 1 is characterized in that: In S1, the flight decision action of the UAV in the state space is discretized to obtain the continuous variables of the UAV flight, including flight speed, direction, and altitude. According to the flight scene and mission requirements, the continuous value range of the variables is divided into multiple fuzzy intervals. Each interval corresponds to a basic discrete action. The basic discrete actions are preliminarily divided, and the rationality of the action combination and the effectiveness of covering the state space are used as the fitness function. The formula is: where p j represents the probability of the jth action combination appearing in the state space, n is the total number of action combinations, is the information entropy of the action combination, R is the rationality score of the action combination, C is the coverage score of the action combination on the state space, ω1, ω2, and ω3 are the corresponding weight coefficients, and the discrete interval is obtained by iteratively adjusting the interval boundary. The discrete actions of different variables are logically associated to form an action space.
5. The method for adaptive optimization of unmanned aerial vehicle based on reinforcement learning training according to claim 1, characterized in that: The S2 builds an action learning model through a proximal strategy optimization algorithm, specifically including: constructing a dual neural network architecture of a strategy network and a value network; the strategy network outputs the action probability distribution in the action space according to the state space of the drone; the value network uses the state information as input to evaluate the value of the current state action; and the action learning model is constructed by alternately updating the parameters of the strategy network and the value network.
6. The method for adaptive optimization of unmanned aerial vehicles based on reinforcement learning training according to claim 5 is characterized in that: The strategy network adopts a multi-layer perceptron architecture. After the input layer receives the state vector S of the state space, it is passed to the hidden layer for feature extraction and transformation. After the feature transformation of multiple hidden layers, the output layer outputs the action feature vector f, which is expressed by the formula: Get the action probability distribution P(a i |s), where H(P(·|s)) is the action probability distribution entropy, a i is the action output, f i is the i-th action feature vector, f l is the lth action feature vector, m is the total number of action feature vectors, and λ is the entropy regularization coefficient. The policy network is used to convert the state vector into action probability distribution.
7. The method for adaptive optimization of unmanned aerial vehicles based on reinforcement learning training according to claim 5 is characterized in that: The value network adopts a double-layer architecture of a convolutional neural network and a long short-term memory network. The convolutional neural network extracts the features of the state information, and the long short-term memory network captures the time dependency of the state information features. The feature vectors processed by the double-layer architecture of the convolutional neural network and the long short-term memory network are concatenated into a value feature vector f k , through the formula: Output state action value V(s), where w T is the weight vector, f z-k is the kth state information of the value feature vector, α k is the attention weight, r is the total number of state information feature vectors, b v is the bias term.
8. The method for adaptive optimization of unmanned aerial vehicles based on reinforcement learning training according to claim 1, characterized in that: The acquisition of dynamic obstacle data in the state space in S3 specifically includes: using a drone to emit a light beam, generating a three-dimensional point cloud by calculating the time difference of the reflected light, and identifying the spatial position and size of the obstacle through the three-dimensional point cloud; acquiring dynamic obstacle image data, identifying the obstacle type, building a dynamic obstacle information database, analyzing the movement law through a spatiotemporal sequence algorithm, predicting the future position and speed change of the obstacle in combination with real-time image data, and acquiring the position, speed and size information of the dynamic obstacle.
9. The method for adaptive optimization of unmanned aerial vehicles based on reinforcement learning training according to claim 1, characterized in that: After the selected action is executed in S4, a reward value is obtained by constructing a multi-factor dynamic reward model; the multi-factor dynamic reward model scores the task completion, flight safety and action efficiency, and outputs a reward value R z , R z =w1R t +w2R s +w3R e , where R t is the task completion reward value, R s is the flight safety bonus value, R e is the action efficiency reward value, w1, w2, w3 are the corresponding weight coefficients respectively.
10. The UAV adaptive optimization method based on reinforcement learning training according to claim 1 is characterized in that: When the data in the experience playback buffer in S6 reaches the set threshold number N, a batch of data B is randomly sampled from the buffer to train the flight learning model, and flight decisions are made by optimizing model parameters, specifically including: Randomly sample a batch of data of size B from the buffer Among them, s i is the status information, a i is the action output, R i is the reward value, s i+1 For the next state, input the flight learning model for training; the flight learning model includes the policy network P θ (a|s) and value network V φ (s), where θ and φ are the model parameters of the policy network and the value network respectively. In each training iteration, the policy network parameters θ and the value network parameters φ are updated respectively by the stochastic gradient descent algorithm until the model training reaches the termination condition and the flight decision is output; the termination condition includes the cumulative reward threshold condition and the number of training iterations condition.
Citation Information
Cited By
Flight training adaptive optimization method, system and device based on big data operation and storage medium
CN120831912A
Flight training adaptive optimization method, system, device and storage medium based on operation big data
CN120831912B
Distribution network field survey path planning method based on near-end strategy optimization PPO
CN121455152A
Proximity policy optimization (ppo) based distribution network field survey path planning method
CN121455152B
Unmanned aerial vehicle visual obstacle avoidance and autonomous navigation method based on improved PPO
CN121477964A