Aircraft task intelligent decision-making method and system based on reinforcement learning
Through intelligent decision-making methods based on reinforcement learning, combined with autoencoder and Q learning framework, the aircraft decisions are dynamically adjusted, and the problem of limited autonomy and efficiency in complex environments is solved, achieving more efficient and safe task execution.
Patent Information
- Application Number
- CN202510372854.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-04
AI Technical Summary
Existing aircraft control systems lack independent optimization decision-making capabilities in complex and dynamic environments, making it difficult to deal with real-time environmental changes and emergencies, resulting in limited mission execution efficiency and safety.
Using an intelligent decision-making method based on reinforcement learning, combined with an autoencoder and a Q learning framework, the aircraft decision-making is optimized through dynamic reward functions and ε-greedy strategies, and the flight strategy is adjusted in real time to improve autonomy and adaptability.
It improves the mission execution efficiency and safety of the aircraft in complex environments, enhances the ability to adapt to dynamic changes, reduces collision risks, and optimizes energy consumption and flight time.
Smart Images

Figure CN120255574A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent decision-making method and system for aircraft missions based on reinforcement learning, belonging to the field of UAV control systems. Background Art
[0002] With the rapid development of UAV technology, UAVs are increasingly widely used in military, logistics, monitoring and other fields. Although existing aircraft control systems and air combat decision-making methods have achieved certain results in some specific scenarios, they usually rely on preset flight routes or static rules and lack the ability to autonomously optimize decisions in complex and dynamic environments. This leads to great limitations of these systems in dealing with real-time environmental changes (such as aircraft states, obstacles, meteorological changes, etc.), especially in dynamic and complex scenarios such as UAV swarm missions or short-range air combat, where it is difficult to cope with uncertainties and emergencies.
[0003] Therefore, there is an urgent need for an intelligent decision-making method and system that can learn in real time and adaptively adjust flight strategies to improve the task execution efficiency, safety and autonomy of aircraft in complex environments. Summary of the Invention
[0004] Object of the Invention: To solve the limitations of the existing technology, the present invention aims to provide an intelligent decision-making method, device, equipment and system for aircraft missions based on reinforcement learning, which uses the basic state data of the aircraft as input, combines real-time environmental changes, and through autonomous learning and optimization, helps the aircraft automatically adjust flight strategies in complex dynamic environments, and improves the autonomy, adaptability and efficiency of task execution.
[0005] Technical Solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0006] An intelligent decision-making method for aircraft missions based on reinforcement learning, comprising the following steps:
[0007] (1) Obtain the state data of the aircraft;
[0008] (2) Establish a dynamic reward function;
[0009] (3) Optimize the aircraft decision-making based on Q-learning;
[0010] (4) Adopt an ε-greedy strategy for aircraft action selection;
[0011] Among them, the expression of the dynamic reward function is:
[0012] R(Z t ,A t ) = α1R eff (Z t ,A t) + βR safety (Z t , A t ) + γR task (Z t , A t ) + δR env (Z t , A t )
[0013]
[0014]
[0015] Where α1, β, γ, δ, α2, β2, γ2, α3, λ1, α4, α5, β5 are weight coefficients, Z t is the state representation at the current time step t, A t is the action at the current time step t, R(Z t , A t ) is the reward function when the action A t is executed in the state Z t . The efficiency function R eff (Z t , A t ) represents the energy consumed by the aircraft when the action A t is executed in the state Z t . The safety function R safety (Z t ) represents the safety of the aircraft when the action A t is executed in the state A t . The task completion function R task (Z t , A t ) represents the task progress of the aircraft when the action A t is executed in the state Z t . The environmental adaptability function R env (Z t , A t ) represents the environmental adaptability of the aircraft when the action A t is executed in the state Z t . d actual is the actual flight distance of the aircraft when the action A t is executed in the state Z t . d shortest is the straight-line distance from the task start point to the end point, Δv avg is the deviation of the instantaneous speed of the aircraft from the average speed within the set time window when the action A t is executed in the state Z t . v avgis the average speed of the aircraft under ideal conditions, and Δa is the change in acceleration when the aircraft performs action A in state Z t under action A t when the acceleration changes, a max is the maximum acceleration of the aircraft, k1 is a non-linear parameter, d obstacle is the distance between the aircraft and the nearest obstacle when the aircraft performs action A in state Z t under action A t ∈1 is a very small number to avoid a zero denominator, and current_progress is the task completion progress when the aircraft performs action A in state Z t under action A t total_progress is the total progress target of the task, task_priority is the priority value of the task, λ2 is the adjustment coefficient of the task priority, P max is the maximum value of the task priority, v actual is the actual speed of the aircraft when the aircraft performs action A in state Z t under action A t v optimal is the optimal speed of the aircraft in an ideal environment, v wind is the wind speed in the environment when the aircraft performs action A in state Z t under action A t δ5 is the coefficient of the influence of wind speed on flight speed.
[0016] Furthermore, the aircraft state data includes the aircraft position coordinates, speed, acceleration, distance to the nearest obstacle, task environment temperature, and wind speed.
[0017] Furthermore, the obtained aircraft state data is encoded and feature-extracted using an autoencoder to obtain the best latent representation of the state data.
[0018] Furthermore, the calculation formula for the learning rate μ of Q-learning is as follows:
[0019]
[0020] where μ0 is the preset initial learning rate, g t is the gradient value of the Q-value function at the current time step t, θ1 and θ2 are weight parameters, g avg is the average value of the gradient of the Q-value function, γ1 and γ2 are hyperparameters, T Q is the total number of training steps of Q-learning, t Q is the number of training steps that have been completed for the current Q-learning.
[0021] Furthermore, the ε-greedy strategy is adopted for aircraft action selection, where the exploration rate ε is calculated as follows:
[0022]
[0023] In the formula, ε0 is the preset initial exploration rate, t Q is the number of training steps that the current Q - learning has completed, T Q is the total number of training steps of Q - learning, τ1 is the hyperparameter that controls the decay of the exploration rate during the training process, σ1, σ2 and σ3 are weight coefficients, ΔR t is the difference between the reward function R(Z t , A t ) at the current time step t and the reward function R(Z t-1 , A t-1 ) at the previous time step t - 1, Z t-1 is the state representation at the previous time step t - 1, A t-1 is the action at the previous time step t - 1, R avg is the average value of the reward function, Var(Z t ) is the variance of the state Z t .
[0024] The present invention also adopts an intelligent decision - making system for aircraft tasks based on reinforcement learning, including one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method as described above.
[0025] The present invention also adopts a computer - readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to execute the method as described above.
[0026] Beneficial effects: Compared with the prior art, the present invention overcomes the limitations of traditional aircraft relying on preset flight paths and fixed control algorithms by introducing an intelligent decision - making method based on reinforcement learning. Different from static control strategies, the present invention enables the aircraft to learn environmental changes in real time and dynamically adjust decisions, so as to autonomously optimize flight strategies in complex and dynamic environments. The specific innovations are reflected in the dimensionality reduction and feature abstraction of flight state data by the auto - encoder technology, the design of the dynamic reward function, and the application of the ε - greedy strategy. This solution not only improves the safety and energy efficiency of the aircraft during task execution, reduces the collision risk, but also enhances the autonomy and adaptability of the aircraft in complex environments, significantly improving the task execution efficiency and success rate. The specific details include:
[0027] · Improve the autonomy and adaptability of the aircraft: Through intelligent decision-making based on reinforcement learning, the aircraft can optimize the decision-making strategy in real time during flight and adapt to the dynamically changing environment (such as changes in obstacles, weather, etc.). This advantage enables the aircraft to perform tasks in complex environments without relying too much on manual intervention.
[0028] · Enhance the safety of the aircraft: The safety assessment in the dynamic reward function can ensure that the aircraft avoids collisions with obstacles or entering dangerous areas, improving the safety of the flight process. In addition, through the strategy optimization of reinforcement learning, the aircraft can select the safest and most efficient flight path.
[0029] · Optimize the execution efficiency of flight tasks: The present invention encourages the aircraft to minimize energy consumption and flight time through the efficiency assessment in the reward function. By learning the shortest path, avoiding obstacles and selecting the best flight altitude, the aircraft can significantly improve the task completion efficiency and reduce unnecessary resource waste.
[0030] · Support the execution of tasks in complex or unknown environments: Different from the prior art, the present invention can cope with unknown obstacles, weather changes or emergencies through continuous learning during flight. The decision-making ability of the aircraft is no longer limited to preset rules, but can make real-time adjustments according to environmental changes, so as to effectively complete tasks in a dynamically complex environment. Brief Description of the Drawings
[0031] Figure 1 is the flowchart of the method of the embodiment of the present invention. Detailed Embodiment
[0032] As Figure 1 shown, an intelligent decision-making method for aircraft tasks based on reinforcement learning of the present invention includes the following steps:
[0033] (1) Collect aircraft state data: Collect the original state data S t from multiple sensor systems of the aircraft, including position coordinates X, Y, Z (unit: meter), speed V x , V y , V z (unit: meter / second), acceleration A x , A y , A z (unit: meter / second^2), the distance D to the nearest obstacle min (unit: meter), environmental temperature T (unit: degree Celsius) and wind speed T (unit: meter / second).
[0034] (2) Encoding and feature abstraction of state data;
[0035] (3) Establish a dynamic reward function;
[0036] (4) Policy optimization based on Q-learning;
[0037] (5) Implement the ε-greedy policy to select actions.
[0038] Preferably, the encoding and feature abstraction method of the state data is as follows:
[0039] Apply the autoencoder E to the state data S t for encoding and dimensionality reduction to generate an optimized state representation Z t . The autoencoder includes an encoder θ e and a decoder θ d , and the training objective is to minimize the reconstruction error:
[0040]
[0041] In the formula, S t is the actual state of the aircraft at time step t. The encoder maps the input state S t to a low-dimensional latent representation Z t through Z e = e(S t ; θ t ), and then the decoder restores the latent representation to reconstructed data that is as close as possible to the original state space.
[0042] (1) Encoder θ e : Map the input aircraft state data x (including information such as position, velocity, and acceleration)
[0043] to a low-dimensional latent space representation Z t , that is:
[0044] Z t = f encoder (x; θ encoder ) = σ(W enc x + b enc )
[0045] Where:
[0046] ·f encoder represents the encoder function;
[0047] ·θ encoder are the parameters of the encoder;
[0048] ·x is the input high-dimensional state data;
[0049] ·Z t is the low-dimensional representation;
[0050] ·W enc is the weight matrix of the encoder;
[0051] ·b enc is the bias term of the encoder;
[0052] ·σ is the activation function.
[0053] (2) Decoder θ d : Restores the low-dimensional latent space representation Z t to the original state data through the decoder
[0054]
[0055] Where:
[0056] ·f decoder represents the decoder function;
[0057] ·θ decoder are the parameters of the decoder;
[0058] · is the reconstructed data output by the decoder;
[0059] ·W dec is the weight matrix of the decoder;
[0060] ·b dec is the bias term of the decoder;
[0061] ·σ is the activation function.
[0062] Preferably, the established dynamic reward function is:
[0063] R(Z t ,A t ) = α1R eff (Z t ,A t ) + βR safety (Z t ,A t ) + γR task (Z t ,A t ) + δR env (Z t ,A t ) where α1, β, γ, δ are weight coefficients.
[0064] Specifically, the efficiency function R eff (Z t ,A t ) : The efficiency function is used to represent the energy consumed when the aircraft executes a certain action, and the unit is joule (J). The goal of this function is to encourage the aircraft to save energy as much as possible when performing tasks. The mathematical expression of the efficiency function is:
[0065]
[0066] ·d actual is the actual flight distance, d shortest is the straight-line distance from the starting point to the ending point;
[0067] ·Δv avg is the change in the average speed of the aircraft, representing the efficiency of speed control;
[0068] ·v avg is the average speed of the aircraft under ideal conditions;
[0069] ·Δa is the change in acceleration, measuring the acceleration fluctuation of the aircraft during flight;
[0070] ·a max is the maximum acceleration of the aircraft;
[0071] ·α2, β2, γ2 are weight coefficients used to adjust the relative importance of path efficiency, speed control, and acceleration control.
[0072] Specifically, the safety function R safety (Z t ,A t ): The safety function measures the safety of the aircraft when performing a certain action, ensuring that the aircraft avoids collisions with obstacles or entering other dangerous areas. The definition of the safety function is as follows:
[0073]
[0074] ·λ1 is the weight of the obstacle threat, adjusting the sensitivity of threat assessment;
[0075] ·k1 is a non-linear parameter used to describe the non-linear relationship between the obstacle distance and the threat;
[0076] ·d obstacle is the distance between the aircraft and the obstacle;
[0077] ·α3 is the weight coefficient, ∈1 is an extremely small number to prevent the denominator from being zero, generally taking 0.01;
[0078] ·This part weights the risk when the distance to the obstacle is relatively close through an exponential function, ensuring that the aircraft can avoid the area where a collision is about to occur.
[0079] Specifically, the task completion function R tsdk (Z t ,A t):The task completion function is used to measure the progress of the aircraft during mission execution. Combining the current task completion and the total task completion, it returns a value between [0,1]. The formula for task completion is as follows:
[0080]
[0081] ·current_progress is the completion progress of the current task;
[0082] ·total_progress is the total progress target of the task;
[0083] ·λ2 is the adjustment coefficient of task priority;
[0084] ·task_priority is the priority value of the task (for example, an urgent task has a higher priority);
[0085] ·P max is the maximum value of task priority;
[0086] ·α4 is the weight coefficient;
[0087] ·This function adjusts the task completion through the sine function, so that high-priority tasks can receive additional rewards.
[0088] Specifically, the environmental adaptability function R env (Z t ,A t ): The adaptability of the aircraft in different environments needs to consider factors such as wind speed, temperature, and air pressure. We take the matching degree between the speed of the aircraft and environmental factors as the basis for adaptability rewards and introduce the influence of wind speed changes.
[0089]
[0090] ·v actual is the actual speed of the aircraft;
[0091] ·v optimal is the optimal speed in the ideal environment;
[0092] ·v wind is the wind speed in the current environment;
[0093] ·δ5 is the coefficient of the influence of wind speed on flight speed;
[0094] ·β5 is the weight coefficient of wind speed adaptability;
[0095] ·α5 is the weight coefficient.
[0096] Preferably, the Q-learning-based policy optimization method is as follows:
[0097] Update the Q-value Q(Z t , A t ) by:
[0098] Q(Z t , A t ) ← Q(Z t , A t ) + μ[R t + γ max a′ Q(Z t+1 , a′) - Q(Z t , A t )]
[0099] Where:
[0100] · Q is the Q-function adopted by Q-learning;
[0101] · μ is the learning rate of Q-learning, calculated by the following formula:
[0102]
[0103] μ0 is the preset initial learning rate of Q-learning;
[0104] g t is the gradient value of the Q-value function at time step t (which can be the norm of the gradient);
[0105] g avg is the average value of the gradient of the Q-value function, representing a smooth gradient level, which can be obtained by moving average calculation;
[0106] The two parameters θ1 and θ2 are used to control the influence degree of different factors on the learning rate, usually ranging from 0 to 1;
[0107] The two hyperparameters γ1 and γ2 are used to control the non-linear influence degree of the gradient magnitude and the training progress on the learning rate, affecting the adjustment rate of the learning rate;
[0108] t Q is the number of training steps (epoch) completed by the current Q-learning;
[0109] T Q is the total number of training steps (epoch) of Q-learning, that is, the maximum number of iterations of training.
[0110] · γ is the discount factor of future rewards in Q-learning;
[0111] · Q(Zt , A t ) is the current state Z t to select action A t of the Q-value;
[0112] · max a′ Q(Z t+1 , a′) is the maximum value among the Q-values of all possible actions in the next state Z t+1 .
[0113] Once the Q-value function is updated and converges, the agent can select the optimal policy through the Q-value. According to the current Q-value function, the agent can select the optimal action in each state:
[0114] π(Z t ) = argmax(Z t , a)
[0115] For each state Z t , the agent will select the action a with the largest Q-value as the execution policy.
[0116] Preferably, the method for selecting an action by implementing the ε-greedy policy is as follows:
[0117] Action A t is selected according to the following policy:
[0118]
[0119] Where:
[0120] · ε is the exploration rate, which determines the frequency of exploring new behaviors in action selection to promote the exploration of the new state space. Its calculation formula is:
[0121]
[0122] Where:
[0123] · ε0 is the preset initial exploration rate, usually set to a relatively large value for strong exploration at the beginning;
[0124] · T Q That is, the total number of training steps (epoch), which is the maximum number of iterations of training.
[0125] · τ1 is a hyperparameter that controls the decay of the exploration rate during training, usually taking values from 0 to 1;
[0126] · t Q is the current training step (epoch) of Q-learning, and the exploration rate is updated synchronously with Q-learning;
[0127] · ΔR tis the reward difference (or profit difference) between the current time step and the previous time step of the reinforcement learning reward function. It reflects the magnitude of the policy change, i.e., |ΔR t | = |R(Z t ) - R(Z t-1 )|;
[0128] ·R avg is the average value of the reinforcement learning reward function R(Z t ) during a training process, which can be represented by the average reward in the entire training process or the moving average reward calculated according to a certain time window;
[0129] ·σ1, σ2 are the weight coefficients that adjust the influence of the reward difference on the exploration rate; σ3 is the coefficient that adjusts the influence of the state space uncertainty on the exploration rate;
[0130] ·Var(Z t ) is the variance of the current state Z t , Var(Z t ) = Var(Q(Z t , a)), where a is the action taken according to the Q value. Var(Z t ) represents the distribution or uncertainty degree of this state. When the state variance is large, the exploration rate is high, indicating uncertainty about the current state and the system needs more exploration.
[0131] To make the content of the present invention easier to understand clearly, the following further elaborates on the present invention in combination with specific embodiments.
[0132] Task description: The drone needs to fly from the distribution center (starting point) of the city to a specific customer location (ending point) to deliver a package. The task environment includes dynamic obstacles such as high-rise buildings and other flying vehicles, as well as changing weather conditions.
[0133] The following will adopt an intelligent decision-making method, device, equipment, and system for aircraft tasks based on reinforcement learning proposed by the present invention:
[0134] 1. Collect flight state data
[0135] · Location: The starting point coordinates are X 0, Y 0, Z0 = (0, 0, 10) meters, and the ending point coordinates are X g, Y f, Z f =
[0136] (1000, 1000, 10) meters.
[0137] · Speed: The initial speed is V x , V y , Vz =(0, 0, 0) m / s.
[0138] · Obstacle distance: The distance to the nearest obstacle D min , starting at 100 m, no direct obstacles.
[0139] · Environmental factors: The starting wind speed W = 5 m / s, eastward.
[0140] 2. State Data Encoding and Feature Abstraction
[0141] The goal is to convert flight state data (including position, velocity, acceleration, environmental parameters, etc.) into a compressed, information-rich low-dimensional feature representation, which helps the subsequent reinforcement learning process to process and make decisions more efficiently.
[0142] Input data preparation:
[0143] (1) Data collection
[0144] · Position X, Y, Z: The current coordinates of the UAV, possibly provided by GPS and other positioning systems.
[0145] · Velocity V x , V y , V z : The velocity vector of the UAV, captured by the IMU (Inertial Measurement Unit).
[0146] · Acceleration A x , A y , D z : The acceleration vector of the UAV, also captured by the IMU.
[0147] · Obstacle distance D min : The distance between the UAV and the nearest obstacle, possibly obtained through a radar or LiDAR system.
[0148] · Environmental factors: Such as wind speed W, obtained through a weather sensor.
[0149] (2) Data preprocessing
[0150] · Normalize all continuous variables, i.e., subtract the mean from each variable and divide by the standard deviation, or perform min-max normalization so that the data falls within a specific range (such as 0 to 1).
[0151] Autoencoder design:
[0152] (1) Encoder θ e Construction:
[0153] · Construct a deep neural network, including several fully connected layers, gradually reducing the number of neurons to achieve dimensionality reduction.
[0154] · Example structure: Input layer (receiving normalized state data), the middle layer gradually reduces the number of neurons (e.g., 128 -> 64 -> 32), and uses the ReLU activation function to promote non-linear learning.
[0155] (2) Decoder θ d Construction:
[0156] · Contrary to the encoder structure, it gradually increases the number of neurons from the low-dimensional feature space to the dimension of the original data.
[0157] · Train the autoencoder through the reconstruction error of the decoding process, with the goal of minimizing the difference between the input data and the reconstructed output.
[0158] Training and optimization:
[0159] · Use the historical flight data of drones to train the autoencoder to ensure that the model can effectively learn how to extract useful low-dimensional features from the original high-dimensional data.
[0160] · Optimization methods such as using the Adam optimizer and the mean squared error (MSE) loss function, adjusting the learning rate and batch size to find the best training parameters.
[0161] 3. Establish a dynamic reward function
[0162] Define the reward function:
[0163] R(Z t , A t ) = α1R eff (Z t , A t ) + βR safety (Z t , A t ) + γR task (Z t , A t ) + δR env (Z t , A t )
[0164] (1) Efficiency function R eff (Z t , A t ):
[0165] Assume the following parameters are known:
[0166] · d actual = 200m, d shortest = 100m
[0167] · Δv avg = 5
[0168] ·v avg = 5
[0169] ·Δa = 5
[0170] ·a max = 10
[0171] ·α2, β2, γ2 = 1
[0172]
[0173] (2) Safety function R safety (Z t ):
[0174] Assume the following parameters are known:
[0175] ·λ1 = 1
[0176] ·k1 = 1
[0177] ·d obstacle = 10
[0178] ·α3 = 1, ∈1 = 0.01
[0179]
[0180] (3) Task completion function R task (Z t , A t ):
[0181] Assume the following parameters are known:
[0182] ·current_progress = 5
[0183] ·total_progress = 5
[0184] ·λ2 = 0
[0185] ·task_priority = 1
[0186] ·P max = 1
[0187] ·α4 = 1
[0188]
[0189] (4) Environmental adaptability function R env (Z t , A t ):
[0190] Assume the following parameters are known:
[0191] ·v actual = 10
[0192] ·v optimal = 15
[0193] ·v wind = 10
[0194] ·δ5 = 1
[0195] ·β5 = 0.5
[0196] ·α5 = 1
[0197]
[0198] 4. Q - learning - based Policy Optimization
[0199] In this step, we use the Q - learning method to update the decision - making policy of the aircraft. Q - learning is a model - free reinforcement learning algorithm that uses a data structure called a Q - table to store and update the Q - values for each state - action pair, which represent the expected total reward obtained by performing a certain action in a given state.
[0200] (1) Calculate the reward R:
[0201] Assume the weights α1 = 1, β = 1, γ = 1.
[0202] R(Z t ) = α1R eff (Z t , A t ) + βR safety (Z t ) + γR task (Z t , A t ) + δR env (Z t , A t )
[0203] = - 3.5 + 0.1 + 1 + 0.71 = - 1.69
[0204] (2) Update the Q - value:
[0205] Adopt the following Q - learning update formula:
[0206] Q(Z t , A t ) ← Q(Z t , A t ) + μ[R t + γmax a′ Q(Z t+1 , a′) - Q(Z t , At )]
[0207] where μ is the learning rate,
[0208] μ0 is the initial learning rate;
[0209] g t is the gradient value at time step t (which can be the norm of the gradient);
[0210] g avg is the average value of the gradient, representing a smoothed gradient level, which can be obtained through moving average calculation;
[0211] The two parameters θ1 and θ2 are used to control the influence degree of different factors on the learning rate, usually ranging from 0 to 1;
[0212] The two hyperparameters γ1 and γ2 are used to control the non - linear influence degree of the gradient magnitude and training progress on the learning rate, affecting the adjustment rate of the learning rate;
[0213] t is the current training step (epoch);
[0214] T is the total number of training steps (epoch), that is, the maximum number of iterations for training.
[0215] γ is the discount factor for future rewards, taking 0.2.
[0216] Assume the initial Q(s,a) is 0, and update the Q value according to the above R value and the possible maximum Q value in the future (assuming the maximum Q value is 0.5).
[0217] 5. Implement the ε - greedy policy to select actions
[0218] In the ε - greedy policy, usually a higher ε value (such as 0.1) is set in the initial stage of the policy. As the policy is continuously optimized, the ε value gradually decreases to increase the frequency of using the known best policy.
[0219] Determine ε:
[0220]
[0221] where:
[0222] · ε0 is the initial exploration rate, usually set to a relatively large value for stronger exploration at the beginning;
[0223] · TQ The total number of training steps (or total number of training iterations), which is used to regulate the progress of training;
[0224] · τ1 is a hyperparameter that controls the decay of the exploration rate during the training process, usually taking values between 0 and 1;
[0225] · t is the current time step (epoch);
[0226] · t Q The number of training steps that the current Q - learning has completed;
[0227] · ΔR t The reward difference (or return difference) between the current time step and the previous time step. It reflects the magnitude of the policy change;
[0228] · R avg The average value of the rewards, which can be represented by the mean reward over the entire training process, or the moving - average reward calculated according to a certain time window;
[0229] · σ1, σ2 are weight coefficients that adjust the influence of the reward difference on the exploration rate;
[0230] · σ3 is a coefficient that adjusts the influence of the state - space uncertainty on the exploration rate;
[0231] · Var(Z t ) The variance of the current state s t represents the degree of distribution or uncertainty of this state. When the state variance is large, the exploration rate is high, indicating uncertainty about the current state and the system needs more exploration.
[0232] In this way, the UAV can maintain a certain exploration rate while, with the accumulation of experience, gradually increasing the utilization of known good policies, so as to achieve more stable and efficient flight control in a dynamic and complex environment. Such implementation details enable the model to not only cope with the initial uncertainty but also adapt to long - term environmental changes and task requirements.
[0233] The present invention optimizes the task execution strategy of an aircraft in a complex environment by integrating an autoencoder and a Q-learning framework. The ε-greedy strategy is introduced, which effectively balances the relationship between exploring new behaviors and exploiting the known best behaviors. This enables the aircraft to explore new action strategies and gradually optimize the known best strategies in a dynamic and complex environment, thereby avoiding premature convergence to local optimal solutions and enhancing decision-making efficiency and autonomy. In addition, the autoencoder technology is used to reduce the dimensionality of the high-dimensional state data of the aircraft, simplifying the data complexity and improving the efficiency of the subsequent decision-making process. Different from traditional reinforcement learning methods that directly process high-dimensional data, the use of an autoencoder makes the system more efficient in processing diverse aircraft state data. In terms of the reward mechanism, a multi-dimensional dynamic reward function is designed, comprehensively considering factors such as operation efficiency, safety, and task completion. In contrast, existing technologies usually focus on single-dimensional optimization and cannot comprehensively improve the task execution ability of the aircraft. Overall, the present invention significantly improves the decision-making ability, safety, and efficiency of the aircraft in a complex environment by introducing the greedy strategy, optimizing data processing, and the reward mechanism, and is particularly applicable to unmanned aircraft and other autonomous aircraft that require high autonomy and adaptability.
[0234] Based on the same technical solution, the present invention also provides an intelligent decision-making system for aircraft tasks based on reinforcement learning, including one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the above-mentioned intelligent decision-making method for aircraft tasks based on reinforcement learning.
[0235] Based on the same technical solution, the present invention also provides a computer-readable storage medium storing one or more programs, the one or more programs including instructions, characterized in that when the instructions are executed by a computing device, the computing device is caused to execute the above-mentioned intelligent decision-making method for aircraft tasks based on reinforcement learning.
[0236] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0237] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0238] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0239] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.
[0240] Obviously, the above embodiments are only examples for clear illustration and not limitations on the usage. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all usage manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. An intelligent decision-making method for aircraft missions based on reinforcement learning, characterized in that, It includes the following steps: (1) Obtain the state data of the aircraft; (2) Establish a dynamic reward function; (3) Optimize the aircraft decision-making based on Q-learning; (4) Adopt the ε-greedy strategy to select the aircraft actions; Among them, the expression of the dynamic reward function is: R(Z t ,A t ) = α1R eff (Z t ,A t ) + βR safety (Z t ,A t ) + γR task (Z t ,A t ) + δR env (Z t ,A t ) Where α1, β, γ, δ, α2, β2, γ2, α3, λ1, α4, α5, β5 are weight coefficients, and Z t is the state representation at the current time step t, A t is the action at the current time step t, R(Z t , A t ) is the reward function when performing action A t in state Z t . The efficiency function R eff (Z t , A t ) represents the energy consumed by the aircraft when performing action A t in state Z t . The safety function R safety (Z t ) represents the safety of the aircraft when performing action A t in state Z t . The task completion function R task (Z t , A t ) represents the task progress of the aircraft when performing action A t in state Z t . The environmental adaptability function A env (Z t , A t ) represents the environmental adaptability of the aircraft when performing action A t in state Z t . d actual is the actual flight distance of the aircraft when performing action A t in state Z t . d shortest is the straight-line distance from the task start point to the end point. Δv avg is the deviation of the instantaneous speed from the average speed within the set time window when the aircraft performs action A t in state Z t . v avg is the average speed of the aircraft under ideal conditions. Δa is the change in acceleration when the aircraft performs action A t in state Z t . a max is the maximum acceleration of the aircraft. k1 is a non-linear parameter. d obstacle is the distance to the nearest obstacle when the aircraft performs action A t in state Z t . ∈1 is a very small number to avoid a zero denominator. current_progress is the progress when performing action A t in state Z t The task completion progress at a certain time, total_progress is the total progress target of the task, task_priority is the priority value of the task, λ2 is the adjustment coefficient of the task priority, P max is the maximum value of the task priority, v actual is the actual speed of the aircraft when performing action A t in state X t v optimal is the optimal speed of the aircraft in an ideal environment, v wind is the aircraft in state Z t when performing action A t The wind speed in the environment at this time, δ5 is the coefficient of the influence of wind speed on flight speed.
2. The method according to claim 1, wherein The aircraft state data includes the aircraft position coordinates, speed, acceleration, distance to the nearest obstacle, task environment temperature, and wind speed.
3. The method according to claim 1, characterized in that It also includes encoding and feature extraction of the obtained aircraft state data by applying an autoencoder to obtain the best latent representation of the state data.
4. The method according to claim 1, characterized in that, The calculation formula of the learning rate μ of Q-learning is as follows: where μ0 is the preset initial learning rate, g t is the gradient value of the Q-value function at the current time step t, θ1 and θ2 are weight parameters, g avg is the average value of the gradient of the Q-value function, γ1 and γ2 are hyperparameters, T Q is the total number of training steps for Q-learning, t Q is the number of training steps that have been completed for the current Q-learning.
5. The method according to claim 1, characterized in that, Adopt the ε-greedy strategy to select the aircraft actions, where the calculation formula of the exploration rate ε is as follows: where ε0 is the preset initial exploration rate, t Q is the number of training steps that have been completed in the current Q-learning, T Q is the total number of training steps of Q-learning, τ1 is the hyperparameter that controls the decay of the exploration rate during the training process, σ1, σ2, and σ3 are weight coefficients, and ΔR t is the difference between the reward function R(Z t , A t ) at the current time step t and the reward function R(Z t-1 , A t-1 ) at the previous time step t - 1, A t-1 is the state representation at the previous time step t - 1, A t-1 is the action at the previous time step t - 1, R avg is the average value of the reward function, and Var(Z t ) is the variance of the state Z t .
6. An intelligent decision-making system for aircraft missions based on reinforcement learning, characterized in that, It includes one or more processors, one or more memories, and one or more programs, where the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing the method according to any one of claims 1 to 5.
7. A computer-readable storage medium storing one or more programs, the one or more programs including instructions, characterized in that, When the instructions are executed by a computing device, the computing device is caused to execute the method according to any one of claims 1 to 5.