An aircraft recovery dynamic scheduling method based on a long short-term memory network
By using a deep neural network architecture based on long short-term memory and a priority experience replay mechanism, the problem of low efficiency in traditional aircraft recovery scheduling is solved, and efficient and safe dynamic scheduling of aircraft recovery is achieved, adapting to complex and dynamic environmental changes.
Patent Information
- Application Number
- CN202511204902.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Traditional aircraft recovery scheduling relies on human experience and is inefficient in complex situations. Existing optimization algorithms are difficult to meet the requirements of real-time response and adaptability to dynamic environments. Deep reinforcement learning faces challenges in aircraft recovery scheduling, such as time-series dependence, massive data processing, and reward mechanism design.
A deep neural network architecture based on long short-term memory is adopted, combined with a priority experience replay and dynamic threshold random judgment mechanism to build an intelligent scheduling model. The model generates retrieval trajectories and updates network parameters through a cyclic strategy process, thereby achieving efficient learning and adaptation to random environments.
It achieves real-time and efficient dynamic scheduling of aircraft recovery, and can deeply understand the temporal dynamics and adapt to random environments, thereby improving the safety and efficiency of scheduling.
Smart Images

Figure CN120725394B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence and automatic scheduling, in particular to an aircraft recovery dynamic scheduling method based on a long short-term memory network. BACKGROUND
[0002] Aircraft recovery scheduling, especially on platforms with extremely limited space and time resources, is a core link to ensure flight safety. The process requires the scheduler to develop a safe and efficient recovery sequence within a very short time according to a large amount of real-time dynamic information of the aircraft, such as remaining fuel, task urgency, aircraft state, and potential recovery failure risk. This is a typical dynamic stochastic scheduling problem, which is complex in many aspects. First, the decision environment is dynamic, and the state of each aircraft, especially the fuel, deteriorates over time. Second, the decision has strong time sequence, and a current recovery decision directly affects the state of all subsequent aircraft and the decision space. Third, the recovery process is accompanied by uncertainty, and the aircraft may need to fly again due to various factors, disrupting the original plan.
[0003] Traditional aircraft recovery mainly relies on the experience of human schedulers. This approach can cope with low-density and routine situations, but in complex situations such as high-intensity, multi-batch, and mixed recovery of multiple aircraft types, the human brain is difficult to handle large information flows and complex constraint relationships, leading to decision delays, low efficiency, and even safety hazards. Some optimization algorithms based on mathematical programming, such as mixed integer programming, can theoretically find the optimal solution, but their computational complexity grows exponentially, making it impossible to meet the demand for real-time response. Heuristic and meta-heuristic algorithms, such as genetic algorithms or simulated annealing, can quickly provide solutions, but their rules are usually static and artificially designed, making it difficult to adapt to dynamic changes in the environment and random events, and their generalization ability is limited.
[0004] In recent years, deep reinforcement learning technology has shown great potential in solving such sequential decision-making problems. However, there are still challenges in applying traditional reinforcement learning methods directly to aircraft recovery scheduling. First, how to effectively represent and handle the time-dependent nature of the problem, i.e., the influence of historical decisions and states on the current optimal choice. Second, how to efficiently learn from vast amounts of experience data and avoid being overwhelmed by a large number of low-value and repetitive experiences. Third, how to design a reward mechanism that accurately reflects the complex scheduling goals, and how to build a high-fidelity training environment that simulates real-world randomness. Therefore, developing an intelligent scheduling method that can deeply understand the time-dependent dynamics, efficiently learn, and adapt to random environments is a key problem that needs to be solved urgently in the current technical field. SUMMARY
[0005] The purpose of the present application is to provide an aircraft recovery dynamic scheduling method based on a long short-term memory network, so as to deeply understand the timing dynamics, realize efficient learning, and adapt to the intelligent scheduling recovery of the random environment.
[0006] To achieve the above purpose, the present application provides the following solutions.
[0007] The present application provides an aircraft recovery dynamic scheduling method based on a long short-term memory network, comprising the following steps.
[0008] Initialize the policy network, the target network and the priority experience replay memory, and set up multiple parallel aircraft recovery simulation environments; the policy network and the target network are both deep neural networks containing a long short-term memory network layer and a multi-layer perceptron layer; the priority experience replay memory loads initial data containing multiple aircrafts to be recovered; the initial data at least includes an aircraft model matrix and a feature matrix of initial attributes of each aircraft.
[0009] In each parallel aircraft recovery simulation environment, generate a recovery trajectory through a loop policy process, and recover the aircraft according to the recovery trajectory in each decision step, and store it in the priority experience replay memory until the recovery task in all parallel aircraft recovery simulation environments is completed.
[0010] When the sample quantity in the priority experience replay memory meets the preset batch quantity, sample a batch of experience data from the priority experience replay memory, and calculate the Q value of the next state of any experience sample in the experience data using the target network, and update the parameters of the policy network .
[0011] According to the parameters of the policy network , the parameters of the target network are updated synchronously , the network update parameters after each update are recorded, and the parameters of the policy network are saved according to the preset saving period ; the network update parameters include a weighted loss function, an average recovery reward and an exploration rate.
[0012] Based on the network update parameters, determine the aircraft recovery dynamic scheduling model according to the updated target network and the updated policy network to perform the recovery task.
[0013] According to the specific embodiments provided by the present application, the present application has the following technical effects.
[0014] The application is carried out in a parallel simulation environment, generates a recovery trajectory through a loop strategy process, recovers the aircraft according to the recovery trajectory in each decision step, and stores it into a priority experience replay memory until all parallel aircraft recovery simulation environments complete the recovery task to achieve a large amount of interactive training; when the sample amount in the priority experience replay memory meets the preset batch quantity, a batch of experience data is sampled from the priority experience replay memory, and the Q value of the next state of any experience sample in the experience data is calculated by using the target network to update the parameters of the strategy network , realizing priority-based network parameter updating; the parameters of the strategy network synchronously update the parameters of the target network to complete the periodic synchronization of the target network; and record the network update parameters after each update, and save the parameters of the strategy network according to the preset saving period , realizing continuous monitoring and saving of the model; finally, based on the network update parameters, the aircraft recovery dynamic scheduling model is determined according to the updated target network and the updated strategy network to execute the recovery task. Through a large amount of interactive training, priority-based network parameter updating, periodic synchronization of the target network, and continuous monitoring and saving of the model, the application finally obtains an aircraft recovery dynamic scheduling model that can respond in real time and make efficient decisions, so as to deeply understand the timing dynamics, realize efficient learning, and adapt to the intelligent scheduling recovery of the random environment. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0016] Figure 1 A flowchart of a long short-term memory network-based aircraft recovery dynamic scheduling method according to an embodiment of the present application. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0018] To make the objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] like Figure 1 As shown in the figure, this application provides a dynamic scheduling method for aircraft recovery based on long short-term memory networks, including the following steps.
[0020] S1: Initialize the policy network, target network, and priority experience replay memory, and set up multiple parallel aircraft recovery simulation environments; the policy network and the target network are both deep neural networks containing long short-term memory network layers and multiple perceptron layers; the priority experience replay memory is loaded with initial data containing multiple aircraft to be recovered; the initial data includes at least an aircraft model matrix and a feature matrix of initial attributes of each aircraft.
[0021] S2: In each parallel aircraft recovery simulation environment, a recovery trajectory is generated through a cyclical strategy process, and the aircraft is recovered according to the recovery trajectory in each decision step and stored in the priority experience replay memory until the recovery tasks in all parallel aircraft recovery simulation environments are completed.
[0022] S3: When the sample size in the priority experience replay memory meets the preset batch size, sample a batch of experience data from the priority experience retrieval memory, and use the target network to calculate the Q value of the next state of any experience sample in the experience data, and update the parameters of the policy network. .
[0023] S4: Based on the parameters of the policy network Synchronously update the parameters of the target network. Record the network update parameters after each update, and save the parameters of the policy network according to a preset saving period. The network update parameters include the weighted loss function, average recovery reward, and exploration rate.
[0024] S5: Based on the network update parameters, determine the aircraft recovery dynamic scheduling model according to the updated target network and the updated policy network to execute the recovery task.
[0025] In practical applications, policy networks are... The target network is The characteristic matrix is The number of aircraft to be recovered is The number of parallel aircraft recovery simulation environments is Priority Experience Replay Memory The capacity is The initial attributes of an aircraft include fuel quantity and inherent recovery success rate.
[0026] Within the preset total number of updates Within, execute S2-S5.
[0027] In an exemplary embodiment, the aircraft is recovered according to the recovery trajectory at each decision step and stored in the priority experience replay memory until the recovery tasks in all parallel aircraft recovery simulation environments are completed, specifically including the following steps.
[0028] S21: Obtain the current state sequence of the current decision step; the current state sequence includes aircraft state tensors for multiple time steps.
[0029] S22: Utilize the policy network Process the current state sequence and output the Q-values corresponding to all possible recycling actions, i.e. ,in, For the current decision-making steps The state sequence, This indicates all recycling actions. , recycling action Corresponding to from Select one of the aircraft to be recovered to perform a recovery attempt.
[0030] S23: Pass - The greedy strategy selects one recovery action from the available recovery actions; the recovery action is to select one of the multiple aircraft to be recovered to perform a recovery attempt.
[0031] S24: Execute the recycling action in the simulation environment, determine the next state sequence, scalar reward, and Boolean flag indicating whether the current recycling task is completed, and determine the quintuple experience; the quintuple experience includes the current state sequence. Recycling action Next state sequence Scalar reward and a Boolean flag indicating whether the current recycling task is complete. .
[0032] S25: Store the quintuple experience in the priority experience replay memory with the highest priority until all parallel aircraft recovery simulation tasks are completed.
[0033] In an exemplary embodiment, the target network is used to calculate the Q-value of the next state for any empirical sample in the empirical data, and the parameters of the policy network are updated accordingly. Specifically, it includes the following steps.
[0034] S31: Calculate the sampling probability of each experience sample according to the priority stored in the priority experience replay memory, and extract batch data containing multiple experience samples, while obtaining the index corresponding to the batch samples in the batch data and importance sampling weight ; the batch data is experience data of the same batch.
[0035] S32: Calculate the Q value of the next state of each experience sample in the batch data using the target network, and construct the time difference target combining the reward and the discount factor.
[0036] S33: Calculate the Q value of the current state using the policy network, and calculate the time difference error of each experience sample according to the Q value of the current state and the time difference target. Wherein, the calculation formula of time difference error .
[0037] S34: Update the sample priority corresponding to the index in the priority experience replay memory according to the time difference error.
[0038] S35: Based on the updated sample priority, calculate the weighted loss function according to the time difference error and the importance sampling weight, and update the parameters of the policy network by gradient descent method .
[0039] In practical applications, at intervals of a preset update period , the parameters of the policy network are completely copied to the parameters of the target network , The value range of is [100, 5000] decision steps.
[0040] After each update, record and save the value of the loss function , the average reward and the exploration rate , and save the parameters of the policy network at a preset saving period , wherein The value range of
[0041] is [50, 500] update times. In an exemplary embodiment, the sampling probability
[0042] is as follows.
[0043] Wherein, is the priority of the experience sample . is a priority index; k is an index of all experience samples in the experience replay pool.
[0044] Upon update, the priority of an experience sample is updated to the absolute value of the time-difference error corresponding to the experience sample plus a small positive number to ensure a non-zero probability, is a positive number with a value range of . In an exemplary embodiment, the time-difference target is as follows.
[0045] In an exemplary embodiment, the time-difference target is as follows.
[0046]
[0047] wherein, is a scalar reward; is a discount factor with a value range of [0.95, 0.995]; is a next state sequence; is a recovery action for the next state; is a parameter of the policy network; is a parameter of the target network.
[0048] In an exemplary embodiment, the weighted loss function is as follows.
[0049]
[0050] wherein, B is a batch of experience data collected by sampling; , is an importance sampling weight of an experience sample , is a current total number of samples in the priority experience replay memory , is an importance sampling index with a value linearly annealed from an initial range of [0.4, 0.6] to 1.0; is a time-difference error , is a current state sequence, is a recovery action for the current state.
[0051] In an exemplary embodiment, the current state sequence is a three-dimensional tensor, which is ; wherein, is a history sequence length with a value range of [4, 16]; is the total number of aircrafts to be recovered; is the feature vector of a single aircraft; the feature vector of the single aircraft at least includes a normalized fuel amount, an intrinsic recovery success rate, and a discrete code representing a current recovery state.
[0052] The corresponding state of the discrete code includes a non-attempted recovery state with a discrete code of 0, a successfully recovered state with a discrete code of 1, and a just-performed-recovery-and-not-yet-to-be-recovered state with a discrete code of 2.
[0053] In an exemplary embodiment, S22 specifically includes the following steps.
[0054] S221: At each time step, concatenate the feature vectors of all aircrafts in the current state sequence to generate an input sequence tensor.
[0055] S222: Input the input sequence tensor into a long short-term memory network layer of the policy network to extract time sequence features and obtain a hidden state of the last time step.
[0056] S223: Input the hidden state into a multi-layer perceptron layer of the policy network to determine the Q values corresponding to all selectable recovery actions under the current state sequence.
[0057] In actual applications, the state sequence is obtained by concatenating the feature vectors of all aircrafts at each time step to form an input sequence tensor; the input sequence tensor is first extracted by a long short-term memory network layer containing layers to obtain a hidden state output of the last time step; then, the hidden state output is input into a multi-layer perceptron containing layers to finally output a Q value vector with a dimension of ; wherein has a value range of [1, 3], has a value range of [2, 4].
[0058] In an exemplary embodiment, the scalar reward is generated by a composite reward function; the composite reward function includes a fuel safety penalty term, an efficiency penalty term, a priority waiting penalty term, and a task completion reward term; wherein, when the fuel amount of any unrecovered aircraft is lower than a preset safety threshold, the fuel safety penalty term is a negative reward proportional to the fuel loss amount.
[0059] In an exemplary embodiment, S24 further includes the following steps.
[0060] S25: At the start of each recovery trajectory, generate a random recovery success rate threshold matrix for multiple aircraft in the current aircraft recovery simulation environment.
[0061] S26: After performing the recovery action at each decision step, determine whether the inherent recovery success rate of the selected aircraft is greater than or equal to the recovery success rate threshold corresponding to the inherent recovery success rate in the recovery success rate threshold matrix; if yes, execute S27, otherwise execute S28.
[0062] S27: The selected aircraft has been successfully recovered.
[0063] S28: Determine that the selected aircraft has failed to be recovered and execute a go-around.
[0064] S29: After the selected aircraft performs a go-around, the inherent recovery success rate of the selected aircraft is increased by a preset ratio in subsequent decision steps.
[0065] In practical applications, the simulated environment uses a random judgment mechanism based on dynamic thresholds to determine whether recycling is successful, which includes the following steps.
[0066] At the start of each recycling trajectory, for the current environment The aircraft generates a random recovery success rate threshold matrix. The generation process of the threshold matrix is based on a preset expected re-flight ratio. A relatively high random threshold is assigned to aircraft with a low inherent recovery success rate to ensure that the overall go-around probability meets expectations.
[0067] At each decision step Execute action Afterwards, compare the selected aircraft Inherent recovery success rate Rather than in the threshold matrix The corresponding threshold, if If the value is greater than or equal to the threshold, the recovery is considered successful; otherwise, the recovery is considered unsuccessful and a go-around is initiated.
[0068] When an aircraft performs a go-around, its inherent recovery success rate In subsequent decision-making steps, it is increased according to a preset ratio.
[0069] The technical solution core of the present application is to construct a complete deep Q network (DQN) reinforcement learning training framework. The agent core of the framework is a strategy network composed of a long short term memory (LSTM) network layer and a multi-layer perceptron (MLP) layer. The LSTM network structure is adopted to specifically capture and process the time series information in the aircraft recovery problem, so that the network can understand the importance of historical state evolution to the current decision. The state observed by the agent is designed as a sequence tensor containing the detailed attributes of all aircrafts in recent time steps, providing rich time series input for the LSTM network.
[0070] In order to improve the training efficiency, the present application adopts a priority experience replay mechanism. Unlike the uniform sampling of traditional experience replay, this mechanism assigns a priority to each experience sample according to its time difference error (i.e. the degree of unexpectedness of model prediction). During training, the agent will preferentially learn from high-priority experiences, thereby concentrating computing resources on the most difficult-to-predict and information-rich decision points, greatly accelerating the convergence speed and performance of the strategy.
[0071] Further, in order to enable the agent to learn how to deal with uncertainties in the recovery process, the present application proposes a novel, dynamic threshold-based random judgment mechanism to simulate the probability of recovery failure. This mechanism can generate a random success rate threshold for each aircraft at the beginning of each training according to the preset expected flyback ratio, so that the flyback events in the simulation environment are both random and statistically expected, thereby training a robust strategy capable of handling unexpected flyback events.
[0072] To make the purpose, technical solution and advantages of the present application clearer, a specific embodiment will be described in detail below to elaborate the aircraft recovery dynamic scheduling method based on long short term memory network proposed by the present application.
[0073] In this embodiment, the execution flow of the entire method follows a complete deep reinforcement learning training paradigm. First, system initialization is performed, at which time the system creates and initializes a strategy network and a target network with the same structure. The core of both networks consists of a long short term memory network layer and a subsequent multi-layer perceptron layer. At the same time, a priority experience replay memory with a preset capacity is established to store the experiences generated by the agent interacting with the environment. The system also loads pre-defined aircraft recovery problem instances, including the model of each aircraft, the initial fuel quantity, and the inherent recovery success rate, and starts multiple parallel simulation environments to accelerate data collection.
[0074] After the training starts, the system enters a macro iteration loop in the unit of update times. In each iteration, the agent interacts with multiple parallel aircraft recovery simulation environments for a complete recovery task, i.e., a "round". At the beginning of a "round", each aircraft recovery simulation environment is reset to the initial state. In particular, at this time, the environment calls its internal random judgment mechanism to generate a random recovery success rate threshold for each aircraft according to the preset expected reflight ratio. This threshold remains unchanged throughout the "round" and is used for subsequent judgment of whether the recovery is successful.
[0075] Subsequently, at each decision time step, the agent begins to make sequential decisions. First, it obtains the current state. This state is not a snapshot at a single time point, but a sequence of aircraft state tensors containing the states of the aircraft at the last several time steps. Each state tensor records in detail the normalized fuel quantity, the priority of each aircraft, and a discrete code representing its current recovery state (e.g., not attempted, successful, or just reflighted). This state sequence is then fed into the policy network. The LSTM layer of the network is responsible for processing sequence information and capturing the dynamic trend of aircraft state changes over time, while the MLP layer calculates a Q value for each aircraft that has not yet been recovered based on the output of the LSTM, which represents the long-term value of choosing this aircraft for recovery. The agent selects an aircraft as the recovery action for this time according to these Q values and an exploration rate that decreases with the training process through an ε-greedy strategy.
[0076] When the action is selected, the simulation environment executes the action. The environment first determines whether this recovery attempt is successful, based on whether the inherent success rate of the selected aircraft is higher than the random threshold generated at the beginning of the "round". If successful, the state of the aircraft is marked as "successful"; if failed, it is marked as "just reflighted", and its inherent success rate is increased by a certain percentage, simulating its accumulation of experience after a failure. Regardless of success or failure, the environment advances the simulation time according to the physical rules (such as the wake separation time, the time required for recovery or reflight), and accordingly deducts the fuel of all aircraft still in the air. Subsequently, the environment calculates a scalar reward value according to the reward function, forming a comprehensive feedback signal.
[0077] The complete experience generated by the decision-making process, i.e., (current state sequence, executed action, obtained reward, next state sequence, task completion flag), is stored in the priority experience replay memory as a five-tuple and is assigned the highest initial priority. This interaction process continues until all aircraft in the environment are successfully recovered, and the data collection of a "round" is completed.
[0078] When the data in the experience replay memory accumulates to a certain amount, the update step of the network parameters is started. The system will sample a batch of data according to the priority of each sample in the memory. The higher the priority of the sample (i.e. the less accurate the previous prediction of the model), the greater the probability of being sampled. For each sample, the system calculates its time difference target using the target network, and then calculates the current Q value using the policy network. The difference between the two is the time difference error. This time difference error is used to update the priority of the sample in the memory, and combined with the importance sampling weight obtained from the sampling process, it forms a weighted loss function. The system updates the parameters of the policy network by minimizing this loss function through gradient descent method. In order to maintain stable training, the parameters of the policy network will be completely copied to the target network every fixed update period. During the entire training process, the performance indicators of the system such as loss value, average reward, etc. are continuously recorded and monitored, and the model parameters with the best performance are periodically saved. Through tens of thousands of iterations or more, the deep neural network model for efficient, safe and intelligent aircraft recovery dynamic scheduling, i.e. the aircraft recovery dynamic scheduling model, is finally trained.
[0079] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present disclosure.
[0080] The principles and implementation manners of the present application are described by using specific examples in the present disclosure, and the above embodiments are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges can be changed. In summary, the content of the present disclosure should not be understood as a limitation of the present application.
Claims
1. An aircraft recovery dynamic scheduling method based on a long short-term memory network, characterized in that, The method comprises the following steps: initializing a policy network, a target network and a priority experience replay memory, and setting up a plurality of parallel aircraft recovery simulation environments; the policy network and the target network are both deep neural networks comprising a long short-term memory network layer and a multi-layer perceptron layer; the priority experience replay memory loads initial data comprising a plurality of aircrafts to be recovered; the initial data at least comprises an aircraft model matrix and a feature matrix of initial attributes of each aircraft; in each parallel aircraft recovery simulation environment, generating a recovery trajectory through a loop policy process, and recovering the aircraft according to the recovery trajectory in each decision step, and storing it into the priority experience replay memory until the recovery task in all parallel aircraft recovery simulation environments is completed; When the amount of samples in the priority experience replay memory meets a preset batch number, a batch of experience data is sampled from the priority experience replay memory, and the Q value of the next state of any experience sample in the experience data is calculated by using the target network to update the parameters of the policy network ; parameters of the policy network synchronously update parameters of the target network record network update parameters after each update, and save parameters of the policy network according to a preset saving period The network update parameters include a weighted loss function, an average recovery reward, and an exploration rate. based on the network update parameters, determining an aircraft recovery dynamic scheduling model according to the updated target network and the updated policy network to perform the recovery task.
2. The aircraft recovery dynamic scheduling method based on long short-term memory network according to claim 1, characterized in that, in each decision step, recovering the aircraft according to the recovery trajectory, and storing it into the priority experience replay memory until the recovery task in all parallel aircraft recovery simulation environments is completed, which specifically comprises: obtaining a current state sequence of a current decision step; the current state sequence comprises aircraft state tensors of a plurality of time steps; processing the current state sequence by using the policy network to output Q values corresponding to all selectable recovery actions; By - a greedy strategy selects one recovery action from the set of optional recovery actions; the recovery action is to select one of a plurality of aircrafts to perform a recovery attempt; executing the recovery action in the simulation environment to determine a next state sequence, a scalar reward and a Boolean flag indicating whether the current recovery task is completed, and determining a five-tuple experience; the five-tuple experience comprises the current state sequence, the recovery action, the next state sequence, the scalar reward and the Boolean flag indicating whether the current recovery task is completed; storing the five-tuple experience in the priority experience replay memory with the highest priority until the recovery task in all parallel aircraft recovery simulation environments is completed.
3. The aircraft recovery dynamic scheduling method based on long short-term memory network according to claim 1, wherein, updating parameters of the policy network using the Q value of the next state of any experience sample in the experience data calculated by the target network , and specifically comprises: calculating a sampling probability of each experience sample according to the priority stored in the priority experience replay memory, and extracting batch data comprising a plurality of experience samples, while obtaining an index and an importance sampling weight corresponding to a batch sample in the batch data; the batch data is experience data of the same batch; calculating a Q value of a next state of each experience sample in the batch data by using the target network, and constructing a time difference target combining the reward and a discount factor; calculating a Q value of a current state by using the policy network, and calculating a time difference error of each experience sample according to the Q value of the current state and the time difference target; updating the sample priority of the corresponding index in the priority experience replay memory according to the time difference error. Based on the updated sample priority, a weighted loss function is calculated according to the time difference error and the importance sampling weight, and the parameters of the policy network are updated by gradient descent method .
4. The aircraft recovery dynamic scheduling method based on long short-term memory network according to claim 3, characterized in that, The sampling probability is: wherein is the priority of the experience sample ; is the priority index; is the priority of the experience sample ; k is the index of all experience samples in the experience replay pool.
5. The aircraft recovery dynamic scheduling method based on long short-term memory network according to claim 4, characterized in that, The time-difference-of-arrival target Is: wherein, is a scalar reward; is a discount factor; is a next state sequence; is a recycled action for the next state; are parameters of the policy network; are parameters of the target network.
6. The aircraft recovery dynamic scheduling method based on long short-term memory network according to claim 5, characterized in that, The weighted loss function is: wherein, B is a set of empirical data collected from a batch of samples; , is an importance sampling weight for an empirical sample , is a prioritized experience replay memory is a current total number of samples in the prioritized experience replay memory is an importance sampling exponent; is a time-difference error , is a current state sequence, is a recycled action for the current state.
7. The aircraft recovery dynamic scheduling method based on long short-term memory network according to claim 2, characterized in that, The current state sequence is a three-dimensional tensor, and the three-dimensional tensor is ; wherein, is a historical sequence length; is a total number of aircrafts to be recovered; is a feature dimension of a single aircraft; the feature vector of the single aircraft at least includes a normalized fuel quantity, an inherent recovery success rate, and a discrete code representing a current recovery state. the state corresponding to the discrete coding comprises an untried recovery state, the discrete coding is 0; a successfully recovered state, the discrete coding is 1; and a just-executed reflight and not-yet-recovered state, the discrete coding is 2.
8. The aircraft recovery dynamic scheduling method based on long short-term memory network according to claim 2, characterized in that, processing the current state sequence by using the policy network to output Q values corresponding to all selectable recovery actions, which specifically comprises: in each decision step, concatenating feature vectors of all aircrafts in the current state sequence to generate an input sequence tensor; inputting the input sequence tensor into a long short-term memory network layer of the policy network to extract time sequence features and obtain a hidden state of a last time step; inputting the hidden state into a multi-layer perception layer of the policy network to determine Q values corresponding to all selectable recovery actions under a current state sequence.
9. The aircraft recovery dynamic scheduling method based on long short-term memory network according to claim 2, wherein, The scalar reward is generated by a composite reward function. The composite reward function includes a fuel safety penalty term, an efficiency penalty term, a priority waiting penalty term and a task completion reward term; wherein when the fuel quantity of any unrecovered aircraft is lower than a preset safety threshold, the fuel safety penalty term is a negative reward proportional to the fuel loss quantity.
10. The aircraft recovery dynamic scheduling method based on long short-term memory network according to claim 2, wherein, The recovery action is executed in a simulation environment to determine a next state sequence, a scalar reward and a Boolean flag indicating whether the current recovery task is completed, and to determine a five-tuple experience, and the method further includes: generating a random recovery success rate threshold matrix for multiple aircraft in the simulation environment for the current aircraft at the beginning of each recovery trajectory; after executing the recovery action at each decision step, determining whether the inherent recovery success rate of the selected aircraft is greater than or equal to the corresponding recovery success rate threshold of the inherent recovery success rate in the recovery success rate threshold matrix; if yes, determining that the selected aircraft recovers successfully; if no, determining that the selected aircraft recovers unsuccessfully and performing a re-fly; after the selected aircraft performs the re-fly, increasing the inherent recovery success rate of the selected aircraft by a preset proportion in subsequent decision steps.
Citation Information
Patent Citations
Unmanned aerial vehicle cluster evolution type simulation training method, system, medium and equipment
CN118171572A
Airport safety management method and system
CN119761832A