Aircraft recovery scheduling sequence generation method and device for aircraft recovery reinforcement learning, and medium
By defining a composite reward function and constructing an optimal recovery scheduling strategy model, the problem of a single reward function in the reinforcement learning method for aircraft recovery scheduling is solved, a balance is achieved among safety, efficiency and task priority, and a comprehensive optimal aircraft recovery scheduling sequence is generated.
Patent Information
- Application Number
- CN202511204854.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing reinforcement learning methods have a single reward function design in aircraft recovery scheduling, which makes it difficult to effectively balance fuel safety, recovery efficiency and task priority, resulting in the scheduling strategy being unable to meet the comprehensive requirements in practical applications.
A composite reward function is defined, including fuel safety penalty, efficiency penalty, priority waiting penalty, and task completion reward. This function is embedded in the reinforcement learning training process to construct an optimal recovery scheduling strategy model and generate an aircraft recovery scheduling sequence that takes multiple objectives into account.
It achieves a balance between safety, efficiency and task priority during the aircraft recovery process, generates a comprehensive optimal scheduling sequence, solves the limitations of traditional scheduling methods under complex dynamic constraints, and realizes the intelligence and optimization of aircraft recovery decisions.
Smart Images

Figure CN120706844A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the intersection of computer technology and aerospace technology, and in particular to a method, device, and medium for generating an aircraft recovery scheduling sequence for aircraft recovery reinforcement learning. Background Art
[0002] Aircraft recovery, especially approach scheduling at busy airports, is an extremely complex dynamic scheduling problem. Dispatchers must make the optimal recovery order decision within a very short timeframe, comprehensively considering a variety of dynamic factors, including the real-time fuel level, airframe health, mission priority, and recovery channel availability of each aircraft in the fleet. This process requires not only extremely high decision-making efficiency but also stringent safety requirements that cannot tolerate errors. Traditional aircraft recovery relies primarily on manual scheduling, with dispatchers making judgments based on extensive experience and rules. However, when faced with large-scale, high-intensity recovery missions, the cognitive load on human operators is enormous, making decision quality difficult to ensure and prone to oversight, posing a potential threat to flight safety. With technological advancements, several decision-making support methods based on traditional operations research or expert systems have been proposed. However, these methods often rely on rigid models and are difficult to adapt to the highly dynamic and highly uncertain civil aviation environment.
[0003] In recent years, reinforcement learning, an artificial intelligence technology that can autonomously learn optimal strategies through interaction with the environment, has provided new insights into solving such complex decision-making problems. In reinforcement learning, the design of the reward function is central to guiding the agent's learning direction and directly determines the quality of the final strategy. Existing reinforcement learning methods applied to scheduling problems often have relatively simple reward function designs or struggle to effectively balance multiple conflicting optimization objectives. For example, overemphasizing fuel safety may lead to low recovery efficiency, while a one-sided pursuit of efficiency may ignore the needs of high-priority tasks, resulting in the resulting scheduling strategy failing to meet the comprehensive requirements of practical applications. Therefore, how to design a reward function that can comprehensively and accurately reflect the complex requirements of aircraft recovery tasks and, based on this, generate a reliable optimal aircraft recovery scheduling sequence is a difficult problem that needs to be solved urgently in the current technical field. Summary of the Invention
[0004] The purpose of this application is to provide an aircraft recovery scheduling sequence generation method, device and medium for aircraft recovery reinforcement learning to solve the problems of a single reward function and the inability of the scheduling strategy to meet the comprehensive requirements in practical applications.
[0005] To achieve the above objectives, this application provides the following solutions.
[0006] In a first aspect, the present application provides an aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning, including the following method.
[0007] A composite reward function is defined in the aircraft recovery process; the composite reward function includes a fuel safety penalty item, an efficiency penalty item, a priority waiting penalty item, and a task completion reward item.
[0008] Based on the composite reward function, a scalar reward is provided to the decision-making agent in each decision step of reinforcement learning; the decision-making agent is used to perform decision-making actions in an environment simulating the dynamic changes of a fleet of aircraft to be recovered; the scalar reward is a reward function used for aircraft recovery reinforcement learning; the reward function is used to guide the decision-making agent to learn a recovery strategy that takes into account fuel safety, recovery efficiency, and task priority.
[0009] The calculation process of the reward function is embedded in the reinforcement learning training process to build an optimal recycling scheduling strategy model.
[0010] An aircraft recovery scheduling sequence is generated according to the optimal recovery scheduling strategy model.
[0011] All aircraft to be recovered are recovered based on the aircraft recovery scheduling sequence.
[0012] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning.
[0013] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning.
[0014] According to the specific embodiments provided in this application, this application has the following technical effects.
[0015] This application defines a composite reward function to provide a scalar reward to a reinforcement learning agent at each decision step. This composite reward function, composed of a fuel safety penalty, an efficiency penalty, a priority waiting penalty, and a task completion reward, comprehensively quantifies the performance of multiple optimization objectives, including safety, efficiency, and priority, during aircraft recovery. The fuel safety penalty penalizes low-fuel states, the efficiency penalty encourages a shorter total recovery time, and the priority waiting penalty prioritizes high-priority or faulty aircraft, effectively balancing multiple conflicting optimization objectives.
[0016] In addition, this application obtains the optimal recovery strategy model by training the decision-making intelligent agent in a simulation environment, and then uses the model to make sequential decisions in actual tasks to generate a comprehensive optimal aircraft recovery scheduling sequence that takes into account multiple objectives, so that the scheduling strategy cannot meet the comprehensive requirements in actual applications, solves the limitations of traditional scheduling methods in dealing with complex dynamic constraints, and realizes the intelligence and optimization of aircraft recovery decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A flowchart of a method for generating an aircraft recovery scheduling sequence for aircraft recovery reinforcement learning provided in one embodiment of the present application. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0020] In order to make the purpose, features and advantages of this application more obvious and easy to understand, this application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0021] like Figure 1 As shown, an embodiment of the present application provides an aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning, including the following steps.
[0022] S1: Define the composite reward function in the aircraft recovery process; the composite reward function includes fuel safety penalty, efficiency penalty, priority waiting penalty and task completion reward. , efficiency penalty , and priority waiting penalty Are all non-positive values, the task completion reward items Is a non-negative value.
[0023] S2: Based on the composite reward function, a scalar reward is provided to the decision-making agent in each decision step of reinforcement learning; the decision-making agent is used to perform decision-making actions in an environment simulating the dynamic changes of a fleet of aircraft to be recovered; the scalar reward is a reward function for aircraft recovery reinforcement learning; the reward function is used to guide the decision-making agent to learn a recovery strategy that takes into account fuel safety, recovery efficiency, and task priority.
[0024] Scalar rewards The calculation formula is as follows.
[0025]
[0026] S3: Embed the calculation process of the reward function into the reinforcement learning training process to build an optimal recycling scheduling strategy model.
[0027] S4: Generate an aircraft recovery scheduling sequence according to the optimal recovery scheduling strategy model.
[0028] S5: Recover all aircraft to be recovered based on the aircraft recovery scheduling sequence.
[0029] In an exemplary embodiment, the following steps are further included before S1.
[0030] Calculating the fuel safety penalty includes the following steps.
[0031] S11: Obtain a queue of all aircraft to be recovered that have not been successfully recovered in the decision step.
[0032] S12: Obtain the current fuel quantity of each aircraft to be recovered in the aircraft to be recovered queue in the decision step.
[0033] S13: Calculating a fuel safety penalty for a single aircraft to be recovered based on the current fuel quantity and the minimum fuel safety threshold.
[0034] S14: Determine the fuel safety penalty item according to the fuel safety penalties of all aircraft to be recovered.
[0035] The fuel safety penalty item establishes a safety bottom line by monitoring the fuel amount of all non-recovered aircraft, imposing a penalty proportional to the fuel loss if the amount falls below the safety threshold, and applying strong negative feedback in extreme cases of fuel depletion.
[0036] In an exemplary embodiment, S13 specifically includes the following steps.
[0037] like ,Sure is 0; among them, For the i The current fuel level of the aircraft to be recovered; is the minimum fuel safety threshold; Fuel safety penalty for a single aircraft to be recovered.
[0038] like ,Sure A huge penalty value for setting up fuel exhaustion Negative value of .
[0039] like ,Sure Equal to the fuel loss amount and the fuel safety penalty coefficient The negative value of the product of all fuel losses is the minimum fuel safety threshold With the i Current fuel level of the aircraft to be recovered The difference.
[0040] Huge penalty for running out of fuel The value range is [500,2000].
[0041] Minimum fuel safety threshold It is [1.1, 1.5] times the total amount of fuel required to complete a standard recovery approach and a missed approach, as preset based on the aircraft model.
[0042] In an exemplary embodiment, the following steps are further included before S1.
[0043] Calculate the efficiency penalty term; the efficiency penalty term as follows.
[0044]
[0045] in, From the previous decision step End to current decision step The length of time it takes to end, is the preset efficiency penalty coefficient.
[0046] The efficiency penalty term is directly related to the time consumption between decision steps, and applies continuous negative feedback to the passage of time to motivate the agent to complete all recycling tasks as quickly as possible.
[0047] In an exemplary embodiment, the following steps are further included before S1.
[0048] Calculate the priority waiting penalty item; the priority waiting penalty item as follows.
[0049]
[0050] in, Aircraft queue for recovery Aircraft index in; For aircraft the relative completeness of For aircraft Relative task priorities; Waiting penalty coefficient for preset priority.
[0051] Fuel safety penalty factor , efficiency penalty coefficient , and the priority waiting penalty coefficient The value range of is [0.1,5.0].
[0052] In practical applications, the relative completeness and relative task priorities Based on the aircraft to be recovered The dimensionless values obtained after normalization of the original fault level and mission level data in the interval [0,1] are used to quantify the scheduling cost of making aircraft with higher fault levels or higher mission levels continue to wait.
[0053] The priority waiting penalty item comprehensively considers the integrity and mission importance of each aircraft to be recovered, and imposes penalties on behaviors that keep high-priority aircraft waiting, ensuring that mission-critical and high-risk aircraft are given priority.
[0054] In an exemplary embodiment, the following steps are further included before S1.
[0055] Completion reward for calculation tasks , specifically including the following steps.
[0056] Determine the queue of aircraft to be recovered Is it an empty set? If so, determine that all aircraft to be recovered have been successfully recovered, and use a fixed positive task completion reward value as the task completion reward item; if not, determine that the task completion reward item is 0.
[0057] Task completion reward value The value range is [100,500].
[0058] The mission completion reward item gives a significant positive reward at the final moment when all aircraft are successfully recovered, providing a clear and ultimate convergence goal for the agent's learning process.
[0059] In practical applications, the preset parameters involved in this application have the following value ranges.
[0060] In an exemplary embodiment, S3 specifically includes the following steps.
[0061] S31: Embed the calculation process of the reward function into the reinforcement learning training process.
[0062] S32: During the reinforcement learning training process, the decision-making agent's internal decision parameters are updated based on the scalar reward, optimized policy function, or value function obtained after the decision-making agent executes each decision action, until the decision-making agent's recovery strategy converges, thereby constructing an optimal recovery scheduling strategy model. The decision action is to select an aircraft for recovery from the queue of aircraft currently waiting to be recovered.
[0063] In an exemplary embodiment, S4 specifically includes the following steps.
[0064] S41: Input the initial state of the queue of aircraft to be recovered into the optimal recovery scheduling strategy model, and sequentially execute the optimal recovery scheduling strategy model, setting it to output a decision on a aircraft to be recovered at each decision step until all aircraft to be recovered are scheduled.
[0065] S42: Arrange the aircraft to be recovered outputted from all decision steps in the order of the corresponding decisions, and generate an aircraft recovery scheduling sequence that achieves the best overall performance among fuel safety, recovery efficiency, and task priority.
[0066] In practical applications, this application uses the above-mentioned reward function design method to train a decision-making agent in a high-fidelity simulation environment. Through continuous trial and error, the decision-making agent receives the composite reward function designed by this application after each decision, and adjusts its internal decision network parameters accordingly, and finally learns an optimal recovery strategy model that can maximize long-term cumulative rewards. Subsequently, in practical applications, the initial state of the fleet of aircraft to be recovered is input into the trained optimal recovery strategy model. The model will make decisions sequentially and step by step based on the current state, and each decision will select an aircraft that should be recovered. Arranging all the decisions made by the model in chronological order constitutes a comprehensive optimal aircraft recovery scheduling sequence, which realizes the automation and intelligence of the entire scheduling process.
[0067] Through the design of a refined multi-objective reward function, this application successfully transforms complex scheduling constraints into numerical signals that can be understood and optimized by decision-making agents. The resulting scheduling strategy achieves a good balance between safety, efficiency and task priority, significantly outperforming traditional methods.
[0068] Taking a typical reinforcement learning framework as an example, the technical solution of this application is further explained.
[0069] In a typical reinforcement learning framework, a decision-making agent interacts with the environment at each decision step. The environment then responds with a reward signal, or reward function, based on the agent's actions. The core of this application lies in the design of this reward signal.
[0070] At each decision step, this application calculates a composite scalar reward value, namely the composite reward function, which is the sum of four components.
[0071] The first component is the fuel safety penalty. To calculate it, the system first determines the set of all aircraft currently in the air awaiting recovery, known as the recovery queue. For each aircraft in this queue, the system obtains its current fuel level and compares it to a preset minimum fuel safety threshold. This threshold is typically set based on the aircraft model and standard approach / go-around fuel consumption, with a safety margin. If the aircraft has sufficient fuel above this threshold, no penalty is incurred. If the fuel level is depleted, a large fixed penalty is imposed to strongly discourage the agent from learning a strategy that would lead to such a catastrophic outcome. If the fuel level is between the minimum safety threshold and zero, the penalty is proportional to the amount of fuel lost, with the greater the loss, the greater the penalty. Finally, the fuel penalties for all recovery aircraft are summed to obtain the total fuel safety penalty for that decision step.
[0072] The second component is the efficiency penalty term. Its purpose is to expedite the recycling process. Its calculation is straightforward: a preset efficiency penalty factor is multiplied by the time elapsed between the previous and current decision steps, with the negative value taken. This means that for every unit of time that passes, the agent incurs a fixed penalty, incentivizing it to take actions that shorten the total time.
[0073] The third component is the priority waiting penalty. This item is designed to ensure that aircraft that are in danger or carrying important missions can be recovered first. For each aircraft that has not yet been recovered, this application first obtains its two key attributes: relative completeness and relative mission priority. These two attributes are usually obtained by normalizing the original fault level and mission level data of the aircraft. The values are between zero and one. The larger the value, the worse the aircraft condition or the more important the mission. The penalty item is calculated by adding the sum of the relative completeness and relative mission priority of all aircraft to be recovered, and then multiplying it by a preset priority waiting penalty coefficient and taking the negative value. In this way, as long as there are high-priority or high-risk aircraft waiting in the queue, a large penalty will continue to be generated, forcing the agent to give priority to them.
[0074] The fourth component is the mission completion reward. This is a sparse but crucial positive reward. At each decision step, the system checks whether the set of aircraft to be recovered is empty. In the vast majority of cases, the set is not empty, and the reward is zero. Only when the last aircraft is successfully recovered and the set becomes empty does the system grant a fixed, large positive reward. This final reward provides a clear endpoint and goal for the agent's learning.
[0075] This application provides an aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning, which is a direct application of the above-mentioned reward function design.
[0076] The aircraft recovery scheduling sequence generation process includes two stages: training and generation. In the training stage, the calculation process of the above-mentioned reward function is embedded in a reinforcement learning training process. A decision-making agent, such as a deep neural network, is trained in a virtual environment that can simulate aircraft flight, fuel consumption and recovery processes. The task of the decision-making agent is to select one aircraft to be recovered from all aircraft to be recovered at each decision moment. After making a choice, the environment will update its state according to its action, and use this application to calculate the compound reward and feedback it to the decision-making agent. Based on this reward signal, the decision-making agent updates its network parameters through algorithms such as policy gradient or value iteration. Its goal is to learn a strategy that can maximize the total reward of the entire recovery task. This training process will be carried out thousands of times until the decision-making agent's strategy converges and stabilizes, forming an optimal recovery scheduling strategy model.
[0077] During the generation phase, when faced with a real recovery task, simply input the initial state information of the fleet of aircraft to be recovered into the trained optimal recovery scheduling strategy model. The model will immediately output the first decision, which is the currently optimal recovery aircraft. After the first aircraft is dispatched, the system state is updated, and the model will then make the second decision, and so on and so forth until all aircraft are dispatched. Arranging these sequentially output decisions together forms a complete, optimized scheduling sequence that achieves a balance between multiple objectives. In this way, this application transforms complex manual scheduling tasks into an automated, data-driven, intelligent decision-making process.
[0078] In an exemplary embodiment, a computer device is provided, comprising a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of each of the above-described method embodiments are implemented. The computer device may be a server or a terminal. The computer device includes a processor, memory, an input / output (I / O) interface, and a communication interface. The processor, memory, and I / O interface are connected via a system bus, and the communication interface is connected to the system bus via the I / O interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operating system and computer program in the non-volatile storage medium. The database of the computer device is configured to store data to be processed. The I / O interface of the computer device is configured to exchange information between the processor and an external device. The communication interface of the computer device is configured to communicate with an external terminal via a network connection. When executed by the processor, the computer program implements a method for generating an aircraft recovery scheduling sequence for aircraft recovery reinforcement learning.
[0079] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0080] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0081] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0082] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by hardware associated with computer program instructions. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to a memory, database, or other medium used in the embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0083] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.
[0084] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0085] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for generating an aircraft recovery scheduling sequence for aircraft recovery reinforcement learning, characterized in that: include: Defining a composite reward function during the aircraft recovery process; the composite reward function includes a fuel safety penalty term, an efficiency penalty term, a priority waiting penalty term, and a task completion reward term; Based on the composite reward function, a scalar reward is provided to the decision-making agent at each decision step of the reinforcement learning; the decision-making agent is used to perform decision actions in an environment simulating a dynamically changing fleet of recoverable aircraft; the scalar reward serves as a reward function for aircraft recovery reinforcement learning; and the reward function is used to guide the decision-making agent to learn a recovery strategy that balances fuel safety, recovery efficiency, and task priority. The calculation process of the reward function is embedded in the reinforcement learning training process to build an optimal recycling scheduling strategy model; generating an aircraft recovery scheduling sequence according to the optimal recovery scheduling strategy model; All aircraft to be recovered are recovered based on the aircraft recovery scheduling sequence.
2. The aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning according to claim 1 is characterized in that: Define the compound reward function during the aircraft recovery process, which previously included: Calculate fuel safety penalties, including: Obtaining a queue of all aircraft to be recovered that have not been successfully recovered in the decision step; Obtaining the current fuel quantity of each aircraft to be recovered in the aircraft to be recovered queue in the decision step; Calculating a fuel safety penalty for a single aircraft to be recovered based on the current fuel quantity and the minimum fuel safety threshold; The fuel safety penalty item is determined based on the fuel safety penalties of all aircraft to be recovered.
3. The aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning according to claim 2 is characterized in that: The fuel safety penalty for a single aircraft to be recovered is calculated based on the current fuel quantity and the minimum fuel safety threshold, specifically including: like ,Sure is 0; among them, For the i The current fuel level of the aircraft to be recovered; is the minimum fuel safety threshold; A single aircraft to be recovered i fuel safety penalties; like ,Sure A huge penalty value for setting up fuel exhaustion Negative value of like ,Sure Equal to the fuel loss amount and the fuel safety penalty coefficient The negative value of the product of all fuel losses is the minimum fuel safety threshold With the i Current fuel level of the aircraft to be recovered The difference.
4. The aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning according to claim 1 is characterized in that: Define the compound reward function during the aircraft recovery process, which previously included: Calculate the efficiency penalty term; the efficiency penalty term for: ; in, From the previous decision step End to current decision step The length of time it takes to end, is the preset efficiency penalty coefficient.
5. The aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning according to claim 2 is characterized in that: Define the compound reward function during the aircraft recovery process, which previously included: Calculate the priority waiting penalty item; the priority waiting penalty item for: ; in, Aircraft queue for recovery Aircraft index in; For aircraft the relative completeness of For aircraft Relative task priorities; Waiting penalty coefficient for preset priority.
6. The aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning according to claim 2 is characterized in that: Define the compound reward function during the aircraft recovery process, which previously included: Completion reward for calculation tasks , specifically including: Determine the queue of aircraft to be recovered Is it an empty set? If so, it is determined that all aircraft to be recovered have been successfully recovered, and the positive task completion reward value is used as the task completion reward item; If not, determine that the task completion reward item is 0.
7. The aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning according to claim 1 is characterized in that: The reward function calculation process is embedded into the reinforcement learning training process to build an optimal recycling scheduling strategy model, which specifically includes: Embedding the calculation process of the reward function into the reinforcement learning training process; In the reinforcement learning training process, the internal decision parameters of the decision agent are updated according to the scalar reward, optimization strategy function or value function obtained after the decision agent executes each decision action until the recycling strategy of the decision agent converges, thereby constructing an optimal recycling scheduling strategy model.
8. The aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning according to claim 1 is characterized in that: Generating an aircraft recovery scheduling sequence according to the optimal recovery scheduling strategy model specifically includes: Inputting the initial state of the queue of recoverable aircraft into the optimal recovery scheduling strategy model, and sequentially executing the optimal recovery scheduling strategy model, setting it to output a decision on a recoverable aircraft at each decision step until all recoverable aircraft are scheduled; The aircraft to be recovered outputted from all decision steps are arranged in the order of the corresponding decisions to generate an aircraft recovery scheduling sequence that achieves the optimal combination of fuel safety, recovery efficiency and task priority.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aircraft recovery scheduling sequence generation method for aircraft recovery reinforcement learning according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating an aircraft recovery scheduling sequence for aircraft recovery reinforcement learning according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Intelligent transportation methods and systems
CA3238745A1
Flight dispatching task planning method, device, equipment and medium
CN113743666A
Multi-park energy scheduling method and system based on deep reinforcement learning
CN114091879A
Auxiliary method for flight maneuvering decision based on reinforcement learning
CN114237267A
Aircraft warehouse-in and warehouse-out transfer scheduling method, device and equipment based on grey wolf algorithm
CN116011724A