Aircraft dynamic recovery sorting method and equipment based on reinforcement learning

By building a dynamic recovery environment model and a multi-dimensional reward function, combined with a reinforcement learning agent, the aircraft recovery sequence is optimized, which solves the coordination issues of fuel safety, task priority and fault handling in aircraft recovery, and improves recovery efficiency and safety.

CN120706845AActive Publication Date: 2025-09-26NAVAL AVIATION UNIV
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202511204863.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-09-26
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively balance fuel safety, mission priority, and fault handling in aircraft recovery sorting. Traditional methods have low learning efficiency, a single reward function design, and a lack of comprehensive consideration of multi-dimensional objectives.

Method used

A dynamic recovery environment model including go-around scenarios is constructed, and a multidimensional reward function and safety cost function are designed. The scheduling strategy is optimized through reinforcement learning agents to output the optimal aircraft recovery sorting sequence, forming a closed-loop learning mechanism to achieve coordination among fuel economy, mission priority, and safety risks.

Benefits of technology

It significantly improves the efficiency and safety of aircraft recovery and achieves coordinated optimization of fuel consumption minimization, mission priority execution and emergency fault handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706845A_ABST
    Figure CN120706845A_ABST
Patent Text Reader

Abstract

The invention discloses an aircraft dynamic recovery sorting method and device based on reinforcement learning, and relates to the technical field of aviation dispatch, and the method comprises the steps: building a dynamic recovery environment model containing a re-flight scene, designing a multi-dimensional reward function and a safety cost function, and achieving the intelligent decision of a reinforcement learning agent for aircraft recovery dispatch; the multi-dimensional reward function integrates four core indexes, namely fuel consumption punishment, fault priority reward, task priority reward and re-flight punishment; the safety cost function quantifies the fuel safety risk of the unlanded aircraft; the reinforcement learning agent optimizes the scheduling strategy under the dynamic balance constraint of maximizing the reward function and minimizing the security cost. According to the method, closed-loop optimization of the aircraft recovery process is realized by periodically updating the environment state and promoting the decision-making stage, the coordination problem among fuel economy, task priority and safety risk is effectively solved, and the aircraft recovery efficiency and safety are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of aviation scheduling technology, and in particular to a method and device for dynamic aircraft recovery sorting based on reinforcement learning. Background Art

[0002] In the field of aviation scheduling, the Aircraft Recovery and Scheduling Problem (ARSP) is a complex dynamic decision-making problem. Traditional methods, primarily based on heuristic rules or mathematical programming, struggle to cope with dynamically changing environments (such as sudden failures and weather changes). Reinforcement learning has been introduced to this field in recent years, but its performance is highly dependent on the design of the reward function. Existing reward function designs are often limited to a single objective (such as minimizing total time) and lack comprehensive consideration of fuel safety, mission priorities, and fault handling. In aircraft recovery scenarios, in particular, multi-dimensional objectives such as go-around risk, fuel constraints, and mission urgency interact with each other. The scheduling policy in the ARSP needs to be translated into an executable recovery sequence, but existing technologies struggle to achieve this balance effectively. Furthermore, traditional methods often employ binary penalty mechanisms to address safety constraints, resulting in low learning efficiency. Therefore, a reinforcement learning reward function design method that can comprehensively quantify multi-dimensional objectives and dynamically adjust policies is urgently needed. Summary of the Invention

[0003] The purpose of this application is to provide a reinforcement learning-based aircraft dynamic recovery sorting method and equipment, which achieves closed-loop optimization of the aircraft recovery process by periodically updating the environmental status and advancing the decision-making stage, effectively solving the coordination problem between fuel economy, mission priority and safety risks, and significantly improving the efficiency and safety of aircraft recovery.

[0004] To achieve the above objectives, this application provides the following solutions: In a first aspect, the present application provides a method for dynamic aircraft recovery sorting based on reinforcement learning, comprising: Build a dynamic recovery environment model including missed approach scenarios; Based on the dynamic recovery environment model, a multidimensional reward function is constructed; the multidimensional reward function is used to describe the fuel consumption penalty, fault priority reward, mission priority reward and go-around penalty in a go-around scenario; Based on the dynamic recovery environment model and the multi-dimensional reward function, a safety cost function is constructed; the safety cost function is used to quantify the total fuel safety risk of all aircraft that have not landed in the aircraft recovery sorting sequence; Based on the multidimensional reward function and the safety cost function, optimizing the scheduling strategy through a reinforcement learning agent to obtain an optimal aircraft recovery sorting sequence; The optimal aircraft recovery sorting sequence is converted into executable instructions and transmitted to the airport dispatch system.

[0005] Optionally, the multidimensional reward function is: ; in, ; ; Where, r is a multidimensional reward function; is the maximum number of decision stages; For the stage The reward function of is the adjustable reward weight coefficient; It is the fuel consumption penalty item; is the rectified linear activation function; For aircraft In stage The landing status indicator variable, Indicates a successful landing; Representation sequence The index of the aircraft waiting to land in the middle stage τ; Priority reward item for failure; It is a task priority reward item; This is the penalty item for missed approach.

[0006] Optionally, the fuel consumption penalty item is: ; Where, Gathering for aircraft to be recovered, Indicates the cardinality of a set; Representation stage time; Representation stage time; Indicates the aircraft to be landed i is the aircraft type Fuel consumption rate at .

[0007] Optionally, the fault priority reward item is: ; Where, Indicates a plane waiting to land Failure probability assessment value; Indicates a plane waiting to land The failure probability assessment value.

[0008] Optionally, the task priority reward item is: ; Where, Indicates a plane waiting to land The mission urgency ranking of the aircraft to be landed is the highest; the mission urgency ranking of the aircraft to be landed is 1.

[0009] Optionally, the go-around penalty term is: ; in, ; Where, Delay phase due to missed approach The time decay factor of is the penalty intensity coefficient; Aircraft selected for the current decision scheduled recycling phase.

[0010] Optionally, the security cost function is: ; in, ; Where, is the security cost function; For the stage Single-stage security cost; For the stage Assembly of aircraft that have not landed; Indicates the safe fuel level threshold of the model; Indicates aircraft exist Stage remaining fuel; in aircraft exist Remaining oil level at this stage Less than the safe fuel level threshold of the model Time, stage Single-stage security cost Proportional to the amount of fuel shortage; is the safety cost conversion coefficient.

[0011] Optionally, based on the multidimensional reward function and the safety cost function, optimizing the scheduling strategy through a reinforcement learning agent to obtain an optimal aircraft recovery ranking sequence specifically includes: By optimizing the scheduling strategy through reinforcement learning agents, a dynamic balance is achieved between maximizing the multi-dimensional reward function and minimizing the safety cost function, and the optimal sequence is output; Determine the multidimensional reward function value corresponding to the optimal sequence as the actual cumulative reward; Determine the security cost function value corresponding to the optimal sequence as the actual security cost; Determine the ratio of actual cumulative rewards to actual security costs as the effectiveness factor; Determine whether the validity factor is lower than a preset validity threshold, and obtain a determination result; If the judgment result is yes, then adjust the adjustable reward weight coefficient in the multidimensional reward function, and return to the step of "optimizing the scheduling strategy through the reinforcement learning agent to achieve a dynamic balance between maximizing the multidimensional reward function and minimizing the safety cost function, and outputting the optimal sequence"; If the judgment result is no, then the optimal sequence is determined to be the optimal aircraft recovery sorting sequence.

[0012] Optionally, after converting the optimal aircraft recovery sorting sequence into executable instructions and transmitting the instructions to the airport dispatch system, the method further includes: executing the executable instructions and recording fuel consumption, risk indicators, and go-around events; Use the fuel status dynamic update formula to update the fuel status and return to step "Build a dynamic recovery environment model including the missed approach scenario"; The fuel status dynamic update formula is: ; Where, Indicates aircraft exist Remaining oil level at the stage; Representation stage With stage The time increment between is the indicator function, when exist Indicator function when the stage has not landed The value is 1.

[0013] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned reinforcement learning-based aircraft dynamic recovery sorting method.

[0014] According to the specific embodiments provided in this application, this application discloses the following technical effects: This application provides a method and device for dynamic recovery sorting of aircraft based on reinforcement learning, constructs a dynamic environment model S including a go-around scenario, and defines a set of execution phases of the sorting sequence. ; Design a multidimensional cumulative reward function r and security cost function for the overall sorting sequence ; Output the optimal recovery sorting sequence A* through the reinforcement learning agent; establish a sorting effectiveness evaluation mechanism and a dynamic adjustment strategy to form a closed-loop learning mechanism of "environmental state perception-action decision-reward feedback-safety assessment-strategy update", and ultimately achieve global optimization of recovery efficiency and safety; This application outputs the optimal aircraft recovery sorting sequence through the reinforcement learning agent, and realizes the coordinated optimization of minimizing fuel consumption, task priority execution, emergency fault handling and safety risk control during aircraft recovery, providing an intelligent solution for the dynamic scheduling of modern aircraft, providing an effective optimization target for the reinforcement learning agent, and realizing the generation of optimal scheduling strategies under multi-objective constraints. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0016] Figure 1 This is a flow chart of a method for dynamic aircraft recovery sorting based on reinforcement learning in one embodiment of the present application. DETAILED DESCRIPTION

[0017] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0018] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0019] In an exemplary embodiment, Figure 1 As shown, a dynamic aircraft recovery sorting method based on reinforcement learning is provided, including: Step 101: Build a dynamic recovery environment model including go-around scenarios .

[0020] Environmental model initialization: Input the model parameters k∈K and initial fuel quantity of the aircraft set I to be recovered and safety thresholds , set the maximum number of decision stages , construct the state space S. Define the execution phase set of the sorting sequence .

[0021] Step 102: Based on the dynamic recovery environment model, a multi-dimensional reward function is constructed. The multi-dimensional reward function is used to evaluate the overall scheduling effect of the sorting sequence and describes the fuel consumption penalty, fault priority reward, mission priority reward, and go-around penalty in the go-around scenario.

[0022] The multidimensional reward function is: .

[0023] in, .

[0024] .

[0025] Where, r is a multidimensional reward function; is the maximum number of decision stages; For the stage The reward function of The adjustable reward weight coefficient is initialized before step 104 and is calculated based on the actual cumulative reward in the optimal sequence effectiveness evaluation stage in step 104. and actual security costs The ratio is adjusted dynamically; is the fuel consumption penalty item; is the rectified linear activation function; For aircraft In the stage The landing status indicator variable, Indicates a successful landing; Representation sequence The index of the aircraft waiting to land in the middle stage τ; Priority reward item for failure; It is a task priority reward item; This is the penalty item for missed approach.

[0026] The fuel consumption penalty items are: .

[0027] Where, Gathering aircraft for recovery, Indicates the cardinality of a set; Representation stage time; Representation stage time; Indicates the aircraft to be landed i is the aircraft type Fuel consumption rate at . Model category identifier, ; A collection of model categories.

[0028] The fault priority reward items are: .

[0029] Where, Indicates a plane waiting to land Failure probability assessment value; Indicates a plane waiting to land The higher the failure probability, the The greater the weight.

[0030] The mission priority reward items are: .

[0031] Where, Indicates a plane waiting to land The mission urgency ranking is used; the weight (mission urgency ranking) of the aircraft to be landed with the highest urgency is the largest and is equal to 1.

[0032] The missed approach penalty items are: .

[0033] in, .

[0034] Where, is the time decay factor; Indicates the number of delay stages caused by the missed approach; Delay phase due to missed approach The time decay factor of is the penalty intensity coefficient; Aircraft selected for the current decision The scheduled recycling phase, Indicates the aircraft selected for the current decision.

[0035] Step 103: Based on the dynamic recovery environment model and the multi-dimensional reward function, a safety cost function is constructed. The safety cost function is used to quantify the total fuel safety risk of all aircraft that have not landed in the aircraft recovery sequence.

[0036] The security cost function is: .

[0037] in, .

[0038] Where, is the security cost function; For the stage Single-stage security cost; For the stage Assembly of aircraft that have not landed; Indicates the safe fuel level threshold of the model; Indicates aircraft exist Stage remaining fuel; in aircraft exist Remaining oil level at this stage Less than the safe fuel level threshold of the model Time, stage Single-stage security cost Proportional to the amount of fuel shortage; is the safety cost conversion coefficient.

[0039] Step 104: Based on the multi-dimensional reward function and the safety cost function, the scheduling strategy is optimized by the reinforcement learning agent to obtain the optimal aircraft recovery sorting sequence.

[0040] Reinforcement learning phase: Optimizing scheduling strategies through reinforcement learning agents ,in For the set of all possible aircraft sorting sequences, under the constraints of multidimensional reward function and safety cost function, a dynamic balance is achieved between maximizing the multidimensional reward function and minimizing the safety cost function, and the optimal sequence is output. ;in Indicates the Aircraft index for stage landing. Dynamic recovery environment model is the state space, which contains all aircraft state information.

[0041] Evaluating the effectiveness of the optimal sequence: Based on the environmental model Simulate the optimal sequence of execution , calculate the actual cumulative reward and actual security costs ,like ( is the preset validity threshold), then the output is executed; otherwise, the policy parameters are adjusted and re-optimized. Specifically: the multi-dimensional reward function value corresponding to the optimal sequence is determined to be the actual cumulative reward; the security cost function value corresponding to the optimal sequence is determined to be the actual security cost; the ratio of the actual cumulative reward to the actual security cost is determined to be the validity factor ; Determine whether the effectiveness factor is lower than the preset effectiveness threshold and obtain a judgment result; if the judgment result is yes, adjust the adjustable reward weight coefficient in the multidimensional reward function and return to the reinforcement learning stage; if the judgment result is no, determine that the optimal sequence is the optimal aircraft recovery sorting sequence.

[0042] The formula for dynamically adjusting the weight coefficient is: .

[0043] in, is the learning rate, express The differential of 、 、 Corresponding to 、 、 Updated coefficients.

[0044] Step 105: Convert the optimal aircraft recovery sorting sequence into executable instructions and transmit them to the airport dispatch system.

[0045] Step 106: Execute the executable instruction and record the fuel consumption , risk indicators and go-around events ; Step 107: Use the fuel status dynamic update formula to update the fuel status, and return to step 101.

[0046] Dynamically update the fuel status while executing the sequence; original Expand to , indicating that the action space is the set of sorted sequences of all aircraft. The dynamic update formula of fuel status is: .

[0047] Where, Indicates aircraft exist Remaining oil level at the stage; Seconds, indicating the stage and The time increment between is the indicator function, when exist Indicator function when the stage has not landed The value is 1.

[0048] This application builds a dynamic environment model including a missed approach scenario , defines the set of execution phases of the sort sequence ; Design a multidimensional cumulative reward function for the overall sorting sequence r , of which fuel consumption penalty Fuel consumption rate by model Dynamic calculation based on detention time, priority reward for failures Based on the probability of failure To achieve normalized weighting, the task priority reward P uses a logarithmic function to quantify the urgency ranking ;, Go-around penalty According to the number of delay stages Establish an exponential decay model; at the same time, construct a security cost function , quantify the total fuel safety risk of all aircraft that have not landed in the sorting sequence; by the fuel shortage The conversion coefficient converts the security risk into a continuous optimizable variable; the reinforcement learning agent uses the strategy π: S → ADecisions are continuously optimized under the dynamic balance constraints of maximizing rewards and minimizing safety costs, forming a closed-loop learning mechanism of "environmental state perception - action decision-reward feedback - safety assessment - strategy update", and ultimately achieving global optimization of recycling efficiency and safety.

[0049] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, memory, an input / output (I / O) interface, and a communication interface. The processor, memory, and I / O interface are connected via a system bus, and the communication interface is connected to the system bus via the I / O interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The I / O interface of the computer device is configured to exchange information between the processor and an external device. The communication interface of the computer device is configured to communicate with an external terminal via a network connection. When executed by the processor, the computer program implements a reinforcement learning-based aircraft dynamic recovery sorting method.

[0050] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0051] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0052] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0053] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0054] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0055] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0056] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A dynamic aircraft recovery sorting method based on reinforcement learning, characterized in that: include: Build a dynamic recovery environment model including missed approach scenarios; Based on the dynamic recovery environment model, a multidimensional reward function is constructed; the multidimensional reward function is used to describe the fuel consumption penalty, fault priority reward, mission priority reward and go-around penalty in a go-around scenario; Constructing a safety cost function based on the dynamic recycling environment model and the multi-dimensional reward function; The safety cost function is used to quantify the total fuel safety risk of all aircraft that have not landed in the aircraft recovery sorting sequence; Based on the multidimensional reward function and the safety cost function, optimizing the scheduling strategy through a reinforcement learning agent to obtain an optimal aircraft recovery sorting sequence; The optimal aircraft recovery sorting sequence is converted into executable instructions and transmitted to the airport dispatch system.

2. The aircraft dynamic recovery sorting method based on reinforcement learning according to claim 1 is characterized in that: The multidimensional reward function is: ; in, ; ; Where, r is a multidimensional reward function; is the maximum number of decision stages; For the stage The reward function of is the adjustable reward weight coefficient; is the fuel consumption penalty item; is the rectified linear activation function; For aircraft In the stage The landing status indicator variable, Indicates a successful landing; Representation sequence The index of the aircraft waiting to land in the middle stage τ; Priority reward item for failure; It is a task priority reward item; This is the penalty item for missed approach.

3. The aircraft dynamic recovery sorting method based on reinforcement learning according to claim 2 is characterized in that: The fuel consumption penalty items are: ; Where, Gathering for aircraft to be recovered, Indicates the cardinality of a set; Representation stage time; Representation stage time; Indicates the aircraft to be landed i is the aircraft type Fuel consumption rate at .

4. The aircraft dynamic recovery sorting method based on reinforcement learning according to claim 3 is characterized in that: The fault priority reward items are: ; Where, Indicates a plane waiting to land Failure probability assessment value; Indicates a plane waiting to land The failure probability assessment value.

5. The aircraft dynamic recovery sorting method based on reinforcement learning according to claim 4 is characterized in that: The task priority reward items are: ; Where, Indicates a plane waiting to land The mission urgency ranking of the aircraft to be landed is the highest; the mission urgency ranking of the aircraft to be landed is 1.

6. The aircraft dynamic recovery sorting method based on reinforcement learning according to claim 5 is characterized in that: The missed approach penalty term is: ; in, ; Where, Delay phase due to missed approach The time decay factor of is the penalty intensity coefficient; Aircraft selected for the current decision scheduled recycling phase.

7. The aircraft dynamic recovery sorting method based on reinforcement learning according to claim 6, characterized in that: The security cost function is: ; in, ; Where, is the security cost function; For the stage Single-stage security cost; For the stage Assembly of aircraft that have not landed; Indicates the safe fuel level threshold of the model; Indicates aircraft exist Stage remaining fuel; in aircraft exist Remaining oil level at this stage Less than the safe fuel level threshold of the model Time, stage Single-stage security cost Proportional to the amount of fuel shortage; is the safety cost conversion coefficient.

8. The aircraft dynamic recovery sorting method based on reinforcement learning according to claim 2, characterized in that: Based on the multi-dimensional reward function and the safety cost function, the optimal aircraft recovery ranking sequence is obtained by optimizing the scheduling strategy through the reinforcement learning agent, specifically including: By optimizing the scheduling strategy through reinforcement learning agents, a dynamic balance is achieved between maximizing the multi-dimensional reward function and minimizing the safety cost function, and the optimal sequence is output; Determine the multidimensional reward function value corresponding to the optimal sequence as the actual cumulative reward; Determine the security cost function value corresponding to the optimal sequence as the actual security cost; Determine the ratio of actual cumulative rewards to actual security costs as the effectiveness factor; Determine whether the validity factor is lower than a preset validity threshold, and obtain a determination result; If the judgment result is yes, then adjust the adjustable reward weight coefficient in the multidimensional reward function, and return to the step of "optimizing the scheduling strategy through the reinforcement learning agent to achieve a dynamic balance between maximizing the multidimensional reward function and minimizing the safety cost function, and outputting the optimal sequence"; If the judgment result is no, then the optimal sequence is determined to be the optimal aircraft recovery sorting sequence.

9. The aircraft dynamic recovery sorting method based on reinforcement learning according to claim 7, characterized in that: After converting the optimal aircraft recovery sequence into executable instructions and transmitting them to the airport dispatch system, the method further includes: executing the executable instructions and recording fuel consumption, risk indicators, and go-around events; Use the fuel status dynamic update formula to update the fuel status and return to step "Building a dynamic recovery environment model including a missed approach scenario"; The fuel status dynamic update formula is: ; Where, Indicates aircraft exist Remaining oil level at the stage; Representation stage With stage The time increment between is the indicator function, when exist Indicator function when the stage has not landed The value is 1.

10. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aircraft dynamic recovery sorting method based on reinforcement learning according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Avionic display systems and methods for generating vertical situation displays

    CN108665731A

  • Aircraft hangar regular maintenance scheduling method, device, equipment and medium

    CN114462136A

  • Reinforced learning unmanned aerial vehicle flight path planning method based on delayed experience-first playback mechanism

    CN116974299A

  • Intranet service quality optimization method and system based on deep reinforcement learning

    CN119496716A

  • Energy sensing unmanned aerial vehicle data acquisition path planning method

    CN119645073A