Aircraft fleet support scheduling method based on discrete-time Markov decision making

By constructing a directed acyclic graph and a discrete-time Markov decision model, the problems of insufficient flexibility and accuracy in aircraft fleet support scheduling are solved, efficient and applicable scheduling solutions are generated, and the aircraft fleet deployment capability is improved.

CN120525293BActive Publication Date: 2025-10-03NAVAL AVIATION UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511012958.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-03
Estimated Expiration
2045-07-23

Smart Images

  • Figure CN120525293B_ABST
    Figure CN120525293B_ABST
Patent Text Reader

Abstract

This application discloses a method for aircraft fleet support scheduling based on discrete-time Markov decision making, relating to the field of aircraft support scheduling. The method comprises: obtaining aircraft fleet support scheduling task data; constructing a directed acyclic graph scheduling model based on the aircraft fleet support scheduling task data; converting the dependencies between the aircraft fleet support scheduling task data into nodes and edges in the directed acyclic graph scheduling model; the directed acyclic graph scheduling model comprises a node set and an edge set; and determining an aircraft fleet support scheduling plan under the constraints of the edge set based on the node set and a discrete-time Markov decision model; the discrete-time Markov decision model is obtained by training a reinforcement learning network. This application can improve the efficiency and applicability of aircraft fleet support scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of aircraft support scheduling, and in particular to an aircraft fleet support scheduling method based on discrete-time Markov decision making. Background Art

[0002] In the field of modern aircraft support and dispatch, the ability to deploy aircraft fleets is a core element supporting the effective operation of formation systems and plays a vital role in their formation and performance. Improving fleet deployment capabilities relies on the efficient implementation of fleet aviation support. Fleet aviation support deeply integrates multiple aspects, including aviation equipment, support resources, and command and decision-making. By implementing real-time, precise, and intelligent dispatch and control, efficient rotation of aircraft fleets can be ensured throughout the entire deployment process, meeting future demands for high intensity, high precision, and high efficiency.

[0003] However, in some cases, the aircraft fleet support scheduling mechanism has many limitations. When formulating the time-series allocation plan for aviation support personnel and service resource stations, commanders rely primarily on past scheduling experience, which makes the plan lack the necessary flexibility and precision. Problems such as insufficient scheduling model refinement, poor operability, and inefficient resource utilization frequently occur, seriously affecting the safe, orderly, and efficient operation of aircraft fleet aviation support. Against the backdrop of countries around the world continuously improving their support capabilities, facing the pressure of multi-type, large-scale, and high-intensity aircraft fleet deployment and recovery, traditional empirical scheduling modeling methods are no longer able to adapt to the new deployment requirements. Therefore, there is an urgent need to promote the transformation and development of aircraft fleet support scheduling modeling methods towards agile intelligence.

[0004] In recent years, artificial intelligence technology has experienced rapid development, achieving significant breakthroughs in numerous fields. Countries are actively leveraging AI to transform traditional deployment and support models to adapt to the evolving nature of aircraft support and scheduling, as well as the profound shifts in operational strategies. As information technology continues to advance and intelligent technologies emerge, leveraging intelligent technology to provide precise and efficient dispatch support and decision-making is crucial for addressing critical challenges in aviation support, such as high-dimensional, multi-variable, and complex coupling constraints that hinder deployment capabilities. This not only lays a solid foundation for enhancing fleet capabilities but also points the way forward for future support transformation and development.

[0005] However, in some cases, this approach presents significant shortcomings in aircraft fleet support scheduling. In different mission scenarios, obtaining the optimal solution often requires repeated iterative calculations and hyperparameter adjustments, which significantly limits its efficiency and applicability. Furthermore, due to design flaws in the algorithm's coding rules and population evolution mechanism, optimization results are poor when faced with complex constraints such as aircraft fleet support. Summary of the Invention

[0006] The purpose of this application is to provide an aircraft fleet support scheduling method based on discrete-time Markov decision-making, which can improve the efficiency and applicability of aircraft fleet support scheduling.

[0007] To achieve the above objectives, this application provides the following solutions.

[0008] In the first aspect, the present application provides an aircraft fleet support scheduling method based on discrete-time Markov decision making, comprising: obtaining aircraft fleet support scheduling task data; constructing a directed acyclic graph scheduling model based on the aircraft fleet support scheduling task data; the directed acyclic graph scheduling model converts the dependency relationship between the aircraft fleet support scheduling task data into nodes and edges; the directed acyclic graph scheduling model comprises: a node set and an edge set; based on the node set and the discrete-time Markov decision model, determining an aircraft fleet support scheduling plan under the constraints of the edge set; the discrete-time Markov decision model is obtained by training a reinforcement learning network.

[0009] Optionally, the data of the aircraft fleet support scheduling task includes the flight support task plan start time vector , No. Aviation support resource matrix and process timing constraint set ;

[0010] in, The start time of the support mission for the first aircraft, For the The start time of the support mission for the aircraft, for and Timing constraints, For the The first aircraft Road guarantee process, For the The first aircraft The first guarantee process The first aircraft Road guarantee process Contains two-dimensional attribute tuples , For the The first aircraft Duration of the road guarantee process, For the The first aircraft The resource demand vector of the guarantee process.

[0011] Optionally, the node set is ,in, For nodes, = , To ensure the number of processes; the edge set is ,in, for Edges connecting to other nodes, for Must be in Completed immediately before.

[0012] Optionally, the training process of the discrete-time Markov decision model specifically includes: constructing a sample directed acyclic graph scheduling model; the sample directed acyclic graph scheduling model includes: a sample node set and a sample edge set; initializing the policy network parameters in the reinforcement learning network; based on the sample node set, determining the sample state vector at the current moment; inputting the sample state vector at the current moment into the policy network, and obtaining the probability of each sample action at the current moment under the constraint of the sample edge set; according to the probability of each sample action at the current moment, using the ε-greedy strategy to screen the sample actions to obtain the sample actions at the current moment; based on the sample state vector at the current moment and the sample action at the current moment, constructing a scheduling objective function, and iteratively optimizing the policy network parameters with the goal of maximizing the scheduling objective function to obtain a discrete-time Markov decision model.

[0013] Optionally, constructing a sample directed acyclic graph scheduling model specifically includes: obtaining sample aircraft fleet support scheduling task data; and constructing a sample directed acyclic graph scheduling model based on the sample aircraft fleet support scheduling task data.

[0014] Optionally, based on the sample state vector at the current moment and the sample action at the current moment, a scheduling objective function is constructed, and with the goal of maximizing the scheduling objective function, the policy network parameters are iteratively optimized to obtain a discrete-time Markov decision model, specifically including: calculating the reward at the current moment based on the sample state vector at the current moment and the sample action at the current moment; calculating the scheduling objective function based on the reward at the current moment; with the goal of maximizing the value of the scheduling objective function, the policy network parameters are iteratively optimized to obtain a discrete-time Markov decision model.

[0015] Optionally, the expression of the scheduling objective function is as follows.

[0016] .

[0017] .

[0018] .

[0019] Among them, among them, is the scheduling objective function; for Discount factor for the moment; is the reward function; for The state vector at the moment; for The action of the moment; is the activation function; As of The maximum completion time at the moment; As of The maximum completion time at the moment; is the initial discount factor; for The discount factor at the moment.

[0020] Optionally, based on the node set and the discrete-time Markov decision model, the aircraft fleet support scheduling scheme is determined under the constraint of the edge set, specifically including: determining the state vector at the current moment based on the node set; the state vector includes the process state vector and the support resource occupancy vector; inputting the state vector at the current moment into the strategy network, using the edge set as a constraint, obtaining the probability of each action at the current moment; according to the probability of each action at the current moment, using the ε-greedy strategy to filter the actions to obtain the action at the current moment; the action is the first The first aircraft The system calculates the state transition probability at the current moment according to the state vector at the current moment and the action at the current moment; determines the state vector at the next moment according to the state transition probability at the current moment; inputs the state vector at the next moment into the strategy network, and obtains the probability of each action at the current moment with the edge set as the constraint, until all the support processes in the node set are scheduled, and the support processes after the scheduling are determined as the aircraft fleet support scheduling plan.

[0021] Optionally, the calculation formula for the state transition probability is as follows.

[0022] .

[0023] Among them, among them, is the state transition probability; for The state vector at the moment; for The state vector at the moment; for The action of the moment; for Execution success rate parameters; for Moment and The difference in time.

[0024] In the second aspect, the present application provides an aircraft cluster support and scheduling system based on discrete-time Markov decision, including: an acquisition module for acquiring aircraft cluster support and scheduling task data; a directed acyclic graph scheduling model construction module for constructing a directed acyclic graph scheduling model based on the aircraft cluster support and scheduling task data; the directed acyclic graph scheduling model converts the dependency relationship between the aircraft cluster support and scheduling task data into nodes and edges; the directed acyclic graph scheduling model includes: a node set and an edge set; an aircraft cluster support and scheduling scheme determination module for determining the aircraft cluster support and scheduling scheme under the constraints of the edge set based on the node set and the discrete-time Markov decision model; the discrete-time Markov decision model is obtained by training a reinforcement learning network.

[0025] According to the specific embodiments provided in this application, this application has the following technical effects:

[0026] The present application provides an aircraft fleet support scheduling method based on discrete-time Markov decision making. The data of aircraft fleet support scheduling tasks are converted into the form of nodes and edges through a directed acyclic graph scheduling model, so that the data of different aircraft fleet support scheduling tasks in various related application scenarios can be converted into the form of node and edge connections, and the dependency relationship between tasks in the data of aircraft fleet support scheduling tasks is quickly obtained, while improving the applicability of multiple scenarios; then, a discrete-time Markov decision model is adopted, and scheduling decisions are made on the node set with the edge set as a constraint, so as to quickly obtain a discrete-time solution for aircraft fleet support scheduling, thereby improving the efficiency and applicability of aircraft fleet support scheduling. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0028] Figure 1 A flowchart of an aircraft fleet support and scheduling method based on discrete-time Markov decision-making is provided in accordance with an embodiment of the present application.

[0029] Figure 2 A schematic diagram of the structure of a directed acyclic graph scheduling model provided in one embodiment of the present application.

[0030] Figure 3 A schematic diagram of the structure of an aircraft fleet support and scheduling system based on discrete-time Markov decision-making is provided in one embodiment of the present application. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0032] In some cases, the methods used to schedule aircraft fleet support can only be used to plan and schedule a certain type of scheduling task. However, in reality, aircraft fleet support scheduling involves many types of tasks. Therefore, traditional methods have poor applicability.

[0033] Based on the above background, this application focuses on the aircraft fleet support scheduling modeling method. Driven by artificial intelligence methods, it deeply analyzes the key elements of aircraft fleet support, including the flight deck environment, aviation support resources and processes. By sorting out and summarizing various constraints, the subject and constraints are transformed into key components of the discrete-time Markov decision process, including state space, action strategy, state transition and reward function. On this basis, this application designs an aircraft fleet support scheduling method based on discrete-time Markov decision, aiming to refine the scheduling modeling process and improve the efficiency and applicability of aircraft fleet support scheduling.

[0034] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0035] In an exemplary embodiment, Figure 1 As shown, a method for aircraft fleet support scheduling based on discrete-time Markov decision-making is provided. The method is executed by a computer device, and specifically can be executed by a computer device such as a terminal or a server alone, or can be executed by a terminal and a server together. In the embodiment of the present application, the method is applied to a server as an example for explanation, including the following steps S1 to S3.

[0036] Step S1: Acquire aircraft fleet support scheduling task data.

[0037] Furthermore, the data of the aircraft fleet support scheduling task includes the flight support task plan start time vector , No. Aviation support resource matrix and process timing constraint set ;in, The start time of the support mission for the first aircraft, For the The start time of the support mission for the aircraft, for and Timing constraints, For the The first aircraft Road guarantee process, For the The first aircraft The first guarantee process The first aircraft Road guarantee process Contains two-dimensional attribute tuples , For the The first aircraft Duration of the road guarantee process, For the The first aircraft The resource demand vector of the guarantee process.

[0038] Specifically, the flight support mission is scheduled to start time Dimension It is determined by the maximum number of parallel support aircraft, that is, the number of aircraft groups for a single support mission. Each element Contains The start time of the support mission for the aircraft , No. Deadline for aircraft support mission Hedi Priority weight of support missions for aircraft .

[0039] Specifically, Aviation support resource matrix correspond Class security resources, Indicates the total amount of this type of resources, satisfying any The first aircraft The first step of the guarantee process All aviation support resources are met .

[0040] Specifically, the process timing constraint set Generated by the process dependency analysis module, if and only if the process The output is Input setting .

[0041] Step S2: Based on the aircraft fleet support scheduling task data, a directed acyclic graph scheduling model is constructed; the directed acyclic graph scheduling model converts the dependency relationship between the aircraft fleet support scheduling task data into nodes and edges; the directed acyclic graph scheduling model includes: a node set and an edge set.

[0042] Furthermore, the node set is ,in, For nodes, = , To ensure the number of processes.

[0043] The edge set is ,in, for Edges connecting to other nodes, for Must be in Completed immediately before.

[0044] In a specific embodiment, if Figure 2 As shown in the figure, a structural diagram of a directed acyclic graph scheduling model for a single aircraft is provided. Sequence numbers 1-21 represent nodes, and the lines between each sequence number represent edges. Sequence number 1 is process 1, sequence number 2 is process 2, sequence number 3 is process 3, sequence number 4 is process 4, sequence number 5 is process 5, sequence number 6 is process 6, sequence number 7 is process 7, sequence number 8 is process 8, sequence number 9 is process 9, sequence number 10 is process 10, sequence number 11 is process 11, sequence number 12 is process 12, sequence number 13 is process 13, sequence number 14 is process 14, sequence number 15 is process 15, sequence number 16 is process 16, sequence number 17 is process 17, sequence number 18 is process 18, sequence number 19 is process 19, sequence number 20 is process 20, and sequence number 21 is process 21. Starting from process number 1, the subsequent processes that can be carried out are obtained according to the edge direction, and the execution order is determined by the decision of the intelligent agent until process number 21 is completed, that is, the support process of a single aircraft is fully executed.

[0045] Step S3: Based on the node set and the discrete-time Markov decision model, the aircraft fleet support scheduling plan is determined under the constraints of the edge set; the discrete-time Markov decision model is obtained by training the reinforcement learning network.

[0046] Furthermore, the training process of the discrete-time Markov decision model specifically includes: constructing a sample directed acyclic graph scheduling model; the sample directed acyclic graph scheduling model includes: a sample node set and a sample edge set; initializing the policy network parameters in the reinforcement learning network; based on the sample node set, determining the sample state vector at the current moment; inputting the sample state vector at the current moment into the policy network, and obtaining the probability of each sample action at the current moment under the constraint of the sample edge set; according to the probability of each sample action at the current moment, using the ε-greedy strategy to screen the sample actions to obtain the sample actions at the current moment; based on the sample state vector at the current moment and the sample action at the current moment, constructing a scheduling objective function, and iteratively optimizing the policy network parameters with the goal of maximizing the scheduling objective function to obtain a discrete-time Markov decision model.

[0047] Specifically, the training set consists of 128 flight support mission instances that are sequentially input into the model for training. During the training phase, the ε-greedy strategy is used to explore the probability ,in, for t Always explore probabilities; is the initial exploration probability, ; is the attenuation rate . Output from the policy network The specific action process to generate is as follows: set a random number u in the range of 0-1 that obeys uniform distribution, when When, from the action space Otherwise, select the action with the highest probability in the current policy network probability distribution.

[0048] Furthermore, a sample directed acyclic graph scheduling model is constructed, specifically including: obtaining sample aircraft fleet support scheduling task data; and constructing a sample directed acyclic graph scheduling model based on the sample aircraft fleet support scheduling task data.

[0049] Furthermore, based on the sample state vector at the current moment and the sample action at the current moment, a scheduling objective function is constructed, and with the goal of maximizing the scheduling objective function, the policy network parameters are iteratively optimized to obtain a discrete-time Markov decision model. Specifically, the method includes: calculating the reward at the current moment based on the sample state vector at the current moment and the sample action at the current moment; calculating the scheduling objective function based on the reward at the current moment; and iteratively optimizing the policy network parameters with the goal of maximizing the value of the scheduling objective function to obtain a discrete-time Markov decision model.

[0050] Furthermore, the expression of the scheduling objective function is as follows.

[0051] .

[0052] .

[0053] .

[0054] in, is the scheduling objective function; for Discount factor for the moment; is the reward function; for The state vector at the moment; for The action of the moment; is the activation function; As of The maximum completion time at the moment; As of The maximum completion time at the moment; is the initial discount factor; for The discount factor at the moment.

[0055] Furthermore, step S3 specifically includes steps S31 to S36.

[0056] Step S31: Based on the node set, determine the state vector at the current moment; the state vector includes the process state vector and the guaranteed resource occupancy vector.

[0057] Step S32: Input the current state vector into the policy network, use the edge set as a constraint, and obtain the probability of each action at the current moment.

[0058] Specifically, the state space , for The process state vector at time t indicates the process is The variable indicating the completed state at all times contains all completed processes. for Moment Aviation support resource occupancy vector.

[0059] Step S33: According to the probability of each action at the current moment, the ε-greedy strategy is used to filter the actions to obtain the action at the current moment; the action is the first The first aircraft Road guarantee process.

[0060] Specifically, the action space , that is, the action is to select a process to start at the current time, where Represents the set of all immediately preceding nodes of a node, determined by the preceding edges of the node, representing an action The condition for being selected is that the process is All the preceding processes at the time have been completed, and the timing constraint relationship C between the processes is satisfied; when the When building the first stage of an aircraft, use the tensor within Is it greater than the current time? This condition replaces the judgment condition of the previous process.

[0061] Specifically, use the greedy strategy in the deployment phase, that is, set , in each round, only the action with the highest probability in the current policy network probability distribution is selected.

[0062] Step S34: Calculate the state transition probability at the current moment based on the state vector at the current moment and the action at the current moment.

[0063] Step S35: Determine the state vector at the next moment according to the state transition probability at the current moment.

[0064] Furthermore, the calculation formula of the state transition probability is as follows.

[0065] .

[0066] in, is the state transition probability; for The state vector at the moment; for The state vector at the moment; for The action of the moment; for Execution success rate parameters; for Moment and The difference in time.

[0067] Step S36: Input the state vector at the next moment into the strategy network, and use the edge set as a constraint to obtain the probability of each action at the current moment until all the support processes in the node set are scheduled, and the support processes after the scheduling are determined as the aircraft fleet support scheduling plan.

[0068] The beneficial effects of the aircraft fleet support scheduling method based on discrete-time Markov decision-making proposed in this application are mainly reflected in the following three aspects.

[0069] (1) This application adopts a unified Markov decision model to solve the aircraft fleet support scheduling problem, so that the data in processing the aircraft fleet support scheduling task can be converted into a directed acyclic graph (DAG) of the aircraft fleet support scheduling problem, which is more flexible and effective, suitable for a variety of related application scenarios, and can overcome the shortcomings of traditional modeling such as weak generalization ability, thereby improving the applicability of multiple scenarios.

[0070] (2) This application takes actual scheduling needs as the background, and proposes an aircraft fleet support scheduling method based on discrete-time Markov decision-making for the complex aircraft support operation scheduling problem with multiple constraints. The purpose is to refine the scheduling modeling process, improve the efficiency and applicability of aircraft fleet support scheduling, and thus better meet the urgent needs for intelligent aircraft support operation scheduling methods in training environments, and provide theoretical and technical support for aircraft support operation scheduling modeling methods on existing and future large-scale offshore platforms, thereby improving the aircraft's dispatch and recovery capabilities and response capabilities, better adapting to complex and changing environments, and improving the overall effectiveness of the formation.

[0071] (3) The aircraft fleet support scheduling problem uses a directed acyclic graph scheduling model to efficiently represent the scheduling underlying environment, and the discrete time Markov decision is good at capturing the feature relationship in the graph structure, so it has a high adaptability. The discrete time Markov decision not only considers the characteristics of nodes and edges, but also can map the state characteristics of aircraft fleet support scheduling in the graph structure, and map these characteristics to the action probability distribution through the neural network, providing strong support for decision-making. The scheduling agent obtained by the neural network can achieve "end-to-end" deployment after pre-training, and generate high-quality scheduling solutions for aircraft fleet support scheduling problems of different scales, showing significant flexibility and adaptability. Through case experiments, this application found that compared with the existing benchmark algorithms, the proposed method showed significant advantages in solution speed, accuracy and generalization ability. Especially when dealing with complex scenarios, compared with the iterative search metaheuristic algorithm that takes minutes, the solution time of this method is only in the second range, and at the same time, it achieves a performance level that is not inferior to the benchmark algorithm, which verifies the advantages of this method.

[0072] Based on the same inventive concept, embodiments of the present application also provide an aircraft fleet support and scheduling system based on discrete-time Markov decision making. The implementation solutions provided by this system are similar to those described in the aforementioned methods. Therefore, the specific limitations of one or more embodiments of the aircraft fleet support and scheduling system based on discrete-time Markov decision making provided below can be found in the limitations of the aircraft fleet support and scheduling method based on discrete-time Markov decision making above, and will not be further elaborated here.

[0073] In an exemplary embodiment, an aircraft fleet support and scheduling system based on discrete-time Markov decision making is provided, which includes the following modules.

[0074] Acquisition module, used to obtain aircraft fleet support scheduling task data;

[0075] A directed acyclic graph scheduling model construction module is used to construct a directed acyclic graph scheduling model based on the aircraft fleet support scheduling task data; the directed acyclic graph scheduling model converts the dependency relationship between the aircraft fleet support scheduling task data into nodes and edges; the directed acyclic graph scheduling model includes: a node set and an edge set;

[0076] The aircraft fleet support scheduling plan determination module is used to determine the aircraft fleet support scheduling plan under the constraints of the edge set based on the node set and the discrete time Markov decision model; the discrete time Markov decision model is obtained by training the reinforcement learning network.

[0077] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0078] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for aircraft fleet support scheduling based on discrete-time Markov decision making, characterized in that: The aircraft fleet support scheduling method based on discrete-time Markov decision-making includes: Obtain aircraft fleet support and scheduling task data; The data of the aircraft fleet support scheduling task includes the flight support task plan start time vector , No. Aviation support resource matrix and process timing constraint set ; in, The start time of the support mission for the first aircraft, For the The start time of the support mission for the aircraft, for and Timing constraints, For the The first aircraft Road guarantee process, For the The first aircraft The first guarantee process The first aircraft Road guarantee process Contains two-dimensional attribute tuples , For the The first aircraft Duration of the road guarantee process, For the The first aircraft The resource demand vector of the guarantee process; Based on the aircraft fleet support scheduling task data, a directed acyclic graph scheduling model is constructed; the directed acyclic graph scheduling model converts the dependency relationship between the aircraft fleet support scheduling task data into nodes and edges; the directed acyclic graph scheduling model includes: a node set and an edge set; Determining an aircraft fleet support scheduling plan under the constraints of an edge set based on a node set and a discrete-time Markov decision model obtained by training a reinforcement learning network; The training process of the discrete-time Markov decision model specifically includes: Constructing a sample directed acyclic graph scheduling model; the sample directed acyclic graph scheduling model includes: a sample node set and a sample edge set; Initialize the policy network parameters within the reinforcement learning network; Based on the sample node set, determine the sample state vector at the current moment; Input the current sample state vector into the policy network, and obtain the probability of each sample action at the current moment under the constraints of the sample edge set; According to the probability of each sample action at the current moment, the ε-greedy strategy is used to filter the sample actions to obtain the sample actions at the current moment; Based on the current sample state vector and the current sample action, a scheduling objective function is constructed. With the goal of maximizing the scheduling objective function, the policy network parameters are iteratively optimized to obtain a discrete-time Markov decision model, which specifically includes: Calculate the reward at the current moment based on the sample state vector at the current moment and the sample action at the current moment; Calculate the scheduling objective function based on the reward at the current moment; With the goal of maximizing the scheduling objective function value, the policy network parameters are iteratively optimized to obtain a discrete-time Markov decision model.

2. The aircraft fleet support scheduling method based on discrete-time Markov decision making according to claim 1 is characterized in that: The node set is ,in, For nodes, = , To ensure the number of processes; The edge set is ,in, for Edges connecting to other nodes, for Must be in Completed immediately before.

3. The aircraft fleet support scheduling method based on discrete-time Markov decision making according to claim 1 is characterized in that: Construct a sample directed acyclic graph scheduling model, including: Obtain sample aircraft fleet support and scheduling task data; Based on the sample aircraft fleet support scheduling task data, a sample directed acyclic graph scheduling model is constructed.

4. The aircraft fleet support scheduling method based on discrete-time Markov decision making according to claim 1 is characterized in that: The expression of the scheduling objective function is: ; ; ; in, is the scheduling objective function; for Discount factor for the moment; is the reward function; for The state vector at the moment; for The action of the moment; is the activation function; As of The maximum completion time at the moment; As of The maximum completion time at the moment; is the initial discount factor; for The discount factor at the moment.

5. The aircraft fleet support scheduling method based on discrete-time Markov decision making according to claim 2, characterized in that: Based on the node set and discrete-time Markov decision model, the aircraft fleet support scheduling plan is determined under the constraints of the edge set, including: Determine the state vector at the current moment based on the node set; the state vector includes a process state vector and a guaranteed resource occupancy vector; Input the current state vector into the policy network, use the edge set as a constraint, and obtain the probability of each action at the current moment; According to the probability of each action at the current moment, the ε-greedy strategy is used to filter the actions to obtain the action at the current moment; the action is the The first aircraft Road guarantee process; Calculate the state transition probability at the current moment based on the current state vector and the current action; Determine the state vector at the next moment based on the state transition probability at the current moment; The state vector at the next moment is input into the strategy network, and the probability of each action at the current moment is obtained with the edge set as the constraint, until all the support processes in the node set are scheduled, and the support processes after the scheduling are determined as the aircraft fleet support scheduling plan.

6. The aircraft fleet support scheduling method based on discrete-time Markov decision making according to claim 5 is characterized in that: The calculation formula of the state transition probability is: ; in, is the state transition probability; for The state vector at the moment; for The state vector at the moment; for The action of the moment; for Execution success rate parameters; for Moment and The difference in time.

7. An aircraft fleet support and dispatching system based on discrete-time Markov decision making, characterized in that: The aircraft fleet support and scheduling system based on discrete-time Markov decision-making includes: Acquisition module, used to obtain aircraft fleet support scheduling task data; The data of the aircraft fleet support scheduling task includes the flight support task plan start time vector , No. Aviation support resource matrix and process timing constraint set ; in, The start time of the support mission for the first aircraft, For the The start time of the support mission for the aircraft, for and Timing constraints, For the The first aircraft Road guarantee process, For the The first aircraft The first guarantee process The first aircraft Road guarantee process Contains two-dimensional attribute tuples , For the The first aircraft Duration of the road guarantee process, For the The first aircraft The resource demand vector of the guarantee process; A directed acyclic graph scheduling model construction module is used to construct a directed acyclic graph scheduling model based on the aircraft fleet support scheduling task data; the directed acyclic graph scheduling model converts the dependency relationship between the aircraft fleet support scheduling task data into nodes and edges; the directed acyclic graph scheduling model includes: a node set and an edge set; An aircraft fleet support scheduling solution determination module is used to determine an aircraft fleet support scheduling solution under the constraints of an edge set based on a node set and a discrete-time Markov decision model obtained by training a reinforcement learning network; The training process of the discrete-time Markov decision model specifically includes: Constructing a sample directed acyclic graph scheduling model; the sample directed acyclic graph scheduling model includes: a sample node set and a sample edge set; Initialize the policy network parameters within the reinforcement learning network; Based on the sample node set, determine the sample state vector at the current moment; Input the current sample state vector into the policy network, and obtain the probability of each sample action at the current moment under the constraints of the sample edge set; According to the probability of each sample action at the current moment, the ε-greedy strategy is used to filter the sample actions to obtain the sample actions at the current moment; Based on the current sample state vector and the current sample action, a scheduling objective function is constructed. With the goal of maximizing the scheduling objective function, the policy network parameters are iteratively optimized to obtain a discrete-time Markov decision model, which specifically includes: Calculate the reward at the current moment based on the sample state vector at the current moment and the sample action at the current moment; Calculate the scheduling objective function based on the reward at the current moment; With the goal of maximizing the scheduling objective function value, the policy network parameters are iteratively optimized to obtain a discrete-time Markov decision model.

Citation Information

Patent Citations

  • Job-shop adaptive scheduling method based on deep reinforcement learning

    CN114707881A

  • Method and system for scheduling tasks

    CN116670684A