Civil aviation fleet maintenance decision-making method based on multi-body distributed hierarchical reinforcement learning

By employing a multi-body distributed hierarchical reinforcement learning method, this study solves the decision-making challenge of civil aviation fleet maintenance strategies in dynamic environments, achieving efficient and economical maintenance scheduling and resource allocation, improving learning efficiency and strategy optimization capabilities, and making it suitable for large-scale aviation maintenance scheduling.

CN121684871APending Publication Date: 2026-03-17ZHEJIANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511863939.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing aircraft maintenance strategies struggle to achieve rapid response and robust decision-making in dynamic disturbance environments, especially in uncertain scenarios such as sensor failures, where they cannot provide effective maintenance solutions. Traditional single-agent reinforcement learning suffers from low learning efficiency and limited policy generalization ability in large-scale civil aviation fleet maintenance.

Method used

A multi-body distributed hierarchical reinforcement learning approach is adopted, which establishes a distributed learner and evolver through the collaborative work of multiple computing nodes. Combined with multi-level task decomposition and local feedback modeling, the fleet maintenance decision is optimized, including fleet management, spare engine management, flight mission and maintenance support sub-models. Distributed reinforcement learning is used to improve learning efficiency and strategy optimization capabilities.

Benefits of technology

In large-scale and complex aviation maintenance scheduling problems, it achieves efficient and economical maintenance scheduling and resource allocation, improves learning efficiency and strategy optimization capabilities, adapts to lifecycle management of multiple airports, multiple tasks, and multiple aircraft, and reduces system feedback costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121684871A_ABST
    Figure CN121684871A_ABST
Patent Text Reader

Abstract

The invention discloses a civil aviation fleet maintenance decision-making method based on multi-body distributed hierarchical reinforcement learning, and the method comprises the steps: constructing a civil aviation fleet maintenance decision-making environment through fleet parametric modeling, and simulating a multi-aircraft multi-task scheduling and maintenance process in real operation; and key factors such as equipment degradation, maintenance resource constraint and task execution priority are fully considered. In the environment, a distributed hierarchical reinforcement learning algorithm is adopted: an upper-layer strategy is used for global maintenance resource allocation and task coordination, and a lower-layer strategy is used for local maintenance actions of a specific body. According to the method, a shared communication mechanism and a reward structure are introduced among multiple agents, so that the cooperation efficiency is improved, and the training convergence is accelerated. Experimental results show that the method is remarkably superior to a traditional heuristic method and a standard reinforcement learning algorithm in the aspects of task completion rate, maintenance cost control and system robustness, and high efficiency and reliability of civil aviation team maintenance are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of civil aviation fleet operation and maintenance, and particularly relates to a civil aviation fleet maintenance decision-making method based on multi-body distributed hierarchical reinforcement learning. BACKGROUND

[0002] With the continuous expansion of modern air transportation scale and the continuous improvement of operation complexity, the maintenance and operation management of civil aviation fleet has gradually become the core link to ensure flight safety, improve flight punctuality rate and control operation cost. The civil aviation fleet usually consists of dozens to hundreds of aircraft, covering different aircraft types, service life and running state, and its running characteristics show the typical characteristics of high frequency, high intensity and high risk. In the long-term high-intensity use process, the power system, structural components and avionics system of the aircraft will inevitably have performance degradation and potential failure. Therefore, how to realize scientific, efficient and economic maintenance scheduling and resource allocation under the premise of safety has become an important problem that the civil aviation industry has long been concerned about.

[0003] Effective maintenance strategy not only relates to flight safety, but also directly affects fleet availability, flight normal rate, maintenance resource utilization rate and the cost structure of the entire aircraft life cycle. According to the relevant statistics of the International Air Transport Association and the International Civil Aviation Organization, maintenance cost has become the second largest item of airline operating expenses next to fuel cost, accounting for more than 10%-15% of total operating costs. In actual operation and maintenance, airlines need to consider the operation of hundreds of aircraft and hundreds of routes, involving different aircraft types, different engine combinations and the procurement, leasing and scheduling of spare engines. This makes the fleet operation problem a typical high-complexity decision problem, which not only involves maintenance task scheduling and route network coordination, but also needs to consider cost control and safety protection.

[0004] The existing aircraft maintenance strategy mainly includes fault repair maintenance, preventive maintenance and state-based maintenance. Among them, fault repair maintenance belongs to non-scheduled maintenance, which usually takes repair action after the equipment fails; preventive maintenance relies on regular inspection and maintenance to prevent potential failures; state-based maintenance uses sensors and monitoring technology to collect data and determines the maintenance time according to the actual state of the equipment. However, these methods mostly rely on static maintenance plans or current system states, and it is difficult to achieve rapid response and robust decision-making in the face of dynamic disturbances such as flight adjustments, climate changes and unexpected events, especially in uncertain scenarios such as sensor failure, which often cannot give effective maintenance solutions.

[0005] In recent years, with the development of emerging technologies such as Internet of Things, cloud computing and big data, the maintenance management of civil aviation fleet is undergoing a transformation from experience-driven to data-driven, from static planning to dynamic optimization. Intelligent decision-making methods, especially reinforcement learning, show great potential in complex scheduling optimization problems. Reinforcement learning can learn long-term optimal strategies through interaction with the environment, and is suitable for high-dimensional state space and multi-stage decision-making scenarios. However, traditional single-agent reinforcement learning often has low learning efficiency and limited strategy generalization ability when facing civil aviation-level super-large state space and multi-task parallel requirements.

[0006] Therefore, distributed reinforcement learning, as a learning method that can train on multiple computing nodes or devices in parallel, has gradually attracted the attention of researchers. Through the collaborative work and strategy sharing of multiple agents, distributed reinforcement learning can effectively break through the computing bottleneck and improve the learning efficiency and convergence performance in large-scale complex tasks. For example, distributed frameworks such as IMPALA and SeedRL have shown superior performance in large-scale decision-making scenarios.

[0007] In the problem of civil aviation fleet operation and maintenance, introducing distributed hierarchical reinforcement learning can not only effectively alleviate the learning burden of a single agent, but also improve the learning efficiency and strategy optimization ability of the overall system in high-dimensional state space and complex dynamic environment through hierarchical task decomposition and local feedback modeling, thereby better meeting the dual demands of high reliability and low cost in actual operation of airlines. SUMMARY

[0008] The purpose of the present application is to overcome the shortcomings of the prior art and provide a civil aviation fleet maintenance decision-making method based on multi-body distributed hierarchical reinforcement learning.

[0009] To achieve the above purpose, the technical scheme adopted by the present application is as follows: As a first aspect of the present application, a civil aviation fleet maintenance decision-making method based on multi-body distributed hierarchical reinforcement learning is provided, comprising the following steps: Collecting civil aviation company fleet operation data and establishing a civil aviation company operation and maintenance model; Based on multiple system processes of multiple computers or establishing a distributed learner and a distributed evolutioner, the distributed evolutioner generates an initial fleet engine state based on a random civil aviation company operation and maintenance model , and formulates an initial maintenance action ; The distributed evolutioner interacts with the civil aviation company operation and maintenance model based on the initial fleet engine state and the initial maintenance action , and obtains a set of state-action value groups: fleet engine state , maintenance action , system feedback cost and the next step in fleet engine status Establish a replay buffer and group the state-action value of each step. Store the data in the replay buffer; select actions in the replay buffer and use them as a training set for training; update the policy network of the learner and the value network of the evolver; update the training set; determine whether the system running step size has reached the set system training step size. If it has, stop training, obtain the trained network, and output the civil aviation company's maintenance optimization strategy. Otherwise, continue to select actions and repeat training until the running step size reaches the training step size.

[0010] Furthermore, the civil aviation company's operation and maintenance model includes a fleet management sub-model, a backup engine management sub-model, a flight mission sub-model, and a maintenance support sub-model. The fleet management sub-model constructs operational differences through multi-aircraft combination settings, where aircraft parameters cover maintenance requirements, range capability, and spare engine type; it sets daily maintenance capacity, aircraft support range, inventory, and transfer speed through airport distribution and maintenance facility capacity settings; it introduces maintenance transfer costs and a centralized-transfer mechanism for spare engines; and it outputs a high-dimensional system network model that satisfies airport capacity constraints and aircraft unique location constraints for maintenance scheduling optimization and inventory configuration analysis. The standby engine management model models the demand for engines of different aircraft types through aircraft type specificity constraints. The inventory variable is constrained by the dynamic balance equation of inventory consumption, return for repair, and cross-airport transfer. Each airport must meet the minimum safety stock threshold and comprehensively consider inventory costs, delay costs, and transfer costs to output a dynamic management strategy that satisfies the optimization of inventory configuration and transfer in a multi-aircraft-type and multi-airport environment. The flight mission model models the route mission through segment scheduling. The segment mission is defined by the departure airport, arrival airport, takeoff time, flight duration, and required aircraft type. The scheduling variables are subject to the constraints of unique mission allocation and time feasibility to ensure that the aircraft's continuous missions do not overlap and the segments are reasonably connected. The maintenance support model models fleet maintenance through maintenance resource constraints and state transition processes. Maintenance tasks are limited by ground conditions and single-trigger constraints; airport maintenance capacity is constrained by man-hours, workstations, and worker allocation; maintenance costs consist of maintenance expenses, delay penalties, and additional losses; engine condition decreases with each flight cycle and recovers to its optimal state with a repair probability, and the state transition function... By combining failure rate and repair rate outputs, a dynamic support optimization mechanism is formed that covers the cumulative effects of flight cycles and maintenance interventions.

[0011] Furthermore, the distributed learner learns from experience by building multiple policy networks and periodically scores the prediction results of all policy networks: ; Indicates from time step Time to step The cumulative rewards for learners; It represents the reward at time step t; the discount factor γ is used to reduce the impact of future rewards on the current decision. Indicates from state to state The change in the predicted value of the policy network is used to measure the difference in the model's predictions between different states; After periodically scoring all policy networks, the parameters of the policy network with the highest score are synchronized to all policy networks.

[0012] Furthermore, the civil aviation company's operation and maintenance model is based on its multi-airport configuration, and the output includes the total maintenance cost and the maintenance cost of each airport.

[0013] Furthermore, when receiving multiple outputs from the civil aviation company's operation and maintenance model, an auxiliary network is introduced for each airport. The input is the local state of the airport. The output is the predicted local reward.

[0014] Furthermore, a total neural network loss function is constructed based on multiple auxiliary networks; the auxiliary networks are used to learn the mapping relationship between airport-level feedback signals and states, and are added to the loss function as auxiliary learning signals: ; in, For the total loss, Loss to the main task This is the i-th auxiliary loss; The weights are used to assist in loss calculations and control the balance between global and local modeling.

[0015] As a second aspect of the present invention, a civil aviation fleet maintenance decision-making system based on multi-body distributed hierarchical reinforcement learning is provided to implement the above-described method, comprising: The Civil Aviation Company Operation and Maintenance Model Building Module is used to build fleet management sub-models, standby management sub-models, flight mission sub-models, and maintenance support sub-models based on the Civil Aviation Company's fleet operation data. The network initialization module is used to initialize the distributed learner and distributed evolver. The maintenance action decision module is used to formulate maintenance strategies based on the current operation and maintenance model of civil aviation companies, so as to avoid the aircraft fleet from becoming highly deteriorated and reduce system feedback costs. The replay buffer module is used to store the state-action value group for each step; The hierarchical training module is used to process the multiple maintenance costs output by the civil aviation company's operation and maintenance model in a hierarchical manner and output the total loss of the neural network. The model training module is used to train the distributed learner and distributed evolver using the total loss of the neural network output by the hierarchical training module.

[0016] The beneficial effects of this invention are that by proposing a multi-agent distributed hierarchical reinforcement learning-based maintenance decision-making method for civil aviation fleets, it can simultaneously consider the lifecycle management problems of multiple airports, multiple tasks, and multiple aircraft in the dynamic maintenance scheduling of large-scale aircraft fleets, and characterize complex constraints such as component degradation and resource coupling. In terms of modeling, this invention introduces a parameterized environment to realistically reproduce multi-level, multi-timescale maintenance decision-making scenarios. In terms of algorithm design, it adopts a two-level policy structure, combined with a shared communication protocol and a collaborative reward mechanism, effectively improving the coordination and learning efficiency among multiple agents. Furthermore, this invention verifies the key role of optimal agent synchronization and auxiliary reward mechanisms through ablation experiments, thereby ensuring the stability and efficiency of the algorithm. This invention exhibits superior performance compared to existing methods, while possessing strong versatility and promotional value. It is suitable for large-scale complex aviation maintenance scheduling problems, is easy to use, and has broad application prospects. Attached Figure Description

[0017] Figure 1 This is a comparison chart of the performance of the multi-body distributed hierarchical reinforcement learning algorithm of this invention; Figure 2 This is a Gantt chart of the fleet operation and maintenance strategy of the present invention; Figure 3 This is a flowchart of the civil aviation company operation and maintenance model of the present invention; Figure 4 This is a flowchart of the multi-body distributed hierarchical reinforcement learning algorithm of the present invention. Detailed Implementation

[0018] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of systems and methods consistent with some aspects of the invention as detailed in the appended claims.

[0019] The present invention will now be described in detail with reference to the accompanying drawings.

[0020] This invention provides a civil aviation fleet maintenance decision-making method based on multi-agent distributed hierarchical reinforcement learning, comprising the following steps: (1) Collect fleet operation data of civil aviation companies and establish operation and maintenance models of civil aviation companies.

[0021] See Figure 3 The civil aviation company's operation and maintenance model includes a fleet management sub-model, a backup engine management sub-model, a flight mission sub-model, and a maintenance support sub-model. Among them: The fleet management sub-model constructs a system model that can be used for operation and maintenance scheduling and resource optimization by characterizing multiple constraints of maintenance capabilities and resource allocation across multiple aircraft types, multiple airports, and multiple levels, combined with parking space capacity, unique aircraft location, maintenance transfer and spare engine allocation mechanisms.

[0022] The backup engine management sub-model constructs a system framework that covers dynamic inventory balancing, minimum safety reserve constraints, and cross-station allocation decisions by modeling specific constraints of engines for different aircraft models, centralized and responsive inventory layout, cross-airport dynamic allocation mechanisms, and multi-dimensional cost functions. This provides a core foundation for subsequent operation and maintenance scheduling optimization and decision support in multi-time periods and multi-uncertainty environments.

[0023] Under a given flight network and scheduling cycle, the flight mission sub-model ensures route feasibility and time continuity by defining segment tasks and aircraft-task matching constraints, and aims to maximize aircraft utilization, providing a scheduling framework for maintenance insertion, engine life management and airport capacity coordination.

[0024] The maintenance support sub-model depicts the engine maintenance process under the constraint of limited maintenance resources. It distinguishes between two maintenance methods: water washing and replacement of the spare engine. It ensures the feasibility of maintenance tasks through constraints on working hours, work positions, and manpower allocation. At the same time, it introduces maintenance costs, delay penalties, and engine state transition processes to comprehensively reflect the performance degradation and maintenance effects caused by flight cycles, providing a dynamic constraint mechanism for resource allocation and life management in fleet operation and maintenance scheduling.

[0025] (1.1) Fleet Management Sub-model: To simulate the operation and maintenance of a large airline's civil aviation fleet, this invention sets the fleet to consist of three typical aircraft types: wide-body aircraft for long-haul trunk routes (such as Boeing 787 and Airbus A350), single-aisle aircraft for medium and short-haul trunk routes (such as Boeing 737 and Airbus A320), and regional turboprop aircraft for regional routes (such as Bombardier Q400). Different aircraft types differ in maintenance requirements, range capability, operating frequency, and spare engine type. In airport modeling, airports are divided into two categories based on geographical location, throughput, and maintenance facility level: one category is a complete MRO airport with engine disassembly and assembly, avionics maintenance, and structural inspection capabilities; the other category can only perform light maintenance and troubleshooting, and cannot perform major engine overhauls. To characterize the capabilities of each airport, the following parameter variable is set: maximum number of parking spaces. Based on this, two types of constraints need to be satisfied: the first is the parking space capacity constraint:

[0026] To ensure that the available parking spaces at the airport are not overloaded; secondly, there is the unique position constraint for the aircraft:

[0027] Ensure that each aircraft is located at only one airport at any given time. Indicates the assembly of aircraft. This indicates meeting at the airport. Represents a time set, As an indicator variable, To specifically assign aircraft, To assign the airport where the aircraft is located. This refers to the specific time at that moment.

[0028] (1.2) Backup Engine Management Sub-model: In the modern civil aviation operation and maintenance system, backup engine management is a key link in ensuring stable flight operation and maintenance efficiency. When an aircraft engine enters a major maintenance cycle, it needs to be replaced by a backup engine to maintain flight missions. Different aircraft types (such as wide-body aircraft, single-aisle mainline aircraft, and regional aircraft) are equipped with different engine models. These engines have characteristics such as strong model specificity, high procurement costs, and slow inventory transfer response. Therefore, in order to improve resource utilization efficiency and reduce flight delays, airlines usually adopt a "centralized + responsive" management model, that is, to centrally deploy a large number of backup engines at primary maintenance hubs, while maintaining a certain amount of dispersed inventory at secondary airports, in order to achieve a balance between transfer speed and transportation costs. Considering that engine transportation relies on specialized equipment, its cross-airport transfer costs and timeliness are often affected by geographical distance and flight density. In order to accurately characterize this process in the mathematical model, the fleet is set to consist of M different aircraft types, each type ∈ {1, 2, …, M} possess An aircraft. Each aircraft type is equipped with a specific type of engine, and the corresponding engine requirement is... (That is, the number of engines required per aircraft, usually 2), the number of engines required per aircraft is The total engine requirement for the fleet is:

[0029] Because engine maintenance cycles are much longer than those of other aircraft components, airlines need to equip each aircraft type with a dedicated spare engine. Its distribution across airports can be represented as follows:

[0030] in express Time machine model At the airport The inventory level. Over time, the inventory is affected by the dynamic balance of replacement, repair, and allocation, which can be described by the update equation:

[0031] in express Always keep the engine in reserve. Indicates repair Always keep backup engines in stock. express From the airport Transferred to The number of spare engines. To ensure operational safety, each airport must maintain a minimum safety stock of spare engines:

[0032] in For model The minimum guarantee factor, For the airport's aircraft type Demand volume. In cost modeling, airlines need to consider three types of costs simultaneously; the first is inventory cost:

[0033] in, Single unit model The inventory cost of spare engines reflects the capital tied up in spare engines and storage expenses; secondly, the allocation cost needs to be calculated. :

[0034] in, Model From the airport Transferred to The unit transportation cost.

[0035] The third factor is the cost of delays: (1) the transportation and management expenses incurred in transferring engines across airports; (2) the cost of delays.

[0036] in, For model The cost of delays caused by a shortage of each spare engine.

[0037] (1.3) Flight Mission Sub-model: Civil aviation fleets execute pre-set flight missions in each flight cycle. These missions are typically based on the company's scheduled flight network, including long-haul trunk missions (such as intercontinental or transoceanic flights), medium- and short-haul trunk missions (such as connections to major city hubs), and regional missions (serving small and medium-sized airports). Different missions are applicable to different aircraft types and have varying degrees of engine wear. Flight mission scheduling not only determines route operations but also directly impacts maintenance windows, engine life, and airport capacity. To effectively simulate flight operations, a flight mission model based on segment scheduling was established. A given scheduling cycle is set... within, share Mission segment numbered The mission for each flight segment is defined as follows:

[0038] in and These are the departure and arrival airports, respectively. For the planned takeoff time, For flight duration, The required aircraft type. Assume the airline has a total of... Airplane, assembly and define scheduling variables , indicating airplane Should the flight segment be executed? Each flight segment must be performed by one and only one aircraft that meets the aircraft type requirements, thus creating a mission allocation constraint:

[0039] To ensure time feasibility between missions, if the aircraft performs the flight segment first... Then it must satisfy:

[0040] in This represents the transfer and preparation time required for connecting flight segments. Furthermore, to ensure that an aircraft performs only one mission at a time, non-overlapping constraints must be introduced:

[0041] This model not only provides a structured task allocation framework for fleet operations, but also lays the foundation for subsequent maintenance schedule insertion, engine availability matching, and airport capacity coordination.

[0042] (1.4) Maintenance Support Sub-model: In the operation and maintenance of an airline fleet, the limited maintenance resources are a key factor affecting the feasibility and efficiency of scheduling. Engine maintenance mainly takes two forms: one is water washing, which can be completed directly at the airport; the other is replacing the spare engine, which requires hangar support, and the replaced engine needs to be sent to a maintenance depot for repair. Since the maintenance facilities at each airport are limited by physical and human resources conditions, such as the number of maintenance workstations, average working hours, number of workers and shift distribution, and the upper limit of the number of aircraft that can be maintained at the same time, the model needs to introduce resource constraints to ensure supply and demand balance. First, when a maintenance task is triggered, the aircraft must be on the ground:

[0043] in Indicates airplane At any moment Is the device in a repairable state? Repair tasks can only be scheduled after a valid trigger time and can only be executed once.

[0044]

[0045] in Indicates airplane At the airport In time Is maintenance type performed? , For maintenance trigger indication function, This is a variable indicating the aircraft's position.

[0046] Airport maintenance capacity is subject to resource constraints:

[0047] in Repair type For the required average working hours, For the airport For repair type Maximum working hours supply capacity This refers to the duration of the time period. To avoid overloading the maintenance area, workstation capacity constraints are also required.

[0048] in Airport The maximum number of maintenance workstations.

[0049] In terms of cost modeling, maintenance costs are defined as:

[0050] in Indicates airplane Perform maintenance type The unit cost. If the repair time limit is exceeded. Failure to complete the task on time will result in a delay penalty:

[0051] in Indicates airplane In time The timing of the maintenance request triggering This is the delay penalty coefficient.

[0052] (2) See Figure 4 Based on multiple system processes across multiple computers, or by establishing distributed learners and distributed evolvers, the distributed evolvers generate initial fleet engine states based on a stochastic civil aviation company operation and maintenance model. Develop initial maintenance procedures .

[0053] (2.1) Based on multiple system processes of multiple computers, each process establishes a reinforcement learning evolver and a reinforcement learning learner, and these processes are connected through the local CPU to establish a distributed evolver and a distributed learner.

[0054] The reinforcement learning evolver consists of a policy neural network whose input is the fleet engine state of the civil aviation company's operations and maintenance model. The output is a maintenance action. .

[0055] A reinforcement learning learner consists of a value neural network and several auxiliary networks. The input to the value neural network is the fleet engine status of the civil aviation company's operation and maintenance model. Maintenance actions System feedback Next state The output is the value neural network's evaluation of the fleet's engine status. With policy probability distribution The number of auxiliary networks equals the number of airports, and the input is the fleet engine status. The output is the predicted local return. .

[0056] (2.2) The distributed evolver generates the initial fleet engine state based on the fleet parameters of the civil aviation company's operation and maintenance model. Develop initial maintenance procedures .

[0057] (3) The distributed evolver is based on the initial fleet engine state. and initial maintenance actions Interacting with the aforementioned civil aviation company's operation and maintenance model, a set of state-action value groups is obtained: system state (i.e., fleet engine state), maintenance actions, system feedback costs, and the next system state; a replay buffer is established to store the state-action value groups for each step. Store the data in the replay buffer; select actions in the replay buffer and use them as a training set for training; update the policy network of the learner and the value network of the evolver; update the training set; determine whether the system running step size has reached the set system training step size. If it has, stop training, obtain the trained network, and output the civil aviation company's maintenance optimization strategy. Otherwise, continue to select actions and repeat training until the running step size reaches the training step size.

[0058] (3.1) The distributed evolver is first based on the initial fleet engine state. With initial maintenance actions By interacting with the civil aviation company's operation and maintenance model, a set of state-action-value quadruples is obtained, which includes system status, maintenance actions, system feedback costs, and the next system status. Subsequently, the quadruple is stored in a distributed replay buffer, which is used to store trajectory data generated by multiple evolvers at different time steps, so that the distributed learner can uniformly sample and train it.

[0059] (3.2) Each learner in the distributed learner extracts training data in batches based on the replay buffer established in step (3.1), and applies the data to the policy according to the value network. Prediction of state under guidance Calculate the advantage function :

[0060] in, The attenuation coefficient is... This is the system feedback cost function.

[0061] The loss function of a value neural network is calculated using the following formula:

[0062] in: For the dominant function, To maintain the policy probability distribution function, For value network prediction, For the value function loss weights, Let be the mathematical expectation.

[0063] In the test environment of this invention, an auxiliary network is introduced for each airport. The input is the local fleet engine status of airport i. The output is the predicted local reward. The auxiliary network is not directly used for decision-making, but rather acts as an evaluator, learning the mapping between airport-level feedback signals and states. The auxiliary network does not directly participate in policy execution, but is added to the loss function as an auxiliary learning signal. The loss function for the auxiliary neural network is:

[0064] For the airport Local engine state, For the airport The maintenance cost; This is the predicted value for the auxiliary neural network.

[0065] During training, the learner's neural network loss function is a combination of the main task loss and the losses of each auxiliary network:

[0066] The weights for the auxiliary loss are used to control the balance between global and local modeling.

[0067] (3.3) The distributed learner learns from experience by building multiple policy networks and periodically scores the prediction results of all policy networks:

[0068] Indicates from time step Time to step Between, the learner's cumulative reward, i.e. the first The scores of each learner are used to compare performance among multiple learners. A higher score indicates that the learner has learned a better policy over a period of time. The starting time step for score calculation; The final time step for calculating the score. This refers to the reward at time step t. The discount factor γ aims to reduce the impact of future rewards on the current decision. Indicates from state to state The change in the predicted value of the policy network is used to measure the difference in predictions between different states.

[0069] After periodically scoring all policy networks, the parameters of the policy network with the highest score are synchronized to all policy networks.

[0070] In a distributed evolver, neural network updates calculate the loss function using the following formula:

[0071] And it is updated using gradient descent through a neural network.

[0072] Repeat steps (3.1)-(3.3) to determine whether the system running step size has reached the set system training step size. If it has, stop training, obtain the trained network, and output the civil aviation company's maintenance optimization strategy. Otherwise, continue to select actions and repeat training until the running step size reaches the training step size.

[0073] This embodiment also provides a civil aviation fleet maintenance decision-making system based on multi-entity distributed hierarchical reinforcement learning, which is used to implement the above embodiments. The terms "module," "unit," etc., used below refer to combinations of software and / or hardware that perform predetermined functions.

[0074] This embodiment provides a civil aviation fleet maintenance decision-making system based on multi-body distributed hierarchical reinforcement learning, including: The Civil Aviation Company Operation and Maintenance Model Building Module is used to build fleet management sub-models, standby management sub-models, flight mission sub-models, and maintenance support sub-models based on the Civil Aviation Company's fleet operation data. The network initialization module is used to initialize the distributed learner and distributed evolver. The maintenance action decision module is used to formulate maintenance strategies based on the current operation and maintenance model of civil aviation companies, so as to avoid the aircraft fleet from becoming highly deteriorated and reduce system feedback costs. The replay buffer module is used to store the state-action value group for each step; The hierarchical training module is used to process the multiple maintenance costs output by the civil aviation company's operation and maintenance model in a hierarchical manner and output the total loss of the neural network. The model training module is used to train the distributed learner and distributed evolver using the total loss of the neural network output by the hierarchical training module.

[0075] Example 1: An embodiment of the present invention was implemented on a machine equipped with an Intel(R) Xeon(R) Gold 6330 CPU, an NVIDIA GTX3090 graphics processor, and 128GB of memory. Using the above implementation method, a fleet model of 10 aircraft at 2 airports was established. Figure 1 The left figure illustrates the performance of this method compared to A2C, PPO, RPPO, TRPO, and rule-based maintenance aggregation comparison algorithms. This method significantly reduces fleet maintenance costs compared to these comparison algorithms. Figure 1 The figure on the right shows a box plot of the average performance of this method compared to the comparison algorithm after training and five actual applications. This method has significantly lower maintenance costs in terms of upper bound, lower bound, and mean.Figure 2 This paper demonstrates the specific maintenance strategies proposed by this method for an aircraft fleet. Results show that this method can accurately model the fleet mathematically and correctly formulate reasonable maintenance strategies for each aircraft's engine.

[0076] The above embodiments are only used to illustrate the design concept and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made based on the principles and design ideas disclosed in the present invention are within the protection scope of the present invention.

Claims

1. A civil aviation fleet maintenance decision-making method based on multi-body distributed hierarchical reinforcement learning, characterized in that, The method comprises the following steps: Collecting civil aviation company fleet operation data to establish a civil aviation company operation and maintenance model; Based on a plurality of computer-based system processes or establishing a distributed learner and a distributed evolutioner, the distributed evolutioner generates an initial fleet engine state based on a random civil aviation company operation and maintenance model , formulating an initial maintenance action ; The distributed evolver is based on the initial fleet engine state. and initial maintenance actions Interacting with the civil aviation company's operation and maintenance model, a set of state-action value groups is obtained: fleet engine status. Maintenance actions System feedback cost and the next step in fleet engine status Establish a replay buffer and group the state-action value of each step. Store the data in the replay buffer; select actions in the replay buffer to use as a training set for training. Updating the policy network of the learner and the value network of the evolutioner; Updating the training set; Determining whether the system running step length reaches the set system training step length, if yes, stopping training, obtaining the trained network, and outputting the civil aviation company maintenance optimization strategy, otherwise, continuing to select actions and repeating training until the running step length reaches the training step length.

2. The method of claim 1, wherein, The civil aviation company operation and maintenance model comprises a fleet management sub-model, a spare engine management sub-model, a flight task sub-model, and a maintenance support sub-model; The fleet management sub-model is constructed by setting up a multi-aircraft type combination to build an operation and maintenance difference, wherein the aircraft type parameters cover maintenance requirements, flight range capabilities, and spare engine types; the daily maintenance capacity, aircraft type support range, inventory, and allocation speed are set by airport distribution and maintenance facility capacity; the maintenance transfer cost and the spare engine concentration-allocation mechanism are introduced; and a high-dimensional system network model satisfying the airport capacity constraint and the aircraft unique position constraint is output, which is used for maintenance scheduling optimization and inventory configuration analysis; The spare engine management model models different aircraft engine requirements by aircraft specificity constraints, wherein the inventory variables are constrained by dynamic balance equations of inventory consumption, repair return flow, and cross-airport allocation; each airport needs to meet the minimum safety inventory threshold and comprehensively considers inventory cost, delay cost, and allocation cost, and outputs a dynamic management strategy that satisfies inventory configuration and allocation optimization in a multi-aircraft and multi-airport environment; The flight task model models the route task by the way of flight segment scheduling, wherein the flight segment task is defined by the departure airport, the arrival airport, the takeoff time, the flight duration, and the required aircraft type; the scheduling variables are limited by the task unique assignment constraint and the time feasibility constraint to ensure that the aircraft continuous task does not overlap and the flight segment connection is reasonable; The maintenance support model models the fleet maintenance through maintenance resource constraints and state transition process, wherein the maintenance task is limited by ground state and single trigger, the airport maintenance capability is constrained by working hours, working position and worker configuration, the maintenance cost is composed of maintenance fee, delay penalty and additional loss, the engine state decreases with flight cycle and is restored to the optimal state with a repair probability, and the state transition function In combination with the failure rate and repair rate output, a dynamic support optimization mechanism is formed which covers flight cycle accumulation and maintenance intervention effect.

3. The method of claim 1, wherein, The distributed learner learns experience by establishing multiple policy networks, and scores all policy network prediction results regularly: ; represents the cumulative reward of the learner between time step and time step ; is the reward at time step t; the discount factor γ is used to let the future rewards decrease the impact on the current decision; represents the predicted value change of the policy network from state to state , which measures the prediction difference of the model between different states; After scoring all policy networks regularly, the parameters of the policy network with the highest score are synchronized to all policy networks.

4. The method of claim 1, wherein, The civil aviation company operation and maintenance model outputs the total maintenance cost and the maintenance cost of each airport based on its multi-airport configuration.

5. The method of claim 4, wherein, When receiving multiple output results of the civil aviation company operation and maintenance model, for each airport, an auxiliary network is introduced , the input is the local state of the airport , and the output is the predicted local return.

6. The method of claim 5, wherein, A total neural network loss function is constructed based on multiple auxiliary networks; the auxiliary network is used to learn the mapping relationship between the airport-level feedback signal and the state, and is added to the loss function as an auxiliary learning signal: ; wherein, is the total loss, is the main task loss, is the ith auxiliary loss; is the weight of the auxiliary loss, used to control the balance between global and local modeling.

7. A multi-agent based distributed hierarchical reinforcement learning system for civil aviation fleet maintenance decision making implementing the method of claim 1, characterized in that, It comprises: A civil aviation company operation and maintenance model construction module for constructing a fleet management sub-model, a spare engine management sub-model, a flight task sub-model, and a maintenance support sub-model based on civil aviation company fleet operation data; A network initialization module for initializing a distributed learner and a distributed evolutioner; A maintenance action decision module for formulating a maintenance strategy according to the current civil aviation company operation and maintenance model to avoid the occurrence of a highly deteriorated state of the aircraft fleet and reduce the system feedback cost; A replay buffer module for storing the state-action-value set of each step; A hierarchical training module for hierarchically processing multiple maintenance costs output by the civil aviation company operation and maintenance model and outputting the total loss of the neural network; The model training module is configured to train the distributed learner and the distributed evolutioner by using the total loss of the neural network output by the hierarchical training module.

Citation Information

Cited By

  • Method and system for predictive assignment decision optimization of air fleet flight missions

    CN122175302A