Aircraft Reinforcement Learning Predictive Maintenance Decision-making Method under Uncertain Remaining Life

Through the adaptive dual-value distribution reinforcement learning method, an aircraft maintenance decision model is built, which solves the problem of "over-maintenance" or "under-maintenance" in regular maintenance, improves the reliability of aircraft predictive maintenance decisions, and optimizes the utilization of maintenance resources.

CN120013530BActive Publication Date: 2025-06-20HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510495315.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-20
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

Among the existing aircraft maintenance decision-making methods, regular maintenance has the problem of "over-maintenance" or "under-maintenance", and the reinforcement learning method has poor reliability in maintenance decision-making when facing the uncertainty of remaining life and overestimation problems.

Method used

Adaptive two-value distribution reinforcement learning method is adopted, and network parameters are optimized to improve decision-making reliability by constructing the first value distribution reinforcement learning model and the second value distribution reinforcement learning model, combining flight mission sequence data.

Benefits of technology

It improves the reliability of aircraft predictive maintenance decisions, reduces the impact of overestimation problems, can handle remaining life uncertainty more accurately, and optimizes the utilization of maintenance resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013530B_ABST
    Figure CN120013530B_ABST
Patent Text Reader

Abstract

The present invention discloses an aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life, and the method includes: generating a training set and a test set according to flight mission sequence data; respectively constructing a first value distribution reinforcement learning model and a second value distribution reinforcement learning model; selecting a flight mission sequence from the training set and inputting it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; calculating the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; updating the network parameters of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model; after the updated first value distribution reinforcement learning model is trained, predicting the aircraft maintenance decision through the trained first value distribution reinforcement learning model, which can improve the reliability of the determined aircraft maintenance decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aircraft maintenance decision-making, and in particular to an aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life. Background Art

[0002] Maintenance is an important part of aircraft operation management. Timely and effective maintenance is a necessary condition to ensure flight safety and reduce operation costs.

[0003] Existing aircraft maintenance decision-making methods are mostly preventive maintenance. The preventive maintenance method means arranging aircraft maintenance according to fixed date intervals or flight hours. Due to the fixed maintenance interval, this method often has problems of "over-maintenance" or "under-maintenance", resulting in waste of maintenance resources or safety hazards. To solve the above preventive maintenance problems, some researchers have combined remaining life prediction with aircraft maintenance decision-making and proposed a predictive maintenance decision-making method. Based on the prediction results of the remaining life, this method quantifies a series of maintenance decision-related factors such as flight missions, faults, maintenance, and spare parts warehousing into increases and decreases in benefits, thereby constructing a mathematical model of maintenance timing and benefits, and then solving for an appropriate future maintenance timing. Current aircraft predictive maintenance decision-making can be divided into dynamic programming methods, heuristic algorithm methods, and reinforcement learning methods. The dynamic programming method and the heuristic method are aimed at a deterministic environment where the flight mission sequence is completely known, constructing a function of maintenance timing and benefits, and then optimizing and solving for the maintenance timing. The reinforcement learning method is aimed at an uncertain environment where only the current flight mission is known, constructing an agent to learn the association rules between maintenance timing and benefits, and then outputting maintenance behaviors. Since most actual maintenance decisions are in an uncertain environment, the currently commonly used aircraft predictive maintenance decision-making method is the reinforcement learning method.

[0004] The reinforcement learning method has achieved certain results in the field of aircraft maintenance decision-making and has achieved relatively satisfactory maintenance decision-making effects. Most current aircraft predictive maintenance decision-making methods use the deterministic remaining life prediction results as input. However, due to the noise in the collection of monitoring data and the performance limitations of the remaining life prediction model, the actually obtained remaining life prediction results are not accurate, resulting in uncertainty of the remaining life and affecting the maintenance decision-making effect. In addition, the greedy algorithm used in the construction of the current reinforcement learning method will cause overestimation problems, making the reinforcement learning agent may make non-optimal maintenance decisions, and the reliability of the made maintenance decisions is poor. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide an aircraft reinforcement learning predictive maintenance decision-making method and device under uncertain remaining life, which can solve the above problems existing in the prior art.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] An embodiment of the present invention provides an aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life, and the method includes:

[0008] Generating a training set and a test set based on flight mission sequence data, where each flight mission sequence includes multiple flight missions;

[0009] Constructing a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively;

[0010] Selecting a flight mission sequence from the training set and inputting it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision;

[0011] Calculating the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model;

[0012] Updating the network parameters of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model;

[0013] After the updated first value distribution reinforcement learning model is trained, predicting the aircraft maintenance decision through the trained first value distribution reinforcement learning model.

[0014] Optionally, the step of updating the network parameters of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model includes:

[0015] Calculating the loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculating the adaptive update interval based on the loss function;

[0016] When the adaptive update interval is satisfied, updating the network parameters of the first value distribution reinforcement learning model to the network parameters of the second value distribution reinforcement learning model.

[0017] Optionally, both the first value distribution reinforcement learning model and the second value distribution reinforcement learning model include: a model input layer dimension, a number of hidden layers, a hidden layer dimension, and an output layer dimension;

[0018] The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed as an adaptive double value distribution reinforcement learning model based on a specific update rule, where the specific update rule is an adaptive update interval related to the training error.

[0019] Optionally, the step of calculating the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model includes:

[0020] Estimate the future maximum benefit corresponding to the aircraft maintenance decision through the second value distribution reinforcement learning model;

[0021] Calculate the current maintenance decision benefit corresponding to the aircraft maintenance decision based on preset rules;

[0022] Calculate the long-term benefit corresponding to the aircraft maintenance decision based on the current maintenance decision benefit, the future maximum benefit, and the discount factor of the future benefit.

[0023] Optionally, the steps of calculating the loss function of the second value distribution reinforcement learning model based on the long-term benefit and calculating the adaptive update interval based on the loss function include:

[0024] Calculate the deviation between the long-term benefit and the future maximum benefit in the form of KL divergence as the loss function of the second value distribution reinforcement learning model;

[0025] Calculate the adaptive update interval based on the loss function.

[0026] An embodiment of the present invention also provides an aircraft reinforcement learning predictive maintenance decision-making device under uncertain remaining life. Among them, the device includes:

[0027] A generation module for generating a training set and a test set based on flight mission sequence data, where each flight mission sequence includes multiple flight missions;

[0028] A construction module for respectively constructing a first value distribution reinforcement learning model and a second value distribution reinforcement learning model;

[0029] A decision generation module for selecting a flight mission sequence from the training set and inputting it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision;

[0030] A calculation module for calculating the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model;

[0031] An update module for updating the network parameters of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model;

[0032] A prediction module for predicting the aircraft maintenance decision through the trained first value distribution reinforcement learning model after the updated first value distribution reinforcement learning model is trained.

[0033] Optionally, the update module includes:

[0034] The first sub-module is used to calculate the loss function of the second value distribution reinforcement learning model based on the long-term return, and calculate the adaptive update interval based on the loss function.

[0035] The second sub-module is used to update the network parameters of the first value distribution reinforcement learning model to the network parameters of the second value distribution reinforcement learning model when the adaptive update interval is satisfied.

[0036] Optionally, both the first value distribution reinforcement learning model and the second value distribution reinforcement learning model include: the dimension of the model input layer, the number of hidden layers, the dimension of the hidden layers, and the dimension of the output layer.

[0037] The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed into an adaptive double value distribution reinforcement learning model based on a specific update rule, where the specific update rule is an adaptive update interval related to the training error.

[0038] Optionally, the calculation module includes:

[0039] The third sub-module is used to estimate the future maximum return corresponding to the aircraft maintenance decision through the second value distribution reinforcement learning model.

[0040] The fourth sub-module is used to calculate the current maintenance decision return corresponding to the aircraft maintenance decision based on a preset rule.

[0041] The fifth sub-module is used to calculate the long-term return corresponding to the aircraft maintenance decision based on the current maintenance decision return, the future maximum return, and the discount factor of the future return.

[0042] Optionally, the first sub-module is specifically used for:

[0043] Calculating the deviation between the long-term return and the future maximum return in the form of KL divergence as the loss function of the second value distribution reinforcement learning model.

[0044] Calculating the adaptive update interval based on the loss function.

[0045] An embodiment of the present invention further provides an electronic device, which is characterized in that it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete mutual communication through the communication bus; the memory is used to store a computer program; when the processor executes the program stored on the memory, it implements the process of any of the above aircraft reinforcement learning predictive maintenance decision methods with uncertain remaining life.

[0046] The aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life disclosed in the embodiments of the present invention generates a training set and a test set based on flight mission sequence data; constructs a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; selects a flight mission sequence from the training set and inputs it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; calculates the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; updates the network parameters of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model; after the updated first value distribution reinforcement learning model is trained, predicts the aircraft maintenance decision through the trained first value distribution reinforcement learning model. The aircraft maintenance decision prediction scheme provided by the embodiments of the present invention constructs two independent value distribution reinforcement learning models, adopts a dual learning framework, and updates based on an adaptive interval related to the training error, realizing the construction of an adaptive dual value distribution reinforcement learning model, which can improve the reliability of aircraft predictive maintenance decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 FIG. is a flowchart showing the steps of an aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life according to an embodiment of the present application;

[0048] Figure 2 FIG. is a schematic diagram of a calculation method for calculating the long-term benefit distribution by projected Bellman update;

[0049] Figure 3 FIG. is a schematic flowchart showing another aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life according to an embodiment of the present application;

[0050] Figure 4 FIG. is a distribution diagram of task benefits and durations of the training set;

[0051] Figure 5 FIG. is a distribution diagram of task benefits and durations of the test set;

[0052] Figure 6 FIG. is a schematic diagram of the change of benefits and adaptive update intervals during the training process;

[0053] Figure 7 FIG. is a schematic diagram of the prediction result of an aircraft reinforcement learning predictive maintenance decision-making model under uncertain remaining life;

[0054] Figure 8 FIG. is a schematic diagram of the result of a fixed update interval model;

[0055] Figure 9 FIG. is a schematic diagram of the result of a fixed update interval and without using a value distribution reinforcement learning model;

[0056] Figure 10Schematic diagram of the result without using the dual learning framework model;

[0057] Figure 11 Schematic diagram of the result comparison in a lower uncertainty scenario;

[0058] Figure 12 Schematic diagram of the result comparison in a benchmark uncertainty scenario;

[0059] Figure 13 Schematic diagram of the result comparison in a higher uncertainty scenario;

[0060] Figure 14 Schematic diagram showing the structure of an aircraft reinforcement learning predictive maintenance decision-making device under remaining life uncertainty according to an embodiment of the present application. Detailed implementation manners

[0061] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0062] The following will, with reference to the accompanying drawings, through specific embodiments and their application scenarios, elaborate in detail on the aircraft reinforcement learning predictive maintenance decision-making method under remaining life uncertainty provided by the embodiments of the present application.

[0063] As shown in the appendix Figure 1 The aircraft reinforcement learning predictive maintenance decision-making method under remaining life uncertainty according to the embodiments of the present application includes the following steps:

[0064] Step 101: Generate a training set and a test set based on flight mission sequence data.

[0065] Among them, each flight mission sequence contains multiple flight missions.

[0066] The training set contains multiple training data for training the value distribution reinforcement learning model to be constructed subsequently, and the test set contains multiple test data for testing the trained value distribution reinforcement learning model to determine the prediction result accuracy of the trained model.

[0067] The specific number of flight mission sequences and the number of tasks included in each flight mission sequence can be flexibly set by those skilled in the art, and no specific limitation is made in the embodiments of the present application. The training set and the test set are divided in any appropriate proportion, for example, divided in a 1:10 ratio, that is, if the flight mission sequence contains 110 flight mission sequences, the training set can be set to contain 10 flight mission sequences and the test set to contain 1000 flight mission sequences.

[0068] Step 102: Construct a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively.

[0069] Both the first value distribution reinforcement learning model and the second value distribution reinforcement learning model include: the dimension of the model input layer, the number of hidden layers, the dimension of the hidden layers, and the dimension of the output layer. The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed into an adaptive double value distribution reinforcement learning model based on a specific update rule, where the specific update rule is an adaptive update interval related to the training error.

[0070] The basic idea of value distribution reinforcement learning is as follows:

[0071] Due to the uncertainty of the remaining life, the results of maintenance decisions may be uncertain, which leads to the benefits of maintenance decisions showing a distribution form. Traditional reinforcement learning methods can only handle the deterministic benefits of maintenance decisions and perform poorly in the case of distributed benefits. To overcome this problem, researchers have proposed value distribution reinforcement learning.

[0072] In value distribution reinforcement learning, the agent learns the long-term benefit distribution of each action and the environment through the Bellman equation .

[0073]

[0074] Since the benefit is represented in a distribution form, the update strategy of value distribution reinforcement learning is different from that of traditional reinforcement learning. The difference lies in two aspects: the Bellman equation and the loss function.

[0075] Value distribution reinforcement learning represents the benefit in a distribution form, which brings the problem that when calculating the long-term reward using Bellman update, the future benefit multiplied by the discount factor has a non-overlapping support set with the current distributed benefit. Therefore, value distribution reinforcement learning needs to calculate the long-term benefit distribution through projected Bellman update, and its calculation method is as Figure 2 shown. In the embodiments of this application, a value distribution reinforcement learning model is constructed based on the basic idea of value distribution reinforcement learning.

[0076] In the embodiments of this application, a dual learning framework is adopted, and the basic principle of the dual learning framework is as follows:

[0077] Traditional reinforcement learning uses the maximum value of the estimated value to replace the estimation of the maximum value, resulting in the problem of overestimating the behavior benefit during decision-making, affecting the learning speed of the agent, and having the risk of selecting a non-optimal strategy. For example, there is a state with different actions, and their true reward is all zero. However, due to the random initialization of the agent, the estimated benefit may be greater than or less than zero, where the maximum overestimation value is around zero, while the estimated value of the maximum value is above zero, that is, the overestimation problem. This can be corrected during the training process because as long as there is enough search time, the state can be selected All possible operations in

[0078] Double Q - learning adopts a cross - validation method to decouple Q - value update and search. This decoupling process is achieved by maintaining two independent Q - value tables, namely and . is used for searching and selecting the best action that reaches the maximum estimated Q - value, and updating with the true Q - value of each selected action. When certain conditions (including the number of updates, probability, etc.) are met, the Q - value in will be updated to the Q - value in

[0079] Step 103: Select a flight mission sequence from the training set and input it into the first value - distribution reinforcement learning model to generate an aircraft maintenance decision.

[0080] Steps 103 to 105 are the training process of the first value - distribution reinforcement learning model and the second value - distribution reinforcement learning model based on the training set. When training, select the flight mission sequences that have not been trained from the training set. For each flight mission among them, with the maintenance environment as the input, the first value - distribution reinforcement learning model determines the aircraft predictive maintenance action (i.e., the aircraft maintenance decision). Among them, the maintenance environment includes but is not limited to: the current remaining life, flight mission revenue, flight mission duration, the number of spare parts, etc.

[0081] Step 104: Calculate the long - term revenue corresponding to the aircraft maintenance decision based on the second value - distribution reinforcement learning model.

[0082] An optionally way to calculate the long - term revenue corresponding to the aircraft maintenance decision based on the second value - distribution reinforcement learning model may include the following sub - steps:

[0083] Sub - step 1: Estimate the future maximum revenue corresponding to the aircraft maintenance decision through the second value - distribution reinforcement learning model;

[0084] Sub - step 2: Calculate the current maintenance decision revenue corresponding to the aircraft maintenance decision based on preset rules;

[0085] Sub - step 3: Calculate the long - term revenue corresponding to the aircraft maintenance decision based on the current maintenance decision revenue, the future maximum revenue, and the discount factor of future revenue.

[0086] This optionally way to calculate the long - term revenue corresponding to the aircraft maintenance decision has a small amount of calculation and high accuracy of calculation results.

[0087] Step 105: Update the network parameters of the first value distribution reinforcement learning model according to the long-term reward and the second value distribution reinforcement learning model.

[0088] In an optional embodiment, the method for updating the network parameters of the first value distribution reinforcement learning model according to the long-term reward and the second value distribution reinforcement learning model includes the following sub-steps:

[0089] Sub-step 1: Calculate the loss function of the second value distribution reinforcement learning model according to the long-term reward, and calculate the adaptive update interval based on the loss function;

[0090] Sub-step 2: When the adaptive update interval is satisfied, update the network parameters of the first value distribution reinforcement learning model to the network parameters of the second value distribution reinforcement learning model.

[0091] When the adaptive update interval is not satisfied, the network parameters of the first value distribution reinforcement learning model are not updated this time, and step 103 is returned to select an untrained flight task sequence from the training set to start the next round of training of the value distribution reinforcement learning model. The training of the value distribution reinforcement learning model is stopped until all the flight task sequences in the training set have been used for training through multiple iterative trainings.

[0092] Step 106: After the updated first value distribution reinforcement learning model is trained, predict the aircraft maintenance decision through the trained first value distribution reinforcement learning model.

[0093] It should be noted that after the first value distribution reinforcement learning model is trained, it can be tested through the flight task sequences in the test set to test the accuracy of the aircraft maintenance decision predicted by the first value distribution reinforcement learning model. When the prediction accuracy of the first value distribution reinforcement learning model meets the preset requirements, it can be applied to the prediction of subsequent aircraft maintenance decisions.

[0094] The aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life provided by the embodiments of the present application generates a training set and a test set based on flight mission sequence data; constructs a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; selects a flight mission sequence from the training set and inputs it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; calculates the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; updates the network parameters of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model; after the updated first value distribution reinforcement learning model is trained, predicts the aircraft maintenance decision through the trained first value distribution reinforcement learning model. The aircraft maintenance decision prediction method provided by the embodiments of the present application constructs two independent value distribution reinforcement learning models, adopts a dual learning framework, and updates based on an adaptive interval related to the training error, realizing the construction of an adaptive dual value distribution reinforcement learning model, which can improve the reliability of aircraft predictive maintenance decisions.

[0095] The following combines Figure 3 a specific example to illustrate the aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life provided by the present application.

[0096] As Figure 3 shown, the reinforcement learning aircraft maintenance decision-making method considering the uncertainty of remaining life provided in this specific example may include the following steps:

[0097] Step 1: Divide multiple flight mission sequence data into a training set and a test set in two parts.

[0098] The data used in this specific example is aircraft flight mission simulation data, which contains a total of 3,100 mission sequences with a length of 50 flight missions. Each flight mission includes flight mission benefits , flight mission duration information. Among them, the flight mission benefits are sampled from a uniform distribution with an upper limit of 100 unit benefits and a lower limit of 25 unit benefits, and the flight mission duration is sampled from a normal distribution with a mean of 20 unit time and a variance of 5 unit time squared.

[0099] The training set and the test set are divided in a ratio of 1:30, that is, the training set contains 100 mission sequences and the test set contains 3,000 mission sequences. The benefit and duration distributions of flight missions in the training set and the test set are as Figure 4 , Figure 5 shown.

[0100] Step 2: Construct the structure of the value distribution reinforcement learning model.

[0101] Among them, the value distribution reinforcement learning model includes: the dimension of the model input layer, the number of hidden layers, the dimension of the hidden layers, and the dimension of the output layer. Based on the designed structure, a first value distribution reinforcement learning model (referred to as Model One) and a second value distribution reinforcement learning model (referred to as Model Two) are constructed. The update rule is determined as the interval related to the training error, and an adaptive double value distribution reinforcement learning model is constructed.

[0102] Step Three: In the training set , select the flight mission sequences that have not been trained. For each flight mission among them, with the maintenance environment (including the current remaining life , the flight mission benefit , the flight mission duration , and the quantity of spare parts) as the input, Model One determines the aircraft predictive maintenance actions .

[0103]

[0104] Step Four: Calculate and generate the long-term benefit corresponding to the maintenance decision based on the current maintenance decision benefit and the future maximum benefit estimated by Model Two. Specifically, it can be calculated through the following formula:

[0105]

[0106] Among them, is the long-term benefit corresponding to the generated maintenance decision, is the current maintenance decision benefit, is the future maximum benefit estimated by Model Two, is the discount factor of the future benefit.

[0107] Calculate the deviation between the long-term benefit of the maintenance decision and the benefit estimated by Model Two in the form of KL divergence, and use it as the loss function of Model Two to train Model Two;

[0108]

[0109] Among them, is the long-term benefit of the maintenance decision, is the benefit estimated by Model Two.

[0110] Step Five: Calculate the adaptive update interval based on the loss function of Model Two.

[0111] In the actual training process, Model Two uses this quantity as the number of maintenance task intervals for covering its own neuron weights to Model One.

[0112]

[0113] For example, if the calculated number of task intervals is 2, after training Model 1 according to the third flight task sequence, update the network parameters of Model 1 using the network parameters of Model 2.

[0114] Step 6: After completing the training for all flight task sequences in the training set, use Model 1 in the test set to make a decision and obtain the aircraft predictive maintenance decision result.

[0115] In this specific example, there are 6 types of maintenance decisions that the aircraft can execute, namely execute the task and order spare parts, execute the task and do not order spare parts, wait for the next task and order spare parts, wait for the next task and do not order spare parts, perform maintenance and order spare parts, and perform maintenance without ordering spare parts. The impacts of each maintenance decision on revenue, remaining life, and the number of spare parts are shown in Table 1:

[0116] Table 1 Impacts of Each Maintenance Decision

[0117]

[0118] In this specific example, the initial expected remaining life is 40 unit times, and the uncertainty of the remaining life is expressed as that its actual value is a normal distribution with the expected value as the mean and 0.2 times the expected value as the variance. The initial number of spare parts is 5 units, 1 unit of spare parts is consumed for maintenance, the number of spare parts ordered each time is 1 unit, and the warehousing cost is 0.3 unit cost per unit of spare parts per unit time.

[0119] Train the aircraft reinforcement learning predictive maintenance decision-making model (i.e., Model 1 and Model 2) with uncertain remaining life constructed in the task sequences of the training set. The decision-making revenue and the calculated adaptive update interval for each training task sequence during the training process are as Figure 6 shown.

[0120] Use the trained aircraft reinforcement learning predictive maintenance decision-making model with uncertain remaining life to make decisions in the task sequences of the test set. Considering that the uncertainty of the remaining life is affected by the number, quality, and computing resources of sensors, the degree of uncertainty may fluctuate within a certain range. To evaluate the stability of the maintenance decision-making model (i.e., Model 1 and Model 2 after training), in this specific example, the maintenance decision-making model was tested under 3 different degrees of uncertainty: low, baseline, and high. Specifically, in the low uncertainty scenario, the variance of the remaining life distribution is 0.1 times the expected remaining life; in the baseline uncertainty scenario, the variance of the remaining life distribution is 0.2 times the expected remaining life; in the high uncertainty scenario, the variance of the remaining life distribution is 0.3 times the expected remaining life.

[0121] In scenarios with three levels of uncertainty, the results of the aircraft reinforcement learning predictive maintenance decision-making model under the uncertainty of remaining life are as follows Figure 7 shown.

[0122] The average returns of the aircraft reinforcement learning predictive maintenance decision-making model under the uncertainty of remaining life in scenarios with low uncertainty level, benchmark uncertainty level, and high uncertainty level are 1257, 1214, and 1119 respectively. The average return in the low uncertainty scenario is 3.54% higher than that in the benchmark uncertainty scenario, and the average return in the high uncertainty scenario is 7.83% lower than that in the benchmark uncertainty scenario.

[0123] To further demonstrate the decision-making effect of the aircraft reinforcement learning predictive maintenance decision-making method under the uncertainty of remaining life, this case compares the proposed method with three methods: fixed update interval, fixed update interval without using the value distribution reinforcement learning model, and without using the double learning framework.

[0124] In scenarios with three levels of uncertainty, the results of the fixed update interval model are as follows Figure 8 shown. The average returns of the fixed update interval model in scenarios with low uncertainty level, benchmark uncertainty level, and high uncertainty level are 1255, 1175, and 995 respectively. The average return in the low uncertainty scenario is 6.81% higher than that in the benchmark uncertainty scenario, and the average return in the high uncertainty scenario is 15.32% lower than that in the benchmark uncertainty scenario.

[0125] In scenarios with three levels of uncertainty, the results of the fixed update interval without using the value distribution reinforcement learning model are as follows Figure 9 shown. The average returns of the fixed update interval without using the value distribution reinforcement learning model in scenarios with low uncertainty level, benchmark uncertainty level, and high uncertainty level are 1203, 1045, and 855 respectively. The average return in the low uncertainty scenario is 15.12% higher than that in the benchmark uncertainty scenario, and the average return in the high uncertainty scenario is 18.19% lower than that in the benchmark uncertainty scenario.

[0126] In scenarios with three levels of uncertainty, the results of the model without using the double learning framework are as follows Figure 10 shown.

[0127] The average returns of the fixed update interval without using the value distribution reinforcement learning model in scenarios with low uncertainty level, benchmark uncertainty level, and high uncertainty level are 1108, 1034, and 845 respectively. The average return in the low uncertainty scenario is 7.16% higher than that in the benchmark uncertainty scenario, and the average return in the high uncertainty scenario is 18.28% lower than that in the benchmark uncertainty scenario.

[0128] In this specific example, the decision-making effects of the aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life and three comparison methods are compared under different degrees of uncertainty scenarios. The comparison results under the lower uncertainty scenario are as Figure 11 shown in Table 2.

[0129] Table 2: Results of the lower uncertainty scenario

[0130]

[0131] It can be seen from the comparison results that the proposed method has higher average revenue and lower quartile revenue than the comparison methods, and the upper quartile revenue and median revenue are close to the highest values in the comparison methods. Thus, it can be obtained that the proposed method has better effects than the comparison methods under the lower uncertainty scenario.

[0132] The comparison results of the proposed method and the three comparison methods under the benchmark uncertainty scenario are as Figure 12 shown in Table 3.

[0133] Table 3: Results of the benchmark uncertainty scenario

[0134]

[0135] It can be seen from the comparison results that the proposed method has higher average revenue, upper quartile revenue, median revenue, and lower quartile revenue than the comparison methods. Thus, it can be obtained that the proposed method has better effects than the comparison methods under the benchmark uncertainty scenario.

[0136] The comparison results of the proposed method and the three comparison methods under the higher uncertainty scenario are as Figure 13 shown in Table 4

[0137] Table 4: Results of the higher uncertainty scenario

[0138]

[0139] It can be seen from the comparison results that the proposed method has higher average revenue, upper quartile revenue, median revenue, and lower quartile revenue than the comparison methods. Thus, it can be obtained that the proposed method has better effects than the comparison methods under the higher uncertainty scenario.

[0140] To further evaluate the stability of the model when affected by different degrees of uncertainty, this case further compares the degree of change in the results of the proposed method and the three comparison methods under different uncertainty degree scenarios.

[0141] Table 5: Comparison of the average revenue change between different uncertainty degree scenarios

[0142]

[0143] It can be seen that in addition to having the highest average return in the low-uncertainty scenario, the benchmark uncertainty scenario, and the high-uncertainty scenario, the proposed method model also has the lowest average return change (increase / decrease) in the low-uncertainty scenario and the high-uncertainty scenario. That is to say, compared with the comparative method, the proposed method has the best and most stable performance and can better cope with uncertainties of different degrees.

[0144] The reinforcement learning aircraft maintenance decision-making method considering remaining life uncertainty provided by this specific example aims at the limitations of the current aircraft predictive maintenance decision-making method in the face of remaining life uncertainty and overestimation problems, and proposes a set of aircraft maintenance decision-making methods based on adaptive double-value distribution reinforcement learning. On the one hand, the value distribution reinforcement learning model adopted can describe the benefits of flight missions in the form of a distribution function to quantify the impact of remaining life uncertainty on maintenance decisions; on the second hand, the dual learning framework adopted can decouple the update and decision-making processes of the reinforcement learning model to reduce the impact of overestimation problems on maintenance decisions; on the third hand, the adopted adaptive update interval can dynamically adjust the update interval of the dual learning framework during the training process to more effectively reduce overestimation problems and improve the optimization efficiency of the value distribution reinforcement learning model.

[0145] Figure 14 It is a structural block diagram of an aircraft reinforcement learning predictive maintenance decision-making device under remaining life uncertainty according to an embodiment of the present application.

[0146] The aircraft reinforcement learning predictive maintenance decision-making device under remaining life uncertainty according to an embodiment of the present application includes the following functional modules:

[0147] A generation module 1401, configured to generate a training set and a test set according to flight mission sequence data, where each flight mission sequence includes multiple flight missions;

[0148] A construction module 1402, configured to respectively construct a first value distribution reinforcement learning model and a second value distribution reinforcement learning model;

[0149] A decision generation module 1403, configured to select a flight mission sequence from the training set and input it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision;

[0150] A calculation module 1404, configured to calculate the long-term return corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model;

[0151] An update module 1405, configured to update the network parameters of the first value distribution reinforcement learning model according to the long-term return and the second value distribution reinforcement learning model;

[0152] A prediction module 1406, configured to predict aircraft maintenance decisions by using the first value distribution reinforcement learning model obtained through training after the training of the first value distribution reinforcement learning model to be updated is completed.

[0153] Optionally, the update module includes:

[0154] A first sub-module, configured to calculate a loss function of the second value distribution reinforcement learning model according to the long-term reward, and calculate an adaptive update interval based on the loss function;

[0155] A second sub-module, configured to update network parameters of the first value distribution reinforcement learning model to network parameters of the second value distribution reinforcement learning model when the adaptive update interval is satisfied.

[0156] Optionally, both the first value distribution reinforcement learning model and the second value distribution reinforcement learning model include: a model input layer dimension, a number of hidden layers, a hidden layer dimension, and an output layer dimension;

[0157] The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed into an adaptive double value distribution reinforcement learning model based on a specific update rule, where the specific update rule is an adaptive update interval related to training error.

[0158] Optionally, the calculation module includes:

[0159] A third sub-module, configured to estimate a future maximum reward corresponding to the aircraft maintenance decision by using the second value distribution reinforcement learning model;

[0160] A fourth sub-module, configured to calculate a current maintenance decision reward corresponding to the aircraft maintenance decision based on a preset rule;

[0161] A fifth sub-module, configured to calculate a long-term reward corresponding to the aircraft maintenance decision based on the current maintenance decision reward, the future maximum reward, and a discount factor of the future reward.

[0162] Optionally, the first sub-module is specifically configured to:

[0163] Calculate a deviation between the long-term reward and the future maximum reward in a form of KL divergence as the loss function of the second value distribution reinforcement learning model;

[0164] Calculate the adaptive update interval based on the loss function.

[0165] The aircraft reinforcement learning predictive maintenance decision-making device under uncertain remaining life provided by the embodiment of the present application constructs two independent value distribution reinforcement learning models, adopts a dual learning framework, and is updated based on an adaptive interval related to the training error, realizing the construction of an adaptive dual value distribution reinforcement learning model, which can improve the reliability of aircraft predictive maintenance decision-making.

[0166] Provided by the embodiment of the present application Figure 14 The aircraft reinforcement learning predictive maintenance decision-making device under uncertain remaining life shown can implement Figure 1 each process implemented by the method embodiment. To avoid repetition, it will not be elaborated here.

[0167] The embodiment of the present invention also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus.

[0168] The memory is used to store a computer program;

[0169] The processor, when executing the program stored on the memory, implements the aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life shown in the above method embodiment.

[0170] The communication bus mentioned above for the terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0171] The communication interface is used for communication between the above terminal and other devices.

[0172] The memory can include a Random Access Memory (RAM), and can also include a non-volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.

[0173] In another embodiment provided by the present invention, a computer-readable storage medium is also provided. Instructions are stored in the computer-readable storage medium. When it runs on an electronic device, the electronic device is enabled to implement the aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life described in any of the above embodiments.

[0174] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which, when running on an electronic device, enables the electronic device to implement the aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life described in any one of the above embodiments.

[0175] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.

[0176] The above is the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A reinforcement learning predictive maintenance decision-making method for aircraft under uncertainty of remaining life, characterized in that: The method comprises: Generate a training set and a test set based on the flight mission sequence data, wherein each flight mission sequence includes multiple flight missions; Construct a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; Selecting a flight mission sequence from the training set and inputting it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; Calculating the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; Calculating a loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculating an adaptive update interval based on the loss function; When the adaptive update interval is met, updating the network parameters of the first value distribution reinforcement learning model to the network parameters of the second value distribution reinforcement learning model; After the training of the updated first value distribution reinforcement learning model is completed, the aircraft maintenance decision is predicted by the first value distribution reinforcement learning model obtained through training.

2. The method according to claim 1, characterized in that: The first value distribution reinforcement learning model and the second value distribution reinforcement learning model both include: model input layer dimension, number of hidden layers, hidden layer dimension and output layer dimension; The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed as adaptive dual-value distribution reinforcement learning models based on a specific update rule, wherein the specific update rule is an adaptive update interval related to a training error.

3. The method according to claim 1, characterized in that The step of calculating the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model includes: estimating the maximum future benefit corresponding to the aircraft maintenance decision by using the second value distribution reinforcement learning model; Calculate the current maintenance decision benefit corresponding to the aircraft maintenance decision based on preset rules; Based on the current maintenance decision benefit, the future maximum benefit and the discount coefficient of the future benefit, the long-term benefit corresponding to the aircraft maintenance decision is calculated.

4. The method according to claim 3, characterized in that The step of calculating the loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculating the adaptive update interval based on the loss function includes: Calculating the deviation between the long-term benefit and the future maximum benefit in the form of KL divergence as the loss function of the second value distribution reinforcement learning model; An adaptive update interval is calculated based on the loss function.

5. An aircraft reinforcement learning predictive maintenance decision-making device under uncertainty of remaining life, characterized in that: The device comprises: A generation module, used to generate a training set and a test set based on flight mission sequence data, wherein each flight mission sequence includes multiple flight missions; A construction module, used to construct a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; A decision generation module, used for selecting a flight mission sequence from the training set and inputting it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; A calculation module, configured to calculate the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; An update module, comprising a first submodule and a second submodule, wherein the first submodule is used to calculate the loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculate the adaptive update interval based on the loss function, and the second submodule is used to update the network parameters of the first value distribution reinforcement learning model to the network parameters of the second value distribution reinforcement learning model when the adaptive update interval is met; The prediction module is used to predict aircraft maintenance decisions by using the first value distribution reinforcement learning model obtained through training after the training of the first value distribution reinforcement learning model to be updated is completed.

6. The device according to claim 5, characterized in that: The first value distribution reinforcement learning model and the second value distribution reinforcement learning model both include: model input layer dimension, number of hidden layers, hidden layer dimension and output layer dimension; The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed as adaptive dual-value distribution reinforcement learning models based on a specific update rule, wherein the specific update rule is an adaptive update interval related to a training error.

7. The device according to claim 5, characterized in that The calculation module comprises: A third submodule is used to estimate the maximum future benefit corresponding to the aircraft maintenance decision through the second value distribution reinforcement learning model; A fourth submodule is used to calculate the current maintenance decision benefit corresponding to the aircraft maintenance decision based on a preset rule; The fifth submodule is used to calculate the long-term benefit corresponding to the aircraft maintenance decision based on the current maintenance decision benefit, the future maximum benefit and the discount coefficient of the future benefit.

8. The device according to claim 7, characterized in that The first submodule is specifically used for: Calculating the deviation between the long-term benefit and the future maximum benefit in the form of KL divergence as the loss function of the second value distribution reinforcement learning model; An adaptive update interval is calculated based on the loss function.

9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the aircraft reinforcement learning predictive maintenance decision-making method under uncertainty of remaining life as described in any one of claims 1 to 4 when executing the program stored in the memory.

Citation Information

Patent Citations

  • Aircraft maintenance path optimization method based on adaptive reinforcement learning

    CN115292959A

  • Reinforcement learning system for maintenance decision making

    EP4407523A1