Aircraft reinforcement learning predictive maintenance decision-making method under residual life uncertainty
By adopting an adaptive two-value distributed reinforcement learning method in aircraft maintenance decisions, the uncertainty problem in regular maintenance and overestimation problem in reinforcement learning are solved, and the reliability and accuracy of maintenance decisions are improved.
Patent Information
- Application Number
- CN202510495315.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Among the existing aircraft maintenance decision-making methods, regular maintenance has the problem of ‘over-maintenance’ or ‘under-maintenance’, and the reinforcement learning method is not effective in dealing with the problem of residual life uncertainty and overestimation.
Adaptive double-value distribution reinforcement learning method is adopted to optimize aircraft maintenance decisions by constructing two independent value distribution reinforcement learning models, and adopting a dual-learning framework and adaptive update interval.
It improves the reliability of aircraft predictive maintenance decisions, reduces overestimation problems, enhances the ability to handle residual life uncertainty, and improves the accuracy and efficiency of maintenance decisions.
Smart Images

Figure CN120013530A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aircraft maintenance decision-making, and in particular to an aircraft reinforcement learning predictive maintenance decision-making method under uncertainty of remaining life. Background Art
[0002] Maintenance is an important part of aircraft operation management. Timely and effective maintenance is a necessary condition to ensure flight safety and reduce operating costs.
[0003] Most of the existing aircraft maintenance decision-making methods are scheduled maintenance. Scheduled maintenance methods refer to scheduling aircraft maintenance according to fixed date intervals or flight hours. Due to the fixed maintenance intervals, this method often has the problem of "over-maintenance" or "under-maintenance", resulting in a waste of maintenance resources or safety hazards. In order to solve the above-mentioned scheduled maintenance problems, researchers have combined the remaining life prediction with aircraft maintenance decisions and proposed a predictive maintenance decision-making method. Based on the prediction results of the remaining life, this method quantifies a series of maintenance decision-making related factors such as flight missions, failures, maintenance, and spare parts storage as the increase or decrease of benefits, thereby constructing a mathematical model of maintenance timing and benefits, and then solving the appropriate future maintenance timing. Current aircraft predictive maintenance decisions can be divided into dynamic programming methods, heuristic algorithm methods, and reinforcement learning methods. Dynamic programming methods and heuristic methods construct a function of maintenance timing and benefits for a deterministic environment where the flight mission sequence is completely known, and then optimize the maintenance timing. Reinforcement learning methods construct an intelligent agent to learn the association rules between maintenance timing and benefits for an uncertain environment where only the current flight mission is known, and then output maintenance behavior. Since most actual maintenance decisions are made in an uncertain environment, the commonly used aircraft predictive maintenance decision-making method is the reinforcement learning method.
[0004] Reinforcement learning methods have achieved certain results in the field of aircraft maintenance decision-making and have achieved relatively satisfactory maintenance decision-making results. Most of the current aircraft predictive maintenance decision-making methods use deterministic remaining life prediction results as input. However, due to the monitoring data acquisition noise and the performance limitations of the remaining life prediction model, the actual remaining life prediction results obtained are not accurate, resulting in remaining life uncertainty and affecting the maintenance decision effect. In addition, the greedy algorithm used in the construction of the current reinforcement learning method will cause over-estimation problems, which may cause the reinforcement learning agent to make non-optimal maintenance decisions, and the maintenance decisions made are unreliable. Summary of the invention
[0005] The purpose of the embodiments of the present invention is to provide a method and device for aircraft reinforcement learning predictive maintenance decision-making under uncertainty of remaining life, which can solve the above-mentioned problems existing in the prior art.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: An embodiment of the present invention provides an aircraft reinforcement learning predictive maintenance decision-making method under uncertainty of remaining life, the method comprising: Generate a training set and a test set based on the flight mission sequence data, wherein each flight mission sequence includes multiple flight missions; Construct a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; Selecting a flight mission sequence from the training set and inputting it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; Calculating the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; According to the long-term benefit and the second value distribution reinforcement learning model, updating the network parameters of the first value distribution reinforcement learning model; After the training of the updated first value distribution reinforcement learning model is completed, the aircraft maintenance decision is predicted by the first value distribution reinforcement learning model obtained through training.
[0007] Optionally, the step of updating the network parameters of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model includes: Calculating a loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculating an adaptive update interval based on the loss function; When the adaptive update interval is met, the network parameters of the first value distribution reinforcement learning model are updated to the network parameters of the second value distribution reinforcement learning model.
[0008] Optionally, the first value distribution reinforcement learning model and the second value distribution reinforcement learning model both include: model input layer dimension, number of hidden layers, hidden layer temperature and output layer dimension; The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed as adaptive dual-value distribution reinforcement learning models based on a specific update rule, wherein the specific update rule is an adaptive update interval related to a training error.
[0009] Optionally, the step of calculating the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model includes: estimating the maximum future benefit corresponding to the aircraft maintenance decision by using the second value distribution reinforcement learning model; Calculate the current maintenance decision benefit corresponding to the aircraft maintenance decision based on preset rules; Based on the current maintenance decision benefit, the future maximum benefit and the discount coefficient of the future benefit, the long-term benefit corresponding to the aircraft maintenance decision is calculated.
[0010] Optionally, the step of calculating a loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculating an adaptive update interval based on the loss function includes: Calculating the deviation between the long-term benefit and the future maximum benefit in the form of KL divergence as the loss function of the second value distribution reinforcement learning model; An adaptive update interval is calculated based on the loss function.
[0011] The embodiment of the present invention further provides an aircraft reinforcement learning predictive maintenance decision-making device under uncertainty of remaining life, wherein the device comprises: A generation module, used to generate a training set and a test set based on flight mission sequence data, wherein each flight mission sequence includes multiple flight missions; A construction module, used to construct a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; A decision generation module, used for selecting a flight mission sequence from the training set and inputting it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; A calculation module, configured to calculate the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; An updating module, configured to update a network parameter of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model; The prediction module is used to predict aircraft maintenance decisions by using the first value distribution reinforcement learning model obtained through training after the training of the first value distribution reinforcement learning model to be updated is completed.
[0012] Optionally, the update module includes: A first submodule, configured to calculate a loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculate an adaptive update interval based on the loss function; The second submodule is used to update the network parameters of the first value distribution reinforcement learning model to the network parameters of the second value distribution reinforcement learning model when the adaptive update interval is met.
[0013] Optionally, the first value distribution reinforcement learning model and the second value distribution reinforcement learning model both include: model input layer dimension, number of hidden layers, hidden layer temperature and output layer dimension; The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed as adaptive dual-value distribution reinforcement learning models based on a specific update rule, wherein the specific update rule is an adaptive update interval related to a training error.
[0014] Optionally, the calculation module includes: A third submodule is used to estimate the maximum future benefit corresponding to the aircraft maintenance decision through the second value distribution reinforcement learning model; A fourth submodule is used to calculate the current maintenance decision benefit corresponding to the aircraft maintenance decision based on a preset rule; The fifth submodule is used to calculate the long-term benefit corresponding to the aircraft maintenance decision based on the current maintenance decision benefit, the future maximum benefit and the discount coefficient of the future benefit.
[0015] Optionally, the first submodule is specifically used for: Calculating the deviation between the long-term benefit and the future maximum benefit in the form of KL divergence as the loss function of the second value distribution reinforcement learning model; An adaptive update interval is calculated based on the loss function.
[0016] An embodiment of the present invention also provides an electronic device, characterized in that it includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; the processor is used to implement any of the above-mentioned aircraft reinforcement learning predictive maintenance decision-making methods under uncertain remaining life when executing the program stored in the memory.
[0017] The aircraft reinforcement learning predictive maintenance decision method under the uncertainty of remaining life disclosed in the embodiment of the present invention generates a training set and a test set according to the flight mission sequence data; constructs a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; selects a flight mission sequence from the training set and inputs it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; calculates the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; updates the network parameters of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model; after the training of the updated first value distribution reinforcement learning model is completed, the aircraft maintenance decision is predicted by the first value distribution reinforcement learning model obtained by training. The aircraft maintenance decision prediction scheme provided by the embodiment of the present invention constructs two independent value distribution reinforcement learning models, adopts a dual learning framework, and updates based on an adaptive interval related to the training error to realize the construction of an adaptive dual value distribution reinforcement learning model, which can improve the reliability of aircraft predictive maintenance decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart showing the steps of a method for predictive maintenance decision-making based on reinforcement learning of an aircraft under uncertainty of remaining life in an embodiment of the present application; Figure 2 Schematic diagram of the computational method for calculating the long-term return distribution for the projected Bellman update; Figure 3 is a flow chart of another aircraft reinforcement learning predictive maintenance decision method under uncertainty of remaining life according to an embodiment of the present application; Figure 4 It is the distribution diagram of training set task benefits and duration; Figure 5 This is the distribution diagram of the test set task benefits and duration; Figure 6 It is a schematic diagram of the training process benefits and the changes in the adaptive update interval; Figure 7 It is a schematic diagram of the prediction results of the aircraft reinforcement learning predictive maintenance decision model under the uncertainty of remaining life; Figure 8 It is a schematic diagram of the fixed update interval model results; Fig. 9 This is a schematic diagram of the results of the reinforcement learning model with a fixed update interval and without using value distribution; Fig.10 This is a schematic diagram of the model results without using the dual learning framework; Fig.11 It is a schematic diagram comparing the results of the lower uncertainty scenario; Fig.12 It is a schematic diagram of the comparison of the results of the benchmark uncertainty scenario; Fig.13 It is a schematic diagram comparing the results of higher uncertainty scenarios; Fig.14 It is a structural schematic diagram of an aircraft reinforcement learning predictive maintenance decision-making device under uncertain remaining life according to an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0020] In the following, in conjunction with the accompanying drawings, the aircraft reinforcement learning predictive maintenance decision-making method under uncertain remaining life provided by the embodiment of the present application is described in detail through specific embodiments and their application scenarios.
[0021] As attached Figure 1 As shown, the aircraft reinforcement learning predictive maintenance decision-making method under the uncertainty of remaining life in the embodiment of the present application includes the following steps: Step 101: Generate a training set and a test set based on the flight mission sequence data.
[0022] Each flight mission sequence includes multiple flight missions.
[0023] The training set contains multiple training data for training the subsequently constructed value distribution reinforcement learning model. The test set contains multiple test data for testing the trained value distribution reinforcement learning model to determine the accuracy of the prediction results of the trained model.
[0024] The specific number of flight mission sequences and the number of tasks contained in each flight mission sequence can be flexibly set by those skilled in the art, and no specific restrictions are imposed on this in the embodiments of the present application. The training set and the test set are divided in any appropriate ratio, such as a 1:10 ratio, that is, if the flight mission sequence contains 110 flight mission sequences, the training set can be set to contain 10 flight mission sequences and the test set to contain 1000 flight mission sequences.
[0025] Step 102: construct a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively.
[0026] The first value distribution reinforcement learning model and the second value distribution reinforcement learning model both include: model input layer dimension, hidden layer number, hidden layer temperature and output layer dimension. The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed as adaptive dual value distribution reinforcement learning models based on specific update rules, wherein the specific update rules are adaptive update intervals related to training errors.
[0027] The basic idea of value distribution reinforcement learning is as follows: Due to the uncertainty of remaining life, the results of maintenance decisions may be uncertain, which causes the benefits of maintenance decisions to be in the form of distribution. Traditional reinforcement learning methods can only handle the deterministic benefits of maintenance decisions and perform poorly when the benefits are in the form of distribution. To overcome this problem, researchers proposed value distribution reinforcement learning.
[0028] In value distribution reinforcement learning, the agent learns the long-term reward distribution of each action and environment through the Bellman equation .
[0029]
[0030] Since the benefits are expressed in a distributed form, the update strategy of value distribution reinforcement learning is different from that of traditional reinforcement learning. The difference lies in two aspects: the Bellman equation and the loss function.
[0031] Value distribution reinforcement learning represents the benefits in the form of distribution, which brings a problem. When Bellman update is used to calculate the long-term reward, the future benefits multiplied by the discount coefficient do not intersect with the current distributed benefits support set. Therefore, value distribution reinforcement learning needs to calculate the long-term benefit distribution through projected Bellman update. The calculation method is as follows: Figure 2In the embodiment of the present application, a value distribution reinforcement learning model is constructed based on the basic idea of value distribution reinforcement learning.
[0032] The dual learning framework is adopted in the embodiment of the present application. The basic principle of the dual learning framework is as follows: Traditional reinforcement learning uses the maximum value of the estimated value instead of the maximum value estimate, which leads to the overestimation of the behavior benefits when making decisions, affects the learning speed of the agent, and has the risk of selecting a non-optimal strategy. For example, there is a state have Different actions, their true value rewards are all zero. However, due to the random initialization of the agent, the estimated payoff It may be greater or less than zero, where the maximum overestimate is around zero and the maximum estimate is above zero, i.e., the overestimation problem. This can be corrected during training, because as long as there is enough search time, the state can be selected However, in traditional Q-learning, since search and update are coupled, the overestimation of the state reward is first propagated to the state before the update, and then decreases as the number of searches increases.
[0033] Dual Q learning uses a cross-validation method to decouple Q value update from search. This decoupling process is achieved by maintaining two independent Q value tables, namely and . It is used to search and select the best action that achieves the maximum estimated Q value, and update the true Q value of each selected action . When certain conditions (including update times, probability, etc.) are met, The Q value in will be updated to Thus, the overestimation problem in the search can be reduced before propagating the update.
[0034] Step 103: Select a flight mission sequence from the training set and input it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision.
[0035] Steps 103 to 105 are the process of training the first value distribution reinforcement learning model and the second value distribution reinforcement learning model based on the training set. During the training, an untrained flight mission sequence is selected from the training set, and for each flight mission, the maintenance environment is used as input, and the aircraft predictive maintenance action (i.e., aircraft maintenance decision) is determined by the first value distribution reinforcement learning model. The maintenance environment includes but is not limited to: the current remaining life, flight mission benefits, flight mission duration, the number of spare parts, etc.
[0036] Step 104: Based on the second value distribution reinforcement learning model, calculate the long-term benefits corresponding to the aircraft maintenance decision.
[0037] An alternative method for calculating the long-term benefits corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model may include the following sub-steps: Sub-step 1: Estimate the maximum future benefit corresponding to the aircraft maintenance decision through the second value distribution reinforcement learning model; Sub-step 2: calculating the current maintenance decision benefit corresponding to the aircraft maintenance decision based on preset rules; Sub-step three: Based on the current maintenance decision benefit, the future maximum benefit and the discount coefficient of the future benefit, the long-term benefit corresponding to the aircraft maintenance decision is calculated.
[0038] This optional method of calculating the long-term benefits corresponding to aircraft maintenance decisions has a small amount of calculation and a high accuracy of the calculation results.
[0039] Step 105: Update the network parameters of the first value distribution reinforcement learning model based on the long-term benefits and the second value distribution reinforcement learning model.
[0040] In an optional embodiment, the method of updating the network parameters of the first value distribution reinforcement learning model according to the long-term benefits and the second value distribution reinforcement learning model includes the following sub-steps: Sub-step 1: calculating the loss function of the second value distribution reinforcement learning model according to the long-term benefits, and calculating the adaptive update interval based on the loss function; Sub-step 2: While satisfying the adaptive update interval, the network parameters of the first value distribution reinforcement learning model are updated to the network parameters of the second value distribution reinforcement learning model.
[0041] If the adaptive update interval is not met, the network parameters of the first value distribution reinforcement learning model are not updated this time, and the process returns to step 103 to select an untrained flight mission sequence from the training set to start the next round of training for the value distribution reinforcement learning model. The training of the value distribution reinforcement learning model is stopped after multiple iterations of training until all the flight mission sequences in the training set are used for training.
[0042] Step 106: After the training of the updated first value distribution reinforcement learning model is completed, the aircraft maintenance decision is predicted by the trained first value distribution reinforcement learning model.
[0043] It should be noted that after the first value distribution reinforcement learning model is trained, it can be tested through the flight mission sequence in the test set to test the accuracy of the aircraft maintenance decision predicted by the first value distribution reinforcement learning model. When the prediction accuracy of the first value distribution reinforcement learning model meets the preset requirements, it can be applied to the prediction of subsequent aircraft maintenance decisions.
[0044] The aircraft reinforcement learning predictive maintenance decision method under the uncertainty of remaining life provided in the embodiment of the present application generates a training set and a test set based on the flight mission sequence data; constructs a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; selects a flight mission sequence from the training set and inputs it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; calculates the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; updates the network parameters of the first value distribution reinforcement learning model based on the long-term benefit and the second value distribution reinforcement learning model; after the training of the updated first value distribution reinforcement learning model is completed, the aircraft maintenance decision is predicted by the first value distribution reinforcement learning model obtained through training. The aircraft maintenance decision prediction method provided in the embodiment of the present application constructs two independent value distribution reinforcement learning models, adopts a dual learning framework, and updates based on an adaptive interval related to the training error to realize the construction of an adaptive dual value distribution reinforcement learning model, which can improve the reliability of aircraft predictive maintenance decisions.
[0045] Combine the following Figure 3 The aircraft reinforcement learning predictive maintenance decision-making method under uncertainty of remaining life provided in this application is illustrated by taking a specific example.
[0046] like Figure 3 As shown, the reinforcement learning aircraft maintenance decision-making method considering the uncertainty of remaining life provided in this specific example may include the following steps: Step 1: Divide multiple flight mission sequence data into training sets and test set Two parts.
[0047] The data used in this specific example is aircraft flight mission simulation data, which contains a total of 3100 mission sequences with a length of 50 flight missions. Each flight mission contains a flight mission benefit , flight mission duration Information on which missions benefit The flight mission duration is sampled from a uniform distribution with an upper limit of 100 units of benefit and a lower limit of 25 units of benefit. Sampled from a normal distribution with mean 20 time units and variance 5 time units squared.
[0048] The training set and test set are split in a ratio of 1:30, that is, the training set contains 100 task sequences and the test set contains 3000 task sequences. The distribution of the revenue and duration of flight missions in the training set and test set is as follows: Figure 4 , Figure 5 shown.
[0049] Step 2: Construct the value distribution reinforcement learning model structure.
[0050] The value distribution reinforcement learning model includes: model input layer dimension, number of hidden layers, hidden layer temperature and output layer dimension. Based on the designed structure, the first value distribution reinforcement learning model (referred to as model one) and the second value distribution reinforcement learning model (referred to as model two) are constructed. The update rule is determined to be an interval related to the training error, and an adaptive dual-value distribution reinforcement learning model is constructed.
[0051] Step 3: In the training set In the training, select the flight mission sequence that has not been trained. For each flight mission, the maintenance environment (including the current remaining life , Flight mission income , flight mission duration , the number of spare parts) as input, and the aircraft predictive maintenance actions are determined by model 1 .
[0052]
[0053] Step 4: Calculate the long-term benefits corresponding to the maintenance decision based on the current maintenance decision benefits and the future maximum benefits estimated by model 2. The specific benefits can be calculated by the following formula:
[0054] in, To generate the long-term benefits corresponding to the maintenance decision, is the current maintenance decision benefit, is the maximum future return estimated by model 2, is the discount factor for future earnings.
[0055] The deviation between the long-term benefit of the maintenance decision and the estimated benefit of model 2 is calculated in the form of KL divergence as the loss function of model 2, and model 2 is trained accordingly.
[0056] in, For the long-term benefits of maintenance decisions, Estimate the returns for Model 2.
[0057] Step 5: Calculate the adaptive update interval based on the loss function of model 2.
[0058] In the actual training process, Model 2 uses this number as the number of maintenance task intervals to cover its own neuron weights to Model 1.
[0059]
[0060] For example, if the calculated number of task intervals is 2, after training model 1 according to the third flight task sequence, the network parameters of model 1 are updated using the network parameters of model 2.
[0061] Step 6: After completing the training of all flight mission sequences in the training set, Model 1 is used to make decisions and obtain the aircraft predictive maintenance decision results.
[0062] In this specific example, there are 6 maintenance decisions that can be executed by the aircraft, namely, execute the mission and order spare parts, execute the mission and do not order spare parts, wait for the next mission and order spare parts, wait for the next mission and do not order spare parts, repair and order spare parts, and repair without ordering spare parts. The impact of each maintenance decision on revenue, remaining life, and the number of spare parts is shown in Table 1: Table 1 Impact of various maintenance decisions
[0063] In this specific example, the initial expected value of the remaining life is 40 units of time, and the uncertainty of the remaining life is manifested in that its actual value is a normal distribution with the expected value as the mean and the expected value 0.2 times as the variance. The initial number of spare parts is 5 units, 1 unit of spare parts is consumed for maintenance, 1 unit of spare parts is ordered each time, and the storage cost is 0.3 units of cost per unit of spare parts per unit of time.
[0064] The aircraft reinforcement learning predictive maintenance decision model (i.e., Model 1 and Model 2) constructed under the uncertainty of remaining life is trained in the task sequence of the training set. During the training process, the decision benefit of each training task sequence and the calculated adaptive update interval are as follows: Figure 6 shown.
[0065] The trained aircraft reinforcement learning predictive maintenance decision model under the uncertainty of remaining life is used to make decisions in the task sequence of the test set. Considering that the uncertainty of remaining life is affected by the number, quality and computing resources of sensors, the degree of uncertainty may fluctuate within a certain range. In order to evaluate the stability of the maintenance decision model after training (i.e., model 1 and model 2 after training), the maintenance decision model was tested under three different degrees of uncertainty: low, baseline and high. Specifically, in the low uncertainty scenario, the variance of the remaining life distribution is 0.1 times the expected value of the remaining life; in the baseline uncertainty scenario, the variance of the remaining life distribution is 0.2 times the expected value of the remaining life; in the high uncertainty scenario, the variance of the remaining life distribution is 0.3 times the expected value of the remaining life.
[0066] In the scenarios with three levels of uncertainty, the results of the aircraft reinforcement learning predictive maintenance decision model under uncertainty of remaining life are as follows: Figure 7 shown.
[0067] The average benefits of the aircraft reinforcement learning predictive maintenance decision model under uncertainty of remaining life in the scenarios of lower uncertainty, baseline uncertainty and higher uncertainty are 1257, 1214 and 1119 respectively. The average benefit of the lower uncertainty scenario is 3.54% higher than that of the baseline uncertainty scenario, and the average benefit of the higher uncertainty scenario is 7.83% lower than that of the baseline uncertainty scenario.
[0068] In order to further demonstrate the decision-making effect of the aircraft reinforcement learning predictive maintenance decision-making method under the uncertainty of remaining life, this case compares the proposed method with three methods: fixed update interval, fixed update interval and no value distribution reinforcement learning model, and no dual learning framework.
[0069] In the three uncertainty scenarios, the fixed update interval model results are as follows: Figure 8 As shown in Figure 2, the average benefits of the fixed update interval model in the lower uncertainty, baseline uncertainty, and higher uncertainty scenarios are 1255, 1175, and 995, respectively. The average benefit of the lower uncertainty scenario is 6.81% higher than that of the baseline uncertainty scenario, and the average benefit of the higher uncertainty scenario is 15.32% lower than that of the baseline uncertainty scenario.
[0070] In the three scenarios with different uncertainty levels, the results of the reinforcement learning model with fixed update interval and without using value distribution are as follows: Fig. 9As shown in the figure, the average benefits of the reinforcement learning model with fixed update interval and no value distribution in the lower uncertainty, baseline uncertainty and higher uncertainty scenarios are 1203, 1045 and 855 respectively. The average benefit of the lower uncertainty scenario is 15.12% higher than that of the baseline uncertainty scenario, and the average benefit of the higher uncertainty scenario is 18.19% lower than that of the baseline uncertainty scenario.
[0071] In the three scenarios with different uncertainty levels, the results of the model without using the dual learning framework are as follows: Fig.10 shown.
[0072] The average benefits of the reinforcement learning model with fixed update interval and no value distribution in the lower uncertainty, baseline uncertainty and higher uncertainty scenarios are 1108, 1034 and 845 respectively. The average benefit of the lower uncertainty scenario is 7.16% higher than that of the baseline uncertainty scenario, and the average benefit of the higher uncertainty scenario is 18.28% lower than that of the baseline uncertainty scenario.
[0073] In this specific example, the decision-making effects of the aircraft reinforcement learning predictive maintenance decision-making method under uncertainty of remaining life and three comparison methods were compared under different degrees of uncertainty scenarios. The comparison results under the lower uncertainty scenario are as follows: Fig.11 As shown in Table 2.
[0074] Table 2: Results for lower uncertainty scenarios
[0075] From the comparison results, it can be seen that the average return and lower quartile return of the proposed method are higher than those of the comparison method, while the upper quartile return and median return are close to the highest value of the comparison method. Therefore, it can be concluded that the proposed method is better than the comparison method in a lower uncertainty scenario.
[0076] The comparison results of the proposed method and the three comparison methods under the benchmark uncertainty scenario are shown in Fig.12 As shown in Table 3.
[0077] Table 3: Baseline uncertainty scenario results
[0078] From the comparison results, it can be seen that the proposed method is higher than the comparison method in average return, upper quartile return, median return, and lower quartile return. Therefore, it can be concluded that the proposed method is better than the comparison method in the benchmark uncertainty scenario.
[0079] The comparison results of the proposed method and the three comparison methods in high uncertainty scenarios are shown in Figure 2. Fig.13 As shown in Table 4 Table 4: Results for higher uncertainty scenarios
[0080] From the comparison results, it can be seen that the proposed method is higher than the comparison method in average return, upper quartile return, median return, and lower quartile return. Therefore, it can be concluded that the proposed method is better than the comparison method in a higher uncertainty scenario.
[0081] In order to further evaluate the stability of the model when it is affected by different degrees of uncertainty, this case further compares the degree of change of the results of the proposed method and the three comparison methods under different uncertainty scenarios.
[0082] Table 5: Comparison of average benefits between scenarios with different degrees of uncertainty
[0083] It can be seen that the proposed method model has the highest average benefits in the lower uncertainty scenario, the baseline uncertainty scenario, and the higher uncertainty scenario, and the lowest average benefit change (increase / decrease) in the lower uncertainty scenario and the higher uncertainty scenario. In other words, compared with the comparison method, the proposed method has the best and most stable performance and can better cope with different degrees of uncertainty.
[0084] This specific example provides a reinforcement learning aircraft maintenance decision-making method that takes into account the uncertainty of remaining life. In view of the limitations of the current aircraft predictive maintenance decision-making method in the face of remaining life uncertainty and over-estimation problems, a set of aircraft maintenance decision-making methods based on adaptive dual-value distribution reinforcement learning is proposed. On the one hand, the adopted value distribution reinforcement learning model can describe the benefits of flight missions in the form of distribution functions, so as to quantify the impact of remaining life uncertainty on maintenance decisions; on the other hand, the adopted dual learning framework can reduce the impact of over-estimation problems on maintenance decisions by decoupling the update and decision-making processes of the reinforcement learning model; on the third hand, the adopted adaptive update interval can dynamically adjust the update interval of the dual learning framework during the training process, so as to more effectively reduce over-estimation problems and improve the optimization efficiency of the value distribution reinforcement learning model.
[0085] Fig.14 A structural block diagram of an aircraft reinforcement learning predictive maintenance decision-making device under uncertainty of remaining life in accordance with an embodiment of the present application.
[0086] The aircraft reinforcement learning predictive maintenance decision-making device under uncertainty of remaining life in the embodiment of the present application includes the following functional modules: A generating module 1401 is used to generate a training set and a test set according to flight mission sequence data, wherein each flight mission sequence includes a plurality of flight missions; A construction module 1402 is used to construct a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; A decision generation module 1403 is used to select a flight mission sequence from the training set and input it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; A calculation module 1404, configured to calculate the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; An updating module 1405, configured to update the network parameters of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model; The prediction module 1406 is used to predict aircraft maintenance decisions by using the first value distribution reinforcement learning model obtained through training after the training of the first value distribution reinforcement learning model to be updated is completed.
[0087] Optionally, the update module includes: A first submodule, configured to calculate a loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculate an adaptive update interval based on the loss function; The second submodule is used to update the network parameters of the first value distribution reinforcement learning model to the network parameters of the second value distribution reinforcement learning model when the adaptive update interval is met.
[0088] Optionally, the first value distribution reinforcement learning model and the second value distribution reinforcement learning model both include: model input layer dimension, number of hidden layers, hidden layer temperature and output layer dimension; The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed as adaptive dual-value distribution reinforcement learning models based on a specific update rule, wherein the specific update rule is an adaptive update interval related to a training error.
[0089] Optionally, the calculation module includes: A third submodule is used to estimate the maximum future benefit corresponding to the aircraft maintenance decision through the second value distribution reinforcement learning model; A fourth submodule is used to calculate the current maintenance decision benefit corresponding to the aircraft maintenance decision based on a preset rule; The fifth submodule is used to calculate the long-term benefit corresponding to the aircraft maintenance decision based on the current maintenance decision benefit, the future maximum benefit and the discount coefficient of the future benefit.
[0090] Optionally, the first submodule is specifically used for: Calculating the deviation between the long-term benefit and the future maximum benefit in the form of KL divergence as the loss function of the second value distribution reinforcement learning model; An adaptive update interval is calculated based on the loss function.
[0091] The aircraft reinforcement learning predictive maintenance decision-making device under the uncertainty of remaining life provided in the embodiment of the present application constructs two independent value distribution reinforcement learning models, adopts a dual learning framework, and updates based on an adaptive interval related to the training error to realize the construction of an adaptive dual value distribution reinforcement learning model, which can improve the reliability of aircraft predictive maintenance decisions.
[0092] The embodiments of the present application provide Fig.14 The aircraft reinforcement learning predictive maintenance decision-making device shown under the uncertainty of remaining life can achieve Figure 1 To avoid repetition, the various processes implemented by the method embodiment are not described here.
[0093] An embodiment of the present invention further provides an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus.
[0094] Memory, used to store computer programs; The processor is used to implement the aircraft reinforcement learning predictive maintenance decision-making method under the uncertainty of remaining life shown in the above method embodiment when executing the program stored in the memory.
[0095] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0096] The communication interface is used for communication between the above terminal and other devices.
[0097] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0098] In another embodiment provided by the present invention, a computer-readable storage medium is also provided, in which instructions are stored. When the computer-readable storage medium is executed on an electronic device, the electronic device implements the aircraft reinforcement learning predictive maintenance decision-making method under the uncertainty of remaining life described in any of the above embodiments.
[0099] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When the computer program product is run on an electronic device, the electronic device implements the aircraft reinforcement learning predictive maintenance decision-making method under the uncertainty of remaining life described in any of the above embodiments.
[0100] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0101] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A reinforcement learning predictive maintenance decision-making method for aircraft under uncertainty of remaining life, characterized in that: The method comprises: Generate a training set and a test set based on the flight mission sequence data, wherein each flight mission sequence includes multiple flight missions; Construct a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; Selecting a flight mission sequence from the training set and inputting it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; Calculating the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; According to the long-term benefit and the second value distribution reinforcement learning model, updating the network parameters of the first value distribution reinforcement learning model; After the training of the updated first value distribution reinforcement learning model is completed, the aircraft maintenance decision is predicted by the first value distribution reinforcement learning model obtained through training.
2. The method according to claim 1, characterized in that: The step of updating the network parameters of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model comprises: Calculating a loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculating an adaptive update interval based on the loss function; When the adaptive update interval is met, the network parameters of the first value distribution reinforcement learning model are updated to the network parameters of the second value distribution reinforcement learning model.
3. The method according to claim 1, characterized in that: The first value distribution reinforcement learning model and the second value distribution reinforcement learning model both include: model input layer dimension, hidden layer number, hidden layer temperature and output layer dimension; The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed as adaptive dual-value distribution reinforcement learning models based on a specific update rule, wherein the specific update rule is an adaptive update interval related to a training error.
4. The method according to claim 1, characterized in that The step of calculating the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model includes: estimating the maximum future benefit corresponding to the aircraft maintenance decision by using the second value distribution reinforcement learning model; Calculate the current maintenance decision benefit corresponding to the aircraft maintenance decision based on preset rules; Based on the current maintenance decision benefit, the future maximum benefit and the discount coefficient of the future benefit, the long-term benefit corresponding to the aircraft maintenance decision is calculated.
5. The method according to claim 2, characterized in that: The step of calculating the loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculating the adaptive update interval based on the loss function includes: Calculating the deviation between the long-term benefit and the future maximum benefit in the form of KL divergence as the loss function of the second value distribution reinforcement learning model; An adaptive update interval is calculated based on the loss function.
6. An aircraft reinforcement learning predictive maintenance decision-making device under uncertainty of remaining life, characterized in that: The device comprises: A generation module, used to generate a training set and a test set based on flight mission sequence data, wherein each flight mission sequence includes multiple flight missions; A construction module, used to construct a first value distribution reinforcement learning model and a second value distribution reinforcement learning model respectively; A decision generation module, used for selecting a flight mission sequence from the training set and inputting it into the first value distribution reinforcement learning model to generate an aircraft maintenance decision; A calculation module, configured to calculate the long-term benefit corresponding to the aircraft maintenance decision based on the second value distribution reinforcement learning model; An updating module, configured to update a network parameter of the first value distribution reinforcement learning model according to the long-term benefit and the second value distribution reinforcement learning model; The prediction module is used to predict aircraft maintenance decisions by using the first value distribution reinforcement learning model obtained through training after the training of the first value distribution reinforcement learning model to be updated is completed.
7. The device according to claim 6, characterized in that The update module includes: A first submodule, configured to calculate a loss function of the second value distribution reinforcement learning model according to the long-term benefit, and calculate an adaptive update interval based on the loss function; The second submodule is used to update the network parameters of the first value distribution reinforcement learning model to the network parameters of the second value distribution reinforcement learning model when the adaptive update interval is met.
8. The device according to claim 6, characterized in that: The first value distribution reinforcement learning model and the second value distribution reinforcement learning model both include: model input layer dimension, hidden layer number, hidden layer temperature and output layer dimension; The first value distribution reinforcement learning model and the second value distribution reinforcement learning model are constructed as adaptive dual-value distribution reinforcement learning models based on a specific update rule, wherein the specific update rule is an adaptive update interval related to a training error.
9. The device according to claim 6, characterized in that The calculation module comprises: A third submodule is used to estimate the maximum future benefit corresponding to the aircraft maintenance decision through the second value distribution reinforcement learning model; A fourth submodule is used to calculate the current maintenance decision benefit corresponding to the aircraft maintenance decision based on a preset rule; The fifth submodule is used to calculate the long-term benefit corresponding to the aircraft maintenance decision based on the current maintenance decision benefit, the future maximum benefit and the discount coefficient of the future benefit.
10. The device according to claim 7, characterized in that The first submodule is specifically used for: Calculating the deviation between the long-term benefit and the future maximum benefit in the form of KL divergence as the loss function of the second value distribution reinforcement learning model; An adaptive update interval is calculated based on the loss function.
11. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the aircraft reinforcement learning predictive maintenance decision-making method under uncertainty of remaining life as described in any one of claims 1 to 5 when executing the program stored in the memory.
Citation Information
Patent Citations
Aircraft maintenance path optimization method based on adaptive reinforcement learning
CN115292959A
Equipment optimal maintenance strategy searching method and system based on reinforcement learning
CN118941275A
Reinforcement learning system for maintenance decision making
EP4407523A1
KR20220059837A
Cited By
Aircraft engine maintenance decision-making method based on residual life prediction and deep reinforcement learning
CN120218911A