Aircraft power supply system health state assessment method based on dynamic weight

By combining the dynamic weight model of the improved analytic hierarchy process and entropy weight method with fuzzy quantization technology, and by constructing a closed-loop management mechanism using LSTM and DQN, the problem of dynamic adjustment and closed-loop management in the health status assessment of aircraft power supply systems is solved, achieving accurate health status assessment and intelligent maintenance decision-making.

CN121809244APending Publication Date: 2026-04-07CHINA AERO POLYTECH ESTAB
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing methods for assessing the health status of aircraft power supply systems have shortcomings in terms of rigid weight determination mechanisms, crude quantitative assessment of status, and disconnect between health assessment and maintenance decisions. They are unable to achieve dynamic adjustment, accurate description of intermediate degradation states, and closed-loop management.

Method used

A dynamic weight determination model based on the combination of improved analytic hierarchy process (IAHP) and entropy weight method is adopted. The weights are fused by adaptive fuzzy reinforcement learning (AFRL) and quantified by trapezoidal fuzzy membership function. A closed-loop management mechanism is constructed by integrating long short-term memory network (LSTM) and deep reinforcement learning (DQN) to realize intelligent support from health status assessment to maintenance decision making.

Benefits of technology

It enables real-time, accurate, and forward-looking health status assessment of aircraft power supply systems, dynamically adjusts weights, precisely quantifies changes in health status, and constructs a closed-loop management system from assessment to decision-making, thereby improving the robustness and accuracy of the assessment system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809244A_ABST
    Figure CN121809244A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power supply, and provides an aircraft power supply system health state evaluation method based on dynamic weight, comprising the following steps: S1, determining the grade of an aircraft power supply system; s2, initializing parameters; s3, a subjective weight determination model and an objective weight determination model are established; s4, obtaining the weight of each index in the finished product parameter level; s5, calculating the health index of the index; s6, judging whether circulation is ended or not; s7, obtaining subjective and objective weights of the layer indexes; s8, obtaining subjective and objective weight difference and a health state change rate; s9, obtaining a subjective and objective weight ratio; s10, obtaining fuzzy subjective and objective weights; s11, fuzzy weighting is adopted to obtain a weight; s12, carrying out normalization on the weight values; s13, returning to the step S5; and S14, obtaining the health level of the power supply system. According to the method, the model and the fuzzy rule table are determined through the subjective and objective weight ratio, and nonlinear self-adaptive fusion is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power supply, and particularly to a method for evaluating the health state of an aircraft power supply system based on dynamic weights. BACKGROUND

[0002] The aircraft power supply system is a key secondary energy system of the aircraft, and its health state is directly related to flight safety and mission execution capability. With the development of the PHM (Predictive and Health Management) technology, it has become possible to monitor the health state, diagnose faults and predict the future of complex equipment. However, when the existing technology is actually applied to complex and dynamic systems such as the aircraft power supply system, there are still several technical bottlenecks to be solved, mainly in the following aspects: Firstly, the weight determination mechanism is rigid and cannot reflect the dynamic characteristics of the system: the existing health state evaluation methods mostly rely on fixed weight models, such as the traditional AHP (Analytic Hierarchy Process). The weights determined by such methods are static and cannot be dynamically adjusted according to the actual operation data of the power supply system, the performance degradation of components or the changes in the task stage. For example, a parameter with a low weight under normal conditions should be given more attention when it experiences severe fluctuations, but the fixed weight model cannot achieve such adaptive adjustment, resulting in a disconnection between the evaluation results and the true state of the system, and making it difficult to accurately identify potential risks.

[0003] Secondly, the state evaluation quantification method is rough and not sensitive to performance degradation. Many evaluation methods use binary (normal / fault) or simple linear quantification indicators, which cannot accurately describe the intermediate degradation state of the system or components between "completely normal" and "completely failed". This all-or-nothing quantification method lacks the ability to capture slow performance degradation, intermittent failures and other situations, resulting in the system failing to trigger effective early warnings when early performance degradation occurs, and missing the best maintenance opportunity.

[0004] Finally, the health evaluation and maintenance decision-making are disconnected, and a closed-loop management is not formed. The existing technology mostly focuses on the evaluation of the current or historical health state, and lacks effective prediction of the future state. Even if the prediction is made, the prediction results cannot be deeply integrated with the optimization of maintenance decision-making. Maintenance decision-making often relies on human experience and lacks data-driven, multi-objective (such as task urgency, maintenance cost, inventory) automated decision support. This makes the maintenance strategy either too conservative causing resource waste or too aggressive bringing safety risks, and cannot achieve precise maintenance based on the state.

[0005] In summary, existing aircraft power supply system health management technologies have significant shortcomings in terms of the dynamism of weights, the precision of status quantification, the closed-loop linkage between assessment, prediction, and decision-making, and the adaptability of information fusion. Therefore, there is an urgent need for a comprehensive solution that can overcome these deficiencies to achieve real-time, accurate, and forward-looking assessment of the health status of aircraft power supply systems and intelligent maintenance decision support. Summary of the Invention

[0006] To address the shortcomings of the prior art, the present invention aims to provide a method for assessing the health status of an aircraft power supply system based on dynamic weights, comprising the following steps: S1, determine the classification of the aircraft's power supply system; S2, parameter initialization; S3, Establish subjective weight determination model and objective weight determination model for indicators; S4, obtain the weight of each indicator in the finished product parameter level; (4); in, The third level of finished product parameters Each indicator in Weight of time, The third level of finished product parameters The subjective weight of each indicator, The third level of finished product parameters Each indicator in Objective weight of time, This is an empirical coefficient; S5, for the first Each indicator in the layer calculates the health index of the indicator. : (5); in, For the first Layer Each indicator in Health index at any time In order to be in The first time is composed of the time components The next level of indicators The weight, for The first time is composed of the time components The next level of indicators The standardized probability value, To form the first The next level of indicators for each indicator; S6, Judgment If the value is 0, proceed to step S14; otherwise, proceed to step S7. S7, for Each indicator in the layer has its subjective weight and objective weight determined according to S2; S8, obtain the difference between subjective and objective weights and the rate of change in health status; S9, determine the proportion of subjective and objective weights for each indicator based on the subjective and objective weight ratio; S10, perform difference fuzzification processing on the subjective and objective weights to obtain the fuzzified subjective and objective weights of each indicator; S11, obtained by using fuzzy weighting. Weights of each metric in the layer: (8); in, and To determine the model's first result based on the ratio of subjective to objective weights Layer The ratio of subjective weight to objective weight for each indicator at time t. For the first Layer Subjective weights after blurring the individual indicators For the first Layer The objective weights of each indicator after fuzzification; S12, for the first The weight values ​​of the same set of indicators in the same layer are normalized; S13, And return to step S5; S14, the health level of the power supply system is obtained based on the power supply system health index.

[0007] Preferably, the S1 aircraft power supply system is hierarchically structured as follows: The zeroth layer is the target layer of the aircraft power supply system; the first layer is the system function level, which includes all functional modules of the power supply system divided according to the system functions; the second layer is the finished product level, which is formed by dividing each functional module according to the finished products used; the third layer is the parameter level, which is formed by monitoring the parameters of each finished product and all monitored parameters.

[0008] Preferably, the initialization of the S2 parameter is specifically as follows: initialization , , and ,in This is the current level number. Indicates the first the health state of the first index of the first layer at the time point, the rate of change of the system health state at the time point, the health state of the first index of the first layer at the time point, the health state of the first index of the first layer at the time point, the health state of the first index of the first layer at the time point, the health state of the first index of the first layer at the time point, the health state of the first index of the first layer at the time point, the health state of the first index of the first layer at the time point, the health state of the first index of the first layer at the time point, the health state of the first index of the first layer at the time point,

[0009] Preferably, the subjective weight determination model and the objective weight determination model in S3 are specifically: the index subjective weight determination model: by obtaining the judgment matrix of the group where the index is located, the subjective weight of the index is obtained by using the improved analytic hierarchy process, and the logical rationality is verified by consistency test, the matrix that does not satisfy the consistency is introduced into correction coefficient iteration optimization, and finally the standardized subjective weight of each index in the same group is obtained by normalization; the index objective weight determination model: by obtaining the monitoring data related to the index, the monitoring data is normalized to a standardized value, the information entropy of the index is calculated, the smaller the entropy value is, the greater the volatility is, and the higher the importance is, and the entropy value is converted into weight by using formula (2): (2) ; wherein, , is the standardized probability; wherein, the information entropy of the first index of the first layer at the time point, the set of the index group where the first index of the first layer is located, the standardized probability value of the first index of the first layer at the time point, the total number of measurement points obtained after the time point of the first index of the first layer, the measurement value of the measurement point at the time point of the first index of the first layer, the objective weight of the first index of the first layer at the time point.

[0010] ​​​​​​​​​​​​​​​​​Preferably, the subjective-objective weight difference and the health state change rate obtained in S8 are specifically as follows: Subjective-objective weight difference : (6) ; Wherein, represents the subjective-objective weight difference of the i-th index of the j-th layer at the time t; ; Health state change rate : (7) ; Wherein, is defined as a time window, i.e. a time step, represents the system health state change rate of the i-th index of the j-th layer at the time t; represents the health state of the i-th index of the j-th layer at the time t; is the health state of the i-th index of the j-th layer at the time t.

[0011] Preferably, the proportion of the subjective-objective weight of each index obtained according to the subjective-objective weight ratio in S9 is determined by a model, and is specifically as follows: S91, constructing a subjective-objective weight proportion determination model; S92, inputting the subjective-objective weight difference, the health state change rate and the health index of each index of the j-th layer into the subjective-objective weight proportion determination model to obtain the proportion of the subjective-objective weight of each index of the j-th layer.

[0012] Preferably, the subjective-objective weight proportion determination model is constructed in S91, and is specifically as follows: The input of the subjective-objective weight ratio determination model is a state space: , wherein, is the difference between the subjective weight and the objective weight, is the health state change rate at the time t, is the health index at the time t, The output of the subjective-objective weight ratio determination model is an action space: ​​​​​​​​​​​​​​​​,in and To obtain new subjective weight proportion coefficients and objective weight proportion coefficients through deep deterministic strategy gradient algorithm optimization; where, and satisfy , The ratio of subjective to objective weights determines the model's reward function: ,in and Hyperparameters that balance prediction accuracy and weight stability; The subjective and objective weight ratio determines the model's network structure: an Actor network is used for output. and The adjustment strategy involves using a Critic network to evaluate the Q-value of state-action pairs, which guides Actor optimization.

[0013] Preferably, step S10 performs difference fuzzification processing on the subjective and objective weights to obtain the fuzzified subjective and objective weights for each indicator, specifically as follows: Obtain the difference between subjective and objective weights and the rate of change in health status As input quantities, the input quantities are fuzzified through the trapezoidal membership function, and the subjective and objective weights of each indicator are obtained according to the fuzzy rule table. The trapezoidal membership function is as follows: Membership degree This indicates a "low" level of difference. Membership degree , This indicates a "medium" level of difference. Membership degree This indicates a "high" level of difference. The rates of change are categorized as follows: This indicates a gentle or gradual transition; , indicating fluctuation; , indicating intense; Fuzzy rule tables are based on the difference between subjective and objective weights. and rate of change in health status The defined fuzzy rules presuppose the proportions of subjective and objective weights.

[0014] Preferably, S12 is for the first The weights of the same set of indicators within the same layer are normalized, specifically as follows: (9); wherein, is the index of the same group, is the index of the same group, is the sum of all indexes belonging to the same group, is the index of the same group, is the index set belonging to the same group.

[0015] Preferably, it further comprises S15, health prediction and maintenance decision optimization: S151, using a health degradation trend prediction model based on an LSTM network, the power supply system health index time series are used as input to predict the power supply system health degradation; S152, using a maintenance decision optimization model based on deep Q network and DQN, the power supply system health index, task urgency, spare parts inventory are used as input to obtain maintenance decision.

[0016] Compared with the prior art, the application has the following beneficial effects: 1. The application breaks through the limitation of fixed weight model: by combining IAHP and entropy weight method, and introducing a dynamic weight fusion mechanism based on adaptive fuzzy reinforcement learning AFRL, the weight of each level index can be adaptively adjusted according to real-time monitoring data and health state change rate, which significantly improves the capture ability of the evaluation model to the dynamic characteristics of the system.

[0017] 2. The application realizes the fine quantization of health state: by introducing a trapezoidal fuzzy membership function to continuously quantify the functional state, the problem of insensitivity of binary index to intermediate decay state is effectively solved, which can more accurately describe the gradual change process of the system from health to failure, and facilitate early detection of performance degradation trend.

[0018] 3. The application constructs a closed-loop management from evaluation to decision: by fusing long short-term memory network LSTM and deep reinforcement learning DQN, a full-process closed-loop management mechanism from real-time health state evaluation, degradation trend prediction to maintenance decision optimization is constructed, realizing data-driven intelligent maintenance and providing key technical support for state-based precision maintenance.

[0019] 4. The application improves the intelligent level of subjective and objective information fusion: by constructing a DDPG-based subjective and objective weight proportion determination model and fuzzy rule table, nonlinear and adaptive fusion of subjective and objective weights is realized, which can make more reasonable trade-off when data and experience conflict, improving the robustness and accuracy of the evaluation system. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 is the flowchart of the aircraft power supply system health state evaluation method based on dynamic weight of the application;​ Figure 2 This is an exemplary diagram of the layered aircraft power supply system in the embodiments of this application; Figure 3 This is a flowchart of the health prediction and maintenance decision optimization steps in the embodiments of this application. Detailed Implementation

[0021] To fully explain the technical content, objectives, and effects of this invention, the embodiments of this invention will be described in detail below with reference to the accompanying drawings.

[0022] This application proposes a dynamic weight-based health status assessment method for aircraft power supply systems. First, a multi-level analysis framework (system function layer, product layer, parameter layer) is used to identify the core functions affecting safety and mission performance. The weights at each level are dynamically optimized using an improved analytic hierarchy process (IAHP) and an entropy weight method. A dynamic weight fusion method based on adaptive fuzzy reinforcement learning (AFRL) is proposed. Second, a trapezoidal fuzzy membership function is introduced to continuously quantify functional states, addressing the lack of sensitivity of binary indicators to intermediate decay states. Finally, a long short-term memory network (LSTM) and deep reinforcement learning (DQN) are integrated to construct a closed-loop management mechanism from health status prediction to maintenance decision optimization. The dynamic weight-based health status assessment method for aircraft power supply systems proposed in this application, such as… Figure 1 As shown, the specific steps are as follows: S1, determine the classification of the aircraft's power supply system.

[0023] By summarizing and analyzing the failure modes, failure frequency, losses, and monitoring data of the aircraft power supply system, and taking the aircraft power supply system as the target layer, the system is further divided into three levels according to function, finished products, and parameters. The aircraft power supply system consists of multiple functional modules, each functional module consists of multiple finished products, and each finished product is monitored by multiple parameters. The aircraft power supply system is designated as the zeroth level target layer. The three levels after this division are as follows: The first layer is the system function level: it includes all functional modules of the power supply system divided according to system functions.

[0024] The second layer is the finished product level: after dividing each functional module according to the finished products used, all finished products constitute the finished product level.

[0025] The third level is the parameter level: parameters are monitored for each finished product, and all monitored parameters constitute the parameter level.

[0026] This embodiment uses a certain type of aircraft power supply system as an example. The aircraft power supply system is illustrated by its layered structure as follows: Figure 2 As shown: Taking the aircraft power supply system as the target layer (layer zero), the aircraft power supply system is divided into three layers: The first layer is the system functional level: from a forward design perspective, it includes power generation, power distribution, uninterrupted power supply, fault warning, and system control capabilities, thus comprising a total of 5 functional modules.

[0027] The second layer is the finished product level: main generator, generator controller, DC-DC converter, AC-DC converter, etc. constitute power generation; primary power distribution, power distribution center, secondary power distribution, etc. constitute power distribution; auxiliary generator, auxiliary generator controller, storage battery, etc. constitute uninterrupted power supply; fault detection unit, storage battery controller, main generator controller, etc. constitute fault warning; control unit, data transmission capacity, etc. constitute system control capability.

[0028] The third layer is the finished product parameter level: each finished product corresponds to several monitoring parameters. For example, the parameters that can be collected by the main generator include output voltage, output current, temperature, noise, insulation resistance, and oil cooling flow.

[0029] S2, parameter initialization.

[0030] Initialize the current level number , , and ;in Indicates the first The first layer Each indicator in The rate of change in the system's health status at any given time. Indicates the first The first layer Each indicator in Constant health status To indicate the first The first layer Each indicator in A person's health status at all times.

[0031] Current level The initial value is set to the layer number minus 1. This is because the weighting algorithm for the bottommost parameter layer is different from that of other layers; the loop starts from the second-to-last layer. Therefore, in this embodiment... , , and The initial values ​​are all set to 0.

[0032] S3, establish a subjective weight determination model and an objective weight determination model for the indicators.

[0033] In this application, each functional module, finished product, and parameter is treated as an indicator of the current layer. According to the classification of the aircraft power supply system, if the indicators of the current layer belong to the same upper-level indicator, then these indicators of the current layer are considered to belong to the same group.

[0034] This paper proposes a method that combines the improved Analytic Hierarchy Process (IAHP) with the entropy weight method to dynamically optimize weights based on adaptive fuzzy reinforcement learning; for the weights of each indicator... All are determined by the subjective weight of the indicators. and objective weight of indicators Composition, weight Indicates the first Layer The weight of each indicator.

[0035] Subjective weight of indicators Model Determination: By obtaining the judgment matrix of the indicator group, the subjective weights of the indicators are obtained using the Improved Analytic Hierarchy Process (IAHP). The logical rationality is verified through consistency checks. For matrices that do not meet consistency requirements, correction coefficients are introduced for iterative optimization. Finally, the standardized subjective weights of each indicator in the same group are obtained through normalization. The judgment matrix is ​​obtained by pairwise importance comparisons and scoring of indicators within the same group. Each element in the judgment matrix ranges from 1 to 9 points, and the scores can be determined by experts or based on experience.

[0036] Objective weight of indicators Model determination: The monitoring data related to the indicators are obtained and normalized to a standardized value of 0 to 1. The information entropy of the indicators is calculated. The smaller the information entropy value, the greater the volatility and the higher the importance. The entropy value is converted into weights using formula (2): (2); In the formula, , This represents the standardized probability.

[0037] In the formula above, For the first The first layer Each indicator in Information entropy at any given moment For the first The first layer The set of indicator groups to which each indicator belongs. For the first The first layer Each indicator in The standardized probability value at time 1. For the first The first layer Each indicator in The total number of measurement points acquired after time step, for example, at time step. After a certain time, it was obtained Data from each measurement point, For the first The first layer Each indicator in The measured value at the measurement point at time t. For the first The first layer Each indicator in Objective weighting of time. At the finished product parameter level, the measured value of the measurement point refers to the normalized physical data obtained through instrument monitoring. For example, voltage is the voltage value obtained from a voltmeter, and temperature is the temperature value obtained from a temperature sensor. At other levels, the measured value of the measurement point refers to the health status of the indicators obtained through calculation.

[0038] When the indicator is at the finished product parameter level, the measured value of the measurement point refers to the value obtained after normalizing the original monitoring value, specifically defined as: (3) in, For the finished product parameter level Each indicator in Normalized measured values ​​of the original monitoring values ​​at any given time. For the finished product parameter level Each indicator in The original monitoring value at that moment, For the finished product parameter level The maximum value of each indicator when it functions normally For the finished product parameter level The minimum value when each indicator functions normally, i.e., the failure threshold.

[0039] According to formula (2), firstly, the information entropy of each indicator is calculated separately, and then the objective weight of each indicator in the same group is obtained by applying formula (2) separately. Regularly updated The monitoring data of each indicator is used to dynamically adjust the objective weights. To ensure an objective reflection of the first The actual importance of each indicator.

[0040] Based on the definitions of subjective and objective indicators, we know that subjective indicators are relatively stable, while objective indicators change in real time according to the detected data. Therefore, the weight of each indicator also changes in real time, realizing real-time monitoring of the health status of the aircraft power supply system.

[0041] S4 retrieves the weight of each indicator in the finished product parameter level.

[0042] In the finished product parameter level, each parameter is an indicator. Parameters belonging to the same finished product are grouped into a set of indicators. Subjective and objective weights are calculated for each finished product's parameters based on S2; then, empirical coefficients are used. By balancing subjective and objective weights, the weight and empirical coefficient of each parameter are obtained. The values ​​are typically between 0.5 and 0.8. The weight of each parameter in the third-level finished product parameter level is as follows: (4); in, The third level of finished product parameters Each parameter index Weight of time, The third level of finished product parameters The subjective weight of each indicator, The third level of finished product parameters Each parameter index Objective weight of time, This is an empirical coefficient.

[0043] Taking the main generator of the finished product in this embodiment as an example, the collected parameters include output voltage, output current, temperature, noise, insulation resistance, and oil cooling flow rate. Therefore, the subjective weights for each parameter of the main generator are: output voltage 0.25, output current 0.20, temperature 0.15, noise 0.10, insulation resistance 0.20, and oil cooling flow rate 0.10. For the output voltage acquisition... After a moment Data from each measurement point, for A specific value is obtained by calculating the information entropy of the output voltage as 0.12 and the objective weight as 0.22 according to formula (2) in step S2; when the empirical coefficient When the value is 0.6, the final weight of the output voltage is obtained as follows: =0.6×0.25 + 0.4×0.22 = 0.238. Similarly, the final weights of the output current, temperature, noise, insulation resistance, and oil cooling flow rate are 0.192, 0.125, 0.159, 0.120, and 0.166, respectively.

[0044] S5, for the first Each indicator in the layer calculates the health index of the indicator. : (5); in, For the first Layer Each indicator in Health index at any time In order to be in The first time is composed of the time components The next level of indicators The weight, for The first time is composed of the time components The next level of indicators The standardized probability value is calculated using the same method as in formula (2). same, To form the first The next level of indicators for each indicator. The value is within the range [0,1].

[0045] In this embodiment, The health index of the main generator at any given time is: = 0.238×0.95 + 0.192×0.88 + 0.125×0.70 +0.159×0.65 +0.120×0.60 +0.166×0.80 =0.79.

[0046] Here This is the set of parameter indicators for the next level belonging to the main generator.

[0047] S6, Judgment If the value is 0, proceed to step S14; otherwise, proceed to step S7.

[0048] S7, based on S2 Each indicator in the layer has its own subjective weight calculated. and objective weight of indicators ;at this time This becomes the formula (2) Values.

[0049] S8 obtains the difference between subjective and objective weights and the rate of change in health status.

[0050] Obtain the difference between subjective and objective weights : (6); in, Indicates the first The first layer Each indicator in Difference between subjective and objective weights at any given moment .

[0051] Obtain the rate of change in health status : (7); in, Defined as a time window, i.e., a time step, such as Hour, Indicates the first The first layer Each indicator in The rate of change in the system's health status at any given time. Indicates the first The first layer Each indicator in Constant health status To indicate the first The first layer Each indicator in Constant health status Time is The point in time preceding the current moment.

[0052] In this embodiment, The difference between the subjective and objective weights of the main generator at any given time is =0.05, the rate of change in health status =0.02.

[0053] S9, determine the proportion of subjective and objective weights for each indicator obtained from the model based on the subjective-objective weight ratio; specifically: S91, Construct a model to determine the proportion of subjective and objective weights.

[0054] The model for determining the weight ratio of subjective and objective factors is implemented using an Actor-Critic network model based on the Deep Deterministic Policy Gradient (DDPG) algorithm, as detailed below: The ratio of subjective to objective weights determines the input to the model as the state space: ,in, This is the difference between subjective weights and objective weights. for Rate of change in health status over time for Health index at any time.

[0055] The ratio of subjective to objective weights determines the model's output as the action space. ,in and To obtain new subjective and objective weight ratios through deep deterministic strategy gradient optimization. and satisfy .

[0056] The ratio of subjective to objective weights determines the model's reward function: ,in and To balance prediction accuracy and weight stability, the preferred hyperparameter is... =0.5, =0.3.

[0057] The subjective and objective weight ratio determines the model's network structure: an Actor network is used for output. and The adjustment strategy involves using a Critic network to evaluate the Q-value of state-action pairs, which guides Actor optimization.

[0058] S92, the first The subjective and objective weight differences, health status change rate, and health index of each indicator in the layer are input into the subjective and objective weight ratio determination model to obtain the first layer. The proportion of subjective and objective weights for each indicator in the layer.

[0059] S10, perform difference blurring processing on subjective and objective weights.

[0060] Obtain the difference between subjective and objective weights and the rate of change in health status As input, the input is also fuzzified using a trapezoidal membership function.

[0061] The trapezoidal membership function is as follows: Membership degree This indicates a "low" level of difference. Membership degree , This indicates a "medium" level of difference. Membership degree This indicates a "high" level of difference.

[0062] The rates of change are categorized as follows: This indicates a gentle or gradual transition; , indicating fluctuation; , indicating intense.

[0063] Fuzzification of input quantities refers to classifying input quantities according to a fuzzy rule table, rather than using actual values ​​for input. In this embodiment, the fuzzy rule table is based on the difference between subjective and objective weights. and rate of change in health status The defined fuzzy rules pre-determine the subjective weights after fuzzification. and objective weight As shown in Table 1: Table 1 Fuzzy Rule Table The subjective and objective weights in Table 1 are only preferred options; the specific values ​​can be set according to actual conditions. According to the fuzzy rule table in this embodiment, if... For high and If the intensity is high, then increase the weighting of objective factors; if... For low and To achieve a smoother transition, subjective weighting should be maintained.

[0064] S11, fuzzy weighting is used instead of linear combination to obtain the weights: (8); in, and To determine the model's first result based on the ratio of subjective to objective weights Layer The ratio of subjective weight to objective weight for each indicator. and These are the fuzzy subjective and objective weight values ​​obtained from S10.

[0065] In this embodiment, For the main generator indicator, the input to the model for determining the weighting of subjective and objective factors is obtained at each moment. ,Will Inputting the data into the subjective and objective weighting ratio determination model yields the following results: =0.8, =0.2.

[0066] At the same time, according to Difference between subjective and objective weights of main generator indicators at any given time =0.05, rate of change in health status =0.02, fuzzification is performed based on the trapezoidal membership function: =0.05 falls within the interval [0, 0.1), therefore its difference level is "low"; =0.02 belongs to The interval is defined as "slow" in terms of its rate of change. Then, by consulting the fuzzy rule table shown in Table 1, at the intersection of "difference level - low" and "rate of change - slow", a subjective weight of 90% and an objective weight of 10% are obtained.

[0067] According to formula (6): .

[0068] S12, normalizes the same set of weight values.

[0069] Based on the weight values ​​obtained after S11 adjustment, it cannot be guaranteed that the weight values ​​of each group of indicators will add up to 1. Therefore, it is necessary to normalize the weight values ​​of the same group.

[0070] (9); in, To and Indicators belonging to the same group To and Add up all the indicators belonging to the same group. For the first Layer The set of indicator groups to which an indicator belongs, here refers to the set of indicators that are related to... A set of indicators belonging to the same group.

[0071] S13, Then return to step S5.

[0072] S14. Evaluate the health level of the power supply system based on the power supply system health index.

[0073] The health index, the top-level indicator, is obtained based on step S5. That is, the health index of the power supply system The health level of the power supply system is determined according to a preset threshold.

[0074] In this embodiment, the five indicators of the first-level system function are obtained iteratively. The times are respectively , , , , The normalized weights of the five indicators are as follows: , , , , Therefore, after another iteration, we obtain At this point, make a judgment This yields the power supply system health index. .

[0075] In this embodiment, the threshold values ​​for classifying the health level of the power supply system are shown in Table 2: Table 2 Health Level Therefore, we can conclude that, The health status level of the power supply system is "qualified", and the health status is "slight performance fluctuations that do not affect task execution".

[0076] S15, Health Prediction and Maintenance Decision Optimization.

[0077] Furthermore, this invention can also obtain a predicted sequence of the power supply system health index using a health degradation trend prediction model based on an LSTM network, based on the power supply system health index time series. This predicted sequence is then input into a maintenance decision optimization model based on a deep Q-network and DQN to obtain maintenance decisions. In other words, by fusing time series prediction and reinforcement learning techniques based on the power supply system health index time series, closed-loop management from health status assessment to maintenance decision-making is achieved. The specific process is as follows: Figure 3 As shown.

[0078] S151, a health degradation trend prediction model based on LSTM network.

[0079] S1511, Model Selection and Principles.

[0080] Long Short-Term Memory (LSTM) networks have been chosen as the core algorithm for health index prediction due to their ability to model long-term dependencies in temporal data. Their gating mechanism (input gate, forget gate, output gate) effectively captures the nonlinear dynamic features and state change patterns during power system degradation by selectively remembering and forgetting temporal features. Compared to traditional RNNs, LSTM avoids the vanishing gradient problem through continuous cell state updates, thus enabling modeling of the entire lifecycle of equipment degradation. Theoretical analysis shows that the forget gate controls the retention ratio of historical information through the sigmoid function, while the input gate filters new features through the hyperbolic tangent function (tanh). The synergistic effect of these two mechanisms significantly improves the accuracy of identifying degradation inflection points.

[0081] The health degradation trend prediction model adopts the following structure: Network architecture: It adopts a stacked LSTM structure, which includes two LSTM layers with 128 units per layer, Dropout=0.2, and one fully connected output layer.

[0082] Loss function: The mean squared error (MSE) is used as the objective function to balance prediction stability and outlier sensitivity.

[0083] Optimization strategy: Adaptive moment estimation optimizer Adam is used, with an initial learning rate of 0.001, and ReduceLROnPlateau is used for dynamic adjustment.

[0084] Regularization method: Early stopping is introduced to monitor the validation set loss. Training is terminated when there is no improvement after 10 consecutive epochs. At the same time, weight decay (L2=1e-4) is used to suppress overfitting.

[0085] S1512, Data Preparation and Processing.

[0086] Input data: Historical power supply system health index , , ..., Time window length The degradation cycle characteristics of the system are determined by autocorrelation function (ACF) analysis.

[0087] Data normalization: Min-Max standardization is used to map the health index to the [0,1] interval.

[0088] Sliding window segmentation: Training samples are generated by scrolling through time steps, with the input being [ , , …, , The output is H(t+T), where T is the prediction step size; for example, T=3 corresponds to the next 3 periods. A time-series segmentation strategy with Shuffle=False is used to maintain data causality. The dataset is divided into training, validation, and test sets in a 7:2:1 ratio to avoid future information leakage. The prediction results are then applied.

[0089] Prediction output: The model outputs the predicted value H(t+T) for the next T steps and its 95% confidence interval, supporting probabilistic risk assessment. For example, when T=3, the predicted value H=0.68±0.05.

[0090] Early warning mechanism: When the predicted value H(t+T) < 0.75, a level 3 early warning is triggered. The threshold is determined based on the ROC curve of historical fault data. If the prediction exceeds the limit three times in a row, a maintenance work order is generated.

[0091] Decision-making closed loop: The prediction results are linked with equipment ledgers and operation and maintenance records, and the degradation model parameters are dynamically corrected through Bayesian updates. A typical case shows that when the baseline H(t) of a power supply system is 0.82, the predicted H=0.68 after 3 cycles is automatically triggered by the system to schedule spare parts and maintenance plans, identifying risks 14 days earlier than traditional threshold alarms.

[0092] S152, a maintenance decision optimization model based on deep Q-networks and DQN.

[0093] S1521, Reinforcement Learning Framework Design.

[0094] The system's state space is as follows: Health Index H(t): The predicted health index of the power supply system is obtained based on the health degradation trend prediction model of the LSTM network in S151, and is a continuous value between 0 and 1.

[0095] Task urgency: Discrete value (high / medium / low), quantified based on task type and time margin.

[0096] Spare parts inventory: a binary variable (sufficient / insufficient). Inventory levels are obtained in real time through the ERP system. If the inventory of critical components is lower than the safety threshold (e.g., ≤2 pieces), it is marked as "insufficient".

[0097] The system's Action space is as follows: α1: Immediate repair, shutdown for maintenance, high cost but significantly improved reliability; α2: Delaying repairs until the next mission results in lower costs, but also accumulates risks; α3: Replace designated components to target predicted failure points, balancing cost and efficiency.

[0098] The system's reward function, Reward, is as follows: (10); in, =0.6, =0.3, =0.1, with weights calibrated by domain experts. This reward function is designed to balance three factors: power supply system task completion efficiency, maintenance costs, and downtime, in order to optimize system health management. By setting different weight values, the function can comprehensively evaluate the system's performance under different health states.

[0099] S1522, DQN model construction; Network Structure: The design comprises three main parts: an input layer, hidden layers, and an output layer. The input layer contains three dimensions representing the system's state vector, specifically including various state features of the environment. The hidden layers consist of two fully connected layers of neurons; the first layer contains 64 neurons and uses the ReLU activation function to extract important feature information from the states. The output layer is a Q-value layer, with three dimensions corresponding to the expected reward for each action. This structure enables the network to effectively evaluate the value of each possible action and make a decision.

[0100] Experience Replay Mechanism: To enhance training efficiency and stability, an ExperienceReplay mechanism is employed. Under this mechanism, the system stores state transition records, including the current state *st*, the executed action *at*, the obtained reward *rt*, and the next state *st+1*. These records are then used for random sampling, breaking down correlations between data and improving training effectiveness. The stored experience data is kept in a buffer of 1,000 records. Each training run randomly selects a batch of samples from this buffer, with each batch size not exceeding 32 records.

[0101] Target Network: This mechanism enhances the stability of the training process. By periodically updating the parameters of the main network in the target network, the target network helps stabilize the Q-value update process, preventing training instability caused by rapid network updates. The target network is not updated with every iteration, but rather at regular intervals, resulting in more stable Q-value estimation and reducing oscillations and instability during training.

[0102] Training and policy optimization: In the training and policy optimization phase, an exploration and exploitation balance approach was adopted.

[0103] In the initial stage, the system explores using an ε-greedy strategy, setting the ε value to 0.9 and gradually decreasing it to 0.1. This approach helps to explore as many possibilities in the environment as possible in the early stages of training, thereby obtaining diverse data and experience and avoiding getting trapped in local optima too early. In the later stages, as learning progresses, the ε value gradually decreases, thereby increasing the utilization of known best strategies, making the maximization of the Q-value the guiding principle for policy selection.

[0104] For updating the Q-value, the classic Q-learning update rule is adopted. Specifically, each update of the state and action follows the formula below: (11); in, The learning rate determines the step size for each update; The immediate reward obtained in the current state; It is a discount factor that controls the degree of influence of future rewards; through this update rule, the Q value is continuously optimized, enabling the system to gradually learn the optimal behavioral strategy in a given environment. s represents the state; in this embodiment, the health index H(t) is used as the state. a represents the action; in the maintenance decision model, the action space includes: a1: immediate maintenance, a2: maintenance delayed until the next task, a3: replacement of a specified component. Indicates the next state. Indicates the next action.

[0105] The system's learning rate is set to 0.01. A low learning rate helps avoid drastic fluctuations during the learning process, ensuring the smoothness of Q-value updates. Discount factor. Setting the Q-value to 0.95 ensures a higher weight for future rewards, further encouraging the system to make more rational decisions in the long run. Finally, the target network is updated every 10 interactions to ensure stable Q-value updates and avoid instability in the training process caused by overly frequent network updates.

[0106] Decision output and application: Two main decision-making strategies were adopted in the strategy output and application stage.

[0107] First, the system uses a real-time decision-making strategy to ensure that the best decision is made based on the latest environmental state at any given moment. Specifically, the system takes the current state s1 as input and outputs the corresponding optimal action a* based on the Q-value function to maximize the current expected reward.

[0108] Secondly, the system dynamically adjusts its strategy by incorporating the prediction results of the LSTM (Long Short-Term Memory) network. By leveraging the temporal prediction capabilities of the LSTM model, the system can adjust its policy based on future states. Specifically, if the LSTM's prediction indicates that the future state is close to the system's minimum threshold, the system will prioritize specific actions α1 or α3, which provide higher stability or optimization near that threshold. This strategy allows the system to not only make decisions based on the current state but also optimize the decision-making process by predicting the future, thereby improving long-term performance and system adaptability.

[0109] For example, for a health degradation trend prediction model based on LSTM networks, given a past N=10 health index sequences, the prediction process for the next T=3 cycles is as follows: Output: H_pred(t+3) = 0.68 ± 0.05; Warning triggered: H_pred<0.75 → generate maintenance work order.

[0110] The DQN maintenance decision-making process is as follows: Status input: H(t)=0.765, Task urgency=High, Spare parts inventory=Sufficient; Action space: α1: Immediate repair, α2: Delayed repair, α3: Replace specified parts; DQN output: Q value is at most α3 → Replace generator controller; Reward value R = 0.6 × (task completion) + 0.3 × (cost control) + 0.1 × (downtime minimization) = 0.72.

[0111] This invention proposes a fuzzy fusion-based method for assessing the health status of power supply systems. By dynamically optimizing weights and using fuzzy quantization techniques, combined with LSTM and deep reinforcement learning (DQN), not only is real-time health status assessment of power supply systems achieved, but also their degradation trends can be predicted and maintenance decisions optimized. This method helps improve the reliability of power supply systems, detect potential failure risks in advance, and provides theoretical support for precise maintenance and the smooth execution of flight missions. In the future, with continuous technological development, this method can be further extended to the health management of other complex systems, promoting the widespread application of intelligent maintenance technology in aviation, energy, and other fields.

[0112] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for assessing the health status of an aircraft power supply system based on dynamic weights, characterized in that: The specific steps are as follows: S1, determine the classification of the aircraft's power supply system; S2, parameter initialization; S3, Establish subjective weight determination model and objective weight determination model for indicators; S4, obtain the weight of each indicator in the finished product parameter level; ; in, The third level of finished product parameters Each indicator in Weight of time, The third level of finished product parameters The subjective weight of each indicator, The third level of finished product parameters Each indicator in Objective weight of time, This is an empirical coefficient; S5, for the first Each indicator in the layer calculates the health index of the indicator. : ; in, For the first Layer Each indicator in Health index at any time In order to be in The first time is composed of the time components The next level of indicators The weight, for The first time is composed of the time components The next level of indicators The standardized probability value, To form the first The next level of indicators for each indicator; S6, Determine If the value is 0, proceed to step S14; otherwise, proceed to step S7. S7, for the first Each indicator in the layer has its subjective and objective weights determined according to S2; S8, obtain the difference between subjective and objective weights and the rate of change in health status; S9, determine the proportion of subjective and objective weights for each indicator based on the subjective and objective weight ratio; S10, perform difference fuzzification processing on the subjective and objective weights to obtain the fuzzified subjective and objective weights of each indicator; S11, obtained by using fuzzy weighting. Weights of each metric in the layer: ; in, and To determine the model's first result based on the ratio of subjective to objective weights Layer The ratio of subjective weight to objective weight for each indicator at time t. For the first Layer Subjective weights after blurring the individual indicators For the first Layer The objective weights of each indicator after fuzzification; S12, for the first The weight values ​​of the same set of indicators in the same layer are normalized; S13, Then return to step S5; S14, the health level of the power supply system is obtained based on the power supply system health index.

2. The method for assessing the health status of an aircraft power supply system based on dynamic weights according to claim 1, characterized in that: The specific hierarchical structure of the S1 aircraft power supply system is as follows: The target layer of the zeroth layer is the aircraft power supply system; the first layer is the system function level, which includes all functional modules of the power supply system divided according to the system functions; the second layer is the finished product level, which is formed by dividing each functional module according to the finished products used; the third layer is the parameter level, which is formed by monitoring the parameters of each finished product.

3. The method for assessing the health status of an aircraft power supply system based on dynamic weights according to claim 1, characterized in that: The initialization of the S2 parameters is as follows: initialization , , and ,in This is the current level number. Indicates the first The first layer Each indicator in The rate of change of the system's health status at any given time. Indicates the first The first layer Each indicator in Constant health status To indicate the first The first layer Each indicator in A state of health at all times.

4. The method for assessing the health status of an aircraft power supply system based on dynamic weights according to claim 1, characterized in that: The subjective weight determination model and the objective weight determination model in S3 are specifically as follows: Subjective weight determination model: By obtaining the judgment matrix of the group to which the indicator belongs, the subjective weight of the indicator is obtained by using the improved analytic hierarchy process, and the logical rationality is verified by the consistency test. For the matrix that does not meet the consistency, the correction coefficient is introduced for iterative optimization, and finally the standardized subjective weight of each indicator in the same group is obtained by normalization. Objective weight determination model: This model obtains weights by acquiring monitoring data related to the indicators. After normalizing the monitoring data into standardized values, the information entropy of the indicators is calculated. A smaller information entropy value indicates greater volatility and higher importance. The entropy value is then converted into weights using the following formula: ; In the formula, , Standardized probability; in, For the first The first layer Each indicator in Information entropy at any given moment For the first The first layer The set of indicator groups to which each indicator belongs. For the first The first layer Each indicator in The standardized probability value at time t. For the first The first layer Each indicator in The total number of measurement points acquired after time step [time]. For the first The first layer Each indicator in The measured value at the measurement point at time t. For the first The first layer Each indicator in Objective weight of time.

5. The method for assessing the health status of an aircraft power supply system based on dynamic weights according to claim 4, characterized in that: Specifically, the difference between subjective and objective weights and the rate of change in health status are obtained in S8 as follows: Difference between subjective and objective weighting : ; in, Indicates the first The first layer Each indicator in Difference between subjective and objective weights at any given moment ; For the first Layer The subjective weight of each indicator, For the first Layer The objective weight of each indicator at time t; Health status change rate : ; in, Defined as a time window, i.e., a time step. Indicates the first The first layer Each indicator in The rate of change of the system's health status at any given time; Indicates the first The first layer Each indicator in Constant health status To indicate the first The first layer Each indicator in A state of health at all times.

6. The method for assessing the health status of an aircraft power supply system based on dynamic weights according to claim 1, characterized in that: In step S9, the proportion of subjective and objective weights for each indicator is determined by the model based on the ratio of subjective to objective weights, specifically as follows: S91, Construct a model to determine the proportion of subjective and objective weights; S92, the first The subjective and objective weight differences, health status change rate, and health index of each indicator in the layer are input into the subjective and objective weight ratio determination model to obtain the first layer. The proportion of subjective and objective weights for each indicator in the layer.

7. The method for assessing the health status of an aircraft power supply system based on dynamic weights according to claim 6, characterized in that: The S91 model for determining the weight ratio of subjective and objective factors is constructed as follows: The ratio of subjective to objective weights determines the input to the model as the state space: ,in, This is the difference between subjective weights and objective weights. for Rate of change in health status at any time for Constant health index The ratio of subjective to objective weights determines the model's output as the action space. ,in and These are the subjective weight proportion coefficients and objective weight proportion coefficients obtained through deep deterministic strategy gradient algorithm optimization, where... and satisfy , The ratio of subjective to objective weights determines the model's reward function: ,in and To balance the hyperparameters of prediction accuracy and weight stability, for Constant health index; The subjective and objective weight ratio determines the model's network structure: an Actor network is used for output. and The adjustment strategy involves using a Critic network to evaluate the Q-value of state-action pairs, which guides Actor optimization.

8. The method for assessing the health status of an aircraft power supply system based on dynamic weights according to claim 1, characterized in that: S10 performs a difference fuzzification process on the subjective and objective weights to obtain the fuzzified subjective and objective weights for each indicator, specifically as follows: Obtain the difference between subjective and objective weights and the rate of change in health status As input quantities, the input quantities are fuzzified through the trapezoidal membership function, and the subjective and objective weights of each indicator are obtained according to the fuzzy rule table. The trapezoidal membership function is as follows: Membership degree This indicates a "low" level of difference. Membership degree , , indicating a "medium" level of difference; Membership degree This indicates a "high" level of difference. The rates of change are categorized as follows: This indicates a smooth or gentle transition; , indicating fluctuation; , indicating intense; Fuzzy rule tables are based on the difference between subjective and objective weights. and rate of change in health status The defined fuzzy rules presuppose subjective and objective weights.

9. The method for assessing the health status of an aircraft power supply system based on dynamic weights according to claim 1, characterized in that: The S12 is for the first The weights of the same set of indicators within the same layer are normalized, specifically as follows: ; in, To and Indicators belonging to the same group To and Add up all the indicators belonging to the same group. For the first The first layer The set of indicator groups to which each indicator belongs. For indicator set One of the indicator serial numbers.

10. The method for assessing the health status of an aircraft power supply system based on dynamic weights according to claim 1, characterized in that: It also includes S15, health prediction and maintenance decision optimization: S151, taking the time series of the power supply system health index as input, uses a health degradation trend prediction model based on LSTM network to predict the health degradation of the power supply system. S152 takes the power supply system health index, task urgency, and spare parts inventory as inputs and uses a maintenance decision optimization model based on deep Q-network and DQN to obtain maintenance decisions.

Citation Information

Patent Citations

  • Power equipment operation state evaluation method and system based on multi-source information fusion

    CN116089897A

  • Method and system for evaluating health state of main equipment of large-scale transformer substation

    CN117454131A

  • Aluminum electrolysis cell health state assessment method based on set pair analysis and evidence theory

    CN118734164A

  • Electric energy storage intelligent management system

    CN120582197A

  • Cable health state assessment method and system

    WO2025241630A1