Method and apparatus for generating training data for PHM model based on digital twin
By constructing a digital twin model with medium fidelity and performing Monte Carlo simulations, training data for the PHM model covering multiple failure modes was generated, solving the problem of data scarcity in aviation systems and improving the accuracy and robustness of fault diagnosis and prediction.
Patent Information
- Application Number
- CN202511156367.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing technologies struggle to efficiently generate PHM model training data covering multiple failure modes in the aviation field, especially in complex failure scenarios where data is scarce, resulting in poor model performance under multiple failure conditions.
A medium-fidelity digital twin model is constructed and calibrated using expert knowledge and historical data. A large amount of state-labeled fault data is generated through Monte Carlo batch simulation, including single fault and compound fault data, covering the coupling relationships and dynamic behaviors of various components of the aerospace system.
It enables the rapid generation of massive and reliable PHM model training data in a virtual environment, covering multiple fault modes, improving the robustness and accuracy of fault diagnosis and prediction, and alleviating the problem of lack of fault samples.
Smart Images

Figure CN120654579B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of aircraft technology, specifically relating to a method and device for generating training data for a PHM model based on digital twins. Background Technology
[0002] Fault prediction and health management (PHM) is one of the key technologies for improving the reliability and safety of aviation equipment. Data-driven PHM models typically rely on a large amount of historical operational and fault data for training to accurately predict and diagnose system health. However, due to the high reliability of aviation systems, the frequency of faults in actual operation is extremely low, making it very difficult to obtain sufficient real-world fault data. Even if faults are deliberately induced through experiments, it is not only costly and poses safety risks, but the resulting sample is still limited and cannot comprehensively cover all fault modes. This data scarcity is particularly prominent in the aviation field, becoming a bottleneck restricting the accuracy and practical application of PHM models.
[0003] To compensate for the lack of real-world data, existing technologies attempt to generate synthetic fault data using simulation methods for model training. However, traditional simulation methods have two limitations: firstly, while low-precision or locally simple models are computationally fast, they cannot realistically reflect the fault evolution process under the coupling effects of various components in complex aerospace systems, resulting in low reliability of simulation data; secondly, while simulations based on high-precision physical models are realistic, their construction and operation costs are extremely high, making them difficult to use for generating large-scale datasets. Digital twin technology offers a new approach to this problem, namely, creating a model in a virtual environment that highly matches the behavior of the real system, allowing for various virtual experiments. However, for complex systems like aerospace, pursuing a completely high-fidelity digital twin model presents difficulties in model construction and computational resource consumption, hindering the efficient generation of massive amounts of data. Conversely, if the model is overly simplified, it is difficult to guarantee the effective representativeness of simulation results for real faults. Therefore, a method that strikes a balance between simulation accuracy and efficiency is urgently needed. Furthermore, existing technologies do not adequately consider the scenario of compound faults (i.e., multiple faults occurring simultaneously), for which data is even scarcer, leading to poor performance of PHM models in handling multi-fault situations. Summary of the Invention
[0004] The present invention aims to at least partially solve one of the technical problems in the aforementioned related technologies.
[0005] Therefore, the purpose of this invention is to provide a method and device for generating training data for a PHM model based on digital twins, which can achieve a balance between simulation accuracy and efficiency.
[0006] To solve the above-mentioned technical problems, the present invention is implemented as follows:
[0007] This invention provides a method for generating training data for a PHM model based on digital twins, the method comprising:
[0008] A medium-fidelity digital twin model is constructed, which comprehensively considers the dynamic characteristics and coupling relationships of various components of the aviation system in order to reproduce the global dynamic behavior;
[0009] The medium-fidelity digital twin model is calibrated based on expert knowledge and historical data to improve accuracy;
[0010] Model simulation is achieved through Monte Carlo batch simulation, and the data is labeled according to the simulation conditions to generate a large amount of PHM model training data.
[0011] In addition, the PHM model training data generation method based on digital twin according to the present invention may also have the following additional technical features:
[0012] In some implementations, the PHM model training data consists of normal data and fault data with state labels, and the fault data includes single fault data and compound fault data.
[0013] In some implementations, the Monte Carlo batch simulation involves conducting multiple simulation tests by randomly changing the initial state, operating conditions, and failure mode parameters of the aerospace system, wherein the failure mode parameters include failure type, occurrence time, and severity; and injecting transient / progressive failures and compound failures.
[0014] In some of these implementations, the composite fault is the simultaneous injection of two or more different types of faults in a single simulation test.
[0015] In some of these implementations, the medium-fidelity digital twin model includes an engine propulsion sub-model, a flight control sub-model, an avionics and sensor sub-model, and an airframe structure dynamics sub-model.
[0016] In some implementations, the medium-fidelity digital twin model preserves interactions with key subsystems of the PHM; including:
[0017] Flight-engine coupling: the effects of aircraft attitude and speed on engine intake and load, and the interaction between engine thrust reaction and flight status;
[0018] Coupling of power generation and consumption: the impact of engine speed on generator output, and the impact of electrical faults on flight control and instruments;
[0019] The impact of structural degradation on performance: the effect of increased drag on the body and structural deformation on handling stability.
[0020] In some of these implementations, the impact of structural degradation on performance is represented with moderate fidelity using parametric perturbations; these include: reducing the slope of the lift curve by introducing a health factor, and characterizing the impact of structural aging on flight performance by increasing the drag coefficient.
[0021] In some of these implementations, the dynamic characteristics of the components include:
[0022] Flight dynamics: Attitude and trajectory are simulated using an open-source 6-DOF nonlinear engine;
[0023] Engine and Propulsion System: Modeling thermodynamics, rotor dynamics, and thrust / fuel consumption using Modelica;
[0024] Power supply and other systems: Modelica was used to build models of the gas supply and hydraulic subsystems to capture key physical evolutions.
[0025] In some implementations, the medium-fidelity digital twin model simplifies non-essential details, including:
[0026] Detailed stress distribution of blades inside the engine;
[0027] The electrical system is simulated using an equivalent pneumatic path model;
[0028] The aerodynamics of the aircraft configuration are based on table lookup data.
[0029] This invention also provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the content of the PHM model training data generation method based on digital twins as described in any of the preceding embodiments.
[0030] Compared with the prior art, the present invention has at least the following beneficial effects:
[0031] In this embodiment of the invention, the PHM model training data generation method based on digital twins can not only reflect the global dynamic behavior under the coupling influence of various components of the aviation system, but also quickly execute a large number of simulation tests to generate training data covering multiple failure modes.
[0032] In this embodiment of the invention, the PHM model training data generation method based on digital twins can obtain a large amount of normal and fault data with state labels in a virtual environment, conduct dangerous and expensive physical fault tests, and significantly alleviate the problem of lack of fault samples in PHM model training.
[0033] In this embodiment of the invention, the PHM model training data generation method based on digital twins provides a modeling method with medium fidelity in terms of simulation efficiency and accuracy. It fully considers the coupling effect between various components of the aviation system, simplifies the model complexity while ensuring a certain level of simulation accuracy, and greatly improves the simulation running efficiency. It supports more than a thousand Monte Carlo batch simulations to generate massive amounts of data.
[0034] In this embodiment of the invention, the PHM model training data generation method based on digital twins provides a way to calibrate the parameters of the digital twin model by introducing expert experience and real historical data in terms of model credibility. This enables the model to accurately reproduce the dynamic behavior characteristics and typical fault symptoms of the actual system, ensuring that the generated data has high credibility and engineering reference value.
[0035] In this embodiment of the invention, the provided method for generating training data for a digital twin-based PHM model has a richness of fault scenarios. The simulation experiments cover a variety of fault modes, including sensor failure, actuator failure, component performance degradation, etc. It can even simulate the occurrence of multiple different faults in a single simulation. Therefore, the generated dataset covers both single and compound fault scenarios, which helps to train a PHM model that can identify complex fault combinations and improves the robustness and accuracy of fault diagnosis and prediction.
[0036] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0037] Figure 1 This is a diagram illustrating the overall architecture of a PHM model training data generation method based on digital twins, as disclosed in an embodiment of the present invention. Detailed Implementation
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0039] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings and specific examples and application scenarios.
[0040] In some embodiments of the present invention, a method for generating PHM model training data based on a medium-fidelity digital twin model of an aviation system is provided, the content of which includes:
[0041] First, a medium-fidelity digital twin model of the aviation system is constructed to reproduce the key components of the actual aviation system, their inter-component coupling relationships, and global dynamic behavior in a virtual environment. This digital twin model is calibrated and refined by incorporating expert knowledge and historical data, enabling it to realistically simulate the aircraft's operational characteristics under both normal and fault conditions.
[0042] Subsequently, large-scale simulation experiments were conducted on the digital twin model using the Monte Carlo simulation method, which involves repeating the simulation multiple times with different initial conditions and operating parameters. In these simulation experiments, some parts operate under normal conditions, while others are simulated according to preset introduction of various fault modes (including single faults and compound faults where multiple faults occur simultaneously), thereby generating rich data on normal and fault behaviors.
[0043] Finally, corresponding state labels (normal or specific fault modes) are added to the acquired simulation data, and the data is aggregated to form a large-scale, multi-label PHM model training dataset. This invention also provides a digital twin system for implementing the above method, including a model building module, a simulation execution module, and a data processing module, used to complete the entire process of establishing and calibrating the digital twin model and generating Monte Carlo simulation data.
[0044] Please see Figure 1 As shown, in some embodiments of the present invention, the digital twin system for generating PHM training data includes core components such as a digital twin model module, a simulation execution module, and a data processing module. The digital twin model module is used to build a virtual model of the aviation system, capable of simulating the working characteristics of the major components of the aircraft and their interactions; the simulation execution module is used to control the operation and experimental configuration of the digital twin model, realizing Monte Carlo batch simulation; and the data processing module is used to collect, process, and annotate the simulation output data.
[0045] Preferably, the system also includes an expert knowledge base and a historical data interface, which are used to store the experience rules of domain experts, failure mode information, and historical data obtained from actual flights or tests. This information is provided to the digital twin model module and the simulation execution module to calibrate the model and guide the simulation test settings so that the model can realistically simulate the normal and failure states of the aviation system.
[0046] In some embodiments of the present invention, the digital twin model employs the Monte Carlo method to perform multiple simulation experiments. Part of the simulation runs under normal operating conditions, while another part introduces preset fault modes, including single faults and composite fault scenarios involving multiple faults simultaneously. The operational data generated from each simulation experiment is collected, and the data is labeled with normal or corresponding fault mode according to the simulation conditions, thus forming a PHM model training dataset.
[0047] First, a digital twin model of the aviation system is established. This model has a medium-fidelity simulation accuracy and comprehensively considers the dynamic characteristics and coupling relationships of the main subsystems of the aircraft. For example, in one embodiment, the digital twin model includes an engine propulsion sub-model, a flight control sub-model, an avionics and sensor sub-model, and an airframe structural dynamics sub-model, which are coupled through internal interfaces and equations to jointly simulate the overall flight dynamics behavior of the aircraft. Each component sub-model can be established based on known physical principles and calibrated using real data: for example, historical flight test data is compared with simulation output results, and the parameters of the engine thrust model are continuously adjusted to ensure that its output under normal operating conditions is consistent with the real performance curve; another example is correcting certain elastic parameters in the structural model according to empirical formulas provided by experts to more accurately reflect the airframe vibration characteristics. Through such calibration measures, the digital twin model achieves a high degree of consistency with the real system in key performance indicators, laying the foundation for subsequent simulations.
[0048] Then, Monte Carlo simulation experiments are conducted based on the calibrated digital twin model. Specifically, the simulation execution module sets the total number of simulations and the parameter range, and randomly selects initial conditions and input parameter combinations at the start of each independent simulation experiment. On the one hand, for normal operating condition simulations, the system randomly changes parameters such as environmental conditions (e.g., ambient temperature, airspeed), flight mission profiles (e.g., load, trajectory), and component performance deviations (e.g., sensor noise, manufacturing tolerances) within a safe range, obtaining data samples covering different normal operating conditions through numerous repeated simulations. On the other hand, for fault condition simulations, while randomizing normal parameters, fault events are introduced into the simulation according to preset probabilities or test plans. Faults can be triggered at random moments on the simulation timeline, or they can be set to exist from the start of the simulation, simulating instantaneous faults or gradual degradation faults as needed. For example, in one simulation, a sensor output can be shifted at a random moment mid-flight to simulate a sensor misalignment fault; or in another simulation, the engine's maximum thrust can be gradually reduced by a certain amount to simulate an engine performance degradation fault. For complex fault scenarios, the simulation execution module can simultaneously inject two or more different types of faults. For example, in the same simulation, both hydraulic system leakage and controller sensor failure can be applied simultaneously to examine their combined impact on aircraft dynamic performance. By systematically and randomly varying various factors and combining different fault modes, this invention can generate simulation data covering a wide range of operating conditions.
[0049] During each simulation, the digital twin model outputs a large amount of data characterizing the system's state and performance, such as sensor readings, operating parameters of key components (temperature, pressure, current, etc.), flight status parameters (speed, altitude, attitude), etc. The data processing module collects and organizes this raw simulation output data in real time. At the end of the simulation, the system records the operational state category based on the simulation settings: if no faults are introduced, it is marked as "normal" data; if the simulation includes a certain fault, it is marked as the corresponding fault type; if multiple faults are included simultaneously, multiple labels or composite labels are created for this combined fault mode. The data processing module associates the labels with the data sequences obtained from the simulation and stores them as samples for training the PHM model. After numerous simulations, the system ultimately compiles a large-scale dataset containing both normal samples and samples of various fault modes. This dataset can be used to train and validate the data-driven PHM algorithm model. For example, the data generated by this invention can be input into machine learning models (such as neural networks, random forests, etc.) for training, allowing them to learn the differences in sensor data under normal and fault conditions, thereby achieving automatic fault detection and identification during deployment. For example, training a prediction model with time-series data that includes the component performance degradation process can improve the accuracy of remaining useful life (RUL) predictions.
[0050] It should be noted that the construction and simulation of the digital twin model of this invention can be achieved using existing mature modeling and simulation platforms, such as flight mechanics simulation software and MATLAB / Simulink modeling environments. It is only necessary to build a corresponding modular model based on the principles of this invention and write fault injection and data acquisition programs. Therefore, the method and system of this invention have strong versatility and can be widely applied to the generation of PHM data for various aircraft (such as airplanes and drones). Furthermore, without departing from the principles of this invention, the specific component divisions, parameter selections, and fault type settings in the above embodiments can be adjusted and transformed according to different application requirements. The scope of protection of this invention is not limited to a specific aircraft or a specific fault type. Through the description of the above embodiments, those skilled in the art can clearly understand the method flow and system architecture of this invention and implement and apply this invention accordingly.
[0051] Example 1:
[0052] This embodiment constructs a digital twin simulation platform to support the training and discovery of evolutionary mechanism-based PHM (Prognostics and Health Management) models in civil aviation aircraft scenarios. The platform emphasizes reproducibility and engineering feasibility, and can integrate the dynamics and maintenance processes of various aircraft subsystems, providing high-quality simulation data for the training and evaluation of PHM algorithms in data-sparse scenarios. The main objectives include:
[0053] 1) Continuous + Discrete Hybrid Simulation: Unifies the simulation of continuous physical processes of aircraft (flight dynamics, engine propulsion, electrical power supply, etc.) and discrete events (maintenance and repair, fault occurrence, scheduling decisions, etc.) to form a system-level hybrid simulation environment.
[0054] 2) Medium-fidelity modeling: A medium-fidelity physical model is adopted, focusing on key coupling mechanisms and global dynamic behavior, rather than an overly detailed CAE model, in order to achieve a balance between accuracy and computational overhead.
[0055] 3) Uncertainty and Monte Carlo: Supports modeling of multi-source uncertainties such as environmental disturbances and component differences, and can perform batch runs of Monte Carlo simulations in multiple scenarios to generate a large amount of diverse data to make up for the lack of real data.
[0056] 4) Fault injection and health evolution: Typical faults (engine thrust loss, electrical power failure, structural fatigue degradation, etc.) can be injected into the simulation to track the performance degradation and remaining life (RUL) evolution of each component.
[0057] 5) Evolutionary PHM training integration: It is tightly integrated with the evolutionary mechanism-based PHM model development process to realize automatic training of model populations, evaluation and selection of fitness based on indicators such as accuracy and "Belief-Reliability", and automatic iterative optimization of models.
[0058] The following details the platform's module design, architecture integration, toolchain selection, and implementation details to ensure that researchers can reproduce and further develop the solution.
[0059] The first aspect is the fusion simulation of continuous physical processes and discrete maintenance events.
[0060] Continuous dynamics simulation: This method uses physical modeling to simulate aircraft flight and continuous system processes. The core components include:
[0061] a. Flight Dynamics: The aircraft's attitude, trajectory, and flight state are simulated using 6-DOF flight dynamics models such as JSBSim. JSBSim is an open-source, nonlinear 6-DoF flight dynamics engine that is data-driven and suitable for batch simulations.
[0062] b. Engine and Propulsion System: Modelica is used to model the thermodynamic and rotor dynamics of aero-engines, as well as thrust output and fuel consumption. As a multi-domain modeling language, Modelica facilitates the creation of one-dimensional, moderately complex models of engines and transmissions, capturing coupling relationships such as speed-thrust-temperature.
[0063] c. Power supply and other physical systems: Modelica can also be used to build electrical system models (engine, APU, battery, etc.) as well as subsystem models such as hydraulics and avionics cooling, to include the physical evolution of key aircraft systems.
[0064] Discrete Events and Maintenance Strategies: This involves introducing maintenance-related discrete behaviors through discrete event simulation (DES) or agent / rule-driven models. The core of these strategies includes:
[0065] a. Fault and Maintenance Events: Simulate random or triggered fault events, as well as maintenance activities such as periodic inspections, repairs, and component replacements. An event scheduling mechanism is used to arrange these discrete events on the timeline and update the system state when an event occurs (e.g., resetting the damage accumulation after replacing a component).
[0066] b. Intelligent Maintenance / Scheduling Strategies: Introduce decision-making algorithms or rules (implemented by agents or policy functions) to trigger maintenance decisions based on system health status and predictive information. For example, preventative maintenance scheduling when the predicted RUL (Relative Usage Limit) is below a threshold, or optimizing overall availability based on multi-fleet scheduling. Simulation platforms allow embedding these strategy algorithms and observing their impact on performance and availability during operation.
[0067] c. Tool Implementation: It is recommended to use multi-method simulation software such as AnyLogic to implement the discrete parts of the model. AnyLogic supports hybrid modeling of multiple methods, including discrete events, agents, and system dynamics, and can easily combine maintenance processes (personnel, resources, task flows) with continuous system state evolution. For example, flowcharts can be used to describe processes such as inspection-repair-release, and linked to continuous degradation processes. If an open-source solution is required, similar functionality can also be achieved using Python's SimPy library or a self-developed event scheduling module.
[0068] Hybrid simulation integration: Enables coordinated operation of continuous and discrete components, with the core components including:
[0069] a. Employing a hybrid scheduling algorithm combining time-stepping and event-driven approaches: Continuous subsystems (flight, engines, etc.) typically integrate with fixed small step sizes (e.g., 0.01~0.1 seconds), while discrete events are sorted in the scheduling queue by trigger time. The proposed hybrid simulation framework checks the next event time in each simulation loop; if no event occurs, it continues the simulation; if an event is imminent, it jumps to the event occurrence time, processes the event logic, and then continues.
[0070] b. Data Exchange: Discrete events (failures) modify the parameters / states of the continuous model. For example, an engine failure event reduces engine thrust (which can be achieved by adjusting the internal efficiency parameters of the engine model). Conversely, the state of the continuous model also affects discrete decisions, such as the maintenance event triggered by the real-time calculated Remaining Life (RUL). A unified data interface needs to be designed to allow the event module to access and modify the key states / parameters of the physical model. A publish-subscribe or global state management mechanism can be used to uniformly manage the state variables of various aircraft subsystems within the simulation. After each event or time step, the state is written to the central storage for other modules to read.
[0071] c. Verify Synchronization: Ensure continuous integration and event handling are strictly synchronized with the simulation clock. For example, use synchronization barriers to dock at each event occurrence point, preventing continuous simulation from exceeding the event time. Tools like AnyLogic internally support synchronization between event scheduling and continuous simulation during hybrid modeling. In custom implementations, events can be actively checked within a Python loop, progressing through time.
[0072] Through the above design, a unified simulation of aircraft operation (continuous dynamics) and maintenance operation (discrete events) is achieved, fully reproducing the actual operating conditions.
[0073] Part Two: Medium-Fidelity Modeling and Preservation of Key Couplings.
[0074] Definition of Medium Fidelity: Medium-fidelity models fall between simple low-fidelity empirical models and complex high-fidelity CAE models. Their goal is to accurately reproduce the system's critical behavior while avoiding excessive computational load. According to Emerson's Digital Twin Simulation Guide, medium-fidelity models typically employ first-order principle physical models with appropriate simplifications to capture mass, energy conservation, and key dynamic characteristics. Compared to purely data-driven models, they offer more reliable extrapolation capabilities for out-of-range conditions.
[0075] Key Coupling Mechanisms: To ensure that PHM-related behaviors are simulated, the model needs to retain the key couplings and global dynamics between subsystems; the core includes:
[0076] a. Flight-Engine Coupling: Aircraft attitude and speed affect engine intake and load, which in turn affect aircraft motion. This interaction needs to be modeled. The proposed solution is to exchange parameters between the JSBSim flight model and the Modelica engine model. For example, current airspeed and air density can be passed to the engine model to calculate pressure ratio and thrust, while the thrust and torque generated by the engine can be fed back to the flight model to calculate acceleration.
[0077] b. Coupling of power generation and consumption: Engine speed affects generator output, and electrical faults can affect flight control / instrumentation systems. Power and speed parameters need to be exchanged between the engine and electrical models. For example, the Modelica electrical system sub-model uses engine speed / torque as input to calculate power generation, outputs bus voltage to supply other systems, and simulates battery charging and discharging.
[0078] c. Structural degradation affects performance: such as increased drag and structural deformation affecting handling stability. This part can be represented with moderate fidelity using parametric perturbation: by introducing a "health factor" to reduce the slope of the lift curve and increase the drag coefficient, the impact of structural aging on flight performance can be simply characterized. In this way, even without detailed finite element analysis, the impact of degradation on global dynamics can be reflected.
[0079] Simplify unnecessary details: Discard minor details that have no significant impact on PHM training. For example:
[0080] a. The detailed blade stress distribution inside the engine does not require CAE modeling; damage progression can be represented by an overall efficiency and flow degradation coefficient.
[0081] b. The electrical system does not use phase-by-phase electromagnetic simulation, but adopts an equivalent circuit model (pressure source + internal resistance, etc.).
[0082] c. Aircraft aerodynamics can be derived from lookup table data (Wind Tunnel database or JSBSim's built-in XML) instead of real-time CFD.
[0083] By focusing on key coupling and degradation effects and simplifying minor details, the model improves simulation speed while maintaining predictive relevance, making it suitable for conducting a large number of Monte Carlo experiments.
[0084] Part Three: Multi-source Uncertainty Modeling and Monte Carlo Data Generation.
[0085] Sources of uncertainty: The platform models multi-source uncertainty through parameter randomization and perturbation injection, including but not limited to:
[0086] a. Environmental uncertainty: The external environment changes randomly, such as atmospheric temperature, air pressure, wind disturbances, and turbulence intensity. During simulation, random disturbances can be applied to the atmospheric model, such as encountering turbulent gusts of varying intensities or temperature deviations, which will affect the engine and aerodynamics.
[0087] b. Individual differences: Manufacturing and wear variations among different aircraft / components. Initial health parameters can be randomly selected, such as a slightly lower initial efficiency for a particular engine, or different initial damage factors for certain structures.
[0088] c. Operational and Load Variations: Pilot operations and mission profiles vary, with slight differences in load, flight path, and flight profile for each flight, affecting component stress spectra. Different flight mission scripts (including climb / cruise / descent time, altitude, and power settings) can be extracted in the simulation.
[0089] d. Measurement and model uncertainty: Sensor noise, model simplification error, etc. can also be reflected by superimposing noise on the output signal.
[0090] Monte Carlo scene generation: Supports batch automated simulation of a large number of scenes, achieving data sampling coverage; core components include:
[0091] a. Parameter combination sampling: Using methods such as Monte Carlo random sampling or Latin hypercube sampling, the aforementioned uncertain parameters are sampled multiple times to generate hundreds or thousands of sets of scenario parameters. Each set of parameters defines a possible aircraft, environment, and usage scenario.
[0092] b. Parallel Simulation: Distribute simulation tasks using Python's parallel libraries (such as Ray). Ray can execute each scene simulation as a remote task, accelerating it through multi-core or cluster parallelism. This allows for the generation of large-scale samples within an acceptable timeframe.
[0093] c. Automatic Data Logging: Each simulation run automatically outputs the required data (sensor timing, fault occurrence time, RUL curve, maintenance action records, etc.). The platform saves the data in a unified format, such as using time-series CSV or SQL database records (refer to the PHM Society UAV framework, which supports automatic recording and verification of simulation data). Ensure that metadata (such as parameter seeds and fault types) is bound to the results so that the training algorithm is traceable.
[0094] By simulating a large number of possible scenarios using Monte Carlo simulations, the PHM model training can obtain a rich and diverse dataset, including different failure modes, different initial health states, and degradation processes under different environmental stresses. This is crucial for improving the robustness and generalization ability of the algorithm in data-sparse scenarios.
[0095] Part Four: Fault Injection Mechanisms and Health Status / RUL Tracking.
[0096] Typical Fault Mode Injection: The platform will incorporate simulation injection mechanisms for various representative aviation system faults; the core includes:
[0097] a. Engine Thrust Loss / Degradation: Thrust loss is simulated by adjusting engine model parameters. For example, gradually reducing compressor and combustion efficiency curves simulates engine performance degradation, leading to a decrease in maximum available thrust. A sharp drop in efficiency can represent sudden failures (such as compressor stall or turbine damage), while a gradual, small decrease represents wear and degradation. NASA's C-MAPSS model uses exponential decay of flow and efficiency parameters to generate degradation failure data.
[0098] b. Electrical system power failure: Simulates situations such as generator failure or bus power outage. At a certain moment in the simulation, a generator in the electrical network may be triggered to zero output, switching the load to the backup system and causing a short-term voltage drop. Alternatively, a rapid battery discharge fault may occur, causing a continuous drop in bus voltage; check if emergency power supply has been activated.
[0099] c. Structural / Component Fatigue Degradation: A fatigue life model (such as the Paris crack growth formula or a simple life deduction model) is set for the structural components. Damage accumulates at each time step based on the stress spectrum; component failure is marked when the damage crosses a threshold. Alternatively, a performance degradation event of a component (such as a hydraulic pump) can be directly triggered at a certain point in the simulation to represent its wear reaching a critical point.
[0100] Therefore, the health index of the component is recorded in the simulation as it decreases over time.
[0101] d. Sensor or control failures: Although the focus is on predicting RUL, sensor misalignment and controller failure can also be simulated to enrich the training scenarios of the PHM algorithm (these are more complex faults and can be included optionally).
[0102] RUL and Health Metric Tracking: Establishing health status variables for each key component and calculating RUL; core components include:
[0103] a. Health Index (HI): Defined as a health index ranging from 0 to 1 or a percentage, where 1 represents a new, undamaged component, and a decrease to 0 indicates failure. The HI is updated in the simulation based on the degradation model. For example, the HI of an engine compressor blade can be estimated based on the efficiency reduction rate (an efficiency decrease of X% corresponds to a HI decrease of Y), while the HI of structural components is calculated based on cumulative fatigue cycle consumption.
[0104] b. Remaining Service Life (RUL): When health indicator trends are predictable, RUL is the estimated remaining operating time (or number of cycles) at the current moment. The true RUL can be calculated directly in simulations using a known degradation model. For example, if the simulation assumes a component will fail when HI drops to 0, the time until HI=0 can be estimated based on the current HI value and the rate of decline as the true RUL. For random failures, the life limit can be randomly selected at the start of the simulation, and then a countdown can be performed during operation.
[0105] c. Performance Evolution Records: In addition to HI and RUL, corresponding performance indicators also need to be tracked, such as the curves of engine thrust and fuel consumption as a function of degradation, and the curves of battery capacity decay as a function of cycles. This can both verify the PHM model's ability to predict performance degradation and make decisions based on this in subsequent maintenance strategy simulations (such as requiring maintenance when performance drops to a threshold).
[0106] d. Data Interface: The platform records the HI and RUL of all components as special outputs over time and provides an interface to the PHM algorithm. This means the PHM model can be trained offline using simulated HI / RUL as labels, or its accuracy can be evaluated by comparing it with the simulated real RUL during online testing.
[0107] Through fault injection and health tracking mechanisms, the simulation platform can generate operational data "tagged with faults and lifetimes," meeting the needs of PHM algorithm development for degradation process and lifetime information. As Saxena et al. pointed out, obtaining complete data on real systems from health to failure is very difficult, and simulation injection provides a feasible way to generate such data for model training and validation.
[0108] Part 5, PHM training closed-loop integration based on evolutionary mechanism.
[0109] Evolutionary PHM model training process: The platform supports embedding the simulation environment into the automated training and optimization loop of the PHM model, forming a closed loop of digital twin + evolutionary algorithm:
[0110] a. Model population initialization: First, a set of candidate PHM models is generated (which can be neural networks with different structures, random forests, or models with different hyperparameters, etc.). These models, as individuals in the population, will be iteratively optimized through evolutionary algorithms.
[0111] b. Model Evaluation Metrics: Define the model fitness decision function, where accuracy (such as the inverse function of RUL prediction error, or fault prediction accuracy, etc.) and metrics such as Belief-Reliability jointly constitute the objective. Belief-Reliability is a reliability evaluation metric that comprehensively considers uncertainty; in this context, it can be understood as the confidence / credibility evaluation of the model in prediction, which may be achieved through the credibility or confidence score of the prediction interval. Therefore, the fitness decision can be set as multiple objectives (maximum accuracy or maximum confidence / reliability) to balance model performance and robustness.
[0112] c. Parallel Simulation Evaluation: Each model is evaluated using data / environment generated by a simulation platform. For example, for each candidate PHM model, a batch of Monte Carlo scenarios (or samples from a previously generated dataset) are run to allow the model to predict RUL or diagnose faults, and then its prediction accuracy and Belief-Reliability metrics are calculated. Due to the potentially large population size and number of scenarios, this step strongly relies on parallel computing. Ray is used to distribute the evaluation tasks for parallel execution, significantly reducing evaluation time.
[0113] d. Selection and propagation: Based on the evaluation results, perform evolutionary algorithm operations on the model population (e.g., select models with high fitness to enter the next generation, eliminate inferior models; perform crossover and mutation on selected models to generate new models). Here, a multi-objective evolutionary algorithm (such as NSGA-II) can be used to select the best among the two objectives of accuracy and reliability, or the two can be weighted into a single score using a simple GA.
[0114] e. Automatic Iteration: The above evaluation and propagation steps are repeated, continuously generating new models and using simulation testing to evaluate and eliminate inferior models until the termination condition (upper limit of the number of generations or performance convergence) is met. The final output is the best-performing PHM model.
[0115] Process integration method: To achieve the above closed loop, it is necessary to solve the interface integration between the simulation platform and the evolutionary algorithm framework:
[0116] a. Training Control Script: A main control script is written in Python, encapsulating the calls to the simulation platform (generating data or running a complete simulation) into easily usable functions / services. For example, a `run_simulation(config) -> results` interface can be provided, returning evaluation results given model and scene parameters.
[0117] b. Evolutionary Algorithm Library: Existing libraries can be used (e.g., DEAP, Nevergrad, or Ray's built-in Tune module supporting evolution, etc.). These libraries allow for convenient population and iteration management in Python. The main control script is responsible for calling the library to generate model parameter sets in each generation, then calling the simulation evaluation function for each parameter set to obtain fitness, and finally feeding the fitness back to the library for selection.
[0118] c. Distributed Architecture: During large-scale evolution, Ray's actor and task models can be used to distribute simulation tasks of different model-scene combinations to multiple workers for parallel processing. Ray can also manage GPU resources to accelerate PHM model inference and training. If the simulation itself is time-consuming, hundreds of concurrent simulations can be easily scheduled to explore the model space through Ray Tune's trial parallel mechanism.
[0119] Results analysis and adaptive iteration: Adaptive mechanisms can be incorporated into the integration process.
[0120] a. If the model exhibits diversity during the monitoring of evolution, and the population converges prematurely, new random individuals can be introduced to increase diversity.
[0121] b. Adjust the distribution of simulation scenarios based on simulation feedback (so-called "targeted evolution"). For example, if the model performs poorly in a certain type of scenario, generate more data for that type of scenario for subsequent focused examination, thereby guiding the algorithm to improve its weak points.
[0122] c. Record the performance and characteristics of each generation of models to analyze the evolution path and reliability improvement process, thereby providing insights into PHM model design (i.e., achieving the "model discovery" mentioned in the problem, and discovering excellent model structures and configurations in simulation experiments).
[0123] Through the aforementioned evolutionary training, the digital twin simulation platform not only provides data but also participates in the closed-loop model optimization, supporting the automatic design and optimization of the PHM model. Introducing the Belief Reliability metric in this process helps the evolutionary algorithm tend to select models that are both accurate and robust, avoiding models that overfit specific data but are not robust to uncertainty.
[0124] Other parts of this invention not described in detail can be referred to in the prior art or are well known to those skilled in the art. This embodiment does not limit these aspects and will not describe them in detail here.
[0125] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.
Claims
1. A method for generating training data for a PHM model based on digital twins, characterized in that, The method includes: A medium-fidelity digital twin model is constructed, which comprehensively considers the dynamic characteristics and coupling relationships of various components of the aviation system in order to reproduce the global dynamic behavior; The medium-fidelity digital twin model is calibrated based on expert knowledge and historical data to improve accuracy; Model simulation is achieved through Monte Carlo batch simulation, and the data is labeled according to the simulation conditions to generate a large amount of PHM model training data. The medium-fidelity digital twin model includes an engine propulsion sub-model, a flight control sub-model, an avionics and sensor sub-model, and an airframe structure dynamics sub-model. The medium-fidelity digital twin model preserves the interactions of key subsystems of PHM, including: Flight-engine coupling: the effects of aircraft attitude and speed on engine intake and load, and the interaction between engine thrust reaction and flight status; Coupling of power generation and consumption: the impact of engine speed on generator output, and the impact of electrical faults on flight control and instruments; The impact of structural degradation on performance: the effect of increased drag on the body and structural deformation on handling stability.
2. The method for generating training data for a PHM model based on digital twins according to claim 1, characterized in that, The training data for the PHM model consists of normal data and fault data with state labels, and the fault data includes single fault data and compound fault data.
3. The method for generating training data for a PHM model based on digital twins according to claim 1, characterized in that, The Monte Carlo batch simulation involves conducting multiple simulation tests by randomly changing the initial state, operating conditions, and failure mode parameters of the aviation system. The failure mode parameters include the failure type, occurrence time, and severity. Transient or progressive failures and compound failures are also injected.
4. The method for generating training data for a PHM model based on digital twins according to claim 3, characterized in that, The composite fault refers to the simultaneous injection of two or more different types of faults in a single simulation test.
5. The method for generating training data for a PHM model based on digital twins according to claim 1, characterized in that, The impact of structural degradation on performance is represented with moderate fidelity using parameter perturbation methods, including: reducing the slope of the lift curve by introducing a health factor, and characterizing the impact of structural aging on flight performance by increasing the drag coefficient.
6. The method for generating training data for a PHM model based on digital twins according to claim 1, characterized in that, The dynamic characteristics of each component include: Flight dynamics: Attitude and trajectory are simulated using an open-source 6-DOF nonlinear engine; Engine and propulsion system: Modeling thermodynamics, rotor dynamics, and thrust or fuel consumption using Modelica; Power supply and other systems: Modelica was used to build models of the gas supply and hydraulic subsystems to capture key physical evolutions.
7. The method for generating training data for a PHM model based on digital twins according to claim 1, characterized in that, The medium-fidelity digital twin model simplifies non-essential details, including: Detailed stress distribution of blades inside the engine; The electrical system is simulated using an equivalent pneumatic path model; The aerodynamics of the aircraft configuration are based on table lookup data.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the content of the PHM model training data generation method based on digital twins as described in any one of claims 1-7.
Citation Information
Patent Citations
Digital twin driven complex equipment fault prediction method
CN111008502A