Data analysis method and device, equipment and storage medium
By normalizing multi-source life data of aero-engines and parallel fitting of various distribution models, the optimal model is automatically selected, solving the bias problem in multi-source data processing and improving the accuracy and efficiency of aero-engine reliability analysis.
Patent Information
- Application Number
- CN202511501799.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies are ill-suited for mixed censoring scenarios involving multi-source aero-engine data, and lack distribution analysis and parameter estimation methods, resulting in systematic biases in reliability analysis results.
By acquiring multi-source life data of aero-engines, performing data processing and normalization, using multiple distribution models for parallel fitting, automatically selecting the optimal distribution model based on the fitting results, and calculating reliability indicators.
It improves the accuracy and efficiency of reliability analysis, supports standardized processing and automated analysis of multi-source data, reduces analytical bias, and enhances the precision of aero-engine safety assessment and operational decision-making.
Smart Images

Figure CN121502189A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of aircraft engine reliability analysis technology, and in particular to a data analysis method, apparatus, device, and storage medium. Background Technology
[0002] Reliability analysis of aircraft engines is a crucial step in ensuring aviation safety and improving operational efficiency. However, in practical analysis, engine maintenance data is often scattered across multiple channels such as fault logs and maintenance records, requiring multi-source data fusion to achieve a comprehensive and accurate assessment. The inventors have discovered that existing technologies suffer from at least the following problems: they are ill-suited to the mixed censoring scenarios common in multi-source data, and lack distribution analysis and parameter estimation methods for such complex data, leading to systematic biases in reliability analysis results.
[0003] Therefore, improving the accuracy of reliability index calculation has become an urgent technical problem to be solved. Summary of the Invention
[0004] The purpose of this application is to provide a data analysis method, apparatus, device, and storage medium that can effectively improve the accuracy of reliability index calculation.
[0005] To achieve the above objectives, a first aspect of this application provides a data analysis method, comprising: Acquire multi-source life data of aero-engines and process the data to obtain a regularized dataset containing component identification, failure time with data type, and failure attributes; Based on the regularized dataset, multiple distribution models are used for parallel fitting, including distribution models with location parameters; Based on the fitting results, the distribution model is selected to obtain the optimal distribution model; The reliability index is calculated based on the optimal distribution model, and the reliability index is output.
[0006] Compared with existing technologies, the data analysis method provided in this application has the following advantages: By acquiring and integrating multi-source lifetime data and processing it into a regular dataset containing component identifiers, data types, failure times, and failure attributes, it solves the problem of scattered and mixed multi-source data and the need for manual processing of mixed and censored data, providing a standardized basis for analysis; at the same time, it uses multiple distribution models, including those containing location parameters, for parallel fitting, and automatically selects the optimal distribution model based on the fitting results, breaking through the limitations of existing technologies that rely on a single distribution and require manual model selection, thus improving the adaptability and accuracy of distribution fitting and parameter estimation; finally, it calculates and outputs reliability indicators based on the optimal model, effectively reducing the analytical bias caused by non-standard data processing and poor model adaptability in existing technologies, improving the accuracy and efficiency of reliability analysis, and more accurately supporting the safety assessment and operational decision-making of aero-engines.
[0007] In some embodiments, the acquisition and processing of multi-source life data of the aero-engine to obtain a regularized dataset containing component identifiers, failure times with data types, and failure attributes includes: Acquire multi-source lifespan data of an aero-engine, wherein the multi-source lifespan data includes at least component identification, failure time, and failure attributes; The failure time of each data point in the multi-source lifetime data is classified, and the data type of each data point is determined to obtain a regularized dataset containing component identifiers, failure times with data types, and failure attributes; wherein, the data type includes at least one of complete data, left-censored data, right-censored data, and interval-censored data.
[0008] In some embodiments, classifying the failure time of each data point in the multi-source lifetime data and determining the data type of each data point includes: When the fault time is the time of component replacement due to fault, it is determined to be complete data; When the fault time is the planned replacement time of the component or normal operating data, it is determined to be right-censored data; When the fault time only indicates that the fault occurred before a certain moment but the specific time cannot be determined, it is identified as left-censored data. When the fault time is clearly defined, and the fault occurs between two specific time points, it is determined to be interval censored data.
[0009] In some embodiments, the multi-source lifetime data is derived from at least one of replacement records, maintenance records, dismantling records, borehole monitoring data, wing monitoring data, and planned maintenance interval data.
[0010] In some embodiments, the multiple distribution models include at least two of the following: exponential distribution, Weibull distribution, normal distribution, log-normal distribution, logistic distribution, Gumbel distribution, Gamma distribution, and mixed Weibull distribution, and at least one of the Weibull distribution, log-normal distribution, logistic distribution, and Gamma distribution is a three-parameter distribution model.
[0011] In some embodiments, the step of selecting the optimal distribution model based on the fitting results includes: The fitting results of each distribution model are initially screened based on the goodness-of-fit test; The distribution models after initial screening are ranked by weighted calculation using AIC, BIC, CC, and LKV indicators, and the distribution model ranked first is selected as the optimal distribution model.
[0012] In some embodiments, the reliability metrics include at least one of reliability-related metrics, lifespan-related metrics, and risk measurement-related metrics; The reliability-related indicators include the failure distribution function (CDF), probability density function (PDF), reliability, and conditional reliability. The life-related metrics include MTTF, MTBF, reliable life, BX life, conditionally reliable life, and mean remaining life. The risk measurement indicators include failure rate and cumulative risk function (CHF).
[0013] To achieve the above objectives, a second aspect of this application provides a data analysis apparatus, the apparatus comprising: The data processing module is used to acquire multi-source life data of aero-engines and process the data to obtain a regularized dataset containing component identification, failure time with data type, and failure attributes. The distribution fitting module is used to perform parallel fitting using multiple distribution models based on the regularized dataset, including distribution models with location parameters. The model selection module is used to select the optimal distribution model based on the fitting results. The indicator output module is used to calculate the reliability indicator based on the optimal distribution model and output the reliability indicator.
[0014] To achieve the above objectives, a third aspect of this application provides an electronic device, the electronic device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the method described in the first aspect.
[0015] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium comprising a stored computer program, wherein the computer program, when executed, controls the device containing the computer-readable storage medium to perform the method described in the first aspect.
[0016] To achieve the above objectives, a fifth aspect of the present application provides a computer program product, which includes a computer program or computer instructions, wherein the computer program or computer instructions, when executed by a processor, implement the method described in the first aspect. Attached Figure Description
[0017] Figure 1 This is a flowchart of a data analysis method provided in an embodiment of this application; Figure 2 yes Figure 1 A flowchart of step S101 in the process; Figure 3 yes Figure 1 A flowchart of step S103 in the process; Figure 4 A schematic diagram of the data analysis device provided in this application embodiment; Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0019] In the description of this application, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0020] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0021] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0022] First, let's analyze some of the terms used in this application: Total Time Since New (TSN): This refers to the cumulative time a part has been in use since it left the factory (new condition) until now.
[0023] Total Cycle Since New (CSN): This refers to the cumulative number of cycles a component has undergone from its manufacturing (new part condition) to the present (such as cycles related to lifting, landing, or heat).
[0024] Time Since Repair (TSR): refers to the cumulative usage time of a component from its most recent repair to the present.
[0025] Cycle Since Repair (CSR): refers to the cumulative number of cycles a component has undergone since its most recent repair.
[0026] Time Since Overhaul (TSO): refers to the cumulative service time of a component from its most recent overhaul to the present.
[0027] Cycle Since Overhaul (CSO): refers to the cumulative number of cycles a component has undergone since its most recent overhaul.
[0028] Aircraft engine reliability analysis is a crucial step in ensuring aviation safety and improving operational efficiency, and its accuracy is critical to the sustainable development of the aviation industry. In practical applications, engine maintenance data is often scattered across multiple channels, including fault logs, replacement records, borescope records, and maintenance records, exhibiting a multi-source characteristic. Only through multi-source data fusion can a comprehensive and accurate assessment of engine reliability be achieved. However, existing technologies have significant limitations in processing such multi-source lifespan data, making it difficult to meet practical analytical needs.
[0029] Specifically, existing technologies have the following prominent problems in the reliability analysis of aero engines: First, the processing of multi-source operation and maintenance data lacks unified standards and specifications. Faced with lifetime data scattered across multiple channels such as fault logs and replacement records, existing technologies lack clear criteria for selecting key data such as CSN, TSN, TSR, and CSR. There is also a lack of unified classification standards for defining right-censored, complete, and interval-censored data, and data cleaning operations lack standardized procedures. This results in mixed data benchmarks across different analysis scenarios, severely impacting the robustness of analysis results and making it difficult to achieve effective alignment and standardized integration of multi-source data.
[0030] Secondly, the ability to estimate the distribution parameters of multi-source fault data is insufficient. In engineering practice, fault data of engine components often includes mixed types such as complete data, right-censored data, left-censored data, and interval-censored data. However, existing technologies lack distribution analysis and parameter estimation methods for such mixed censored data, which can easily lead to systematic biases in reliability analysis results. At the same time, existing technologies do not adequately support models such as logistic distribution with location parameters, log-normal distribution, and three-parameter Gamma distribution, which cannot adapt to the analysis needs of components with different failure mechanisms, limiting the realization of refined reliability assessment, and thus affecting the accuracy of remaining service prediction and maintenance decisions.
[0031] Third, the analysis process lacks standardized and scientific guidance. Existing technologies and related products provide weak guidance for the analysis process, especially lacking customized guidance for the aero-engine field, making it difficult to effectively implement the technology. In the selection of distribution models and analysis results, there is an over-reliance on human experience, and it is impossible to achieve automated and scientific selection through goodness-of-fit tests and comprehensive evaluation of multiple indicators, which restricts the guiding value of the analysis results for production practice.
[0032] For example, the technical solution disclosed in patent CN116611251A only analyzes a single fault mode. The data input is limited to complete observation data or right-censored data for a specific fault. Left-censored data and interval censored data need to be manually removed or transformed, and automated processing cannot be achieved. In terms of distribution model selection, this technology always uses a single distribution (such as a two-parameter Weibull distribution or exponential distribution). Parameter estimation is only solved once by maximum likelihood or least squares fitting. The output results cannot be flexibly switched to other distributions or censoring types. When facing multiple fault modes or multiple censoring scenarios, repeated operations are required, resulting in extremely poor flexibility.
[0033] Furthermore, while the PosWeibull reliability analysis software in related technologies supports the analysis of various distribution types and different data types, it has significant shortcomings in data processing: although it can import "complete data", "complete + right-censored data" and "single interval censored data", it requires separate fitting through separate modules and cannot process all censoring types at once; in terms of parameter estimation methods, it only relies on proprietary numerical optimizers (such as the quasi-Newton method), the iterative optimization scheme is singular, and it is prone to computational lag; in the fitting process, users need to manually select the distribution and censoring type, and after the system runs multiple sets of fittings, users have to compare indicators such as AIC and BIC to select the optimal solution, which relies too much on human experience and lacks guidance.
[0034] In summary, existing technologies generally suffer from problems such as inability to handle mixed scenarios of multiple types of censored data, poor adaptability of distributed models, and reliance on human experience in the analysis process, which restrict the accuracy and efficiency of aero-engine reliability analysis.
[0035] Therefore, improving the accuracy of reliability index calculation has become an urgent technical problem to be solved.
[0036] Please see Figure 1 , Figure 1 This is an optional flowchart of the data analysis method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S104.
[0037] Step S101: Obtain multi-source life data of the aero-engine and process the data to obtain a regularized dataset containing component identification, failure time with data type, and failure attributes. Step S102: Based on the regular dataset, multiple distribution models are used for parallel fitting, including distribution models with location parameters; Step S103: Based on the fitting results, the distribution model is selected to obtain the optimal distribution model; Step S104: Calculate the reliability index based on the optimal distribution model and output the reliability index.
[0038] Steps S101 to S104, as illustrated in this embodiment, acquire and integrate multi-source lifetime data and process it into a regular dataset containing component identifiers, data types, failure times, and failure attributes. This solves the problem of scattered and mixed multi-source data and the need for manual processing of mixed censored data, providing a standardized foundation for analysis. Simultaneously, multiple distribution models, including those containing location parameters, are used for parallel fitting, and the optimal distribution model is automatically selected based on the fitting results. This overcomes the limitations of existing technologies that rely on a single distribution and require manual model selection, improving the adaptability and accuracy of distribution fitting and parameter estimation. Finally, reliability indicators are calculated and output based on the optimal model, effectively reducing analytical biases caused by non-standard data processing and poor model adaptability in existing technologies. This improves the accuracy and efficiency of reliability analysis and can more accurately support the safety assessment and operational decisions of aero-engines.
[0039] Prior to step S101 in some embodiments, because an aircraft engine may contain one or more failure modes, the data analysis method further includes determining the type of engine failure mode: If it is a single type of fault mode, then proceed directly to step S101; If there are multiple fault modes and there is no need to consider fault coupling relationships, then in S104, a mixed Weibull distribution model can be used for fitting. If there are multiple failure modes and the failure coupling relationship needs to be considered, the system reliability analysis is performed by combining the Reliability Block Diagram (RBD) technology, and the corresponding reliability indicators are finally output.
[0040] Therefore, the data analysis method of this application is applicable to data analysis of single-type failure modes and data analysis of multiple failure modes without considering complex coupling relationships.
[0041] In step S101 of some embodiments, multi-source life data refers to a dataset reflecting the life status of aero-engine components, which comes from multiple channels such as engine replacement records, maintenance records, disassembly records, borehole monitoring data, and on-wing monitoring data.
[0042] Please see Figure 2 In some embodiments, step S101 may include, but is not limited to, steps S201 to S202: Step S201: Obtain multi-source life data of the aero-engine. The multi-source life data includes at least component identification, failure time, and failure attributes. Step S202: Classify the failure time of each data point in the multi-source lifetime data, determine the data type of each data point, and obtain a regularized dataset containing component identifiers, failure times with data types, and failure attributes.
[0043] In step S201 of some embodiments, the multi-source lifetime data is derived from at least one of replacement records, maintenance records, disassembly records, borehole monitoring data, wing monitoring data, and planned maintenance interval data.
[0044] It should be noted that the component identifier (i.e., component / accessory system identifier) can be a unique identifier used to distinguish different engine components. The failure time can be the specific time the component failed (e.g., flight hours) or the number of cycles (e.g., flight cycles). Failure attributes can include information such as failure type (e.g., fracture, creep) and failure cause (e.g., material fatigue, improper maintenance). By integrating multi-channel data generated during engine operation and maintenance, information including component identifiers, failure times, and failure attributes is collected to ensure that the data covers complete lifespan-related dimensions.
[0045] It should be noted that engine replacement records, maintenance records, replacement records, borescope monitoring data, on-wing monitoring data, and planned maintenance interval data can all be used to obtain fault time point information. Among them, engine replacement records record the issuance time and reason for the issuance of engine components in the fleet; engine maintenance record data includes the maintenance time, replacement time, maintenance level, and maintenance facility information of specific components; replacement records clarify the replacement time of faulty parts; borescope monitoring anomaly data are records related to component anomalies discovered through borescope detection; real-time on-wing monitoring anomaly data are anomaly data monitored in real time when the engine is running on the wing; and planned maintenance intervals for components are pre-set standard maintenance time intervals for components.
[0046] In step S202 of some embodiments, the data type includes at least one of complete data, left-censored data, right-censored data, and interval-censored data. The failure time of each data point in the multi-source lifetime data is classified to determine the data type of each data point, including: When the failure time is the time when the component fails and is replaced, it is determined to be complete data; When the failure time coincides with the planned replacement time of the component or normal operating data, it is determined to be right-censored data; When the fault time only indicates that the fault occurred before a certain moment but the specific time cannot be determined, it is identified as left-censored data. When the fault time is clear and the fault occurs between two specific time points, it is determined to be interval censored data.
[0047] It should be noted that the failure times in the normalized dataset are typically any combination of four data types: complete data, left-censored data, interval-censored data, and right-censored data. The data analysis method in this application supports lifetime analysis for three different data combinations: complete data only; a combination of complete data and right-censored data; and a comprehensive analysis of complete data, left-censored data, interval-censored data, and right-censored data. For specific aero-engine operation data, reasonable classification can be performed based on the business scenario to adapt to the corresponding analysis logic.
[0048] Therefore, by classifying the failure time of each data point in the multi-source lifetime data and determining the data type of each data point, a standardized dataset containing component identification, failure time with data type, and failure attributes can be obtained. It should be noted that the standardized dataset covers at least three dimensions: component / accessory system identification, failure time, and failure attributes. Specifically, this standardized dataset includes standardized failure information, operational status data, maintenance history data (including parameters such as TSN, CSN, TSR, CSR, TSO, and CSO), and real-time recorded normal operation status data.
[0049] In practical engine engineering, various types of censored fault data exist. For example, a high-pressure turbine blade in an engine may simultaneously contain complete data (clearly recording the specific time when the component failed), right-censored data (the component was still operating normally at a certain time, but the failure time was not observed), left-censored data (it is known that the component failed before a certain time, but the specific failure time is unknown), and interval-censored data (clearly indicating that the failure occurred between two observation times). Therefore, based on the recording characteristics of fault times in multi-source lifetime data, the fault time of each data point in multi-source lifetime data can be classified into one or more of the following types according to the aforementioned definition: complete data, left-censored data, right-censored data, and interval-censored data. For example, please refer to Table 1, which is an example of the classification of fault times in multi-source lifetime data in one embodiment.
[0050]
[0051] In some embodiments, after obtaining a regularized dataset containing component identifiers, failure times with data types, and failure attributes, step S101 further includes: preprocessing the regularized dataset. The preprocessing includes: verifying the data validity of the regularized dataset, removing empty data, abnormal data, and non-target failure data; unifying the data baseline of the regularized dataset to ensure consistency in component operating conditions, configurations, and materials corresponding to the data; and determining the data measurement type of the regularized dataset, where the measurement type includes time data or cyclic data.
[0052] Specifically, firstly, the basic validity of the regularized dataset needs to be verified. The dataset must meet the basic data volume requirements for parameter distribution fitting, and null values, outliers, and non-target fault data caused by accidental factors (such as FOD or bird strikes) must be removed. Secondly, it must be ensured that the components corresponding to the data are consistent in operating conditions, configuration, and materials to eliminate interference from differences in basic attributes on the analysis results. Finally, the data measurement type (time data or cyclic data) must be clearly defined, depending on the fault type: cyclic data is used for faults strongly correlated with takeoff and landing and thermal cycles; time data can be used for component faults weakly correlated with aircraft takeoff and landing; in some scenarios, a combination of both is required for analysis.
[0053] In step S102 of some embodiments, the distribution model can be a mathematical model describing the probability distribution of lifetime data, including exponential distribution, Weibull distribution, log-normal distribution, Gamma distribution, etc. Distribution models with location parameters refer to distributions that include location parameters in addition to shape and scale parameters (such as three-parameter Weibull distribution, three-parameter Gamma distribution, log-logistic distribution with location parameters, etc.), which can more accurately fit specific failure patterns.
[0054] The multiple distribution models include at least two of the following: exponential distribution, Weibull distribution, normal distribution, log-normal distribution, logistic distribution, Gumbel distribution, Gamma distribution, and mixed Weibull distribution, and at least one of the following: Weibull distribution, log-normal distribution, logistic distribution, and Gamma distribution is a three-parameter distribution model.
[0055] Based on a regularized dataset, multiple distribution models are used for parameter estimation simultaneously. A unified likelihood function is constructed to handle mixed censored data, resulting in multiple sets of fitting results. These distribution models include: exponential distribution (single parameter), exponential distribution (two parameters), Weibull distribution (two parameters), Weibull distribution (three parameters), normal distribution (two parameters), lognormal distribution (two parameters), lognormal distribution (three parameters), loglogistic distribution (two parameters), loglogistic distribution (three parameters), Gumbel distribution (two parameters), Gamma distribution (two parameters), Gamma distribution (three parameters), and mixed Weibull distribution.
[0056] In step S103 of some embodiments, the fitting result can be multiple fitting results formed by simultaneously estimating parameters based on regular datasets and multiple distribution models.
[0057] Please see Figure 3 In some embodiments, step S103 may include, but is not limited to, steps S301 to S302: Step S301: Perform preliminary screening of the fitting results of each distribution model based on the goodness-of-fit test; Step S302: Combine AIC, BIC, CC and LKV indices for weighted calculation, sort the distribution models after initial screening, and select the distribution model with the highest ranking as the optimal distribution model.
[0058] In step S301 of some embodiments, the fitting results of various distribution models are initially screened based on goodness-of-fit tests (such as KS test, AD test, etc.), eliminating models that obviously cannot fit the data features and retaining potential suitable models. In step S302 of some embodiments, for the models after initial screening, a weighted calculation is performed by combining AIC (Akaike Information Criterion, which measures the balance between model fit and complexity), BIC (Bayesian Information Criterion, which focuses more on penalizing complex models), CC (Correlation Coefficient, which reflects the correlation between model predictions and actual data) and LKV (Log Likelihood Value, which reflects the model's ability to interpret data) indicators (the weights can be dynamically adjusted according to the analysis scenario), and the models are ranked according to the comprehensive score, and the model ranked first is determined as the optimal distribution model.
[0059] It should be noted that for the cleaned and well-organized dataset, distribution recommendation provides initial guidance from two levels: first, it recommends potentially suitable distributions from the perspective of fault mechanisms; second, it presents direct probability indicators through nonparametric estimation. In the specific distribution fitting stage, it first checks whether the data volume meets the fitting requirements of the corresponding distribution; then, it uses goodness-of-fit tests to initially screen some distributions; finally, it combines weighted calculations using indicators such as AIC, BIC, CC, and LKV to rank the fitting results and ultimately recommend the optimal distribution.
[0060] It should be noted that parallel fitting of multiple distribution models (including models with location parameters) based on a regularized dataset essentially ensures that different distribution hypotheses are fully matched with the data characteristics, avoiding the loss of optimal solutions due to the limitations of a single distribution hypothesis. Secondly, by selecting the optimal model through the fitting results (combining goodness-of-fit tests and indicators such as AIC and BIC), the rationality of the model can be verified from both statistical significance and practical adaptability. Finally, the distribution model with the strongest explanatory power and best generalization ability for the current data is selected from multiple candidate models. Therefore, the embodiments of this application not only ensure the full mining of data features but also avoid subjective selection bias through objective indicators, making it particularly suitable for complex lifetime data analysis scenarios containing multiple types of censoring.
[0061] In step S104 of some embodiments, the reliability index includes at least one of reliability-related index, life-related index, and risk measurement-related index. Reliability-related metrics include failure distribution function (CDF), probability density function (PDF), reliability, and conditional reliability. Lifetime-related metrics include MTTF, MTBF, reliable life, BX life, conditionally reliable life, and mean remaining life. Risk measurement indicators include failure rate and cumulative risk function (CHF).
[0062] It should be noted that, given sufficient data, the aforementioned distributions will calculate the following reliability-related indicators. As shown in Table 2, reliability indicators can be categorized as follows, which can be used by system users for comprehensive purposes, such as providing relevant fault warnings, setting and optimizing the gradient of inspection intervals, etc.
[0063]
[0064] It should be noted that the Cumulative Distribution Function (CDF) (unreliability) describes the product's lifespan as not exceeding [a certain threshold]. The probability of the product in The probability of failure occurring before a certain time is also known as unreliability in reliability.
[0065] Probability Density Function (PDF): For continuous-lifetime random variables, it describes the probability density function of a system at a specific time. The probability density of failure.
[0066] Reliability (no specific abbreviation, commonly used) (Indicates) The product is within the specified time. The probability of performing the specified function under certain conditions is related to the unreliability as follows: Failure distribution function .
[0067] Conditional Reliability (no specific abbreviation, commonly used) (Representation): The known system is in If it has not expired before the specified time, from arrive The probability that no failure will occur during this period.
[0068] Mean Time to Failure (MTTF) and Mean Time Between Failures (MTBF) are the mathematical expectation of a product's lifespan, representing the average operating time. MTTF refers to the mean time to failure before failure for unrepairable products, while MTBF refers to the mean time between failures for repairable products.
[0069] Reliable Life (no specific abbreviation, commonly used) (Note: To ensure product reliability is not lower than) Working hours; of which The reliable lifespan at that time is called the median lifespan. BX lifetime has a failure probability of ); Lifespan, i.e. .
[0070] Conditional Reliable Life (no specific abbreviation, commonly used) (Representation): The known system is in If the time limit is not expired, from The remaining lifespan at which the system can continue to operate (the duration corresponding to a specific reliability requirement).
[0071] Mean Residual Life (MRL): Given the current survival time At that time, the expected value of the product's remaining lifespan.
[0072] Failure rate (also known as hazard rate, no specific abbreviation, commonly used) (Indicated): Work until Products that have never failed are, The probability of failure occurring within a unit of time after a given time.
[0073] Cumulative Hazard Function (CHF): Describes the system's risk over time. The total amount of risk accumulated previously, that is, from 0 to The sum of failure rates.
[0074] In one specific embodiment, a reliability analysis system for multi-source aero-engine life data has been developed based on the data analysis method of this application. First, the system facilitates the input and selection of multi-source aero-engine life data through an interactive interface, resolving data alignment issues and demonstrating its ability to fuse and standardize multi-source data. Then, the system can filter distribution models through the interactive interface to obtain the optimal distribution model. Furthermore, the visualization area in the middle of the interactive interface can be used to display graphs such as probability density function (PDF) and failure distribution function (CDF). Combined with multi-distribution dynamic parallel fitting technology, it supports fitting and automatic recommendation of various common distributions, accurately presenting the distribution characteristics under different failure mechanisms and meeting the needs of mixed analysis in multiple business scenarios. The indicator areas on the right and bottom of the interactive interface can be used to calculate and display a series of reliability indicators such as reliability and average life, helping to assess the risk of unplanned replacement and optimize maintenance resources. Finally, through a digital and visual interface design, engineers can efficiently obtain analysis results without coding, based on a standardized process transformed from expert knowledge. This promotes the implementation of aero-engine reliability analysis from technical solutions to practical systems, improving the scientific and economical nature of maintenance decisions. For example, Table 3 shows a sample of multi-source heterogeneous data related to the lifespan of aero-engines being entered into the reliability analysis system, which is used to demonstrate the system's compatibility and processing capabilities for complex data such as multi-type censored data.
[0075]
[0076] This application first standardizes multi-source lifespan data, such as fault logs and lifespan data, through multi-source data fusion technology. It innovatively unifies the encoding of various types of censored data and fits them to the same scenario, solving data alignment problems, achieving automated and standardized processing, reducing human error, and laying the foundation for component analysis adapted to different failure mechanisms. Based on this, it supports mixed analysis of data from multiple business scenarios. Utilizing multi-distribution dynamic parallel fitting technology, it automatically selects the optimal distribution by combining nonparametric estimation and indicators such as AIC / BIC, achieving accurate distribution fitting and parameter estimation. Simultaneously, it standardizes expert knowledge through zero-code operation, couples maintenance business requirements through highly integrated design, and enhances engineering practicality with a digital visualization interface and user-friendly interaction. Ultimately, it can calculate a series of reliability indicators, effectively reducing unplanned replacement rates, optimizing maintenance resources and costs, and ensuring the economic efficiency of aviation operations and flight safety.
[0077] Please see Figure 4 This application also provides a data analysis apparatus that can implement the above-described data analysis method. The apparatus includes: The data processing module 401 is used to acquire multi-source life data of the aero-engine and process the data to obtain a regularized dataset containing component identification, failure time with data type, and failure attributes. The distribution fitting module 402 is used to perform parallel fitting of multiple distribution models based on a regular dataset, including distribution models with location parameters. The model selection module 403 is used to select the distribution model based on the fitting results and obtain the optimal distribution model. The indicator output module 404 is used to calculate the reliability index based on the optimal distribution model and output the reliability index.
[0078] The specific implementation of this data analysis device is basically the same as the specific implementation of the data analysis method described above, and will not be repeated here.
[0079] Thirdly, embodiments of this application provide an electronic device, see [link to relevant documentation]. Figure 5 The diagram shown is a structural schematic of an electronic device provided in this application.
[0080] like Figure 5 As shown, the device includes: Memory 31 is used to store computer programs; Processor 32 is used to execute computer programs; When the processor 32 executes a computer program, it implements the data analysis method as described in any of the above embodiments.
[0081] For example, a computer program may be divided into one or more modules / units, one or more of which are stored in memory 31 and executed by processor 32 to complete this application. One or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in an electronic device.
[0082] The processor 32 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0083] The memory 31 can be used to store computer programs and / or modules. The processor 32 implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory 31 and calling the data stored in the memory 31. The memory 31 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 31 may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0084] It should be noted that the aforementioned electronic devices include, but are not limited to, processors and memory, as will be understood by those skilled in the art. Figure 5 The structural diagram is merely an example of the electronic device described above and does not constitute a limitation on the electronic device. It may include more components than shown in the diagram, or combine certain components, or use different components.
[0085] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed, implements the data analysis method of any of the above embodiments.
[0086] It should be understood that all or part of the processes in the above-described data analysis method can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described data analysis method. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content of the computer-readable medium can be appropriately added to or subtracted according to the requirements of legislation and patent practice in the relevant jurisdiction. For example, in some relevant jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0087] Fifthly, embodiments of this application also provide a computer program product, which is stored in a storage medium and executed by at least one processor to implement the data analysis method of any of the above embodiments.
[0088] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0089] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.
Claims
1. A data analysis method, characterized in that, include: Acquire multi-source life data of aero-engines and process the data to obtain a regularized dataset containing component identification, failure time with data type, and failure attributes; Based on the regularized dataset, multiple distribution models are used for parallel fitting, including distribution models with location parameters; Based on the fitting results, the distribution model is selected to obtain the optimal distribution model; The reliability index is calculated based on the optimal distribution model, and the reliability index is output.
2. The method as described in claim 1, characterized in that, The process of acquiring and processing multi-source lifespan data of aero-engines to obtain a regularized dataset containing component identifiers, failure times with data types, and failure attributes includes: Acquire multi-source lifespan data of an aero-engine, wherein the multi-source lifespan data includes at least component identification, failure time, and failure attributes; The failure time of each data point in the multi-source lifetime data is classified, and the data type of each data point is determined to obtain a regularized dataset containing component identifiers, failure times with data types, and failure attributes; wherein, the data type includes at least one of complete data, left-censored data, right-censored data, and interval-censored data.
3. The method as described in claim 2, characterized in that, The process of classifying the failure time of each data point in the multi-source lifetime data and determining the data type of each data point includes: When the fault time is the time of component replacement due to fault, it is determined to be complete data; When the fault time is the planned replacement time of the component or normal operating data, it is determined to be right-censored data; When the fault time only indicates that the fault occurred before a certain moment but the specific time cannot be determined, it is identified as left-censored data. When the fault time is clearly defined, and the fault occurs between two specific time points, it is determined to be interval censored data.
4. The method as described in claim 1, characterized in that, The multi-source lifetime data comes from at least one of the following: replacement records, maintenance records, disassembly records, borehole monitoring data, wing monitoring data, and planned maintenance interval data.
5. The method as described in claim 1, characterized in that, The multiple distribution models include at least two of the following: exponential distribution, Weibull distribution, normal distribution, log-normal distribution, logistic distribution, Gumbel distribution, Gamma distribution, and mixed Weibull distribution, and at least one of the Weibull distribution, log-normal distribution, logistic distribution, and Gamma distribution is a three-parameter distribution model.
6. The method as described in claim 1, characterized in that, The step of selecting the optimal distribution model based on the fitting results includes: The fitting results of each distribution model are initially screened based on the goodness-of-fit test; The distribution models after initial screening are ranked by weighted calculation using AIC, BIC, CC, and LKV indicators, and the distribution model ranked first is selected as the optimal distribution model.
7. The method as described in claim 1, characterized in that, The reliability indicators include at least one of reliability-related indicators, life-related indicators, and risk measurement-related indicators. The reliability-related indicators include the failure distribution function (CDF), probability density function (PDF), reliability, and conditional reliability. The life-related metrics include MTTF, MTBF, reliable life, BX life, conditionally reliable life, and mean remaining life. The risk measurement indicators include failure rate and cumulative risk function (CHF).
8. A data analysis device, characterized in that, include: The data processing module is used to acquire multi-source life data of aero-engines and process the data to obtain a regularized dataset containing component identification, failure time with data type, and failure attributes. The distribution fitting module is used to perform parallel fitting using multiple distribution models based on the regularized dataset, including distribution models with location parameters. The model selection module is used to select the optimal distribution model based on the fitting results. The indicator output module is used to calculate the reliability indicator based on the optimal distribution model and output the reliability indicator.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the data analysis method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the data analysis method as described in any one of claims 1 to 7.