Flight parameter co-occurrence fault prediction method and device based on naive Bayes algorithm

By applying the Naive Bayes algorithm in the symbiotic fault prediction of flight parameters, the problem of large amount of data and high participation of experts in the existing technology is solved, and efficient and accurate fault prediction is achieved, reducing costs.

CN119962357APending Publication Date: 2025-05-09CHENGDU HANLAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510026010.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing flight failure prediction technologies require a large amount of training data, are difficult to model, and require continuous participation of experts, resulting in high costs and low efficiency.

Method used

The flight parameter symbiotic fault prediction method based on the Naive Bayes algorithm is adopted. By collecting all the fault combinations in the flight parameters, a continuous sample space is constructed, the symbiotic fault information is converted into unsupervised learning training data, and a naive Bayes fault prediction model is constructed to achieve probability prediction of target faults.

Benefits of technology

Reduces dependence on large amounts of training data and expert participation, simplifies the model construction process, reduces overall construction and maintenance costs, and improves the efficiency and accuracy of fault prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962357A_ABST
    Figure CN119962357A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of flight parameter data processing, in particular to a flight parameter co-occurrence fault prediction method and device based on a naive Bayesian algorithm, and the method comprises the steps: collecting all fault combinations in flight parameters, and constructing a continuous sample space with a time attribute; symbiotic fault information in the flight parameters is converted into training data of unsupervised learning, and a naive Bayes fault prediction model is constructed; obtaining target fault information to be predicted; and inputting the target fault information into the naive Bayesian fault prediction model, and outputting the probability of occurrence of the target fault. According to the method, the naive Bayesian algorithm is taken as a basic idea, faults occurring in flight at the same time are counted, and a target fault in flight parameters and other faults in symbiotic association with the target fault are converted into training data which can be learned by a machine learning model by combining'labels' and'features' in machine learning; the method for predicting the occurrence probability of the target fault by taking the historical symbiotic fault as the correlation factor is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of flight parameter data processing, and in particular to a flight parameter symbiotic fault prediction method and device based on a naive Bayes algorithm. Background Art

[0002] In the aircraft flight parameters (hereinafter referred to as flight parameters), multiple faults may exist simultaneously in a certain period of time, and there may be correlations between the faults that occur simultaneously. Due to the complexity of the aircraft system itself and the volatility of the external environment, this fault correlation can hardly be described using detailed and accurate rules. A fault information is actually an abstraction of the state of a part of the aircraft's related components. This abstraction actually encapsulates two aspects of information: data and the interactive behavior between internal and external related components. Therefore, correlation modeling and analysis at the fault level is actually a way to understand the aircraft status in a more dynamic, modular and holistic way.

[0003] Usually a complex system consists of several levels and different subsystems, and there are different degrees of penetration or association between different levels and subsystems. When the system structure reaches a certain complexity and is placed in another complex system, such as when an aircraft is placed in a complex atmospheric system, the subsystems or components that were not originally associated at the aircraft system design level will also have indirect associations due to the connection of complex factors. The impact of an association on the system itself is not determined by whether the association is direct or indirect, nor by whether the association is explainable. In practical applications, the priority of predicting possible failures and taking preventive measures in advance is much higher than understanding the causal logic of a failure. Determined causal logic usually needs to start from several elements with fixed attribute states, such as the tree and state of several components in a subsystem. However, for a running system, the environment in which each element is located and the system itself are constantly changing. The causal logic framework established in the initial state may become invalid with the changes of each element. At this time, what is needed is not static certainty, but the judgment that is constantly revised as the state information (data) is continuously updated. This is the core idea of ​​Bayes' theorem.

[0004] Currently, there are data-driven methods, physical model-based methods, and statistical analysis methods for in-flight fault prediction. The data-driven method uses real-time sensor data and information from monitoring equipment to accurately monitor and analyze aircraft status. For example, Airbus's Skywise system uses big data analysis and machine learning technology to monitor and analyze aircraft sensor data in real time to predict faults and optimize maintenance plans. By using machine learning algorithms, large amounts of data can be processed to identify potential failure modes. However, this method requires a large amount of data and data labeling to train the model, and the quantity and quality of labeled data have a great impact on the model, and the final effect is uncertain. Each round of iterative training of the model requires a high cost of manpower and material resources, and it takes a long time to obtain a new model and evaluate the model effect.

[0005] The physical model-based method uses the physical characteristics of the system for prediction and has a certain degree of interpretability. For example, MathWorks' Simulink is a tool for system modeling and simulation that can be used to develop physical models of aircraft systems and perform fault prediction and evaluation. Through mathematical modeling and simulation technology, we can better understand the operating mechanism of aircraft systems and help accurately predict potential faults. However, establishing an accurate physical model requires accurate system parameters and models, and requires a high level of understanding of the system structure. This work requires a lot of manpower investment. In addition, in actual applications, aircraft systems are usually very complex, and accurate modeling is a challenge that requires long-term support from experts and is costly.

[0006] Methods based on statistical analysis are relatively simple and easy to implement. They can use historical failure data for statistical analysis to identify the probability distribution and trend of failures. For example, IBM's Predictive Maintenance and Quality (PMQ) uses machine learning and statistical analysis techniques to perform fault prediction and maintenance optimization solutions.

[0007] However, the above methods and products are highly complex. The early implementation of functions requires a lot of expert knowledge, and the subsequent maintenance and upgrades also require continuous participation of experts, resulting in high overall construction and maintenance costs. In addition, the use of large platform products requires a lot of migration work to drive the implementation of fault prediction functions, and it is difficult to easily migrate or integrate single-point functions into existing systems. Summary of the invention

[0008] In view of this, the purpose of the present invention is to provide a flight parameter symbiotic fault prediction method and device based on the Naive Bayes algorithm, so as to at least solve the problems existing in the existing flight fault prediction technology, such as the need for a large amount of training data, difficulty in modeling, and the need for continuous participation of experts.

[0009] The present invention solves the above technical problems by the following technical means:

[0010] In a first aspect, an embodiment of the present invention provides a flight parameter symbiotic fault prediction method based on a naive Bayes algorithm, comprising the following steps:

[0011] Collect all fault combinations of flight parameters to construct a continuous sample space, wherein the fault combination contains at least two faults and the continuous sample space has a time attribute;

[0012] Based on the continuous sample space, the co-occurrence fault information in the flight parameters is converted into training data for unsupervised learning, and a naive Bayes fault prediction model is constructed;

[0013] Obtain target fault information to be predicted;

[0014] The target fault information is input into a naive Bayes fault prediction model, and the probability of the target fault occurring is output.

[0015] In combination with the first aspect, in some possible implementations, based on the continuous sample space, the symbiotic fault information in the flight parameters is converted into training data for unsupervised learning to construct a naive Bayes fault prediction model, including:

[0016] Perform multiple random sampling on the fault data in the continuous sample space and split them into training sets and test sets according to the preset ratio;

[0017] Convert the symbiotic fault information in the flight parameters of the training set into training data for unsupervised learning;

[0018] Use the training data for unsupervised learning training and establish a naive Bayes fault prediction model;

[0019] The naive Bayes fault prediction model is validated using the test set.

[0020] In combination with the first aspect, in some possible implementations, the preset ratio is 8:2.

[0021] In combination with the first aspect, in some possible implementations, the symbiotic fault information in the flight parameters of the training set is converted into training data for unsupervised learning as follows:

[0022] The faults that occurred simultaneously in the flight parameter data frame are regarded as associated faults. In the flight parameter data of the training set, faults that have only occurred alone are excluded. For a single fault X, there may be one or more faults that occurred simultaneously as the associated fault set S of X. X is taken as the object to be predicted, and the fault data in S is taken as the basis for prediction, that is, X is taken as the label to be predicted, and the fault information in S is taken as the feature of X.

[0023] In combination with the first aspect, in some possible implementations, when the target fault corresponds to an associated fault, the probability calculation formula for the target fault to occur is as follows:

[0024]

[0025] Among them, fault A is the target fault, P(A|B) is the conditional probability that needs to be solved, that is, the probability of fault A occurring after fault B occurs; P(B) represents the statistical probability of fault B occurring independently; P(A∩B) represents the statistical probability of A and B occurring at the same time.

[0026] In combination with the first aspect, in some possible implementations, when the target fault corresponds to two associated faults, the probability calculation formula for the target fault to occur is as follows:

[0027]

[0028] Among them, fault A is the target fault, P(A|B∩C) represents the probability of fault A occurring when faults B and C occur simultaneously;

[0029] P(A) represents the proportion of the number of frames in which fault A occurs in the flight parameters to the total number of all data frames;

[0030] P(B|A) represents the ratio of the number of frames where fault B occurs to the number of frames where fault A occurs;

[0031] P(C|A) represents the proportion of frames where fault C occurs to the number of frames where fault A occurs;

[0032] It indicates the ratio of the number of frames in which fault A does not occur in the flight parameters to the total number of all data frames;

[0033] It indicates the ratio of frames where fault B occurs to frames where fault A does not occur.

[0034] Indicates the ratio of frames where fault C occurs to frames where fault A does not occur.

[0035] In combination with the first aspect, in some possible implementations, when the target fault corresponds to any associated faults, assuming that according to a large number of historical flight parameters, there are n faults {M1, M2, ... Mn} that are all related to the occurrence of the target fault A, the occurrence probability P(A│M1∩M2∩...Mn) of the target fault A is to be calculated, that is, when these faults occur at the same time, the probability calculation formula of the occurrence of fault A is as follows:

[0036]

[0037] Among them, P(M 1 |A) indicates that in the number of frames where fault A occurs, fault M 1 The proportion of frames that occurred;

[0038] P(M 2 |A) indicates that in the number of frames where fault A occurs, fault M 2 The proportion of frames that occurred;

[0039] P(M n |A) indicates that in the number of frames where fault A occurs, fault M n The proportion of frames that occurred;

[0040] It indicates the ratio of the number of frames in which fault A does not occur in the flight parameters to the total number of all data frames;

[0041] P(A) represents the proportion of the number of frames in which fault A occurs in the flight parameters to the total number of all data frames;

[0042] It indicates the ratio of the number of frames in which fault M1 occurs to the number of frames in which fault A does not occur;

[0043] It indicates the ratio of the number of frames in which fault M2 occurs to the number of frames in which fault A does not occur;

[0044] It indicates the ratio of the number of frames where fault Mn occurs to the number of frames where fault A does not occur.

[0045] In the second aspect, an embodiment of the present invention further provides a flight parameter symbiotic fault prediction device based on a naive Bayes algorithm, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the symbiotic fault prediction method described in the first aspect when executing the computer program.

[0046] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the symbiotic fault prediction method described in the first aspect are implemented.

[0047] One or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages:

[0048] The flight parameter co-occurring fault prediction method based on the naive Bayes algorithm of the present invention targets the fault information contained in the flight parameters, takes the naive Bayes algorithm as the basic idea, statistically learns the fault information that has appeared simultaneously in the flight, and combines the basic concepts of "label" and "features" in machine learning to convert the target fault prediction in the flight parameters and other faults with symbiotic associations into training data that can be learned by the machine learning model, thereby realizing a method for predicting the probability of occurrence of the target fault with historical symbiotic faults as correlation factors.

[0049] The flight parameter symbiotic fault prediction method based on the Naive Bayes algorithm of the present invention uses historical fault data as training data, which eliminates a large amount of labeling work that requires the participation of experts. Only a small amount of preprocessing of the original data is required to obtain a large amount of training data. The basic idea of ​​the data processing and model building method mentioned in the present invention is simple, and the early implementation does not require a lot of expert knowledge. The update and maintenance of functions are more related to data updates, and do not require continuous expert participation. The overall construction and maintenance costs are low. In addition, the versatility and data-oriented characteristics of its underlying method enable it to model data of different types and formats, and can easily migrate or integrate the constructed new functions into the existing system. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flow chart of the flight parameter symbiotic fault prediction method based on the Naive Bayes algorithm;

[0051] Figure 2 It is a conceptual diagram of symbiotic failure. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0053] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of objects. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" refers to two or more, for example, multiple processing units refer to two or more processing units, etc., multiple elements refer to two or more elements, etc.

[0054] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0055] Artificial Intelligence (AI): It is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0056] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0057] Natural Language Processing (NLP): It is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0058] Machine Learning (ML): It is a multi-disciplinary interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0059] Machine learning is one of the important branches of artificial intelligence, which can be divided into supervised learning, unsupervised learning and semi-supervised learning. Among them, unsupervised learning is a type of machine learning method that does not require labeled data and completes learning tasks by analyzing the inherent structure and patterns of data. Unlike supervised learning, unsupervised learning does not rely on labeled data, but models it through the distribution and characteristics of the data itself. Unsupervised learning mainly includes the following tasks:

[0060] (1) Clustering: grouping similar data points to reveal the intrinsic structure and patterns of the data.

[0061] (2) Dimensionality reduction: projecting high-dimensional data into a low-dimensional space while maintaining the main features of the data to facilitate data visualization and subsequent analysis.

[0062] (3) Anomaly Detection: Identify abnormal points or outliers in the data to discover potential abnormal situations or erroneous data.

[0063] (4) Association Rule Mining, which discovers the associations and patterns between data items, is often used in areas such as market basket analysis and inventory management.

[0064] In the embodiments of the present application, the artificial intelligence technologies mainly involved include the above-mentioned machine learning, natural language processing (NLP) and other directions. For example, it may involve text processing, semantic understanding, etc. in natural language processing; it may also involve deep learning (deeplearning) in machine learning (ML), including autoencoders, artificial neural networks, etc., which are not specifically described in the embodiments of the present application.

[0065] As a multivariate data frame that is continuous on a macroscopic time scale but discrete on a microscopic time scale, the flight parameter data stream provides a dynamic sample space containing multi-dimensional information, which is very suitable for the Naive Bayes algorithm. Therefore, the Naive Bayes algorithm can be used to model and analyze the correlation between faults, and to predict the probability of a subsequent fault based on historical fault co-occurrence data at a specific moment.

[0066] The flight parameter co-occurring fault prediction method based on the naive Bayes algorithm of the present invention targets the fault information contained in the flight parameters, takes the naive Bayes algorithm as the basic idea, statistically learns the fault information that has appeared simultaneously in the flight, and combines the basic concepts of "label" and "features" in machine learning to convert the target fault prediction in the flight parameters and other faults with symbiotic associations into training data that can be learned by the machine learning model, thereby realizing a method for predicting the probability of occurrence of the target fault with historical symbiotic faults as correlation factors.

[0067] For details, please refer to Figure 1 The flight parameter symbiotic fault prediction method based on the naive Bayes algorithm of the present invention comprises the following steps:

[0068] Step 100, collect all fault combinations in the flight parameters and construct a continuous sample space, where the fault combination contains at least two faults and the continuous sample space has a time attribute.

[0069] Step 200, based on the continuous sample space, convert the co-occurrence fault information in the flight parameters into unsupervised learning training data, and construct a naive Bayes fault prediction model.

[0070] Step 300: Obtain target fault information to be predicted.

[0071] Step 400: input the target fault information into the Naive Bayes fault prediction model, and output the probability of the target fault occurring.

[0072] The naive Bayes algorithm is used to analyze the correlation between faults in the flight data, and the probability of the occurrence of a fault at the next moment is predicted based on the associated fault of an unreported fault. It can be understood as using the associated fault of a certain fault as "evidence" to predict the occurrence or non-occurrence of the fault, and this evidence has both positive and negative aspects. For example, when a fault A has co-appeared with the three faults X, Y, and Z in the historical flight data, then XYZ can be regarded as the feature of fault A, and the reporting or non-reporting of fault A can be regarded as a label or annotation (0 or 1). The co-occurrence here not only refers to the appearance of A, X, Y, and Z in the same frame of flight data, but can be any combination of A and the three. If A and Z faults only appear together once, and Z fault appears in large numbers when A fault is not reported, then Z fault is both favorable evidence that A fault may occur and reverse evidence that A fault will not occur. However, since Z appears more often when A fault is not reported, the existence of Z fault will only increase the predicted probability of A fault after calculation according to the naive Bayes algorithm. If in subsequent flight parameters, a large number of data frames in which Z and A appear simultaneously begin to appear from a certain moment, then the strength of Z as positive evidence for predicting the appearance of A will begin to increase. This dynamic method of continuously adjusting the prediction value based on new evidence also reflects the core idea of ​​the Bayesian correlation algorithm.

[0073] In the actual modeling process, using a single associated fault as a feature to predict another fault is one-sided, but using multiple associated faults can capture more reliable correlations between faults. The Naive Bayes algorithm simplifies the probability calculation of different combinations when there are multiple pieces of evidence, and can make better predictions for specified labels when the amount of data is insufficient.

[0074] Since the processing and generation of training data is relatively easy, and the model training method is easy to understand, the present invention can form a module with both versatility and independence, which can be applied to different aircraft models and different flight parameter format data. The present invention uses historical fault data as training data, which eliminates a lot of labeling work that requires the participation of experts, and only needs to do a small amount of preprocessing on the original data to obtain a large amount of training data.

[0075] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0076] In step 100, all possible fault combinations in flight are used as sample space, which contains at least two faults, and the fault information contained in each frame of flight parameter data is used as a single sample point constituting the sample space. Combined with the characteristics of flight parameters being continuous on a macroscopic time scale, discrete on a microscopic scale, and continuously accumulating in quantity, it is abstracted into a continuously expanding continuous sample space containing time attributes, and the basic conceptual framework for algorithm application is established.

[0077] In step 200, based on the continuous sample space, the co-occurrence fault information in the flight parameters is converted into training data for unsupervised learning, and a naive Bayes fault prediction model is constructed, including:

[0078] Step 210 , perform multiple random sampling on the fault data in the continuous sample space, and split them into a training set and a test set according to a preset ratio of 8:2.

[0079] Step 220 , converting the symbiotic fault information in the training set flight parameters into training data for unsupervised learning.

[0080] Step 230: Perform unsupervised learning training using the training data to establish a naive Bayes fault prediction model.

[0081] Step 240: Use the test set to verify the naive Bayes fault prediction model.

[0082] In step 220, the faults that have occurred simultaneously in the flight parameter data frame are regarded as associated faults. In a certain volume of flight parameter data, faults that have only occurred alone are excluded. For a single fault X, there may be one or more faults that have occurred simultaneously as an associated fault set S of X. At this time, if X is taken as the object to be predicted and the fault data in S is taken as the basis for prediction, then conceptually X is taken as the label to be predicted and the fault information in S is taken as the feature of X.

[0083] Based on the above abstract logic, the symbiotic fault information in the flight parameters is converted into training data that can be used for unsupervised learning. For example, suppose there are the following 10 frames of flight parameter data, the horizontal t is different time, the vertical is the fault data of different faults, 0 represents no fault report, 1 represents fault report, and the statistics are shown in Table 1:

[0084]

[0085]

[0086] Table 1

[0087] In Table 1, for each fault from A to G, data can be established for prediction. Taking fault C as an example, how to understand and establish training data for fault C is shown in Table 2:

[0088] t1 t2 t3 t4 t5 t6 t7 t8 t9 t10 A 0 0 0 1 1 1 1 0 0 0 B 1 0 1 1 1 0 0 0 1 1 C 1 1 1 1 0 0 0 1 0 0 D 0 0 0 0 0 0 1 1 1 1 E 0 0 0 0 0 0 0 0 0 1 F 0 1 1 1 0 0 0 0 1 1 G 1 0 0 0 0 0 0 1 1 1

[0089] Table 2

[0090] When C is used as the label to be predicted, each frame is used as a label data of 0 or 1 for C. In each frame, the faults other than C that are 1 can be used as the features of C reporting or not reporting a fault. Thus, 10 pieces of data can be formed, as shown in Table 3:

[0091]

[0092]

[0093] Table 3

[0094] The first column in the above table is the label value for reporting or not reporting the C fault, which has two values ​​0 and 1. The corresponding second column is the characteristic value corresponding to the entry data, which is actually other faults with a fault word label value of 1 in the corresponding flight parameter data frame.

[0095] In step 240, when the data is limited, such as when there is only one flight parameter, the data can be split, such as 80% of the data is used for training and 20% for testing. When the data is relatively sufficient, such as when there are data from multiple flights, some of the flight data can be retained to verify the prediction effect.

[0096] For each frame of data, for a single fault, other reported faults in this frame of data are used as model input to make predictions and obtain predicted probability values. A threshold can be set to convert the probability values ​​into labels of 0 and 1. For example, when it is greater than 65%, it is set to 1, and when it is less than 35%, it is set to 0. Data in the range of 35%-65% are not compared. In this way, the accuracy of the model prediction can be calculated, and more complex model evaluation criteria can also be used to evaluate the prediction effect.

[0097] Apply the basic form of Bayes' theorem to the associated faults in the flight parameters. Assume that A and B are a pair of associated faults in the flight parameters. If we want to express "the probability of fault A occurring when fault B occurs", we can express it according to the basic form of Bayes' theorem as follows:

[0098]

[0099] In formula (1), P(A|B) is the conditional probability to be solved, that is, the probability of A occurring after B occurs, also known as the posterior probability; P(B|A) is the probability of B occurring when fault A occurs, also known as the prior probability, obtained based on the actual flight parameter data; P(A) and P(B) are the statistical probabilities of fault A and fault B occurring independently, respectively; P(B|A), P(A) and P(B) are the basis for predicting the probability of fault A occurring.

[0100] When a target fault has two associated faults, but these two associated faults have never appeared together before, but appear together in the current frame flight parameters, how to calculate the joint probability of the two based on the total probability formula and the "naive" assumption in the naive Bayes algorithm and based on the correlation between the two and the target fault.

[0101] (I) The probability calculation when the target fault A corresponds to an associated fault is as follows:

[0102] Assuming that there is a correlation between fault B and fault A, the probability of fault A occurring when fault B occurs can also be expressed as the probability of fault A occurring given the condition that fault B occurs, that is, P(A|B) described above. According to the conditional probability formula:

[0103]

[0104] It can be calculated as: the number of data frames where faults A and B occur simultaneously, divided by the number of data frames where fault B occurs.

[0105] (2) In the formula, P(A∩B) represents the probability of faults A and B occurring simultaneously.

[0106] (II) The probability calculation when the target fault A corresponds to two associated faults is as follows:

[0107] Assume that both faults B and C are correlated with fault A. When only one of faults B and C occurs, the method for calculating the conditional probability of fault A and their respective conditions is the same as mentioned above. When calculating the probability of fault A occurring when both faults B and C occur at the same time, the calculation method is different.

[0108] In this case, it is known that both faults B and C are related to the occurrence of A, and P(A|B) and P(A|C) can be calculated based on historical flight parameter data. Then when faults B and C occur at the same time, the probability of fault A is expressed as P(A|B∩C), where B∩C represents the case where faults B and C occur at the same time. Using the conditional probability calculation formula mentioned above, we get:

[0109]

[0110] Using the Bayesian formula:

[0111]

[0112] In the conditional probability formula, event A∩B∩C may not have a corresponding situation in the actual flight data, that is, the situation where the three faults ABC occur simultaneously has not appeared in the flight data. Event B∩C may also not have occurred (see Figure 2 ), according to the statistical data, the value of P(B∩C) is 0, but obviously the denominator cannot be 0. Since the occurrence of faults B and C are related to fault A, when faults B and C occur at the same time, the probability of fault A occurring should be higher than when faults B or C occur alone. At this time, a way is needed to estimate a reasonable probability of faults B and C occurring at the same time. One way is to directly obtain P(B) and P(C) by statistically analyzing the flight parameters and then multiplying the two. The premise for this operation is that faults B and C are independent events, that is, the occurrence or non-occurrence of fault B has no effect on the probability of C occurring, and vice versa. However, faults B and C are both related to A, so this premise is not true, and such calculation loses information related to fault A. At this time, the full probability calculation method is required:

[0113]

[0114] Among them, {Xn:n=1,2,3,...} is a finite or countably infinite partition of a probability space (i.e., Xn is a complete event group), and each set Xn is a measurable set. Assuming that there are only two types of faults A and B in the flight parameters, then there are two Xn here, one is the probability space P(B) when fault B occurs, and the other is the probability space when B does not occur. The two situations where B occurs or does not occur constitute a complete sample space related to B, which can be expressed mathematically as follows:

[0115]

[0116] In this case, the probability of A failure occurring P(A) can be calculated as:

[0117]

[0118] When the situation is extended to the case of two associated faults B and C, then:

[0119]

[0120] In formula (8), An refers to A and non-A, that is, the case where fault A occurs and the case where fault A does not occur. After expansion, we get:

[0121]

[0122] At this point, a reasonable assumption can be made for the calculation of P(B∩C|A): assuming that under the premise that A occurs, or in the sample space of the event set containing A, event B and event C are independent events. This assumption is the "naive" assumption made in the naive Bayes algorithm. Because this independent assumption retains information related to A, it is naive but reasonable. Based on this assumption, the joint probability of B and C under the condition that A occurs can be obtained using the direct product method:

[0123] P(B∩C|A)=P(B|A)·P(C|A) (10)

[0124] Similarly, if A is replaced by non-A, that is, under the condition that A does not occur, there is also:

[0125]

[0126] Substitute the above two formulas into the total probability formula for calculating P(B∩C)

[0127]

[0128] get:

[0129]

[0130] A and They represent the situations where fault A occurs and fault A does not occur, respectively, forming a complete sample space. Based on the naive assumption that fault B and fault C are independent events in the two sample spaces where fault A occurs and does not occur, we can obtain the probability of fault B and fault C occurring simultaneously in a complete sample space.

[0131] According to the Bayesian formula that needs to be calculated initially:

[0132]

[0133] get:

[0134]

[0135] In formula (14), P(A) is the ratio of the number of frames in which A occurs in the flight parameter data to the total number of all data frames; P(B|A) is the ratio of the number of frames in which B occurs in the number of frames in which A occurs; P(C|A) is the ratio of the number of frames in which C occurs in the number of frames in which A occurs; is the ratio of the number of frames in which A does not occur in the flight parameter data to the total number of all data frames; is the ratio of frames where B occurs to frames where A does not occur; is the ratio of frames where C occurs to frames where A does not occur.

[0136] All of the above items can be obtained at a certain moment based on historical flight parameter data statistics. This provides a method for calculating the predicted probability of two associated faults for the target fault.

[0137] (III) The probability calculation when the target fault corresponds to any associated fault is as follows:

[0138] For the case where the target fault has any associated faults, the future occurrence probability of the target fault is predicted based on the historical occurrence data of the co-occurring faults.

[0139] Assume that based on a large amount of historical flight parameter data, there are n faults {M1, M2, …Mn} that are all related to the occurrence of the target fault A. To calculate P(A│M1∩M2∩…Mn), that is, the probability of fault A occurring when these faults occur simultaneously, the formula needs to be expanded item by item based on the above four points and expressed as:

[0140]

[0141] In formula (15), P(M 1 |A) indicates that in the number of frames where fault A occurs, fault M 1 The proportion of frames that occurred; P(M 2 |A) indicates that in the number of frames where fault A occurs, fault M 2 The proportion of frames that occurred; P(M n |A) indicates that in the number of frames where fault A occurs, fault M n The proportion of frames that occurred; It indicates the ratio of the number of frames in which the fault A does not occur in the flight parameters to the total number of all data frames; P(A) indicates the ratio of the number of frames in which the fault A occurs in the flight parameters to the total number of all data frames; It indicates the ratio of the number of frames in which fault M1 occurs to the number of frames in which fault A does not occur; It indicates the ratio of the number of frames in which fault M2 occurs to the number of frames in which fault A does not occur; It indicates the ratio of the number of frames where fault Mn occurs to the number of frames where fault A does not occur.

[0142] The above technical solution is described below with examples:

[0143] The flight parameters may contain dozens or hundreds of faults. You can select specific faults for model training or train all faults as needed. After training a certain amount of data, the model can be used, such as:

[0144] (1) Early warning of the probability of unreported faults

[0145] Assume that training has been conducted for 26 faults A, B, C...X, Y, Z. In the flight parameters at a certain moment, the values ​​of 10 faults A, B, C...J are 0, and the values ​​of the other 16 faults are 1. Then, for each fault from A to J, the corresponding associated faults can be found from the remaining 16 faults. For example, the associated faults of fault A are C, F, H, X, Y, Z. Since the fault values ​​of C and F in the current frame are 0, the characteristic faults related to A that can be collected in this frame of data are H, X, Y, Z. These four features are passed as input to the model algorithm used to predict fault A, and the probability of A occurring can be calculated based on the associated faults.

[0146] If a prediction model is trained for n faults in a selected fault set, then at any time when p faults in n have not occurred, the probability of occurrence of each fault in P can be predicted. The predicted set of probability values ​​can be used in multiple ways, such as sorting and displaying them according to the probability size, or setting a threshold value such as 75%, and issuing early warning displays or reminders for faults that exceed the set threshold.

[0147] (2) Analysis of high-risk symbiotic failure modes

[0148] For all associated faults of fault X, a finite number of associated fault combinations can be calculated. The model can be used to calculate the predicted probability value of each combination for X, and then high-risk combinations can be extracted for fault correlation analysis. The results may serve as a basis for fault prevention or aircraft design improvement.

[0149] The core of the flight parameter symbiotic fault prediction method based on the naive Bayes algorithm of the present invention lies in: first, it is to convert the "fault" in the flight parameter into a label in the training data, and convert each frame of such time-series-based data into a single sample, and convert all frames within a period of time into a conceptual abstract process of sample space. Second, how to establish a corresponding relationship between each element in the naive Bayes algorithm and the information of each dimension of the flight parameter, and finally construct the training data and a specific model training method.

[0150] In fact, this method is not only applicable to flight parameters, but any data with fault information can be modeled in this way, and specific system levels or modules can be selected for modeling. For example, the parameters with fault information of a key component in a complex system can be modeled to independently implement fault prediction or maintenance suggestions for the component. The overall fault status data of the system can also be modeled to obtain prediction information at a higher level of abstraction. Furthermore, this conversion idea is not limited to fault information, but can also be used for other data with key information output, as long as this information can be automatically extracted, such as in a parameter that does not contain fault information, a condition is set for a specific variable to mark a state, which can be bad (such as a certain type of fault) or good (such as a certain indicator reaching above the standard value).

[0151] Another embodiment of the present invention provides a flight parameter symbiotic fault prediction device based on a naive Bayesian algorithm, comprising: a processor, a memory, and a computer program in the memory that can be run on the processor, such as a flight parameter symbiotic fault prediction method program based on a naive Bayesian algorithm. When the processor executes the computer program, the steps in each of the above-mentioned flight parameter symbiotic fault prediction method embodiments based on a naive Bayesian algorithm are implemented, such as Figure 1 steps.

[0152] Exemplarily, the above-mentioned computer program can be divided into one or more modules / units, one or more modules / units are stored in a memory and executed by a processor to complete the present invention. One or more modules / units can be a series of computer program instruction segments that can complete specific functions. The instruction segments are used to describe the execution process of the computer program in the flight parameter symbiotic fault prediction device based on the naive Bayes algorithm. For example, the computer program can be divided into a sample space module, a model building module, a target fault acquisition module, and a probability prediction module. The specific functions of each module are as follows:

[0153] The sample space module is used to collect all fault combinations in the flight parameters and construct a continuous sample space. The fault combination contains at least two faults, and the continuous sample space has a time attribute.

[0154] The model building module is used to convert the co-occurrence fault information in the flight parameters into training data for unsupervised learning based on the continuous sample space and to build a naive Bayes fault prediction model.

[0155] The target fault acquisition module is used to obtain target fault information to be predicted.

[0156] The probability prediction module is used to input the target fault information into the naive Bayes fault prediction model and output the probability of the target fault occurring.

[0157] The flight parameter symbiotic fault prediction device based on the naive Bayes algorithm can be a computing device such as a desktop computer, a notebook, a handheld computer, a cloud server, etc. The flight parameter symbiotic fault prediction device based on the naive Bayes algorithm can include, but is not limited to, a processor and a memory, for example, it can also include an output device, a network access device, a bus, etc.

[0158] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the flight parameter symbiotic fault prediction device based on the naive Bayesian algorithm, and uses various interfaces and lines to connect various parts of the flight parameter symbiotic fault prediction device based on the naive Bayesian algorithm.

[0159] The memory can be used to store computer programs and / or modules. The processor implements various functions of the flight parameter symbiotic fault prediction device based on the naive Bayes algorithm by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory.

[0160] If the module / unit integrated in the flight parameter symbiotic fault prediction device based on the naive Bayes algorithm is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of each embodiment of the above-mentioned flight parameter symbiotic fault prediction method based on the naive Bayes algorithm.

[0161] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention is described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should be included in the scope of the claims of the present invention. The techniques, shapes, and structural parts not described in detail in the present invention are all known technologies.

Claims

1. A flight parameter symbiotic fault prediction method based on the naive Bayes algorithm, characterized in that: The following steps are involved: Collect all fault combinations of flight parameters to construct a continuous sample space, wherein the fault combination contains at least two faults and the continuous sample space has a time attribute; Based on the continuous sample space, the co-occurrence fault information in the flight parameters is converted into training data for unsupervised learning, and a naive Bayes fault prediction model is constructed; Obtain target fault information to be predicted; The target fault information is input into a naive Bayes fault prediction model, and the probability of the target fault occurring is output.

2. The flight parameter symbiotic fault prediction method based on the naive Bayes algorithm according to claim 1 is characterized in that: Based on the continuous sample space, the symbiotic fault information in the flight parameters is converted into training data for unsupervised learning, and a naive Bayes fault prediction model is constructed, including: Perform multiple random sampling on the fault data in the continuous sample space and split them into training sets and test sets according to the preset ratio; Convert the symbiotic fault information in the flight parameters of the training set into training data for unsupervised learning; Use the training data for unsupervised learning training and establish a naive Bayes fault prediction model; The naive Bayes fault prediction model is validated using the test set.

3. The flight parameter symbiotic fault prediction method based on the naive Bayes algorithm according to claim 2 is characterized in that: The preset ratio is 8:

2.

4. The flight parameter symbiotic fault prediction method based on the naive Bayes algorithm according to claim 2 is characterized in that: The symbiotic fault information in the training set flight parameters is converted into training data for unsupervised learning as follows: The faults that occurred simultaneously in the flight parameter data frame are regarded as associated faults. In the flight parameter data of the training set, faults that have only occurred alone are excluded. For a single fault X, there may be one or more faults that occurred simultaneously as the associated fault set S of X. X is taken as the object to be predicted, and the fault data in S is taken as the basis for prediction, that is, X is taken as the label to be predicted, and the fault information in S is taken as the feature of X.

5. The flight parameter symbiotic fault prediction method based on the naive Bayes algorithm according to claim 2 is characterized in that: When the target fault corresponds to an associated fault, the probability calculation formula of the target fault occurrence is as follows: Among them, fault A is the target fault, P(A|B) is the conditional probability that needs to be solved, that is, the probability of fault A occurring after fault B occurs; P(B) represents the statistical probability of fault B occurring independently; P(A∩B) represents the statistical probability of A and B occurring at the same time.

6. The flight parameter symbiotic fault prediction method based on the naive Bayes algorithm according to claim 2 is characterized in that: When the target fault corresponds to two associated faults, the probability calculation formula of the target fault occurrence is as follows: Among them, fault A is the target fault, P(A|B∩C) represents the probability of fault A occurring when faults B and C occur simultaneously; P(A) represents the proportion of the number of frames in which fault A occurs in the flight parameters to the total number of all data frames; P(B|A) represents the ratio of the number of frames where fault B occurs to the number of frames where fault A occurs; P(C|A) represents the proportion of frames where fault C occurs to the number of frames where fault A occurs; It indicates the ratio of the number of frames in which fault A does not occur in the flight parameters to the total number of all data frames; It indicates the ratio of frames where fault B occurs to frames where fault A does not occur. Indicates the ratio of frames where fault C occurs to frames where fault A does not occur.

7. The flight parameter symbiotic fault prediction method based on the naive Bayes algorithm according to claim 2 is characterized in that: When the target fault corresponds to any associated fault, assuming that according to a large number of historical flight parameters, there are n faults {M1, M2, ... Mn} that are related to the occurrence of the target fault A, the occurrence probability of the target fault A is calculated P(A│M1∩M2∩...Mn), that is, when these faults occur at the same time, the probability calculation formula of the occurrence of fault A is as follows: Wherein, P(M1|A) represents the proportion of frames in which fault M1 occurs to the number of frames in which fault A occurs; P(M2|A) represents the proportion of frames in which fault M2 occurs to the number of frames in which fault A occurs; P(M n |A) indicates that in the number of frames where fault A occurs, fault M n The proportion of frames that occurred; It indicates the ratio of the number of frames in which fault A does not occur in the flight parameters to the total number of all data frames; P(A) represents the proportion of the number of frames in which fault A occurs in the flight parameters to the total number of all data frames; It indicates the ratio of the number of frames in which fault M1 occurs to the number of frames in which fault A does not occur; It indicates the ratio of the number of frames in which fault M2 occurs to the number of frames in which fault A does not occur; It indicates the ratio of the number of frames where fault Mn occurs to the number of frames where fault A does not occur.

8. A flight parameter symbiotic fault prediction device based on a naive Bayesian algorithm, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the symbiotic fault prediction method according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the symbiotic fault prediction method according to any one of claims 1 to 7 are implemented.