Aeroengine Gas Path Fault Diagnosis Method Considering Class Imbalance
By fusing the self-training method with the ACGAN method, aeronautical engine fault data sets are generated and balanced, and a semi-supervised ELM fault diagnosis model is established, which solves the problem of class imbalance and high-cost annotation data, and achieves higher fault diagnosis accuracy and generalization performance.
Patent Information
- Application Number
- CN202410698706.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-05-31
AI Technical Summary
There is a class imbalance problem in aircraft engine failure data, which affects the accuracy of the self-training method. In addition, traditional deep learning models rely on a large amount of labeled data, resulting in high manual labeling costs and high fault simulation experiment costs.
The self-training method is fused with the ACGAN method, and data is generated through the ACGAN model before self-training, and the unbalanced data set is balanced, a semi-supervised ELM fault diagnosis model is established, and the unsupervised data is used to improve the accuracy of the model.
Through data balance processing, the impact of data set imbalance is avoided, the accuracy of the model is improved, unsupervised data can be used more efficiently, and the generalization performance of the model is improved.
Smart Images

Figure CN118482935B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of aeroengine fault diagnosis, and particularly relates to a method for diagnosing aeroengine gas path faults considering class imbalance. Background Art
[0002] Aeroengines operate in a harsh environment, often under variable and high-load operating conditions, with diverse fault modes, resulting in maintenance costs accounting for more than 40% of the total engine maintenance cost. Among all fault modes of aeroengines, gas path system faults account for more than 90%, and their repair costs reach 60% of the overall engine repair cost. Gas path system fault modes generally include erosion, fouling, corrosion, interblade wear, foreign object damage, etc. Accurate diagnosis of gas path system faults in the engine can achieve rapid fault location, effectively reduce maintenance costs, and avoid major economic losses and safety accidents.
[0003] With the development of computer and sensor technologies, deep learning methods have gained the favor of most scholars. However, traditional deep learning models are usually trained based on a large amount of labeled fault mode data, which makes these methods have certain limitations. Manually annotating aero-engine fault data requires huge human and time costs. At the same time, the cost of fault simulation experiments is too high to be carried out, and only a small number of fault samples annotated by experts can be used for fault diagnosis. To address this problem, fault diagnosis methods based on semi-supervised learning have certain advantages. Semi-supervised learning establishes the internal connection between a small number of labeled fault samples and a large amount of unlabeled data, and uses the information of both labeled and unlabeled data for diagnosis. The self-training algorithm is a simple and powerful semi-supervised method. It trains an initial classifier with labeled data, obtains high-confidence pseudo-labeled data from unlabeled data to update the classifier. In this process, the classifier can obtain the information of both labeled and unlabeled data simultaneously, and has better generalization ability. ZHANG et al. published a paper titled "Semisupervised Momentum Prototype Network for Gearbox Fault Diagnosis Under Limited Labeled Samples" in the journal IEEE Transactions on Industrial Informatics, proposing a self-learning semi-supervised diagnosis method based on the prototype network, and using threshold selection based on Monte Carlo uncertainty to increase the confidence of pseudo-labels. ZHENG et al. published a paper titled "A Self-Adaptive Temporal-Spatial Self-Training Algorithm for Semisupervised Fault Diagnosis of Industrial Processes" in the journal IEEE Transactions on Industrial Informatics, introducing the time identity into the measurement of confidence, making the self-training algorithm applicable to industrial processes. JIAN et al. published a paper titled "Self-Training Reinforced Adversarial Adaptation for Machine Fault Diagnosis" in the journal IEEE Transactions on Industrial Electronics, using the self-training algorithm to use the information of unlabeled data to enhance domain adaptation. In the self-training algorithm, the performance of the initial classifier is crucial. A classifier with poor performance is likely to introduce incorrect pseudo-labels into the training, reducing the performance of the model. For aero-engines, there is a major difficulty affecting the establishment of the initial classifier, namely data imbalance.
[0004] In industrial equipment failures, it often occurs that the number of fault samples is less than that of normal samples, the classifier is more inclined to normal samples, and the effect of fault diagnosis is poor. At the same time, for aeroengines, there is also a class imbalance between different faults. For example, the proportion of dirt is 70-75%, and the proportion of erosion is 5%. GAN learns the data distribution characteristics from the original samples and generates new samples with a similar distribution, which is widely used in the classification of imbalanced data. ZHOU et al. published a paper titled "Deep learning fault diagnosis method based on globaloptimization GAN for unbalanced data" in the journal Knowledge-Based Systems, using a global optimization scheme to generate discriminant fault samples, selecting the generated samples, and filtering out unqualified samples. SHAO et al. published a paper titled "Generative adversarial networks for data augmentation in machine faultdiagnosis" in the journal Computers in Industry, proposing an Auxiliary Classifier Generative Adversarial Network (ACGAN). Based on GAN, the real label signal is introduced into the generation of synthetic signals, and synthetic signals with categories can be generated as augmented data for further application in mechanical fault diagnosis. It can be seen that in the prior art, there are a large number of unlabeled data in aeroengine flight data, and common supervised training methods are restricted. At the same time, there is a class imbalance problem in aeroengine fault data, which affects the accuracy of self-training methods. Summary of the Invention
[0005] To solve the above problems, the present invention combines the self-training method and the ACGAN method. Before self-training, data is generated through the ACGAN model to balance the imbalanced data set, and then a self-training model is established, which can avoid the influence brought by the imbalanced data set and can utilize a large amount of unsupervised data to improve the accuracy of the model.
[0006] An aeroengine gas path fault diagnosis method considering class imbalance provided by the present invention includes the following steps:
[0007] Step 1: Data collection: collect the operation data of the aircraft engine as sample data. The operation data is the state monitoring parameters of each section, including temperature, pressure, and speed. The collected samples are divided into normal data samples, fault data samples, and unlabeled samples. At the same time, the fault type of the collected samples is marked. Each sample consists of a period of operation data;
[0008] Step 2: Preprocess the sample data set;
[0009] Step 3: For the data preprocessed in step 2, the fault type is used as a label, and the state monitoring parameters of each section corresponding to each fault type are used as input to obtain a fault data set; the data without fault type labels are used as unlabeled data sets; and the fault data set is divided into a training set, a validation set, and a test set according to a certain ratio;
[0010] Step 4: In the first stage, random noise, real data, and the fault type label corresponding to the real data are used as input to construct the generator and discriminator of the auxiliary classifier generative adversarial network ACGAN respectively, where the generator outputs synthetic data, and the discriminator determines the probability that the synthetic data is real data and the fault type of the synthetic data. When the generator and the discriminator reach Nash equilibrium, the auxiliary classifier generative adversarial network ACGAN model Net1 is constructed. The real data is the collected samples. In the second stage, Net1 is used to balance the sample data corresponding to each fault type in the training set. By generating a few types of data and adding them to the training set, the proportion of each fault type data is the same. Then, each section state monitoring parameter is used as input, and the aircraft engine fault type is used as output to construct the extreme learning machine ELM model, and self-training is performed with an unlabeled data set to obtain a semi-supervised ELM fault diagnosis model Net2.
[0011] Step 5: After the training is completed, two-stage neural network models Net 1 and Net 2 are obtained, where Net 1 is a data generation model and Net 2 is a semi-supervised fault diagnosis model;
[0012] Step 6: Fault diagnosis: Input the parameters collected by the onboard sensors of the aircraft during flight into the model Net 2 to obtain the diagnosis results.
[0013] Furthermore, the step 2 specifically includes:
[0014] Step 2.1: Clear incomplete data and remove outliers;
[0015] Step 2.2: Normalize all data. The normalization method is as follows: Among them, x is the original data; x' is the normalized data; x min is the minimum value of the original data; x max is the maximum value of the original data.
[0016] Preferably, the fault data set is divided into a training set, a validation set, and a test set according to a ratio of 7:2:1.
[0017] Furthermore, the training of the ACGAN model Net 1 and the ELM fault diagnosis model Net 2 is performed according to the following steps:
[0018] Step 4.1: In the first stage, build the ACGAN generation model Net 1. The inputs are random noise, real data, and the fault type labels corresponding to the real data, and the output is realistic synthetic data. ACGAN is an extension of the generative adversarial network GAN. GAN consists of two competing neural networks, including a generator G and a discriminator D. The generator receives random noise z and generates data G(z) similar to the real data x. The discriminator D is used to distinguish whether the current input is real or generated by the generator and backpropagates the discriminant gradient information to the generator G to guide data generation. The model parameters of D and G are alternately updated through adversarial training until the Nash equilibrium is reached. The optimization direction of the generator G is to minimize the loss function L(G), and the optimization direction of the discriminator is to maximize the loss function L(D). Therefore, overall, the objective function L(D,G) of GAN is:
[0019]
[0020] Among them, P r (x) and P z (z) are the distributions of the real data x and the random noise z; D(x) is the probability that the discriminator determines that the real data x comes from the real data; G(z) is the output of the generator; E is the mathematical expectation; embedding the class label information c into the random noise vector z, the objective function of ACGAN consists of a discriminant loss and a classification loss. The discriminant loss L S is expressed as follows:
[0021]
[0022] L S is used to confirm the authenticity of the data, and the classification loss L C is expressed as follows:
[0023]
[0024] L C is used to measure the accuracy of the output samples; P(class = c|x) represents the probability that the discriminator correctly judges the class when the input is the real data x; P(class = c|G(z)) represents the probability that the discriminator correctly judges the class when the input comes from the generator; for the discriminator, the optimization goal is to maximize L C +L S; For the generator, the optimization goal is to minimize L C -L S ;
[0025] Step 4.2: In the second stage, the fault diagnosis model Net 2 is trained using the data balanced by the model Net1. The input of the model Net 2 is the parameter of the balanced fault data set, and the output is the fault type; the ELM calculation process is expressed as:
[0026]
[0027] where x j is the input vector of the j-th sample; L is the number of hidden nodes; N is the number of training samples; β i is the weight vector between the i-th hidden node and the output; ω i is the weight vector between the i-th hidden node and the input; g is the activation function; b i is the threshold of the i-th hidden node, and the above formula is simplified as:
[0028] T = Hβ
[0029] where H is the output matrix of the hidden layer, which is invertible; T is the target matrix of the training set; β is the weight vector matrix, which are respectively expressed as:
[0030]
[0031]
[0032]
[0033] where m is the number of outputs, and the optimization function is expressed as:
[0034]
[0035] where is the Moore-Penrose generalized inverse of H.
[0036] Furthermore, the steps for training the corresponding ELM fault diagnosis model Net 2 are specifically as follows:
[0037] Step 4.2.1: Generate minority class data through the ACGAN model, perform class balancing operation on the training set, and then train the ELM model based on the balanced training set. Select the model Model with the highest accuracy rate on the validation set for the next step;
[0038] Step 4.2.2: Use the model Model to test the unlabeled data, and the prediction result is used as the pseudo label;
[0039] Step 4.2.3: Select unlabeled samples based on the confidence of pseudo-labels, add them to the training set, and retrain to obtain model Model 1 ;
[0040] Step 4.2.4: Use the model Model obtained in Step 4.2.3 1 as the input of Step 4.2.2, iterate until all unlabeled samples participate in the training, and evaluate the final Model 1 using the test set.
[0041] The beneficial technical effects of the present invention are as follows: The present invention combines the self-training method and the ACGAN method. Before self-training, data is generated through the ACGAN model to balance the class-imbalanced dataset, and then a self-training model ELM is established for fault diagnosis. This method can more accurately diagnose the gas path faults of aero-engines, more efficiently utilize unsupervised data to improve the generalization performance of the model, and at the same time use the ACGAN model to generate minority-class data to prevent the classifier from tending to the data with a large proportion, and has a higher accuracy than directly performing fault diagnosis. Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 is a flowchart of a method for diagnosing gas path faults of aero-engines considering class imbalance provided by an embodiment of the present invention.
[0044] Figure 2 is a schematic diagram of the fusion of the ACGAN model and the self-training model ELM provided by an embodiment of the present invention. Detailed Embodiments
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0046] As Figure 1 shown, a method for diagnosing gas path faults of aero-engines considering class imbalance includes the following steps:
[0047] Step 1: Data acquisition. During multiple flight phases of an aeroengine or in software simulation, multiple different types of sensors are used to obtain the monitoring parameters of the engine's cross-section states such as temperature, pressure, and rotational speed at a sampling frequency within a specified range. Specifically, it includes: high-pressure rotor rotational speed, low-pressure rotor rotational speed, high-pressure compressor pressure, low-pressure compressor pressure, high-pressure compressor temperature, low-pressure compressor temperature, high-pressure turbine temperature, low-pressure turbine temperature, high-pressure turbine pressure, and low-pressure turbine pressure. The collected samples are divided into normal data samples, fault data samples, and unlabeled samples, and their fault types are marked at the same time. Each sample consists of the operation data for a period of time.
[0048] Step 2: Preprocess the sample dataset. Specifically:
[0049] Step 2.1: Remove incomplete data and eliminate outliers. The outliers are data points that are significantly different or deviate from the normal pattern compared with other data points in the dataset.
[0050] Step 2.2: Normalize all data. The normalization method is: where x is the original data; x' is the normalized data; x min is the minimum value of the original data; x max is the maximum value of the original data.
[0051] Step 3: For the data preprocessed in Step 2, use the fault type as the label and the monitoring parameters of each cross-section state corresponding to each fault type as the input to obtain a fault dataset; use the data without a fault type label as an unlabeled dataset, and divide the fault dataset into a training set, a validation set, and a test set according to a certain ratio.
[0052] Preferably, divide the fault dataset into a training set, a validation set, and a test set according to a ratio of 7:2:1.
[0053] Step 4: In the first stage, with random noise, real data, and the corresponding fault type labels of the real data as inputs, the generator and discriminator of the auxiliary classifier generative adversarial network ACGAN are respectively constructed. The generator outputs synthetic data, and the discriminator judges the probability that the synthetic data is real data and the fault type of the synthetic data. When the generator and the discriminator reach the Nash equilibrium, the auxiliary classifier generative adversarial network ACGAN model Net1 is constructed. Among them, the real data is the collected samples. In the second stage, Net1 is used to balance the sample data corresponding to each fault type in the training set. By generating minority type data and adding it to the training set, the proportion of data of each fault type is made the same. Then, with the monitoring parameters of each cross-section state as inputs and the aero-engine fault type as the output, an extreme learning machine ELM model is constructed, and self-training is carried out with the unlabeled data set to obtain the semi-supervised ELM fault diagnosis model Net2, with the optimization goal of improving the diagnostic accuracy.
[0054] More specifically, as Figure 2 shown, the training of the ACGAN generative model Net 1 and the ELM fault diagnosis model Net 2 is carried out according to the following steps:
[0055] Step 4.1: In the first stage, build the ACGAN generative model Net 1, with random noise, real data, and the corresponding fault type labels of the real data as inputs, and the output is realistic synthetic data; ACGAN is an extension of the generative adversarial network GAN. GAN consists of two competing neural networks, including a generator G and a discriminator D. The generator receives random noise z and generates data G(z) similar to the real data x. The discriminator D is used to distinguish whether the current input is real or generated by the generator, and backpropagates the discriminant gradient information to the generator G to guide data generation. The model parameters of D and G are alternately updated through adversarial training until the Nash equilibrium is reached. The optimization direction of the generator G is to minimize the loss function L(G), and the optimization direction of the discriminator is to maximize the loss function L(D). Therefore, overall, the objective function L(D,G) of GAN is:
[0056]
[0057] where, P r (x) and P z (z) are the distributions of the real data x and the random noise z; D(x) is the probability that the discriminator judges that the real data x comes from the real data; G(z) is the output of the generator; E is the mathematical expectation; embedding the class label information c into the random noise vector z, the objective function of ACGAN consists of the discriminant loss and the classification loss, and the discriminant loss L S is expressed as follows:
[0058]
[0059] L S For verifying the authenticity of data, the classification loss L C is expressed as follows:
[0060]
[0061] L C is used to measure the accuracy of the output samples; P(class = c|x) represents the probability that the discriminator correctly judges the class when the input is the real data x; P(class = c|G(z)) represents the probability that the discriminator correctly judges the class when the input comes from the generator; for the discriminator, the optimization objective is to maximize L C +L S ; for the generator, the optimization objective is to minimize L C -L S ;
[0062] Step 4.2: In the second stage, the fault diagnosis model Net 2 is trained using the data balanced by the model Net1. The input of the model Net 2 is the parameter of the balanced fault data set, and the output is the fault type; ELM is a feedforward neural network with good generalization performance and extremely fast learning ability. ELM sets the weights through the Moore-Penrose generalized inverse. Specifically, the ELM calculation process is expressed as:
[0063]
[0064] where, x j is the input vector of the j-th sample; L is the number of hidden nodes; N is the number of training samples; β i is the weight vector between the i-th hidden node and the output; ω i is the weight vector between the i-th hidden node and the input; g is the activation function; b i is the threshold of the i-th hidden node. The above formula is simplified as:
[0065] T = Hβ
[0066] where, H is the output matrix of the hidden layer, which is invertible; T is the target matrix of the training set; β is the weight vector matrix, which are respectively expressed as:
[0067]
[0068]
[0069]
[0070] where, m is the number of outputs, and the optimization function is expressed as:
[0071]
[0072] Among them, is the Moore-Penrose generalized inverse of H.
[0073] Furthermore, the steps of training the corresponding ELM fault diagnosis model Net 2 are specifically as follows:
[0074] Step 4.2.1: Generate minority class data through the ACGAN model, perform class balancing operation on the training set, and then train the ELM model based on the balanced training set. Select the model Model with the highest accuracy rate on the validation set for the next step;
[0075] Step 4.2.2: Use the model Model to test the unlabeled data, and the prediction result is used as the pseudo-label;
[0076] Step 4.2.3: Select the unlabeled samples according to the confidence of the pseudo-labels, add them to the training set, and retrain to obtain the model Model 1 ;
[0077] Step 4.2.4: Use the model Model obtained in Step 4.2.3 1 as the input of Step 4.2.2, iterate until all unlabeled samples participate in the training, and evaluate the final Model 1 using the test set.
[0078] Step 5: The training is completed, and two-stage neural network models Net 1 and Net 2 are obtained, where Net 1 is the data generation model and Net 2 is the semi-supervised fault diagnosis model.
[0079] Step 6: Fault diagnosis, input the parameters collected by the on-board sensors during the flight of the aircraft into the model Net 2 to obtain the diagnosis result.
[0080] For better understanding and implementation, the following uses simulation data combined with the accompanying drawings to give specific embodiments to illustrate in detail the method used in the present invention.
[0081] The fault types are shown in Table 1:
[0082] Table 1: Fault Types
[0083]
[0084]
[0085] First, clean the collected data of outliers and normalize it.
[0086] On this basis, an ACGAN generation model Net 1 is built. The inputs are random noise, real data, and the fault type labels corresponding to the real data, and the output is synthetic data similar to the corresponding fault data. The ACGAN parameter settings are shown in Table 2.
[0087] Table 2: ACGAN Parameter Settings
[0088]
[0089] On this basis, which is the second stage, a self-training ELM neural network model Net 2 is built in the same programming software. The input is the state monitoring parameters of each section of the fault data set, and the output is the engine fault type. The model structure has 120 neurons in the hidden layer, and the ReLU function is selected as the activation function. Self-training is carried out according to Step 4.2 to establish a semi-supervised model. The model is selected with the validation set, and the fault diagnosis accuracy is evaluated with the test set data. The specific diagnosis results obtained are shown in Table 3. It can be seen that this model can handle the class imbalance situation and also has a high accuracy when the number of labeled samples is small.
[0090] Table 3: Diagnosis Results of Fault Diagnosis Model Net 2 (Mean %)
[0091]
[0092]
[0093] In this embodiment, the ACGAN model is used to balance the unbalanced data, and the self-training ELM model is used for fault diagnosis. With 10 measurable flight parameters as the input, the model is trained with simulation data. The results show that this model has high precision and stability, and can diagnose the fault modes of aero-engines when the labeled data is small and the data is relatively unbalanced, and the accuracy still remains at a high level.
[0094] The above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A method for diagnosing aeroengine gas path faults taking into account class imbalance, characterized in that: The method comprises: Step 1: Data collection: collect the operation data of the aircraft engine as sample data. The operation data is the state monitoring parameters of each section, including temperature, pressure, and speed. The collected samples are divided into normal data samples, fault data samples, and unlabeled samples. At the same time, the fault type of the collected samples is marked. Each sample consists of a period of operation data; Step 2: Preprocess the sample data set, including clearing incomplete data and removing outliers; normalize all data using the following normalization method: Among them, x is the original data; x' is the normalized data; x min is the minimum value of the original data; x max is the maximum value of the original data; Step 3: For the data preprocessed in step 2, the fault type is used as a label, and the state monitoring parameters of each section corresponding to each fault type are used as input to obtain a fault data set; the data without fault type labels are used as unlabeled data sets; and the fault data set is divided into a training set, a validation set, and a test set according to a certain ratio; Step 4: In the first stage, random noise, real data, and the fault type label corresponding to the real data are used as input to construct the generator and discriminator of the auxiliary classifier generative adversarial network ACGAN respectively, where the generator outputs synthetic data, and the discriminator determines the probability that the synthetic data is real data and the fault type of the synthetic data. When the generator and the discriminator reach Nash equilibrium, the auxiliary classifier generative adversarial network ACGAN model Net1 is constructed, where the real data is the collected samples. In the second stage, Net1 is used to balance the sample data corresponding to each fault type in the training set, and a minority of types of data are generated and added to the training set to make the proportion of each fault type data the same. Then, each section state monitoring parameter is used as input, and the aircraft engine fault type is used as output to construct the extreme learning machine ELM model, and self-train with an unlabeled data set to obtain the semi-supervised ELM fault diagnosis model Net2; Step 5: After the training is completed, two-stage neural network models Net 1 and Net 2 are obtained, where Net 1 is a data generation model and Net 2 is a semi-supervised fault diagnosis model; Step 6: Fault diagnosis: input the parameters collected by the onboard sensors of the aircraft during flight into the model Net 2 to obtain the diagnosis results; Among them, the training of ACGAN model Net 1 and ELM fault diagnosis model Net 2 is performed according to the following steps: Step 4.1: In the first stage, build the ACGAN generation model Net 1, with random noise, real data and the fault type label corresponding to the real data as input, and realistic synthetic data as output; ACGAN is an extension of the generative adversarial network GAN. GAN consists of two competing neural networks, including the generator G and the discriminator D. The generator receives random noise z and generates data G(z) similar to the real data x. The discriminator D is used to distinguish whether the current input is real or generated by the generator, and back-propagates the discriminant gradient information to the generator G to guide data generation. The model parameters of D and G are updated alternately through adversarial training until the Nash equilibrium is reached. The optimization direction of the generator G is to minimize the loss function L(G), and the optimization direction of the discriminator is to maximize the loss function L(D). Therefore, in general, the objective function L(D, G) of GAN is: Among them, P r (x) and P z (z) is the distribution of real data x and random noise z; D(x) is the probability that the discriminator judges that the real data x comes from the real data; G(z) is the output of the generator; E is the mathematical expectation; the class label information c is embedded into the random noise vector z. The objective function of ACGAN consists of the discriminant loss and the classification loss. The discriminant loss L S It is expressed as follows: L S Used to confirm the authenticity of the data, the classification loss L C It is expressed as follows: L C Used to measure the accuracy of output samples; P(class=c|x) represents the probability that the discriminator correctly judges its category when the input is real data x; P(class=c|G(z)) represents the probability that the discriminator correctly judges its category when the input comes from the generator; for the discriminator, the optimization goal is to maximize L C +L S ; For the generator, the optimization goal is to minimize L C -L S ; Step 4.2: In the second stage, the fault diagnosis model Net 2 is trained using the data balanced by the model Net 1. The input of the model Net 2 is the balanced fault data set parameters, and the output is the fault type. The ELM calculation process is expressed as: Among them, x j is the input vector of the jth sample; L is the number of hidden nodes; N is the number of training samples; β i is the weight vector between the i-th hidden node and the output; ω i is the weight vector between the i-th hidden node and the input; g is the activation function; b i is the threshold of the i-th hidden node, and the above formula is abbreviated as: T+Hβ Among them, H is the hidden layer output matrix, which is reversible; T is the training set target matrix; β is the weight vector matrix, which are expressed as: Where m is the number of outputs and the optimization function is expressed as: in, is the Moore-Penrose generalized inverse of H.
2. The method for diagnosing aero-engine gas path faults taking into account class imbalance according to claim 1, characterized in that: The fault data set is divided into training set, validation set and test set in the ratio of 7:2:
1.
3. The method for diagnosing aero-engine gas path faults taking into account class imbalance according to claim 1, characterized in that: The steps for training the corresponding ELM fault diagnosis model Net 2 are as follows: Step 4.2.1: Generate minority class data through the ACGAN model, perform class balancing on the training set, train the ELM model based on the balanced training set, and select the model with the highest accuracy on the validation set to proceed to the next step; Step 4.2.2: Use the model to test the unlabeled data and use the predicted results as pseudo labels; Step 4.2.3: Select unlabeled samples based on the confidence of the pseudo-label, add them to the training set, and retrain to obtain model Model1; Step 4.2.4: Use the model Model1 obtained in step 4.2.3 as the input of step 4.2.2, iterate until all unlabeled samples are involved in the training, and use the test set to evaluate the final Model1.
Citation Information
Patent Citations
Fatigue crack growth prediction
CN110431395A
Battery fault classification method based on unbalanced semi-supervised adversarial training framework
CN117076871A