A Fault Diagnosis Method for Unknown Faults of Bogies Based on Zero-Shot Learning Strategy of Generative Model

Through the zero-sample learning strategy based on the generative model, unknown fault data characteristics are generated and fault diagnosis classification models are trained, the problem of data imbalance in high-speed train bogie fault diagnosis is solved, and higher diagnostic accuracy and generalization are achieved.

CN119358375BActive Publication Date: 2025-05-27SOUTHWEST JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411297685.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-05-27
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

The prior art has data imbalance in the diagnosis of high-speed train bogies, which leads to the model tending to predict fault types with large data samples, making it difficult to accurately identify unknown faults.

Method used

Using a zero-sample learning strategy based on the generative model, a network of fault attribute semantic matrices and residual feature extraction is constructed, unknown fault data features are generated, and fault diagnosis classification models are trained in combination with known fault data features.

Benefits of technology

It improves the accuracy and generalization of high-speed train bogie fault diagnosis, can effectively identify known and unknown faults, and enhances the robustness and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119358375B_ABST
    Figure CN119358375B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of high-speed train bogie fault detection, and discloses a bogie unknown fault diagnosis method based on a zero-shot learning strategy of a generative model. The multi-body dynamics analysis software SIMPACK is used to conduct bogie mechanical fault simulation experiments on electric locomotives to collect normal and fault data; a fault attribute semantic matrix is built for the influencing factors of train operation; feature extraction is performed according to data attributes and fault data; an unknown fault data feature sample generation model based on a diffusion model is built, and the data features of unknown faults are generated by combining random noise and unknown fault attributes; a fault diagnosis classification model is constructed to realize the mapping from data features to fault attributes and then to fault categories. The present invention completes the diagnosis of known and unknown faults of high-speed train bogies, improves the accuracy and generalization of fault diagnosis, and provides comprehensive guarantee for the safe use of high-speed train bogies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of high-speed train bogie fault detection, and particularly to a method for diagnosing unknown faults of a bogie based on a zero-shot learning strategy of a generative model. Background Art

[0002] With the acceleration of urbanization and economic globalization, high-speed trains are playing an increasingly important role in improving regional connectivity and promoting economic development. However, the characteristics of high-speed operation also require extremely high reliability and safety of the train system, which highlights the importance of research on high-speed train fault diagnosis. As a key component of high-speed trains, the bogie is responsible for the vibration damping suspension system of the train, directly affecting the running stability and riding comfort of the train, and its faults may seriously affect the train operation. The key components of the bogie include the key vibration damping component air spring AS, the secondary lateral damper LD, and the anti-hunting damper AD for reducing the vibration between the wheel and the car body. The fault diagnosis of the high-speed train bogie vibration damping system mainly focuses on these damping components. The air spring is a damping component that uses air as an elastic medium, and its material may degrade due to long-term use and external environmental factors, resulting in reduced elasticity, cracking, and leakage. The lateral damper and the anti-hunting damper are mainly used to reduce the lateral movement and hunting movement of the train car body, and their components may crack and leak oil during long-term high-speed operation, resulting in failure faults.

[0003] The existing bogie fault diagnosis methods can generally be divided into model-based and data-based. The basic idea of the former is to calculate the residuals under the system operating state using a mathematical model and diagnose faults by analyzing these residuals. However, in practical applications, the non-linearity and complexity of the bogie system limit the accuracy of the analytical model, making the model-based diagnosis method have certain limitations. The latter data-based fault diagnosis method mainly uses various sensor data collected during the system operation process to deeply excavate the equipment state information and fault characteristics hidden in the data. Through advanced data analysis and machine learning techniques, fault patterns can be effectively identified and potential fault trends can be predicted. At present, the field of bogie fault diagnosis is developing in the direction of data-driven, and combining advanced data analysis techniques can effectively overcome the limitations of the model-based method and improve the accuracy and timeliness of bogie fault diagnosis.

[0004] Although data-driven methods offer many conveniences, they also face significant challenges in practical working conditions. The effectiveness of data-driven methods depends to a large extent on the quality of the input data. If there are too many missing values or errors in the data, it will directly affect the accuracy of fault diagnosis. In the engineering practice of bogie faults in high-speed trains, the normal operation data samples are often much more than the fault samples, and the probability of a single fault occurring is often much higher than that of a compound fault. However, unbalanced data has a great impact on data-driven fault diagnosis. Unbalanced data can easily lead to the drift of the weights of the training model, causing the model to tend to predict the fault type with a large amount of data samples. For different speed intervals and different fault locations, how to achieve accurate identification of fault categories will become the top priority. However, for fault types that have never occurred, how to use machine learning and deep neural networks to help improve data balance and enhance the accurate identification of specific fault categories becomes the key difficulty. The zero-shot learning strategy can well help complete the task of unknown class identification. This is a strategy that uses auxiliary information to help achieve cross-modal knowledge transfer and thus infer unknown class information. The generalized zero-shot learning strategy is to distinguish unknown classes from known classes based on auxiliary information such as known class data and attributes. The zero-shot learning strategy can currently be divided into three categories: the attribute prediction method, the embedding space learning method, and the generative model method. Currently, most of the work in the field of fault diagnosis focuses on attribute prediction and the embedding space. However, these two methods are extremely vulnerable to the influence of data imbalance, and the research on bogies does not require rich domain knowledge. On the contrary, the research on the generative model method, which can well help balance data and has good generalization ability and robustness, is relatively less.

[0005] From the above background, it can be clearly obtained that the three key points that must be solved for using data-driven methods to diagnose bogie faults are: (1) The method must be able to help the high-speed train bogie with numerous components achieve accurate identification of unknown class faults and known class faults, replacing manual diagnosis to achieve intelligent diagnosis; (2) The method can well balance the actual engineering situation of data imbalance and has good generalization and robustness; (3) The method model needs to have both accuracy and stability, can integrate information source data from multiple sensors, adapt to the complex and dynamic characteristics of the system, and ensure the stable operation of the train. Summary of the Invention

[0006] The purpose of the present invention is to provide a bogie unknown fault diagnosis method based on the zero-shot learning strategy of the generative model to complete the diagnosis of known class faults and unknown class faults for high-speed train bogies, improve the accuracy and generalization of fault diagnosis, and provide comprehensive protection for the safe use of high-speed train bogies.

[0007] In order to achieve the above purpose, the following technical solutions are adopted:

[0008] A fault diagnosis method for unknown faults of a bogie based on a zero - sample learning strategy of a generative model, the method comprising:

[0009] Obtain train operation data;

[0010] According to the fault position and running speed on the train bogie, construct a fault attribute semantic matrix of the train bogie, and sample the train operation data based on the fault attribute semantic matrix to obtain sampled data, where the sampled data includes known - class fault data and unknown - class fault data;

[0011] Use a residual feature extraction network to sample the known - class fault data and unknown - class fault data to obtain known - class fault data features and unknown - class fault data features;

[0012] Construct a generative model for unknown - class fault data features based on a diffusion model, train the generative model for unknown - class fault data features using random noise and known - class fault data features, and input the unknown - class fault data features into the trained generative model for unknown - class fault data features to obtain unknown - class generated fault data features;

[0013] Construct a fault diagnosis classification model, train the fault diagnosis classification model using the known - class fault data features and unknown - class generated fault data features, and realize the diagnosis of unknown faults of the bogie based on the trained fault diagnosis classification model.

[0014] Further, obtaining train operation data includes:

[0015] Establish a train model, where the train model includes a car body, two bogies and four wheel sets, and the train model has multiple degrees of freedom;

[0016] Use the LMA economic tread, the first track and the second track as model excitations, and obtain train operation data based on a set sampling frequency.

[0017] Further, the fault attribute semantic matrix of the train bogie is classified according to the fault position and running speed conditions. For the fault position, if a fault occurs at the corresponding position, it is represented by 1, and if no fault occurs at the corresponding position, it is represented by 0; for the running speed condition, the train running speed is divided into multiple non - overlapping speed intervals, and if it runs in the corresponding speed interval, it is represented by 1, and if it does not run in the corresponding speed interval, it is represented by 0.

[0018] Further, construct a residual feature extraction network by the following method:

[0019] Build a one-dimensional convolutional feature extraction module; the one-dimensional convolutional feature extraction module includes a first convolutional layer and a second convolutional layer, and performs convolutional operations through the first convolutional layer and the second convolutional layer to extract the structural features of adjacent domains;

[0020] Construct two residual blocks. Based on the extracted structural features of adjacent domains, use the two residual blocks alternately twice to extract deep fault mode features, and directly map to connect different layers of the network;

[0021] Build an adaptive average pooling layer applied to one-dimensional signals. The adaptive average pooling layer adjusts the input deep fault model features to a set output size to obtain symmetric low-dimensional data features; among them, the symmetric low-dimensional data features include known-class fault data features and unknown-class fault data features.

[0022] Furthermore, the unknown fault data feature generation model based on the diffusion model includes a noise prediction model, an encoding module and a decoding module using skip connections; among them, the noise prediction model includes two embedding layers and a fusion layer. The two embedding layers are used to transform the noise addition step and the attribute encoding label features to facilitate the fusion of features. The fusion layer is used to splice and fuse the features with the noise addition step features and the attribute encoding label features to obtain fused features; the encoding module includes a Conv convolutional layer, a ReLU activation function layer and a MaxPool pooling layer for feature extraction, and the decoding module includes a ConvTrans deconvolutional layer and a ReLU activation function layer for upsampling.

[0023] Furthermore, use random noise and known-class fault data features to train the unknown fault data feature generation model, and input the unknown-class fault data features into the trained unknown fault data feature generation model to obtain unknown-class generated fault data features, including:

[0024] Gradually add noise to the original data to obtain noise data and gradually transition from the clear state to the completely random noise state to form a Markov chain:

[0025]

[0026]

[0027] In the formula, represents the state data at the time step, represents the state data at the time step, represents the normal distribution, represents the noise parameter at the time step, denotes the identity matrix, denotes the Gaussian distribution at the th time step, with a mean of and a variance of , , denotes all state sequences from the 1st to the th time step, denotes the joint probability distribution of the entire sequence { , , , ..., } starting from denotes the index of the current time step, denotes the total number of time steps, denotes the conditional probability given at ;

[0028] Let and , and through the reparameterization trick and inverse reparameterization, we get:

[0029]

[0030]

[0031] In the formula, denotes the parameter at the th time step, denotes the product of all from time step 1 to , and are coefficients used to combine and the noise , denotes the noise term sampled from the standard normal distribution;

[0032] During the training process, what is learned is the noise prediction function for each step. The sampling process starts from the standard Gaussian distribution, subtracts the predicted noise, and through continuous iteration, finally generates the feature of the unknown class generated fault data.

[0033] Furthermore, starting from the standard Gaussian distribution, subtracting the predicted noise, and through continuous iteration, finally generating the feature of the unknown class generated fault data is expressed as:

[0034]

[0035]

[0036]

[0037]

[0038]

[0039] wherein, represents the joint probability distribution of the model under the parameter covering the entire process from to ; represents all state sequences from the 0th to the th time step; represents the state at the time step of the initial probability distribution; represents the conditional probability distribution of predicting the previous state given ; represents the conditional probability distribution of predicting the previous state at the time step given and its prediction and variance ; represents the mean value at the time step calculated based on the parameter ; represents the parameter of the model; represents the parameter at the time step ; represents the noise prediction function at the time step ; represents the covariance matrix at the time step ; represents the standard deviation parameter at the time step ; represents the noise term sampled from the standard Gaussian distribution; represents the previous state generated according to the state at the time step .

[0040] Furthermore, when training the fault diagnosis and classification model by using the known class fault data features and the unknown class generated fault data features, binary cross-entropy is used as the loss function.

[0041] Furthermore, the fault diagnosis classification model includes an integration layer, a fully connected layer, and an attribute layer; wherein, the integration layer is used to integrate the known class fault data features and the unknown class generated fault data features, and output feature information to the fully connected layer, and the fully connected layer is used to map the feature information to the attribute label space for data attribute prediction and fault classification. The output of the fully connected layer is expressed as:

[0042]

[0043] In the formula, represents the output of the th layer, represents the input of the th layer, f represents the activation function, D i , b i both represent learning parameters, which are the weight matrix and bias term in the th layer;

[0044] The attribute layer is used to convert the output of the fully connected layer into the probability value of the corresponding class.

[0045] Furthermore, the attribute layer is used to convert the output of the fully connected layer into the probability value of the corresponding class through the following formula:

[0046]

[0047] In the formula, softmax represents the softmax function, which is used to convert the output of the fully connected layer into the probability of each class. softmax(y) i represents the probability value of the th class output by the function, represents the original output value of the th class, represents the original output value of the th class, represents the exponentiation of the original output value of the th class, represents the exponentiation of the original output value of the th class, j represents the class index, i represents the current class index, and n represents the total number of classes.

[0048] The beneficial effects of the present invention are:

[0049] 1. By introducing the generalized zero-shot strategy, the present invention constructs a dedicated attribute description matrix for the bogie.

[0050] 2. The present invention innovatively applies the diffusion generation model with remarkable effects in the field of images to the generation of low-dimensional data, which can significantly improve the adaptability and generation ability of the model in generalized zero-shot learning, enhance the generalization ability of the model, and enrich the generated data set.

[0051] 3. The present invention designs a generalized zero-shot fault diagnosis problem framework based on a conditional diffusion generation model, improves the residual network to extract the low-dimensional data features of the bogie, and completes the ternary mapping relationship of data-attribute-feature. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 The flowchart of a bogie unknown fault diagnosis method based on a zero-shot learning strategy of a generation model according to an embodiment of the present invention is shown.

[0053] Figure 2 The schematic diagram of the fault-exclusive semantic matrix of the high-speed train bogie according to an embodiment of the present invention is shown.

[0054] Figure 3 The schematic diagram of the residual network feature extraction framework according to an embodiment of the present invention is shown.

[0055] Figure 4 The schematic diagram of the principle of the diffusion model according to an embodiment of the present invention is shown.

[0056] Figure 5 The schematic diagram of the diffusion model noise learning framework according to an embodiment of the present invention is shown.

[0057] Figure 6 The schematic diagram of the fault diagnosis confusion matrix of all categories according to an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0059] The following further describes in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments.

[0060] The embodiment of the present invention provides a bogie unknown fault diagnosis method based on a zero-shot learning strategy of a generation model, as Figure 1 shown. The method includes the following steps:

[0061] S1. Obtain the train operation data.

[0062] In some embodiments, step S1 specifically includes:

[0063] S11. Establish a train model, where the train model includes a car body, two bogies, and four wheel sets, and the train model has multiple degrees of freedom;

[0064] S12. Using the LMA economic profile, the first track, and the second track as model excitations, based on the set sampling frequency, obtain train operation data.

[0065] In this embodiment, since the faulty bogie cannot run at high speed for a long time to collect data, simulation experiments are carried out through SIMPACK to collect experimental data. SIMPACK is a multi-body dynamics analysis software widely used to simulate and analyze the dynamic performance of mechanical and mechatronic systems. The train model consists of a car body, two bogies, and four wheel sets, with 62 degrees of freedom. Fully considering the non-linearity of wheel-rail contact, the LMA economic profile, the CN60 track, and the Wuhan-Guangzhou line track (i.e., the first track and the second track) are set as model excitations, and the sampling frequency is set to 243 Hz to obtain experimental data. 220 s of each type of data is collected. Considering that there is more noise in the initial stage of operation, the first 3 seconds of data are removed. 58 sensors are installed on the bogie, so the data information size per second is 243 * 58. This simulation method can not only accurately simulate the dynamic performance of the bogie in normal and faulty states, but also test various hypothetical fault conditions without affecting the actual operation safety and efficiency, repeat the experiments under different operating conditions and fault scenarios to ensure a comprehensive data set is obtained, and these data are crucial for subsequent analysis and training of the fault diagnosis model.

[0066] S2. According to the fault location and operating speed on the train bogie, construct a fault attribute semantic matrix of the train bogie, and sample the train operation data based on the fault attribute semantic matrix to obtain sampled data, where the sampled data includes known-class fault data and unknown-class fault data.

[0067] In this embodiment, a fault attribute semantic matrix exclusive to the high-speed train bogie is built according to the main fault locations and operating speeds on the high-speed train bogie. The designed semantic matrix is as Figure 2As shown in the table, the classification is mainly carried out according to the fault location and the operating speed conditions. For the fault location, the occurrence or non-occurrence of the fault is represented by numbers, where 1 indicates that a fault occurs at the corresponding location, and 0 indicates that no fault occurs at that location. The fault may be a single-component fault or a composite-component fault. However, since the probability of all three parts failing is almost 0, the compound fault of the three parts is not studied. For the normal operating speed range of high-speed trains from 150 km / h to 210 km / h, it is evenly divided into three speed ranges, where 1 indicates operation in this speed range, and 0 indicates non-operation in this speed range.

[0068] For the operation data obtained in step S1, if the corresponding category can be found in the fault attribute semantic matrix, it is determined as known-class fault data; if the corresponding category cannot be found in the fault attribute semantic matrix, it is determined as unknown-class fault data.

[0069] S3. Use the residual feature extraction network to sample the known-class fault data and the unknown-class fault data to obtain the known-class fault data features and the unknown-class fault data features.

[0070] In this embodiment, due to the characteristics of the diffusion model network itself, fault data of the size of 243*58, which is asymmetric, is not conducive to direct generation. To further improve the diagnostic accuracy of the fault diagnosis framework, deeply explore the hidden fault mode information behind the data, and at the same time facilitate the data generation of the diffusion model, a feature extraction network is designed to extract the data fault features. The generally used hierarchical deep convolutional network makes the close relationship between the extracted features and the high-level semantics of the data through multiple convolutional layers and pooling layers. The multi-layer network brings about the problem of network degradation. At this time, the training effect of the shallow network is far better than that of the deep network, and at the same time, the data information contained in the network decreases as the number of layers deepens. Using direct mapping to connect different layers of the network, a residual network is introduced to solve this problem. The data passes through the Figure 3 As shown in the network architecture, two types of residual blocks are designed. The mapping of the residual blocks effectively reduces the learning difficulty and speeds up the model convergence speed. The learning rate of the feature extraction network framework is adjusted by the Adam optimizer with an initial learning rate of 1e-3. The batch size of data training is 100, and the number of iterations is 150. The final classification network is based on the residual network, and a fully connected layer, a linear mapping layer, and a Softmax layer are added for classification.

[0071] In some embodiments, the residual feature extraction network is constructed through the following steps:

[0072] S31. Build a one-dimensional convolutional feature extraction module for the first time. The input data size is 243 sampling frequencies * 58 sensor channels. The first part of the residual feature extraction network is convolutional layers C1 and C2, which perform convolutional operations to extract the structural features of adjacent domains.

[0073] S32. In combination with the principle of residual connection, two types of residual blocks are alternately used twice to extract deep fault mode features, and direct mapping is used to connect different layers of the network to solve the problem of network degradation. This step can obtain a feature map.

[0074] S33. Finally, an adaptive average pooling layer applied to one-dimensional signals is built to adjust the input feature map to a specified output size, helping the output features to be symmetric low-dimensional data features, which is convenient for the subsequent diffusion model to generate.

[0075] S4. Construct a generation model for unknown fault data features based on a diffusion model, train the generation model for unknown fault data features using random noise and known-class fault data features, and input the unknown-class fault data features into the trained generation model for unknown fault data features to obtain unknown-class generated fault data features.

[0076] In some embodiments, step S4 specifically includes the following steps:

[0077] S41. Build a noise prediction model based on Unet (semantic segmentation network). First, two embedding layers are used to convert the noise addition step and the attribute encoding label features for convenient feature fusion, and then a concatenation fusion (Cat) layer is used to splice and fuse the features with the noise addition step features and the attribute encoding label features as the model input.

[0078] S42. Build an encoding module and a decoding module with a Unet structure using skip connections in two layers. The encoding module includes a Conv convolutional layer, a ReLU activation function layer, and a MaxPool pooling layer for feature extraction, and the decoding module includes a ConvTrans deconvolutional layer and a ReLU activation function layer for upsampling.

[0079] S43. Train the noise prediction model using random noise and the attributes and low-dimensional features of known fault classes.

[0080] It should be noted that the known-class fault data features include the attributes and low-dimensional features of known fault classes; among them, the attributes of known fault classes are determined according to the fault attribute semantic matrix, and the low-dimensional features are extracted by the residual feature extraction network.

[0081] S44: Generate unknown-class fault data features using random noise, unknown-class fault attributes, and the trained noise prediction model.

[0082] In some embodiments, a generation model for unknown fault data features based on a diffusion model is constructed by the following method:

[0083] Build a diffusion model, and the specific framework and working principle are as Figure 4As shown. The diffusion model is a generative model that has performed well recently in deep generative models, especially in image generation. Its structure mainly consists of two processes, the forward diffusion process and the reverse generation process, both of which are parameterized Markov chains. The forward diffusion process gradually adds Gaussian noise to the data features until the data becomes random noise, and the model gradually adds noise to the original data and transforms the data into noise data close to the Gaussian distribution through multiple steps . The data features gradually transition from a clear state to a completely random noise state, forming a Markov chain:

[0084]

[0085]

[0086] In the formula, represents the state data at the th time step, represents the state data at the th time step, represents the normal distribution, represents the noise parameter at the th time step, represents the identity matrix, represents the th time step given when the Gaussian distribution, whose mean is and the variance is . represents all state sequences from the 1st to the th time step, represents the entire sequence{ starting from , , ..., } of the joint probability distribution, represents the index of the current time step, represents the total number of time steps, represents the conditional probability when given and .

[0087] Define and . Through the reparameterization trick and inverse reparameterization, we can obtain:

[0088]

[0089]

[0090] In the formula, Denotes the parameter at the th time step, Denotes the product from time step 1 to all , and is the coefficient used to combine and the noise , Denotes the noise term sampled from the standard normal distribution.

[0091] The derivation formula reflects that can be sampled based on the linear combination of the original data and the random noise , while and is also called the combination coefficient.

[0092] Contrary to the diffusion process with noise addition, the generation process (reverse) is a denoising process. The process of gradually denoising from the random noise is also a Markov chain, which is a series of Gaussian distributions parameterized by neural networks.

[0093]

[0094]

[0095]

[0096]

[0097] In the above formula , while is a parameterized Gaussian distribution, and the mean and variance can be learned by the training network , that is, replacing the Gaussian noise of the posterior distribution with the noise prediction function , and the evolving variance does not need to be learned. Finally, we get:

[0098]

[0099] In the formula, denotes the joint probability distribution of the model under the parameter , covering the entire process from to , denotes all state sequences from the 0th to the th time step, denotes the initial probability distribution of the state at the time step , denotes at the time step Given , predict the conditional probability distribution of the previous state . Denote the time step Given and its prediction and variance , the Gaussian distribution Denote the mean at time step calculated based on the parameter . Denote the parameters of the model Denote the parameters at time step . Denote the noise prediction function at time step . Denote the covariance matrix at time step . Denote the standard deviation parameter at time step . Denote the noise term sampled from the standard Gaussian distribution Denote the previous state generated according to the state at time step .

[0100] In summary, what the training process learns is the noise prediction function at each step . The sampling process starts sampling from the standard Gaussian distribution, then subtracts the predicted noise, and iterates continuously to finally generate new samples

[0101] This unknown fault data feature generation model based on the diffusion model is essentially a noise learning model. It is based on the Unet architecture. Unet is an effective convolutional neural network commonly used in the field of image segmentation. Its feature is that through the symmetric structure of the encoder and decoder, plus skip connections, it can effectively capture and utilize multi-scale context information. The specific network design is as Figure 5As shown in the framework. In this embodiment, the network is used to generate symmetric low-dimensional fault data features. The input of the model is processed through an embedding layer, which includes two parts: one is the encoding of the noise addition step, and the other is the encoding of the fault attributes. The noise addition step encoding is used to represent a specific collective stage in the diffusion process. Each noise addition step corresponds to a time point in the diffusion model, indicating the degree of transition from the noise-free state to the full-noise state. This encoding helps the model understand the current noise level of the data, so as to appropriately adjust the denoising process. The attribute encoding, on the other hand, defines the specific characteristics of the fault, such as the fault location and operating speed, etc. These attribute auxiliary information are used to guide the addition of noise and the recovery of features of different fault types, ensuring that the generated data conforms to the expected fault characteristics. These encoded features are concatenated through the Cat layer, ensuring that information from different sources can be effectively fused in the model.

[0102] In the encoding module during the model learning stage, the convolutional layer (Conv) and the pooling layer (MaxPool) work together to extract deep features of the input data. The ReLU activation function provides non-linear processing in this process, helping the model capture more complex data patterns. Then, the data flows to the decoding module, which includes a transposed convolutional layer (ConvTrans) for upsampling the feature map to a higher resolution. The ReLU activation function is also used in this process to ensure the non-linear expression of the features. The value of each step is determined by linear interpolation between (0.00001, 0.005), and the total number of noise addition steps is set to 50. The mean squared error (MSE) is used as the loss function to measure the difference between the predicted noise and the true noise in each step. The optimizer used is the stochastic gradient descent (SGD) with periodic decay. A total of 150,000 training iterations are set. The initial learning rate is 0.01, and the learning rate is multiplied by 0.9 every 5000 epochs for adjustment.

[0103] Through such an encoding-decoding process, Unet can reconstruct fault features with specific attributes from the initial noise data. The training of this process depends on the attributes of known fault classes and low-dimensional feature data. After training, the model can accept new random noise inputs and generate corresponding fault data in combination with the attribute labels of unknown fault classes. These synthetic data not only enhance the diversity of the training set but also help the model learn how to perform accurate fault diagnosis in the case of unseen faults, effectively expanding the capabilities of the fault diagnosis system, enabling it to handle a wider range of fault scenarios, thereby improving the overall performance and reliability of the system.

[0104] S5. Construct a fault diagnosis classification model, train the fault diagnosis classification model using the fault data features of known classes and the generated fault data features of unknown classes, and implement the diagnosis of unknown faults of the bogie based on the trained fault diagnosis classification model.

[0105] In this embodiment, the generalized zero-shot learning problem is converted into a supervised classification problem. A fault diagnosis classification model is built and trained using the generated fault data features of unknown classes and the real data features of known classes. The integration (Flatten) layer is used to convert the low-dimensional features into one-dimensional features. Finally, the extracted feature information is converted into the attribute label space through the fully connected layer to complete data classification. The final output data is converted into the probability value of the corresponding class by softmax to predict the fault class.

[0106] Specifically, based on the fault diagnosis classification model, the fault features of the unknown fault class generated by the generation model and the fault features of the known fault class are used together to perform attribute prediction and class training on the final fault. The fault diagnosis classification model first integrates the features extracted by the model through the Flatten layer, and maps the extracted feature information to the attribute label space through the fully connected layer for data attribute prediction and fault classification. Its output can be expressed as:

[0107]

[0108] In the formula, represents the output of the th layer, represents the input of the th layer, f represents the activation function, D i , b i both represent learning parameters, which are the weight matrix and bias term in the th layer.

[0109] The last layer of the model is a Dense layer with 6 neurons, and the sigmoid activation function is used to output 6 values, each value corresponding to the predicted value of an attribute. The range of these 6 output values is between 0 and 1. The binary cross-entropy is used as the loss function, which is suitable for multi-label classification tasks or binary classification tasks of multiple attributes. The Adam optimizer is used for optimization to evaluate the accuracy. It is trained for 100 rounds, and if the monitored loss does not improve within 5 training rounds, the learning rate is adjusted to be reduced by 10 times.

[0110] Finally, the attribute layer is converted into the probability value of the corresponding class by softmax, and its expression is:

[0111]

[0112] In the formula, softmax represents the softmax function, which is used to convert the output of the fully connected layer into the probability of each category, and softmax(y) i represents the probability value of the th category of the function output, represents the original output value of the th category, represents the original output value of the th category, represents the exponentiation of the original output value of the th category, represents the exponentiation of the original output value of the th category. Here, j represents the category index, i represents the current category index, and n represents the total number of categories.

[0113] The confusion matrix is one of the important tools for evaluating the performance of a classification model. It is a square matrix that shows the comparison between the prediction results and the actual results of the classification model. Each row of the matrix represents the actual category, and each column represents the predicted category. Through the confusion matrix, the classification situation of the model in each category can be understood, and the advantages and disadvantages of the model can be identified. Randomly select 4 groups as unknown faults, and the others as known faults. The diagnostic accuracy of known class faults is higher than 95%, the accuracy of known faults in unknown classes exceeds 70%, and the harmonic mean of the two exceeds 80%. The specific diagnostic results are shown in Figure 6 .

[0114] In summary, the overall solution of the present invention has a complete diagnostic process, providing an effective method for detecting unknown faults in the bogie of high-speed trains.

[0115] Aiming at the situation of incomplete unknown class fault data, a dedicated bogie fault attribute knowledge base is constructed, combining the advantages of the diffusion generative model network and the generalized zero-shot strategy based on the generative model, deeply mining hidden fault modes, and using a deep neural network to learn the mapping relationship between fault data, fault features, fault attributes, and fault categories, taking into account accuracy, generalization, and robustness. The visualization results of the confusion matrix show that the designed diagnostic network framework is very sensitive to the fault features of the bogie, can effectively distinguish between known and unknown faults of the high-speed train bogie, and meets the diagnostic requirements.

[0116] The above embodiments are only used to illustrate the present invention, rather than to limit the present invention. Those of ordinary skill in the relevant technical fields can also make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention, and the patent protection scope of the present invention should be defined by the claims.

Claims

1. A bogie unknown fault diagnosis method based on a generative model zero-shot learning strategy, characterized in that: The method comprises: Obtain train operation data; According to the fault location and running speed of the train bogie, a fault attribute semantic matrix of the train bogie is constructed, and based on the fault attribute semantic matrix, the train running data is sampled to obtain sampled data, wherein the sampled data includes known class fault data and unknown class fault data; The known class fault data and the unknown class fault data are sampled by using a residual feature extraction network to obtain known class fault data features and unknown class fault data features; Constructing an unknown fault data feature generation model based on a diffusion model, training the unknown fault data feature generation model using random noise and known class fault data features, inputting the unknown class fault data features into the trained unknown fault data feature generation model to obtain unknown class generated fault data features; Constructing a fault diagnosis classification model, training the fault diagnosis classification model using the known class fault data features and the unknown class generated fault data features, and realizing unknown bogie fault diagnosis based on the trained fault diagnosis classification model; The residual feature extraction network is constructed by the following method: Building a one-dimensional convolution feature-based one-time extraction module; wherein the one-dimensional convolution feature-based one-time extraction module includes a first convolution layer and a second convolution layer, and a convolution operation is performed through the first convolution layer and the second convolution layer to extract structural features of adjacent domains; Two residual blocks are constructed. Based on the extracted structural features of the adjacent domains, the two residual blocks are used alternately twice to extract deep fault mode features and directly mapped to connect different layers of the network. Building an adaptive average pooling layer applied to a one-dimensional signal, wherein the adaptive average pooling layer adjusts the input deep fault model features to a set output size to obtain symmetrical low-dimensional data features; wherein the symmetrical low-dimensional data features include known class fault data features and unknown class fault data features; The unknown fault data feature generation model based on the diffusion model includes a noise prediction model, an encoding module using jump connections, and a decoding module; wherein the noise prediction model includes two embedding layers and a fusion layer, the two embedding layers are used to convert the noise step and attribute coding label features to facilitate the fusion of features, and the fusion layer is used to concatenate and fuse the features with the noise step features and the attribute coding label features to obtain the fusion features; the encoding module includes a Conv convolution layer, a ReLU activation function layer, and a MaxPool pooling layer for extracting features, and the decoding module includes a ConvTrans deconvolution layer and a ReLU activation function layer for upsampling.

2. The bogie unknown fault diagnosis method based on the generative model zero-sample learning strategy according to claim 1 is characterized in that: Obtain train operation data, including: Establishing a train model, wherein the train model includes a car body, two bogies and four wheel pairs, and the train model has multiple degrees of freedom; The LMA economical tread, the first track and the second track are used as model excitations, and the train operation data is obtained based on the set sampling frequency.

3. The bogie unknown fault diagnosis method based on the generative model zero-sample learning strategy according to claim 1 is characterized in that: The fault attribute semantic matrix of the train bogie is classified according to the fault location and the operating speed condition. For the fault location, a fault at the corresponding location is represented by 1, and a no fault at the corresponding location is represented by 0. For the operating speed condition, the train operating speed is divided into multiple non-overlapping speed intervals. Running in the corresponding speed interval is represented by 1, and not running in the corresponding speed interval is represented by 0.

4. The bogie unknown fault diagnosis method based on the generative model zero-sample learning strategy according to claim 1, characterized in that: The unknown fault data feature generation model is trained using random noise and known class fault data features, and the unknown class fault data features are input into the trained unknown fault data feature generation model to obtain the unknown class generated fault data features, including: Gradually add noise to the original data x0 to obtain the noise data x T , and gradually transition from a clear state to a completely random noise state, forming a Markov chain: In the formula, x t represents the state data at time step t, x t-1 represents the state data of the t-1th time step, represents normal distribution, β t represents the noise parameter at the tth time step, I represents the identity matrix, Indicates that at time step t, given x t-1 x t The Gaussian distribution of The variance is β t I, x 1:T represents all state sequences from the 1st to the Tth time step, q(x 1:T |x0) represents the entire sequence {x1,x2,...,x T }, t represents the index of the current time step, T represents the total number of time steps, q(x t |x t-1 ) means that given x t-1 Time t The conditional probability of Let α t =1-β t and After reparameterization and anti-reparameterization, we get: In the formula, α i represents the parameters of the i-th time step, represents all α from time step 1 to t i The product of and is the coefficient used to combine x0 and noise ∈, ∈ represents the noise term sampled from the standard normal distribution; The training process learns the noise prediction function at each step. The sampling process starts with sampling from the standard Gaussian distribution, subtracts the predicted noise, and iterates continuously to finally generate the unknown class fault data features.

5. The bogie unknown fault diagnosis method based on the generative model zero-sample learning strategy according to claim 4 is characterized in that: Sampling starts from the standard Gaussian distribution, subtracting the prediction noise, and iterating continuously to finally generate the unknown class to generate the fault data feature representation: In the formula, p θ (x 0:T ) represents the joint probability distribution of the model under parameter θ, covering from x0 to x T The whole process, x 0:T represents all state sequences from the 0th to the Tth time step, p(x T ) represents the state x at time step T T The initial probability distribution, p θ (x t-1 |x t ) means that at time step t, given x t In the case of t-1 The conditional probability distribution of Denotes time step t given x t and its prediction μ θ and variance∑ θ Gaussian distribution under θ (x t ,t) represents the mean of time step t calculated based on parameter θ, θ represents the parameter of the model, α t represents the parameter of time step t, ∈ θ (x t , t ) represents the noise prediction function at time step t, ∑ θ (x t ,t) represents the covariance matrix at time step t, σ t represents the standard deviation parameter at time step t, z represents the noise term sampled from a standard Gaussian distribution, and x t-1 Represents the state x at time step t t The previous state generated.

6. The bogie unknown fault diagnosis method based on the generative model zero-sample learning strategy according to claim 1, characterized in that: When the fault diagnosis classification model is trained using the known class fault data features and the unknown class generated fault data features, binary cross entropy is used as a loss function.

7. The bogie unknown fault diagnosis method based on the generative model zero-sample learning strategy according to claim 1, characterized in that: The fault diagnosis classification model includes an integration layer, a fully connected layer and an attribute layer; wherein the integration layer is used to integrate the known class fault data features and the unknown class generated fault data features, and output feature information to the fully connected layer, and the fully connected layer is used to map the feature information to the attribute label space to perform data attribute prediction and fault classification. The output of the fully connected layer is expressed as: In the formula, represents the output of layer l, represents the input of the l-1th layer, f represents the activation function, D i 、b i All represent learning parameters, which are the weight matrix and bias term in the lth layer; The attribute layer is used to convert the output of the fully connected layer into a probability value of the corresponding category.

8. The bogie unknown fault diagnosis method based on the generative model zero-sample learning strategy according to claim 7, characterized in that: The attribute layer is used to convert the output of the fully connected layer into the probability value of the corresponding category through the following formula: In the formula, softmax represents the softmax function, which is used to convert the output of the fully connected layer into the probability of each category, softmax(y) i Represents the probability value of the i-th category output by the function, y i Represents the original output value of the i-th category, y j represents the original output value of the jth class, represents the exponentialization of the original output value of the i-th class, Represents the index of the original output value of the jth category, j represents the category index, i represents the current category index, and n represents the total number of categories.

Citation Information

Patent Citations

  • Fault propagation diagnosis method for complex electromechanical system

    CN115374606A

  • Rail transit train bearing fault diagnosis method and device based on meta-learning algorithm

    CN117668689A