A Few-shot Fault Diagnosis Method Based on Temporal Features and Federated Distillation
By combining the improved TimeGAN model and federal distillation algorithm, one-dimensional timing fault samples are generated and data sharing is carried out, the problems of insufficient samples and unbalanced distribution in elevators and other equipment are solved, and efficient and accurate fault diagnosis and safe maintenance of intelligent equipment are achieved.
Patent Information
- Application Number
- CN202210907519.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The prior art is difficult to achieve efficient and accurate fault diagnosis in the case of insufficient samples and unbalanced distribution, especially in equipment such as elevators, where data island problems and traditional generative models have poor effects, resulting in difficult model training.
Combined with the improved TimeGAN model and the federal distillation algorithm, time sequence features are extracted through the central server and distributed to the client, the knowledge distillation idea is used for model training and parameter sharing, one-dimensional timing fault samples are generated, sample inadequate samples and distribution imbalance problems, and data sharing and calculation decentralization are carried out on the premise of protecting privacy.
It realizes efficient and accurate fault diagnosis results, reduces communication costs and privacy risks, improves model accuracy and generalization performance, and realizes the safe maintenance of intelligent equipment.
Smart Images

Figure CN115526226B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of fault diagnosis, and particularly relates to a small-sample fault diagnosis method based on temporal features and federated distillation. Background Art
[0002] Currently, with the development of big data and artificial intelligence technologies, fault diagnosis has also been widely studied. However, artificial intelligence algorithm models are data-driven. Therefore, combining relevant algorithms to achieve fault diagnosis and prediction often requires a large amount of data, especially fault data. For many daily-use devices, such as elevators, it is very difficult to collect fault data, which cannot support model training. Therefore, it is necessary to generate model-generated fault samples. Traditional generative models are mostly used in image augmentation and restoration, and their effects in one-dimensional fault sample generation are very poor and cannot achieve the effect of sample augmentation.
[0003] Chinese Patent with application number 202210388562.2 discloses a small-sample fault diagnosis method based on an improved TimeGAN model. This method combines the TimeGAN model and the least square loss function in the Least Square Generative Adversarial Networks (LSGAN) to obtain the LS-Time GAN model, that is, the improved TimeGAN model, including: an Embedding network, a Recovery network, a sequence generator, and a sequence discriminator, which can fully consider the temporal correlation between one-dimensional signals and still retain the dynamic correlation between samples while generating samples, thereby realizing sample augmentation and solving the problem of insufficient samples.
[0004] In addition, different manufacturers or users of the same device often do not share data for reasons such as privacy protection, which results in the problem of data islands. As a distributed machine learning method, federated learning can achieve the decentralization of computing and data resources, aggregate models and share samples while ensuring the privacy of user data, indirectly expanding the samples and improving the model training efficiency. However, traditional federated learning has problems such as user heterogeneity and high communication costs. Summary of the Invention
[0005] The purpose of the present invention is to propose a small-sample fault diagnosis method based on temporal features and federated distillation for the above problems, which can solve the problems of insufficient samples and distribution imbalance, and obtain efficient and accurate fault diagnosis results, so as to realize the safety maintenance of intelligent devices.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows: A small-sample fault diagnosis method based on temporal features and federated distillation proposed by the present invention includes the following steps:
[0007] S1. Collect elevator operation data to form a data set, where the data set includes normal samples and fault samples;
[0008] S2. Input the fault samples into the improved TimeGAN model to generate new fault samples, and store the new fault samples in the data set to make the number of fault samples in the data set equal to that of normal samples, and divide the data set into a training set and a test set;
[0009] S3. Establish a federated distillation model, where the federated distillation model includes a central server and several clients, and input the training set into the federated distillation model for training, including:
[0010] S31. Use the central server to extract the temporal features of all samples in the training set and distribute them to the local models θ k of each client for training. Extracting the temporal features of all samples in the training set is achieved by learning the sequence generator G ω of the improved TimeGAN model. The temporal features include the latent static representation h S and the latent temporal representation h t at time t, where:
[0011] The sequence generator G ω satisfies the following objective function:
[0012]
[0013] In the formula, ω is the generator parameter, is the empirical risk of the approximate value of the true prior distribution p(y) of the output label y, is the empirical risk of h s , h t input into G ω (h s , h t |y), G ω (h s , h t |y) is the posterior probability distribution of h s , h t under the output label y, l is the loss function, σ is the activation function of the fully connected layer of the local model θ k , g is the output function of the fully connected layer of the local model θ k , is the prediction module parameter of the local model of the k-th client, k = 1, …, K, and K is the total number of clients;
[0014] The latent static representation h s and the latent temporal representation h tMeet the following conditions:
[0015]
[0016] Where e S is the embedding network for temporal static features, and e χ is the embedding network for temporal dynamic features. h t-1 is the latent temporal representation at time t - 1, s is the static feature vector, and x t is the temporal feature vector at time t. H S and H χ are the corresponding latent vector spaces of the static feature vector space S and the temporal feature vector space χ respectively;
[0017] S32. Introduce the noise vector into the sequence generator G ω , and meet the following conditions:
[0018]
[0019] S33. Send the sequence generator with the introduced noise vector to the client based on knowledge distillation, so that each local model θ k samples from the sequence generator G ω to obtain the enhanced representation h s , h t ~G ω (·|y) on the feature space, and construct the following objective function:
[0020]
[0021] Where is the average loss during training , and the formula is as follows:
[0022]
[0023] Where is the observed data of the domain , is the original data set formed by collecting elevator operation data, is and the empirical risk between h s , h t ~G ω (·|y). p is the prediction module of the local model, and f is the feature extraction module of the local model. is the parameter of the feature extraction module of the local model of the k-th client, x i is the i-th input sample, and c * is the true label function;
[0024] S34. Share the prediction module parameters of the local model and keep the feature extraction module parameters of the local model localized;
[0025] S4. Input the test set into the trained federated distillation model for testing, and evaluate according to the fault diagnosis results to obtain the final federated distillation model.
[0026] Preferably, the local model adopts a CNN neural network model.
[0027] Preferably, the loss function l adopts the maximum likelihood estimation method to supervise the loss and satisfies the following formula:
[0028]
[0029] In the formula, is the expectation of the probability distribution P of the random variable ζ and the temporal feature vector x at time t, μ t is the temporal random vector at time t. t
[0030] Preferably, the approximation of the true prior distribution p(y) of the label satisfies the following formula:
[0031]
[0032] In the formula, is the empirical risk of, is the observed data of the domain , is the original data set formed by collecting elevator operation data, I(·) is the indicator function, x is the input sample, y is the output label, and c * is the true label function.
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows: This method combines an improved TimeGAN model with a federated distillation algorithm. The improved TimeGAN model is used to generate one-dimensional time-series fault samples, solving the problems of insufficient samples and unbalanced distribution. Federated learning is used to decentralize computing and data resources, aggregating the models and shared data of different users while protecting privacy to achieve sample enhancement. A time-series feature extraction module is introduced into the central server of the federated distillation model. The sequence generator is used to learn the time-series features of the elevator operation data of each client and summarize them, and then distribute them to each client, enabling each client to learn the time-series features of the global data, thereby improving the model accuracy. Finally, using the idea of knowledge distillation, each local model in federated learning is used as a teacher network for training. For example, a CNN neural network model is used to achieve feature extraction and prediction, and the prediction part of the model obtained by training is aggregated into the global model, that is, the student network (sequence generator), generating a time-series feature representation containing the data distributions of all local models, and then distributing it to each local model for training, realizing data sharing and enhancement, thereby reducing communication costs, improving generalization performance, obtaining efficient and accurate fault diagnosis results, and thus realizing the safety maintenance of intelligent devices. Description of the Drawings
[0034] Figure 1 is a flowchart of the federated distillation small-sample fault diagnosis method based on time-series features of the present invention;
[0035] Figure 2 is a schematic structural diagram of the federated distillation model of the present invention. Detailed Embodiments
[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0037] It should be noted that unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the description of this application herein are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0038] As Figure 1-2 shown, a federated distillation small-sample fault diagnosis method based on time-series features includes the following steps:
[0039] S1. Collect elevator operation data to form a data set, and the data set includes normal samples and fault samples.
[0040] S2. Input the fault samples into the improved TimeGAN model to generate new fault samples, and store the new fault samples in the dataset to make the number of fault samples and normal samples in the dataset equal. Then divide the dataset into a training set and a test set.
[0041] S3. Establish a federated distillation model, which includes a central server and several clients, and input the training set into the federated distillation model for training, including:
[0042] S31. Use the central server to extract the temporal features of all samples in the training set and distribute them to the local models θ of each client k for training. Extract the temporal features of all samples in the training set by learning the sequence generator G of the improved TimeGAN model ω to achieve. The temporal features include the latent static representation h S and the latent temporal representation h at time t t , where:
[0043] The sequence generator G ω satisfies the following objective function:
[0044]
[0045] In the formula, ω is the generator parameter, is the approximate value of the true prior distribution p(y) of the output label y of the empirical risk, is h s , h t input into G ω (h s , h t |y) of the empirical risk, G ω (h s , h t |y) is the posterior probability distribution of h s , h t under the output label y, l is the loss function, σ is the activation function of the fully connected layer of the local model θ k , g is the output function of the fully connected layer of the local model θ k , is the prediction module parameter of the local model of the k-th client, k = 1, …, K, and K is the total number of clients;
[0046] The latent static representation h s and the latent temporal representation h at time t t satisfy the following conditions:
[0047]
[0048] where, e S is the embedding network for the sequential static features, and e χ is the embedding network for the sequential dynamic features, h t-1 is the latent temporal representation at time t-1, s is the static feature vector, and x t is the temporal feature vector at time t. H S and H χ are the corresponding latent vector spaces of the static feature vector space S and the temporal feature vector space χ respectively.
[0049] In practical applications, each client is independently distributed and does not communicate directly with each other. The corresponding elevator operation data collected is used to generate new fault samples through the improved TimeGAN model to expand the dataset, making the number of fault samples and normal samples in the dataset equal. Then, the expanded dataset is divided into a training set and a test set. The local model is trained using the training set on the client side to obtain the prediction module parameters of the local model, and the prediction module parameters are sent to the central server. Then, the central server extracts the sequential features of all samples in the training set and distributes them to the local models of each client for iterative training.
[0050] In one embodiment, the local model adopts a CNN neural network model. The CNN neural network model includes a feature extraction module and a prediction module. The prediction module is the output layer, such as a fully connected layer with a Sigmoid activation function, etc. The feature extraction module is other modules except the output layer, and the CNN neural network model can adopt any model in the prior art, which will not be elaborated here.
[0051] In one embodiment, to minimize the difference between the sequential features and the original data, the loss function l uses the maximum likelihood estimation method for supervised loss and satisfies the following formula:
[0052]
[0053] where, is the expectation of the probability distribution P of the random variable ζ and the temporal feature vector x t at time t, and μ t is the temporal random vector at time t.
[0054] In one embodiment, the approximation of the true prior distribution p(y) of the label satisfies the following formula:
[0055]
[0056] where, is the empirical risk of , and is the observed data of the domain . The original dataset formed by collecting elevator operation data, I(·) is the indicator function, x is the input sample, y is the output label, and c * is the true label function. That is, predictable data with causal relationships over time, corresponding to the time series data in this article.
[0057] S32. Introduce the noise vector into the sequence generator G ω , satisfying the following conditions:
[0058]
[0059] To diversify the output of G(·|y), introduce the noise vector into the sequence generator. Given any target output label y, this generator can generate the time series feature representation h s , h t ~G ω (·|y), and obtain the ideal prediction from the local model set, that is, the sequence generator generates a feature representation approximately consistent with the user data distribution of the client. Figure 2 where z in it is the time series feature h s , h t .
[0060] S33. Send the sequence generator with the introduced noise vector to the client based on knowledge distillation, so that each local model θ k samples from the sequence generator G ω to obtain the enhanced representation h s , h t ~G ω (·|y) in the feature space, and construct the following objective function:
[0061]
[0062] where, is the average loss during training , and the formula is as follows:
[0063]
[0064] In the formula, is the observed data of the domain , is the original dataset formed by collecting elevator operation data, is and the empirical risk between h s , h t ~G ω (·|y), p is the prediction module of the local model, and f is the feature extraction module of the local model. are the parameter of the feature extraction module of the local model for the k-th client, x i is the i-th input sample, c * is the true label function.
[0065] Knowledge distillation, also known as the teacher-student model, can use one or more powerful teacher networks to train lightweight student models and has been widely used in federated learning. By constructing the above objective function, the probability of generating ideal prediction results for augmented samples can be maximized. In this embodiment, each local model in federated learning is used as the teacher network for training. For example, a CNN neural network model is used to implement feature extraction and prediction, and the model prediction part obtained by training is aggregated into the global model, that is, the student network (sequence generator), to generate a temporal feature representation containing the data distributions of all local models, and then distributed to each local model for training to achieve data sharing and enhancement, thereby reducing communication costs and improving generalization performance, and obtaining efficient and accurate fault diagnosis results.
[0066] S34. Share the prediction module parameters of the local model for parameter sharing and keep the parameter of the feature extraction module of the local model localized.
[0067] Traditional federated learning algorithms share the entire local model, but this cannot guarantee communication costs and privacy security. On the one hand, advanced networks with deep feature extraction layers usually contain millions of parameters, which brings a huge burden to communication. On the other hand, for some actual federated learning application scenarios (such as the healthcare or financial fields, etc.), sharing the entire local model parameters may bring considerable privacy risks. Therefore, this application alleviates these problems by only sharing the prediction module parameters of the local model while keeping the parameter of the feature extraction module of the local model localized. Compared with the strategy of sharing the entire local model in the prior art, this partial sharing mode is more efficient and less vulnerable to data leakage.
[0068] S4. Input the test set into the trained federated distillation model for testing, and evaluate according to the fault diagnosis results to obtain the final federated distillation model. Inputting the sample to be recognized into the final federated distillation model can obtain the fault diagnosis result of the elevator. Evaluating the final federated distillation model through the test set is a conventional means for those skilled in the art and will not be elaborated here.
[0069] This method combines the improved TimeGAN model with the federated distillation algorithm. It uses the improved TimeGAN model to generate one-dimensional time-series fault samples to solve the problems of insufficient samples and unbalanced distribution. It also uses federated learning to decentralize computing and data resources, aggregating the models and shared data of different users while protecting privacy to achieve sample enhancement. Moreover, a time-series feature extraction module is introduced into the central server of the federated distillation model. The sequence generator is used to learn the time-series features of the elevator operation data of each client and summarize them, and then distribute them to each client so that each client can learn the time-series features of the global data, thereby improving the model accuracy. Finally, using the idea of knowledge distillation, each local model in federated learning is used as a teacher network for training. For example, a CNN neural network model is used to implement feature extraction and prediction, and the model prediction part obtained by its training is aggregated into the global model, that is, the student network (sequence generator), to generate a time-series feature representation containing the data distributions of all local models, and then distribute it to each local model for training to achieve data sharing and enhancement, thereby reducing communication costs, improving generalization performance, obtaining efficient and accurate fault diagnosis results, and thus realizing the safety maintenance of intelligent devices.
[0070] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0071] The above-described embodiments only express the embodiments of the present application that are described in more specific and detailed terms, but should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A few-shot fault diagnosis method based on temporal features by federal distillation, characterized in that: The small-sample fault diagnosis method based on temporal features comprises the following steps: S1. Collect elevator operation data to form a data set, where the data set includes normal samples and fault samples; S2. Input the fault samples into the improved TimeGAN model to generate new fault samples, and store the new fault samples in the data set to make the number of fault samples in the data set equal to that of normal samples, and divide the data set into a training set and a test set; S3. Establish a federated distillation model, where the federated distillation model includes a central server and several clients, and input the training set into the federated distillation model for training, including: S31. Use the central server to extract the temporal features of all samples in the training set and distribute them to the local models θ of each client for training. The extraction of the temporal features of all samples in the training set is achieved by learning to improve the sequence generator G of the TimeGAN model. The temporal features include the latent static representation h and the latent temporal representation h at time t, where: k The training is carried out by learning to improve the sequence generator G of the TimeGAN model. The extraction of the temporal features of all samples in the training set is achieved by learning to improve the sequence generator G of the TimeGAN model. The temporal features include the latent static representation h and the latent temporal representation h at time t, where: ω The temporal features include the latent static representation h S and the latent temporal representation h at time t t , where: The sequence generator G ω satisfies the following objective function: where ω is the generator parameter, is an approximation of the true prior distribution p(y) of the output label y of the empirical risk, is h s , h t input G ω (h s , h t |y) of the empirical risk, G ω (h s , h t |y) is the posterior probability distribution of h s , h t under the output label y, l is the loss function, σ is the activation function of the fully connected layer of the local model θ k and g is the output function of the fully connected layer of the local model θ k ; is the prediction module parameter of the local model of the k-th client, k = 1, …, K, and K is the total number of clients; The potential static representation h s and the potential temporal representation h at time t t satisfy the following conditions: where, e S is the embedding network of the temporal static feature, and e χ is the embedding network of the temporal dynamic feature, h t-1 is the latent temporal representation at time t-1, s is the static feature vector, and x t is the temporal feature vector at time t, and H S , H χ are the corresponding latent vector spaces of the static feature vector space S and the temporal feature vector space χ respectively; S32. Introduce the noise vector into the sequence generator G ω , and satisfy the following conditions: S33. Send the sequence generator that introduces the noise vector to the client based on knowledge distillation, so that each local model θ k Samples from the sequence generator G ω To obtain the enhanced representation h on the feature space s , h t ~G ω (·|y), and construct the following objective function: Among them, is the average loss during training, and the formula is as follows: In the formula, is the observed data of the domain , is the original data set formed by collecting elevator operation data, is and h s , h t ~G ω (·|y) empirical risk, p is the prediction module of the local model, f is the feature extraction module of the local model, is the parameter of the feature extraction module of the local model of the k-th client, x i is the i-th input sample, c * is the true label function; S34. Share the parameter of the prediction module of the local model and keep the parameter of the feature extraction module of the local model localized; S4. Input the test set into the trained federated distillation model for testing, and obtain the final federated distillation model according to the evaluation of the fault diagnosis results.
2. The method for few-shot fault diagnosis based on temporal features according to claim 1, wherein: The local model adopts a CNN neural network model.
3. The method for few-shot fault diagnosis based on temporal features as described in claim 1, characterized in that: The loss function l uses the maximum likelihood estimation method for supervised loss and satisfies the following formula: In the formula, is the temporal feature vector x of the random variable ζ and at time t t the expectation of the probability distribution P, μ t is the temporal random vector at time t.
4. The federated distillation few-shot fault diagnosis method based on temporal features according to claim 1, wherein: The approximation of the true prior distribution p(y) of the label satisfies the following formula: In the formula, is empirical risk, is the observed data of domain , is the original data set formed by collecting elevator operation data, I(·) is the indicator function, x is the input sample, y is the output label, and c * is the true label function.
Citation Information
Patent Citations
A small sample fault diagnosis method based on improved TimeGAN model
CN114692506B