Small sample scene shield main bearing rolling body scratch damage sample generation method
The IWVAE-GAN model generates scratch damage samples of the main bearing rolling body of the shield machine, which solves the problem of many healthy samples and scarce damage samples, and achieves high-quality damage samples generation, which improves diagnostic accuracy and efficiency.
Patent Information
- Application Number
- CN202510320994.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art faces the problems of health sample diversity and scarcity of damage samples, difficulty in status monitoring and fault marking, small sample conditions and sample imbalance in the generation of scratch damage samples of main bearings of shield machines, making it difficult to effectively generate high-quality damage samples.
The importance weighted variational autoencoder generation adversarial network (IWVAE-GAN) model based on contrast learning is adopted, and the characteristics of healthy and impaired samples are extracted through shared feature encoder and private feature encoder, and pseudo-samples are generated by combining generators and discriminators. The healthy samples are used to enhance impaired samples generation and alleviate sample imbalance problem.
Generating high-quality rolling element scratch damage samples under conditions of very few damage samples improves the accuracy and efficiency of diagnosis, alleviates sample imbalance problem, and enhances the diversity and authenticity of damage characteristics.
Smart Images

Figure CN120277407A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault prediction and health management of shield machine main bearings, and particularly to a method for generating samples of rolling element scratch damage of a shield main bearing in a small sample scenario. Background Art
[0002] The health state of the main bearing of a shield machine is directly related to the operation safety of the entire shield machine and the tunneling efficiency of the shield. Scratch damage of the rolling elements in the main bearing is a typical damage of the shield machine main bearing. During the operation of the shield machine, since the rolling elements of the main bearing are enclosed in the raceway, the scratch damage occurs in the internal structure of the main bearing that is difficult to observe with the naked eye. Therefore, in actual shield tunneling operations, even under monitoring conditions such as vibration and oil, it is particularly difficult to identify the early stage of rolling element scratch damage and generate samples of such damage.
[0003] The existing technology faces several significant challenges in dealing with the generation of samples of rolling element scratch damage of the main bearing: (1) Diversity of healthy samples and scarcity of damaged samples: In the normal operation data of the shield machine, samples in the healthy state account for the vast majority, while due to regular maintenance and engineering safety considerations, the actually obtained scratch damage samples are relatively few. (2) Difficulty in state monitoring and fault marking: Due to the special structure and working environment of the main bearing, it is very difficult to directly extract and correlate the scratch damage state of the rolling elements from the state monitoring signals. When the main bearing needs to be disassembled for remanufacturing, relatively serious damage has often occurred, and this late-stage fault information is of limited help for early diagnosis and prevention; (3) Small sample condition: During the tunneling operation of the shield machine main bearing, although scratch damage of the rolling elements does exist as a typical fault form and occurs frequently, due to its usually weak and hardly noticeable early manifestations, and the difficulty of marking scratch damage samples, the labeled scratch damage samples are very scarce. This characteristic makes the samples of rolling element scratch damage of the shield machine main bearing show the nature of small samples, and traditional sample generation techniques often cannot effectively reproduce the scratch damage characteristics of the shield machine main bearing, making it difficult to achieve effective generation of scratch damage samples, especially in the case of lack of sufficient scratch damage data support; (4) Sample imbalance: In the state monitoring data of the shield machine main bearing, the samples in the healthy state are much more than the scratch damage samples, resulting in serious data imbalance. This imbalance is a huge challenge for deep learning models based on statistics, because the model is likely to be biased towards the majority class (i.e., healthy samples), thus ignoring the minority but critical scratch damage samples. Summary of the Invention
[0004] To solve the problems existing in the generation of samples for processing the scratch damage of bearing rolling elements in the above-mentioned existing technologies, the present invention proposes a method for generating scratch damage samples of the main bearing rolling elements of a shield tunneling machine in a small-sample scenario to solve the above problems.
[0005] The present application discloses a method for generating scratch damage samples of the main bearing rolling elements of a shield tunneling machine in a small-sample scenario, including the following steps:
[0006] S1. Collect the vibration data of the main bearing health state and the vibration data of the scratch damage state of the shield tunneling machine, perform standardization processing on the vibration data, and use the time window strategy to cut to obtain scratch damage samples and healthy samples;
[0007] S2. Construct an importance-weighted variational autoencoder generative adversarial network model based on contrastive learning, including a shared feature encoder, a private feature encoder, a generator, a discriminator, and a classifier;
[0008] S3. Use the shared feature encoder to encode the scratch damage samples and healthy samples into the shared feature latent space, and use the private feature encoder to encode the scratch damage samples into the private feature latent space;
[0009] S4. Sample in the private feature latent space to obtain the private feature latent representation of the scratch damage samples, and sample in the shared feature latent space to obtain the shared feature latent representation of the scratch damage samples and the shared feature latent representation of the healthy samples;
[0010] S5. Concatenate the private feature latent representation of the scratch damage samples and the shared feature latent representation of the scratch damage samples to obtain the scratch damage sample latent representation, and concatenate the shared feature latent representation of the healthy samples with a zero matrix to obtain the healthy sample latent representation;
[0011] S6. Input the scratch damage sample latent representation into the generator to generate pseudo-scratch damage samples, and input the healthy sample latent representation into the generator to generate pseudo-healthy samples;
[0012] S7. Input the pseudo-scratch damage samples, pseudo-healthy samples, scratch damage samples, and healthy samples into the discriminator and classifier, use the discriminator to discriminate the nature of the samples, and use the classifier to identify the state of the samples;
[0013] S8. Calculate the loss function of the importance-weighted variational autoencoder generative adversarial network model based on contrastive learning, including KL divergence loss, reconstruction loss, discriminator loss, generator loss, and state recognition loss;
[0014] S9. Calculate the gradient based on the loss obtained in S8, and alternately update the parameters of the shared feature encoder, private feature encoder, generator, discriminator, and classifier;
[0015] S10. Repeat S3 - S9 until the training epochs of the importance - weighted variational auto - encoder generative adversarial network model based on contrastive learning reach a preset value;
[0016] S11. Randomly sample the shared - feature latent representation and the private - feature latent representation from the standard normal distribution, concatenate them and input them into the trained generator to generate shield main bearing rolling - element scratch damage samples.
[0017] Preferably, the network architectures of the shared - feature encoder and the private - feature encoder are based on the Informer encoder, the network architecture of the generator is based on the Informer decoder, and the network architectures of the discriminator and the classifier are CNN - LSTM hybrid architectures.
[0018] Preferably, S3 includes the following steps:
[0019] S31. Use the private - feature encoder to extract the private features of the scratch - damage samples to obtain the distribution parameters of the private - feature latent representation of the scratch - damage samples. At the same time, use the shared - feature encoder to extract the shared features in the scratch - damage samples to obtain the distribution parameters of the shared - feature latent representation of the scratch - damage samples:
[0020]
[0021] where the subscript s corresponds to the shared features, the subscript z corresponds to the private features, and i corresponds to the sample index. is the scratch - damage sample, is the mean of the private - feature latent representation of the scratch - damage sample, is the standard deviation of the private - feature latent representation of the scratch - damage sample, is the mean of the shared - feature latent representation of the scratch - damage sample, is the standard deviation of the shared - feature latent representation of the scratch - damage sample, is the private - feature encoder, is the shared - feature encoder;
[0022] S32. Use the shared - feature encoder to extract the shared features of the healthy samples to obtain the distribution parameters of the shared - feature latent representation of the healthy samples:
[0023]
[0024] where, is the healthy sample, is the mean of the shared - feature latent representation of the healthy sample, is the standard deviation of the shared - feature latent representation of the healthy sample.
[0025] Preferably, S4 includes the following steps:
[0026] S41. Sample N times in the latent space of the private feature encoder to obtain the private feature latent representation of the scratch damage sample:
[0027]
[0028] where, is the private feature latent representation of the n-th sampled scratch damage sample, and ∈ is the noise term of the standard normal distribution ;
[0029] S42. Calculate the importance weights of the private feature latent representations of each scratch damage sample:
[0030]
[0031] where, is the private feature encoder, is and 's joint probability distribution, is the variational posterior distribution of the private feature encoder, indicating the probability of the latent representation given the input sample ;
[0032] S43. Sample N times in the latent space of the shared feature encoder to obtain the shared feature latent representation of the scratch damage sample:
[0033]
[0034] where, is the shared feature latent representation of the n-th sampled scratch damage sample;
[0035] S44. Sample once in the latent space of the shared feature encoder to obtain the shared feature latent representation of the healthy sample:
[0036]
[0037] where, is the shared feature latent representation of the healthy sample.
[0038] Preferably, the said S5 includes the following steps:
[0039] S51. Concatenate the private feature latent representation of the scratch damage sample and the shared feature latent representation of the scratch damage sample to obtain N latent representations of the scratch damage samples:
[0040]
[0041] where, is the potential representation of the nth scratch damage sample;
[0042] S52. Concatenate the potential representation of the shared features of the healthy samples and the zero matrix to form the potential representation of the healthy samples:
[0043]
[0044] where, is the potential representation of the healthy samples.
[0045] Preferably, S6 includes the following steps:
[0046] S61. Generate pseudo scratch damage samples. Input the potential representations of N scratch damage samples into the generator to generate N pseudo scratch damage samples:
[0047]
[0048] where, is the nth pseudo scratch damage sample, and g θ is the generator;
[0049] S62. Generate healthy samples. Input the potential representation of the healthy samples into the generator to generate pseudo healthy samples:
[0050]
[0051] where, is the pseudo healthy sample.
[0052] Preferably, S7 includes the following steps:
[0053] S71. Input the generated samples and the real samples and into the discriminator to obtain the discrimination result on whether the sample belongs to the real sample or the generated sample. The feedback of the discriminator is used to optimize the generator to ensure that the distribution of the generated samples is closer to that of the real samples:
[0054]
[0055] where, d i is the determination result of the discriminator, X f is the set of real damage samples, X h is the set of real healthy samples, and h θ is the discriminator;
[0056] S72. Input the generated samples and the real samples and Input the classifier to obtain the recognition result of whether the sample belongs to the healthy state or the damaged state:
[0057]
[0058] Among them, c i is the recognition result of the classifier, and f θ is the classifier.
[0059] Preferably, the S8 includes the following steps:
[0060] Calculate the KL divergence loss:
[0061]
[0062] Among them, is the KL divergence loss of the generated pseudo-scratch damage sample, is the KL divergence loss of the generated pseudo-healthy sample, is the total KL divergence loss, is the private feature latent representation of the nth pseudo-scratch damage sample, s is the shared latent representation, and are the KL loss weight coefficients;
[0063] Calculate the reconstruction loss:
[0064]
[0065] Among them, is the reconstruction loss of the generated pseudo-scratch damage sample, is the reconstruction loss of the generated pseudo-healthy sample, is the total reconstruction loss, is the expectation of sampling s from the variational distribution , and are the reconstruction loss weight coefficients;
[0066] Calculate the discriminant loss:
[0067]
[0068]
[0069] Among them, is the discriminant loss of the generated pseudo-healthy sample, is the discriminant loss of the generated scratch damage sample, is the total discriminant loss, and are the discriminant loss weight coefficients;
[0070] Calculate the generator loss:
[0071]
[0072] Among them, is the total loss of the generator, and are the loss weight coefficients of the generator;
[0073] Calculate the status recognition loss:
[0074]
[0075] Among them, is the status recognition loss of the generated pseudo-healthy samples, is the status recognition loss of the generated scratch damage samples, is the total status recognition loss, and are the status recognition loss weight coefficients.
[0076] Preferably, S9 includes the following steps:
[0077] S91. Freeze the shared feature encoder, private feature encoder, generator, and classifier, and update the parameters of the discriminator using the gradient descent method:
[0078]
[0079] Among them, η is the learning efficiency, is the parameter of the discriminator in the t-th round, is the gradient of the discriminator;
[0080] S92. Unfreeze the shared feature encoder, private feature encoder, generator, and classifier, freeze the discriminator, and update the parameters of the shared feature encoder, private feature encoder, generator, and classifier:
[0081]
[0082] Among them, is the parameter of the shared feature encoder in the t-th round, is the gradient of the shared feature encoder, is the parameter of the private feature encoder in the t-th round, is the gradient of the private feature encoder, is the parameter of the generator in the t-th round, is the gradient of the generator, is the parameter of the classifier in the t-th round, is the gradient of the classifier.
[0083] Preferably, S11 includes the following steps:
[0084] Randomly sample the latent representation s from the standard normal distribution f and z h , combine s f and z h to obtain the latent representation ω of the scratch damage sample f = [s f , z h , input ω f into the trained generator to obtain the scratch damage sample of the shield main bearing rolling element:
[0085]
[0086] wherein, ω f is the latent representation of the scratch damage sample, s f and z h are two latent representations randomly sampled from the normal distribution.
[0087] Advantages of the present invention:
[0088] (1) High-quality scratch damage sample generation ability: The present invention utilizes the powerful generation capabilities of the importance-weighted variational autoencoder (IWVAE) and the generative adversarial network (GAN) to effectively generate high-quality rolling element scratch damage samples under the condition of extremely few scratch damage samples.
[0089] (2) Generation of damaged samples enhanced by healthy samples: The present invention makes full use of abundant healthy state samples to support the generation of scarce scratch damage samples. By disentangling the healthy features and damaged features in the scratch damage samples and combining contrast learning, adversarial learning, and multi-task learning techniques, the extraction ability of the generation model for damaged features is enhanced.
[0090] (3) Alleviating the sample imbalance problem and enhancing diversity and authenticity: The present invention strengthens the generation of scratch damage samples through the strategies of adversarial learning and multi-task learning, while keeping the statistical characteristics of the generated samples matched with the real data set. Description of the drawings
[0091] Figure 1 is the flowchart of the method for generating scratch damage samples of the shield main bearing rolling element in the small sample scenario of the embodiment of the present invention;
[0092] Figure 2 is the data flow display diagram of the IWVAE-GAN model in the embodiment of the present invention;
[0093] Figure 3 is the schematic diagram of the encoder network architecture based on Informer in the embodiment of the present invention;
[0094] Figure 4Schematic diagram of the decoder network architecture based on Informer according to an embodiment of the present invention;
[0095] Figure 5 Schematic diagram of the network architecture based on the hybrid connection of double - layer CNN - LSTM according to an embodiment of the present invention. Detailed implementation manners
[0096] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the following examples are given with reference to the accompanying drawings to further elaborate on the present application in detail.
[0097] The embodiment of the present application discloses a method for generating samples of the scratch damage of the rolling elements of the shield machine main bearing in a small - sample scenario. The present invention makes full use of rich healthy - state samples to support the generation of scarce scratch - damage samples. By disentangling the healthy features and damage features in the scratch - damage samples and combining contrast learning, adversarial learning, and multi - task learning techniques, the extraction ability of the generation model for damage features is enhanced. This method can not only learn and reproduce damage features from limited scratch - damage data but also utilize the ability of the generation model to make up for the scarcity of rolling - element scratch - damage samples, thereby improving the accuracy and efficiency of the diagnosis of the scratch damage of the rolling elements of the shield machine main bearing. The process is as Figure 1 shown and includes the following steps:
[0098] S1. Data collection and pre - processing. Collect the vibration data of the healthy state of the shield machine main bearing and the vibration data of the scratch - damage state, and use the z - score standardization method to standardize the vibration data to eliminate the influence of the dimension. Use the time - window strategy to cut to obtain the time - window samples of the scratch - damage state, that is, the scratch - damage samples \(x\) f and the time - window samples of the healthy state, that is, the healthy samples \(x\) h .
[0099] S2. Construct an Importance - Weighted Variational Autoencoder - Generate Adversarial Networks (IWVAE - GAN) model based on contrast learning. As Figure 2 shown, it includes a shared - feature encoder a private - feature encoder a generator (\(g\) θ ), a discriminator (\(h\) θ ) and a classifier (\(f\) θ ).
[0100] S3. Feature extraction by double encoders. Use the shared - feature encoder to encode the scratch - damage samples and healthy samples into the shared - feature latent space, and use the private - feature encoder to encode the scratch - damage samples into the private - feature latent space.
[0101] The IWVAE-GAN network structure contains two different encoders, which extract shared features (healthy features) and private features (scratch damage features) respectively. Considering the long time series characteristics of the samples, the encoder architecture of Informer as shown in Figure 3 is selected as the encoder architecture of the IWVAE-GAN model. The encoder of Informer can accurately capture the complex relationships between time steps, avoid information loss, and thus extract and separate the features in the healthy state and the scratch damage state more effectively.
[0102] S31. Use the private feature encoder to extract the private features of the scratch damage samples, and obtain the distribution parameters of the private feature latent representation of the scratch damage samples. At the same time, use the shared feature encoder to extract the shared features in the scratch damage samples, and obtain the distribution parameters of the shared feature latent representation of the scratch damage samples:
[0103]
[0104] Among them, the subscript s corresponds to the shared features, the subscript z corresponds to the private features, and i corresponds to the sample index. is the scratch damage sample. is the mean of the private feature latent representation of the scratch damage sample. is the standard deviation of the private feature latent representation of the scratch damage sample. is the mean of the shared feature latent representation of the scratch damage sample. is the standard deviation of the shared feature latent representation of the scratch damage sample. is the private feature encoder. is the shared feature encoder.
[0105] S32. Use the shared feature encoder to extract the shared features of the healthy samples, and obtain the distribution parameters of the shared feature latent representation of the healthy samples:
[0106]
[0107] Among them, the superscript h corresponds to the healthy samples. is the healthy sample. is the mean of the shared feature latent representation of the healthy sample. is the standard deviation of the shared feature latent representation of the healthy sample.
[0108] S4. Reparameterized sampling. Sample in the private feature latent space to obtain the private feature latent representation of the scratch damage samples, and sample in the shared feature latent space to obtain the shared feature latent representation of the scratch damage samples and the shared feature latent representation of the healthy samples.
[0109] S41. Calculate the latent representation of the private features of the scratch damage samples Perform N samplings in the latent space of the private feature encoder based on IWAE , that is, sample N latent representations from the variational distribution where
[0110]
[0111] is the latent representation of the private features of the scratch damage samples for the n-th sampling, and ∈ is the noise term of the standard normal distribution .
[0112] S42. Calculate the importance weight α of the latent representation of the private features of each scratch damage sample i,n , and this weight corresponds to the importance in the subsequent generation of each pseudo-scratch damage sample in the data generation process
[0113]
[0114] where is the private feature encoder, is the joint probability distribution of and is the variational posterior distribution of the private feature encoder, indicating the probability of the latent representation given the input sample .
[0115] S43. Calculate the latent representation of the shared features of the scratch damage samples Perform N samplings in the latent space of the shared feature encoder based on VAE , that is, sample N latent representations from the variational distribution where
[0116]
[0117] is the latent representation of the shared features of the scratch damage samples for the n-th sampling .
[0118] S44. Calculate the latent representation of the shared features of the healthy samples. Perform a single sampling in the latent space of the shared feature encoder based on VAE , that is, sample one latent representation from the variational distribution where
[0119]
[0120] Among them, is the potential representation of the shared features of healthy samples.
[0121] S5. Concatenate the potential representation of the private features of the scratch-damaged samples and the potential representation of the shared features of the scratch-damaged samples to obtain the potential representation of the scratch-damaged samples, and concatenate the potential representation of the shared features of the healthy samples with a zero matrix to obtain the potential representation of the healthy samples.
[0122] S51. Concatenate the potential representation of the private features of the scratch-damaged samples and the potential representation of the shared features of the scratch-damaged samples to obtain N potential representations of the scratch-damaged samples:
[0123]
[0124] Among them, is the potential representation of the nth scratch-damaged sample.
[0125] S52. Concatenate the potential representation of the shared features of the healthy samples and a zero matrix to form the potential representation of the healthy samples:
[0126]
[0127] Among them, is the potential representation of the healthy samples.
[0128] S6. Generation of pseudo-healthy samples and pseudo-scratch-damaged samples. Input the potential representation of the scratch-damaged samples into the generator to generate pseudo-scratch-damaged samples, and input the potential representation of the healthy samples into the generator to generate pseudo-healthy samples.
[0129] The network architecture of the generator is as Figure 4 shown. It is an architecture based on the Informer decoder. Compared with other neural network architectures, the architecture of the Informer decoder has stronger long-time series generation ability, and in the case of small samples, it has excellent anti-overfitting ability and generalization performance.
[0130] S61. Generation of pseudo-scratch-damaged samples. Input the N potential representations of the scratch-damaged samples into the generator to generate N scratch-damaged samples:
[0131]
[0132] Among them, is the nth pseudo-scratch-damaged sample, and g θ is the generator.
[0133] S62. Generation of pseudo-healthy samples. Input the final potential representation of the healthy samples into the generator to generate healthy samples:
[0134]
[0135] Among them, is a pseudo-healthy sample.
[0136] S7. Input the scratched damage sample and the healthy sample into the discriminator and the classifier, use the discriminator to distinguish the nature of the sample, and use the classifier to identify the state of the sample.
[0137] The discriminator needs to be able to distinguish the difference between real data and generated data. In the generated time series data, subtle time series differences may indicate whether the data comes from the generator. Select, for example, Figure 5 the CNN-LSTM hybrid architecture shown as the network architecture of the discriminator. This architecture combines the local feature extraction ability of the Convolutional Neural Network (CNN) and the global time series modeling ability of the Long Short-Term Memory (LSTM), enabling the model to capture both short-term time series changes and understand long-term dependencies when processing time series data. The architecture based on the double-layer CNN-LSTM hybrid can not only effectively extract local patterns in the input sequence but also capture long-term dependencies in the time series, efficiently judging the authenticity of the data.
[0138] The state recognizer needs to be able to distinguish the difference between the scratched damage sample and the healthy sample. Select, for example, Figure 5 the CNN-LSTM hybrid architecture shown as the network architecture of the discriminator as the network architecture of the classifier.
[0139] S71. Input the generated sample and the real sample and into the discriminator to obtain the discrimination result on whether the sample belongs to the real sample or the generated sample. The feedback of the discriminator is used to optimize the generator to ensure that the distribution of the generated sample is closer to that of the real sample:
[0140]
[0141] Among them, d i is the determination result of the discriminator, X f is the set of real damaged samples, X h is the set of real healthy samples, and h θ is the discriminator.
[0142] S72. Input the generated sample and the real sample and Input the classifier to obtain the recognition result of whether the sample belongs to the healthy state or the damaged state:
[0143]
[0144] Among them, c i is the recognition result of the classifier, and f θ is the classifier.
[0145] S8. Calculate the loss of IWVAE - GAN. Calculate the loss function of the IWVAE - GAN model. In the IWVAE - GAN model, since it combines the IWVAE, VAE, and GAN models, the loss function includes multiple components (KL divergence loss, reconstruction loss, discriminator loss, generator loss, state recognition loss) to ensure that the generated scratch - damaged samples reach the best in terms of quality and diversity.
[0146] Calculate the KL divergence loss. Since the IWVAE - GAN model uses two different encoders (IWAE and VAE), the KL divergence loss includes the weighted KL divergence loss of IWAE and the standard KL divergence loss of VAE. The KL divergence loss is used to regularize the latent space to ensure that the latent representation conforms to the prior distribution. By minimizing the difference between the true latent distribution and the inferred distribution, the smoothness of the latent space is maintained. The IWVAE - GAN includes the KL divergence loss for the generated healthy samples and the KL divergence loss for the generated scratch - damaged samples
[0147]
[0148] Among them, is the KL divergence loss for the generated pseudo - scratch - damaged samples, is the KL divergence loss for the generated pseudo - healthy samples, is the total KL divergence loss, is the private feature latent representation of the nth pseudo - scratch - damaged sample, s is the shared latent representation, and are the KL loss weight coefficients, which are used to balance the contributions of the KL losses of the pseudo - healthy samples and the pseudo - scratch - damaged samples in the total KL divergence loss.
[0149] Calculate the reconstruction loss. The reconstruction loss is used to measure the difference between the samples output by the generator and the real samples to ensure that the generated samples can reconstruct the original samples as much as possible. The IWVAE - GAN includes the reconstruction loss for the generated healthy samples and the reconstruction loss for the generated scratch - damaged samples
[0150]
[0151]
[0152] wherein, is the reconstruction loss of the generated pseudo-scratch damage samples, is the reconstruction loss of the generated pseudo-healthy samples, is the total reconstruction loss, is the expectation of sampling s from the variational distribution . and are the reconstruction loss weight coefficients, used to balance the contributions of the reconstruction losses of the pseudo-healthy samples and the pseudo-scratch damage samples to the total reconstruction loss.
[0153] Calculate the discriminator loss, which is used to train the discriminator to effectively distinguish real samples from generated samples. The goal of the discriminator is to maximize the discrimination accuracy for real samples while minimizing the discrimination accuracy for generated samples. IWVAE-GAN includes the discriminator loss for real and generated healthy samples and the discriminator loss for real and generated scratch damage samples
[0154]
[0155] wherein, is the discriminator loss of the generated pseudo-healthy samples, is the discriminator loss of the generated pseudo-scratch damage samples, is the total discriminator loss, and are the discriminator loss weight coefficients, used to balance the contributions of the discriminator losses of the pseudo-healthy samples and the pseudo-scratch damage samples to the total discriminator loss.
[0156] Calculate the generator loss. The goal of the generator is to deceive the discriminator by minimizing the output of the discriminator, making the generated samples increasingly close to real samples. IWVAE-GAN includes the generator loss for generated healthy samples and scratch damage samples.
[0157]
[0158] wherein, is the total generator loss, and are the generator loss weight coefficients, used to balance the contributions of the generation losses of the healthy samples and the scratch damage samples to the total generator loss.
[0159] Calculate the state recognition loss, which is used to train the state recognition classifier (also known as the fault diagnosis classifier) to correctly identify healthy samples and damaged samples. The goal of the classifier is to correctly label real and generated healthy and damaged samples as healthy or damaged samples.
[0160]
[0161] Among them, is the state recognition loss of the generated pseudo-healthy samples, is the state recognition loss of the generated pseudo-scratch damaged samples, is the total state recognition loss, and are the state recognition loss weight coefficients, which are used to balance the contributions of the state recognition losses of pseudo-healthy samples and pseudo-scratch damaged samples in the total state recognition loss.
[0162] S9. Update the parameters of each module of the IWVAE-GAN. Calculate the gradients based on the losses obtained in S8, and alternately update the parameters of the shared feature encoder, private feature encoder, generator, discriminator, and classifier.
[0163] During the optimization process of the IWVAE-GAN, the shared feature encoder, private feature encoder, generator, discriminator, and classifier are alternately trained, and each module uses its corresponding loss function during optimization. The specific optimization steps are as follows:
[0164] S91. Freeze the shared feature encoder, private feature encoder, generator, and classifier, and use the gradient descent method to update the parameters of the discriminator:
[0165]
[0166] Among them, η is the learning efficiency, is the discriminator parameter in the t-th round, is the gradient of the discriminator;
[0167] S92. Unfreeze the shared feature encoder, private feature encoder, generator, and classifier, freeze the discriminator, and update the parameters of the shared feature encoder, private feature encoder, generator, and classifier:
[0168]
[0169] Among them, is the shared feature encoder parameter in the t-th round, is the gradient of the shared feature encoder, is the private feature encoder parameter in the t-th round, is the gradient of the private feature encoder, is the generator parameter for the t-th round, is the gradient of the generator, is the classifier parameter for the t-th round, is the gradient of the classifier.
[0170] S10. Repeat S3 - S9 until the training cycle of the IWVAE - GAN model reaches a preset value. In this embodiment, the preset value of the training cycle of the IWVAE - GAN model is 20.
[0171] S11. Generate the scratch damage samples of the main bearing rolling elements of the shield machine. Randomly sample the latent representation s and z f from the standard normal distribution h , and combine s f and z h to obtain the latent representation ω f of the scratch damage sample ω = [s f , z h . Input ω f into the generator (g θ ) of the trained IWVAE - GAN to obtain the generated scratch damage samples of the main bearing rolling elements of the shield machine
[0172]
[0173] where ω f is the latent representation of the scratch damage sample, s f and z h are two latent representations randomly sampled from the normal distribution.
[0174] The present invention makes full use of rich healthy state samples to support the generation of scarce scratch damage samples. By untangling the healthy features and damage features in the scratch damage samples and combining contrastive learning, adversarial learning, and multi - task learning techniques, the extraction ability of the generation model for damage features is enhanced. This method breaks through the limitation of the traditional method that only relies on damage data to generate damage data, and realizes a wider sample coverage and higher sample authenticity. Preliminary experiments show that the average division error of the scratch damage samples of the main bearing rolling elements of the shield machine generated by the present invention is reduced by 16.42% compared with the samples generated by the CSMOTE (Synthetic Minority Over - sampling Technique) method, by 10.11% compared with the samples generated by the VAE method, and by 20.68% compared with the samples generated by the GAN method.
[0175] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A method for generating a sample of scratch damage to the rolling elements of a shield main bearing in a small-sample scenario, characterized in that, It includes the following steps: S1. Collect the vibration data of the main bearing health status and scratch damage status of the shield machine, and use the time window strategy to cut to obtain scratch damage samples and healthy samples; S2. Construct an importance weighted variational autoencoder generative adversarial network model based on contrastive learning, including a shared feature encoder, a private feature encoder, a generator, a discriminator, and a classifier; S3. Use the shared feature encoder and the private feature encoder to encode the samples obtained in S1 into the latent space; S4. Sample in the latent space to obtain the private feature latent representation of the scratch damage sample, the shared feature latent representation of the scratch damage sample, and the shared feature latent representation of the healthy sample; S5. Concatenate the private feature latent representation and the shared feature latent representation of the scratch damage sample to obtain the scratch damage sample latent representation, and concatenate the shared feature latent representation of the healthy sample with a zero matrix to obtain the healthy sample latent representation; S6. Input the scratch damage sample latent representation and the healthy sample latent representation into the generator to generate pseudo scratch damage samples and pseudo healthy samples; S7. Based on the pseudo scratch damage samples, pseudo healthy samples, scratch damage samples, and healthy samples, use the discriminator to discriminate the nature of the samples and use the classifier to identify the status of the samples; S8. Calculate the model losses, including KL divergence loss, reconstruction loss, discriminator loss, generator loss, and status recognition loss; S9. Calculate the model gradients and alternately update the parameters of the shared feature encoder, private feature encoder, generator, discriminator, and classifier; S10. Repeat S3 - S9 until the training cycle reaches the preset value; S11. Use the trained generator to generate scratch damage samples of the shield main bearing rolling elements.
2. The method for generating a sample of scratch damage to the rolling elements of the shield main bearing in a small-sample scenario according to claim 1, wherein The network architectures of the shared feature encoder and the private feature encoder are based on the Informer encoder, the network architecture of the generator is based on the Informer decoder, and the network architectures of the discriminator and the classifier are CNN - LSTM hybrid architectures.
3. The method for generating a sample of scratch damage of the rolling elements of the shield main bearing in a small sample scenario according to claim 2, wherein The S3 includes the following steps: S31. Use the private feature encoder to extract the private features of the scratch damage sample to obtain the distribution parameters of the private feature latent representation of the scratch damage sample. At the same time, use the shared feature encoder to extract the shared features in the scratch damage sample to obtain the distribution parameters of the shared feature latent representation of the scratch damage sample: Among them, the subscript s corresponds to the shared feature, the subscript z corresponds to the private feature, and i corresponds to the sample index. is the scratch damage sample. is the mean of the latent representation of the private features of the scratch damage sample. is the standard deviation of the latent representation of the private features of the scratch damage sample. is the mean of the latent representation of the shared features of the scratch damage sample. is the standard deviation of the latent representation of the shared features of the scratch damage sample. is the private feature encoder. is the shared feature encoder. S32. Use the shared feature encoder to extract the shared features of the healthy sample to obtain the distribution parameters of the shared feature latent representation of the healthy sample: Among them, is a healthy sample, is the mean of the latent representation of the shared features of the healthy samples, is the standard deviation of the latent representation of the shared features of the healthy samples.
4. The method for generating a sample of scratch damage to the rolling elements of the shield main bearing in a small-sample scenario according to claim 3, wherein The S4 includes the following steps: S41. Sample N times in the latent space of the private feature encoder to obtain the private feature latent representation of the scratch damage sample: Among them, is the latent representation of the private features of the scratch damage sample for the nth sampling, and ∈ is the noise term of the standard normal distribution ; S42. Calculate the importance weights of the private feature latent representation of each scratch damage sample: Among them, is a private feature encoder, is the joint probability distribution with ; is the variational posterior distribution of the private feature encoder, indicating the probability of the latent representation given the input sample . S43. Sample N times in the latent space of the shared feature encoder to obtain the shared feature latent representation of the scratch damage sample: Among them, is the shared feature latent representation of the nth sampled scratch damage sample; S44. Sample once in the latent space of the shared feature encoder to obtain the shared feature latent representation of the healthy sample: Among them, is the potential representation of the shared features of healthy samples.
5. The method for generating a sample of scratch damage to the rolling elements of the shield main bearing in a small-sample scenario according to claim 4, wherein The S5 includes the following steps: S51. Concatenate the private feature latent representation of the scratch damage sample and the shared feature latent representation of the scratch damage sample to obtain the latent representations of N scratch damage samples: Among them, is the potential representation of the nth scratch damage sample; S52. Concatenate the shared feature latent representation of the healthy sample and the zero matrix to form the latent representation of the healthy sample: Among them, is the potential representation of healthy samples.
6. The method for generating a sample of scratch damage to the rolling elements of the shield main bearing in a small sample scenario according to claim 5, wherein The S6 includes the following steps: S61. Pseudo-scratch damage sample generation. Input the latent representations of N scratch damage samples into the generator to generate N pseudo-scratch damage samples: Among them, is the nth pseudo-scratch damage sample, g θ is a generator; S62. Healthy sample generation. Input the latent representation of the healthy sample into the generator to generate pseudo-healthy samples: Among them, is a pseudo-healthy sample.
7. The method for generating a sample of scratch damage of rolling elements of a shield main bearing in a small-sample scenario according to claim 6, wherein The S7 includes the following steps: S71. Generate a sample and a real sample and input them into a discriminator to obtain the discrimination result on whether the sample belongs to a real sample or a generated sample. The feedback of the discriminator is used to optimize the generator to ensure that the distribution of the generated sample is closer to that of the real sample: where d i is the judgment result of the discriminator, X f is the set of real damage samples, X h is the set of real healthy samples, h θ is the discriminator; S72. Generate samples and real samples and input them into a classifier to obtain the recognition results of whether the samples belong to the healthy state or the damaged state: Among them, c i is the classifier recognition result, and f θ is the classifier.
8. The method for generating a sample of scratch damage to the rolling elements of a shield main bearing in a small-sample scenario according to claim 7, characterized in that The S8 includes the following steps: Calculate the KL divergence loss: Among them, is the KL divergence loss of the generated pseudo-scratch damage samples, is the KL divergence loss of the generated pseudo-healthy samples, is the total KL divergence loss, is the private feature latent representation of the nth pseudo-scratch damage sample, s is the shared latent representation, and are the KL loss weight coefficients; Calculate the reconstruction loss: Among them, is the reconstruction loss of the generated pseudo-scratch damage sample, is the reconstruction loss of the generated pseudo-healthy sample, is the total reconstruction loss, is the expectation of sampling s from the variational distribution , and are the reconstruction loss weight coefficients; Calculate the discriminant loss: Among them, is the discriminative loss of the generated pseudo-healthy samples, is the discriminative loss of the generated scratch damage samples, is the total discriminative loss, and are the discriminative loss weight coefficients; Calculate the generator loss: Among them, is the total loss of the generator, and is the loss weight coefficient of the generator; Calculate the state recognition loss: Among them, is the state recognition loss of the generated pseudo-healthy samples, is the state recognition loss of the generated scratch damage samples, is the total state recognition loss, and are the state recognition loss weight coefficients.
9. The method for generating a sample of scratch damage of the rolling elements of the shield main bearing in a small-sample scenario according to claim 8, wherein The S9 includes the following steps: S91. Freeze the shared feature encoder, private feature encoder, generator, and classifier, and use the gradient descent method to update the parameters of the discriminator: where η is the learning efficiency, is the discriminator parameter at the t-th round, is the gradient of the discriminator; S92. Unfreeze the shared feature encoder, private feature encoder, generator, and classifier, freeze the discriminator, and update the parameters of the shared feature encoder, private feature encoder, generator, and classifier: Among them, is the parameter of the shared feature encoder in the t-th round, is the gradient of the shared feature encoder, is the parameter of the private feature encoder in the t-th round, is the gradient of the private feature encoder, is the parameter of the generator in the t-th round, is the gradient of the generator, is the parameter of the classifier in the t-th round, is the gradient of the classifier.
10. The method for generating a sample of scratch damage to the rolling elements of the main bearing of a shield tunneling machine in a small-sample scenario according to claim 9, wherein, The S11 includes the following steps: From the standard normal distribution Randomly sample latent representation s in f and z h , merge f and z h Get the potential representation ω of the scratch damage sample f =[s f ,z h ], and ω f Input the trained generator to obtain the scratch damage sample of the rolling element of the shield main bearing: where, ω f is the latent representation of the scratch damage sample, s f and z h are two latent representations randomly sampled from a normal distribution.