Bearing small sample data expansion method and system based on variational auto-encoder

Through the data expansion method based on the variational autoencoder, combining uniform manifold approximation and projection, regularized Gaussian hybrid model and radial basis function, the traditional generative model is solved in the problem of feature blurring and distortion in the bearing small sample data, and high-quality bearing defect data is generated, which improves the quality of the data set and the generalization ability of the model.

CN120180084AActive Publication Date: 2025-06-20HARBIN ENG UNIV

Patent Information

Application Number
CN202510255506.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Traditional generative models have limitations in bearing small sample data, which are prone to feature blur and distortion, and the data set is poor quality.

Method used

The data expansion method based on the Variational Autoencoder (VAE) is adopted to map high-dimensional data to the potential space through the encoder, and the uniform manifold approximation and projection algorithm are used to reduce the dimensionality, and the Gaussian hybrid model optimized by regularization and particle swarm algorithm is used to expand the data, and the dimensionality is increased through the radial basis function, and finally the high-quality defective data is reconstructed through the decoder.

Benefits of technology

High-quality, fusion features, and realistic bearing defect data are generated, which improves the generation quality of high-resolution bearing images, avoids blur and detail loss, and significantly improves the model's ability to extract effective defect information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180084A_ABST
    Figure CN120180084A_ABST
Patent Text Reader

Abstract

The invention provides a bearing small sample data expansion method and system based on a variational auto-encoder, and belongs to the field of deep learning and data enhancement. The problems that a traditional generation model has limitation in bearing small sample data, feature fuzziness and distortion are prone to occurring, and the data set quality is poor are solved. According to the method, a deep VAE framework is constructed, and a dimension reduction module, a data expansion module and a dimension raising module are used in a potential space; dimensionality reduction is performed on high-dimensional data by adopting a UMAP algorithm, so that the topological structure of the data is effectively reserved, and the extraction efficiency and quality of data features are improved; a Gaussian mixture model combining regularization and particle swarm optimization optimization is used for fitting distribution of scattered small sample data, new data with fusion features are expanded through sampling, and data diversity is increased; a radial basis function is used for nonlinear data dimension raising, new data can be ensured to be accurately mapped back to a high-dimensional space, meanwhile, the relation between features is reserved, and defect data with fusion features is reconstructed through a decoder.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and data augmentation, and in particular, to a method and system for augmenting small-sample bearing data based on a variational autoencoder. Background Art

[0002] In the fields of mechanical engineering and industrial manufacturing, bearings, as key mechanical components, the detection of surface defects is crucial for ensuring the operating efficiency and service life of equipment. Although traditional defect detection methods such as manual inspection and image processing are widely used, they have many limitations in terms of efficiency, accuracy, and stability. In recent years, deep learning techniques, especially deep convolutional neural networks, have made remarkable progress in image recognition and object detection tasks. However, these models usually rely on a large amount of labeled data for training, and in actual industrial applications, it is very difficult to obtain sufficient defect samples, resulting in the problems of limited data sample quantity and unbalanced sample distribution. This limits the full learning of defect features by the model, thus affecting the detection accuracy and the generalization ability of the model.

[0003] In the field of few-shot learning, to address the problem of data scarcity, a variety of generative model methods have been proposed. Goodfellow et al. (2020) proposed the generative adversarial network (GAN) to expand the dataset by generating new samples. Ruyu Wang (2022) further proposed the Defect Transfer GAN (DT-GAN), which uses style transfer and noise generation techniques to generate diverse synthetic defect samples in the case of limited defect samples, thus effectively alleviating the overfitting problem of GAN on small-sample datasets. However, the GAN architecture has high hardware requirements, an unstable training process, and requires reaching the Nash equilibrium, which makes it perform poorly in the task of generating small-sample bearing data.

[0004] As an unsupervised learning neural network, the autoencoder (AE) reconstructs input data through an encoder and a decoder to learn latent representations. Compared with GANs, the training process of autoencoders is more stable and more adaptable to data. The variational autoencoder (VAE) proposed by Kingma et al. (2013) is an improved version of the autoencoder, which generates latent representations of data through variational inference and samples to generate new data. Chen Z et al. (2017) proposed a bearing fault diagnosis method combining a sparse autoencoder and a deep belief network for VAE, which collects and statistically analyzes signals through multiple accelerometers to identify fault signals. Chadebec C et al. (2021) proposed a novel sampling method based on the geometric structure of the latent space, which integrates Riemannian geometric elements into the VAE and explores in the latent space through a random walk algorithm to generate new data points. Experimental results on the OASIS database show that this method can significantly improve the balanced accuracy. VQ-VAE (van den Oord et al., 2017) effectively utilizes the latent space to model data features by learning discrete latent representations.

[0005] However, existing VAE networks still face challenges when dealing with high-dimensional small-sample bearing defect datasets. Due to the overlapping and non-obvious regularities of the features of bearing data, the samples generated by VAE often show a scattered and overlapping distribution in the latent space, and there are multiple extreme data points, resulting in artifacts and unrealistic situations in the generated data. The Gaussian Mixture Model (GMM) has unique advantages in modeling and sampling scattered data. GMM can not only perform soft clustering analysis on data, but also fit a distribution based on existing data to sample new data points that conform to the probability distribution. However, when dealing with high-dimensional data, especially when the sample size is small, the covariance matrix estimation of GMM becomes complex, and at the same time, VAE is prone to blur and distortion in high-resolution data generation, affecting the quality of the generated defects.

[0006] Therefore, it is necessary to perform efficient dimensionality reduction and restoration on the data to retain the key features of the data and reduce the data dimension. Current dimensionality reduction algorithms include PCA (Abdi, H et al., 2010), t-SNE (van der Maaten L J P et al., 2008), and UMAP (McInnes L et al., 2018). As a linear dimensionality reduction method, PCA can achieve dimensionality reduction and restoration, but its effect is limited on complex datasets. While t-SNE and UMAP, as non-linear dimensionality reduction techniques, perform well in data visualization but do not support reverse restoration. Summary of the Invention

[0007] The technical problem to be solved by the present invention is:

[0008] To solve the limitations of traditional generation models in bearing small sample data, where feature blur and distortion are prone to occur and the quality of the data set is poor.

[0009] The technical solution adopted by the present invention to solve the above technical problems:

[0010] The present invention provides a method for augmenting bearing small sample data based on a variational autoencoder, comprising the following steps:

[0011] S100. Map the input high-dimensional data into the latent space through an encoder to obtain a latent representation;

[0012] S200. Data dimensionality reduction. In the latent space, use the uniform manifold approximation and projection algorithm to reduce the dimensionality of the data while maintaining the topological structure of the data; that is, reduce the dimensionality of the high-dimensional data generated by the encoder through a data dimensionality reduction module, and while retaining the high-dimensional feature similarity, transform it into a low-dimensional space;

[0013] S300. Data augmentation. Use a Gaussian mixture model optimized by combining regularization and particle swarm optimization to augment the data after dimensionality reduction in step S200 to generate new data points; that is, generate new meaningful data points in the low-dimensional space through a data augmentation module to solve the overfitting caused by data scarcity and the extreme values and outliers brought about by uneven data distribution;

[0014] S400. Data dimensionality increase. Use a radial basis function to increase the dimensionality of the data after augmentation in step S300 so that it is accurately mapped back to the high-dimensional space; that is, map the augmented low-dimensional data back to the high-dimensional space through a data dimensionality increase module;

[0015] S500. Reconstruct high-quality defect data with fused features through a decoder.

[0016] Further, in step S200, specifically,

[0017] S210. Place the high-dimensional data h i generated by the encoder in a D-dimensional space for preprocessing; by calculating the distances between these data points and based on a preset number of nearest neighbor points, adjust the connection number of each data point to determine its nearest neighbor data points, and construct a high-dimensional adjacency graph based on similarity calculation. The edges in the adjacency graph represent the similarity between data points; the similarity s ij is calculated as follows:

[0018]

[0019] In the formula, d(h i , h j ) is the high-dimensional data point h iand h j The distance between them, ρ i is the distance from point h i to its nearest neighbor point, σ i is the smoothing parameter for adjusting the similarity weight;

[0020] S220. Map the high-dimensional data to an initial low-dimensional space by spectral embedding; first calculate the similarity ws ij between high-dimensional data points, and construct the similarity matrix W. Then calculate the degree matrix D, and construct the Laplacian matrix La by combining it with the identity matrix I:

[0021]

[0022] Perform eigenvalue decomposition on the Laplacian matrix La to obtain the eigenvectors v k and their corresponding eigenvalues λ k . Select the eigenvectors corresponding to the first d smallest non-zero eigenvalues to form the matrix V d ; In the low-dimensional space, the representation of the initial data point l i is given by the i-th row of the V d matrix as:

[0023] Lav k =λ k v k ,V d =[v1,v2...,v d ,l i =V d [i,:] (3)

[0024] S230. In the low-dimensional space, construct a low-dimensional adjacency graph by calculating the similarity between data points; the similarity between the low-dimensional data points l i and l j is expressed as:

[0025] u ij =(1+a·d(l i ,l j ) 2b ) -1 (4)

[0026] where a and b are positive hyperparameters, and d(l i ,l j ) is the distance between the points l i and l j in the low-dimensional space;

[0027] S240. Optimize by minimizing the cross-entropy of the high-dimensional and low-dimensional similarities, and the objective function L is:

[0028]

[0029] S250. During the optimization process, the stochastic gradient descent method is used to minimize the loss function, thereby adjusting the low-dimensional embedding to obtain the optimal low-dimensional data set. Among them, d is the dimension after data dimensionality reduction to better reflect the structural characteristics of high-dimensional data.

[0030] Furthermore, in step S300, it specifically includes

[0031] S310. During the training process of the Gaussian mixture model, set the starting points of the model parameters, including the mean μ j , the mixing weight φ j and the covariance matrix Cov j , and add a small regularization term while initializing the covariance matrix;

[0032] Calculate the probability that the data point belongs to each component, that is, the responsibility degree w ij It is expressed as:

[0033]

[0034] In the formula, the data point l i belongs to the j-th component with a responsibility degree of w ij , and the probability density function of the multivariate Gaussian distribution under the mean μ j and the covariance matrix Cov j is N(l i |μ j , Cov j ), and the initial mixing weight is K is the total number of Gaussian components;

[0035] S320. Optimize the mean μ generated in the expectation step, and use the particle swarm optimization algorithm to search for the optimal mean in the parameter space. The optimization function is shown in the following formula (7):

[0036]

[0037] S330. Through the maximization step, update the model parameters to maximize the likelihood of the data. The update formula is as follows:

[0038]

[0039] By iteratively executing the expectation step, PSO, and maximization step until the iteration termination condition is met, the model is fitted;

[0040] S340. Select a Gaussian distribution according to the mixing weight and randomly sample from this distribution to generate new data points, which is expressed as:

[0041] l n ~N(μ j ,Σ j )(10)

[0042] Use the known n d-dimensional original data sets to estimate the parameters of the GMM, which are subsequently used to generate m interpolated d-dimensional low-dimensional data sets

[0043] Furthermore, in step S400, specifically,

[0044] S410. Calculate the Euclidean distance from each point l to be interpolated ni to each known low-dimensional data point l j to obtain an m×n distance matrix D = [d ij , where d ij = ||l ni - l j ||;

[0045] S420. Feed the distance matrix obtained in step S410 into the radial basis function for processing to obtain an m×n weight matrix R = [r ij , where θ is the set shape parameter;

[0046] S430. Normalize the weight matrix so that the sum of each row is 1 to obtain the normalized weight matrix R n = [r nij , where

[0047] S440. Multiply the normalized weight matrix by the original high-dimensional data to obtain the final interpolated data, represented as an m×D matrix H n = [h ni , where

[0048] Furthermore, in step S100, the input high-dimensional data is processed by an encoder; the encoder is implemented using a convolutional neural network, including multiple convolutional layers and batch normalization layers, and the activation function is GELU; the encoder maps the input data into the latent space to obtain a latent representation; the input data is an RGB image of 512×512×3, the encoder contains five convolutional layers and batch normalization layers, the number of convolutional kernels is 32, 64, 128, 256, 512 in sequence, the convolutional kernel size is 3×3, the stride is 2, the padding is 1, and finally the latent space data is output through a fully connected layer, and the dimension of the latent space is 100.

[0049] Further, in step S500, corresponding to the encoder, the input of the decoder is a latent representation of size 100, which is extended to a feature map of 1024×8×8 through a reshape layer; then through five transposed convolution layers and batch normalization layers, the number of convolution kernels is 1024, 512, 256, 128, 64, 32 in sequence, the convolution kernel size is 3×3, the stride is 2, the padding is 1, the activation function is GELU, and finally the Sigmoid activation function is used to output a reconstructed image with a size of 512×512×3.

[0050] A bearing small sample data augmentation system based on a variational autoencoder, the system has program modules corresponding to the above steps, and executes the steps in the above bearing small sample data augmentation method based on a variational autoencoder when running.

[0051] A computer-readable storage medium, the computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the bearing small sample data augmentation method based on a variational autoencoder when called by a processor.

[0052] Compared with the prior art, the beneficial effects of the present invention are:

[0053] The present invention successfully solves the problem of generating small sample data sets by comprehensively using VAE, UMAP, Gaussian mixture model and RBF, and generates high-quality, feature-fused, and realistic bearing defect data;

[0054] Aiming at the problem of generating small sample bearing defect images, a model based on the variational autoencoder (VAE) network framework is proposed; by designing a deeper network structure and an efficient activation function, the generation quality of high-resolution bearing images is improved, and blurring and detail loss are avoided;

[0055] Aiming at the problem of high-dimensional data feature learning, a latent space dimensionality reduction module based on UMAP is proposed, which integrates the dimensionality reduction ability of UMAP into the latent space of VAE, effectively reduces the gap between dimensions and the number of samples, enhances the model's ability to extract effective defect information, and excludes redundant information;

[0056] In order to augment the rich defect features in scattered data, a latent space augmentation module based on the Gaussian mixture model (GMM) is proposed; the mean value in GMM is optimized by particle swarm optimization (PSO), and a small regularization term is added to solve the problems of data anomalies and overfitting, thereby generating new data that integrates diverse features;

[0057] In order to ensure the complete restoration of the data after dimensionality reduction, a dimensionality increase module based on the radial basis function (RBF) is proposed, which embeds RBF into the decoder of VAE, realizes the accurate restoration of features and the recovery of diversity, generates more realistic data and reduces the appearance of artifacts;

[0058] Compared with multiple existing generation algorithms, this model performs better in terms of indicators such as FID, KID, IS, perceptual similarity, and PRD, demonstrating a high degree of matching between the generated data and the real data, as well as the high quality and diversity of the generated images. In addition, the comparative experiments of the dimensionality reduction and dimensionality increase modules further verify the effectiveness of the algorithm. The public dataset also verifies the versatility of the model. The results of the overfitting experiment show that when the ratio of real data to synthetic data is 1:3, overfitting can be effectively prevented, and the generalization ability and data diversity of the model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a flowchart of a method for augmenting small-sample bearing data based on a variational autoencoder in an embodiment of the present invention;

[0060] Figure 2 It is a schematic structural diagram of a data dimensionality reduction module in an embodiment of the present invention;

[0061] Figure 3 It is a schematic structural diagram of a data augmentation module in an embodiment of the present invention;

[0062] Figure 4 It is a schematic structural diagram of a data dimensionality increase module in an embodiment of the present invention;

[0063] Figure 5 It is a legend of bearing defect data in an embodiment of the present invention;

[0064] Figure 6 It is a comparison chart of generation algorithms in an embodiment of the present invention;

[0065] Figure 7 It is a comparison result chart of dimensionality reduction methods in an embodiment of the present invention;

[0066] Figure 8 It is a box plot of the FID index of the dimensionality reduction effect in an embodiment of the present invention;

[0067] Figure 9 It is a box plot of the IS index of the dimensionality reduction effect in an embodiment of the present invention;

[0068] Figure 10 It is a comparison result chart of dimensionality increase functions in an embodiment of the present invention;

[0069] Figure 11 It is a test result chart of dimensionality increase functions in an embodiment of the present invention;

[0070] Figure 12 It is a generated map of defect deduction of a public dataset in an embodiment of the present invention;

[0071] Figure 13 It is a comparison chart of training accuracy curves at different ratios in an embodiment of the present invention. DETAILED DESCRIPTION

[0072] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0073] Specific implementation plan 1: Combine Figures 1 to 4 As shown, the present invention provides a bearing small sample data expansion method based on variational autoencoder, comprising the following steps:

[0074] S100, mapping the input high-dimensional data into a latent space through an encoder to obtain a latent representation;

[0075] The input high-dimensional data is first processed by the encoder; the encoder is implemented using a convolutional neural network, including multiple convolutional layers and batch normalization layers, and the activation function is GELU; the encoder maps the input data into the latent space to obtain a latent representation; for example, for an RGB image of size 512×512×3, the encoder contains five convolutional layers and batch normalization layers, the number of convolution kernels is 32, 64, 128, 256, 512, the convolution kernel size is 3×3, the step size is 2, and the padding is 1. Finally, the latent space data is output through the fully connected layer, and the latent space dimension is 100;

[0076] Suppose the high-dimensional data set is: Where n is the number of data points in the dataset, and D is the dimension of the decoder output;

[0077] S200, combined Figure 2 As shown, in the latent space, the uniform manifold approximation and projection (UMAP) technology is used to reduce the dimension of the data and maintain the topological structure of the data; that is, the high-dimensional data generated by the encoder is effectively reduced in dimension through the data dimension reduction module, and the high-dimensional feature similarity is retained while converting it to a low-dimensional space; UMAP constructs a high-dimensional adjacency graph by calculating the similarity between data points, and maps the data to a low-dimensional space by spectral embedding; for example, the high-dimensional data generated by the encoder is placed in a D-dimensional space for preprocessing, and the distance between data points is calculated and the number of connections of each data point is adjusted based on the preset number of nearest neighbor points to determine its nearest data point;

[0078] This process significantly reduces the data dimension, reduces the interference of redundant background information, and enhances the feature discrimination, thereby improving the learning efficiency of the model; Dimensionality reduction also provides an efficient basis for subsequent data expansion steps, making the expansion process more accurate and efficient; specifically,

[0079] S210, the high-dimensional data h generated by the encoder iIt is placed in a D-dimensional space for preprocessing; by calculating the distances between these data points and based on a preset number of nearest neighbors, the number of connections of each data point is adjusted to determine its nearest data points, and a high-dimensional adjacency graph is constructed based on similarity calculation. The edges in the adjacency graph represent the similarity between data points; the similarity s ij is calculated as follows:

[0080]

[0081] where d(h i , h j ) is the distance between the high-dimensional data points h i and h j , ρ i is the distance from the point h i to its nearest neighbor, and σ i is a smoothing parameter for adjusting the similarity weight;

[0082] S220. Map the high-dimensional data to an initial low-dimensional space by spectral embedding; first calculate the similarity ws ij between high-dimensional data points and construct a similarity matrix W, then calculate the degree matrix D and construct the Laplacian matrix La by combining it with the identity matrix I:

[0083]

[0084] Perform eigenvalue decomposition on the Laplacian matrix La to obtain the eigenvectors v k and their corresponding eigenvalues λ k . Select the eigenvectors corresponding to the first d smallest non-zero eigenvalues to form the matrix V d ; in the low-dimensional space, the representation of the initial data point l i is given by the i-th row of the V d matrix as:

[0085] Lav k = λ k v k , V d = [v1, v2..., v d , l i = V d [i, :] (3)

[0086] S230. In the low-dimensional space, construct a low-dimensional adjacency graph by calculating the similarity between data points; the similarity between the low-dimensional data points l i and l j can be expressed as:

[0087] u ij = (1 + a·d(l i , lj ) 2b ) -1 (4)

[0088] where a and b are positive hyperparameters, and d(l i , l j ) is the distance between points l i and l j in the low-dimensional space;

[0089] S240. To retain the structural features of the high-dimensional adjacency graph as much as possible, it is finally optimized by minimizing the cross-entropy of the high-dimensional and low-dimensional similarities. The objective function L is:

[0090]

[0091] S250. During the optimization process, the Stochastic Gradient Descent (SGD) method is used to minimize the loss function, thereby adjusting the low-dimensional embedding to obtain the optimal low-dimensional dataset where d is the dimension after data dimensionality reduction to better reflect the structural characteristics of the high-dimensional data;

[0092] S300. As shown in Figure 3 , the Gaussian Mixture Model (GMM) optimized by combining regularization and the Particle Swarm Optimization algorithm is used to expand the dimensionality-reduced data to generate new data points; that is, new meaningful data points are generated in the low-dimensional space through the data augmentation module to solve the overfitting caused by data scarcity and the extreme values and outliers brought about by uneven data distribution;

[0093] The traditional Gaussian Mixture Model (GMM) is unstable in dealing with these problems. Therefore, the GMM is innovatively improved: a regularization term is introduced to enhance the stability of the algorithm, prevent overfitting, and allow the covariance matrix to be singular to ensure stable operation in situations where the traditional GMM is difficult to handle; in addition, by combining the Particle Swarm Optimization (PSO) algorithm, a global search strategy is realized to avoid falling into local optima and effectively complete data augmentation;

[0094] S310. During the training process of the Gaussian Mixture Model (GMM), first, the starting points of the model parameters are set, including the mean μ j , the mixing weight φ j , and the covariance matrix Cov j , and a small regularization term is added while initializing the covariance matrix;

[0095] Then, enter the expectation step, that is, calculate the probability that the data points belong to each component, that is, the responsibility degree w ijThis process can be expressed as:

[0096]

[0097] In the formula, the responsibility degree of data point l i belonging to the j-th component is w ij , and under the mean μ j and the covariance matrix Cov j , the probability density function of the multivariate Gaussian distribution is N(l i |μ j , Cov j ), and the initial mixing weight is where K is the total number of Gaussian components;

[0098] S320. Optimize the mean μ generated in the expectation step, and use the particle swarm optimization (PSO) algorithm to search for the optimal mean in the parameter space, which can avoid the problem of local means; the optimization function is shown in the following formula (7):

[0099]

[0100] S330. Through the maximization step, update the model parameters to maximize the likelihood of the data, and the update formula is as follows:

[0101]

[0102] By iteratively executing the expectation step, PSO, and maximization step until the iteration termination condition is met, the model can be fitted;

[0103] S340. Select a Gaussian distribution according to the mixing weight, and randomly sample from this distribution to generate new data points, which can be expressed as:

[0104] l n ~N(μ j , Σ j )(10)

[0105] Through these steps, first use the known n d-dimensional original data sets to estimate the parameters of the GMM, and these parameters are subsequently used to generate m interpolated d-dimensional low-dimensional data sets

[0106] S400. Combine Figure 4As shown, the augmented data is dimensionally elevated through Radial Basis Function (RBF) to accurately map it back to the high-dimensional space; that is, the augmented low-dimensional data is mapped back to the high-dimensional space through the data dimensional elevation module; this step is crucial for the entire data processing flow because it solves the problem that non-linear dimensionality reduction methods are difficult to fully restore the data, and at the same time overcomes the situation where the performance of deep learning networks is limited when the data is insufficient, as well as the challenge of poor effects that may be brought about by linear dimensional elevation techniques; this module performs interpolation dimensional elevation based on Radial Basis Function (RBF) to ensure that diverse data can be accurately restored, providing the necessary data structure for the decoder to ensure the smooth progress of the decoding process and the integrity of the data; specifically,

[0107] S410. Calculate the Euclidean distance from each point \(l\) to be interpolated to each known low-dimensional data point \(l\), obtaining an \(m\times n\) distance matrix \(D = [d_{ij}]\), where \(d_{ij}=\|l_i - l_j\|\); ni to each known low-dimensional data point \(l\) j to obtain an \(m\times n\) distance matrix \(D = [d_{ij}]\), where ij \(d_{ij}\) ij =\(\|l_i - l_j\|\); ni - \(l_j\) j \|;

[0108] S420. Feed the distance matrix obtained in step S410 into the Radial Basis Function (RBF) for processing to obtain an \(m\times n\) weight matrix \(R = [r_{ij}]\), where ij \(\theta\) is the set shape parameter, which is set to 2 in the present invention; \(\theta\) is the set shape parameter, which is set to 2 in the present invention;

[0109] S430. Normalize the weight matrix so that the sum of each row is 1, obtaining the normalized weight matrix \(R' = [r'_{ij}]\), where n =\([r'_{ij}]\), where nij ;

[0110] S440. Multiply the normalized weight matrix by the original high-dimensional data to obtain the final interpolated data, expressed as an \(m\times D\) matrix \(H = [h_{ij}]\), where n =\([h_{ij}]\), where ni ;

[0111] S500. Reconstruct high-quality defective data with fused features through the decoder;

[0112] The input to the decoder is a latent representation of size 100, which is expanded to a feature map of 1024×8×8 through a reshape layer. Then, through five transposed convolution layers and batch normalization layers, the number of convolution kernels is 1024, 512, 256, 128, 64, 32 in sequence, the convolution kernel size is 3×3, the stride is 2, the padding is 1, the activation function is GELU, and finally the Sigmoid activation function is used to output a reconstructed image with a size of 512×512×3.

[0113] Specific implementation method 2: A bearing small sample data augmentation system based on a variational autoencoder according to the present invention, which has program modules corresponding to the above steps and executes the steps in the above-mentioned bearing small sample data augmentation method based on a variational autoencoder when running.

[0114] The other combinations and connection relationships in this implementation are the same as those in the first specific implementation.

[0115] Specific implementation method 3: A computer-readable storage medium according to the present invention, wherein the computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the bearing small sample data augmentation method based on a variational autoencoder when called by a processor.

[0116] The other combinations and connection relationships in this implementation are the same as those in the first specific implementation.

[0117] Simulation experiment

[0118] I. Dataset and experimental configuration

[0119] In this experiment, a bearing dataset produced by Harbin Bearing Group was used. The data was collected from the bearing production line. To ensure data quality, the experiment was carried out in an experimental shed to control the light source and reduce external interference. Figure 5 Shows an example of the dataset, which contains the following types: Outer Surface Normal (ON), Outer Surface Rust (OR), Outer Surface Scratch (OS), Side Surface Normal (SN), Side Surface Scratch (SS). 40 representative samples were selected from each type for the experiment.

[0120] The experiments were conducted on the Ubuntu 20.04 operating system, based on the open-source deep learning framework PyTorch, using versions Torch1.8.0 and Torchvision 0.8.0. The computing resources configured for the experiments were an NVIDIA GeForce RTX1080 GPU with 20 GB of memory. In the experiments, the number of data augmentations m was 40, the high dimension D was set to 100, the low dimension d was set to 50, and the total number of Gaussian components K used was 10.

[0121] II. Network Comparative Experiments

[0122] To verify the effectiveness of the model of the present invention in bearing data generation, comparative experiments were conducted on various generation algorithms, including GAN (Goodfellow et al., 2020), GMVAE (Dilokthanakul et al., 2017), GVAE (Chadebec, C et al., 2021), VQVAE (van den Oord et al., 2017), RHVAE (Chadebec, C et al., 2020), and Stable Diffusion (SD, Rombach et al., 2021). Table 1 lists the evaluation index results of each generation algorithm, including FID (↓), KID (↓), IS (↑), DS (↑), Fβ (↑), and PS (↓).

[0123] Table 1 Test Index Table of Generation Algorithms

[0124]

[0125] Figure 6 The visual results of different generation algorithms are shown. It can be seen that due to insufficient latent space processing ability, GAN and GMVAE have poor generation effects on the high-dimensional and small-sample bearing data set, with many artifacts; although GVAE introduces Riemannian geometry and random walk algorithms, there are still too many artifacts in the generated data. The data generated by VQVAE and RHVAE is of relatively good quality, but the details are still blurred. In contrast, Stable Diffusion generates better details, but there are still problems of excessive background redrawing and shape distortion overall. The model of the present invention performs best in terms of authenticity, detail retention, and artifact control.

[0126] III. Comparative Experiments on Dimensionality Reduction Modules

[0127] The comparative experiments on dimensionality reduction modules aim to evaluate the impact of different dimensionality reduction methods on the quality of generated data. Four methods, namely direct dimensionality reduction, PCA dimensionality reduction, t-SNE dimensionality reduction, and the dimensionality reduction of the present invention, were used for testing, while keeping other modules unchanged. Figure 7Shows partial data generated by different dimensionality reduction methods. It can be clearly seen that the data generated by the dimensionality reduction of the present invention has clearer details, fewer artifacts, and higher overall quality.

[0128] The evaluation results of the FID and IS metrics for the dimensionality reduction process are presented through Figure 8 and Figure 9 box plots. The scatter points of different colors in the figure represent the results of different dimensionality reduction methods. The box contains 50% of the data points, and the median and mean of each group are also marked in the figure. The dimensionality reduction method of the present invention performs best in terms of median, mean, and overall FID and IS, indicating that the generated data after dimensionality reduction using the present invention has higher quality and diversity. At the same time, the height of the box of the algorithm of the present invention in the figure is smaller, indicating better stability and concentration of the generated data.

[0129] IV. Comparative Experiments on the Dimensionality Increase Module

[0130] The comparison of the dimensionality increase module tests the effects of different functions in the data dimensionality increase module by ensuring the consistency of the augmented data after each dimensionality reduction method. Eight functions, namely multivariate quadratic, inverse multivariate quadratic, linear, fifth-degree function, Gaussian function, thin plate spline, inverse quadratic, and cubic function, are tested respectively, and the generated results are evaluated.

[0131] As Figure 10 shows, after the dimensionality reduction module of the present invention, some representative data generated by different dimensionality increase functions are selected. It can be seen that the inverse multivariate quadratic and multivariate quadratic functions perform excellently in terms of the quality of the data after dimensionality increase. They can not only well restore and decode the data with multiple features sampled by the data augmentation module, but also generate combinations of defects that do not exist in the original data. At the same time, the combinations of these defect features look very natural, without the chaotic situation as shown in the thin plate spline interpolation data in the figure, nor the artifact phenomenon caused by the rigid fusion of features in the dimensionality increase of the Gaussian function as in Figure 10 . These results indicate that the use of the multivariate quadratic function has good effects on this small sample dataset.

[0132] As Figure 11 shows the FID metric results of different functions of the dimensionality increase module. It can be clearly seen from Figure 11 that the multivariate quadratic, inverse multivariate quadratic, linear, and fifth-degree functions all have good effects. In particular, the data generated by dimensionality increase using the multivariate quadratic function not only has high quality but also can adapt to a variety of different dimensionality reduction methods.

[0133] V. Experiments on Public Datasets

[0134] To further verify the generality of the model, experiments are conducted on public datasets. Figure 12Shows the generation results of three groups of steel surface defects in the FSC-20 large-scale small-sample classification dataset of Northeastern University. Each category in this dataset contains 200 images of 200x200, and the deduction process of defect features can be clearly observed, verifying the effectiveness of the model of the present invention in generating diverse defects.

[0135] In addition, generation tests were conducted on the public dataset CelebAHQ256 and compared with mainstream generation algorithms in recent years. The results are shown in Table 2. Although the FID of the model of the present invention is not the lowest among VAE-type networks, it outperforms many models in recent years. The algorithm of the present invention is for generation of small-sample datasets, and the effect on this large-sample dataset is not good.

[0136] Table 2 Test Index Table of Generation Algorithm

[0137]

[0138]

[0139] VI. Overfitting Experiment

[0140] To deeply analyze the impact of the generated data mixed with the original dataset on the classification and detection performance, the following experiment was designed. Combine the generated data with the original data in different proportions to construct multiple mixed datasets, and use these datasets to train a classifier based on ResNet18 and test its performance. The experiment is divided into five groups, and each group uses different virtual-real data proportions, which are 1:0, 1:1, 1:2, 1:3, and 1:4 respectively. Here, "1" represents the original data, and "n" represents the proportion of the generated data.

[0141] Figure 13 Shows the training results under different data combination proportions. The horizontal axis is the training cycle (from 0 to 300), and the vertical axis is the accuracy. Each line represents the training or test accuracy of a certain proportion, where the solid line represents the training accuracy and the dashed line represents the test accuracy. As the training cycle increases, the accuracy of all proportions gradually increases and stabilizes after about 150 cycles. The difference between the training and test accuracies is represented by different colors to intuitively reflect the influence of different combination proportions.

[0142] Table 3 shows the specific values of the training and test accuracies of each proportion dataset under different training rounds.

[0143] Table 3 Training Accuracy Table

[0144]

[0145] According to the data in Table 3, as the number of training rounds increases, the overall accuracy shows an upward trend. Especially when the ratio of real to virtual data is 1:3, both the training and test accuracies of the classifier are excellent, and the gap between them is the smallest, indicating that this ratio effectively controls overfitting. Although the training accuracy is the highest at the 1:1 ratio after 300 rounds of training, the gap between the training and test accuracies for the 1:0, 1:1, and 1:2 ratios is relatively large, suggesting a risk of overfitting for these ratios.

[0146] In contrast, the 1:3 and 1:4 ratios effectively suppress the overfitting phenomenon. However, at the 1:4 ratio, although the test accuracy is relatively high, the decrease in both the training and test accuracies indicates that excessive generated data may introduce interference. In other training rounds, the 1:3 ratio also performs stably in mitigating overfitting. Therefore, the 1:3 ratio of real to virtual data combination is considered the optimal choice for the generated dataset.

[0147] Although the present invention is disclosed as above, the scope of protection of the present invention is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present disclosure, and these changes and modifications will all fall within the scope of protection of the present invention.

Claims

1. A bearing small sample data expansion method based on variational autoencoder, characterized in that: The following steps are involved: S100, mapping the input high-dimensional data into a latent space through an encoder to obtain a latent representation; S200, data dimensionality reduction, in the latent space, uniform manifold approximation and projection algorithm are used to reduce the dimensionality of the data to maintain the topological structure of the data; that is, the high-dimensional data generated by the encoder is reduced in dimension through the data dimensionality reduction module, while retaining the similarity of high-dimensional features, it is converted into a low-dimensional space; S300, data expansion, using a Gaussian mixture model optimized by combining regularization and particle swarm algorithm to expand the data after dimensionality reduction in step S200 to generate new data points; that is, new meaningful data points are generated in the low-dimensional space through the data expansion module to solve the problems of overfitting caused by data scarcity and extreme values ​​and outliers caused by uneven data distribution; S400, data dimension upgrading, upgrading the data expanded in step S300 by radial basis function, so that it can be accurately mapped back to the high-dimensional space; that is, mapping the expanded low-dimensional data back to the high-dimensional space by the data dimension upgrading module; S500, reconstruct high-quality defect data with fusion features through a decoder.

2. According to claim 1, a bearing small sample data expansion method based on variational autoencoder is characterized by: In step S200, specifically including: S210, the high-dimensional data h generated by the encoder i Place them in a D-dimensional space for preprocessing; by calculating the distance between these data points and adjusting the number of connections of each data point based on the preset number of nearest neighbor points, to determine its nearest data point, and construct a high-dimensional adjacency graph based on similarity calculation. The edges in the adjacency graph represent the similarity between data points; similarity s ij The calculation of is as follows: In the formula, d(h i ,h j ) is a high-dimensional data point h i and h j The distance between i It's point h i The distance to the nearest neighbor, σ i is a smoothing parameter that adjusts the similarity weight; S220, by means of spectral embedding, the high-dimensional data is mapped to an initial low-dimensional space; first, the similarity ws between the high-dimensional data points is calculated ij , and construct the similarity matrix W, then calculate the degree matrix D, and combine it with the identity matrix I to construct the Laplace matrix La: Perform eigendecomposition on the Laplace matrix La to obtain the eigenvector v k and its corresponding eigenvalue λ k , select the eigenvectors corresponding to the first d smallest non-zero eigenvalues ​​to form the matrix V d ; In low-dimensional space, the initial data point l i is represented by V d The i-th row of the matrix is ​​represented as: Lav k =λ k v k ,V d =[v1,v2...,v d ],l i =V d [i,:] (3) S230, in low-dimensional space, construct a low-dimensional adjacency graph by calculating the similarity between data points; low-dimensional data point l i With l j The similarity between them is expressed as: u ij =(1+a·d(l i ,l j ) 2b ) -1 (4) Where a and b are positive hyperparameters, d(l i ,l j ) is the midpoint l in the low-dimensional space i and l j The distance between S240, optimize by minimizing the cross entropy of high-dimensional and low-dimensional similarities, the objective function L is: S250, during the optimization process, the stochastic gradient descent method is used to minimize the loss function, thereby adjusting the low-dimensional embedding to obtain the optimal low-dimensional data set Among them, d is the dimension after data dimensionality reduction to better reflect the structural characteristics of high-dimensional data.

3. The bearing small sample data expansion method based on variational autoencoder according to claim 1 is characterized in that: In step S300, specifically including: S310, during the training process of the Gaussian mixture model, setting the starting point of the model parameters, including the mean μ j , mixed weight φ j And the covariance matrix Cov j , and add a small regularization term while initializing the covariance matrix; Calculate the probability that a data point belongs to each component, i.e., the degree of responsibility w ij It is expressed as: In the formula, data point l i The responsibility degree of the jth component is w ij , at the mean μ j and the covariance matrix Cov j The probability density function of the multivariate Gaussian distribution under is N(l i ∣μ j ,Cov j ), the initial mixing weight is K is the total number of Gaussian components; S320, optimize the mean μ generated by the expected step, and use the particle swarm optimization algorithm to search for the optimal mean in the parameter space. The optimization function is shown in the following formula (7): S330, through the maximization step, update the model parameters to maximize the likelihood of the data, the update formula is as follows: The model is fitted by iteratively executing the expectation step, PSO and maximization step until the iteration termination condition is met; S340, select a Gaussian distribution according to the mixed weight, and randomly sample from the distribution to generate new data points, expressed as: l n ~N(μ j ,S j (10) Using a known n d dimensional original dataset To estimate the parameters of the GMM, which is then used to generate m interpolated d-dimensional low-dimensional data sets 4. The bearing small sample data expansion method based on variational autoencoder according to claim 1 is characterized in that: In step S400, specifically including: S410, calculate each point l to be interpolated ni To each known low-dimensional data point l j Euclidean distance, and obtain an m×n distance matrix D=[d ij ], where d ij =||l ni -l j ||; S420: Send the distance matrix obtained in step S410 to the radial basis function for processing to obtain an m×n weight matrix R=[r ij ],in, θ is the set shape parameter; S430, normalize the weight matrix so that the sum of each row is 1, and obtain the normalized weight matrix R n =[r nij ],in, S440, multiply the normalized weight matrix by the original high-dimensional data to obtain the final interpolation data, represented as an m×D matrix H n =[h ni ], 5. The bearing small sample data expansion method based on variational autoencoder according to claim 1 is characterized in that: In step S100, the input high-dimensional data is processed by the encoder; the encoder is implemented using a convolutional neural network, including multiple convolutional layers and batch normalization layers, and the activation function is GELU; the encoder maps the input data to the latent space to obtain a latent representation; the input data is a 512×512×3 RGB image, the encoder contains five convolutional layers and batch normalization layers, the number of convolution kernels is 32, 64, 128, 256, and 512 respectively, the convolution kernel size is 3×3, the step size is 2, and the padding is 1. Finally, the latent space data is output through the fully connected layer, and the latent space dimension is 100.

6. The bearing small sample data expansion method based on variational autoencoder according to claim 5 is characterized in that: In step S500, corresponding to the encoder, the input of the decoder is a potential representation of size 100, which is expanded to a feature map of 1024×8×8 through the reshape layer; then through five deconvolution layers and batch normalization layers, the number of convolution kernels is 1024, 512, 256, 128, 64, and 32 respectively, the convolution kernel size is 3×3, the step size is 2, the padding is 1, the activation function is GELU, and finally the Sigmoid activation function is used to output a reconstructed image of size 512×512×3.

7. A bearing small sample data expansion system based on variational autoencoder, characterized by: The system has a program module corresponding to the steps of any one of claims 1 to 6 above, and executes the steps in the above-mentioned bearing small sample data expansion method based on variational autoencoder when running.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the bearing small sample data expansion method based on variational autoencoder according to any one of claims 1 to 6 when called by a processor.

Citation Information

Patent Citations

  • Intelligent diagnosis method for vibration signal of rotating equipment under unbalanced data

    CN114897022A

  • Fluid pipeline leakage identification method based on multi-classification G-WLSTSVM model

    CN116150687A

  • Method for generating continuous pictures by long text based on diffusion model

    CN117521672A

  • Methods for prognosing mechanical systems

    US20100023307A1

  • Multi-directional scene text recognition method and system based on multi-element attention mechanism

    US20220121871A1

Cited By

  • Multi-view feature fusion multi-mode process virtual sample generation method

    CN120705711A

  • Multi-mode process virtual sample generation method with multi-view feature fusion

    CN120705711B