Bearing fault data generation method

By performing Mel spectrogram processing on the bearing fault data set and building an adaptive sampling feature extraction and data generation model of the generator module, the problem of uneven generation of bearing fault data in the existing technology is solved, and the generated data is more balanced, improving the performance and detection accuracy of the fault diagnosis algorithm.

CN120180339APending Publication Date: 2025-06-20WUXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510499606.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The data categories generated by the existing bearing fault data generation methods are uneven, which makes it difficult for the generated fault data to fully cover the characteristics of different states of the equipment, and is weak in generalization and cannot face complex and changeable fault conditions.

Method used

A bearing failure data generation method is proposed. By processing the bearing failure data set, a Mel spectrogram is obtained, and a bearing failure data generation model including an adaptive sampling feature extraction module and a generator module is constructed. Using the adaptive furthest point sampling method and the adaptive density function, fault data are generated uniformly distributed in different states of the bearing.

Benefits of technology

Improve the balance of the generated bearing fault data, ensure the balanced distribution of data samples in each state, avoid uneven distribution of generated data caused by insufficient data or data imbalance in certain states, and thus optimize the performance and detection accuracy of the fault diagnosis algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180339A_ABST
    Figure CN120180339A_ABST
Patent Text Reader

Abstract

The invention provides a bearing fault data generation method, and relates to the technical field of fault data generation. The method comprises the following steps: firstly, taking a fault data set of a bearing as a first data set, and processing the first data set to obtain a Mel spectrogram of the first data set; thirdly, constructing a bearing fault data generation model; the bearing fault data generation model comprises a self-adaptive sampling feature extraction module and a generator module; information of a Mel spectrogram is coded into point cloud data, a self-adaptive farthest point sampling method in a self-adaptive sampling feature extraction module is utilized to select sampling points so as to cover point cloud data space of feature distribution of different states of the bearing, and global feature vectors of uniform distribution of the different states of the bearing are extracted; and inputting the first data set and the global feature vector into a generator module, and generating a second data set in which different states of the bearing are uniformly distributed by using an adaptive density function, namely generating bearing fault data. According to the method, the constructed bearing fault data generation model is used on the whole, and the balance of the generated bearing fault data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fault data generation, and more specifically, relates to a method for generating bearing fault data. Background Art

[0002] Fault data generation is a process of generating a part of data with fault characteristics through simulation or synthesis techniques when it is impossible to directly obtain or the existing fault samples are insufficient. In the application of real industrial scenarios, since the probability of failure of key equipment inside some machines is relatively low, it is difficult to collect real fault sample data, and the types of real fault sample data are few. When conducting fault diagnosis, there are also certain requirements for the quantity and category of fault data for training the fault diagnosis algorithm based on deep learning. Therefore, the research on the imbalance of fault data categories is particularly important for the field of industrial equipment fault diagnosis. Through the method of data augmentation, the quantity of samples can be effectively supplemented, thereby optimizing the performance of the fault diagnosis algorithm, improving the detection accuracy and coverage of the fault diagnosis algorithm, and greatly enhancing the reliability of the fault diagnosis algorithm in actual application scenarios.

[0003] The existing fault data generation mainly uses the Diffusion model and the generative adversarial network GAN. The Diffusion model is based on the idea of a stochastic process and generates data through two stages: diffusion and denoising. During the diffusion process, the model gradually adds noise to the original data, making it gradually transform into an approximate Gaussian noise distribution; subsequently, during the denoising process, the original data is gradually recovered from the noise through reverse learning, and finally new samples are generated. Due to the staged operation and progressive learning mechanism of this method, the training process is stable and the generated quality is relatively high, especially good at modeling complex data distributions. However, the Diffusion model has a strong dependence on high-quality training data and is limited in application scenarios with limited resources or insufficient data. The adversarial network GAN consists of two parts: a generator and a discriminator. The generator receives random noise to generate simulation data, and the discriminator discriminates the authenticity of the data. The two continuously improve the authenticity of the generated data during adversarial training, thereby expanding the dataset at low cost. However, when the original data distribution is unbalanced, the traditional generative adversarial network GAN will inherit and amplify this imbalance, resulting in the generated samples still being mainly normal data, and the fault samples cannot be effectively supplemented due to the scarcity of the original data. Eventually, the generated data is difficult to comprehensively cover the characteristics of different states of the equipment, the sample distribution of each state is unbalanced, and the generalization of the generated fault data is weak, unable to face complex and changeable fault conditions. Summary of the Invention

[0004] In order to solve the problem of unbalanced data categories generated by the existing bearing fault data generation method, the present invention proposes a method for generating bearing fault data to improve the balance of the generated fault data.

[0005] To achieve the above technical effects, the technical solution of the present invention is as follows:

[0006] Using the bearing fault data set as the first data set, process the first data set to obtain the Mel spectrogram of the first data set;

[0007] Construct a bearing fault data generation model; the bearing fault data generation model includes: an adaptive sampling feature extraction module and a generator module;

[0008] Encode the information of the Mel spectrogram into the point cloud data, and use the adaptive farthest point sampling method in the adaptive sampling feature extraction module to select sampling points to cover the point cloud data space of different state characteristics of the bearing, and extract the global feature vectors uniformly distributed in different states of the bearing;

[0009] Input the first data set together with the global feature vectors into the generator module, and use the adaptive density function to generate a second data set uniformly distributed in different states of the bearing, that is, generate bearing fault data.

[0010] Further, the process of processing the first data set is as follows:

[0011] S11: Use a first-order high-pass filter to pre-emphasize the signals in the fault data set. The expression of the pre-emphasized signal is:

[0012] y[n] = x[n] - α * x[n - 1]

[0013] Where α represents the coefficient of the high-pass filter, x[n] represents the signal in the bearing data set, y[n] represents the pre-emphasized signal, and n represents the sampling points of the signal;

[0014] S12: Perform frame segmentation on the pre-emphasized signal;

[0015] S13: Add a window function to each frame of the signal after frame segmentation to obtain the windowed signal; the expression of the window function is:

[0016]

[0017] Where ω[n] represents the window function, n represents the current sampling point position, n = 0, 1,..., N - 1, and N represents the length of the window function;

[0018] S14: Perform a fast Fourier transform on the windowed signal, calculate the power spectrum, and extract the Mel frequency band energy through the Mel filter bank to obtain the Mel spectrogram of the first data set.

[0019] Further, the process of encoding the information of the Mel spectrogram into the point cloud data is as follows:

[0020] First, map the non-zero elements in the Mel spectrogram to a point in the point cloud, and the expression is:

[0021] P i =(x i ,y i ,z i )

[0022] In the formula, P i represents the point in the point cloud, x i represents the time of the i-th point, y i represents the frequency of the i-th point, and z i represents the spectral energy value of the i-th point;

[0023] Next, perform normalization processing on the point cloud data, and set all points in the point cloud data within a unit sphere. The process is as follows:

[0024] Calculate the centroid of the point cloud data, and the expression is:

[0025]

[0026] In the formula, C represents the centroid, that is, the average position of all point cloud data points, and N represents the number of points in the point cloud data;

[0027] Translate the centroid of the point cloud data to the origin; scale the point cloud data points after translation so that all points in the point cloud data are located within a unit sphere.

[0028] Furthermore, using the adaptive farthest point sampling method, the expression for selecting sampling points is:

[0029]

[0030] In the formula, c j represents the selected sampling point, P represents the point in the point cloud data, C represents the set of selected sampling points, and d(p,c k ) represents the distance from point p to the selected sampling point c k ;

[0031] Furthermore, the process of extracting the global feature vector uniformly distributed in different states of the bearing is as follows: using the sphere neighborhood with the sampling point as the center and radius r, use the PointNet network to extract the local multi-scale features of the sphere neighborhood, introduce the density adaptation weight to adjust the importance of the local multi-scale features, and finally output the global feature vector uniformly distributed in different states of the bearing through the max pooling layer;

[0032] Among them, the expression of the local multi-scale feature is:

[0033]

[0034] In the formula, F multi represents the local multi-scale feature, ∪ r∈R represents the union operation for all scales r, ∪ c∈C performs the union operation on the key sampling points, p i represents a point within the neighborhood N(c,r) with the sampling point as the center of the sphere and radius r, and N represents a set of points with c as the sampling point.

[0035] Furthermore, the generator module includes: an encoder, a style encoder, and a decoder;

[0036] Input the Mel spectrogram into the encoder, downsample the feature map through three layers of convolution to extract the features of the Mel spectrogram;

[0037] Input the global feature vector and the target domain label into the style encoder, concatenate the global feature vector and the target domain label, and output the style vector after passing through a fully connected layer. The expression is:

[0038] s = FC(Concat(t, F global ))

[0039] In the formula, s represents the style vector, F global represents the global feature vector, Concat represents the concatenation operation, FC represents the fully connected layer, and t represents the target domain label;

[0040] Use two layers of transposed convolution in the decoder for upsampling operations, and finally, through one layer of convolution and the Tanh activation function operation, generate a second dataset with a uniform distribution of different states, which is the generated bearing fault data.

[0041] Furthermore, in the generator module, an adaptive instance normalization AdaIN is introduced to dynamically adjust the scaling factor and bias of the features extracted by the encoder. The calculation expression of the adaptive instance normalization AdaIN is:

[0042]

[0043] In the formula, x in represents the features extracted by the encoder, γ represents the scaling factor, β represents the bias, ⊙ represents the main channel multiplication, μ represents the mean, and σ represents the variance.

[0044] Furthermore, the bearing fault data generation model further includes: a discriminator module; constructing a loss function for the generator module using the adaptive density function and the judgment result of the discriminator module, and training the generator module through the loss function;

[0045] The loss function of the generator module includes: adversarial loss function, style loss function, cycle consistency loss function, adaptive density loss function, and maximum mean discrepancy loss function;

[0046] Among them, the expression of the adversarial loss function is:

[0047]

[0048] In the formula, denotes the adversarial loss function, D real denotes the score of the discriminator for the data source, and G denotes the generator;

[0049] The expression of the maximum mean discrepancy loss function is:

[0050]

[0051] In the formula, T represents the distribution of the mel spectrogram, Q represents the target domain data, n represents the number of samples drawn from the distribution of the mel spectrogram, m represents the number of samples drawn from the target domain data, x a , x a′ represent the ath and a'th sample points in the mel spectrogram, y b , y b′ represent the bth and b'th sample points in the target domain data, and k represents the kernel function;

[0052] The expression of the generator module loss function is:

[0053]

[0054] In the formula, denotes the generator module loss function, denotes the adversarial loss function, denotes the style loss function, denotes the cycle consistency loss function, denotes the adaptive density loss function, denotes the maximum mean discrepancy loss function, denotes the hyperparameter.

[0055] Furthermore, the construction process of the adaptive density loss function is as follows:

[0056] Calculate the local density of each pixel in the mel spectrogram of the first dataset, and the expression is:

[0057] ρ(p i ) = (I * K)(p i )

[0058] In the formula, p iDenote the \(i\)-th pixel in the Mel spectrogram, \(I\) represents the Mel spectrogram, \(*\) represents the convolution operation, and \(K\) represents the convolution kernel;

[0059] Calculate the average density of the Mel spectrogram, and the expression is:

[0060]

[0061] In the formula, \(N\) represents the total number of pixels in the Mel spectrogram, represents the global average density of the Mel spectrogram;

[0062] Construct an adaptive sampling loss function according to the square difference between the local density of each pixel and the global average density, and the expression is:

[0063]

[0064] In the formula, represents the adaptive sampling loss function.

[0065] Furthermore, in the discriminator module: Input the Mel spectrogram of the first data set and the second data set with a uniform distribution of different states of the generated bearing into the discriminator. Through three-layer convolution, LeakyReLU activation function and InstanceNorm normalization operation in the discriminator, downsample the features, and compress them into a 256-dimensional feature vector through the global average pooling operation; In the judgment branch, use the fully connected layer to output a 1-dimensional Sigmoid result to distinguish the data source, and in the fault category branch, output the \(K\)-dimensional fault category probability distribution through the fully connected layer.

[0066] Compared with the prior art, the beneficial effects of this method are:

[0067] The present invention provides a bearing fault data generation method. First, process the bearing fault data set and convert it into a Mel spectrogram. Then, construct a bearing fault data generation model including an adaptive sampling feature extraction module and a generator module. Encode the information of the Mel spectrogram into the point cloud data, and use the adaptive farthest point sampling method in the adaptive sampling feature extraction module to select sampling points covering the feature distributions of different fault states of the bearing in the point cloud data, and extract the global feature vector; Through the generator module, use the original data and the global feature module to finally generate bearing fault data with balanced state distribution. The present invention generally uses the constructed bearing fault data generation model to improve the balance of the generated bearing fault data. Brief Description of the Drawings

[0068] Figure 1 represents the flow chart of the bearing fault data generation method proposed in the embodiment of the present invention;

[0069] Figure 2An example diagram showing the Mel spectrogram of the first data set proposed in the embodiments of the present invention;

[0070] Figure 3 An example diagram showing the generated bearing fault data proposed in the embodiments of the present invention. Detailed implementation manners

[0071] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the present patent;

[0072] To better illustrate this embodiment, some parts of the accompanying drawings are omitted, enlarged or reduced, and do not represent the actual size;

[0073] For those skilled in the art, it is understandable that some well-known content descriptions in the accompanying drawings may be omitted.

[0074] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0075] The description of the positional relationship in the accompanying drawings is only for illustrative purposes and should not be construed as limiting the present patent;

[0076] Embodiment 1

[0077] This embodiment proposes a method for generating bearing fault data. As shown in the flowchart of the method, the method proposed in this embodiment generally includes the following steps: Figure 1 As shown in the flowchart of the method, the method proposed in this embodiment generally includes the following steps:

[0078] S1: Using the fault data set of the bearing as the first data set, processing the first data set to obtain the Mel spectrogram of the first data set;

[0079] S2: Constructing a bearing fault data generation model; the bearing fault data generation model includes: an adaptive sampling feature extraction module and a generator module;

[0080] S3: Encoding the information of the Mel spectrogram into the point cloud data, using the adaptive farthest point sampling method in the adaptive sampling feature extraction module to select sampling points to cover the point cloud data space of different state characteristics of the bearing, and extracting the global feature vectors evenly distributed in different states of the bearing;

[0081] S4: Inputting the first data set and the global feature vector into the generator module, and using the adaptive density function to generate a second data set evenly distributed in different states of the bearing, that is, generating bearing fault data.

[0082] In this embodiment, using the fault data set of the bearing as the first data set, processing the first data set to obtain the Mel spectrogram of the first data set, as Figure 2An example diagram of the Mel spectrogram of the first data set shown. Encoding the information of the Mel spectrogram into point cloud data, first converting the Mel spectrogram into the data form of a point cloud, regarding the Mel spectrogram as a two-dimensional plane, where the pixel intensity in the two-dimensional plane corresponds to the point density in the point cloud. By sampling, the high-density area of the image is converted into a dense point group in the point cloud data, and the low-density area is converted into a sparse point group in the point cloud data; the process is as follows:

[0083] First, map the non-zero elements in the Mel spectrogram to a point in the point cloud, and the expression is:

[0084] P i =(x i ,y i ,z i )

[0085] In the formula, P i represents the point in the point cloud, x i represents the time of the i-th point, y i represents the frequency of the i-th point, and z i represents the spectral energy value of the i-th point;

[0086] Next, perform normalization processing on the point cloud data, and set all points in the point cloud data within a unit sphere. The process is as follows:

[0087] Calculate the centroid of the point cloud data, and the expression is:

[0088]

[0089] In the formula, C represents the centroid, that is, the average position of all point cloud data points, and N represents the number of points in the point cloud data;

[0090] Translate the centroid of the point cloud data to the origin, and subtract the position of the centroid from the position of each point; scale the translated point cloud data points so that all points in the point cloud data are located within a unit sphere. Not only located inside a unit sphere, but also has a zero mean as a whole.

[0091] In this embodiment, using the adaptive farthest point sampling method, the expression for selecting the sampling points is:

[0092]

[0093] In the formula, c j represents the selected sampling point, P represents the point in the point cloud data, C represents the set of selected sampling points, and d(p, c k ) represents the distance from point p to the selected sampling point c k ;

[0094] In this embodiment, a set of sampling points C is constructed. An arbitrary point is selected as the initial sampling point and stored in the set of sampling points. The farthest point sampling (FPS) method is used. By iteratively selecting the point farthest from the set of sampling points, the selected sampling points evenly cover the entire data space, fully characterizing the feature distributions of the normal state, fault state, and transition state. The farthest point sampling method adaptively adjusts the sampling density for the non-uniformity of the data distribution, preferentially captures the key features in the low-density regions, and avoids missing local features, thereby improving the model's discrimination ability and generalization performance for multiple fault modes.

[0095] In this embodiment, the process of extracting the global feature vectors uniformly distributed in different states of the bearing is as follows: taking the sampling points as the centers of spheres with a radius of r, using the PointNet network to extract the local multi-scale features of the spherical neighborhoods, introducing density adaptation weights to adjust the importance of the local multi-scale features, and finally outputting the global feature vectors uniformly distributed in different states of the bearing through the max pooling layer;

[0096] Among them, the expression of the local multi-scale features is:

[0097]

[0098] In the formula, F multi represents the local multi-scale features, ∪ r∈R represents the union operation for all scales r, ∪ c∈C represents the union operation for the sampling points, p i represents a point within the neighborhood N(c, r) with the sampling point c as the center of the sphere and a radius of r, and N represents a set of points with c as the sampling point.

[0099] Exemplarily, the radius of the spherical neighborhood is adjusted according to the actual situation. In this embodiment, the radius of the spherical neighborhood is set to 2.

[0100] In this embodiment, when using the PointNet network to extract the features of the spherical neighborhood, the coordinates of all points within the spherical neighborhood are translated relative to the sampling point c j of this region to better learn the local geometric structure. The local region point set is input into the Pointnet network to extract local features. Features at different scales are extracted by the sizes of the spherical neighborhoods at different scales and the sampling point selection strategy, forming multi-scale features, so that the features of each state are evenly distributed and comprehensively extracted.

[0101] In this embodiment, the generator module includes: an encoder, a style encoder, and a decoder;

[0102] The mel spectrogram is input into the encoder, and the feature map is downsampled through three layers of convolution to extract the features of the mel spectrogram;

[0103] Input the global feature vector and the target domain label into the style encoder, concatenate the global feature vector and the target domain label, and output the style vector through a fully connected layer. The expression is as follows:

[0104] s = FC(Concat(t, F global ))

[0105] In the formula, s represents the style vector, F global represents the global feature vector, Concat represents the concatenation operation, FC represents the fully connected layer, and t represents the target domain label;

[0106] Use two transposed convolutions in the decoder for upsampling operations, and finally, through a convolution and a Tanh activation function operation, generate a second dataset with a uniform distribution of different states, which is the generated bearing fault data.

[0107] In this embodiment, input the first dataset together with the global feature vector into the generator module, and use the adaptive density function to generate a second dataset with a uniform distribution of different states of the bearing, that is, generate bearing fault data, as Figure 3 shown in the example diagram of the generated bearing fault data.

[0108] In the generator module, introduce the adaptive instance normalization AdaIN to dynamically adjust the scaling factor and bias of the features extracted by the encoder. The calculation expression of the adaptive instance normalization AdaIN is as follows:

[0109]

[0110] In the formula, x in represents the features extracted by the encoder, γ represents the scaling factor, β represents the bias, ⊙ represents the main channel multiplication, μ represents the mean, and σ represents the variance.

[0111] In this embodiment, the target domain label is used to indicate that the generator switches to a certain target domain, and the global feature vector specifies the style details during image conversion, realizing a smooth transition between different states and increasing the diversity and generalization of the generated data. The generator module can not only use normal data to generate normal data and fault data to generate fault data, but also use normal data to generate fault data and fault data to generate normal data. Therefore, the diversity and generalization of the data are improved.

[0112] In this embodiment, the specific process of the encoder is as follows: The encoder receives a 128×128×3 Mel spectrogram, and gradually downsamples it to a 32×32×256 feature map through three layers of convolution. After each layer of convolution, InstanceNorm instance normalization and the ReLU activation function are connected, and triple residual blocks (including 3×3 convolution and skip connections) are stacked to strengthen deep feature extraction. The activation function is used to increase the non-linearity of the network, and the normalization layer is connected to stabilize and accelerate the training process of the deep neural network and reduce variable offset. The spatial dimension is reduced through downsampling and the computational complexity is reduced. In addition, residual block connections are also set in the overall architecture to promote information flow.

[0113] The generator module generates bearing fault data, generating data with an equal distribution of data samples in each state, avoiding uneven distribution of the generated data caused by insufficient existing data or unbalanced data in some states. At the same time, it effectively supplements the small amount or missing data quantity, ensures that the important features of all states are comprehensive and balanced, further optimizes the performance of the fault diagnosis algorithm, improves the detection accuracy and coverage of the fault diagnosis algorithm, and greatly enhances the reliability of the actual application scenario.

[0114] Exemplarily, the original bearing fault data has 4 categories (cltot = 4), and the four categories are normal data (cli = 1), roller fault data (cli = 2), inner ring fault data (cli = 3), and outer ring fault data (cli = 4). The total amount of data is set to 200, and the data volume of each category is D cli ={D1, D2, D3, D4} = {120, 25, 30, 25}. The bearing fault data generation model constructed according to this embodiment generates bearing fault data, and the proportion of the number of categories of the generated bearing fault data is similar to that of the original data.

[0115] Embodiment 2

[0116] In this embodiment, the process of using the bearing fault data set as the first data set and processing the first data set to obtain the Mel spectrogram of the first data set will be described in detail.

[0117] The fault data set of the bearing in this embodiment uses the CWRU (Case Western Reserve University) data set. In the CWRU data set, there are ten classification data, namely: inner ring fault of the bearing with a loss diameter of 0.1778 mm, inner ring fault of the bearing with a loss diameter of 0.3556 mm, inner ring fault of the bearing with a loss diameter of 0.5334 mm, outer ring fault of the bearing with a loss diameter of 0.1778 mm, outer ring fault of the bearing with a loss diameter of 0.3556 mm, outer ring fault of the bearing with a loss diameter of 0.5334 mm, outer ring fault of the bearing with a loss diameter of 0.1778 mm, outer ring fault of the bearing with a loss diameter of 0.3556 mm, outer ring fault of the bearing with a loss diameter of 0.3556 mm, and data of the normal bearing state.

[0118] In this embodiment, the data set of the 12k drive end (DE) bearing is selected as the first data set, the signals in the first data set are read, and the first data set is processed to obtain the Mel spectrogram of the first data set.

[0119] In this embodiment, the processing process of the first data set is as follows:

[0120] S11: Use a first-order high-pass filter to pre-emphasize the signals in the fault data set. The expression of the pre-emphasized signal is:

[0121] y[n] = x[n] - α * x[n - 1]

[0122] In the formula, α represents the coefficient of the high-pass filter, x[n] represents the signals in the bearing data set, y[n] represents the pre-emphasized signal, and n represents the sampling points of the signal.

[0123] The original vibration signal is pre-emphasized by a first-order high-pass filter to suppress low-frequency components, enhance high-frequency features, and balance the spectral energy distribution.

[0124] S12: Perform frame segmentation on the pre-emphasized signal;

[0125] In this embodiment, through frame segmentation, the pre-emphasized signal is divided into several parts, each part represents a frame, and the relationship between frequency and time is obtained. Based on the high-frequency characteristics of the bearing fault characteristic frequency, the frame length is set to 2048 sampling points, while capturing important frequency components, the local time variation information is retained. The frame shift is selected as 512 sampling points, which not only maintains the time resolution but also reduces the calculation amount. The signal after frame segmentation ensures good time resolution while capturing the frequency characteristics of the bearing fault.

[0126] S13: Add a window function to each frame of the signal after frame segmentation to obtain the windowed signal; the expression of the window function is:

[0127]

[0128] In the formula, ω[n] represents the window function, n represents the position of the current sampling point, n = 0, 1, …, N - 1, and N represents the length of the window function.

[0129] Adding a window function to reduce the boundary effect caused by signal truncation, smoothly reducing the signal value to zero at the beginning and end of the signal, thereby avoiding the phenomenon of spectral leakage.

[0130] S14: Perform a fast Fourier transform on the windowed signal, calculate the power spectrum, extract the Mel frequency band energy through the Mel filter bank, and obtain the Mel spectrogram of the first data set.

[0131] In this embodiment, a fast Fourier transform (FFT) is performed on each frame of the windowed signal to obtain the complex spectrum, the modulus square of the complex result of each frequency component is taken to obtain the power spectrum; the power spectrum is passed through the Mel filter bank, and the output of each filter is integrated to obtain the energy of each Mel frequency band, and the Mel spectrogram of the first data set is obtained.

[0132] The Mel spectrogram maps the power spectrum to the Mel frequency scale, then discretizes and accumulates it to form a two-dimensional Mel spectrogram. The horizontal axis represents time, the vertical axis represents the Mel frequency, and the color intensity of the image represents the energy of each frequency component. The Mel frequency scale is a non-linear transformation, which helps to identify and analyze features that may be ignored by the linear frequency. Finally, the Mel filter bank is weighted and rearranged through the Mel frequency scale to obtain an image reflecting the frequency characteristics of the signal.

[0133] Embodiment 3

[0134] In this embodiment, the bearing fault data generation model further includes: a discriminator module; constructing a loss function of the generator module using the adaptive density function and the judgment result of the discriminator module, and training the generator module through the loss function.

[0135] The loss function of the generator module includes: an adversarial loss function, a style loss function, a cycle consistency loss function, an adaptive density loss function, and a maximum mean error loss function;

[0136] Among them, the expression of the adversarial loss function is:

[0137]

[0138] In the formula, represents the adversarial loss function, D real represents the score of the discriminator for the data source, and G represents the generator;

[0139] The expression of the maximum mean error loss function is:

[0140]

[0141] Wherein, T represents the distribution of the Mel spectrogram, Q represents the target domain data, n represents the number of samples extracted from the distribution of the Mel spectrogram, m represents the number of samples extracted from the target domain data, and x a 、x a′ represent the a-th and a'-th sample points in the Mel spectrogram, and y b 、y b′ represent the b-th and b'-th sample points in the target domain data, and k represents the kernel function;

[0142] The expression of the loss function of the generator module is:

[0143]

[0144] Wherein, represents the loss function of the generator module, represents the adversarial loss function, represents the style loss function, represents the cycle consistency loss function, represents the adaptive density loss function, represents the maximum mean error loss function, represents the hyperparameter.

[0145] The hyperparameter is used to balance the importance of each loss, and the hyperparameter is adjusted accordingly according to the requirements of the actual situation and the data set to achieve the best effect of generating fault data.

[0146] In this embodiment, by constructing an adaptive density loss function, the global feature uniformity distribution is adapted. For two-dimensional Mel spectrogram data, first calculate the pixel density of the image, and according to the loss function, a penalty mechanism is applied to the over-generated dense area to enhance the attention to the sparse area, and the regional density area of the generated image is adjusted so that the global features are fully concerned, and image data with a uniform distribution of each state is generated.

[0147] The construction process of the adaptive density loss function is as follows:

[0148] Calculate the local density of each pixel in the Mel spectrogram of the first data set, and the expression is:

[0149] ρ(p i ) = (I * K)(p i )

[0150] Wherein, p i represents the i-th pixel in the Mel spectrogram, I represents the Mel spectrogram, * represents the convolution operation, and K represents the convolution kernel;

[0151] The local density of each pixel is calculated using a convolution operation, and the local density is represented as the number of non-zero pixels around each pixel. The convolution kernel K is used to calculate the local density of the image I.

[0152] Calculate the average density of the Mel spectrogram, and the expression is:

[0153]

[0154] In the formula, N represents the total number of pixels of the Mel spectrogram, represents the global average density of the Mel spectrogram;

[0155] According to the square difference between the local density of each pixel and the global average density, the sum of the squares of the differences is used as the loss value to construct an adaptive sampling loss function, and the expression is:

[0156]

[0157] In the formula, represents the adaptive sampling loss function.

[0158] In the discriminator module: The Mel spectrogram of the first data set and the second data set with a uniform distribution of different states of the generated bearing are input into the discriminator. Through three-layer convolution, LeakyReLU activation function and InstanceNorm normalization operations in the discriminator, the features are downsampled, and compressed into a 256-dimensional feature vector through global average pooling operation; in the judgment branch, a 1-dimensional Sigmoid result is output using a fully connected layer to discriminate the data source, and in the fault category branch, a K-dimensional fault category probability distribution is output through a fully connected layer.

[0159] The judgment branch discriminates the difference between the generated data and the real data, outputs a authenticity score, optimizes the quality of the generated data, and helps to adjust the data in time to conform to the style and attributes of the target domain.

[0160] Construct the loss function of the discriminator module, calculate the accuracy of the fault category probability of the generated data and the judgment classification of the actual category label, optimize the discriminator module, and improve the judgment and classification accuracy.

[0161] The loss function of the discriminator module includes: an adversarial loss function and a classification cross-entropy loss function;

[0162] Among them, the expression of the classification cross-entropy loss function is:

[0163]

[0164] In the formula, represents the classification cross-entropy loss function, k represents the fault category, y kOne-hot encoding representing the true label, D class,k Represents the probability that the output of the discriminator module belongs to the k-th type of fault;

[0165] The expression of the loss function of the discriminator module is:

[0166]

[0167] In the formula, Represents the loss function of the discriminator module, Represents a hyperparameter, Represents a hyperparameter, λ adv Represents a hyperparameter, λ class Represents a hyperparameter.

[0168] According to the calculation result of the loss function of the discriminator module, update the weights of the convolutional layer, InstanceNorm parameters, and fully connected layer parameters through backpropagation, and alternately train with the generator. Set the training frequency of the discriminator module and the generator module to 1:2, and the discriminator module improves the ability to distinguish true and false and the accuracy of fault classification.

[0169] The embodiments are only examples for clearly explaining the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for generating bearing fault data, characterized in that: The following steps are involved: Taking the bearing fault data set as the first data set, processing the first data set to obtain a Mel frequency spectrum diagram of the first data set; Construct a bearing fault data generation model; The bearing fault data generation model includes: an adaptive sampling feature extraction module and a generator module; The information of the Mel spectrum map is encoded into the point cloud data. The sampling points are selected using the adaptive farthest point sampling method in the adaptive sampling feature extraction module to cover the point cloud data space of the bearing's different state feature distributions, and the global feature vectors of the bearing's different states are evenly distributed are extracted. The first data set is input into the generator module together with the global feature vector, and the second data set with uniform distribution in different bearing states is generated by using the adaptive density function, that is, the bearing fault data is generated.

2. A method for generating bearing fault data according to claim 1, characterized in that: The process of processing the first data set is: S11: Use a first-order high-pass filter to pre-emphasize the signal in the fault data set. The expression of the pre-emphasized signal is: y[n]=x[n]-α*x[n-1] In the formula, α represents the coefficient of the high-pass filter, x[n] represents the signal in the bearing data set, y[n] represents the pre-emphasized signal, and n represents the sampling point of the signal; S12: performing frame processing on the pre-emphasized signal; S13: adding a window function to each frame signal after the frame processing to obtain a windowed signal; the expression of the window function is: Wherein, ω[n] represents the window function, n represents the current sampling point position, n=0,1,…,N-1, and N represents the length of the window function; S14: Perform fast Fourier transform on the windowed signal, calculate the power spectrum, extract the Mel frequency band energy from the power spectrum through a Mel filter bank, and obtain a Mel frequency spectrum diagram of the first data set.

3. A method for generating bearing fault data according to claim 1, characterized in that: The process of encoding the information of the Mel spectrum map into the point cloud data is: First, the non-zero elements in the Mel spectrum map are mapped to a point in the point cloud. The expression is: P i =(x i ,y i ,z i ) Where P i represents a point in the point cloud, x i represents the time of the i-th point, y i represents the frequency of the ith point, z i Represents the spectrum energy value of the i-th point; Next, the point cloud data is normalized and all points in the point cloud data are set within a unit sphere. The process is as follows: Calculate the centroid of the point cloud data, the expression is: In the formula, C represents the centroid, that is, the average position of all point cloud data points, and N represents the number of points in the point cloud data; The centroid of the point cloud data is translated to the origin; the translated point cloud data points are scaled so that all points in the point cloud data are located in a unit sphere.

4. A method for generating bearing fault data according to claim 3, characterized in that: Using the adaptive farthest point sampling method, the expression for selecting the sampling point is: In the formula, c j represents the selected sampling point, P represents the point in the point cloud data, C represents the set of selected sampling points, d(p,c k ) represents the distance from point p to the selected sampling point c k distance.

5. A method for generating bearing fault data according to claim 4, characterized in that: The process of extracting the uniformly distributed global feature vectors of the bearing in different states is as follows: taking the sampling point as the sphere with a radius of r, using the PointNet network to extract the local multi-scale features of the sphere neighborhood, introducing density adaptive weights to adjust the importance of the local multi-scale features, and finally outputting the uniformly distributed global feature vectors of the bearing in different states through the maximum pooling layer; Wherein, the expression of the local multi-scale feature is: In the formula, F multi represents the local multi-scale features, U r∈R Indicates the union operation of all scales r, U c∈C Indicates the union operation of the sampling points, p i represents a point in the neighborhood N(c,r) with the sampling point as the center and radius r, and N represents a point set with c as the sampling point.

6. A method for generating bearing fault data according to claim 1, characterized in that: The generator module includes: an encoder, a style encoder and a decoder; The Mel-spectrogram is input into the encoder, and the feature map is downsampled through three layers of convolution to extract the features of the Mel-spectrogram; The global feature vector and the target domain label are input into the style encoder, the global feature vector and the target domain label are concatenated, and the style vector is output through the fully connected layer. The expression is: s=FC(Concat(t,F global )) In the formula, s represents the style vector, F global represents the global feature vector, Concat represents the concatenation operation, FC represents the fully connected layer, and t represents the target domain label; The upsampling operation is performed using two layers of transposed convolution in the decoder, and finally a layer of convolution and Tanh activation function operation is performed to generate a second data set with uniform distribution in different states, which is the generated bearing fault data.

7. A method for generating bearing fault data according to claim 6, characterized in that: In the generator module, adaptive instance normalization AdaIN is introduced to dynamically adjust the scaling factor and bias of the encoder to extract features. The calculation expression of adaptive instance normalization AdaIN is: In the formula, x in Represents the features extracted by the encoder, γ represents the scaling factor, β represents the bias, ⊙ represents the main channel multiplication, μ represents the mean, and σ represents the variance.

8. A method for generating bearing fault data according to claim 1, characterized in that: The bearing fault data generation model further includes: a discriminator module; a loss function of the generator module is constructed using the adaptive density function and the judgment result of the discriminator module, and the generator module is trained through the loss function; The loss functions of the generator module include: adversarial loss function, style loss function, cycle consistency loss function, adaptive density loss function and maximum mean error loss function; Among them, the expression of the adversarial loss function is: In the formula, represents the adversarial loss function, D real represents the score of the discriminator on the data source, and G represents the generator; The expression of the maximum mean error loss function is: Where T represents the distribution of the Mel-spectrogram, Q represents the target domain data, n represents the number of samples extracted from the distribution of the Mel-spectrogram, m represents the number of samples extracted from the target domain data, and x a 、x a′ Indicates the ath and a'th sample points in the Mel spectrum graph, y b ,y b′ represents the bth and b'th sample points in the target domain data, and k represents the kernel function; The expression of the generator module loss function is: In the formula, represents the generator module loss function, represents the adversarial loss function, represents the style loss function, represents the cycle consistency loss function, represents the adaptive density loss function, represents the maximum mean error loss function, Represents a hyperparameter.

9. A method for generating bearing fault data according to claim 8, characterized in that: The construction process of the adaptive density loss function is: Calculate the local density of each pixel in the Mel spectrum of the first data set. The expression is: p(p i )=(I*K)(p i ) In the formula, p i represents the i-th pixel in the Mel spectrum map, I represents the Mel spectrum map, * represents the convolution operation, and K represents the convolution kernel; Calculate the average density of the Mel spectrum graph, the expression is: Where N represents the total number of pixels in the Mel spectrum graph. Represents the global average density of the Mel-spectrogram; According to the square difference between the local density of each pixel and the global average density, an adaptive sampling loss function is constructed, which is expressed as: In the formula, represents the adaptive sampling loss function.

10. A method for generating bearing fault data according to claim 8, characterized in that: In the discriminator module: the Mel spectrum of the first data set and the generated second data set with uniform distribution of different bearing states are input into the discriminator, and the features are downsampled through three layers of convolution, LeakyReLU activation function and InstanceNorm normalization operation in the discriminator, and compressed into a 256-dimensional feature vector through global average pooling operation; in the judgment branch, the fully connected layer is used to output the 1-dimensional Sigmoid result to identify the data source, and in the fault category branch, the fully connected layer outputs the K-dimensional fault category probability distribution.