Bearing data dimension raising method and system based on radial basis function interpolation

By using the method based on radial basis function interpolation, the data after dimensionality reduction in small sample bearing defect detection is increased, which solves the problem of recovery of high-dimensional feature structures and retains feature connections, and improves the detection performance and generalization ability of the model.

CN120180082AActive Publication Date: 2025-06-20HARBIN ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510255501.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

In the detection of small sample bearing defects, how to effectively increase the dimension reduction data, restore the high-dimensional feature structure of the data, retain the connection between features, provide richer feature information for model training, and improve the detection performance and generalization capabilities of the model.

Method used

The bearing data dimension-raising method based on radial basis function interpolation is adopted. By obtaining the dimensionality-reduced data set, calculating the Euclidean distance, processing the radial basis function, normalizing the weight matrix, and finally multiplying the normalized weight matrix with the original high-dimensional data to obtain the high-dimensional data set after dimensionality-reduced.

Benefits of technology

The high-dimensional feature structure of the data after dimensionality reduction is effectively restored, the connection between features is retained, richer feature information is provided for model training, and the detection performance and generalization ability of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180082A_ABST
    Figure CN120180082A_ABST
Patent Text Reader

Abstract

The invention provides a bearing data dimension raising method and system based on radial basis function interpolation, and belongs to the field of data dimension raising. The objective of the invention is to solve the problem of how to effectively raise the dimension of dimension-reduced data, recover the high-dimensional feature structure of the data, maintain the relation between features, provide richer feature information for model training and improve the detection performance and generalization ability of a model in small sample bearing defect detection. In order to ensure complete restoration of data after dimension reduction, a dimension raising module based on a radial basis function is provided, accurate restoration and diversity restoration of features are achieved, generated data are more vivid, and artifacts are reduced; according to the method, the radial basis function interpolation technology is adopted to perform dimension raising on the data after dimension reduction, the high-dimensional feature structure of the data is recovered, the relation between features is reserved, richer feature information is provided for model training, and the detection performance and generalization ability of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data dimensionality elevation, and in particular, to a bearing data dimensionality elevation method and system based on radial basis function interpolation. Background Art

[0002] In the field of bearing defect detection, especially in the case of small samples, the obtained bearing data is often high-dimensional and the number of samples is limited. In order to reduce the data dimension and calculation cost, the data is usually processed by dimensionality reduction. However, during the dimensionality reduction process, some feature information will inevitably be lost, resulting in the destruction of the high-dimensional feature structure of the data and the weakening of the connection between features. This makes it difficult for the model to fully learn the complex features of the data during subsequent model training, affecting the detection performance and generalization ability of the model. Traditional dimensionality elevation methods such as linear interpolation have certain limitations in dealing with non-linear feature relationships and cannot well restore the high-dimensional feature structure of the data. Summary of the Invention

[0003] The technical problem to be solved by the present invention is:

[0004] To solve the problem of how to effectively elevate the dimensionality of the dimensionality-reduced data, restore the high-dimensional feature structure of the data, retain the connection between features, provide richer feature information for model training, and improve the detection performance and generalization ability of the model in the detection of bearing defects with small samples.

[0005] The technical solution adopted by the present invention to solve the above technical problem:

[0006] The present invention provides a bearing data dimensionality elevation method based on radial basis function interpolation, including the following steps:

[0007] S100. Obtain the dimensionality-reduced bearing data set, which contains multiple low-dimensional data points, and each data point represents the low-dimensional feature vector of a bearing sample;

[0008] S200. Calculate the Euclidean distance, calculate the Euclidean distance from each point to be interpolated to each known low-dimensional data point, and obtain a distance matrix;

[0009] S300. Process the radial basis function, send the distance matrix obtained in step S200 into the radial basis function for processing, and obtain a weight matrix;

[0010] S400. Normalize the weight matrix obtained in step S300 so that the sum of each row is 1, and obtain the normalized weight matrix;

[0011] S500. Calculate the elevated-dimensional data, multiply the normalized weight matrix by the original high-dimensional data, and obtain the elevated-dimensional high-dimensional data set.

[0012] Further, in step S100, the bearing dataset is processed by a manifold approximation and projection dimensionality reduction method.

[0013] Further, the processing process of the manifold approximation and projection dimensionality reduction method includes

[0014] S110. Map the input high-dimensional data into the latent space through an encoder to obtain a latent representation, and place the high-dimensional data h i obtained by the encoder in a D-dimensional space for preprocessing; by calculating the distances between these data points and based on a preset number of nearest neighbor points, adjust the connection number of each data point to determine its nearest neighbor data points, and construct a high-dimensional adjacency graph based on similarity calculation. The edges in the adjacency graph represent the similarity between data points; the similarity s ij is calculated as follows:

[0015]

[0016] In the formula, d(h i , h j ) is the distance between the high-dimensional data points h i and h j , ρ i is the distance from the point h i to the nearest neighbor point, and σ i is a smoothing parameter for adjusting the similarity weight;

[0017] S120. Map the high-dimensional data into an initial low-dimensional space through spectral embedding; first calculate the similarity ws ij between high-dimensional data points and construct a similarity matrix W, then calculate the degree matrix D, and construct the Laplacian matrix La by combining with the identity matrix I:

[0018]

[0019] Perform eigenvalue decomposition on the Laplacian matrix La to obtain the eigenvectors v k and their corresponding eigenvalues λ k , select the eigenvectors corresponding to the first d smallest non-zero eigenvalues to form a matrix V d ; in the low-dimensional space, the representation of the initial data point l i is represented by the i-th row of the V d matrix as:

[0020] Lav k = λ k v k , V d = [v1, v2..., v d , l i = V d[i,:] (3)

[0021] S130. In the low-dimensional space, construct a low-dimensional adjacency graph by calculating the similarity between data points; the similarity between low-dimensional data points l i and l j is expressed as:

[0022] u ij =(1 + a·d(l i , l j )) 2b ) -1 (4)

[0023] In the formula, a and b are positive hyperparameters, and d(l i , l j ) is the distance between points l i and l j in the low-dimensional space;

[0024] S140. Optimize by minimizing the cross-entropy of high-dimensional and low-dimensional similarities. The objective function L is:

[0025]

[0026] S150. During the optimization process, use the stochastic gradient descent method to minimize the loss function, thereby adjusting the low-dimensional embedding to obtain the optimal low-dimensional data set where d is the dimension after data dimensionality reduction to better reflect the structural characteristics of high-dimensional data.

[0027] Furthermore, in step S200, calculate the Euclidean distance from each point l ni to be interpolated to each known low-dimensional data point l j to obtain an m×n distance matrix D = [d ij , where d ij = ||l ni - l j ||.

[0028] Furthermore, in step S300, the obtained weight matrix is an m×n weight matrix R = [r ij , where θ is a set shape parameter.

[0029] Furthermore, in step S400, the obtained normalized weight matrix is R n = [r nij , where

[0030] Furthermore, in step S500, after calculating the upsampled data, an m×D matrix H n= [h ni wherein

[0031]

[0032] A bearing data dimensionality increase system based on radial basis function interpolation, the system having program modules corresponding to the above steps and executing the steps in the above bearing data dimensionality increase method based on radial basis function interpolation when running.

[0033] A computer-readable storage medium storing a computer program configured to implement the steps of a bearing data dimensionality increase method based on radial basis function interpolation when called by a processor.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] A bearing data dimensionality increase method and system based on radial basis function interpolation of the present invention have the following beneficial effects:

[0036] Restore high-dimensional feature structure: Through the radial basis function interpolation technique, the high-dimensional feature structure of the data after dimensionality reduction can be effectively restored, the connection between features is retained, and the data is closer to the original high-dimensional data distribution;

[0037] Provide rich feature information: The dataset after dimensionality increase contains more feature information, which can provide a more comprehensive and accurate feature representation for model training, improving the learning effect and detection performance of the model;

[0038] Improve the generalization ability of the model: The restored high-dimensional feature structure helps the model better understand and learn the complex feature relationships of the data, improving the generalization ability and adaptability of the model in practical applications;

[0039] Strong adaptability: This method is applicable to various types of bearing data, has strong adaptability and a wide range of applications, and can meet the data dimensionality increase requirements in different scenarios; BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a flowchart of a bearing data dimensionality increase method based on radial basis function interpolation in an embodiment of the present invention;

[0041] Figure 2 is a legend of bearing defect data in an embodiment of the present invention;

[0042] Figure 3 is a comparison result graph of dimensionality reduction methods in an embodiment of the present invention;

[0043] Figure 4 is a box plot of the FID index result of the dimensionality reduction effect in an embodiment of the present invention;

[0044] Figure 5 This is the box plot of the IS index result for the dimensionality reduction effect in the embodiments of the present invention;

[0045] Figure 6 This is the comparison result graph of the dimensionality increase functions in the embodiments of the present invention;

[0046] Figure 7 This is the comparison graph of the test results of the dimensionality increase functions in the embodiments of the present invention. Specific Embodiments

[0047] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings.

[0048] Specific Embodiment 1: In combination with Figure 1 As shown, the present invention provides a bearing data dimensionality increase method based on radial basis function interpolation, including the following steps:

[0049] S100. Obtain the dimensionality-reduced bearing data set, which contains multiple low-dimensional data points, and each data point represents the low-dimensional feature vector of a bearing sample; the data set can be processed by the uniform manifold approximation and projection (UMAP) dimensionality reduction method. After dimensionality reduction, the data dimension is lower, and the feature information is more concentrated;

[0050] That is, the data dimensionality reduction module effectively reduces the high-dimensional data generated by the encoder. While retaining the high-dimensional feature similarity, it is transformed into a low-dimensional space; UMAP constructs a high-dimensional adjacency graph by calculating the similarity between data points and maps the data to a low-dimensional space through spectral embedding; for example, the high-dimensional data generated by the encoder is placed in a D-dimensional space for preprocessing. By calculating the distance between data points and based on the preset number of nearest neighbor points, the connection number of each data point is adjusted to determine its nearest neighbor data points;

[0051] This process significantly reduces the data dimension, reduces the interference of redundant background information, enhances the feature discrimination, and thus improves the learning efficiency of the model; dimensionality reduction also provides an efficient basis for the subsequent data augmentation step, making the augmentation process more accurate and efficient; specifically,

[0052] S110. Place the high-dimensional data h i generated by the encoder in a D-dimensional space for preprocessing; by calculating the distance between these data points and based on the preset number of nearest neighbor points, the connection number of each data point is adjusted to determine its nearest neighbor data points, and a high-dimensional adjacency graph is constructed based on the similarity calculation. The edges in the adjacency graph represent the similarity between data points; the similarity s ij is calculated as follows:

[0053]

[0054] In the formula, d(h i ,h j ) is the distance between the high-dimensional data points h i and h j , ρ i is the distance from the point h i to its nearest neighbor, and σ i is a smoothing parameter for adjusting the similarity weights;

[0055] S120. Map the high-dimensional data to an initial low-dimensional space by spectral embedding; first calculate the similarity ws ij between the high-dimensional data points, and construct a similarity matrix W, then calculate the degree matrix D, and construct a Laplacian matrix La by combining with the identity matrix I:

[0056]

[0057] Perform eigenvalue decomposition on the Laplacian matrix La to obtain the eigenvector v k and its corresponding eigenvalue λ k . Select the eigenvectors corresponding to the first d smallest non-zero eigenvalues to form a matrix V d ; in the low-dimensional space, the representation of the initial data point l i is represented by the i-th row of the V d matrix as:

[0058] Lav k = λ k v k , V d = [v1, v2..., v d , l i = V d [i, :] (3)

[0059] S130. In the low-dimensional space, construct a low-dimensional adjacency graph by calculating the similarity between data points; the similarity between the low-dimensional data points l i and l j can be expressed as:

[0060] u ij = (1 + a·d(l i , l j ) 2b ) -1 (4)

[0061] In the formula, a and b are positive hyperparameters, and d(l i , l j ) is the distance between the points l i and l j in the low-dimensional space;

[0062] S140. To preserve the structural features of the high-dimensional adjacency graph as much as possible, it is finally optimized by minimizing the cross-entropy of the high-dimensional and low-dimensional similarities. The objective function L is as follows:

[0063]

[0064] S150. During the optimization process, the Stochastic Gradient Descent (SGD) method is used to minimize the loss function, thereby adjusting the low-dimensional embedding to obtain the optimal low-dimensional dataset. Among them, d is the dimension after data dimensionality reduction, which can better reflect the structural characteristics of the high-dimensional data.

[0065] S200. Calculate the Euclidean distance. Calculate the Euclidean distance from each point l to be interpolated to each known low-dimensional data point l, and obtain an m×n distance matrix D = [d]. Among them, d = ||l - l||. The Euclidean distance is a commonly used method to measure the difference between two data points. By calculating the distance matrix, the similarity degree and spatial relationship between the data points to be upscaled and the known data points can be understood, providing a basis for subsequent radial basis function interpolation. ni to each known low-dimensional data point l j to obtain an m×n distance matrix D = [d ij , where d ij = ||l ni -l j ||; The Euclidean distance is a commonly used method to measure the difference between two data points. By calculating the distance matrix, the similarity degree and spatial relationship between the data points to be upscaled and the known data points can be understood, providing a basis for subsequent radial basis function interpolation.

[0066] S300. Process the radial basis function. Send the distance matrix obtained in step S200 into the radial basis function (RBF) for processing to obtain an m×n weight matrix R = [r]. Among them ij , where θ is the set shape parameter, which is set to 2 in the present invention.

[0067] The radial basis function is a function centered at the origin and symmetric along the radial direction, with good interpolation performance and non-linear mapping ability. In the present invention, using the radial basis function to process the distance matrix can convert the distance information into weight information, and the size of the weight reflects the similarity degree and spatial proximity between the data points to be upscaled and the known data points.

[0068] S400. Normalize the weight matrix obtained in step S300 so that the sum of each row is 1 to obtain the normalized weight matrix R n = [r nij , where

[0069] Normalization can ensure the rationality and consistency of the weights, making the weight distribution of each data point to be upscaled more balanced and avoiding deviations in the interpolation results caused by too large or too small weights.

[0070] S500, Dimensionality - raising data calculation: Multiply the normalized weight matrix by the original high - dimensional data to obtain a high - dimensional data set after dimensionality - raising, represented as an m×D matrix H n =[h ni where Restore the high - dimensional feature structure of the data;

[0071] In this way, the low - dimensional data points can be mapped back to the high - dimensional space, retaining the connections between features, providing richer feature information for model training; the data set after dimensionality - raising can better reflect the original feature distribution and structural characteristics of the data, helping to improve the detection performance and generalization ability of the model.

[0072] Specific implementation plan two: A bearing data dimensionality - raising system based on radial basis function interpolation according to the present invention. This system has program modules corresponding to the above steps and executes the steps in the above - mentioned bearing data dimensionality - raising method based on radial basis function interpolation when running.

[0073] The other combinations and connection relationships in this implementation plan are the same as those in the first specific implementation plan.

[0074] Specific implementation plan three: A computer - readable storage medium according to the present invention. The computer - readable storage medium stores a computer program, and the computer program is configured to implement the steps of the bearing data dimensionality - raising method based on radial basis function interpolation when called by a processor.

[0075] The other combinations and connection relationships in this implementation plan are the same as those in the first specific implementation plan.

[0076] Simulation experiment

[0077] Experimental data set:

[0078] This study used a bearing data set produced by Harbin Bearing Group. The data was collected from a bearing production line. To ensure data quality, the experiment was carried out in a customized experimental shed to control the light source and reduce external interference. Figure 2 Shows an example of the data set. The data set includes the following types: Outer Surface Normal (ON), Outer Surface Rust (OR), Outer Surface Scratch (OS), Side Surface Normal (SN), Side Surface Scratch (SS). 40 representative samples were selected from each type for the experiment.

[0079] Experimental settings:

[0080] The experiment was conducted on the Ubuntu 20.04 operating system, based on the open-source deep learning framework PyTorch, using versions Torch1.8.0 and Torchvision 0.8.0. The computing resources configured for the experiment were an NVIDIA GeForce RTX1080 GPU with 20GB of memory. In the experiment, the number of data augmentations m was 40, the high dimension D was set to 100, the low dimension d was set to 50, and the total number of Gaussian components K used was 10.

[0081] Comparative experiment of the dimensionality reduction module

[0082] The comparative experiment of the dimensionality reduction module aims to evaluate the impact of different dimensionality reduction methods on the quality of the generated data. Four methods, namely direct dimensionality reduction, PCA dimensionality reduction, t-SNE dimensionality reduction, and the dimensionality reduction method of this paper, were used for testing respectively, while keeping other modules unchanged. Figure 3 Some of the data generated by different dimensionality reduction methods are shown. It can be clearly seen that the data generated by the dimensionality reduction method of this paper has clearer details, fewer artifacts, and higher overall quality.

[0083] The evaluation results of the FID and IS metrics for the dimensionality reduction process are presented through Figure 4 and Figure 5 box plots. The scatter points of different colors in the figure represent the results of different dimensionality reduction methods. The box contains 50% of the data points, and the median and mean of each group are also marked in the figure. The dimensionality reduction method of the present invention performs best in terms of median, mean, and overall FID and IS, indicating that the quality and diversity of the generated data after dimensionality reduction using the present invention are higher. At the same time, the height of the box of the algorithm of the present invention in the figure is smaller, indicating that the stability and centrality of the generated data are better.

[0084] Comparative experiment of the dimensionality increase module

[0085] The comparative experiment of the dimensionality increase module tested the effects of different functions in the data dimensionality increase module by ensuring the consistency of the augmented data after each dimensionality reduction method. Eight functions, namely multivariate quadratic, inverse multivariate quadratic, linear, fifth-degree function, Gaussian function, thin plate spline, inverse quadratic, and cubic function, were tested respectively to evaluate the generation results.

[0086] As Figure 6It shows that after passing through the dimensionality reduction module of the present invention, some different dimensionality increase functions are selected to generate representative data. It can be seen that the inverse multi-variate quadratic and multi-variate quadratic functions perform excellently in terms of the data quality after dimensionality increase. They can not only well restore and decode the data with various features sampled by the data augmentation module, but also generate defect combinations that do not exist in the original data. At the same time, the combinations of these defect features look very natural, without the chaotic situation like the thin plate spline interpolation data in the figure, nor the artifact phenomenon caused by the rigid fusion of features caused by dimensionality increase with the Gaussian function in the figure. These results indicate that the use of the multi-variate quadratic function has good effects on this small sample data set.

[0087] As Figure 7 It shows the FID index results of different functions of the dimensionality increase module. It can be clearly seen from the figure that the multi-variate quadratic, inverse multi-variate quadratic, linear, and fifth-degree functions all have good effects. Especially, the data generated by dimensionality increase using the multi-variate quadratic function not only has high quality, but also can adapt to various different dimensionality reduction methods.

[0088] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will all fall within the protection scope of the present invention.

Claims

1. A method for increasing the dimension of bearing data based on radial basis function interpolation, characterized in that: The following steps are involved: S100, obtaining a bearing data set after dimensionality reduction, where the data set includes multiple low-dimensional data points, each data point representing a low-dimensional feature vector of a bearing sample; S200, calculating the Euclidean distance, calculating the Euclidean distance from each point to be interpolated to each known low-dimensional data point, and obtaining a distance matrix; S300, processing the radial basis function, sending the distance matrix obtained in step S200 into the radial basis function for processing to obtain a weight matrix; S400, normalizing the weight matrix obtained in step S300 so that the sum of each row is 1, thereby obtaining a normalized weight matrix; S500, dimensionality-enhanced data calculation, multiplying the normalized weight matrix with the original high-dimensional data to obtain an upgraded high-dimensional data set.

2. The method for increasing the dimension of bearing data based on radial basis function interpolation according to claim 1, characterized in that: In step S100, the bearing data set is processed using manifold approximation and projection dimensionality reduction methods.

3. The method for increasing the dimension of bearing data based on radial basis function interpolation according to claim 2, characterized in that: The manifold approximation and projection dimensionality reduction method processing process includes: S110, the input high-dimensional data is mapped to the latent space through the encoder to obtain the latent representation, and the high-dimensional data h generated by the encoder is mapped to the latent space i Place them in a D-dimensional space for preprocessing; by calculating the distance between these data points and adjusting the number of connections of each data point based on the preset number of nearest neighbor points, to determine its nearest data point, and construct a high-dimensional adjacency graph based on similarity calculation. The edges in the adjacency graph represent the similarity between data points; similarity s ij The calculation of is as follows: In the formula, d(h i ,h j ) is a high-dimensional data point h i and h j The distance between i It's point h i The distance to the nearest neighbor, σ i is a smoothing parameter that adjusts the similarity weight; S120, by means of spectral embedding, high-dimensional data is mapped to an initial low-dimensional space; first, the similarity ws between high-dimensional data points is calculated ij , and construct the similarity matrix W, then calculate the degree matrix D, and combine it with the identity matrix I to construct the Laplace matrix La: Perform eigendecomposition on the Laplace matrix La to obtain the eigenvector v k and its corresponding eigenvalue λ k , select the eigenvectors corresponding to the first d smallest non-zero eigenvalues ​​to form the matrix V d ; In low-dimensional space, the initial data point l i is represented by V d The i-th row of the matrix is ​​represented as: Lav k =λ k v k ,V d =[v1,v2...,v d ],l i =V d [i,:] (3) S130, in low-dimensional space, construct a low-dimensional adjacency graph by calculating the similarity between data points; low-dimensional data point l i With l j The similarity between them is expressed as: u ij =(1+a·d(l i ,l j ) 2b ) -1 (4) Where a and b are positive hyperparameters, d(l i ,l j ) is the midpoint l in the low-dimensional space i and l j The distance between S140, optimize by minimizing the cross entropy of high-dimensional and low-dimensional similarities, the objective function L is: S150. During the optimization process, the stochastic gradient descent method is used to minimize the loss function, thereby adjusting the low-dimensional embedding to obtain the optimal low-dimensional data set. Among them, d is the dimension after data dimensionality reduction to better reflect the structural characteristics of high-dimensional data.

4. The method for increasing the dimension of bearing data based on radial basis function interpolation according to claim 3 is characterized in that: In step S200, each point l to be interpolated is calculated ni To each known low-dimensional data point l j Euclidean distance, and obtain an m×n distance matrix D=[d ij ], where d ij =||l ni -l j ||.

5. The method for increasing the dimension of bearing data based on radial basis function interpolation according to claim 4, characterized in that: In step S300, the obtained weight matrix is ​​an m×n weight matrix R=[r ij ],in θ is the set shape parameter.

6. A method for increasing the dimension of bearing data based on radial basis function interpolation according to claim 5, characterized in that: In step S400, the normalized weight matrix obtained is R n =[r nij ],in 7. A method for increasing the dimension of bearing data based on radial basis function interpolation according to claim 6, characterized in that: In step S500, the dimension-upgraded data is calculated to obtain an m×D matrix H n =[h ni ]in 8. A bearing data dimension-increasing system based on radial basis function interpolation, characterized in that: The system has a program module corresponding to the steps of any one of claims 1 to 7, and executes the steps in the above-mentioned bearing data dimensionality enhancement method based on radial basis function interpolation when running.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the bearing data dimensionality enhancement method based on radial basis function interpolation according to any one of claims 1 to 7 when called by a processor.

Citation Information

Patent Citations

  • Mining area geological landslide displacement prediction method based on radial basis function neural network

    CN118820941A

  • System and method for dimensionality reduction using multidimensional data learning through collaborative filtering

    FR3144360A1