Intelligent bearing health monitoring method based on MSE-SPP-KAResnet
By adopting the MSE-SPP-KAResnet method in bearing health monitoring, the problem of data set imbalance and limited single-channel signal information is solved, and efficient and accurate diagnosis of bearing failures is achieved.
Patent Information
- Application Number
- CN202510094185.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art faces problems in bearing fault diagnosis of imbalance, limited single-channel signal information, and difficult to describe complexity and uncertainty characteristics in traditional time-frequency feature extraction.
Using an intelligent bearing health monitoring method based on MSE-SPP-KAResnet, the multi-scale features of multi-channel data are extracted through multi-scale sample entropy, spatial pyramid pooling layer and improved residual module, and the model parameters are updated through cross-entropy loss function and backpropagation to achieve fault diagnosis.
It effectively solves the accuracy problems under multi-channel data fusion and imbalanced data sets, improves the accuracy and robustness of bearing fault identification, and enhances the overall accuracy of fault diagnosis.
Smart Images

Figure CN119935554A_ABST
Abstract
Description
Technical Field
[0002] The present invention relates to the technical field of bearing health monitoring, and in particular to an intelligent bearing health monitoring method based on MSE-SPP-KAResnet. Background Art
[0004] In industrial environments, rolling bearings are key components of rotating machinery, and their health is directly related to the safe and stable operation of the equipment. However, bearings often work under harsh conditions and are easily affected by various factors and fail. Bearing failures will not only cause equipment downtime and affect production efficiency, but may also cause safety accidents and cause huge economic losses. With the development of industrial automation and intelligence, the demand for bearing health monitoring is becoming increasingly urgent. In recent years, with the rapid development of sensor technology and communication technology, a large amount of bearing operation data has been collected. This data provides a valuable source of information for equipment health status monitoring and fault diagnosis. At the same time, the rise of deep learning technology has also brought new solutions to the field of fault diagnosis. Deep learning technology does not need to rely on expert experience, and has the characteristics of fast processing speed and high accuracy. It has become one of the core technologies in the field of fault diagnosis.
[0005] 1. In actual bearing fault diagnosis, the amount of fault data is usually much less than the amount of normal state data, resulting in an unbalanced data set. The deep learning model tends to favor majority class samples during training, making it difficult to fully learn and identify the features of minority class fault samples, thus affecting the accuracy of fault diagnosis.
[0006] 2. Traditional fault diagnosis methods are mostly based on feature extraction based on single-channel data. However, the information of single-channel signals is limited and it is difficult to fully reflect the health status of the bearing. Multi-channel data fusion can make up for the problem of insufficient single-channel signal information, but how to effectively fuse multi-channel data and extract representative and distinguishing features is a technical challenge currently faced;
[0007] 3. When bearings work in complex environments, their fault characteristic signals often have high complexity and uncertainty. Traditional time-frequency feature extraction methods are difficult to fully describe these characteristics. How to accurately describe and extract the complexity and uncertainty characteristics in the signal to improve the robustness and accuracy of fault diagnosis is a technical problem that needs to be solved urgently. Summary of the invention
[0009] The purpose of the present invention is to provide an intelligent bearing health monitoring method based on MSE-SPP-KAResnet to solve the problems raised in the above background technology.
[0010] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0011] The intelligent bearing health monitoring method based on MSE-SPP-KAResnet includes the following steps:
[0012] Step 1: Use multiple three-axis acceleration sensors to collect data from the equipment bearings to obtain multi-channel raw signal data to provide basic data support for subsequent fault diagnosis;
[0013] Step 2: Use a sliding window to split the data and divide the data set into a training set and a test set, with a ratio of 7:3 between the training set and the test set, to prepare for model training and testing.
[0014] Step 3: Use multi-scale sample entropy (MSE) to process the training set and test set separately. The scale of MSE is 40, the template length is 1, and the matching threshold is 0.2. The features describing the complexity and uncertainty of the signal are extracted, and the model parameters are initialized. The features are extracted through the convolution layer, and the ReLu activation function is used for nonlinear transformation.
[0015] Step 4: Multi-scale feature fusion of multi-channel data is performed through the spatial pyramid pooling layer to capture spatial information of different scales and enhance the model's feature extraction capability for multi-channel heterogeneous data.
[0016] Step 5, design the KARes module and construct the corresponding KA layer module through the Kolmogorov-Arnold representation theorem to improve the residual module for feature processing;
[0017] Step 6: Use the cross entropy loss function to measure the classification error and update the model parameters through back propagation and gradient descent.
[0018] Step 7: After the training process is completed, the fault diagnosis capability of the model is saved and completed, and the diagnosis results are output.
[0019] A further improvement of the technical solution of the present invention is that in step 1, the process of obtaining multi-channel original signal data includes:
[0020] Step 11, identify the bearings to be monitored and the equipment they are located in, and select a three-axis acceleration sensor with an appropriate range, frequency response range and accuracy based on the size, speed and load parameters of the bearing. The three-axis acceleration sensor can simultaneously measure the vibration acceleration in three mutually perpendicular directions, providing data support for comprehensive monitoring of the vibration state of the bearing;
[0021] Step 12, install the three-axis acceleration sensor near the bearing seat or on a component directly connected to the bearing. The installation position should be as close to the bearing as possible to more accurately capture the vibration signal of the bearing. At the same time, avoid installing it on other components with large vibration to avoid introducing noise interference, and connect the output signal line of the three-axis acceleration sensor to the input port of the data acquisition device;
[0022] Step 13, according to the working conditions and monitoring requirements of the bearing, set the sampling rate, sampling time and sampling period parameters of the data acquisition device. The sampling rate should be higher than twice the highest frequency component of the bearing vibration signal to avoid aliasing. The sampling time should be long enough to collect enough vibration signal data for subsequent analysis and processing, and then start the sensor for data acquisition, capture the bearing vibration signal in real time, and convert it into an electrical signal for transmission;
[0023] Step 14, using multiple three-axis acceleration sensors, simultaneously collects raw signal data of multiple channels, including signals in the radial vibration, axial vibration and tangential vibration directions of the bearing. The collected raw signal data will be recorded in real time and stored in a designated data storage device for subsequent fault diagnosis and analysis.
[0024] A further improvement of the technical solution of the present invention is that in step 2, the data set division process includes:
[0025] Step 21, according to the vibration characteristics of the bearing and the sampling rate of the data, determine the size of the sliding window and the sliding step size, wherein the window size is 4096 data points, the distance of each sliding is 2048 data points, the window size should be large enough to include one or more cycles of the bearing vibration signal, and the sliding step size should be smaller than the window size to ensure overlap between windows and avoid information loss;
[0026] Step 22, using a sliding window to segment the data of each channel, starting from the starting position of the data, according to the determined window size and sliding step, sequentially intercepting data segments to form multiple window data, and segmenting each channel data of each type of health status respectively, dividing into 200 data segments, each segment corresponds to a window;
[0027] Step 23, storing the segmented window data in a data list, each window data contains its corresponding health status label and channel information for subsequent model training and testing;
[0028] Step 24: According to the requirements of model training and testing, the ratio of the training set to the test set is determined to be 7:3, that is, 70% of the data is used to train the model and 30% of the data is used to test the performance of the model. In each health status category, 70% of the window data is randomly selected as the training set, and the remaining 30% of the window data is used as the test set to ensure that the data of each category is representative in the training set and the test set to avoid the deviation of the model training and test results due to category imbalance. The result is that the number of training sets and test sets for each channel data of each health status is 140 and 60 respectively.
[0029] Step 25, store the divided training set and test set in different files, namely training set file and test set file, which contain the characteristics and label information of the data, so as to facilitate the subsequent model training and testing reading and use.
[0030] A further improvement of the technical solution of the present invention is that in step 3, the process of performing nonlinear transformation using MSE and ReLU activation functions includes:
[0031] Step 31, respectively load the data of the training set and the test set, including the features and label information of each window data, and set the parameters of MSE, wherein the scale parameter of MSE is determined to be 40, which means that when calculating the sample entropy, the original signal sequence is divided into 40 scale subsequences for analysis, and the template length is set to 1, indicating the number of signal points for each comparison, and the matching threshold is 0.2 to judge the similarity between signal points;
[0032] Step 32, for each window data in the training set and the test set, a coarsened sequence is generated, the original signal sequence is divided according to the scale parameter 40 to obtain multiple subsequences, each subsequence is composed of the average value of 40 adjacent signal points, thereby generating a coarsened signal sequence;
[0033] Step 33, based on the coarsened sequence, calculate the MSE. For each coarsened sequence, use the sample entropy algorithm with a template length of 1 to calculate the similarity between the signal points in the sequence. By comparing the distance between the signal points with the matching threshold, determine the number of similar signal point pairs, and then calculate the sample entropy value. The calculated MSE value is used as a feature vector to extract features that describe signal complexity and uncertainty.
[0034] Step 34, storing the calculated MSE features together with the label information of the original data, obtaining an entropy training set and an entropy test set for subsequent processing;
[0035] Step 35, use He initialization to set the model weights according to the number of input units To initialize the weights, calculate the variance value initialized by He, which is , and according to the calculated variance value, randomly generate the weight parameters of the model from the normal distribution or uniform distribution to complete the initialization of the model parameters;
[0036] Step 36, design a neural network KAN that displays the parameterized equation, set the convolution kernel size, number and step size parameters, input the processed training set data into the convolution layer, extract the features of the local area of the input data through the convolution kernel, and obtain the feature map. In KAN, each layer is defined as , The values of are determined by the input and output corresponding to the layer;
[0037] Step 37, applying the ReLU activation function to the feature map output by the convolution layer for nonlinear transformation, and after being processed by the convolution layer and the ReLU activation function, outputting a feature map containing the extracted feature information.
[0038] A further improvement of the technical solution of the present invention is that in step 4, the process of multi-scale feature fusion includes:
[0039] Step 41, input the multi-channel feature map processed by the convolution layer and the nonlinear transformation (ReLU activation function) to the spatial pyramid pooling (SPP) layer. The multi-channel feature map contains local feature information of different channels, has different spatial dimensions and semantic information, and sets the pooling scale of the SPP layer according to the requirements of the model and the characteristics of the data. The pooling scale determines the number and size of the pooling operations performed by the SPP layer at different scales. Among them, three pooling scales can be selected, namely 1×1, 2×2 and 4×4, corresponding to global pooling, medium-scale pooling and local pooling, respectively;
[0040] Step 42, the SPP layer divides the feature map into multiple spatial regions of different sizes, performs a maximum pooling operation in each region, extracts the most significant features in the region, and captures spatial information of different scales through a multi-level pooling strategy;
[0041] Step 43, the pooled features of different scales are fused to obtain comprehensive multi-scale features, and the fused multi-scale features are vectorized to convert them into feature vectors of fixed length, wherein the feature vector contains multi-scale and multi-level spatial information, thereby enhancing the model's feature extraction capability for multi-channel heterogeneous data.
[0042] A further improvement of the technical solution of the present invention is that in step 5, the process of improving the residual module to perform feature processing includes:
[0043] Step 51, design the KARes module, improve the residual module through the Kolmogorov-Arnold representation theorem, and enhance feature extraction and information flow;
[0044] Step 52, in the residual module, according to the Kolmogorov-Arnold representation theorem, construct a corresponding KA layer module, including multiple unary function modules and combination modules, perform nonlinear transformation and combination on the feature map, and design multiple unary function modules, each module operates on a channel in the feature map, and the unary function module can use a simple nonlinear function, such as ReLU activation function, sigmoid activation function or tanh activation function, etc., to perform nonlinear transformation of the ReLU activation function on the feature channel to extract the feature information in the channel;
[0045] Step 53, designing a combination module, combining the outputs of multiple unary function modules by weighted summation, and then integrating the feature information of different channels to form a complex feature representation;
[0046] Step 54, based on the KA layer module, a residual connection is designed to add the input feature map and the output feature map of the KA layer module to form the final output of the residual module. The residual connection can effectively alleviate the gradient vanishing and gradient exploding problems, improve the training effect of the model, and use the KA layer module to replace the fully connected layer to perform nonlinear processing on the data;
[0047] Step 55, through the synergy of the KA layer module and the residual connection, the feature extraction and information flow are enhanced, and the gradient flow and nonlinear mapping capabilities of the model are optimized through the improved residual module, thereby improving the recognition accuracy of the model on the imbalanced data set.
[0048] A further improvement of the technical solution of the present invention is that the design of the KARes module includes the following steps:
[0049] Step 511, obtaining input features: the input features correspond to the obtained multi-channel fusion features, and the fusion features capture key information of the bearing health status by processing and fusing multi-channel signals;
[0050] Step 512, KA layer and layer normalization: The KA layer can effectively process complex data structures, has better expression ability, and improves the stability and convergence speed of model training. Layer normalization is added after the KA layer, and layer normalization is used to standardize the output of each layer, reduce internal covariance offset, and improve training effect;
[0051] Step 513, fully connected layer and normalized layer processing: as another branch in the residual module, the input feature is a multi-channel fusion feature, and the linear features in the original data are extracted through the fully connected layer and the normalized layer;
[0052] Step 514, fully connected layer and normalized layer processing: using the fully connected layer and normalized layer processing to further abstract and optimize the input features, so that the network can learn more complex health status features;
[0053] Step 515, feature addition: applying the conventional residual module operation to feature addition. The residual module effectively avoids gradient vanishing and explosion by adding the input features to the processed output features, so that the model can be trained deeper and faster.
[0054] Step 516, output features: the step of outputting features for KA layer processing.
[0055] A further improvement of the technical solution of the present invention is that in step 6, the process of measuring classification errors includes:
[0056] Step 61, using the softmax function to calculate the probability of each category, outputting the fault diagnosis result, and determining whether the bearing is in a normal or faulty state;
[0057] Step 62, input the input data (feature map after multi-scale pooling and improved residual module processing) to the last layer of the model, obtain the predicted probability of each category, and obtain the corresponding true label from the data set, and use the cross entropy loss function to measure the error between the model prediction value and the true label, where the true label is a one-hot encoded vector indicating the category to which the sample belongs;
[0058] Step 63, determine whether the model iteration has reached a preset number of times. If so, save the trained model parameters to ensure that the model can perform fault diagnosis tasks. Otherwise, update the parameters until the preset number of training times is reached, wherein back propagation and gradient descent are used to update the model parameters to reduce losses and improve model performance.
[0059] A further improvement of the technical solution of the present invention is that the parameters of the model are updated by back propagation and gradient descent, and the specific process is as follows:
[0060] Step 631, starting from the calculated cross entropy loss value, initialize the back propagation process and calculate the gradient of the output layer (softmax layer), where the gradient is the partial derivative of the loss function with respect to the output layer weights and biases;
[0061] Step 632, starting from the output layer, calculate the gradient of each layer forward layer by layer. For each layer, use the chain rule to pass the gradient of the current layer to the previous layer, and calculate the gradient of the previous layer until the input layer is calculated. In the process of calculating the gradient, accumulate the gradient value of each parameter to prepare for the subsequent parameter update;
[0062] Step 633, based on the requirements of stochastic gradient descent, set the learning rate, and use the calculated gradient value and the set learning rate to update the parameters (weights and biases) of the model layer by layer according to the update rule of the optimization algorithm. The learning rate determines the step size of each parameter update. A learning rate that is too large may cause unstable training, and a learning rate that is too small may cause slow convergence.
[0063] Step 634, in each training round, repeat the steps of calculating the loss, back-propagating the gradient, and updating the model parameters until a preset number of training times is reached;
[0064] Step 635, after each training round, use the test set to evaluate the performance of the model, calculate the loss value and accuracy on the test set, monitor the training effect of the model, and after the training process is completed, save the trained model parameters for use in practical applications.
[0065] A further improvement of the technical solution of the present invention is that in step 7, the process of outputting the diagnosis result includes:
[0066] Step 71, select a model saving format according to the deep learning framework used, and use the saving function provided by the deep learning framework to save the parameters and structure of the model to a file;
[0067] Step 72, after the training is completed, use the test set to comprehensively evaluate the model, calculate the accuracy, precision, recall rate and F1 score of the model on the test set, evaluate the classification performance and fault diagnosis ability of the model, and conduct in-depth analysis of the evaluation results to understand the performance of the model in different categories, identify the advantages and disadvantages of the model, analyze the recognition accuracy of the model on minority class fault samples, and ensure that the model can effectively diagnose various fault types;
[0068] Step 73, in a scenario where fault diagnosis is required, load the saved model, and deploy the model to a server, edge device, or embedded system according to actual application requirements to perform fault diagnosis in real time;
[0069] Step 74, preprocess the data to be diagnosed to make it meet the input requirements of the model, input the preprocessed data into the model to perform fault diagnosis, the model outputs a prediction result indicating the fault type, and outputs a fault diagnosis report based on the prediction result of the model.
[0070] Due to the adoption of the above technical solution, the present invention has the following technical advances compared with the prior art:
[0071] 1. The present invention provides an intelligent bearing health monitoring method based on MSE-SPP-KAResnet. By integrating multi-scale feature extraction, KA optimization strategy and MSE module, the accuracy problem of traditional fault diagnosis methods under multi-channel data fusion and unbalanced data sets is effectively solved. Compared with the prior art, the spatial information of different scales is captured by the SPP layer, which improves the feature extraction capability of multi-channel heterogeneous data. The KA optimization residual module is used to enhance the gradient flow and nonlinear mapping capabilities, effectively improving the recognition accuracy of minority class samples. The introduced MSE module improves the robustness to signal complexity and uncertainty, and finally ensures the efficiency and accuracy of bearing fault diagnosis in an unbalanced data environment.
[0072] 2. The present invention provides an intelligent bearing health monitoring method based on MSE-SPP-KAResnet, introduces the residual module of the KA optimization strategy, optimizes the gradient flow and nonlinear mapping capabilities, and especially improves the performance of the model on unbalanced data sets. It can effectively alleviate the problem of bias towards majority class samples in traditional methods, and improve the recognition accuracy of minority class fault samples, thereby enhancing the overall accuracy of fault diagnosis.
[0073] 3. The present invention provides an intelligent bearing health monitoring method based on MSE-SPP-KAResnet. The MSE module is adopted to provide a robust method for the model to describe the complexity and uncertainty of the signal. By integrating MSE into SPP-KAResNet, the proposed model can extract entropy-based features to replace traditional time-frequency features, and ultimately improve its diagnostic performance in different imbalance scenarios.
[0074] 4. The present invention provides an intelligent bearing health monitoring method based on MSE-SPP-KAResnet, which adopts the SPP module to perform spatial feature fusion of multi-channel signals at multiple scales. The SPP module enhances the model's feature extraction capability for heterogeneous data, enabling the model to better capture complex fault modes, thereby improving diagnostic performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0077] Figure 1 It is a schematic diagram of the process flow of the intelligent bearing health monitoring method based on MSE-SPP-KAResnet of the present invention;
[0078] Figure 2 This is the design diagram of the KARes module of the present invention;
[0079] Figure 3 This is the structural diagram of KAN of the present invention;
[0080] Figure 4 It is the structural diagram of the MSE-SPP-KANResnet model of the present invention;
[0081] Figure 5 It is a normalized scale factor (NSF) parameter selection result diagram of the present invention;
[0082] Figure 6 The effects of different parameters in the KANlayer model of the present invention on the classification performance, (a) is the result diagram of dataset A when NSF is 20, (b) is the result diagram of dataset A when NSF is 30, (c) is the result diagram of dataset B when NSF is 20, (d) is the result diagram of dataset B when NSF is 30;
[0083] Figure 7 is the fault diagnosis accuracy of the proposed MSE-SPP-KANResnet under two datasets with different imbalance ratios (UR);
[0084] Figure 8 (a) is the confusion matrix of data set A when UR=1, (b) is the confusion matrix of data set B when UR=1, (c) is the confusion matrix of data set A when UR=10, and (d) is the confusion matrix of data set B when UR=10;
[0085] Fig. 9 Table 1;
[0086] Fig.10 Table 2;
[0087] Fig.11 Table 3. DETAILED DESCRIPTION
[0089] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0090] Embodiment 1, as Figures 1 to 8 As shown, the present invention provides a smart bearing health monitoring method based on MSE-SPP-KAResnet, comprising the following steps:
[0091] Step 1: Use multiple three-axis acceleration sensors to collect data from the equipment bearings, obtain multi-channel raw signal data, provide basic data support for subsequent fault diagnosis, clarify the bearings to be monitored and the equipment they are located in, and select three-axis acceleration sensors with appropriate range, frequency response range and accuracy according to the size, speed and load parameters of the bearings. The three-axis acceleration sensor can simultaneously measure the vibration acceleration in three mutually perpendicular directions to provide data support for comprehensive monitoring of the vibration state of the bearing. Install the three-axis acceleration sensor near the bearing seat or on the component directly connected to the bearing. The installation position should be as close to the bearing as possible to more accurately capture the vibration signal of the bearing. At the same time, avoid installing it on other components with large vibration to avoid introducing noise interference, and the output of the three-axis acceleration sensor should be The output signal line is connected to the input port of the data acquisition device. According to the working conditions and monitoring requirements of the bearing, the sampling rate, sampling time and sampling period parameters of the data acquisition device are set. The sampling rate is higher than twice the highest frequency component of the bearing vibration signal to avoid aliasing. The sampling time should be long enough to collect enough vibration signal data for subsequent analysis and processing, and then the sensor is started for data acquisition to capture the vibration signal of the bearing in real time and convert it into an electrical signal for transmission. Using multiple three-axis acceleration sensors, the original signal data of multiple channels are collected at the same time, including the radial vibration, axial vibration and tangential vibration direction of the bearing. The collected original signal data will be recorded in real time and stored in the specified data storage device for subsequent fault diagnosis and analysis;
[0092] Step 2: Use a sliding window to segment the data and divide the data set into a training set and a test set. The ratio of training set to test set is 7:3 to prepare for model training and testing. According to the vibration characteristics of the bearing and the sampling rate of the data, determine the size of the sliding window and the sliding step. The window size is 4096 data points, and the distance of each sliding is 2048 data points. The window size should be large enough to include one or more cycles of the bearing vibration signal. The sliding step should be smaller than the window size to ensure overlap between windows and avoid information loss. Use a sliding window to segment the data of each channel. Starting from the starting position of the data, according to the determined window size and sliding step, intercept data segments in turn to form multiple window data, and segment the data of each channel of each health state separately, dividing 200 data segments, each segment corresponds to a window, and store the segmented window data in In the data list, each window data contains its corresponding health status label and channel information for subsequent model training and testing. According to the requirements of model training and testing, the division ratio of training set and test set is determined to be 7:3, that is, 70% of the data is used to train the model, and 30% of the data is used to test the performance of the model. In each health status category, 70% of the window data is randomly selected as the training set, and the remaining 30% of the window data is used as the test set to ensure that the data of each category is representative in the training set and the test set to avoid the deviation of model training and test results due to category imbalance. The result is that the number of training sets and test sets for each channel data of each health status is 140 and 60 respectively. The divided training set and test set are stored in different files, divided into training set files and test set files. The files contain the characteristics and label information of the data, which is convenient for subsequent model training and testing to read and use;
[0093] Step 3: Use MSE to process the training set and the test set separately. The scale of MSE is 40, the template length is 1, and the matching threshold is 0.2. Features describing the complexity and uncertainty of the signal are extracted, and the model parameters are initialized. Features are extracted through the convolution layer, and nonlinear transformation is performed using the ReLu activation function. The data of the training set and the test set are loaded separately, including the features and label information of each window data, and the parameters of MSE are set. Among them, determining the scale parameter of MSE to be 40 means that when calculating the sample entropy, the original signal sequence is divided into 40 scale subsequences for analysis. The template length is set to 1, indicating the number of signal points for each comparison, and the matching threshold is 0.2 to judge the similarity between signal points. For each window data in the training set and the test set, a coarsened sequence is generated. The original signal sequence is divided according to the scale parameter 40 to obtain multiple subsequences, each of which consists of adjacent subsequences. The average value of 40 signal points is used to generate a coarsened signal sequence. On the basis of the coarsened sequence, the MSE is calculated. For each coarsened sequence, the sample entropy algorithm with a template length of 1 is used to calculate the similarity between the signal points in the sequence. By comparing the distance between the signal points with the matching threshold, the number of similar signal point pairs is determined, and then the sample entropy value is calculated. The calculated MSE value is used as a feature vector to extract features that describe the complexity and uncertainty of the signal. The feature vector can reflect the complexity and uncertainty of the signal at different scales. The larger the sample entropy value, the higher the complexity and uncertainty of the signal, which provides important feature information for subsequent model training and fault diagnosis. The calculated MSE features are stored together with the label information of the original data, and the entropy training set and the entropy test set are obtained for subsequent processing. The model weights are set using He initialization according to the number of input units. To initialize the weights, calculate the variance value initialized by He, which is , and according to the calculated variance value, randomly generate the weight parameters of the model from the normal distribution or uniform distribution, complete the initialization of the model parameters, design a neural network KAN that displays the parameterized equation, set the convolution kernel size, number and step parameters, input the processed training set data into the convolution layer, and extract the features of the local area of the input data through the convolution kernel to obtain the feature map. In KAN, each layer is defined as , The values of are determined by the input and output corresponding to the layer. The ReLU activation function is applied to the feature map output by the convolution layer for nonlinear transformation. After being processed by the convolution layer and the ReLU activation function, the feature map containing the extracted feature information is output, which provides a basis for subsequent model training and fault diagnosis. Among them, the ReLU activation function can set the negative values in the feature map to zero and retain the positive values, thereby increasing the nonlinear ability of the model. It also has the advantages of simple calculation and fast training speed.
[0094] The structure of KAN is as follows Figure 3 As shown, Figure 3 There are n features as input in the input layer. The input layer first introduces these features into the bottom layer, which contains the parameter matrix , used to calculate the intermediate feature representation, the bottom layer generates the intermediate representation by applying the spline function to the input features, and then passes it to the top layer, which contains the parameter matrix , responsible for processing the intermediate representation obtained from the bottom layer and then calculating the final output;
[0095] In order to avoid excessive data volume in the multi-channel data processing process, MSE is used to process the data of each channel appropriately. The data of a single channel is , MSE resamples and generates signal sequences of different scales by setting NSF, and MSE finally combines the entropy values of each scale into a feature vector , used to characterize the randomness and complexity of signals at different scales;
[0096] The dimension of a single-channel input signal sample is , 64 is the batch size, 2048 is the signal data length, and MSE further extracts the input feature map into Input data, for the number of channels After performing MSE processing on the multi-channel data one by one, we can get The feature map is used for subsequent processing. Represents the number of time series channels (such as the number of sensors), and NSF represents the number of MSE features extracted on each channel;
[0097] Step 4: Multi-scale feature fusion of multi-channel data is performed through the spatial pyramid pooling layer to capture spatial information of different scales and enhance the model's feature extraction capability for multi-channel heterogeneous data.
[0098] Step 5, design the KARes module and construct the corresponding KA layer module through the Kolmogorov-Arnold representation theorem to improve the residual module for feature processing;
[0099] Step 6: Use the cross entropy loss function to measure the classification error and update the model parameters through back propagation and gradient descent.
[0100] Step 7: After the training process is completed, the fault diagnosis capability of the model is saved and completed, and the diagnosis results are output;
[0101] The specific summary is as follows: vibration signals are collected from bearings in different health states through multi-channel sensors, and the collected data are divided into windows to construct training and test data sets. In order to enhance the representativeness of the data, the unbalanced ratio is applied to reconstruct the training and test data so that it can better reflect the data distribution under different fault conditions; the training data with known states are then input into the MSE-SPP-KAResNet model for processing and training. First, the MSE method is used to preprocess the data. In the training stage, SPP-KAResNet processes multi-channel data through layer-by-layer convolution and SPP layers, extracts features and performs dimensionality reduction, thereby effectively capturing fault features. The KAres layer further emphasizes relevant features, which helps to deeply understand the fault type and its severity; the test data with unknown states is then input into the trained MSE-SPP-KAResNet model for health monitoring and fault diagnosis. The health status of the bearing is diagnosed based on the intelligent diagnosis method of MSE-SPP-KAResNet, and the results under different models and parameter settings are compared. Through ablation experiments and accuracy evaluation (such as Polito The accuracy of UR=10 reaches 95%) to verify the effectiveness and robustness of the model.
[0102] Embodiment 2, as Figures 1 to 8 As shown, based on Example 1, the present invention provides a technical solution: Preferably, in step 4, the process of multi-scale feature fusion includes:
[0103] The multi-channel feature map after the convolution layer and nonlinear transformation is input to the SPP layer. The multi-channel feature map contains local feature information of different channels, with different spatial dimensions and semantic information. According to the needs of the model and the characteristics of the data, the pooling scale of the spatial pyramid pooling layer is set. The pooling scale determines the number and size of the pooling operations performed by the SPP layer at different scales. Among them, three pooling scales can be selected, namely 1×1, 2×2 and 4×4, corresponding to global pooling, medium-scale pooling and local pooling, respectively. The SPP layer divides the feature map into multiple spatial regions of different sizes, performs the maximum pooling operation in each region, extracts the most significant features in the region, captures spatial information of different scales through a multi-level pooling strategy, fuses the pooling features of different scales, obtains comprehensive multi-scale features, and vectorizes the fused multi-scale features to convert them into fixed-length feature vectors. Among them, the feature vector contains multi-scale and multi-level spatial information, which enhances the model's feature extraction capability for multi-channel heterogeneous data.
[0104] Furthermore, the basic residual module helps prevent information loss during transmission. Its greatest feature is that when the input and output dimensions are different, linear mapping can be performed through the residual path to ensure the effectiveness of information transmission. Combining the KAN Layer on the residual path can further enhance the fluidity of information. In particular, the nonlinearity of the KAN Layer allows the network to capture and transmit more complex relationships between different features.
[0105] Furthermore, the main branch of the KARes module consists of a fully connected layer, a batch normalization layer, and a ReLu activation function. The process is expressed by the following formula:
[0106] ;
[0107] For the residual path, when the channels of the input and output are not equal, the input is adjusted to a dimension that matches the output through KANLayer. This process is expressed by the following formula:
[0108]
[0109] Eventually and Add them together to get the output of the KARes module;
[0110] Mapped in the form of the following formula:
[0111]
[0112] is a multivariate continuous function, Represents the outer function, Represents the outer function;
[0113] In addition, in the fusion of multi-channel signal features, the SPP layer provides an effective way to integrate the spatial features of multiple channels, so that the network can extract and integrate the feature information of multiple channels at different spatial scales. Specifically, the SPP layer performs pooling operations on the feature maps of each channel through pooling windows of multiple scales, thereby capturing information at global and local scales. For input feature dimensions of For the multi-channel feature map of the image, SPP can pool each channel at multiple scales separately and integrate the pooling results of these scales together to form a multi-scale feature representation containing global and local spatial information;
[0114] SPP can be Multi-scale pooling is performed on two dimensions, NSF and , to fuse feature information. First, use The pooling window is used for global pooling. Extract a feature value from the dimension, such as the formula As shown, this eigenvalue represents the overall complexity of all temporal channels and multi-scale entropy features, which is used to capture global information;
[0115] SPP then uses The pooling window of The space of is divided into regions. Specifically, the pooling operation will The timing channels are divided into two groups (i.e., each group channels), similarly The entropy features of each sample are divided into two regions. After pooling operation in each region, 4 eigenvalues can be obtained, which represent the local information of each region. In this way, the model can capture the medium-scale relationship between different channel groups and different entropy feature regions. In addition, SPP fine-grained local pooling (4×4 pooling) can extract very local features under this fine-grained division, revealing the entropy features of each group of time series channels in different small-scale regions.
[0116] In step 5, the process of improving the residual module for feature processing includes:
[0117] The KARes module is designed to improve the residual module through the Kolmogorov-Arnold representation theorem, enhance feature extraction and information flow. The design of the KARes module is as follows:
[0118] Get input features: The input features correspond to the obtained multi-channel fusion features. The fusion features capture the key information of the bearing health status through the processing and fusion of multi-channel signals. KA layer and layer normalization: The KA layer can effectively process complex data structures, has better expression ability, and improves the stability and convergence speed of model training. Layer normalization is added after the KA layer. The output of each layer is standardized by layer normalization to reduce internal covariance offset and improve training effect. Fully connected layer and normalization layer processing: As another branch in the residual module, the input feature is multi-channel fusion feature. The linear features in the original data are extracted through the fully connected layer and normalization layer. Fully connected layer and normalization layer processing: Use the fully connected layer and normalization layer processing to further abstract and optimize the input features, so that the network can learn more complex health status features. Feature addition: Apply conventional residual module operations to feature addition. The residual module effectively avoids gradient vanishing and explosion by adding the input features to the processed output features, so that the model can be trained deeper and faster. Output features: The step of outputting features for KA layer processing
[0119] In the residual module, according to the Kolmogorov-Arnold representation theorem, the corresponding KA layer module is constructed, including multiple unary function modules and combination modules, and the feature map is subjected to nonlinear transformation and combination of the ReLU activation function. Multiple unary function modules are designed, and each module operates on a channel in the feature map. The unary function module can use simple nonlinear functions, such as ReLU activation function, sigmoid activation function or tanh activation function, to perform nonlinear transformation on the feature channel and extract the feature information in the channel. A combination module is designed to combine the outputs of multiple unary function modules in a weighted summation manner, and then integrate the feature information of different channels to form a complex feature representation. On the basis of the KA layer module, a residual connection is designed to connect the input The input feature map is added to the output feature map of the KA layer module to form the final output of the residual module. The residual connection can effectively alleviate the gradient vanishing and gradient exploding problems, improve the training effect of the model, and use the KA layer module to replace the fully connected layer to perform nonlinear processing on the data. The synergy of the KA layer module and the residual connection can enhance feature extraction and information flow. The KA layer module can capture the complex patterns and nonlinear relationships in the feature map, and the residual connection can retain the important information in the input feature map, so that the model can better learn and represent the features in the data. The improved residual module optimizes the gradient flow and nonlinear mapping capabilities of the model, improves the recognition accuracy of the model on unbalanced data sets, and enhances the feature learning and recognition capabilities of minority fault samples, thereby improving the overall fault diagnosis performance;
[0120] The overall network framework of MSE-SPP-KAResNet is as follows Figure 4 As shown in the figure, MSE-SPP-KAResNet is mainly composed of MSE and SPP-KAResNet. MSE is a data preprocessing module that processes signals from the perspective of multi-scale and entropy to capture the complexity and randomness of data at different scales. SPP-KAResNet is the backbone structure of MSE-SPP-KAResNet, which uses components such as convolutional layers, SPP layers, KARes modules, and KAN layers to fuse channel features, perform nonlinear mapping, and finally achieve decision-making. The complete structure of the model is shown in the figure. Figure 4 As shown in Figure 2, the information of each layer parameter is as follows: Fig. 9 As shown in Table 1;
[0121] In step 6, the process of measuring classification error includes:
[0122] Use the softmax function to calculate the probability of each category, output the fault diagnosis result, and decide whether the bearing is in a normal or faulty state. Input the input data (feature map after multi-scale pooling and improved residual module processing) to the last layer of the model to obtain the predicted probability of each category, and obtain the corresponding true label from the data set. Use the cross entropy loss function to measure the error between the model prediction value and the true label, where the true label is a one-hot encoded vector indicating the category to which the sample belongs. Determine whether the model iteration has reached the preset number of times. If so, save the trained model parameters to ensure that the model can perform fault diagnosis tasks. Otherwise, update the parameters until the preset number of training times is reached. Use back propagation and gradient descent to update the model parameters to reduce losses and improve model performance.
[0123] Among them, the expression of the cross entropy loss function is:
[0124] ;
[0125] In the formula, is the true label, is the probability predicted by the model, is the number of categories. The cross entropy loss measures the difference between the probability distribution predicted by the model and the true distribution. The smaller the loss value, the more accurate the model prediction.
[0126] Furthermore, back propagation and gradient descent are used to update the parameters of the model. The specific process is as follows:
[0127] Starting from the calculated cross entropy loss value, initialize the back propagation process and calculate the gradient of the output layer (softmax layer). The gradient is the partial derivative of the loss function with respect to the weights and biases of the output layer. Starting from the output layer, calculate the gradient of each layer forward layer by layer. For each layer, use the chain rule to pass the gradient of the current layer to the previous layer and calculate the gradient of the previous layer until the input layer is calculated. In the process of calculating the gradient, accumulate the gradient value of each parameter to prepare for the subsequent parameter update. Based on the requirements of stochastic gradient descent, set the learning rate, and use the calculated gradient value and the set learning rate. According to the update rule of the optimization algorithm, update the parameters (weights and biases) of the model layer by layer. The learning rate determines the step size of each parameter update. A learning rate that is too large may cause unstable training, and a learning rate that is too small may cause slow convergence. In each training round, repeat the steps of calculating loss, back propagating gradients, and updating model parameters until the preset number of training times is reached. After each training round, use the test set to evaluate the performance of the model, calculate the loss value and accuracy on the test set, monitor the training effect of the model, and save the trained model parameters after the training process for use in practical applications.
[0128] In step 7, the process of outputting the diagnosis result includes:
[0129] According to the deep learning framework used, select the model saving format, use the save function provided by the deep learning framework to save the model parameters and structure to the file, after the training is completed, use the test set to comprehensively evaluate the model, calculate the accuracy, precision, recall rate and F1 score of the model on the test set, evaluate the classification performance and fault diagnosis ability of the model, and conduct in-depth analysis of the evaluation results to understand the performance of the model in different categories, identify the advantages and disadvantages of the model, analyze the recognition accuracy of the model on minority class fault samples, ensure that the model can effectively diagnose various fault types, load the saved model in the scenario where fault diagnosis is required, and deploy the model to the server, edge device or embedded system according to the actual application needs to perform fault diagnosis in real time, preprocess the data to be diagnosed to make it meet the input requirements of the model, input the preprocessed data into the model for fault diagnosis, and the model outputs the prediction result representing the fault type, and outputs the fault diagnosis report based on the prediction result of the model.
[0130] Embodiment 3, as Figures 1 to 8 As shown, on the basis of Example 1 and Example 2, the present invention provides a technical solution: preferably, relevant tests are carried out through the Southeast University test bench, and the test bench obtains bearing data and gear data through a transmission system dynamic simulator, and the data are collected under two different working conditions where the speed system load is set to 20HZ-0V or 30HZ-2V. At a sampling frequency of 5.12KHz, each state has a total of 8 channel signals, which are motor vibration signals; vibration signals in three directions of x, y and z of planetary gears; motor torque; vibration signals in three directions of x, y and z of reducer; the data set contains 4 bearing fault states and 1 healthy state, and the bearing fault states are rolling element cracks, inner ring cracks, outer ring cracks, and inner and outer ring cracks respectively;
[0131] In order to ensure that the length of each sample fully reflects the number of points contained in a complete rotation of the bearing, the data of each channel is divided into 200 samples using a sliding window method with a window size of 4096 and a sliding distance of 2048. Fig.10 As shown in Table 2:
[0132] Embodiment 4, as Figures 1 to 8As shown, on the basis of Example 1 and Example 2, the present invention provides a technical solution: preferably, the general overview of the bearing test bench of the Polytechnic University of Turin is shown, the test bench is mainly composed of a high-speed main shaft and a secondary shaft driven to rotate by the main shaft, the main shaft bearing is fixed on a high-rigidity single bracket located on a huge steel bottom plate, and the experimental data is mainly measured by an XYZ three-axis sensor to obtain different acceleration information in three spatial directions on the bearing bracket and the main shaft, specifically the acceleration information on the support of the faulty bearing B1 and the support of the bearing B2, the acceleration sensor is a three-axis IEPE type, the frequency three dimensions are 1-12000Hz (amplitude ±5%, phase ±10°), the nominal resonant frequency is 55kHz, and the nominal sensitivity is 1mV / ms-2. The roller bearing used in the experiment has an inner ring connected to a short hollow shaft specially designed for a maximum of 35000rpm;
[0133] In the experiment, data with a sampling frequency of 51.2KHZ, a shaft speed of 100HZ (6000rpm), and no-load conditions were used. The data set includes three health states: inner ring fault, roller fault, and normal state. According to the different fault severity, the inner ring and roller faults are each divided into three degrees. Therefore, the data set has a total of 7 states. Each fault type has 6 channels of acceleration information, and each channel has 512,000 data points. The division of the data set is consistent with the method of data set A. The specific content is as follows: Fig.11 As shown in Table 3:
[0134] Embodiment 5, as Figures 1 to 8 As shown, on the basis of Embodiment 1 to Embodiment 4, the present invention provides a technical solution: preferably, the algorithm parameters are the key factors for achieving good classification within the MSE-SPP-KAResnet framework, and the parameters are adjusted for different classification tasks;
[0135] The coarse-grained parameters of MSE, the number of grid points in KANlayer, and the order of piecewise polynomials are discussed. The detailed experimental results under different NSFs are shown in Figure 5As shown in the figure, from the overall trend, the classification accuracy of dataset A under different NSF values (8, 10, 20, 30, 40, 50) almost reaches 100%, indicating that for dataset A, the change of NSF parameters has very little effect on the classification performance of the model. In contrast, the classification accuracy of dataset B shows obvious fluctuations with the change of NSF parameters. When the NSF value is low (such as 8), the accuracy is 94.05%, which is relatively low; as the NSF value increases, the classification accuracy gradually increases, reaching 99.76% at NSF=20, and then slowly decreases with the increase of NSF; This shows that dataset B is more sensitive to the NSF parameter, and a higher NSF value helps to improve its classification accuracy. It can be inferred that the NSF value may be related to the granularity of feature extraction. A higher NSF can capture richer feature information, thereby improving the classification effect;
[0136] Figure 6 The effect of different parameters in the KANlayer model on the classification performance is shown. The four sub-figures correspond to different experimental conditions, marked as (a), (b), (c) and (d) respectively: Figure 6 In (a), the experiment is conducted on dataset A, and the NSF parameter is set to 20. The bar graph shows the classification accuracy of the model under different parameter combinations, where the height of the bar represents the classification accuracy. It can be seen from the figure that different parameter combinations have a great impact on the classification effect, especially under specific parameter combinations, the accuracy can be significantly improved; Figure 6 (b) shows the classification performance of dataset A when the NSF parameter is set to 20, so as to observe the consistency of the same dataset in different experiments or the difference in effect after slight parameter adjustment; Figure 6 (c) and (d) show the classification performance of the B dataset with NSF parameters of 20 and 30. The results of Figure 6(c) show that the classification effects of the B dataset under different parameter combinations are obviously different. The fluctuation of the column height in the histogram shows that the adjustment of the parameters has a greater impact on the classification accuracy, while Figure 6 In (d), the performance of different parameter combinations is further verified. The classification accuracy change trend of dataset B is similar to that of dataset A. Although the datasets are different, the fluctuation range of the classification effect is still obvious. These results show that the parameters in KANlayer have a significant impact on the experimental results, especially in datasets A and B. Therefore, in order to obtain the best classification performance, the grid search method is used throughout the experiment to tune the KANlayer parameters to ensure the best classification effect on different datasets.
[0137] Embodiment 6, as Figures 1 to 8As shown, based on Examples 1 to 5, the present invention provides a technical solution: Preferably, in order to verify the effectiveness of the proposed method, three different URs are set for the two sets of data sets, specifically UR=1, UR=5, and UR=10. The results are as follows Figure 7 As shown by Figure 7 It can be seen that MSE-SPP-KAResNet achieved an accuracy of more than 99.5% and 95% on these two datasets, respectively, which fully demonstrated the ability of this method to effectively distinguish bearing samples with different health conditions under unbalanced conditions;
[0138] In order to further analyze the classification of each category of data under different imbalance ratios, the confusion matrix under UR=1 and UR=10 is added, as shown in Figure 8 As shown, when UR=1 ( Figure 8 (a) and Figure 8 (b)), the classification accuracy of the model for all categories is close to 100%, when UR=10 ( Figure 8 (c) and Figure 8 (d)), although the accuracy of B1 and B3 categories has decreased, the overall performance is still stable, and the accuracy of most categories (such as B0, B2, B4, B5, B6) is close to 100%, which shows that the performance of the model under imbalanced data has limited decline and still has strong classification ability and adaptability. In general, the model still shows high robustness under imbalanced data and has practical application value.
[0139] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. The intelligent bearing health monitoring method based on MSE-SPP-KAResnet is characterized by: The following steps are involved: Step 1: Use multiple three-axis acceleration sensors to collect data at the bearing of the equipment to obtain multi-channel raw signal data; Step 2: Use a sliding window to split the data and divide the data set into a training set and a test set; Step 3: Use multi-scale sample entropy to process the training set and test set separately, initialize the model parameters, extract features through the convolution layer, and use the ReLu activation function for nonlinear transformation; Step 4: Multi-scale feature fusion of multi-channel data is performed through the spatial pyramid pooling layer to capture spatial information of different scales; Step 5, design the KARes module and construct the corresponding KA layer module through the Kolmogorov-Arnold representation theorem to improve the residual module for feature processing; Step 6: Use the cross entropy loss function to measure the classification error and update the model parameters through back propagation and gradient descent; Step 7: After the training process is completed, the fault diagnosis capability of the model is saved and completed, and the diagnosis results are output.
2. The intelligent bearing health monitoring method based on MSE-SPP-KAResnet according to claim 1 is characterized in that: In step 1, the process of obtaining multi-channel original signal data includes: Step 11, identify the bearings to be monitored and the equipment they are located in, and select a three-axis acceleration sensor with appropriate range, frequency response range and accuracy based on the size, speed and load parameters of the bearing; Step 12, installing a three-axis acceleration sensor near the bearing seat or on a component directly connected to the bearing, and connecting an output signal line of the three-axis acceleration sensor to an input port of a data acquisition device; Step 13, according to the working conditions and monitoring requirements of the bearing, set the sampling rate, sampling time and sampling cycle parameters of the data acquisition device, and then start the sensor to collect data, capture the vibration signal of the bearing in real time, and convert it into an electrical signal for transmission; Step 14, using multiple three-axis acceleration sensors, simultaneously collects raw signal data of multiple channels, including signals in the radial vibration, axial vibration and tangential vibration directions of the bearing. The collected raw signal data will be recorded in real time and stored in a designated data storage device.
3. The smart bearing health monitoring method based on MSE-SPP-KAResnet according to claim 2 is characterized in that: In step 2, the data set division process includes: Step 21, determining the size of the sliding window and the sliding step length according to the vibration characteristics of the bearing and the sampling rate of the data, wherein the window size is 4096 data points and the distance of each sliding is 2048 data points; Step 22, using a sliding window to segment the data of each channel, starting from the starting position of the data, according to the determined window size and sliding step, sequentially intercepting data segments to form multiple window data, and segmenting each channel data of each type of health status respectively, dividing into 200 data segments, each segment corresponds to a window; Step 23, storing the segmented window data in a data list, each window data includes its corresponding health status label and channel information; Step 24, according to the requirements of model training and testing, determine the division ratio of the training set to the test set to be 7:
3. In each health status category, randomly select 70% of the window data as the training set, and the remaining 30% of the window data as the test set. The result is that the number of training sets and test sets for each channel data of each health status is 140 and 60 respectively. Step 25, store the divided training set and test set in different files, namely, training set files and test set files, and the files contain data features and label information.
4. The intelligent bearing health monitoring method based on MSE-SPP-KAResnet according to claim 3 is characterized in that: In step 3, the process of using the ReLu activation function to perform nonlinear transformation includes: Step 31, respectively load the data of the training set and the test set, including the features and label information of each window data, and set the parameters of the multi-scale sample entropy, wherein the scale parameter of the multi-scale sample entropy is determined to be 40, the template length is set to 1, and the matching threshold is set to 0.2; Step 32, for each window data in the training set and the test set, a coarsened sequence is generated, and the original signal sequence is divided according to the scale parameter 40 to obtain multiple subsequences; Step 33, based on the coarsened sequence, multi-scale sample entropy is calculated. For each coarsened sequence, a sample entropy algorithm with a template length of 1 is used to calculate the similarity between signal points in the sequence. The number of similar signal point pairs is determined by comparing the distance between the signal points with the matching threshold, and then the sample entropy value is calculated. The calculated multi-scale sample entropy value is used as a feature vector to extract features describing signal complexity and uncertainty. Step 34, storing the calculated multi-scale sample entropy features together with the label information of the original data to obtain an entropy training set and an entropy test set; Step 35, use He initialization to set the model weights according to the number of input units To initialize the weights, calculate the variance value initialized by He, which is , and according to the calculated variance value, randomly generate the weight parameters of the model from the normal distribution or uniform distribution to complete the initialization of the model parameters; Step 36, design a neural network KAN that displays the parameterized equation, set the convolution kernel size, number and step size parameters, input the processed training set data into the convolution layer, extract the features of the local area of the input data through the convolution kernel, and obtain the feature map. In KAN, each layer is defined as , The values of are determined by the input and output corresponding to the layer; Step 37, applying the ReLU activation function to the feature map output by the convolution layer for nonlinear transformation, and after being processed by the convolution layer and the ReLU activation function, outputting a feature map containing the extracted feature information.
5. The smart bearing health monitoring method based on MSE-SPP-KAResnet according to claim 4 is characterized in that: In step 4, the process of multi-scale feature fusion includes: Step 41, inputting the multi-channel feature map processed by the convolution layer and the nonlinear transformation into the spatial pyramid pooling layer, and setting the pooling scale of the spatial pyramid pooling layer according to the requirements of the model and the characteristics of the data, the multi-channel feature map contains local feature information of different channels, has different spatial dimensions and semantic information, and sets the pooling scale of the spatial pyramid pooling layer according to the requirements of the model and the characteristics of the data, and the pooling scale determines the number and size of the pooling operations performed by the SPP layer at different scales, wherein three pooling scales can be selected, namely 1×1, 2×2 and 4×4, corresponding to global pooling, medium-scale pooling and local pooling, respectively; Step 42, the spatial pyramid pooling layer divides the feature map into multiple spatial regions of different sizes, performs a maximum pooling operation in each region, extracts the most significant features in the region, and captures spatial information of different scales through a multi-level pooling strategy; In step 43, the pooled features of different scales are fused to obtain comprehensive multi-scale features, and the fused multi-scale features are vectorized to be converted into feature vectors of fixed length.
6. The smart bearing health monitoring method based on MSE-SPP-KAResnet according to claim 5 is characterized in that: In step 5, the process of improving the residual module to perform feature processing includes: Step 51, design the KARes module, improve the residual module through the Kolmogorov-Arnold representation theorem, and enhance feature extraction and information flow; Step 52, in the residual module, according to the Kolmogorov-Arnold representation theorem, construct a corresponding KA layer module, including multiple unary function modules and combination modules, and perform nonlinear transformation and combination of the ReLU activation function on the feature map; Step 53, designing a combination module, combining the outputs of multiple unary function modules by weighted summation, and then integrating the feature information of different channels to form a complex feature representation; Step 54, based on the KA layer module, a residual connection is designed to add the input feature map and the output feature map of the KA layer module to form the final output of the residual module, and the KA layer module is used to replace the fully connected layer to perform nonlinear processing on the data; Step 55, through the synergy of the KA layer module and the residual connection, the feature extraction and information flow are enhanced, and the gradient flow and nonlinear mapping capabilities of the model are optimized through the improved residual module, thereby improving the recognition accuracy of the model on the imbalanced data set.
7. The smart bearing health monitoring method based on MSE-SPP-KAResnet according to claim 6 is characterized in that: The design of the KARes module includes the following steps: Step 511, obtaining input features: the input features correspond to the obtained multi-channel fusion features, and the fusion features capture key information of the bearing health status by processing and fusing multi-channel signals; Step 512, KA layer and layer normalization: layer normalization is added after the KA layer, and the output of each layer is standardized by layer normalization to reduce the internal covariance offset; Step 513, fully connected layer and normalized layer processing: as another branch in the residual module, the input feature is a multi-channel fusion feature, and the linear features in the original data are extracted through the fully connected layer and the normalized layer; Step 514, fully connected layer and normalized layer processing: using the fully connected layer and normalized layer processing to further abstract and optimize the input features; Step 515, feature addition: applying conventional residual module operation to feature addition; Step 516, output features: the step of outputting features for KA layer processing.
8. The smart bearing health monitoring method based on MSE-SPP-KAResnet according to claim 7 is characterized in that: In step 6, the process of measuring classification errors includes: Step 61, using the softmax function to calculate the probability of each category, outputting the fault diagnosis result, and determining whether the bearing is in a normal or faulty state; Step 62, input the input data to the last layer of the model, obtain the predicted probability of each category, and obtain the corresponding true label from the data set, and use the cross entropy loss function to measure the error between the model prediction value and the true label; Step 63, determine whether the model iteration reaches a preset number of times, if so, save the trained model parameters, if not, update the parameters until the preset number of training times is reached, wherein back propagation and gradient descent are used to update the model parameters.
9. The smart bearing health monitoring method based on MSE-SPP-KAResnet according to claim 8 is characterized in that: The specific process of using back propagation and gradient descent to update the parameters of the model is as follows: Step 631, starting from the calculated cross entropy loss value, initialize the back propagation process and calculate the gradient of the output layer; Step 632, starting from the output layer, calculate the gradient of each layer forward layer by layer. For each layer, use the chain rule to pass the gradient of the current layer to the previous layer, and calculate the gradient of the previous layer until the input layer is calculated. In the process of calculating the gradient, the gradient value of each parameter is accumulated; Step 633, based on the requirements of stochastic gradient descent, set the learning rate, and use the calculated gradient value and the set learning rate to update the parameters of the model layer by layer according to the update rule of the optimization algorithm; Step 634, in each training round, repeat the steps of calculating the loss, back-propagating the gradient, and updating the model parameters until a preset number of training times is reached; Step 635, after each training round, use the test set to evaluate the performance of the model, calculate the loss value and accuracy on the test set, monitor the training effect of the model, and save the trained model parameters after the training process is completed.
10. The intelligent bearing health monitoring method based on MSE-SPP-KAResnet according to claim 9 is characterized in that: In step 7, the process of outputting the diagnosis result includes: Step 71, select a model saving format according to the deep learning framework used, and use the saving function provided by the deep learning framework to save the parameters and structure of the model to a file; Step 72, after the training is completed, use the test set to comprehensively evaluate the model, calculate the accuracy, precision, recall and F1 score indicators of the model on the test set, evaluate the classification performance and fault diagnosis ability of the model, and conduct in-depth analysis of the evaluation results to identify the advantages and disadvantages of the model and analyze the recognition accuracy of the model on minority class fault samples; Step 73, in a scenario where fault diagnosis is required, load the saved model, and deploy the model to a server, edge device, or embedded system according to actual application requirements to perform fault diagnosis in real time; Step 74, preprocess the data to be diagnosed to make it meet the input requirements of the model, input the preprocessed data into the model to perform fault diagnosis, the model outputs a prediction result indicating the fault type, and outputs a fault diagnosis report based on the prediction result of the model.
Citation Information
Cited By
Intelligent power grid-oriented underground power transmission cable fault position discrimination method and system
CN121253993A