Contact ball bearing loss degree identification method and system
Through multimodal data fusion technology, the feature fusion network of image and voiceprint signals and the improvement of Transformer model are solved, and a high-accurate loss degree identification and grading result output are achieved.
Patent Information
- Application Number
- CN202510421095.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, single mode data is difficult to fully reflect the loss status of contact ball bearings, resulting in inaccurate identification results.
Multimodal data fusion technology is adopted to acquire bearing surface image sequences and array microphones through industrial cameras to collect voiceprint signals, build a multimodal feature fusion network, use a cross-modal attention mechanism to fuse image features and voiceprint features, and build a loss assessment model based on improved Transformer.
It realizes a comprehensive description of the operating status of contact ball bearings, improves the accuracy of identification of loss degree, outputs the graded results of bearing loss degree, and supports equipment maintenance and management.
Smart Images

Figure CN120217207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mechanical equipment fault diagnosis, and specifically relates to a method and system for identifying the wear degree of a ball bearing. Background Art
[0002] As a key component in mechanical equipment, the wear degree of a ball bearing directly affects the operating performance and service life of the equipment. Traditional methods for identifying the wear degree of bearings mainly rely on vibration signal analysis. However, single-modal data often fails to comprehensively reflect the wear state of the bearing. For example, vibration signals may be affected by various factors such as environmental noise and equipment structure, resulting in inaccurate identification results. In recent years, with the development of sensor technology and artificial intelligence, multi-modal data fusion technology has gradually become a research hotspot. By combining different types of sensor data, such as images and sounds, the operating state of the bearing can be more comprehensively described, thereby improving the accuracy of wear degree identification. However, how to effectively fuse multi-modal data and construct an efficient identification model remains a technical problem currently faced. Summary of the Invention
[0003] Aiming at the above-mentioned technical deficiencies, the purpose of the present invention is to provide a method and system for identifying the wear degree of a ball bearing, so as to solve the problem that single-modal data in the prior art is difficult to comprehensively reflect the wear state of the bearing and the identification result is inaccurate.
[0004] To solve the above technical problems, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a method for identifying the wear degree of a ball bearing, and the method includes: Step S100: Synchronously collect multi-modal data during the operation of the ball bearing, where the multi-modal data includes a sequence of bearing surface images obtained by an industrial camera and a voiceprint signal collected by an array microphone; Step S200: Perform adaptive illumination compensation and denoising processing on the sequence of bearing surface images, and perform frame division and windowing processing on the voiceprint signal and convert it into a Mel spectrogram; Step S300: Based on the denoised sequence of bearing surface images and the Mel spectrogram, construct a multi-modal feature fusion network, extract image features and voiceprint features, and fuse the image features and voiceprint features using a cross-modal attention mechanism; Step S400: Construct a wear assessment model based on an improved Transformer to process spatio-temporal features, where the spatio-temporal features are the fused image features and voiceprint features; Step S500: Output the classification of the bearing wear degree according to the result of the wear assessment model, including four levels: normal, slight wear, moderate wear, and severe wear.
[0005] Preferably, in a possible implementation manner of the first aspect, the frame rate of the industrial camera is not less than 60 times the rotation speed of the bearing, and an annular polarized light source is equipped; the sampling rate of the array microphone is not less than 10 times the bearing voiceprint frequency.
[0006] Preferably, in a possible implementation manner of the first aspect, the cross-modal attention mechanism satisfies:
[0007] where is a learnable modal weight coefficient, dynamically calculated through a fully connected layer, is the sigmoid activation function, represents the Hadamard product, and are the image feature vector and the voiceprint feature vector respectively.
[0008] Preferably, in a possible implementation manner of the first aspect, the improved Transformer adopts a spatio-temporal cross-attention mechanism:
[0009] where is the query matrix generated by the image features, is the key-value matrix generated by mapping the voiceprint features to different semantic spaces through two independent linear transformations, is the time domain mask matrix, is the position matrix at each time step, n is the time step of the image features, m is the time step of the voiceprint features, d is the feature dimension, and T represents the transpose operation.
[0010] Preferably, in a possible implementation manner of the first aspect, the encoder structure of the loss evaluation model is: The image encoder adopts a deep learning-based EfficientNet-B7 classification model, and outputs a 2560-dimensional image feature vector through its last layer of convolution; The voiceprint encoder adopts a deep learning-based ConvNeXt-XL classification model, and outputs a 2048-dimensional voiceprint feature vector through its last layer of convolution; The fusion layer uniformly maps the image feature vector and the voiceprint feature vector to a 512-dimensional fusion feature space through a multi-layer perceptron.
[0011] Preferably, in a possible implementation manner of the first aspect, the loss function of the loss evaluation model is:
[0012] where is the cross-entropy loss, and are the true and predicted bearing wear grades respectively, is the mean square error of bearing clearance prediction, and are the measured and predicted values of the clearance, is the regularization term of the model parameters.
[0013] Preferably, in a possible implementation manner of the first aspect, the optimizer of the loss evaluation model adopts AdamW, and the parameter update satisfies:
[0014] where is the model parameter at the current time step, is the model parameter at the previous time step, and are the first-order momentum and the second-order momentum respectively. The decay rate of the first-order momentum is , and the decay rate of the second-order momentum is , is the numerical stability term, is the learning rate, is the weight decay coefficient.
[0015] Preferably, in a possible implementation manner of the first aspect, the loss evaluation model outputs a wear depth quantization value , the bearing clearance change and the dynamic load distribution anomaly coefficient . The loss evaluation model uses the scoring formula:
[0016] where , .
[0017] Preferably, in a possible implementation manner of the first aspect, the bearing loss degree is classified into four grades: normal, slight wear, moderate wear, and severe wear according to the score and the preset grading threshold.
[0018] In a second aspect, the present invention provides a contact ball bearing loss degree identification system, which includes: A data acquisition module that synchronously acquires multi-modal data during the operation of the contact ball bearing. The multi-modal data includes a sequence of bearing surface images obtained by an industrial camera and a voiceprint signal collected by an array microphone; A preprocessing module that performs adaptive illumination compensation and denoising processing on the sequence of bearing surface images, and performs frame division, windowing processing on the voiceprint signal and converts it into a Mel spectrogram; A feature fusion module constructs a multimodal feature fusion network based on the denoised bearing surface image sequence and Mel spectrogram, extracts image features and voiceprint features, and uses a cross-modal attention mechanism to fuse the image features and voiceprint features; A loss evaluation module constructs a loss evaluation model based on an improved Transformer to process spatio-temporal features, which are the fused image features and voiceprint features; A result output module outputs the classification of the bearing loss degree according to the result of the loss evaluation model, including four levels: normal, slight wear, moderate wear, and severe wear.
[0019] The beneficial effects of the present invention are as follows: By synchronously collecting multimodal data during the operation of the ball bearing, including the bearing surface image sequence obtained by an industrial camera and the voiceprint signal collected by an array microphone, a comprehensive description of the bearing operation state is realized. By performing adaptive illumination compensation and denoising processing on the image sequence, and performing frame division and windowing processing on the voiceprint signal and converting it into a Mel spectrogram, the quality and usability of the data are improved.
[0020] Constructing a multimodal feature fusion network and using a cross-modal attention mechanism to fuse image features and voiceprint features effectively utilizes the complementarity of multimodal data and improves the accuracy of feature expression. Constructing a loss evaluation model based on an improved Transformer to process spatio-temporal features further improves the accuracy of loss degree recognition. In addition, the present invention also outputs the classification result of the bearing loss degree, including four levels: normal, slight wear, moderate wear, and severe wear, providing strong support for the maintenance and management of equipment. In summary, the present invention has the advantages of high recognition accuracy, strong reliability, and good practicability, and is of great significance for improving the operation performance and service life of mechanical equipment. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 This application provides a flowchart of a method for identifying the loss degree of a ball bearing.
[0023] Figure 2 This application provides a structural diagram of a system for identifying the loss degree of a ball bearing.
[0024] Description of the reference numerals: 1 - data acquisition module, 2 - preprocessing module, 3 - feature fusion module, 4 - loss assessment module, 5 - result output module. Specific embodiments
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] Embodiment 1: As Figure 1 shown, the present invention provides a method for identifying the loss degree of a contact ball bearing, including: Step S100: Synchronously collect multi-modal data during the operation of the contact ball bearing. The multi-modal data includes a sequence of bearing surface images obtained by an industrial camera and a voiceprint signal collected by an array microphone.
[0027] In this embodiment, the industrial camera uses a FLIR OX5-3260 global shutter industrial camera in combination with a Brüel&Kjær 4958-A-011 array microphone. When the bearing speed is 3000 rpm, the bearing rotation frequency is 50 Hz. The frame rate of the industrial camera is not less than 60 times the bearing speed. In this embodiment, the industrial camera is set to a frame rate of 3400 fps, and a circularly polarized light source with a wavelength of 630 nm is configured to eliminate metal reflection interference. The synchronous trigger signal is rigidly connected to the bearing rotating shaft through an encoder to achieve phase synchronization. The array microphone uses a sampling rate of 48 kHz, and the voiceprint signal is collected through a cross-shaped four-microphone array. The sampling rate is not less than 10 times the bearing voiceprint frequency. The system uses the IEEE 1588 protocol to achieve time synchronization of the image and audio signals, and the measured time alignment error is less than 10 microseconds.
[0028] The data synchronization mechanism realizes hardware-level triggering through an FPGA. The camera exposure signal is used as the main clock source to trigger the FPGA to generate a voiceprint acquisition start pulse. The pulse delay is less than 1 ms after calibration. The original data is stored in binary format. The image sequence is saved in 12-bit RAW format, and the voiceprint signal is stored in 24-bit PCM format. All data is embedded with time stamps and transmitted to an industrial computer.
[0029] Step S200: Perform adaptive illumination compensation and denoising processing on the bearing surface image sequence, and perform frame division and windowing processing on the voiceprint signal and convert it into a Mel spectrogram.
[0030] In this embodiment, image preprocessing needs to solve the problems of uneven illumination, metal reflection, and motion blur. First, the multi-scale illumination compensation algorithm based on the Retinex theory is executed in three steps: (1) Separate the incident light component through Gaussian pyramid decomposition; (2) Apply adaptive gamma correction on each layer of the pyramid to equalize the illuminance; (3) When reconstructing the reflection component, use guided filtering to retain edge details. For surface defect enhancement, use locally contrast-limited adaptive histogram equalization.
[0031] The denoising algorithm uses an improved non-local means method: the search window is expanded to 21×21 pixels, the similarity window is 7×7 pixels, and the filtering parameter h is adaptively calculated according to the local noise level. For the highlight area, introduce morphological top-hat transformation to extract bright noise, and then smooth it through bilateral filtering. The final image is verified by Canny edge detection to process the effect.
[0032] The voiceprint signal processing uses frame segmentation and windowing and Mel spectrum conversion techniques. The signal first passes through a pre-emphasis filter to enhance the high-frequency components, and then is framed with a Hamming window. Each frame of data is converted into a power spectrum through FFT. The Mel filter bank is designed with 128 triangular filters, with a frequency range of 80 Hz to 8 kHz. The output of each filter is logarithmically compressed to generate a logarithmic Mel spectrum. To eliminate device differences, the spectrogram is globally mean-normalized: the means of each frequency band are statistically calculated from 100 hours of standard bearing data and standard deviations , and the input spectrum is executed transformation, and the final spectrogram is resampled to 224×224 pixels.
[0033] Step S300: Based on the denoised bearing surface image sequence and Mel spectrogram, construct a multi-modal feature fusion network, extract image features and voiceprint features, and use a cross-modal attention mechanism to fuse image features and voiceprint features.
[0034] In this embodiment, the multi-modal feature fusion network architecture adopts a dual encoder-decoder structure. Among them, the image encoder is improved based on the EfficientNet-B7 model. The input image size is adjusted to 512×512 pixels to meet the high-resolution detection requirements of the bearing surface. When the model performs forward propagation, a deformable convolution module is inserted after the Conv4_3 layer. This module contains 3×3 deformable convolution kernels, and the offset of each kernel is dynamically generated by the auxiliary network, enhancing the ability to extract local deformation features such as raceway spalling and cage fracture. The image encoder finally outputs a 2560-dimensional image feature vector. The voiceprint encoder adopts the ConvNeXt-XL architecture, inputs the Mel spectrogram converted from the voiceprint signal sampled at 192kHz, and introduces a bidirectional gated recurrent unit at the network Stage 3 level, with a hidden layer dimension of 256, which is used to capture the periodic impact components caused by bearing wear in the voiceprint signal. The voiceprint encoder finally outputs a 2048-dimensional voiceprint feature vector.
[0035] The physical equation solution of the vibration model-assisted feature extraction module is executed in parallel with deep learning: During the operation of the bearing, the axial vibration acceleration signal is collected in real time, and the differential equation is numerically solved by the fourth-order Runge-Kutta method
[0036] where the equivalent mass m is dynamically loaded from the bearing model parameter library, and the damping coefficient , and the stiffness coefficient k is updated online according to the measured value of the bearing clearance. The coefficient of the wear force disturbance term , , , and the vibration feature vector calculated and output is concatenated with the image and voiceprint features at the fusion layer.
[0037] The cross-modal attention mechanism includes two stages: spatial-channel collaborative attention and dynamic weight allocation. First, the image feature and the voiceprint feature are respectively reduced to 512 dimensions through 1×1 convolution, and then the fusion formula is executed:
[0038] where the modal weight coefficient is dynamically generated by the vibration feature through a three-layer fully connected network, ensuring that when the vibration signal is abnormal, the weight of the voiceprint feature is increased by 20%-40%. The fused feature and the vibration feature are subjected to channel attention weighting. The specific process is as follows: is input into the Squeeze-and-Excitation module to generate a 512-dimensional channel weight vector, which is combined with Multiply channel by channel and finally output the joint features .
[0039] The network training adopts a hybrid loss function:
[0040] where is the cross-entropy loss, and are the true and predicted bearing wear grades respectively, is the mean square error of bearing clearance prediction, and are the measured and predicted values of the clearance, is the regularization term of the model parameters. The weight coefficient 0.6 of the cross-entropy term in the loss function corresponds to the classification task priority, and the weight 0.3 of the mean square error term ensures that the clearance prediction error is controlled within . When loading the training data, random occlusion and HSV color gamut perturbation are applied to the images, and time-domain masking and frequency-domain random filtering are applied to the voiceprint signals to improve the robustness of the model. The parameter update of the optimizer AdamW satisfies:
[0041] where and are the first-order momentum and the second-order momentum respectively. The decay rate of the first-order momentum , the decay rate of the second-order momentum , is the numerical stability term, the initial learning rate is set to , the weight decay coefficient , and a linear warm-up strategy is adopted during the training process (the learning rate increases from to in the first 10 rounds), and is reduced to in cooperation with the cosine annealing schedule.
[0042] Step S400: Construct a loss evaluation model based on the improved Transformer to process spatio-temporal features, where the spatio-temporal features are the fused image features and voiceprint features.
[0043] In this embodiment, a spatio-temporal Transformer model with physical constraints is constructed. The model input is the joint feature , and first, the dimensions are unified through three layers of MLP to generate the serialized feature (the time step T = 64 corresponds to the number of bearing rotation cycles). The encoder layer adopts a spatio-temporal cross-attention mechanism and introduces position encoding guided by vibration features: the standard sine position encoding is replaced by the phase angle solved from the vibration equation The generated rotational position encoding, the position matrix at each time step is calculated by such that the key-value pairs in the attention calculation contain the actual operating phase information of the bearing.
[0044] The spatio-temporal cross-attention is divided into two stages: cross-modal interaction and physical constraint. In the cross-modal interaction stage, the query matrix Q is generated from the image feature sequence, and the key-value matrices K and V are generated from the voiceprint feature sequence. The attention score calculation uses an improved formula:
[0045] where is the query matrix generated from the image features, is the key-value matrix generated by mapping the voiceprint features to different semantic spaces through two independent linear transformations, is the time-domain mask matrix, is the position matrix at each time step, n is the time step of the image features, m is the time step of the voiceprint features, d is the feature dimension, T represents the transpose operation, and the time-domain mask matrix is set as a lower triangular matrix. In the physical constraint stage, the wear depth predicted by the vibration model is used as prior knowledge and converted into a 256-dimensional feature vector through differentiable layering, and gate-added to the attention output features:
[0046] The decoder part uses causal convolution to enhance the temporal modeling ability: After the output of the Transformer encoder, 4 layers of causal dilated convolution are connected, the convolution kernel size of each layer is 3, and the dilation coefficients are 1, 2, 4, and 8 in turn, and the number of channels remains 256 unchanged. The final state is input into three parallel regression heads after global average pooling: (1) The wear depth prediction head, using 3 layers of MLP, with an output dimension of 1, and the activation function Sigmoid is constrained to 0 - 0.1 mm; (2) The clearance change prediction head, using 2 layers of MLP, with an output dimension of 1, and linear activation; (3) The anomaly coefficient prediction head, using the Mahalanobis distance to calculate the deviation of the feature distribution.
[0047] The model training adopts a multi-stage optimization strategy: In the first 50 rounds, the parameters of the feature extraction network are fixed, and only the Transformer part is trained, and the learning rate is set to In the loss function, the term only acts on the clearance prediction; in the subsequent 150 rounds, all parameters are jointly optimized, and the learning rate is reduced to , the KL divergence loss between the predicted values of the vibration model and the deep learning model is introduced to force the wear depth prediction errors of the two modalities to be less than 5%. The training data is divided into training set, validation set, and test set in the ratio of 8:1:1, and the early stopping mechanism monitors the comprehensive metrics on the validation set:
[0048] When the validation score does not improve for 20 consecutive rounds, automatic stopping is triggered and rolled back to the best checkpoint.
[0049] Step S500: Output the bearing loss degree classification according to the results of the loss assessment model, including four levels: normal, slight wear, moderate wear, and severe wear.
[0050] In this embodiment, the classification logic adopts a dual decision-making mechanism: the main scoring system calculates the comprehensive index according to the formula where , , and the thresholds are normal as , slight wear , moderate wear , severe wear ).
[0051] The auxiliary system verifies the multi-modal feature consistency: calculate the cosine similarity S of the image and voiceprint features. If then manual review is triggered. The specific classification rules are as follows: Normal: and and , and ; Slight wear: or , and ; Moderate wear: or , and ; Severe wear: or or , or .
[0052] The exception handling mechanism includes: (1) When and conflict across levels (such as corresponding to moderate, corresponding to severe), take the highest level; (2) Automatically trigger historical data backtracking for severe wear cases and retrieve the operation trend in the last 100 hours; (3) The output result is transmitted to the PLC through the Modbus-TCP protocol, driving the relay contacts to trigger the equipment to stop.
[0053] The system self-calibration is performed once every 8 hours: the built-in reference bearing runs for 10 minutes, and the data is collected to verify the model output. If the absolute error of continuously exceeds for three consecutive times, then online incremental learning is started: the Transformer layer is frozen, and the feature fusion network is fine-tuned for 50 rounds, with a learning rate of The data buffer retains the most recent 1000 groups of samples. The calibration log is recorded in the SQL database, and PDF reports are supported for export through the web interface.
[0054] Embodiment 2: As Figure 2 shown, the present invention provides a system for identifying the loss degree of a contact ball bearing, including: Data acquisition module 1: Synchronously acquire multi-modal data during the operation of the bearing, including high-resolution image sequences and multi-channel voiceprint signals, realize the time synchronization of cross-modal data through a triggering mechanism, and at the same time integrate vibration sensors to obtain mechanical dynamic responses.
[0055] Preprocessing module 2: Perform dynamic light compensation and noise suppression processing on the images to enhance surface defect features; perform frame addition and windowing processing on the voiceprint signals and convert them into Mel spectrograms to extract acoustic fingerprints related to the bearing state.
[0056] Feature fusion module 3: Use a deep network to extract high-dimensional features of images and voiceprints respectively, dynamically allocate modal weights through an attention mechanism, and fuse them through a multi-modal feature fusion network to generate a joint representation to capture cross-modal correlations.
[0057] Loss evaluation module 4: Process the fused features based on the time series modeling framework of the improved Transformer, analyze the wear depth, clearance change and abnormal vibration mode, and predict the comprehensive loss state of the bearing through multi-scale feature interaction.
[0058] Result output module 5: Combine the physical model and the deep learning prediction results to generate the grading of the bearing loss degree and maintenance suggestions. The grading of the bearing loss degree includes four levels: normal, slight wear, moderate wear, and severe wear, and supports real-time visual feedback and historical data backtracking analysis.
[0059] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.
Claims
1. A method for identifying the degree of loss of a contact ball bearing, characterized in that: The method comprises: Step S100: synchronously collecting multimodal data of the contact ball bearing during operation, wherein the multimodal data includes a bearing surface image sequence acquired by an industrial camera and a voiceprint signal acquired by an array microphone; Step S200: performing adaptive illumination compensation and denoising processing on the bearing surface image sequence, performing frame division and windowing processing on the voiceprint signal and converting it into a Mel spectrum diagram; Step S300: Based on the denoised bearing surface image sequence and Mel spectrum, a multimodal feature fusion network is constructed to extract image features and voiceprint features, and the image features and voiceprint features are fused using a cross-modal attention mechanism; Step S400: constructing a loss assessment model based on an improved Transformer to process spatiotemporal features, where the spatiotemporal features are fused image features and voiceprint features; Step S500: Outputting the bearing wear degree classification according to the result of the loss assessment model, including four levels: normal, slight wear, moderate wear, and severe wear.
2. A contact ball bearing wear degree identification method as claimed in claim 1, characterized in that: The frame rate of the industrial camera is not less than 60 times the bearing rotation speed and is equipped with a circular polarized light source; the sampling rate of the array microphone is not less than 10 times the bearing soundprint frequency.
3. A contact ball bearing wear degree identification method as claimed in claim 1, characterized in that: The cross-modal attention mechanism satisfies: in is a learnable modal weight coefficient, which is dynamically calculated through the fully connected layer. is the sigmoid activation function, represents the Hadamard product, and They are image feature vector and voiceprint feature vector respectively.
4. A contact ball bearing wear degree identification method as claimed in claim 1, characterized in that: Improve Transformer to adopt spatiotemporal cross attention mechanism: in The query matrix generated for the image features, The key-value matrix generated by mapping the voiceprint features to different semantic spaces through two independent linear transformations, is the time domain mask matrix, is the position matrix for each time step, is a real number matrix, n is the time step of the image feature, m is the time step of the voiceprint feature, d is the feature dimension, and T represents the transpose operation.
5. A contact ball bearing wear degree identification method as claimed in claim 4, characterized in that: The encoder structure of the loss estimation model is: The image encoder uses the EfficientNet-B7 structure, and outputs a 2560-dimensional image feature vector through its last convolution layer; The voiceprint encoder adopts the ConvNeXt-XL structure, and outputs a 2048-dimensional voiceprint feature vector through its last layer convolution; The fusion layer uses a multi-layer perceptron to uniformly map the image feature vector and voiceprint feature vector to a 512-dimensional fusion feature space.
6. A contact ball bearing wear degree identification method as claimed in claim 1, characterized in that: The loss function of the loss assessment model is: in is the cross entropy loss, and are the actual and predicted bearing wear levels, is the mean square error of bearing clearance prediction, and is the measured value and predicted value of clearance, is the regularization term of the model parameters.
7. A contact ball bearing wear degree identification method as claimed in claim 1, characterized in that: The optimizer of the loss evaluation model adopts AdamW, and the parameter update satisfies: in are the model parameters at the current time step, are the model parameters at the previous time step, and They are the first-order momentum and the second-order momentum respectively. The decay rate of the first-order momentum is , the decay rate of the second-order momentum is , is a numerical stability term, is the learning rate, is the weight decay coefficient.
8. A contact ball bearing wear degree identification method as claimed in claim 1, characterized in that: The wear assessment model output includes a quantitative value of the wear depth , bearing clearance change and dynamic load distribution abnormality coefficient , the loss assessment model uses the scoring formula: in , .
9. A contact ball bearing wear degree identification method as claimed in claim 8, characterized in that: According to the score and preset classification threshold, the bearing loss degree is classified into four levels: normal, slight wear, moderate wear and severe wear.
10. A contact ball bearing wear degree identification system, characterized in that: The system comprises: The data acquisition module synchronously collects multimodal data of the contact ball bearing during operation. The multimodal data includes the bearing surface image sequence acquired by the industrial camera and the voiceprint signal collected by the array microphone; The preprocessing module performs adaptive illumination compensation and denoising on the bearing surface image sequence, performs frame division and windowing on the voiceprint signal and converts it into a Mel spectrum. The feature fusion module builds a multimodal feature fusion network based on the denoised bearing surface image sequence and Mel spectrum, extracts image features and voiceprint features, and uses a cross-modal attention mechanism to fuse image features and voiceprint features; The loss assessment module builds a loss assessment model based on the improved Transformer to process spatiotemporal features. The spatiotemporal features are the fused image features and voiceprint features. The result output module outputs the bearing loss degree classification according to the results of the loss assessment model, including four levels: normal, slight wear, moderate wear, and severe wear.
Citation Information
Cited By
Bearing clearance dynamic measurement method and system
CN120747039A