Equipment fault detection method based on multi-source data
By collecting and fusing visual and audio data and utilizing adversarial training techniques, the problems of misjudgment and missed judgment in multi-source data fault detection under complex environments have been solved. This has enabled accurate identification and grading of equipment faults, reduced the false alarm rate, and made it suitable for intelligent operation and maintenance of industrial equipment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHEASTERN UNIV CHINA
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-28
Smart Images

Figure CN121935752A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of equipment fault diagnosis technology, and in particular to a method for equipment fault detection based on multi-source data. Background Technology
[0002] In the field of industrial equipment monitoring, early detection of equipment failures is crucial for ensuring production safety and reducing maintenance costs. Traditional fault detection mainly relies on periodic manual inspections or single-sensor monitoring, which often suffer from latency and struggle to promptly capture abnormal equipment conditions. With advancements in sensor technology and data analysis methods, multi-source data fusion-based fault detection methods have gradually become a research hotspot. By integrating multi-dimensional information from different sensors, the accuracy and real-time performance of fault detection can be improved.
[0003] However, existing multi-source data fault detection technologies still have many shortcomings. First, most existing front-end devices are purely visual cameras, relying on single image data, which are prone to false positives and false negatives in complex environments such as obstruction, strong light, and fog. Existing algorithms have limited functionality and fail to achieve joint detection of dynamic behavior and static hazards, nor do they perform correlation verification and cross-analysis of visual data with environmental sensor data (such as temperature, humidity, gas, vibration, etc.), resulting in a persistently high false alarm rate (the industry average exceeds 30%). Summary of the Invention
[0004] In view of this, the present invention provides a device fault detection method based on multi-source data, which obtains fused features through sound data and visual data, and realizes device fault diagnosis through fused features.
[0005] Therefore, the present invention provides the following technical solution: A device fault detection method based on multi-source data, characterized in that it includes: Collect visual and audio data from the device's operation; The visual and audio data are preprocessed; Visual features are extracted from preprocessed visual data, and sound features are extracted from preprocessed sound data. A comprehensive fused feature is obtained by fusing visual and auditory features through adversarial training; Based on the comprehensive fusion of features, combining visual and acoustic features, the equipment fault category is determined, and the corresponding fault level is identified.
[0006] Furthermore, the visual and audio data are preprocessed, including: The video frames are sequentially subjected to grayscale conversion, Gaussian filtering, edge detection, and region of interest cropping. The voiceprint signal is sequentially subjected to pre-emphasis, Hamming window framing, fast Fourier transform, and Mel filter bank.
[0007] Furthermore, the visual features are extracted based on the preprocessed visual data using the YOLOv5-s model; Fault feature frequencies are extracted from preprocessed sound data using MFCC feature vector analysis as the sound features.
[0008] Furthermore, the method of obtaining comprehensive fused features by fusing visual and auditory features through adversarial training includes: Visual confidence is calculated based on visual features, and sound confidence is calculated based on sound features. Visual confidence and sound confidence are input into a two-layer multilayer perceptron and the Sigmoid function is used to output dynamic weights. Preliminary fusion features are obtained by weighted fusion based on dynamic weights; Adversarial training is performed on visual and auditory features to fuse them and obtain pre-fused features; The comprehensive fusion feature is obtained by fusing the preliminary fusion feature and the pre-fusion feature.
[0009] Furthermore, the method of determining the equipment fault category based on comprehensive fusion features combined with visual and acoustic features includes: When both visual and acoustic features meet the preset threshold of the target level, and the integrated features match the equipment fault category of the target level, the operating condition is determined to be at the target level.
[0010] Furthermore, the sound features include: Outer ring failure frequency, inner ring failure frequency, and cage failure frequency; The visual features include: bearing end cap vibration displacement, shaft extension swing amplitude, and lubricating oil leakage traces.
[0011] Furthermore, the visual feature threshold division includes: The bearing is considered to be operating normally if the vibration displacement of the bearing end cover is ≤2 pixels / frame, the swing amplitude of the shaft extension is ≤0.5°, and there are no signs of lubricating oil leakage. The bearing end cover vibration displacement is 2-5 pixels / frame, the shaft extension swing amplitude is 0.5°-1°, and there are no traces of lubricating oil leakage, indicating that the bearing is slightly worn. The bearing end cover vibration displacement is 5-8 pixels / frame, the shaft extension swing amplitude is 1°-2°, and the oil stain area ratio is ≤5%, indicating that the bearing is moderately worn. If the bearing end cover vibration displacement is ≥8 pixels / frame, the shaft extension swing amplitude is ≥2°, and the oil stain area ratio is >5%, it indicates severe bearing wear or peeling. Vibration displacement of the bearing end cap, no swing amplitude of the shaft extension, and oil stain area accounting for more than 3% indicate insufficient bearing lubrication.
[0012] Furthermore, the sound feature threshold division includes: If the sound characteristic has a fault-free frequency peak, then the bearing is operating normally. A signal-to-noise ratio (SNR) of 15-25 dB for the outer ring fault frequency or the inner ring fault frequency indicates slight bearing wear. A signal-to-noise ratio of 25-40dB for the outer ring fault frequency or 25-40dB for the inner ring fault frequency indicates moderate bearing wear. If the signal-to-noise ratio of the outer ring fault frequency or the signal-to-noise ratio of the inner ring fault frequency is ≥40dB, it indicates severe wear or spalling of the bearing.
[0013] Furthermore, the calculation of visual confidence based on visual features includes:
[0014] in, The visual confidence level has a value range of [0,1]. This refers to the vibration displacement of the bearing end cover; This refers to the swing amplitude of the shaft extension; The percentage of the area showing signs of lubricant leakage. For feature normalization function; The consistency coefficient is used to verify the abnormal synergy of multiple visual features: if two or more features exceed the threshold simultaneously, the consistency coefficient is 0.2; if one feature is abnormal, the consistency coefficient is 0.05; if there are no abnormal features, the consistency coefficient is 0.1. , , and For weight parameters, .
[0015] Furthermore, the calculation of sound confidence based on sound features includes:
[0016]
[0017] in, The sound confidence level, with a value range of [0,1]. For the outer ring fault frequency, For inner ring fault frequency, To reduce cage failure frequency; , , For the weight parameters, satisfying ; , , For the signal-to-noise ratio weights at each frequency, satisfying ; The significance of the fault frequency peak is used to quantify the difference between the fault frequency peak and the background noise. The formula is:
[0018] In the formula, The signal peak value corresponds to the fault frequency. This represents the mean of the background noise. The normalized signal-to-noise ratio for each fault frequency is given by the following formula:
[0019] In the formula, 40dB is the signal-to-noise ratio threshold for severe faults, ensuring that the confidence does not overflow after exceeding the threshold; The characteristic energy distribution verification coefficient is 0.2 if the energy dispersion is ≤30%; 0.1 if the energy dispersion is between 30% and 60%; and 0.05 if the energy dispersion is >60% or the peak value of the fault-free frequency.
[0020] Advantages and positive effects of the present invention: This method improves the completeness of fault feature extraction by simultaneously acquiring visual and acoustic data and combining it with multi-dimensional status information of the equipment, thus avoiding the limitations of single sensor data and improving the accuracy of fault diagnosis. The use of adversarial training fusion technology effectively suppresses the impact of noise, lighting changes, and other interferences in the industrial environment on the features, enhancing the robustness of the fused features and ensuring stable detection performance under complex operating conditions.
[0021] By fusing and collaboratively analyzing visual and acoustic features, this method can not only accurately identify fault categories but also quantify fault levels, shorten troubleshooting time, and reduce false alarm rates. Visual confidence focuses on multi-feature collaborative verification to address the issues of visual data being susceptible to environmental interference and misjudgment based on a single feature; acoustic confidence focuses on the significance of fault frequency and energy purity to address the issues of acoustic data being susceptible to noise interference and low fault signal recognition, avoiding homogenization between the two confidence calculation methods. Visual confidence uses a consistency coefficient to verify feature synergy, while acoustic confidence uses an energy distribution coefficient to verify signal purity; both are designed to address the deficiencies of their respective data types, improving the reliability of confidence calculation. Both confidence calculations include a reserved interface for dynamic weight adjustment, allowing adaptation to diverse industrial scenarios based on the varying sensitivities of visual / acoustic features across different devices.
[0022] This method does not rely on historical fault databases, is suitable for new equipment or rare fault scenarios, has strong generalization ability, and can be widely used in the field of intelligent operation and maintenance of industrial equipment. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a device fault detection method based on multi-source data. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0027] This invention provides a device fault detection method based on multi-source data, which collects visual and sound data of device operation; extracts features from the visual and sound data, and fuses the visual and sound features based on adversarial training; and determines the device fault category and fault level by combining the fused features with the visual and sound features.
[0028] Combination Figure 1 As shown, a device fault detection method based on multi-source data includes: S1. Collect visual and audio data from the device's operation; S2. Extract the visual and acoustic features from the collected visual and acoustic data; S3, based on adversarial training that fuses visual and auditory features; S4. By combining visual and acoustic features, the fault type and fault level of the equipment can be determined.
[0029] Example This method is used to test the 6308 deep groove ball bearing at the front end of a Y2-160M-4 type three-phase AC motor with a power of 11kW and a rated speed of 1480r / min.
[0030] S1. Collect equipment operation data.
[0031] 1. Visual data: In this embodiment, visual data is collected by a DH-IPC-POE200W industrial camera with a lens installed 30cm in front of the motor bearing end cover, perpendicular to the connection between the bearing end cover and the shaft extension. The acquisition parameters are 1080P resolution, 30fps frame rate, exposure time 1 / 500s and ISO sensitivity 400 to ensure no motion blur in motion scenes.
[0032] 2. Sound data: In this embodiment, data is collected by a WN1601B industrial-grade sound sensor that is attached to the top of the bearing end cover and is ≤5cm away from the outer ring of the bearing, while avoiding the resonance interference of the motor base; the acquisition parameters are: 48kHz sampling rate, 16-bit depth, acquisition frequency band 20Hz-10kHz, and 50Hz power frequency interference is removed by hardware filtering.
[0033] S2. Preprocess the collected data.
[0034] 1. Synchronization Processing: The FPGA synchronization circuit, combined with the GPS / BeiDou timing module, adds μs-level timestamps to the visual frames and audio sampling points, ensuring precise alignment between a single frame image (33ms period) and the corresponding audio data for that time period, with a time deviation ≤5. .
[0035] 2. Visual preprocessing: The video frames are sequentially processed by grayscale conversion, Gaussian filtering, edge detection, and region of interest cropping, preserving the bearing end cap and shaft extension area while removing background interference.
[0036] In this embodiment, the Gaussian filter kernel size is 5×5; the threshold of the Canny operator in edge detection is 100-200; and the size of the bearing end cover and shaft extension area is 800×600 pixels.
[0037] 3. Sound preprocessing: The voiceprint signal is sequentially subjected to pre-emphasis, Hamming window framing, fast Fourier transform, and Mel filter bank.
[0038] In this embodiment, the pre-emphasis is a high-pass filter with a cutoff frequency of 1kHz; the frame length of the Hamming window framing is 25ms, and the frame shift is 10ms; the number of points in the fast Fourier transform is 1024; and the Mel filter bank includes 40 filters. S3. After data preprocessing, feature extraction is performed.
[0039] Visual features were extracted based on visual data, including: bearing end cover vibration displacement, shaft extension swing amplitude, and lubricating oil leakage traces.
[0040] Sound features are extracted based on sound data, including: outer ring fault frequency, inner ring fault frequency, and cage fault frequency.
[0041] Environmental features are extracted based on auxiliary parameters, including: real-time load current of the motor, ambient temperature around the bearing, and relative humidity.
[0042] 1. Visual Feature Extraction: Using the YOLOv5-s model, visual features are extracted, including: Bearing end cover vibration displacement: The average displacement of the edge pixels of the end cover is calculated by the inter-frame difference method (unit: pixels / frame). Shaft extension swing amplitude: Fit the shaft extension edge profile and calculate the maximum offset angle (unit: °) between the axis center line and the baseline in each frame. Lubricating oil leakage traces: Identify the area percentage (in %) of brown oil stains by color space conversion.
[0043] 2. Sound Feature Extraction: By using MFCC eigenvector analysis (Mel-Frequency Cepstral Coefficients), characteristic frequencies of bearing failure are captured, including: Outer ring failure frequency, inner ring failure frequency, and cage failure frequency.
[0044] In this embodiment, the outer ring fault frequency is 148Hz; the inner ring fault frequency is 237Hz; and the cage fault frequency is 23Hz.
[0045] S4. Calculate visual confidence and sound confidence based on the extracted visual and sound features.
[0046] 1. Visual confidence focuses on "feature reliability and synergy of changes," employing a calculation method of "normalized feature weighting + anomaly consistency coefficient." This method is adapted to the characteristics of visual data being susceptible to interference from lighting and occlusion. The formula is expressed as:
[0047] in, The visual confidence level ranges from [0,1]. The closer the value is to 1, the clearer the visual feature fault indication and the higher the reliability. Vibration displacement of the bearing end cap (pixels / frame); The swing amplitude of the shaft extension (°); The percentage of area showing signs of lubricating oil leakage (%) Let be the feature normalization function. Max-min normalization is used to map the feature values to the interval [0, 0.8]. The formula is as follows:
[0048] In the formula, This represents the upper limit of the severe fault threshold corresponding to this feature. This is the lower limit of the normal state threshold.
[0049] The feature consistency coefficient ranges from 0 to 0.2 and is used to verify the abnormal synergy of multiple visual features: if two or more features exceed the threshold simultaneously, the consistency coefficient is 0.2; if one feature is abnormal, the consistency coefficient is 0.05; and if there are no abnormal features, the consistency coefficient is 0.1. , , and For weight parameters, Default value , =0.3、 =0.25、 =0.15, the feature weight can be dynamically adjusted according to the device type.
[0050] 2. Sound confidence focuses on fault frequency identification and energy concentration, employing a composite calculation method of peak significance + signal-to-noise ratio weighting + energy distribution verification. This method is adapted to the characteristics of sound data being sensitive to mechanical vibration and easily affected by environmental noise, creating a logical difference from visual confidence. The formula is expressed as follows:
[0051]
[0052] in, The sound confidence level ranges from [0,1]. The closer the value is to 1, the stronger the fault directionality of the sound feature and the better the anti-interference ability. For the outer ring fault frequency, For inner ring fault frequency, To reduce cage failure frequency; , , For the weight parameters, satisfying Default value: =0.35、 =0.5、 =0.15, suitable for industrial environment noise interference scenarios, and can be dynamically adjusted according to the on-site noise intensity.
[0053] The significance of the fault frequency peak (range [0, 0.4]) is used to quantify the difference between the fault frequency peak and the background noise. The formula is:
[0054] In the formula, The signal peak value corresponds to the fault frequency. The ratio represents the mean of background noise; a larger ratio indicates a more significant fault frequency. The normalized signal-to-noise ratio (dB) for each fault frequency, ranging from [0,1], is given by the following formula:
[0055] In the formula, 40dB is the signal-to-noise ratio threshold for severe faults, ensuring that the confidence does not overflow after exceeding the threshold; , , For the signal-to-noise ratio weights at each frequency, satisfying ;default =0.4、 =0.4、 =0.2; This is the characteristic energy distribution verification coefficient, with a value range of [0, 0.2], used to determine whether the fault energy is concentrated in the target frequency band. If the pure fault frequency energy is concentrated in the corresponding characteristic frequency band, that is, the energy dispersion is ≤30%, the coefficient is 0.2; the energy is concentrated, indicating a fault signal. If the energy dispersion is between 30% and 60%, the coefficient is 0.1; however, some interference requires consideration of other characteristics. If the energy dispersion is greater than 60% or the peak value of the fault-free frequency is reached, the coefficient is 0.05; if the interference is severe, the confidence level decreases.
[0056] S5, integration of dynamic weight allocation and adversarial training.
[0057] 1. Input visual and auditory confidence scores into a two-layer multilayer perceptron and use a sigmoid function to output dynamic weights; then perform weighted fusion based on these dynamic weights to obtain preliminary fused features, using the following formula:
[0058] in, For the updated visual feature weights, Updated sound feature weights; For visual feature vectors, This is the sound feature vector.
[0059] 2. Integration of adversarial training: 1) The encoder compresses the dimensions of both the visual and acoustic feature vectors to 64 dimensions. 2) The generator generates an aligned feature matrix based on compressed features to simulate the feature distribution under normal or fault conditions; 3) The discriminator analyzes the feature matrix through a sliding time window, outputs a truth score, and feeds it back to optimize the generator; 4) Initial fusion characteristics Pre-fusion features of adversarial training output To obtain comprehensive integration characteristics .
[0060] S6. The deep learning model outputs fault categories by integrating and fusing features, including: The bearing is normal, the bearing is slightly worn, the bearing is moderately worn, the bearing is severely worn or peeling off, and the bearing is insufficiently lubricated.
[0061] S7. Determine the equipment fault level based on integrated features, visual features, and sound features.
[0062] When visual and acoustic features simultaneously meet the preset thresholds for their respective levels, and the integrated features match the corresponding equipment fault category, the equipment is judged to be at the preset fault level.
[0063] 1. When the visual characteristics meet the following criteria: bearing end cover vibration displacement ≤ 2 pixels / frame, shaft extension swing amplitude ≤ 0.5°, and no traces of lubricating oil leakage; and the sound characteristics meet the following criteria: no fault frequency peak; and the integrated fusion features match the normal model; and all three criteria are met simultaneously, the bearing is judged to be operating normally, and the fault level is level 1.
[0064] 2. When the visual characteristics meet the following criteria: the bearing end cover vibration displacement is 2-5 pixels / frame, the shaft extension swing amplitude is 0.5°-1°, and there are no traces of lubricating oil leakage; and the sound characteristics meet the following criteria: the signal-to-noise ratio of the outer ring fault frequency or the signal-to-noise ratio of the inner ring fault frequency is 15-25dB; and the comprehensive fusion feature matches the mild fault model; and all three criteria are met simultaneously, the bearing is judged to have mild wear, and the fault level is 2.
[0065] 3. When the visual characteristics meet the following conditions: the bearing end cover vibration displacement is 5-8 pixels / frame, the shaft extension swing amplitude is 1°-2°, and the oil stain area ratio is ≤5%; and the sound characteristics meet the following conditions: the signal-to-noise ratio of the outer ring fault frequency is 25-40dB or the signal-to-noise ratio of the inner ring fault frequency is 25-40dB; and the comprehensive fusion of features indicates a moderate fault model; then the bearing has moderate wear and the fault level is 3.
[0066] 4. When the visual characteristics meet the following criteria: bearing end cover vibration displacement ≥ 8 pixels / frame, shaft extension swing amplitude ≥ 2°, oil stain area ratio > 5%; and the sound characteristics meet the following criteria: outer ring fault frequency signal-to-noise ratio or inner ring fault frequency signal-to-noise ratio ≥ 40dB; and the comprehensive fusion features match the severe fault model; then the bearing is severely worn or peeling off, and the fault level is 4.
[0067] 5. When the visual characteristics meet the following conditions: no obvious displacement and no abnormal swaying, oil stain area ratio >3%; and the sound characteristics have no specific fault frequency; and the comprehensive fusion characteristics match the insufficient lubrication model; then the bearing is insufficiently lubricated, and the fault level is level 2.
[0068] A fault warning is triggered for a fault level of 2; an alarm is triggered for a fault level of 3; and an emergency shutdown is triggered for a fault level of 4.
[0069] In this embodiment, the integrated features can be input into a deep learning model to obtain the equipment fault category.
[0070] The terminal outputs alarms of corresponding levels via audible and visual alarms: Level 1: No alarm; Level 2: Yellow light and intermittent buzzer; Level 3: Red light and continuous buzzer; Level 4: Red light, rapid buzzer, and DI / DO interface triggers motor stop signal.
[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for equipment fault detection based on multi-source data, characterized in that, include: Collect visual and audio data from the device's operation; The visual and audio data are preprocessed; Visual features are extracted from preprocessed visual data, and sound features are extracted from preprocessed sound data. A comprehensive fused feature is obtained by fusing visual and auditory features through adversarial training; Based on the comprehensive fusion of features, combining visual and acoustic features, the equipment fault category is determined, and the corresponding fault level is identified.
2. The method according to claim 1, characterized in that, Preprocessing of the visual and audio data includes: The video frames are sequentially subjected to grayscale conversion, Gaussian filtering, edge detection, and region of interest cropping. The voiceprint signal is sequentially subjected to pre-emphasis, Hamming window framing, fast Fourier transform, and Mel filter bank.
3. The method according to claim 2, characterized in that, The visual features are extracted based on the preprocessed visual data using the YOLOv5-s model. Fault feature frequencies are extracted from preprocessed sound data using MFCC feature vector analysis as the sound features.
4. The method according to claim 1, characterized in that, The method of obtaining comprehensive fused features by fusing visual and auditory features through adversarial training includes: Visual confidence is calculated based on visual features, and sound confidence is calculated based on sound features. Visual confidence and sound confidence are input into a two-layer multilayer perceptron and the Sigmoid function is used to output dynamic weights. Preliminary fusion features are obtained by weighted fusion based on dynamic weights; Adversarial training is performed on visual and auditory features to fuse them and obtain pre-fused features; The comprehensive fusion feature is obtained by fusing the preliminary fusion feature and the pre-fusion feature.
5. The method according to claim 1, characterized in that, The method of determining the equipment fault category based on comprehensive fusion features combined with visual and acoustic features includes: When both visual and acoustic features meet the preset threshold of the target level, and the integrated features match the equipment fault category of the target level, the operating condition is determined to be at the target level.
6. The method according to claim 1, characterized in that, The sound features include: Outer ring failure frequency, inner ring failure frequency, and cage failure frequency; The visual features include: bearing end cap vibration displacement, shaft extension swing amplitude, and lubricating oil leakage traces.
7. The method according to claim 6, characterized in that, The visual feature threshold division includes: The bearing is considered to be operating normally if the vibration displacement of the bearing end cover is ≤2 pixels / frame, the swing amplitude of the shaft extension is ≤0.5°, and there are no signs of lubricating oil leakage. The bearing end cover vibration displacement is 2-5 pixels / frame, the shaft extension swing amplitude is 0.5°-1°, and there are no traces of lubricating oil leakage, indicating that the bearing is slightly worn. The bearing end cover vibration displacement is 5-8 pixels / frame, the shaft extension swing amplitude is 1°-2°, and the oil stain area ratio is ≤5%, indicating that the bearing is moderately worn. If the bearing end cover vibration displacement is ≥8 pixels / frame, the shaft extension swing amplitude is ≥2°, and the oil stain area ratio is >5%, it indicates severe bearing wear or peeling. Vibration displacement of the bearing end cap, no swing amplitude of the shaft extension, and oil stain area accounting for more than 3% indicate insufficient bearing lubrication.
8. The method according to claim 6, characterized in that, The sound feature threshold division includes: If the sound characteristic has a fault-free frequency peak, then the bearing is operating normally. A signal-to-noise ratio (SNR) of 15-25 dB for the outer ring fault frequency or the inner ring fault frequency indicates slight bearing wear. A signal-to-noise ratio of 25-40dB for the outer ring fault frequency or 25-40dB for the inner ring fault frequency indicates moderate bearing wear. If the signal-to-noise ratio of the outer ring fault frequency or the signal-to-noise ratio of the inner ring fault frequency is ≥40dB, it indicates severe wear or spalling of the bearing.
9. The method according to claim 4, characterized in that, The calculation of visual confidence based on visual features includes: in, The visual confidence level has a value range of [0,1]. This refers to the vibration displacement of the bearing end cover; This refers to the swing amplitude of the shaft extension; The percentage of the area showing signs of lubricant leakage. For feature normalization function; The consistency coefficient is used to verify the abnormal synergy of multiple visual features: if two or more features exceed the threshold simultaneously, the consistency coefficient is 0.2; if one feature is abnormal, the consistency coefficient is 0.05; if there are no abnormal features, the consistency coefficient is 0.
1. , , and For weight parameters, .
10. The method according to claim 4, characterized in that, The calculation of sound confidence based on sound features includes: in, The sound confidence level, with a value range of [0,1]. For the outer ring fault frequency, For inner ring fault frequency, To reduce cage failure frequency; , , For the weight parameters, satisfying ; , , For the signal-to-noise ratio weights at each frequency, satisfying ; The significance of the fault frequency peak is used to quantify the difference between the fault frequency peak and the background noise. The formula is: In the formula, The signal peak value corresponds to the fault frequency. This represents the mean of the background noise. The normalized signal-to-noise ratio for each fault frequency is given by the following formula: In the formula, 40dB is the signal-to-noise ratio threshold for severe faults, ensuring that the confidence does not overflow after exceeding the threshold; The characteristic energy distribution verification coefficient is 0.2 if the energy dispersion is ≤30%; 0.1 if the energy dispersion is between 30% and 60%; and 0.05 if the energy dispersion is >60% or the peak value of the fault-free frequency.