Non-contact micro-vibration measurement and fault diagnosis method based on dynamic visual signal alignment

By using the methods of dynamic visual signal alignment and feature clustering, event cameras and deep neural networks, the problems of non-contact micro-vibration measurement and fault diagnosis of rotating machinery are solved, and efficient fault identification and monitoring are achieved.

CN119719979BActive Publication Date: 2025-10-03XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411770214.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-10-03
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

Existing contact sensors have problems such as wear, poor applicability, and unstable performance in high-temperature and high-pressure environments in the condition monitoring of rotating machinery. The event data of non-contact event cameras lacks detailed feature information, making it difficult to effectively reflect equipment vibration and perform fault diagnosis.

Method used

A method based on dynamic visual signal alignment is adopted to collect event data of rotating machinery using an event camera. Through signal alignment and feature clustering, combined with a deep neural network, fault diagnosis is performed, the time domain micro-vibration signal of each pixel point is extracted, and the fault mode is calibrated and identified.

Benefits of technology

It realizes non-contact, global vibration monitoring and fault diagnosis, improves the monitoring applicability and fault identification accuracy of rotating machinery, and enhances the robustness and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719979B_ABST
    Figure CN119719979B_ABST
Patent Text Reader

Abstract

A non-contact micro-vibration measurement and fault diagnosis method based on dynamic visual signal alignment. The method first uses a dynamic visual sensor-event camera to non-contactly collect event data of rotating machinery vibration, and uses an event data characterization method to convert the event stream into an equidistant, non-overlapping and continuous event frame sequence. Then, a deep neural network is used to extract the micro-vibration information of each pixel in the event frame sequence, and a signal alignment method is used to reduce the difference in vibration signals of adjacent pixels. The contact sensing data is used as a reference for optimizing model performance through a feature clustering method. Finally, the obtained micro-vibration data is used for intelligent fault diagnosis. The present invention improves the applicability of non-contact sensors in the field of vibration monitoring and fault diagnosis of rotating machinery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of mechanical fault monitoring and diagnosis, and in particular relates to a non-contact micro-vibration measurement and fault diagnosis method based on dynamic visual signal alignment. Background Art

[0002] In modern industry, rotating machinery plays a key role in tasks such as manufacturing and transportation. The stable and precise operation of such machinery is essential to ensuring production quality, efficiency, and safety. However, complex working environments significantly affect machine performance, and in severe cases, can lead to equipment failure or even shutdown. Therefore, it is very important to monitor the status of rotating machinery in a timely and effective manner.

[0003] At present, the condition monitoring of rotating machinery adopts various types of signals, including vibration, temperature, pressure, current, sound signals, etc. Among them, vibration signals are particularly important and effective for machine condition monitoring. Since it is highly sensitive to slight changes and faults in the mechanical structure, vibration signals can well reveal the operating characteristics of rotating machinery.

[0004] Many studies use contact sensors, such as displacement, velocity, and acceleration sensors. Although these sensors have high sensitivity and a wide frequency response range, they also have certain limitations: 1) Long-term installation of contact sensors will cause wear on the equipment surface and are not applicable in many industrial scenarios; 2) Rotating machinery is usually in extreme working environments such as high temperature and high pressure, which places high demands on contact sensors; 3) The signals collected by contact sensors are usually concentrated at a single measurement point. In order to reflect the large range of the machinery monitoring area, multiple sensors are required. At the same time, the installation location also needs to be appropriately selected.

[0005] To overcome the limitations of contact sensors, a new type of non-contact dynamic vision sensor—the event camera—is beginning to be applied in monitoring and diagnostics. While standard cameras generate image frames containing brightness information for all pixels within their field of view at a constant sampling frequency, event cameras capture relevant information from pixels whose brightness changes exceed a specified threshold and output an asynchronous event stream. This signal acquisition method endows event cameras with unique dynamic characteristics, such as high temporal resolution, wide dynamic range, and low power consumption, making them ideal for capturing dynamic visual signals. Consequently, event cameras have enormous potential for application in industrial fields such as edge tracking, autonomous driving, and high-speed positioning.

[0006] While event cameras offer numerous advantages, they also present challenges in microvibration measurement and fault diagnosis applications. Event data primarily captures the pixel location and corresponding time of vibration, but lacks detailed characteristic information. Therefore, event data cannot accurately reflect the precise vibration of the equipment. Furthermore, events occur at random times and locations. Therefore, converting event data into analyzable and practical microvibration signals and applying these estimates to fault diagnosis is a complex and difficult task. Summary of the Invention

[0007] In order to overcome the defects of the above-mentioned prior art, the purpose of the present invention is to provide a non-contact micro-vibration measurement and fault diagnosis method based on dynamic visual signal alignment, which uses an event camera to collect event data of rotating machinery vibration and extract the time domain micro-vibration signal of each pixel point from it. After signal alignment and feature clustering of the extracted signals, a deep neural network is used to realize the fault classification of the rotating machinery, thereby improving the applicability of non-contact sensors in the field of vibration monitoring and fault diagnosis of rotating machinery.

[0008] In order to achieve the above object, the technical solution adopted by the present invention is:

[0009] A non-contact micro-vibration measurement and fault diagnosis method based on dynamic visual signal alignment first uses a dynamic visual sensor - an event camera to non-contactly collect event data of rotating machinery vibration, and uses an event data characterization method to convert the event stream into an equally spaced, non-overlapping and continuous event frame sequence; then, a deep neural network is used to extract the micro-vibration information of each pixel point in the event frame sequence, and the difference in vibration signals of adjacent pixels is narrowed by a signal alignment method. The contact sensing data is used as a reference for optimizing model performance through a feature clustering method; finally, the obtained micro-vibration data is used for intelligent fault diagnosis.

[0010] A non-contact micro-vibration measurement and fault diagnosis method based on dynamic visual signal alignment includes the following steps:

[0011] Step 1, event data acquisition: using an event camera to collect asynchronous event streams of rotating machinery vibration;

[0012] Step 2, event data representation: First, the event stream is reconstructed into an event vibration sequence, which is then used as a sample for extracting vibration signals. The collected event stream data is reconstructed using a two-dimensional representation method for event data. A segment of event stream data is divided into equal-length, non-overlapping, and continuous sub-event stream data. In the sub-event stream data, for each pixel, the polarity of each event is superimposed, and the sub-event stream is compressed into a two-dimensional event frame. The entire event stream is reconstructed into an event frame sequence, specifically represented as where p irepresents the polarity of the event, (x, y) represents the location of the event, Ω represents the two-dimensional field of view of the event frame, and n d Indicates the specified time period t d The total number of events at that pixel;

[0013] The data of the contact sensor is used to calibrate the vibration signal extracted from the event frame sequence and perform feature difference analysis. First, the data of the contact sensor corresponding to each operating state of the rotating machinery is collected, and then the relevant features are extracted as calibration labels. The i-th event data sample is represented as s i =(r i ,h i ,j i ), where r i =[r i1 , r i2 ,…,r in ] represents an event frame sequence, n represents the number of consecutive event frames contained in the event frame sequence; h i Indicates the machine health status label; j i A calibration label representing the acceleration signal characteristics under the corresponding state;

[0014] Step 3: Establish a model of a micro-vibration monitoring and fault diagnosis network framework based on event signals. The model consists of two parts: a vibration signal extractor and a fault pattern recognizer. The vibration signal extractor extracts the time domain vibration signal of each pixel from the input event frame sequence, and then feeds it into the fault pattern recognizer to identify its fault mode.

[0015] Step 4, vibration signal alignment: Assume that the object monitored by the event camera is a rigid body, and each pixel vibrates synchronously with the adjacent pixels. The time domain vibration signals of these pixels are similar. This assumption is used as the basis for the deviation correction of the time domain vibration signals of adjacent pixels.

[0016]

[0017] Among them, L d Indicates the corresponding loss; (x, y) indicates the coordinates of the collected pixels; S (x,y) Represents the time domain vibration signal corresponding to the selected pixel point; Represents the time domain vibration signal of adjacent pixel points; a and b represent an even number of adjacent points in the X and Y directions respectively;

[0018] Step 5, vibration feature clustering: For the samples in the training set, calculate the time domain features of their vibration signals and cluster them according to the working conditions. The cluster center is defined as the time domain features calculated based on the data of the contact sensor. The intra-class loss of each class is calculated based on the cluster center, and the losses of all classes are aggregated to form the final cluster loss.

[0019]

[0020] Q j =[q j1 ,q j2 ,…,q jr ]

[0021] q jr =[A,B,C]

[0022]

[0023]

[0024] Where, L e represents the total clustering loss, w i represents the intra-class loss of each class, that is, the Euclidean distance between the feature vector of the samples contained in each class and the central feature vector; K represents the total number of classes; Q acc represents the central eigenvector; N e Indicates the total number of samples in the class; Q j Represents the feature vector, which contains the time domain feature q of each vibration signal jr ; Among them, A, B and C represent the peak index, margin index and skewness index in the time domain characteristics respectively, and these indexes are dimensionless parameters;

[0025] Step 6, model training: Initialize the network parameters and input the training samples into the model to calculate the classification loss, alignment loss, and clustering loss to optimize the network parameters; after each epoch of training is completed, discard the samples of that epoch until the model converges; finally, input unlabeled test data into the model to evaluate the test performance.

[0026] Step 1: For a single pixel, the event camera records the brightness of the pixel at each moment with a resolution of microseconds and compares it with the brightness at the previous moment:

[0027] ±C=logI(x,y,t)-logI(x,y,t-Δt)

[0028] Where C represents contrast sensitivity, i.e., the threshold; I(x,y,t) represents the brightness of the pixel at coordinate (x,y) in the two-dimensional resolution region at time t; Δt represents the time interval; when the logarithmic change in brightness exceeds the threshold, the event camera will mark and output the event associated with the pixel; where the change exceeds the positive threshold, a positive event is output; where the change exceeds the negative threshold, a negative event is output; the format of the event data is a four-element vector e = [t,x,y,p], where t represents the timestamp; x and y represent the horizontal and vertical coordinates of the pixel in the resolution region, respectively; p represents the polarity of the event: p = +1 represents that the logarithmic change in brightness exceeds the positive threshold, and p = -1 represents that the logarithmic change in brightness exceeds the negative threshold; the event stream represents the set of all events captured by the event camera over a period of time, expressed as Among them, e i represents the i-th event, n e is the time period t e The total number of events recorded in the .

[0029] Step 3 introduces Gabor transform to extract texture features of different scales and directions in the image, thereby obtaining more accurate local motion information; Gabor transform is based on Gabor function, which is the product of complex sine wave and Gaussian function, specifically expressed as,

[0030]

[0031] Where g(x,y,λ,θ,ψ,σ,γ) represents the Gabor filter; (x,y) is the two-dimensional coordinate of the selected pixel; (x′,y′) is the two-dimensional coordinate after rotation; λ is the wavelength of the sinusoidal plane wave, θ is the direction of the parallel stripes of the filter; ψ is the phase offset, σ is the standard deviation of the Gaussian kernel, and γ is the spatial aspect ratio;

[0032] Assuming that the intensity of the image frame at the moment is E(x, y, t), the image frame is convolved with the Gabor filter to obtain the frequency domain information of the image frame:

[0033]

[0034] Where e(x,y,t) represents the frequency domain information of the image frame;

[0035] In the time period Δt, if there is local motion at a certain location (x, y) in the image, the intensity of the image frame before and after the motion is E(x, y, t) and E(x+Δx, y+Δy, t+Δt), respectively. Assuming that the parallel stripe direction of the Gabor filter is 0°, the frequency domain information before and after the motion is expressed in double integral form as follows:

[0036]

[0037] By analyzing the above two equations, we can get the phase difference The relationship between it and the horizontal displacement Δx is:

[0038]

[0039] A hybrid method of Gabor transform and deep neural network is adopted. The convolution kernel generated by Gabor transform replaces the convolution kernel of the first layer of deep neural network to extract the underlying features of edge and texture of the image. For each event frame sequence sample of the feature extractor, the first layer of the network adopts a Gabor filter with fixed parameters to perform layered convolution on each event frame in the sample. Each event frame corresponds to a Gabor filter, so that the number of channels and event frame order of the sample after convolution are the same as the number of channels and event frame order of the sample before convolution. Then the texture features are input into the deep neural network for more detailed feature extraction. The deep neural network contains two residual blocks, each of which adopts the convolution kernel k s The samples are layered convolved and activated using the LeakyReLU function. Next, the texture features are superimposed with the features output by the deep neural network to achieve subtle adjustments to the texture features. Ultimately, a sequence of frames is obtained, where each frame contains the vibration characteristics of each pixel.

[0040] Based on the extracted vibration features, the time-domain vibration signals of all pixels are obtained, thereby realizing the vibration estimation of multiple pixels of the rotating machinery. First, a pixel point is selected, and this pixel point needs to be a pixel point within the rotating machinery area in the frame. Second, the pixel values ​​in each frame of the frame sequence are extracted to form a time series. The first point is determined as the reference point, and the other points are calibration points. Then, the difference between each calibration point and the reference point is calculated in sequence to realize the relative displacement of the pixel points at these two moments. Through iteration, the time-domain vibration signal of the pixel point is finally obtained.

[0041] The obtained time-domain vibration signal is input into the fault pattern recognizer, which consists of a one-dimensional deep neural network containing two residual blocks. The convolution layer in each residual block uses a convolution kernel of length 9. The first residual block increases the dimension of the signal to 20 dimensions to extract its high-dimensional features and reduces its length by half. The second residual block further reduces its length by half. After flattening, the sample features are connected to two fully connected layers, each with 64 neurons. The number of neurons indicates the number of relevant health conditions. Finally, the softmax function is used for classification.

[0042] Compared with the prior art, the present invention has the following beneficial effects:

[0043] This paper proposes a non-contact micro-vibration measurement and fault diagnosis method based on dynamic visual signal alignment. By analyzing visual data captured by an event camera, the method extracts a time-domain vibration signal from each pixel in the event stream. This extracted time-domain signal is then used to identify the health status of the machine, achieving non-contact, global vibration monitoring and fault diagnosis. To enhance the robustness of the model, a signal alignment method is used to minimize the differences in vibration signals between adjacent pixels. Furthermore, a feature clustering method is introduced, using contact sensing data as a reference to optimize model performance. Experimental results demonstrate that the proposed method effectively extracts vibration features from dynamic visual data under various operating conditions and accurately identifies machine fault types, suggesting a promising future for industrial applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Schematic diagram of a flow chart of an embodiment of the present invention.

[0045] Figure 2 Schematic diagram of a vibration frequency domain signal converted from an event signal according to an embodiment of the present invention.

[0046] Figure 3 4 is a confusion matrix diagram of the embodiment of the present invention and the comparative method.

[0047] Figure 4 TSNE visualization diagrams of the embodiment of the present invention and the comparative method. DETAILED DESCRIPTION

[0048] The present invention is described in further detail below with reference to the embodiments and accompanying drawings.

[0049] Reference Figure 1 , a non-contact micro-vibration measurement and fault diagnosis method based on dynamic visual signal alignment, comprising the following steps:

[0050] Step 1: Event data acquisition: Use an event camera to capture the asynchronous event stream of rotating machinery vibration. For a single pixel, the event camera records the brightness of the pixel at each moment with microsecond resolution and compares it with the brightness at the previous moment:

[0051] ±C=logI(x,y,t)-logI(x,y,t-Δt)

[0052] Where C represents contrast sensitivity, i.e., the threshold; I(x,y,t) represents the brightness of the pixel at coordinate (x,y) in the two-dimensional resolution region at time t; Δt represents the time interval; when the logarithmic change of brightness exceeds the threshold, the event camera will mark and output the event associated with the pixel; where the change exceeds the positive threshold, a positive event is output; where the change exceeds the negative threshold, a negative event is output; the format of the event data is a four-element vector e = [t,x,y,p], where t represents the timestamp; x and y represent the horizontal and vertical coordinates of the pixel in the resolution region, respectively; p represents the polarity of the event: p = +1 represents that the logarithmic change of brightness exceeds the positive threshold, and p = -1 represents that the logarithmic change of brightness exceeds the negative threshold; for the event stream, it represents the set of all events captured by the event camera over a period of time, expressed as Among them, e i represents the i-th event, n e is the time period t e The total number of events recorded in

[0053] Step 2, event data representation: It is difficult to extract vibration information directly from the event stream, because these data do not represent the actual vibration signal of the rotating machinery, and the event data cannot be directly divided into samples. To address this problem, the event stream is first reconstructed into an event vibration sequence, and then used as a sample to extract the vibration signal; specifically, the collected event stream data is reconstructed using the two-dimensional representation method of event data. A section of event stream data is divided into equal-length, non-overlapping, and continuous sub-event stream data; in the sub-event stream data, for each pixel, the polarity of each event is superimposed, so the sub-event stream is compressed into a two-dimensional event frame; and the entire event stream is reconstructed into an event frame sequence, specifically expressed as where p i represents the polarity of the event, (x, y) represents the location of the event, Ω represents the two-dimensional field of view of the event frame, and n d Indicates the specified time period t d The total number of events at the pixel within the event stream; this representation method can effectively integrate the event stream in the time and space dimensions, making the characteristics of the event stream more prominent. At the same time, it also converts complex event stream data into typical frame data, which facilitates the subsequent extraction of micro-vibration signals from the event frame;

[0054] In machine condition monitoring, contact vibration sensors such as accelerometers usually have accurate measurements and stable performance. Therefore, the data from contact sensors can be used to calibrate the vibration signals extracted from the event frame sequence. Since the two signals do not correspond one to one, feature difference analysis is required. Taking the acceleration data of the accelerometer as an example, the acceleration data corresponding to each operating state of the machine is first collected, and then the relevant features are extracted as calibration labels. Therefore, the i-th event data sample can be expressed as s i =(ri ,h i ,j i ), where r i =[r i1 , r i2 ,…,r in ] represents an event frame sequence, n represents the number of consecutive event frames contained in the event frame sequence; h i Indicates the machine health status label. i A calibration label representing the acceleration signal characteristics under the corresponding state;

[0055] Step 3: Establish a model of a micro-vibration monitoring and fault diagnosis network framework based on event signals. The model mainly consists of two parts: a vibration signal extractor and a fault pattern recognizer. The vibration signal extractor extracts the time domain vibration signal of each pixel from the input event frame sequence, and then feeds it into the fault pattern recognizer to identify its fault mode.

[0056] In order to extract the local motion phase from the event frame sequence, the Gabor transform is introduced. The Gabor transform is a short-time windowed Fourier transform with good time domain localization and frequency domain bandpass characteristics. It can well extract texture features of different scales and directions in the image, thereby obtaining more accurate local motion information. The Gabor transform has a wide range of applications in image processing, pattern recognition and other fields. The Gabor transform is based on the Gabor function, which is the product of a complex sine wave and a Gaussian function. It is specifically expressed as,

[0057]

[0058] where g(x,y,λ,θ,ψ,σ,γ) represents the Gabor filter; (x,y) are the two-dimensional coordinates of the selected pixel; (x′,y′) are the two-dimensional coordinates after rotation; λ is the wavelength of the sinusoidal plane wave, θ is the direction of the parallel stripes of the filter; ψ is the phase offset, σ is the standard deviation of the Gaussian kernel, and γ is the spatial aspect ratio.

[0059] Assuming that the intensity of the image frame at the moment is E(x, y, t), the image frame is convolved with the Gabor filter to obtain the frequency domain information of the image frame:

[0060]

[0061] Where e(x,y,t) represents the frequency domain information of the image frame;

[0062] In the time period Δt, if there is local motion at a certain location (x, y) in the image, the intensity of the image frame before and after the motion is E(x, y, t) and E(x+Δx, y+Δy, t+Δt), respectively. Assuming that the parallel stripe direction of the Gabor filter is 0°, the frequency domain information before and after the motion is expressed in double integral form as follows:

[0063]

[0064] By analyzing the above two equations, we can get the phase difference The relationship between it and the horizontal displacement Δx is:

[0065]

[0066] In order to improve the performance of Gabor transform, a hybrid method of Gabor transform and deep neural network is adopted. The convolution kernel generated by Gabor transform replaces the convolution kernel of the first layer of deep neural network, which can more effectively extract the underlying features of the image such as edges and textures, thereby improving the perception ability. Specifically, for each event frame sequence sample of the feature extractor, the first layer of the network uses a Gabor filter with fixed parameters to perform layered convolution on each event frame in the sample. Each event frame corresponds to a Gabor filter, so that the number of channels and event frame order of the sample after convolution are the same as the number of channels and event frame order of the sample before convolution. Then the texture features are input into the deep neural network for more detailed feature extraction. The deep neural network contains two residual blocks, each of which uses the convolution kernel k s The samples are layered convolved and activated using the LeakyReLU function. Next, the texture features are superimposed with the features output by the deep neural network to achieve subtle adjustments to the texture features. Ultimately, a sequence of frames is obtained, where each frame contains the vibration characteristics of each pixel.

[0067] Based on the extracted vibration features, the time-domain vibration signals of all pixels can be obtained, thereby realizing the vibration estimation of multiple pixel points of the rotating machinery. First, a suitable pixel point is selected, which needs to be a pixel point within the rotating machinery area in the frame. Second, the pixel values ​​in each frame of the frame sequence are extracted to form a time series. The first point is determined as the reference point, and the other points are calibration points. Then, the difference between each calibration point and the reference point is calculated in sequence. After that, the relative displacement of the pixel points at these two moments is realized. Through iteration, the time-domain vibration signal of the pixel point is finally obtained.

[0068] The resulting time-domain vibration signal is fed into a fault pattern recognizer, which consists of a one-dimensional deep neural network consisting of two residual blocks. Each residual block uses a convolutional kernel of length 9. The first residual block increases the signal's dimensionality to 20 to extract high-dimensional features and reduces its length by half. The second residual block further reduces its length by half. After flattening, the sample features are connected to two fully connected layers, each with 64 neurons, where the number of neurons represents the number of health conditions. Finally, a softmax function is used for classification.

[0069] Step 4, vibration signal alignment: Assume that the object monitored by the event camera is a rigid body, and each pixel vibrates synchronously with the adjacent pixels. The time domain vibration signals of these pixels are similar. This assumption is used as the basis for the deviation correction of the time domain vibration signals of adjacent pixels.

[0070]

[0071] Among them, L d Indicates the corresponding loss; (x, y) indicates the coordinates of the collected pixels; S (x,y) Represents the time domain vibration signal corresponding to the selected pixel point; Represents the time domain vibration signal of adjacent pixel points; a and b represent an even number of adjacent points in the X and Y directions respectively;

[0072] Step 5, vibration feature clustering: For different samples collected under the same working condition and healthy state, the time domain signals obtained through the network will be different, and it is very difficult to calculate the error of each point. However, the time domain features of these signals have certain similarities, so these features can be clustered to reduce the intra-class loss. Specifically, for the samples in the training set, the time domain features of their vibration signals are calculated and clustered according to the working conditions. The cluster center is defined as the time domain features calculated based on the accelerometer acquisition signal of this category. The intra-class loss of each class is calculated based on the cluster center, and the losses of all classes are aggregated to form the final cluster loss.

[0073]

[0074] Q j =[q j1 ,q j2 ,…,q jr ]

[0075] q jr =[A,B,C]

[0076]

[0077] Where, L e represents the total clustering loss, w irepresents the intra-class loss of each class, that is, the Euclidean distance between the feature vector of the samples contained in each class and the central feature vector; K represents the total number of classes; Q acc represents the central eigenvector; N e Indicates the total number of samples in the class; Q j Represents the feature vector, which contains the time domain feature q of each vibration signal jr ; Where A, B, and C represent the peak index, margin index, and skew index in the time domain characteristics, respectively. These indices are dimensionless parameters;

[0078] Step 6, model training: Initialize the network parameters and input the training samples into the model to calculate the classification loss, alignment loss, and clustering loss to optimize the network parameters; after each epoch of training is completed, discard the samples of that epoch until the model converges; finally, input unlabeled test data into the model to evaluate the test performance.

[0079] The embodiment takes a rolling bearing in a rotating machine as an example and verifies the effectiveness of the method of the present invention based on rolling bearing experimental data.

[0080] The experimental platform uses a motor to drive the shaft through a coupling, and the shaft is equipped with an ER-16K rolling bearing. The parameters are shown in Table 1. An event camera is placed in front of the bearing to collect bearing vibration event data. This experiment uses a Prophesee 3.1 event camera, and the parameters are shown in Table 2. An accelerometer is placed above the bearing to synchronously collect acceleration vibration signals for comparative analysis. The sampling rate of the accelerometer is 12.8kHz. The experiment considers four typical rolling bearing health states: healthy, outer ring fault, inner ring fault, and rolling element fault. Experimental data is collected at five different speeds. The specific experimental conditions are shown in Table 3. The hyperparameters of the model in this embodiment are shown in Table 4.

[0081] Table 1 Rolling bearing parameters

[0082]

[0083] Table 2 Event camera parameters

[0084]

[0085]

[0086] Table 3 Experimental conditions

[0087]

[0088] Table 4 Model hyperparameter settings

[0089]

[0090] Various methods are evaluated to illustrate the effectiveness and superiority of the proposed fault diagnosis approach within the cross-modal learning framework:

[0091] 1) NoA method (no alignment): In order to evaluate the advantages of the proposed feature alignment method, the NoA method is implemented. In this case, the feature alignment method is not considered, thereby ignoring the alignment loss of vibration signals at different pixels of the same sample during the model optimization process. All other parameters remain the same as the proposed method.

[0092] 2) NoC method (no clustering): The performance enhancement of the feature clustering method is demonstrated by the NoC method, in which the feature clustering method is not used during model optimization, only the feature alignment method is used, while all other parameters are generally kept the same as those of the proposed method.

[0093] 3) NoAC methods (No Alignment and Clustering): Deep neural network methods represent the standard approach for fault diagnosis modeling, following the supervised learning paradigm, in which only the cross-entropy loss function of dynamic visual data is considered in model optimization, while feature alignment and clustering methods are excluded from the proposed network architecture.

[0094] 4) NoGF method (without Gabor filter): In order to evaluate the effectiveness of Gabor filter for event frame vibration feature extraction, the Gabor filter was removed from the comparison method, and the first layer of the feature extractor was replaced with a standard convolution kernel, which was randomly initialized and continuously optimized throughout the training process, instead of using a Gabor filter with fixed parameters.

[0095] The time domain micro-vibration signal extracted by the embodiment is converted into a frequency domain signal through Fourier transform, and its spectrum is as follows: Figure 2 As shown, for different rotation speeds, the rotation frequency of the device and its harmonics can be clearly found in the spectrum, indicating that the embodiment method can effectively extract the micro-vibration signal of each pixel from the event data, laying the foundation for the subsequent identification of different fault modes.

[0096] Table 5 Classification accuracy of different methods

[0097]

[0098] Table 5 lists the overall experimental results for various fault diagnosis tasks using different methods. It can be seen that the embodiment method achieved a test accuracy exceeding 98%. In contrast, alternative event data processing methods (NoA, NoC, and NoAC) showed competitive performance, with test accuracy generally exceeding 90%. However, these methods showed a slight decrease in accuracy compared to the embodiment method, indicating that time-domain signal alignment and feature clustering help improve model performance. In contrast, NoGF achieved a relatively low classification accuracy of less than 50%, indicating that basic deep neural networks are unable to effectively capture real-world time-domain vibration signals. These results confirm the feasibility of using vibration signals extracted from dynamic visual data for fault diagnosis.

[0099] In order to show the detailed test results of various machine failures, Figure 3 The confusion matrices of the embodiment method and the comparison method are shown. The comparison shows that without the Gabor filter, the model can only accurately identify rolling element faults at 30Hz and 40Hz, and cannot distinguish other fault modes and healthy states. In contrast, the embodiment method has an identification accuracy of more than 98% for each fault mode, exceeding the performance of other comparison methods in the corresponding category.

[0100] The T-SNE technique is used to visualize and evaluate the features learned by various methods, with a special focus on the data representation of the fully connected layer of the fault pattern recognizer. The results are shown in Figure 2. Figure 4 As shown in the visualization results generated by the NoGF method, it can be clearly seen that most samples are clustered in the same area with minimal separation. Although the NoAC method exhibits high accuracy, its cross-category clustering effect is less than ideal, with significant overlap at category boundaries. In contrast, the embodiment method shows significantly improved clustering results for samples representing different machine health states, with clearer category boundaries and enhanced separation. This demonstrates the effectiveness and superiority of the embodiment method in identifying machine failure modes based on event data.

Claims

1. A non-contact micro-vibration measurement and fault diagnosis method based on dynamic visual signal alignment, characterized by: First, a dynamic visual sensor—an event camera—is used to non-contactly collect event data of rotating machinery vibration. Event data characterization methods are used to convert the event stream into an evenly spaced, non-overlapping, and continuous sequence of event frames. A deep neural network is then used to extract micro-vibration information from each pixel in the event frame sequence. Signal alignment methods are used to narrow the differences in vibration signals between adjacent pixels. Feature clustering methods are used to use the contact sensing data as a reference for optimizing model performance. Finally, the resulting micro-vibration data is used for intelligent fault diagnosis. The method comprises the following steps: Step 1, event data acquisition: using an event camera to collect asynchronous event streams of rotating machinery vibration; Step 2, event data representation: First, the event stream is reconstructed into an event vibration sequence, which is then used as a sample for extracting vibration signals. The collected event stream data is reconstructed using a two-dimensional representation method for event data. A segment of event stream data is divided into equal-length, non-overlapping, and continuous sub-event stream data. In the sub-event stream data, for each pixel, the polarity of each event is superimposed, and the sub-event stream is compressed into a two-dimensional event frame. The entire event stream is reconstructed into an event frame sequence, specifically represented as where p i represents the polarity of the event, (x, y) represents the location of the event, Ω represents the two-dimensional field of view of the event frame, and n d Indicates the specified time period t d The total number of events at that pixel; The data of the contact sensor is used to calibrate the vibration signal extracted from the event frame sequence and perform feature difference analysis. First, the data of the contact sensor corresponding to each operating state of the rotating machinery is collected, and then the relevant features are extracted as calibration labels. The i-th event data sample is represented as s i =(r i ,h i ,j i ), where r i =[r i1 , r i2 ,…,r in ] represents an event frame sequence, n represents the number of consecutive event frames contained in the event frame sequence; h i Indicates the machine health status label; j i A calibration label representing the acceleration signal characteristics under the corresponding state; Step 3: Establish a model of a micro-vibration monitoring and fault diagnosis network framework based on event signals. The model consists of two parts: a vibration signal extractor and a fault pattern recognizer. The vibration signal extractor extracts the time domain vibration signal of each pixel from the input event frame sequence, and then feeds it into the fault pattern recognizer to identify its fault mode. Step 4, vibration signal alignment: Assume that the object monitored by the event camera is a rigid body, and each pixel vibrates synchronously with the adjacent pixels. The time domain vibration signals of these pixels are similar. This assumption is used as the basis for the deviation correction of the time domain vibration signals of adjacent pixels. Among them, L d Indicates the corresponding loss; (x, y) indicates the coordinates of the collected pixels; S (x,y) Represents the time domain vibration signal corresponding to the selected pixel point; Represents the time domain vibration signal of adjacent pixel points; a and b represent an even number of adjacent points in the X and Y directions respectively; Step 5, vibration feature clustering: For the samples in the training set, calculate the time domain features of their vibration signals and cluster them according to the working conditions. The cluster center is defined as the time domain features calculated based on the data of the contact sensor. The intra-class loss of each class is calculated based on the cluster center, and the losses of all classes are aggregated to form the final cluster loss. Q j =[q j1 ,q j2 ,…,q jr ] q jr =[A,B,C] Where, L e represents the total clustering loss, w i represents the intra-class loss of each class, that is, the Euclidean distance between the feature vector of the samples contained in each class and the central feature vector; K represents the total number of classes; Q acc represents the central eigenvector; N e Indicates the total number of samples in the class; Q j Represents the feature vector, which contains the time domain feature q of each vibration signal jr ; Among them, A, B and C represent the peak index, margin index and skewness index in the time domain characteristics respectively, and these indexes are dimensionless parameters; Step 6, model training: Initialize the network parameters and input the training samples into the model to calculate the classification loss, alignment loss, and clustering loss to optimize the network parameters; after each epoch of training is completed, discard the samples of that epoch until the model converges; finally, input unlabeled test data into the model to evaluate the test performance.

2. The method according to claim 1, wherein: Step 1: For a single pixel, the event camera records the brightness of the pixel at each moment with a resolution of microseconds and compares it with the brightness at the previous moment: ±C=logI(x,y,t)-logI(x,y,t-Δt) Where C represents contrast sensitivity, i.e., the threshold; I(x,y,t) represents the brightness of the pixel at coordinate (x,y) in the two-dimensional resolution region at time t; Δt represents the time interval; when the logarithmic change in brightness exceeds the threshold, the event camera will mark and output the event associated with the pixel; where the change exceeds the positive threshold, a positive event is output; where the change exceeds the negative threshold, a negative event is output; the format of the event data is a four-element vector e = [t,x,y,p], where t represents the timestamp; x and y represent the horizontal and vertical coordinates of the pixel in the resolution region, respectively; p represents the polarity of the event: p = +1 represents that the logarithmic change in brightness exceeds the positive threshold, and p = -1 represents that the logarithmic change in brightness exceeds the negative threshold; the event stream represents the set of all events captured by the event camera over a period of time, expressed as Among them, e i represents the i-th event, n e is the time period t e The total number of events recorded in the .

3. The method according to claim 1, wherein: Step 3 introduces Gabor transform to extract texture features of different scales and directions in the image, thereby obtaining more accurate local motion information; Gabor transform is based on Gabor function, which is the product of complex sine wave and Gaussian function, specifically expressed as, Where g(x,y,λ,θ,ψ,σ,γ) represents the Gabor filter; (x,y) is the two-dimensional coordinate of the selected pixel; (x′,y′) is the two-dimensional coordinate after rotation; λ is the wavelength of the sinusoidal plane wave, θ is the direction of the parallel stripes of the filter; ψ is the phase offset, σ is the standard deviation of the Gaussian kernel, and γ is the spatial aspect ratio; Assuming that the intensity of the image frame at the moment is E(x, y, t), the image frame is convolved with the Gabor filter to obtain the frequency domain information of the image frame: Where e(x,y,t) represents the frequency domain information of the image frame; In the time period Δt, if there is local motion at a certain location (x, y) in the image, the intensity of the image frame before and after the motion is E(x, y, t) and E(x+Δx, y+Δy, t+Δt), respectively. Assuming that the parallel stripe direction of the Gabor filter is 0°, the frequency domain information before and after the motion is expressed in double integral form as follows: By analyzing the above two equations, we can get the phase difference The relationship between it and the horizontal displacement Δx is: A hybrid method of Gabor transform and deep neural network is adopted. The convolution kernel generated by Gabor transform replaces the convolution kernel of the first layer of deep neural network to extract the underlying features of edge and texture of the image. For each event frame sequence sample of the feature extractor, the first layer of the network adopts a Gabor filter with fixed parameters to perform layered convolution on each event frame in the sample. Each event frame corresponds to a Gabor filter, so that the number of channels and event frame order of the sample after convolution are the same as the number of channels and event frame order of the sample before convolution. Then the texture features are input into the deep neural network for more detailed feature extraction. The deep neural network contains two residual blocks, each of which adopts the convolution kernel k s The samples are layered convolved and activated using the LeakyReLU function. Next, the texture features are superimposed with the features output by the deep neural network to achieve subtle adjustments to the texture features. Ultimately, a sequence of frames is obtained, where each frame contains the vibration characteristics of each pixel. Based on the extracted vibration features, the time-domain vibration signals of all pixels are obtained, thereby realizing the vibration estimation of multiple pixels of the rotating machinery. First, a pixel point is selected, and this pixel point needs to be a pixel point within the rotating machinery area in the frame. Second, the pixel values ​​in each frame of the frame sequence are extracted to form a time series. The first point is determined as the reference point, and the other points are calibration points. Then, the difference between each calibration point and the reference point is calculated in sequence to realize the relative displacement of the pixel points at these two moments. Through iteration, the time-domain vibration signal of the pixel point is finally obtained. The obtained time-domain vibration signal is input into the fault pattern recognizer, which is composed of a one-dimensional deep neural network. The network contains two residual blocks. The convolution layer in each residual block adopts a convolution kernel of length 9. The first residual block increases the dimension of the signal to 20 dimensions to extract its high-dimensional features and reduces its length by half. The second residual block further reduces its length by half. After flattening, the sample features are connected to two fully connected layers, each with 64 neurons. The number of neurons represents the number of related health conditions. Finally, the softmax function is used for classification.

Citation Information

Patent Citations

  • CNN color characteristic pattern-based method for identifying fault of bearing under rated operation

    CN110261108A

  • Non-contact mechanical vibration monitoring and fault diagnosis method based on event camera

    CN116734980A