Video micro-vibration signal correction method
By hybridizing the CNN-LSTM-FFT deep learning architecture and combining time domain and frequency domain features, a multimodal fusion calibration model was constructed to solve the problem of low precision in video micro-vibration signal monitoring and achieve high-precision correction of video micro-vibration signals.
Patent Information
- Application Number
- CN202510861710.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing video micro-vibration signal monitoring methods are affected by environmental factors, noise and on-site operations in actual engineering applications, resulting in low accuracy and distortion of time domain and frequency domain information extraction, making it difficult to achieve high-precision structural vibration monitoring.
A hybrid CNN-LSTM-FFT deep learning architecture is adopted, time domain and frequency domain features are combined, a multimodal fusion calibration model is constructed, and signal correction is performed using the Fourier spectrum logarithmic amplitude difference to achieve dual correction of video micro-vibration signals in the time domain and frequency domain.
The data accuracy of video micro-vibration signal monitoring is improved, and the accuracy of non-contact video monitoring is improved through point sensors, ensuring the consistency of key modal frequencies and signal correction effects.
Smart Images

Figure CN120705481A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and in particular relates to a method for correcting video micro-vibration signals. Background Art
[0002] Hydraulic structures such as dams, spillways, and aqueducts play a vital role in water resource management and infrastructure safety. During long-term operation, these structures are subject to complex dynamic loads caused by continuous hydraulic shock, sediment erosion, and operational vibration. Accumulated structural damage (such as concrete cracking and steel corrosion) and potential resonance caused by environmental excitations (such as flood discharges and traffic loads) can seriously compromise their operational safety. Structural health monitoring (SHM) systems are crucial for early damage detection and risk mitigation, ensuring the long-term safe operation of structures.
[0003] Traditional structural health monitoring systems rely primarily on point sensors (such as accelerometers, strain gauges, or displacement sensors) to measure responses at critical locations. However, deploying such sensors on large or geometrically complex hydraulic structures often results in sparse spatial distribution for economic reasons, and the limited number of measurement points is unable to capture detailed, full-field vibration patterns. In recent years, video-based non-contact vibration monitoring has emerged as a promising alternative, leveraging computer vision techniques to extract full-field displacement data from monitoring videos. In particular, video vibration information extraction techniques that combine the phase-based optical flow method with phase-based video magnification offer advantages such as enhanced robustness, the absence of structural markers, sub-pixel displacement resolution, and the ability to amplify micron-level vibrations visible to the naked eye. This theoretically achieves monitoring accuracy comparable to that of traditional point sensors, making this approach feasible for use in harsh or inaccessible environments and on large structures. Over the past decade, this method has been continuously refined, expanding the monitoring area from localized motion to the entire capture domain. Vibration information, such as time history, frequency spectrum, and structural modes, has evolved from requiring indirect conversion to being directly extracted from video, and the computational efficiency of the extraction process has also continued to improve. However, most current research has conducted experimental validation and method reliability assessments using simplified scaled structural models in laboratory settings. The monitored objects are generally small or have clear textures, and the influencing factors of the monitoring environment are minimized. The primary obstacle to the advancement of this method from theoretical (or carefully designed experimental environments) to practical applications is the reliability of the extracted structural vibration information under the influence of these combined factors. For example, existing techniques have investigated the potential for information distortion in the algorithm itself and combined deep learning or denoising algorithms to further improve vibration information extraction accuracy. Existing techniques have also analyzed the potential impact of the monitoring environment and, in real-world monitoring environments, have attempted to address data accuracy issues caused by inherent camera vibration or instability.
[0004] Previous studies have shown that distortion in extracting vibration information manifests itself in both the time and frequency domains. In the time domain, raw pixel motion extracted from video is susceptible to noise, scale drift, and low-frequency dominance, resulting in inaccurate time-domain waveforms. In the frequency domain, the spectrum is easily distorted due to limitations in frame rate, resolution, and algorithmic approximation. Commonly used methods for improving the accuracy of video vibration measurement often focus only on isolated domains. Time-domain methods (such as Kalman filtering and sliding averaging) struggle to recover high-frequency components lost during video acquisition. Frequency-domain methods (such as bandpass filtering) often destroy phase continuity, resulting in waveform distortion. Summary of the Invention
[0005] In response to the above-mentioned deficiencies in the prior art, the present invention provides a video micro-vibration signal correction method that solves the problem of low accuracy caused by environmental factors, noise, on-site operations, etc. when directly extracting micro-vibration information from monitoring videos.
[0006] In order to achieve the above objectives, the technical solution adopted by the present invention is: a video micro-vibration signal correction method, comprising: Video capture: Place the vibration sensor on the structure to be tested to capture the video of the structure to be tested; Feature extraction: Based on the video capture results, the time domain features and frequency domain features of the vibration signal are extracted; Constructing a multimodal fusion calibration model: Using time domain features and frequency domain features to train and construct a multimodal fusion calibration model. The multimodal fusion calibration model is a hybrid CNN-LSTM network architecture. Correction result output: A spectral loss term based on the Fourier spectrum logarithmic amplitude difference is introduced, and the constructed multimodal fusion calibration model is used to obtain the dual correction target results of time domain waveform reconstruction and frequency domain key mode alignment to complete the correction of the video micro-vibration signal.
[0007] The present invention has the following beneficial effects: It proposes a video micro-vibration signal correction method based on a hybrid CNN-LSTM-FFT deep learning architecture. This method targets signals acquired by point sensors and combines the physical characteristics of structural vibration with the signal's time-frequency domain information to achieve dual correction of video micro-vibration signals in both the time and frequency domains. First, the signal's time- and frequency-domain characteristics, such as displacement, velocity, acceleration, power spectral density (PSD), and short-term statistics, are extracted to construct a signal time-frequency feature matrix. Second, a hybrid CNN-LSTM model is trained. The CNN layer extracts local motion patterns and denoises the signal, while the LSTM layer captures long-term correlations and modal trends in the vibration sequence, learning the mapping relationship between the video-extracted signal and the sensor-acquired signal. Third, a spectral loss term based on the Fourier spectrum logarithmic amplitude difference is introduced to ensure consistency of key modal frequencies in the corrected signal. This method improves the data accuracy of non-contact video micro-vibration monitoring of structures using point sensors.
[0008] Furthermore, the time domain features include: Dynamic features:
[0009]
[0010] in, Indicates that the vibration signal is n The first-order difference of the sampling points, Indicates that the vibration signal isn The second-order difference of the sampling points, represents the displacement change of the vibration signal, represents the sampling time, 、 and Respectively represent the vibration signal in n 、 n -1. n -2 displacement of sampling points, represents the time step; Sliding window statistics:
[0011]
[0012]
[0013] in, represents the sliding mean, represents the sliding window length, i represents the local index sampling point within the sliding window, Indicates the i The displacement value of the sampling point, i∈[1,n], represents the variance, Indicates the window length The maximum value of the inner sampling point n; Cumulative displacement:
[0014] in, It represents the energy drift characteristics of the quantized signal through time accumulation; Autocorrelation function:
[0015] in, Indicates calculation lag q The signal self-similarity at step , Total signal length.
[0016] The beneficial effect of the above further solution is that, based on the calculation of the above time domain characteristics, time domain characteristic parameters are provided for signal calibration.
[0017] Furthermore, the frequency domain features include: Fourier amplitude spectrum:
[0018]
[0019]
[0020] in, represents the amplitude spectrum eigenvector, Represents the complex amplitude of the kth frequency point of discrete Fourier transform DFT, To obtain the frequency domain amplitude spectrum of the complex modulus value, K max Indicates the preset maximum effective frequency index, represents the windowed signal, j represents the imaginary unit, represents the Hamming window, N represents the total length of the signal, and n represents the sampling point; Power spectral density:
[0021] in, represents the power spectral density, T represents the sampling time, represents the energy in the time domain, Represents the transpose of x after Fourier transformation.
[0022] The beneficial effect of the above further solution is that, through the above calculation of the frequency domain characteristics, frequency domain characteristic parameters are provided for signal calibration.
[0023] Furthermore, the multimodal fusion calibration model is constructed by training using time domain features and frequency domain features, which is specifically as follows: Deep features of vibration signals are automatically extracted by using a hierarchical one-dimensional convolutional neural network (1D CNN); Fusion of vibration signal depth features with time domain features and frequency domain features; Based on the fused feature vector, a two-layer long short-term memory network (LSTM) is used to establish a long-term dependency branch model to model the long-term dependency relationship, so as to learn the high-dimensional nonlinear mapping relationship between the video extraction signal and the point sensor acquisition signal, and complete the construction of the multimodal fusion calibration model.
[0024] The beneficial effect of the above further solution is: using 1D CNN to identify local details and global trends of the signal respectively, and then using LSTM to solve the problem that traditional CNN is insufficient in modeling periodic features.
[0025] Furthermore, the vibration signal deep features automatically extracted by using the hierarchical one-dimensional convolutional neural network 1D CNN are specifically: For the input vibration displacement signal sequence, a symmetrically padded one-dimensional convolution kernel is used in the short-time dynamic feature extraction layer to capture high-frequency components and transient impacts through local weighted summation and nonlinear mapping: Based on the captured high-frequency components and transient impacts, batch normalization is performed to obtain normalized high-frequency transient features; In the long-term trend modeling layer, the low-frequency envelope and slow-varying drift features are extracted through the diffusion convolution kernel and regularization to obtain the low-frequency drift features. The deep features of the vibration signal include high-frequency transient features and low-frequency drift features. The deep features of the vibration signal are spliced with the extracted time domain features and frequency domain features along the channel dimension to obtain a fused feature vector, wherein the fused feature vector is input into a two-layer long short-term memory network LSTM.
[0026] Furthermore, the expression of the high-frequency transient characteristic is as follows:
[0027]
[0028]
[0029] in, Represents high-frequency transient characteristics, and Both represent affine transformation parameters, represents the 𝑘th eigenvalue of the lth convolution output, represents the variance, represents the mean, represents a numerical stability constant, d represents the feature index dimension, D represents a dataset, represents the convolution kernel weight, represents the linear rectification activation function, Represents the effective frequency index, c represents the number of convolution channels, Represents the input vibration signal, Indicates time, represents the bias term of the first layer; N represents the total length of the signal; The expression of the low-frequency drift characteristic is as follows:
[0030] in, Represents the low-frequency drift characteristics, represents the regularization operation, represents a trainable weight tensor, represents the input feature, that is, the high-frequency transient feature, Represents the bias term of the second layer.
[0031] The beneficial effect of the above further solution is that the accuracy of the high-frequency range is improved through the above calculation.
[0032] Furthermore, the expression of the fused feature vector is as follows:
[0033] in, represents the fused feature vector, Indicates the depth characteristics of the vibration signal, represents the time domain characteristics, represents the frequency domain characteristics, R represents feature splicing, 、 and Both represent feature dimensions.
[0034] The beneficial effect of the above further solution is that the time domain and frequency domain features of the signal are fused as input to the training model.
[0035] Furthermore, the use of a two-layer long short-term memory network (LSTM) to establish a long-term dependency branch modeling long-term dependency relationship has the following specific requirements: The first-layer long short-term memory network (LSTM) receives the fused dimensional feature vector, filters key timing patterns through a gating mechanism, and outputs a hidden state sequence containing temporal context information. The first-layer long short-term memory network (LSTM) can learn and maintain the long-term spatiotemporal dependencies of the current vibration sequence. The second-layer long short-term memory network LSTM is used to model high-order temporal evolution with 256-dimensional hidden states, and output the final high-order temporal feature representation; among them, based on the processing results of the first and second-layer short short-term memory networks LSTM, the modal energy distribution information of the frequency domain power spectral density is further spliced to obtain the high-dimensional nonlinear mapping relationship between the video extraction signal and the point sensor signal.
[0036] The beneficial effect of the above further solution is: through the above design, the correction model is trained to obtain a more accurate correction result.
[0037] Furthermore, the loss function of the multimodal fusion calibration model is expressed as follows:
[0038]
[0039]
[0040] in, represents the loss function of the multimodal fusion calibration model, represents the time domain weighted mean square error loss, represents the frequency domain spectrum loss, and Represent the weight coefficients of time domain loss and frequency domain loss respectively, represents the L2 regularization coefficient, represents the L2 norm of the multimodal fusion calibration model, represents the measured video micro-vibration signal spectrum amplitude, represents the predicted video micro-vibration signal spectrum amplitude, and denote the mean of the measured and predicted spectrum amplitudes, represents the total sample time, represents the sampling time, represents the time domain weight, represents the spectrum amplitude of the measured video micro-vibration signal, represents the predicted signal spectrum amplitude, k Indicates the valid frequency index.
[0041] The beneficial effect of the above further solution is that the video signal can be corrected by calculating the above loss function. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Flow chart of the method of the present invention.
[0043] Figure 2 This is a flowchart of the processing of the present invention. DETAILED DESCRIPTION
[0044] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0045] Example Based on the above background technology, this paper proposes a video micro-vibration signal correction method based on a hybrid CNN-LSTM-FFT deep learning architecture. This method targets signals acquired by point sensors and combines the physical characteristics of structural vibration with the signal's time-frequency domain information to achieve dual correction of video micro-vibration signals in both the time and frequency domains. First, the signal's time- and frequency-domain characteristics, such as displacement, velocity, acceleration, power spectral density (PSD), and short-term statistics, are extracted to construct a signal time-frequency feature matrix. Second, a hybrid CNN-LSTM model is trained. The CNN layer extracts local motion patterns and denoises the signal, while the LSTM layer captures long-term correlations and modal trends in the vibration sequence, learning the mapping relationship between the extracted video signal and the sensor-acquired signal. Third, a spectral loss term based on the Fourier spectrum logarithmic amplitude difference is introduced to ensure that the key modal frequencies of the corrected signal are consistent. Finally, the method's reliability is verified using an old aqueduct in a field environment, laying the foundation for the application of accurate video micro-vibration monitoring in large-scale hydraulic structures.
[0046] like Figure 1 As shown, the present invention provides a method for correcting video micro-vibration signals, and its implementation method is as follows: Video capture: Place the vibration sensor on the structure to be tested to capture the video of the structure to be tested; Feature extraction: Based on the video capture results, the time domain features and frequency domain features of the vibration signal are extracted; Constructing a multimodal fusion calibration model: Using time domain features and frequency domain features to train and construct a multimodal fusion calibration model, where the multimodal fusion calibration model is a hybrid CNN-LSTM network architecture. Correction result output: A spectral loss term based on the Fourier spectrum logarithmic amplitude difference is introduced, and the constructed multimodal fusion calibration model is used to obtain the dual correction target results of time domain waveform reconstruction and frequency domain key mode alignment to complete the correction of the video micro-vibration signal.
[0047] In this embodiment, a multimodal fusion calibration model is constructed by training time domain features and frequency domain features, which is specifically as follows: Deep features of vibration signals are automatically extracted by using a hierarchical one-dimensional convolutional neural network (1D CNN); Fusion of vibration signal depth features with time domain features and frequency domain features; Based on the fused feature vector, a two-layer long short-term memory network (LSTM) is used to establish a long-term dependency branch model to model the long-term dependency relationship, so as to learn the high-dimensional nonlinear mapping relationship between the video extraction signal and the point sensor acquisition signal, and complete the construction of the multimodal fusion calibration model.
[0048] In this embodiment, the vibration signal depth features automatically extracted by using a hierarchical one-dimensional convolutional neural network 1D CNN are specifically: For the input vibration displacement signal sequence, a symmetrically padded one-dimensional convolution kernel is used in the short-time dynamic feature extraction layer to capture high-frequency components and transient impacts through local weighted summation and nonlinear mapping: Based on the captured high-frequency components and transient impacts, batch normalization is performed to obtain normalized high-frequency transient features; In the long-term trend modeling layer, the low-frequency envelope and slow-varying drift features are extracted through the diffusion convolution kernel and regularization to obtain the low-frequency drift features. The deep features of the vibration signal include high-frequency transient features and low-frequency drift features. The deep features of the vibration signal are spliced with the extracted time domain features and frequency domain features along the channel dimension to obtain a fused feature vector, wherein the fused feature vector is input into a two-layer long short-term memory network LSTM.
[0049] In this embodiment, a two-layer long short-term memory network (LSTM) is used to establish a long-term dependency branch model for long-term dependency relationships. The specific requirements are as follows: The first-layer long short-term memory network (LSTM) receives the fused dimensional feature vector, filters key timing patterns through a gating mechanism, and outputs a hidden state sequence containing temporal context information. The first-layer long short-term memory network (LSTM) can learn and maintain the long-term spatiotemporal dependencies of the current vibration sequence. The second-layer long short-term memory network LSTM is used to model high-order temporal evolution with 256-dimensional hidden states, and output the final high-order temporal feature representation; among them, based on the processing results of the first and second-layer short short-term memory networks LSTM, the modal energy distribution information of the frequency domain power spectral density is further spliced to obtain the high-dimensional nonlinear mapping relationship between the video extraction signal and the point sensor signal.
[0050] In this embodiment, a two-layer long short-term memory (LSTM) network architecture is designed to model long-term temporal dependencies, as follows: The first LSTM layer is used to capture fundamental long-term patterns and filter information. Its functions include: receiving a sequence of fused feature vectors; and modeling long-term dependencies. Leveraging its cell state and gating mechanisms (forget gate, input gate, and output gate), the first LSTM layer can selectively retain, update, or forget historical information spanning long time steps. The forget gate determines which old information to discard, and the input gate determines which new information to incorporate into the cell state. This enables the first LSTM layer to learn and maintain long-term spatiotemporal dependency patterns that are crucial for current predictions (for example, the potential impact of a specific vibration pattern in the early stages of equipment operation on the current fault state). The first LSTM layer outputs a sequence of hidden states that contain temporal context information.
[0051] The second-layer LSTM is used for high-order temporal evolution modeling and feature deepening. Its function is to receive the hidden state sequence output by the first-layer LSTM; long-term dependency modeling mechanism: Based on the temporal patterns extracted by the first layer, the second-layer LSTM further learns more complex and abstract high-order temporal dynamics and evolution laws. It inherits the long-term dependency modeling capability of the first-layer LSTM and captures dependencies spanning longer time ranges in a higher-level feature space (for example, capturing the feature evolution pattern of the entire gradual process from fault initiation to development). Its hidden state dimension is set to 256, providing sufficient capacity to characterize these complex long-term evolution features; the second-layer LSTM outputs the final high-order temporal feature representation (256-dimensional hidden state sequence).
[0052] Feature fusion enhancement: To construct a more comprehensive feature representation, the high-order time series features (representing long-term evolution patterns) output by the second-layer LSTM are concatenated with the modal energy distribution information extracted from the frequency-domain power spectral density (PSD). The frequency-domain features provide static distribution information of signal energy at different frequency components, complementing the dynamic long-term evolution features captured by LSTM.
[0053] Final mapping relationship: The fused feature vector (long-term temporal dependency features + frequency domain energy features) is input into subsequent modules such as the fully connected layer. This high-dimensional fusion feature integrates the long-term dynamic evolution of the video vibration signal in the time dimension and the energy distribution characteristics in the frequency dimension, thereby more effectively learning the complex high-dimensional nonlinear mapping relationship between the video extraction signal (containing spatial and temporal information) and the point sensor signal (containing precise physical quantity information), providing strong feature support for the final diagnosis or prediction task.
[0054] In this embodiment, Figure 2 As shown, the present invention mainly relates to the following contents: (1) Video capture: The vibration sensor is placed on the structure to be tested. The video image contains the location of the sensor. The sensor signal is used as the reference signal, and the vibration signal extracted from the pixel point at the sensor position in the video is used as the signal to be corrected.
[0055] (2) Feature extraction: Extract the time domain and frequency domain feature indicators of the signal and construct a feature indicator library. The time domain indicators include displacement, velocity, acceleration, sliding window function (sliding mean, variance, maximum value, cumulative displacement) and autocorrelation function; the frequency domain features include Fourier amplitude spectrum and power spectrum density.
[0056] (3) Joint deep analysis in the time-frequency domain: Based on the deep features of the vibration signal automatically extracted by the hierarchical one-dimensional convolutional neural network (1D CNN), a feature expression that complements the local motion details and the global dynamic response is constructed; based on the two-layer long short-term memory network (LSTM), a long-term dependency branch is established to model the long-term dependency relationship, in which the forward propagation captures the vibration propagation and the back propagation models the lag effects such as the structural damping oscillation; the modal energy distribution information of the frequency domain power spectral density (PSD) is further spliced to form a cross-domain joint feature matrix.
[0057] (4) Multimodal fusion calibration model: Design a hybrid CNN-LSTM network architecture, use the layered convolution kernel of CNN to extract multi-scale spatiotemporal motion patterns, use the LSTM layer to model the long-term spatiotemporal dependence and modal evolution of the vibration sequence, and establish a high-dimensional nonlinear mapping relationship between the video extraction signal and the point sensor signal.
[0058] (5) Time-frequency domain loss optimization: A spectral consistency constraint based on the logarithmic amplitude difference of the Fourier spectrum is introduced. By jointly optimizing the time-frequency domain loss function, the dual correction goals of time domain waveform reconstruction and frequency domain key mode alignment are achieved.
[0059] In this embodiment, manual feature extraction is as follows: To achieve the time-frequency domain collaborative representation of vibration signals, the present invention extracts the features from the discrete signal sequence. Constructing joint eigenvectors in , including two complementary feature sets in time domain and frequency domain. In this embodiment, the time domain features characterize the physical behavior of the vibration signal through the dynamic evolution of the signal waveform and local statistical characteristics, and extract four key features: dynamic characteristics, sliding window statistics, integral characteristics and autocorrelation function.
[0060] In this embodiment, the dynamic characteristics are determined by the first-order difference (velocity) of the signal. and second-order difference (acceleration) Characterization is used to capture the changing trend of vibration waveform and high-frequency impact effect. Its calculation formula is:
[0061]
[0062] in, Indicates that the vibration signal is n The first-order difference (velocity) of the sampling points, Indicates that the vibration signal is n The second-order difference (acceleration) of the sampling points, represents the displacement change of the vibration signal, represents the sampling time, 、 and Respectively represent the vibration signal in n 、 n -1. n -2 displacement of sampling points, Indicates the time step.
[0063] In this embodiment, the sliding window statistics are calculated by The sliding window calculates the sliding mean ,variance and maximum value , suppress noise interference and quantify the transient response:
[0064]
[0065]
[0066] in, represents the sliding mean, represents the sliding window length, i represents the local index within the sliding window, Indicates the i The displacement value of the sampling point, i∈[1,n], represents the variance, Indicates the window length Internal sampling point n The maximum value of .
[0067] In this embodiment, the integral characteristic (accumulated displacement) The energy drift characteristics of the signal are quantified by time accumulation:
[0068] in, It represents the energy drift characteristic of the quantized signal accumulated over time. represents the time step, i Represents the local index within the sliding window.
[0069] In this embodiment, the autocorrelation function Through hysteresis q Similarity analysis of the steps reveals a periodic component:
[0070] in, Indicates calculation lag q The signal self-similarity at step , Total signal length (number of sampling points), i Represents the local index within the sliding window.
[0071] In this embodiment, to effectively characterize the frequency domain characteristics of non-stationary signals, the present invention uses Fourier transforms to extract two types of features: power spectral density (PSD) and amplitude spectrum. The power spectral density reflects the frequency band distribution of signal energy through statistical averaging, suppressing noise interference; the amplitude spectrum preserves the amplitude details of instantaneous frequency components. The combination of the two provides complementary frequency domain information for deep learning models.
[0072] In this embodiment, the windowed signal , compute the discrete Fourier transform:
[0073] Before Retention The amplitude spectrum eigenvector of the effective frequency point is:
[0074] The power spectral density is estimated using the Welch method:
[0075] in, represents the amplitude spectrum eigenvector, represents the complex amplitude of the kth frequency point of the discrete Fourier transform DFT, X[k] is the frequency domain amplitude spectrum of the complex modulus value, Indicates the preset maximum effective frequency index, represents the windowed signal, j represents the imaginary unit, represents the Hamming window, represents the power spectral density, T represents the sampling time, represents the energy in the time domain, Represents the transpose of x after Fourier transformation.
[0076] In this embodiment, deep feature analysis based on 1D-CNN: To overcome the limitations of traditional time domain analysis in multi-scale vibration feature characterization, the present invention designs a hierarchical one-dimensional convolutional neural network architecture. By integrating signal processing theory and deep learning methods, it realizes cross-scale feature extraction from transient impact to low-frequency drift. This design is based on local convolution operation ( K =5) and the dilated convolution kernel ( K =15) as the core, combined with the regularization strategy to build a two-level structure, its mathematical form and physical meaning are as follows: Input vibration signal sequence , the short-term dynamic feature extraction layer uses a symmetrically filled one-dimensional convolution kernel ( K =5, c =64), capturing high-frequency components (>50Hz) and transient impacts through local weighted summation and nonlinear mapping:
[0077] The output features are batch normalized (BN) to suppress covariate shift:
[0078] in, , represents a trainable weight tensor, represents the linear rectification activation function, defined as , a represents the linear weighted sum result of the convolution operation, and Both represent affine transformation parameters, represents the 𝑘th eigenvalue of the lth convolution output, represents the variance, represents the mean, represents the numerical stability constant, which is 10 -5 , d represents the feature index dimension, D represents a dataset, represents a trainable weight tensor, represents the linear rectification activation function, Indicates the effective frequency index c represents the number of convolution channels, Represents the input vibration signal, Indicates time, represents the bias term of the first layer, N represents the signal length, To ensure numerical stability. This layer is equivalent to 5 sampling intervals (Δ t eff =5Δ t ) differential filter to enhance the sensitivity to transient responses such as mechanical collision.
[0079] In this embodiment, the long-term trend modeling layer expands the convolution kernel ( K =15, c =128) and Dropout regularization (ratio 0.2) to extract low-frequency envelope (<10Hz) and slow-changing drift features:
[0080] in, Represents the low-frequency drift characteristics, represents the regularization operation, represents a trainable weight tensor, represents the input feature, that is, the high-frequency transient feature, represents the bias term of the second layer, W (2) ∈R 15×64×128 , Dropout (ratio 0.2) enhances generalization ability by randomly masking neurons. The receptive field of this layer is extended to 15 sampling intervals (Δ t eff =3×5Δ t ), its integral characteristic ( ) and the first-level differential operation ( d X / dt ) complements the multi-resolution analysis used in simulating vibration signal processing. The number of channels is increased to 128 to enhance feature representation, and Dropout prevents overfitting to noise by randomly blocking 20% of neurons.
[0081] Two-stage convolution achieves feature decoupling through a multi-scale collaborative mechanism: short kernels focus on local details (high-frequency transients), while long kernels model global trends (low-frequency drifts). Its parameter count (126,080) is two orders of magnitude lower than that of a fully connected network. This design strikes a balance between computational efficiency and feature completeness, providing a high signal-to-noise ratio time-domain embedding for subsequent LSTM time series modeling and significantly improving the physical consistency of vibration signal calibration.
[0082] In this embodiment, frequency domain feature splicing: In order to achieve the multimodal characterization synergy of time domain dynamic characteristics and frequency domain energy distribution, the present invention proposes a channel dimension fusion strategy to combine the deep time domain features extracted by the hierarchical one-dimensional convolutional neural network 1D-CNN , Indicates the total sample time, manual time domain statistical features And frequency domain characteristics Joint stitching along the channel dimension.
[0083] For each time , the fused feature vector is constructed by concatenation:
[0084] in, represents the fused feature vector, Indicates the depth characteristics of the vibration signal, represents the time domain characteristics, represents the frequency domain characteristics, R represents feature splicing, 、 and Both represent feature dimensions. 、 as well as , the combined representation after fusion Layer-wise normalization and nonlinear projection ( , represents the dimension after projection, represents the weight matrix of the nonlinear projection layer) and then inputs the bidirectional LSTM module.
[0085] In this embodiment, in order to accurately model the long-range spatiotemporal dependency of video vibration signals, the present invention uses a two-layer long short-term memory network (LSTM) to construct a temporal dynamic mapping model. The network input is a fusion feature sequence The network architecture utilizes a stacked LSTM design: the first LSTM layer receives a 256-dimensional feature vector and uses a gating mechanism to filter key temporal patterns. The second LSTM layer models high-order temporal evolution using a 256-dimensional hidden state, and the output gate uses the tanh activation function, which exhibits superior saturation properties. To enhance generalization, a dropout rate of 0.2 is set between layers, and gradient clipping is used to ensure stability during long sequence training.
[0086] In this embodiment, in order to jointly optimize the time domain waveform matching and the frequency domain energy distribution consistency, the present invention designs a composite loss function, which is composed of the time domain weighted mean square error, the frequency domain spectrum correlation loss and the L2 regularization term.
[0087] The time-domain weighted mean square error enhances the calibration accuracy for transient shock events through dynamic weight coefficients:
[0088] The frequency domain spectral correlation loss constrains the energy distribution consistency by normalizing the covariance:
[0089] The total loss function integrates the time-frequency domain supervision signal and the regularization constraint:
[0090] in, represents the loss function of the multimodal fusion calibration model, represents the time domain weighted mean square error loss, Frequency domain spectrum loss, and Represent the weight coefficients of time domain and frequency domain losses respectively, represents the L2 regularization coefficient, represents the L2 norm (weight decay term) of the multimodal fusion calibration model, k Indicates the valid frequency index. represents the measured signal spectrum amplitude, represents the predicted signal spectrum amplitude, and denote the mean of the measured and predicted spectrum amplitudes, represents the total sample time, represents the sampling time, Represents the time domain weight, which is dynamically adjusted with the signal derivative, giving a higher penalty weight in the vibration peak area (such as the moment of mechanical collision). represents the measured amplitude of the signal at time t, Represents the predicted amplitude of the signal at time t.
Claims
1. A video micro-vibration signal correction method, characterized in that: include: Video capture: Place the vibration sensor on the structure to be tested to capture the video of the structure to be tested; Feature extraction: Based on the video capture results, the time domain features and frequency domain features of the vibration signal are extracted; Constructing a multimodal fusion calibration model: Using time domain features and frequency domain features to train and construct a multimodal fusion calibration model. The multimodal fusion calibration model is a hybrid CNN-LSTM network architecture. Correction result output: A spectral loss term based on the Fourier spectrum logarithmic amplitude difference is introduced, and the constructed multimodal fusion calibration model is used to obtain the dual correction target results of time domain waveform reconstruction and frequency domain key mode alignment to complete the correction of the video micro-vibration signal.
2. The video micro-vibration signal correction method according to claim 1, characterized in that: The time domain features include: Dynamic features: in, Indicates that the vibration signal is n The first-order difference of the sampling points, Indicates that the vibration signal is n The second-order difference of the sampling points, represents the displacement change of the vibration signal, represents the sampling time, 、 and Respectively represent the vibration signal in n 、 n -1. n -2 displacement of sampling points, represents the time step; Sliding window statistics: in, represents the sliding mean, represents the sliding window length, i represents the local index sampling point within the sliding window, Indicates the i The displacement value of each sampling point, i∈[1,n], represents the variance, Indicates the window length The maximum value of the inner sampling point n; Cumulative displacement: in, It represents the energy drift characteristics of the quantized signal through time accumulation; Autocorrelation function: in, Indicates calculation lag q The signal self-similarity at step , Total signal length.
3. The video micro-vibration signal correction method according to claim 2, characterized in that: The frequency domain features include: Fourier amplitude spectrum: in, represents the amplitude spectrum eigenvector, Represents the complex amplitude of the kth frequency point of discrete Fourier transform DFT, To obtain the frequency domain amplitude spectrum of the complex modulus value, K max Indicates the preset maximum effective frequency index, represents the windowed signal, j represents the imaginary unit, represents the Hamming window, N represents the total length of the signal, and n represents the sampling point; Power spectral density: in, represents the power spectral density, T represents the sampling time, represents the energy in the time domain, Represents the transpose of x after Fourier transformation.
4. The video micro-vibration signal correction method according to claim 1, characterized in that: The multimodal fusion calibration model is constructed by training with time domain features and frequency domain features, which is specifically as follows: Deep features of vibration signals are automatically extracted by using a hierarchical one-dimensional convolutional neural network (1D CNN); Fusion of vibration signal depth features with time domain features and frequency domain features; Based on the fused feature vector, a two-layer long short-term memory network (LSTM) is used to establish a long-term dependency branch model to model the long-term dependency relationship, so as to learn the high-dimensional nonlinear mapping relationship between the video extraction signal and the point sensor acquisition signal, and complete the construction of the multimodal fusion calibration model.
5. The video micro-vibration signal correction method according to claim 4, characterized in that: The vibration signal deep features automatically extracted by using a hierarchical one-dimensional convolutional neural network 1D CNN are specifically: For the input vibration displacement signal sequence, a symmetrically padded one-dimensional convolution kernel is used in the short-time dynamic feature extraction layer to capture high-frequency components and transient impacts through local weighted summation and nonlinear mapping: Based on the captured high-frequency components and transient impacts, batch normalization is performed to obtain normalized high-frequency transient features; In the long-term trend modeling layer, the low-frequency envelope and slow-varying drift features are extracted through the diffusion convolution kernel and regularization to obtain the low-frequency drift features. The deep features of the vibration signal include high-frequency transient features and low-frequency drift features. The deep features of the vibration signal are spliced with the extracted time domain features and frequency domain features along the channel dimension to obtain a fused feature vector, wherein the fused feature vector is input into a two-layer long short-term memory network LSTM.
6. The video micro-vibration signal correction method according to claim 5, characterized in that: The expression of the high-frequency transient characteristic is as follows: in, Represents high-frequency transient characteristics, and Both represent affine transformation parameters, represents the 𝑘th eigenvalue of the lth convolution output, represents the variance, represents the mean, represents a numerical stability constant, d represents the feature index dimension, D represents a dataset, represents the convolution kernel weight, represents the linear rectification activation function, Represents the effective frequency index, c represents the number of convolution channels, Represents the input vibration signal, Indicates time, represents the bias term of the first layer; N represents the total length of the signal; The expression of the low-frequency drift characteristic is as follows: in, Represents the low-frequency drift characteristics, represents the regularization operation, represents a trainable weight tensor, represents the input feature, that is, the high-frequency transient feature, Represents the bias term of the second layer.
7. The video micro-vibration signal correction method according to claim 4, characterized in that: The expression of the fused feature vector is as follows: in, represents the fused feature vector, Indicates the depth characteristics of the vibration signal, represents the time domain characteristics, represents the frequency domain characteristics, R represents feature splicing, 、 and Both represent feature dimensions.
8. The video micro-vibration signal correction method according to claim 4, characterized in that: The two-layer long short-term memory network LSTM is used to establish a long-term dependency branch model for long-term dependency relationships. The specific requirements are: The first-layer long short-term memory network (LSTM) receives the fused dimensional feature vector, filters key timing patterns through a gating mechanism, and outputs a hidden state sequence containing temporal context information. The first-layer long short-term memory network (LSTM) can learn and maintain the long-term spatiotemporal dependencies of the current vibration sequence. The second-layer long short-term memory network LSTM is used to model high-order temporal evolution with 256-dimensional hidden states, and output the final high-order temporal feature representation; among them, based on the processing results of the first and second-layer short short-term memory networks LSTM, the modal energy distribution information of the frequency domain power spectral density is further spliced to obtain the high-dimensional nonlinear mapping relationship between the video extraction signal and the point sensor signal.
9. The video micro-vibration signal correction method according to claim 1, characterized in that: The loss function of the multimodal fusion calibration model is expressed as follows: in, represents the loss function of the multimodal fusion calibration model, represents the time domain weighted mean square error loss, represents the frequency domain spectrum loss, and Represent the weight coefficients of time domain loss and frequency domain loss respectively, represents the L2 regularization coefficient, represents the L2 norm of the multimodal fusion calibration model, represents the measured video micro-vibration signal spectrum amplitude, represents the predicted video micro-vibration signal spectrum amplitude, and denote the mean of the measured and predicted spectrum amplitudes, represents the total sample time, represents the sampling time, represents the time domain weight, represents the spectrum amplitude of the measured video micro-vibration signal, represents the predicted signal spectrum amplitude, k Indicates the valid frequency index.
Citation Information
Patent Citations
Structural micro-amplitude vibration working mode analysis method based on optical flow method
CN114187330A
Data processing method and device, equipment and storage medium
CN116959473A
Video vibration calibration method integrating video micro-vibration monitoring and sensor
CN118261001A
Visual and vibration monitoring data fused arch bridge dynamic displacement identification method and system
CN118520413A
Real-time cable force identification method and device for inhaul cable based on deep learning
CN119942322A