A vibration signal space-time reconstruction method based on a multi-modal conditional diffusion model

The spatiotemporal reconstruction method of vibration signals using a multimodal conditional diffusion model solves the problem of signal loss caused by limited sensor deployment in complex engineering structures, achieving high-precision and high-consistency signal reconstruction and improving the data reliability and availability of structural health monitoring.

CN121502240BActive Publication Date: 2026-05-08HUAQIAO UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAQIAO UNIVERSITY
Filing Date
2026-01-13
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively recover the spatiotemporal correlation and accuracy of multi-sensor vibration signals in complex engineering structures. In particular, when sensor density is insufficient or signal acquisition is interrupted, traditional methods are unable to maintain the global trend, local features, and spatiotemporal correlation of the signal, resulting in insufficient reconstruction accuracy and generalization ability.

Method used

A spatiotemporal reconstruction method for vibration signals based on a multimodal conditional diffusion model is adopted. Iterative denoising and inversion are performed through a pre-trained diffusion model. Combined with Gaussian noise filling and multimodal conditional embedding, the signal amplitude and temporal characteristics of the missing region are gradually restored. In self-supervised learning, the model is guided to learn spatiotemporal related features and missing modes.

Benefits of technology

It significantly improves the reconstruction accuracy and physical consistency of multi-channel vibration signals in complex missing scenarios, provides high-integrity data support, and provides reliable data support for structural modal identification, damage diagnosis and safety assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502240B_ABST
    Figure CN121502240B_ABST
Patent Text Reader

Abstract

The application provides a vibration signal space-time reconstruction method based on a multi-modal condition diffusion model, and relates to the technical field of vibration signal reconstruction. First, multi-sensor is used to collect structural vibration response, a multi-dimensional vibration signal matrix is constructed, and spatial continuous missing and time random missing areas are automatically identified. Then, adaptive multi-scale interpolation is used for coarse reconstruction of the missing areas to restore the basic trend and frequency band characteristics of the signal. Further, a pseudo-missing mask is applied to the complete data, training samples are constructed through a self-supervised strategy, and the model is guided to learn space-time correlation characteristics and missing patterns. In the training stage, the diffusion model is used as a generation framework, Gaussian noise disturbance is applied to the missing areas, and four types of condition embedding, including time, space, trend and frequency domain, are introduced in the denoising inversion process to respectively depict the periodicity of the signal, the multi-sensor space coupling, the low-frequency change and the physical frequency spectrum structure, and realize high-fidelity signal reconstruction under the joint constraint of multi-modal information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vibration signal reconstruction technology, and specifically to a spatiotemporal reconstruction method for vibration signals based on a multimodal conditional diffusion model. Background Technology

[0002] In the field of structural health monitoring and vibration signal analysis, acquiring complete and continuous multi-sensor vibration response data is fundamental for modal parameter identification, structural condition assessment, and damage diagnosis. However, in practical engineering applications, due to harsh environmental conditions, complex structural layouts, or insufficient sensor reliability, it is often difficult to deploy a sufficient number of sensors at all critical locations or maintain stable data acquisition during long-term monitoring. For example, in complex scenarios such as deep foundation pits, long-span bridges, offshore wind turbine towers, or high-rise buildings, some areas may experience insufficient sensor density or signal acquisition interruptions due to space constraints, drastic temperature and humidity changes, strong electromagnetic interference, or poor installation conditions, resulting in missing structural vibration signals in both time and space dimensions. This missing signal not only manifests as random interruptions in the time series of a single sensor but may also present as continuous spatial absences of multiple sensors within a certain time period, severely affecting the accuracy and reliability of subsequent structural modal identification, frequency domain analysis, and safety assessment. Furthermore, traditional data interpolation or filtering methods struggle to simultaneously recover the global trend, local features, and spatiotemporal correlation of the signal, especially in multi-sensor high-dimensional signal scenarios, where their reconstruction accuracy and generalization capabilities are significantly insufficient.

[0003] For dynamic monitoring of complex engineering structures, the integrity of vibration signals directly affects the accuracy of subsequent modal identification, damage diagnosis, and structural health assessment. Currently, reconstruction methods for missing or abnormal vibration signals mainly include traditional signal processing techniques such as interpolation-based time-domain reconstruction, filtering and smoothing, low-rank decomposition, and wavelet transform, as well as deep learning methods that have been gradually introduced in recent years, such as generative adversarial networks (GANs) and compressed sensing (CS). While time-domain interpolation methods can restore some signal continuity in locally missing regions, they struggle to maintain the global trend of the signal and the spatial correlation between multiple channels. Filtering and frequency-domain reconstruction methods improve signal smoothness to some extent, but they easily lose high-frequency details and instantaneous modal features. Low-rank decomposition and wavelet transform have limited reconstruction accuracy when facing large-scale missing signals, spatially continuous missing signals, or random temporal missing signals, and are difficult to adapt to complex signal structures with high dimensions and multiple channels. While deep learning methods have improved the realism and continuity of reconstructed signals to some extent and enhanced the recovery of local missing data, existing GAN models still fall short in their ability to model complex spatiotemporal dependencies, suffer from unstable training processes, and exhibit poor adaptability to missing patterns. Although compressed sensing methods utilize signal sparsity to achieve reconstruction under undersampling conditions and have certain advantages for underdetermined problems, their reconstruction accuracy and robustness still need improvement for structural response signals with significant dynamic changes or complex multimodal features.

[0004] In summary, existing methods struggle to reliably reconstruct spatiotemporal vibration signals under complex structures while preserving the global trend, local details, and spatial correlation of the signal across multiple sensors.

[0005] In view of the above, this application is hereby submitted. Summary of the Invention

[0006] This invention provides a spatiotemporal reconstruction method for vibration signals based on a multimodal conditional diffusion model, which can at least partially improve the above-mentioned problems.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A spatiotemporal reconstruction method for vibration signals based on a multimodal conditional diffusion model, comprising:

[0009] The missing vibration signals collected by the sensor components deployed on the engineering structure are acquired. The vibration signals are input into a pre-trained diffusion model and spatiotemporal reconstruction is performed to obtain a complete multi-sensor vibration signal matrix. The missing regions of the vibration signals are filled with Gaussian noise and used as the initial input. Under the constraints of multimodal embedding, the amplitude and temporal characteristics of the missing region signals are gradually recovered through iterative denoising and inversion.

[0010] The complete multi-sensor vibration signal matrix is ​​quantitatively evaluated and physically verified to ensure that the reconstructed signal conforms to the laws of structural dynamics.

[0011] In summary, this method utilizes multiple sensors to collect structural vibration signals, constructs a multi-dimensional vibration signal matrix, and determines spatially continuous and temporally random missing regions based on sensor sampling states. For the missing regions, an adaptive multi-scale interpolation method is used for coarse reconstruction to restore the basic trend and frequency band characteristics of the signal. Subsequently, complete and missing data are distinguished, and a pseudo-missing mask is applied to the non-missing data. Training samples are constructed through a self-supervised learning strategy to guide the model to learn spatiotemporally related features and missing patterns. During model training, a diffusion model is used as the signal generation framework, Gaussian noise perturbation is applied to the missing regions, and multimodal conditional feature embedding is introduced during the denoising and inversion process, including temporal feature embedding, spatial feature embedding, trend feature embedding, and frequency domain feature embedding. Among them, temporal embedding is used to characterize the periodicity and temporal dependence of the signal, spatial embedding captures the spatial correlation between multiple sensors through convolution and graph structure modeling, trend embedding uses convolution and gated recurrent networks to extract the low-frequency variation patterns of the signal, and frequency domain embedding combines physical spectrum features with lightweight convolutional networks to achieve spectral structure perception. By jointly constraining the aforementioned multimodal conditional information, the diffusion model is guided to generate a high-fidelity reconstructed signal that conforms to physical laws during the inversion process. This invention's method can effectively recover vibration signals from multi-sensor structures under complex spatiotemporal conditions, improving the accuracy of signal reconstruction and providing high-integrity vibration signal data support for engineering scenarios where complex structures or harsh environmental conditions limit sensor deployment.

[0012] This method can effectively recover the amplitude, phase and spectral characteristics of multi-channel vibration signals in complex missing scenarios, significantly improving reconstruction accuracy and physical consistency, and providing complete and reliable data support for structural modal identification, damage diagnosis and safety assessment. Attached Figure Description

[0013] Figure 1 This is a schematic flowchart of the spatiotemporal reconstruction method for vibration signals based on a multimodal conditional diffusion model provided in an embodiment of the present invention.

[0014] Figure 2 This is a flowchart of the spatiotemporal reconstruction method for vibration signals based on a multimodal conditional diffusion model provided in this embodiment of the invention.

[0015] Figure 3 This is a schematic diagram of the cantilever beam structure model provided in the embodiment of the present invention (a structural diagram of the experimental verification device).

[0016] Figure 4This is a comparison diagram of the reconstruction effect of the cantilever beam structure without changing the structure provided in the embodiment of the present invention (the reconstruction effect diagram under the condition that the cantilever beam structure generates a stable signal).

[0017] Figure 5 This is a comparison diagram of the reconstruction effect of a cantilever beam structure over time provided by an embodiment of the present invention (the reconstruction effect diagram under unstable signal generation of the cantilever beam structure).

[0018] Figure 6 This is a bar chart comparing the cantilever beam structure reconstruction algorithm provided in this embodiment of the invention with the structure unchanged (comparison experiment with other algorithms at 5% and 10% missing rates).

[0019] Figure 7 This is a bar chart comparing the reconstruction algorithms for cantilever beam structures over time provided in this embodiment of the invention (comparison experiment with other algorithms at 5% and 10% missing rates).

[0020] Figure 8 This is a frequency domain comparison diagram of cantilever beam structure reconstruction without structural change provided in the embodiments of the present invention (comparing whether the frequency domain information before and after signal reconstruction is consistent; the stable signal uses a frequency domain diagram).

[0021] Figure 9 This is a comparison chart of the reconstructed frequency domain of a cantilever beam structure over time, provided by an embodiment of the present invention (comparing whether the frequency domain information before and after signal reconstruction is consistent; for unstable signals, a time-frequency domain chart is used). Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0023] refer to Figures 1-2 As shown, the first embodiment of the present invention discloses a spatiotemporal reconstruction method for vibration signals based on a multimodal conditional diffusion model, which can be executed by a spatiotemporal reconstruction device for vibration signals based on a multimodal conditional diffusion model (hereinafter referred to as the reconstruction device), specifically, by one or more processors within the reconstruction device, to implement the following method:

[0024] S1, acquire the missing vibration signal collected by the sensor components deployed on the engineering structure, input the vibration signal into the pre-trained diffusion model, perform spatiotemporal reconstruction processing, and obtain a complete multi-sensor vibration signal matrix. In this case, Gaussian noise is filled into the missing region of the vibration signal and used as the initial input. Under the multi-modal condition embedding constraint, the amplitude and temporal characteristics of the signal in the missing region are gradually recovered through iterative denoising and inversion.

[0025] Specifically, step S1 further includes: acquiring the missing vibration signal collected by the sensor assembly deployed on the engineering structure, and gradually adding Gaussian noise to it to fill it with Gaussian noise.

[0026] In the reverse denoising process, a pre-trained diffusion model is used to predict noise after filling, and then denoise the noise to obtain the final reconstructed vibration signal X. recon ;

[0027] For the reconstructed vibration signal X recon By performing inverse normalization, the signal is restored to the physical dimensions of the engineering structure, resulting in a complete multi-sensor vibration signal matrix. ,in, To represent element-wise multiplication, The standard deviation of each channel of the original signal before normalization. This represents the mean of all channels of the original signal before normalization. This step involves denormalizing the signal after reconstruction to the order of magnitude of the original data.

[0028] Specifically, in this embodiment, a sensor assembly consisting of multiple acceleration or displacement sensors is arranged on the surface of the engineering structure to continuously collect vibration signals, forming an original multi-channel vibration signal matrix. Due to environmental interference or equipment failure, some spatiotemporal missing regions exist in this matrix. The missing vibration signals are received and marked, and then Gaussian noise is applied to the missing regions. That is, random values ​​generated by a zero-mean Gaussian distribution are used to fill the missing positions in the original signal, forming an initial input signal matrix with noise. This filling method avoids the information bias caused by traditional zero-value or mean filling, provides a reasonable random initial state for the reverse denoising process of the diffusion model, and helps to improve the naturalness and continuity of the reconstructed signal.

[0029] Next, the padded signal is input into a pre-trained diffusion model. This model, having learned the spatiotemporal distribution characteristics of the signal through self-supervised learning during training, possesses the ability to respond to multimodal conditional information. During the reverse denoising process, the model predicts noise based on the current time step information and multimodal conditional embeddings (including temporal, spatial, trend, and frequency domain features), and removes Gaussian noise from the signal layer by layer, gradually restoring the amplitude and temporal characteristics of the missing regions. Through multiple iterative inversions, the model can generate a reconstructed signal with a distribution and structural characteristics consistent with the original signal, achieving high-fidelity recovery of the missing data.

[0030] Ultimately, the model outputs a reconstructed signal in the normalized domain. To give it practical physical meaning, this signal needs to be denormalized: using the mean and standard deviation of each channel saved during the training phase, a linear transformation is performed on the signal to restore its original amplitude range and dimensions, resulting in a complete multi-sensor vibration signal matrix. This matrix exhibits good waveform continuity and spectral consistency between the missing and non-missing regions, effectively supporting subsequent tasks such as structural modal analysis and damage identification, and significantly improving data availability and analytical reliability under complex missing conditions.

[0031] S2 performs quantitative evaluation and physical verification of the complete multi-sensor vibration signal matrix to ensure that the reconstructed signal conforms to the laws of structural dynamics.

[0032] Specifically, step S2 further includes: performing quantitative evaluation and physical verification on the complete multi-sensor vibration signal matrix, and calculating its mean square error, mean absolute error and cos similarity index;

[0033] Quantify the reconstruction accuracy and fitting effect, analyze the signal spectrum characteristics and mode shapes to ensure that the reconstructed signal conforms to the laws of structural dynamics.

[0034] Specifically, in this embodiment, the process proceeds to quantitative evaluation and physical verification to ensure that the reconstruction results are not only numerically reliable but also reasonable in terms of structural dynamics. First, in the quantitative evaluation stage, the mean square error (RMSE) and mean absolute error (MAE) between the reconstructed signal and the corresponding reference signal are calculated to measure amplitude deviation; simultaneously, the cosine similarity is calculated to assess the directional consistency of the two signal segments in vector space. The physical verification section focuses on the conformity to structural dynamic laws. Spectral analysis is performed on the reconstructed complete signal matrix to extract the natural frequencies, damping ratios, and principal modes, and these are compared with design parameters or historical health monitoring benchmarks. If the peak position, bandwidth, and energy distribution of the reconstructed signal are consistent with the existing modal parameters of the structure, it indicates that it is physically reliable.

[0035] Quantitative assessment and physical verification form a dual "numerical-physical" checkpoint: the former ensures that the reconstructed signal is statistically accurate enough, while the latter ensures that its spectrum and mode shape conform to the actual structural dynamics. The combination of the two effectively eliminates the risk of "seemingly good fit but containing false modes." As a result, engineers can confidently use the verified complete vibration signal matrix for modal identification, damage detection, and safety assessment, significantly improving the reliability and engineering usability of structural health monitoring data in sensor-deficient scenarios.

[0036] Preferably, before calling the trained diffusion model, the method further includes:

[0037] S01, acquire the vibration signal collected by the sensor assembly deployed on the engineering structure, preprocess it, and extract the original vibration signal matrix, wherein the preprocessing includes denoising, normalization and missing labeling.

[0038] Specifically, step S01 further includes: acquiring multi-channel vibration signals collected by multiple acceleration sensors or displacement sensors deployed on the engineering structure; stitching the acquired multi-channel vibration signals together to construct a vibration signal matrix containing vibration signals from all sensor channels. Where M is the number of sensors and T is the length of the time series;

[0039] The system automatically identifies missing regions in the vibration signal matrix, including continuous spatial missing regions and random temporal missing regions, and marks the missing parts of the vibration signal matrix with NaN values ​​to construct a missing region mask matrix M. observe This is used to identify whether each sensor has valid data at each time point. In the missing mask matrix, non-missing positions are marked as 1, and missing positions are marked as 0.

[0040] Calculate the mean of the non-missing data for each sensor channel, and fill in the missing locations with this mean to obtain a preliminary complete signal matrix X. filled ;

[0041] The filled signal matrix is ​​standardized, and the standard deviation of each channel is calculated. Data with a standard deviation less than a threshold are subjected to division by zero prevention according to the formula. The data from each channel is processed to generate a normalized signal matrix. This yields the original vibration signal matrix, ensuring that each channel is subjected to subsequent feature extraction and training at the same scale.

[0042] In this embodiment, the raw vibration records collected on-site undergo a series of preprocessing steps to lay the foundation for the data quality of subsequent model inputs. First, vibration responses are continuously collected using acceleration or displacement sensor components deployed at key parts of the engineering structure, obtaining multi-channel time series data. These channels are then stitched together according to a fixed sampling time sequence to form an initial vibration signal matrix, ensuring that all measurement point information is included within the same spatiotemporal framework. This matrix is ​​automatically scanned to identify continuous spatial gaps. For detected gaps, NaN is used to mark them, and a gap mask matrix is ​​generated simultaneously, recording 1 for valid data and 0 for gaps. This ensures that each subsequent processing step accurately distinguishes between "true values" and "areas to be restored," avoiding information confusion.

[0043] To mitigate the impact of the entire NaN region on statistical calculations, the mean of the non-missing data for each channel is further calculated, and this mean is used to backfill the missing cells of the corresponding channels, resulting in a preliminary complete signal matrix. This operation only serves as a "skeleton" support in the preprocessing stage and does not change the actual observed values, while ensuring the smooth execution of subsequent standardization steps. Next, the standard deviation of the preliminary complete signal matrix is ​​calculated for each channel. If the standard deviation is less than a set threshold, zero-prevention processing is performed to prevent noise amplification during normalization. Then, the mean and standard deviation are subtracted and divided in a cascaded manner for each channel, finally outputting a normalized signal matrix with a mean of 0 and a variance of 1, which is the original vibration signal matrix, and can be directly used as input for training or inference of the diffusion model.

[0044] S02, an adaptive multi-scale interpolation method is used to initially reconstruct the missing regions in the detected original vibration signal matrix, and a rough reconstructed signal is generated by integrating three types of constraint information: time, frequency and space.

[0045] Specifically, step S02 further includes: processing the signal matrix. Frequency domain analysis was performed by converting the signals from each sensor channel to the frequency domain using a Fast Fourier Transform (FFT) to obtain a complex spectrum matrix. Statistical features of the amplitude spectrum and energy spectrum are extracted from the complex spectrum matrix to construct a frequency domain feature matrix. , For frequency domain feature mapping operators;

[0046] For signal matrix Perform window partitioning and extract the time tensor within each window. Mask tensor and frequency domain tensor K is the number of sensor channels, and L is the window length. Frequency domain feature dimension;

[0047] In the spatial dimension, based on the weighted smoothing compensation of missing data from adjacent sensors, and by integrating the three types of constraint information—time, frequency, and space—a rough reconstruction signal is generated.

[0048] In this embodiment, an adaptive multi-scale interpolation method is used to initially reconstruct the detected missing regions to restore the overall trend and local details of the signal. This includes: preserving short-term fluctuation characteristics through sliding window interpolation in the time dimension; applying global spectral constraints in the frequency dimension to maintain the spectral consistency of the signal; and compensating for missing data based on weighted smoothing from adjacent sensors in the spatial dimension, thereby generating a coarsely reconstructed signal by integrating the three types of constraint information: time, frequency, and space.

[0049] Specifically, adaptive multi-scale interpolation is used to "coarsely repair" missing regions, providing an initial solution for the subsequent diffusion model that combines global trends and local features. First, a Fast Fourier Transform is performed channel-by-channel along the time axis on the normalized signal matrix to obtain a complex spectrum matrix. The mean and variance of the amplitude and energy spectra are extracted from this matrix and concatenated to form a frequency domain feature matrix. This matrix serves as a frequency domain mapping operator, repeatedly called in subsequent interpolation processes to constantly remind the algorithm "which frequency components should be preserved," avoiding high-frequency energy leakage caused by simple smoothing.

[0050] Next, the complete signal matrix is ​​slidably cut into several spatiotemporal windows of fixed length. Within each window, the time tensor, mask tensor, and the aforementioned frequency domain tensor are simultaneously extracted, giving the same window three perspectives: temporal fluctuations, missing locations, and spectral structure. For NaN points appearing within the window, the algorithm initiates spatial dimension weighted smoothing: centered on the missing sensor, weights are calculated based on its Euclidean distance and spectral similarity with neighboring measurement points, and the observations of adjacent channels are weighted and averaged to compensate for missing samples. Simultaneously, short-term fluctuations are preserved using local linear interpolation in time, and the energy distribution provided by the frequency domain feature matrix is ​​used as a soft constraint in frequency to prevent false peaks in the interpolation results that do not match the structural response.

[0051] By using window sliding and adaptive weight updates, time, frequency, and spatial information are uniformly incorporated into the same objective function, and a coarsely reconstructed signal is output after block-by-block optimization. Although the signal has not yet achieved high fidelity, it has successfully recovered the approximate amplitude range, main frequency components, and spatial continuity of the missing segments, significantly reducing the search space of subsequent diffusion models. At the same time, due to the "supervision" of the frequency domain feature matrix, the coarse restoration result is consistent with the real record in terms of the position and bandwidth of the main formant peaks, effectively suppressing the "oversmoothing" or "pseudo-oscillation" phenomena common in traditional interpolation, laying a reliable initial field for subsequent fine denoising and inversion.

[0052] S03, apply an artificial missing mask to the original vibration signal to simulate continuous spatial missing and random temporal missing, and record the missing position and preliminary interpolation results. Use the pseudo-missing sample and the corresponding complete signal as input for self-supervised training of the diffusion model, so that the model can fully capture the spatiotemporal features and missing patterns of the signal during the learning process, and separate the real data used for training and the pseudo-missing signal used for verification.

[0053] Specifically, step S03 further includes: processing the signal matrix. Perform missing detection and use a mask matrix. This indicates that when the value is 0, the signal matrix is ​​within the current mask matrix. The position is missing; when the value is 1, it indicates that the signal matrix is ​​in the current mask matrix. The positions are not missing. The missing status of each position in the signal is indicated by 1 for non-missing signals and 0 for missing signals. The collected raw data is divided into subsets of missing signals. Non-missing signal It consists of two parts. The mask matrix represents the non-missing signals;

[0054] The non-missing signal subset is masked again to obtain the pseudo-missing signal. This is used to calculate the loss function during subsequent validation. This is a pseudo-missing mask.

[0055] In this embodiment, firstly, missing value detection is performed on the original acquired vibration signal matrix: the program iterates through each channel and at each time step, recording 0 for NaN and 1 for valid values, thus obtaining a mask matrix. To further expand the verification scenario, a "pseudo-missing" mask is applied again to the non-missing signal subset: spatially continuous missing blocks and temporally random missing points are randomly generated according to the same proportion as the actual line drops, resulting in a new pseudo-missing mask. After multiplying the pseudo-missing mask point by point with the intact segment, artificial gaps appear in the originally continuous waveform, forming a pseudo-missing signal. This pseudo-missing signal and its corresponding intact segment constitute a pair of "missing-true" samples, specifically used for loss function calculation in the verification stage, enabling the model to obtain numerical feedback immediately after each inverse denoising without waiting for real missing data from the field.

[0056] Through a dual-channel strategy of "real missing data + pseudo missing data," the training and validation sets are strictly separated: the real missing data subset is used to enable the network to perceive complex disconnection patterns in real-world engineering scenarios, while the pseudo missing data subset provides a large number of repeatable and comparable evaluation samples. The two complement each other, ensuring efficient data utilization while avoiding inflated performance due to training-validation confusion. Ultimately, the diffusion model, within a self-supervised framework, fully exploits spatiotemporal features and missing data patterns, laying a robust sample foundation for subsequent high-fidelity reconstruction.

[0057] S04, Gaussian noise is filled into the signal in the missing region as the initial state for denoising inversion. Multimodal conditional features are extracted, and the four types of features are fused to generate the side information input of the diffusion model, which is used to guide the diffusion model for denoising and reconstruction, so that the model can learn the spatiotemporal distribution law of the missing signal in self-supervised training. The multimodal conditional features include: temporal feature embedding to capture temporal dependence and periodic information, spatial feature embedding to extract the correlation between sensors, trend feature embedding to capture low-frequency trends, and frequency domain feature embedding to extract the frequency domain structure by combining physical spectrum indicators and lightweight convolution.

[0058] Specifically, step S04 further includes: processing the signal matrix. The missing part is initialized with zero-mean Gaussian noise to obtain the initial input matrix for denoising inversion, which provides a random starting point for the stepwise denoising of the diffusion model to be trained.

[0059] Extracting the temporal features of the signal matrix within each window For each time point t, a position encoding vector PE(t) is constructed to represent time or phase information, and sine and cosine are used for encoding, with the formula being: , , `dt` is an integer, representing the temporal embedding dimension, used to iterate through and extract temporal information from different dimensions. For each window or entire sequence, construct its temporal embedding tensor. , as a time-conditional injection-diffusion model;

[0060] Extracting the spatial features of the signal matrix within each window The design utilizes a convolutional network to capture the spatial correlation and local coupling features between vibration signals from multiple sensors. It extracts local spatial responses through depthwise separable convolution, where depthwise convolution independently extracts temporal patterns within each channel, and pointwise convolution achieves linear blending mapping between channels. , It is a non-linear activation function with a kernel size of 3. For batch normalization, For the signal matrix within each window, For convolutional features, This involves passing through a one-dimensional convolutional layer with a kernel size of 1×1.

[0061] To further capture the global topological dependencies between sensors, an adjacency matrix based on the sensor spatial layout is constructed. The formula for the ij-th adjacency matrix is: , e raised to the power of Let i be the spatial location of the i-th sensor. Let j be the spatial location of the j-th sensor. For scale parameters;

[0062] Neighborhood information diffusion is achieved through graph convolution smoothing, and the formula is as follows: , The graph smoothing features are then fused with the convolutional features to obtain the final spatial embedding representation. , As the weight matrix, this spatial feature embedding module comprehensively utilizes local convolution and global topological information to achieve spatial collaborative modeling among multi-sensor responses;

[0063] Extract the trend features of the signal matrix within each window. One-dimensional convolution is applied to the time series of each sensor for local smoothing to obtain the smoothed signal. Its formula is The convolution kernel size is 5, and it operates only in the time dimension L to eliminate short-term noise fluctuations, preserve the main trend of the signal, and smooth the signal. Input one-way GRU module to extract global time trend T is the transpose. The global time trend is processed by linear mapping and nonlinear activation function to obtain the final trend embedding. , This is the weight matrix. The bias vector is a variable parameter during training, which provides low-frequency trend constraints for multimodal conditional inputs, enabling the diffusion model to maintain the overall temporal structure and long-term variation law of the signal during denoising and reconstruction.

[0064] Extract the frequency domain features of the signal matrix within each window. For frequency domain feature matrix Key physical frequency domain features, including the spectral centroid, are extracted based on the amplitude spectrum. Significance of the main peak With mid-frequency energy concentration , and These are the two largest peaks in the amplitude spectrum. Let f be the amplitude at frequency f, and F represent the upper limit of the sampling frequency;

[0065] The spectral centroid, main peak significance, and mid-frequency energy concentration are spliced ​​together to obtain the spliced ​​features. The amplitude spectrum is preprocessed using a lightweight convolutional neural network to extract high-dimensional spectral features. And obtain the specified dimension through linear mapping. The frequency domain embedding is obtained above. , feature F phys The features from the CNN are fused using gating to obtain the final frequency domain feature representation of each sensor within that window. As a diffusion model, it provides frequency domain constraints. , Here, represents the weights for gated fusion, and i represents the number of frequency sub-bands obtained by further slicing a time series L. This is the frequency domain feature embedding vector within the i-th frequency domain sub-band of the window. Let be the physical information feature vector within the i-th frequency domain sub-band of any window. Let be the feature vector extracted after one-dimensional convolution within the i-th frequency domain sub-band of any window. This is the multiplication operator. To flatten a multidimensional vector into a one-dimensional vector, This is a one-dimensional pooling operation; For amplitude, This indicates the number of frequency domain features extracted within a window;

[0066] Time characteristics Spatial features Trend characteristics Frequency domain characteristics By concatenating the data, the lateral information input of the diffusion model is obtained. .Will As a side information input to the diffusion model, it enables the model to simultaneously consider time, space, trend, frequency domain and missing information during the denoising and reconstruction process, thus achieving spatiotemporal signal recovery under multimodal constraints.

[0067] In this embodiment, four types of features are extracted in parallel for each spatiotemporal window, and then concatenated into a unified conditional tensor, which serves as the second input to the network backbone besides the noisy graph. In the time dimension, a sine-cosine position encoding formula is constructed for each sampling point t within the window, consistent with the classic Transformer form: even indices use sin, and odd indices use cos, with the frequency decaying exponentially with the dimension. The resulting temporal embedding tensor T_embed is injected into the model, enabling the denoising network to perceive the inherent periodicity and phase order of the signal, avoiding the generation of future information that violates temporal causality in missing segments. In the spatial dimension, depthwise separable convolution is first used to capture local responses: depthwise convolution slides independently within each sensor channel to extract single-point temporal patterns; pointwise convolution then performs cross-channel linear mixing to output convolutional features. Simultaneously, an adjacency matrix is ​​constructed based on the sensor's physical coordinates, and graph smoothing features are obtained through graph convolution smoothing, ensuring that each missing node can receive reliable observation information from its topological nearest neighbors.

[0068] The trend branch uses one-dimensional convolution to locally smooth the single-channel signal. The smoothed result is fed into a unidirectional GRU to update the hidden state step by step, extracting long-term evolution patterns. The hidden state is then linearly mapped and activated by GELU to obtain the trend embedding, which is used to maintain the low-frequency contour of the signal during denoising, preventing the model from mistakenly removing structural drifts that should be slowly changing as noise. The frequency domain branch takes into account both "physical indicators + data-driven" perspectives: on the one hand, it directly calculates the centroid, main peak significance, and mid-frequency energy concentration on the amplitude spectrum to form 3D physical features. On the other hand, it feeds the entire amplitude spectrum into a two-layer lightweight CNN to extract high-dimensional spectral features. The two types of features are fused through a gating mechanism, with the gating weights autonomously generated by the CNN features. This allows the model to respect physical spectral peaks while adaptively correcting details, effectively avoiding the information loss that may occur with traditional hand-crafted features.

[0069] After four-way feature extraction, the four features are concatenated along the channel dimension to form the side information input. This tensor is cascaded with the current noisy signal in each reverse denoising step of the diffusion model and fed into the U-Net backbone to guide the network in predicting the noise components to be subtracted. Since the side information input integrates temporal phase, spatial topology, low-frequency trends, and spectral structure, the model is always subject to multimodal joint constraints during the denoising process of missing regions. The generated signal segments conform to both local statistical characteristics and overall physical laws, thereby significantly improving reconstruction accuracy and engineering interpretability.

[0070] S05, the coarsely reconstructed signal and the fused multimodal conditional feature input diffusion model, are trained by progressively denoising and inverting under a self-supervised training framework. The parameters are continuously optimized through pseudo-missing samples until the model reaches the preset requirements, and the training ends, resulting in a well-trained diffusion model.

[0071] Specifically, step S05 further includes: during model training, in order to enable the diffusion model to learn the ability to recover the true signal from noise, Gaussian noise is gradually added based on the spurious missing signal through a forward diffusion process, the formula of which is: ,in, This is the cumulative noise attenuation coefficient. Let i be the noise intensity at time step i. Gaussian noise conforming to a standard normal distribution;

[0072] During the backdiffusion phase, through conditional networks The noise is predicted using the corresponding time step information, and the model-predicted noise is obtained by formulating the following formula: Using model-predicted noise The formula for predicting missing signals is as follows: This yields an instantaneous estimate of the original noise-free signal from the model, where, For noisy vibration signal data at time step t, The noise is the result of the model prediction during the denoising process. Based on the current noise data and predicted noise The estimated raw data;

[0073] During the iteration process, optimization is performed using a joint noise loss function and a frequency domain loss function, the formula of which is: , , , and All are weighting coefficients. For the total loss function, Let be the noise loss function. For frequency domain loss function, Let time step t be Gaussian noise that conforms to a standard normal distribution. Mathematical expectation under influence For mathematical expectation, The frequency is obtained by performing a Fourier transform on the response signal. It is the square of the L2 norm.

[0074] In this embodiment, the coarsely reconstructed signal and the fused multimodal conditional features are input into the diffusion model, and progressive denoising and inversion are performed under a self-supervised training framework. Specifically, the multimodal conditional features are fed into the diffusion model, and self-supervised training is initiated. The training loop adopts a classic diffusion-based noise addition and denoising framework: in the forward phase, Gaussian noise is added to the pseudo-missing signal step by step according to time step i. This process only needs to be performed at observation positions that have been marked as "valid," ensuring that NaN segments are not updated incorrectly, while enabling the model to learn to recover true fluctuations from pure noise in subsequent inversions. In the reverse phase, the U-Net backbone receives the current noisy signal, time step encoding, and side information, and outputs predicted noise. The instantaneous estimate of the original noiseless signal can be obtained using the standard update rules of the diffusion model. Thanks to the common constraints of the four types of embeddings in the side information—time, space, trend, and frequency domain—the noise in the missing region can simultaneously take into account local fluctuations and global modes, avoiding the pseudo-solution of "texture correct but dynamic error" common in traditional generative networks.

[0075] Furthermore, within the self-supervised framework, the model does not rely on external labels but instead performs self-comparative learning using pseudo-missing samples. The loss function consists of two parts: noise prediction loss and frequency domain consistency loss. Weighting coefficients are used to balance the optimization objectives in the time and frequency domains. Parameter optimization employs a joint loss function, where the noise loss function constrains the model's reconstruction of the continuity and local fluctuation characteristics of the time-domain signal, while the frequency domain loss function ensures that the recovered signal retains the spectral structure and modal characteristics of the original vibration signal. This allows the model to recover the spatiotemporal dependencies of missing data while preserving the global trend and local features of the signal.

[0076] As iterations proceed, the RMSE and cosine similarity on the validation set improve synchronously, indicating that the model has captured the spatiotemporal distribution patterns of structural vibration signals. Training automatically terminates when the validation loss fails to decrease for two consecutive epochs, resulting in a diffusion model with fixed weights. Model parameters are updated using the backpropagation algorithm, and the optimal model is obtained through continuous optimization iterations. Thanks to this self-supervised framework, the model demonstrates excellent generalization ability for continuous spatial or random temporal missing data unseen in real-world engineering. It can then be directly used with offline data from the field to output high-fidelity reconstructed signals that conform to structural dynamics, significantly improving the data integrity and analytical reliability of the monitoring system in sensor-constrained scenarios.

[0077] Please see Figures 3-9 In summary, this method addresses the problem of high-dimensional and complex mode loss in structural health monitoring caused by limited sensor deployment or data acquisition interruptions. It proposes a cascaded framework of "adaptive multi-scale interpolation - self-supervised diffusion inversion - multi-modal conditional embedding." By filling the missing regions with Gaussian noise and fusing temporal, spatial, trend, and frequency domain information during the preprocessing stage, the diffusion model can simultaneously adhere to the triple constraints of local waveform continuity, spectral structure preservation, and multi-sensor spatial collaboration during the inverse denoising process, thereby generating a reconstructed signal highly consistent with the actual structural response. The entire process requires no additional annotation; self-supervised training can be completed using only existing intact data from the field to construct pseudo-missing samples, significantly reducing implementation costs. Experimental verification shows that this method can still recover a complete vibration matrix with correct amplitude, phase, and modal characteristics even in scenarios with both spatially continuous and temporally random missing data. This provides a highly reliable and usable data foundation for subsequent modal identification, damage detection, and safety assessment, effectively improving the robustness and intelligence of structural health monitoring systems in complex engineering environments.

[0078] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for spatiotemporal reconstruction of vibration signals based on a multimodal conditional diffusion model, characterized in that, include: The missing vibration signals collected by the sensor components deployed on the engineering structure are acquired. The vibration signals are input into a pre-trained diffusion model and spatiotemporal reconstruction is performed to obtain a complete multi-sensor vibration signal matrix. The missing regions of the vibration signals are filled with Gaussian noise and used as the initial input. Under the constraints of multimodal embedding, the amplitude and temporal characteristics of the missing region signals are gradually recovered through iterative denoising and inversion. The complete multi-sensor vibration signal matrix is ​​quantitatively evaluated and physically verified to ensure that the reconstructed signal conforms to the laws of structural dynamics. Before calling the trained diffusion model, the process also includes: filling the signal in the missing region with Gaussian noise as the initial state for denoising inversion, and extracting multimodal conditional features, wherein the multimodal conditional features include: frequency domain feature embedding combined with physical spectrum indicators and lightweight convolution to extract frequency domain structure.

2. The method for spatiotemporal reconstruction of vibration signals based on a multimodal conditional diffusion model according to claim 1, characterized in that, The vibration signal is input into a pre-trained diffusion model for spatiotemporal reconstruction processing to obtain a complete multi-sensor vibration signal matrix, specifically: The missing vibration signals collected by the sensor components deployed on the engineering structure are acquired, and Gaussian noise is gradually added to them to fill the Gaussian noise. In the reverse denoising process, a pre-trained diffusion model is used to predict noise after filling, and then denoise the noise to obtain the final reconstructed vibration signal X. recon ; For the reconstructed vibration signal X recon By performing inverse normalization, the signal is restored to the physical dimensions of the engineering structure, resulting in a complete multi-sensor vibration signal matrix. ,in, To represent element-wise multiplication, The standard deviation of each channel of the original signal before normalization. This represents the mean of all channels of the original signal before normalization.

3. The method for spatiotemporal reconstruction of vibration signals based on a multimodal conditional diffusion model according to claim 1, characterized in that, The complete multi-sensor vibration signal matrix is ​​quantitatively evaluated and physically verified to ensure that the reconstructed signal conforms to the laws of structural dynamics. Specifically: The complete multi-sensor vibration signal matrix is ​​quantitatively evaluated and physically verified, and its mean square error, mean absolute error and cosine similarity index are calculated. Quantify the reconstruction accuracy and fitting effect, analyze the signal spectrum characteristics and mode shapes to ensure that the reconstructed signal conforms to the laws of structural dynamics.

4. The method for spatiotemporal reconstruction of vibration signals based on a multimodal conditional diffusion model according to claim 1, characterized in that, Before calling the trained diffusion model, the following is also included: Vibration signals collected by sensor components deployed on the engineering structure are acquired and preprocessed to extract the original vibration signal matrix. The preprocessing includes denoising, normalization, and missing labeling. An adaptive multi-scale interpolation method is used to initially reconstruct the missing regions in the detected original vibration signal matrix. By integrating time, frequency, and spatial constraint information, a coarse reconstructed signal is generated. An artificial missing mask is applied to the original vibration signal to simulate continuous spatial missing and random temporal missing, and the missing position and preliminary interpolation results are recorded. The pseudo-missing sample and the corresponding complete signal are used as input to separate the real data for training and the pseudo-missing signal for verification. The signal in the missing region is filled with Gaussian noise as the initial state for denoising inversion. Multimodal conditional features are extracted, and the four types of features are fused to generate the side information input of the diffusion model, which is used to guide the diffusion model for denoising and reconstruction. This allows the model to learn the spatiotemporal distribution pattern of the missing signal in self-supervised training. The multimodal conditional features also include: temporal feature embedding to capture temporal dependence and periodic information, spatial feature embedding to extract the correlation between sensors, and trend feature embedding to capture low-frequency trends. The coarsely reconstructed signal and the fused multimodal conditional features are input into the diffusion model. Under a self-supervised training framework, the model is gradually denoised and inverted. The parameters are continuously optimized using pseudo-missing samples until the model meets the preset requirements. The training ends, and the trained diffusion model is obtained.

5. The method for spatiotemporal reconstruction of vibration signals based on a multimodal conditional diffusion model according to claim 4, characterized in that, The vibration signals collected by the sensor assembly deployed on the engineering structure are acquired, preprocessed, and the original vibration signal matrix is ​​extracted, specifically as follows: The system acquires multi-channel vibration signals from multiple accelerometers or displacement sensors deployed on the engineering structure, and then stitches these signals together to construct a vibration signal matrix containing vibration signals from all sensor channels. Where M is the number of sensors and T is the length of the time series; The system automatically identifies missing regions in the vibration signal matrix, including continuous spatial missing regions and random temporal missing regions, and marks the missing parts of the vibration signal matrix with NaN values ​​to construct a missing region mask matrix M. observe This is used to identify whether each sensor has valid data at each time point. In the missing mask matrix, non-missing positions are marked as 1, and missing positions are marked as 0. Calculate the mean of the non-missing data for each sensor channel, and fill in the missing locations with this mean to obtain a preliminary complete signal matrix X. filled ; The filled signal matrix is ​​standardized, and the standard deviation of each channel is calculated. Data with a standard deviation less than a threshold are subjected to division by zero prevention according to the formula. The data from each channel is processed to generate a normalized signal matrix. The original vibration signal matrix is ​​obtained.

6. The method for spatiotemporal reconstruction of vibration signals based on a multimodal conditional diffusion model according to claim 5, characterized in that, An adaptive multi-scale interpolation method is used to initially reconstruct the missing regions in the detected original vibration signal matrix. By integrating time, frequency, and spatial constraint information, a coarse reconstructed signal is generated, specifically: For signal matrix Frequency domain analysis was performed by converting the signals from each sensor channel to the frequency domain using a Fast Fourier Transform (FFT) to obtain a complex spectrum matrix. Statistical features of the amplitude spectrum and energy spectrum are extracted from the complex spectrum matrix to construct a frequency domain feature matrix. , For frequency domain feature mapping operators; For signal matrix Perform window partitioning and extract the time tensor within each window. Mask tensor and frequency domain tensor K is the number of sensor channels, and L is the window length. Frequency domain feature dimension; In the spatial dimension, based on the weighted smoothing compensation of missing data from adjacent sensors, and by integrating the three types of constraint information—time, frequency, and space—a rough reconstruction signal is generated.

7. The method for spatiotemporal reconstruction of vibration signals based on a multimodal conditional diffusion model according to claim 6, characterized in that, An artificial missing mask is applied to the original vibration signal to simulate continuous spatial missingness and random temporal missingness. The missing locations and preliminary interpolation results are recorded. The pseudo-missing samples and their corresponding complete signals are used as input to separate the real data for training and the pseudo-missing signals for validation. Specifically: For signal matrix Perform missing detection and use a mask matrix. This indicates that when the value is 0, the signal matrix is ​​within the current mask matrix. The position is missing; when the value is 1, it indicates that the signal matrix is ​​in the current mask matrix. The location is not missing; the collected raw data is divided into subsets of missing signals. Non-missing signal It consists of two parts. The mask matrix represents the non-missing signals; The non-missing signal subset is masked again to obtain the pseudo-missing signal. This is used to calculate the loss function during subsequent validation. This is a pseudo-missing mask.

8. The method for spatiotemporal reconstruction of vibration signals based on a multimodal conditional diffusion model according to claim 7, characterized in that, The signal in the missing region is filled with Gaussian noise as the initial state for denoising inversion. Multimodal conditional features are extracted, and the four types of features are fused to generate the side information input of the diffusion model, specifically: For signal matrix The missing part is initialized with zero-mean Gaussian noise to obtain the initial input matrix for denoising inversion, which serves as a random starting point for the stepwise denoising of the diffusion model to be trained. Extracting the temporal features of the signal matrix within each window For each time point t, a position encoding vector PE(t) is constructed to represent time or phase information, and sine and cosine are used for encoding, with the formula being: , , Let dt be an integer and dt be the temporal embedding dimension. For each window or integer sequence, construct its temporal embedding tensor. , as a time-conditional injection-diffusion model; Extracting the spatial features of the signal matrix within each window The design utilizes a convolutional network to capture the spatial correlation and local coupling features between vibration signals from multiple sensors. It extracts local spatial responses through depthwise separable convolution, where depthwise convolution independently extracts temporal patterns within each channel, and pointwise convolution achieves linear blending mapping between channels. , It is a non-linear activation function. For batch normalization, For the signal matrix within each window, For convolutional features, This involves passing through a one-dimensional convolutional layer with a kernel size of 1×1. Constructing an adjacency matrix based on sensor spatial layout The formula for the ij-th adjacency matrix is: , To represent e raised to the power of Let i be the spatial location of the i-th sensor. Let j be the spatial location of the j-th sensor. For scale parameters; Neighborhood information diffusion is achieved through graph convolution smoothing, and the formula is as follows: , The graph smoothing features are then fused with the convolutional features to obtain the final spatial embedding representation. , This is the weight matrix; Extract the trend features of the signal matrix within each window. One-dimensional convolution is applied to the time series of each sensor for local smoothing to obtain the smoothed signal. Its formula is and the smoothed signal Input one-way GRU module to extract global time trend T is the transpose. The global time trend is processed by linear mapping and nonlinear activation function to obtain the final trend embedding. , This is the weight matrix. Here, represents the bias vector, and represents the variable parameters during the training process. Extract the frequency domain features of the signal matrix within each window. For frequency domain feature matrix Key physical frequency domain features, including the spectral centroid, are extracted based on the amplitude spectrum. Significance of the main peak With mid-frequency energy concentration , and These are the two largest peaks in the amplitude spectrum. Let f be the amplitude at frequency f, and F represent the upper limit of the sampling frequency; The spectral centroid, main peak significance, and mid-frequency energy concentration are spliced ​​together to obtain the spliced ​​features. The amplitude spectrum is preprocessed using a lightweight convolutional neural network to extract high-dimensional spectral features. And obtain the specified dimension through linear mapping. The frequency domain embedding is obtained above. , feature F phys The features from the CNN are fused using gating to obtain the final frequency domain feature representation of each sensor within that window. As a diffusion model, it provides frequency domain constraints. , Here, represents the weights for gated fusion, and i represents the number of frequency sub-bands obtained by further slicing a time series L. This is the frequency domain feature embedding vector within the i-th frequency domain sub-band of the window. Let be the physical information feature vector within the i-th frequency domain sub-band of any window. Let be the feature vector extracted after one-dimensional convolution within the i-th frequency domain sub-band of any window. This is the multiplication operator. To flatten a multidimensional vector into a one-dimensional vector, This is a one-dimensional pooling operation; For amplitude, This indicates the number of frequency domain features extracted within a window; Time characteristics Spatial features Trend characteristics Frequency domain characteristics By concatenating the data, the lateral information input of the diffusion model is obtained. .

9. The method for spatiotemporal reconstruction of vibration signals based on a multimodal conditional diffusion model according to claim 8, characterized in that, The coarsely reconstructed signal and the fused multimodal conditional features are input into the diffusion model. Under a self-supervised training framework, noise is progressively denoised and inverted for training. Parameters are continuously optimized using pseudo-missing samples until the model meets preset requirements, at which point training ends, yielding a trained diffusion model. Specifically: During model training, Gaussian noise is gradually added through a forward diffusion process based on the spurious missing signal, as shown in the formula: ,in, This is the cumulative noise attenuation coefficient. Let i be the noise intensity at time step i. Gaussian noise conforming to a standard normal distribution; During the backdiffusion phase, through conditional networks Using the corresponding time step information to predict noise, the noise predicted by the model after conditional information fusion is obtained, and its formula is: Noise predicted by the model after fusion of conditional information The formula for predicting missing signals is as follows: This yields an instantaneous estimate of the original noise-free signal from the model, where, For noisy vibration signal data at time step t, The noise is the result of the model prediction during the denoising process. Based on the current noise data and predicted noise The estimated raw data; During the iteration process, optimization is performed using a joint noise loss function and a frequency domain loss function, the formula of which is: , , , and All are weighting coefficients. For the total loss function, Let be the noise loss function. For frequency domain loss function, Let time step t be Gaussian noise that conforms to a standard normal distribution. Mathematical expectation under influence For mathematical expectation, The frequency is obtained by performing a Fourier transform on the response signal. It is the square of the L2 norm.

Citation Information

Patent Citations

  • Traffic data interpolation method based on time-frequency feature fusion and conditional diffusion model

    CN120653898A

  • Method, device and equipment for identifying modal parameters of space-time missing vibration signals

    CN121210857A