Sensor data reconstruction method based on multi-source time sequence generative adversarial network
Through multi-source timing generation, the nonlinear relationship of sensor data is learned by adversarial network learning, the data loss problem caused by sensor damage is solved, and the high-precision data reconstruction of the bridge health monitoring system is realized, which improves the reliability and accuracy of the system.
Patent Information
- Application Number
- CN202510463265.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-29
AI Technical Summary
In the existing bridge health monitoring system, sensor damage leads to data loss, and traditional data reconstruction methods are not effective in large-scale or complex modes, affecting the reliability and accuracy of the monitoring system.
The method of generating adversarial networks based on multi-source timing is adopted, and complex nonlinear relationships of sensor data are learned through deep learning technology, generators and discriminators are built, and adaptive weighted loss functions and composite loss functions are used for training to generate realistic reconstruction data.
It realizes high-precision data reconstruction of multi-sensors, which is suitable for bridge structures in complex operating environments, has high robustness and noise immunity, simplifies data processing flow, and improves the accuracy and reliability of data reconstruction.
Smart Images

Figure CN120386988A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bridge structure health monitoring, and particularly relates to a sensor data reconstruction method based on a multi-source time series generative adversarial network. Background Art
[0002] As an important transportation infrastructure, the health status of bridges is crucial for ensuring traffic safety and smoothness; traditional bridge health monitoring mainly relies on various sensors (such as acceleration sensors) installed on the bridge structure to collect data for evaluating the health status of the bridge. However, sensors may be damaged for various reasons, resulting in data loss, thus affecting the accurate evaluation of the bridge health status; therefore, how to effectively reconstruct the data of damaged sensors is an urgent problem to be solved currently.
[0003] In modern bridge health monitoring systems, the data collected by sensors is the key to analyzing the structural integrity of bridges and predicting potential problems. This data helps engineers and maintenance teams evaluate the current state of the bridge and formulate repair or reinforcement plans. Traditionally, bridge health monitoring systems rely on a large number of sensors, such as accelerometers, strain gauges, and displacement sensors, which can provide real-time feedback on various parts of the bridge under different environmental and load conditions. However, these sensors are often exposed to harsh environments, such as extreme temperature changes, humidity, and mechanical wear, which may all cause sensor failures, resulting in incomplete or missing data.
[0004] Existing data reconstruction methods mainly include linear interpolation, polynomial interpolation [Yan Jie. Multiple Imputation of Missing Data [M]. Chongqing: Chongqing University Press, 2017] and methods based on statistics [Jin Yongjin, Shao Jun. Statistical Processing of Missing Data [M]. Beijing: China Statistics Press, 2009]. These methods are relatively effective in dealing with small-scale data missing, but they perform poorly when facing large-scale or continuous data missing. In addition, these methods usually assume that there is a linear relationship between data, ignoring the complex non-linear patterns that may appear in actual monitoring data. Therefore, when the scale of data missing is large or the data has complex patterns, these traditional methods cannot accurately reconstruct the data, thus affecting the reliability and accuracy of the entire monitoring system. Summary of the Invention
[0005] In view of the above, the present invention provides a sensor data reconstruction method based on a multi-source time series generative adversarial network, which effectively improves the accuracy and reliability of data reconstruction by using deep learning technology to learn and simulate the complex non-linear relationship of sensor data.
[0006] A sensor data reconstruction method based on a multi-source time series generative adversarial network includes the following steps:
[0007] (1) Obtain sensor data of the bridge structure under different load excitations at different times through sensors in the bridge health monitoring system, and slice the sensor data to obtain a large number of data samples;
[0008] (2) Perform a masking operation on the data samples, and use the data samples before masking as labels, that is, real data;
[0009] (3) Build a multi-source time-series generative adversarial network, including:
[0010] A generator for reconstructing the input data samples to generate reconstructed data;
[0011] A discriminator for discriminating the authenticity of real data and reconstructed data;
[0012] (4) Input the masked data samples and their labels into the multi-source time-series generative adversarial network for training;
[0013] (5) Use the trained multi-source time-series generative adversarial network to reconstruct the sensor data with partial data missing to restore it to complete data.
[0014] Further, the sensor data is an n×l data matrix, where n is the number of sensors and l is the length of the data collected by the sensors, that is, the number of sampling points; in step (1), a sliding window is used to slice the sensor data, the window length is m (1 < m < l), and the sliding step is s (s ≤ m), to obtain multiple data samples of size n×m.
[0015] Further, in step (2), a masking operation is performed on the data samples, that is, assuming some damaged sensors, the data of these sensors is set to 0.
[0016] Further, the multi-source time-series generative adversarial network adopts a 2D-DCGAN (2D-Deep Convolutional Generative Adversarial Networks, two-dimensional deep convolutional generative adversarial network) architecture; compared with the traditional generative adversarial network architecture, it can better learn the deep features and relationships between data and is more suitable for data reconstruction tasks.
[0017] Furthermore, the generator adopts an encoder-decoder structure. The encoder therein is used for feature extraction and analysis of data samples to achieve the encoding function. The encoder consists of five cascaded convolutional layers. Except for the first convolutional layer, the output of each of the other convolutional layers is activated through batch normalization operation and the LeakyReLU function. The decoder is used for performing deconvolution operations on the features generated by the encoder, decoding the features and reconstructing them into time series data. The decoder consists of five cascaded deconvolutional layers. Except for the last deconvolutional layer, the output of each of the other deconvolutional layers is activated through batch normalization operation and the ReLU function. The discriminator consists of five cascaded convolutional layers D1 to D5. The output of D1 is only activated by the LeakyReLU function. The outputs of D2 to D4 are all activated through batch normalization operation and the LeakyReLU function. The output of D5 is not subjected to batch normalization operation and function activation.
[0018] Furthermore, during the training process of the multi-source time series generative adversarial network in step (4), the loss function of the generator adopts an adaptive weighted loss function, introducing corresponding weight coefficients for each sensor, dynamically calculating them during training to balance sensor accuracy and error. The loss function of the discriminator adopts a composite loss function, balancing its performance with that of the generator during training to avoid problems of gradient disappearance and explosion, and achieving fast convergence and efficient training.
[0019] Furthermore, the calculation expression of the loss function of the generator is as follows:
[0020]
[0021] Where: Loss G is the loss function of the generator, W(p, q) is the Wasserstein distance between p and q, p and q are the probability distributions of the real data and the reconstructed data respectively, λ i represents the weight coefficient of the i-th sensor, X i and X' i are respectively the real data and the reconstructed data of the i-th sensor, MSE(X i , X' i ) is the mean square error between X i and X' i , and n is the number of sensors.
[0022] Furthermore, the calculation expression of the weight coefficient λ i is as follows:
[0023]
[0024] Where: abs() represents taking the absolute value of each element in the vector, and mean() represents calculating the average value of all elements in the vector.
[0025] Furthermore, the calculation expression of the loss function of the discriminator is as follows:
[0026] Loss D = W(p, q) + L MSE (X, X′)
[0027] where: Loss D is the loss function of the discriminator, W(p, q) is the Wasserstein distance between p and q, p and q are the probability distributions of the real data and the reconstructed data respectively, and L MSE (X, X′) is the mean square error between X and X′, and X and X′ are the real data and the reconstructed data respectively.
[0028] Furthermore, the multi-source time-series generative adversarial network learns the data relationship between sensors using the sensor data during normal working periods. Due to the masking operation, the generator cannot know the masked sensor data and can only reconstruct based on the existing normally working sensor data to generate reconstructed data; the discriminator obtains the real data of the assumed partially damaged sensors and the reconstructed data generated by the generator, and judges the authenticity of these two types of data; the generator continuously generates reconstructed data to try to deceive the discriminator into identifying it as true, while the discriminator tries to correctly distinguish the authenticity of the reconstructed data. Through this adversarial process, the generator and the discriminator continuously improve their respective performances, and finally enable the generator to generate sufficiently realistic reconstructed data, which is almost identical to the real data.
[0029] Compared with the prior art, the present invention has the following beneficial technical effects:
[0030] 1. The sensor data reconstruction method of the present invention is used for reconstructing and restoring missing data of civil engineering structures, especially bridge structures, and can simultaneously achieve high-precision data reconstruction of multiple sensors.
[0031] 2. Before data reconstruction using the method of the present invention, operations such as denoising, filtering, and normalization are not required, thus ensuring the true characteristics of the original data to the greatest extent.
[0032] 3. The method of the present invention can be applied to real long-span complex structure system bridges in complex operating environments, such as cable-stayed bridges and suspension bridges, and has high robustness, anti-noise ability, and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a schematic flow chart of the sensor data reconstruction method based on the multi-source time-series generative adversarial network of the present invention.
[0034] Figure 2(a) is a schematic diagram of the structure of a three-span continuous beam bridge with sensors installed in an embodiment of the present invention.
[0035] Figure 2(b) is a schematic diagram of the finite element model of the three-span continuous beam bridge structure in the embodiment of the present invention.
[0036] Figure 3(a) is a time-domain waveform diagram of the acceleration signal obtained by reconstructing the sensor data of the three-span continuous beam bridge structure using the method of the present invention.
[0037] Figure 3(b) is a frequency-domain waveform diagram of the acceleration signal obtained by reconstructing the sensor data of the three-span continuous beam bridge structure using the method of the present invention.
[0038] Figure 4 It is a schematic diagram for comparing the reconstruction accuracies of the method of the present invention and two other existing data reconstruction methods under different training epochs. Detailed implementation manners
[0039] In order to describe the present invention more specifically, the technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0040] As Figure 1 shown, the sensor data reconstruction method of the present invention based on the multi-source time series generative adversarial network includes the following steps:
[0041] (1) Design the network structure and build a multi-source time series generative adversarial network.
[0042] According to the number of sensors to be analyzed, design the network structures of the generator and discriminator to complete the construction of the multi-source time series generative adversarial network; the multi-source time series generative adversarial network adopts the two-dimensional deep convolutional generative adversarial network 2D-DCGAN. The generator adopts an encoder-decoder structure, where the encoder extracts and analyzes the features of the data to achieve the encoding function; the decoder performs a deconvolution operation on the encoder to decode and reconstruct the encoded features into time series data, aiming to generate time series data as consistent as possible with the real data. The discriminator is used to evaluate the authenticity of the reconstructed time series data and, at the same time, compete with the generator to improve the performance of both.
[0043] In this embodiment, the encoder consists of five convolutional layers and corresponding batch normalization layers. Except for the first convolutional layer, batch normalization is performed after each convolution, and the subsequent activation function selects the LeakyReLU function. The decoder consists of five deconvolution layers and batch normalization layers. Batch normalization is performed after each deconvolution, and the subsequent activation function selects the ReLU function. Batch normalization and activation are not performed on the last layer.
[0044] The discriminator consists of five convolutional layers. Batch normalization is not performed on the first and the last layers. The LeakyReLU function is used for activation after the first four layers, and no activation function calculation is performed on the last layer. The discriminator finally outputs a real number. Different from a general discriminator that uses a sigmoid activation function in the last layer to output a result between 0 and 1, no activation is performed here to facilitate the calculation of the subsequent proposed composite loss function.
[0045] (2) Collect structural response data and perform data slicing to obtain data samples.
[0046] Using the structural health monitoring system arranged on the structure, collect the response signals of its acceleration sensors at different time periods to obtain structural response data; the collected data should cover as long a time range as possible, including the responses of the structure under various loads (wind loads of different magnitudes, temperature loads, vehicle loads, etc.), so that the network can learn the relationships between the structural response data under different environmental load excitations, enhancing the robustness and generalization ability of the network.
[0047] Construct the original data matrix of size n×l from the response signals, where n is the number of sensors and l is the total length of the data collected under the corresponding working conditions, that is, the number of sampling points.
[0048] Use a sliding window to perform data slicing on the original data matrix. The window length is m (m>l), and the sliding step is s (s≤m) to obtain multiple data samples of size n×m.
[0049] (3) Mask operation and sample label setting.
[0050] Perform a mask operation on the data samples, set the assumed damaged sensor data to 0, and use the data before masking as the sample label for the calculation of the generator loss function, and at the same time use it together with the reconstructed data generated by the generator for the training of the discriminator.
[0051] (4) Construct the loss functions of the generator and the discriminator.
[0052] The generator adopts an adaptive weighted loss function, introducing corresponding weight coefficients for each sensor, which are dynamically calculated during training to balance sensor accuracy and error; the discriminator adopts a composite loss function to balance its performance with that of the generator during training, avoiding the problems of gradient vanishing and explosion, and achieving fast convergence and efficient training.
[0053] The adaptive weighted loss function aims to balance the error between the reconstructed data and the real data, and at the same time takes into account the differences in the importance of different sensor data to achieve the effective reconstruction of multi-sensor data. Specifically, first, the mean square error (MSE) is calculated separately for the reconstructed data of each sensor, and then a weight coefficient is introduced to perform weighted summation on the MSE to obtain the final loss function. The specific calculation formula is as follows:
[0054]
[0055] loss i = MSE(X i , X' i )
[0056] where: X i and X' i represent the real data and the reconstructed data of the i-th acceleration sensor respectively, n is the number of sensors, λ i is the weight coefficient, loss i is the loss function of each sensor data, calculated using the mean square error.
[0057] The loss function of the discriminator adopts a composite loss function based on the Wasserstein distance and the mean square error loss. The calculation formula of the composite loss function is as follows:
[0058] Loss D = W(p, q) + L MSE
[0059] where: W(p, q) represents the Wasserstein distance, L MSE is the MSE loss.
[0060] The calculation formula of the Wasserstein distance is as follows:
[0061]
[0062] where: inf represents the infimum, that is, the minimum value among all possible cases, Π(p, q) represents all possible joint distributions γ, whose marginal distributions are p and q respectively; E (x,y)~γ [∥x - y∥] represents the expected value of the distance between the random variables x (from distribution p) and y (from distribution q) under the joint distribution γ, and ∥x - y∥ usually adopts the Euclidean distance.
[0063] (5) Joint training of the generator and the discriminator.
[0064] Use the data of all sensors during normal working periods for network training to learn the relationships between sensor data. Since masking operations were performed in step (3), the generator cannot predict the masked sensor data. It reconstructs based on the data of the normally working sensors, and the generated reconstructed data will be compared with the sample labels set in step (3) to calculate the loss function.
[0065] The discriminator receives the real data of the supposedly damaged sensors and the reconstructed data generated by the generator and judges their authenticity. The generator continuously generates reconstructed data to try to deceive the discriminator into identifying it as real; while the discriminator tries to correctly distinguish between real and fake data. Through this adversarial process, the generator and the discriminator can continuously improve their performance, and ultimately enable the generator to generate sufficiently realistic reconstructed data that is almost identical to the real data.
[0066] Hyperparameters including learning rate, batch size, number of training epochs, and optimizer need to be set during training; generally, a relatively small value such as 0.0002 is used for the learning rate to avoid gradient vanishing and gradient explosion during training; the batch size is selected according to the device performance, and generally 32 or 64 can be chosen; the number of training epochs is generally 200 - 500 epochs to ensure sufficient network training and convergence of the loss function; generally, Adam is chosen as the optimizer, and its advantage lies in that it combines the characteristics of momentum and adaptive learning rate, and can effectively handle sparse gradients and different parameter update frequencies, thus improving the convergence speed and performance when training deep learning models.
[0067] (6) Use the trained network for data reconstruction.
[0068] Utilize the trained multi - source time - series generative adversarial network to reconstruct the data samples during the data - missing period. When one or several sensors are damaged and cause data loss, input the data of the normally working sensors into the trained multi - source time - series generative adversarial network. Based on the learned relationships between sensor data, reconstruct the data of the damaged sensors to recover the lost data.
[0069] Using the above steps to reconstruct the data of the damaged acceleration sensor can achieve the recovery of missing data.
[0070] Embodiment
[0071] In this embodiment, a three - span continuous beam bridge is taken as an example. As shown in Figure 2(a), data reconstruction is performed on the sensors installed on the bridge to achieve the recovery of missing data of the faulty and damaged sensors. The specific process is as follows:
[0072] Step S1: Establish a finite - element model of the three - span continuous beam bridge structure and install 5 acceleration sensors on the bridge, as shown in Figure 2(b).
[0073] Step S2: Collect the response data of each acceleration sensor on the structure.
[0074] In this embodiment, n = 5 acceleration sensors are installed on the three-span continuous beam bridge, and l = 200,000 data points are collected for each working condition. Therefore, the size of the data set for each working condition is 5 × 200,000.
[0075] Step S3: Perform data slicing to obtain data samples.
[0076] In this implementation, the sliding window method is used for data slicing. The window length is selected as m = 512, and the step size is s = 128. The window length is selected as an integer power of 2 to facilitate subsequent training and feature extraction of the convolutional autoencoder. The sliding step size is selected as one-fourth of the window length, which not only ensures the difference of data samples but also plays a role in making full use of the data and expanding the sample size. After data slicing, data samples of size 5 × 512 are obtained.
[0077] Step S4: Mask operation and sample label setting.
[0078] Perform a mask operation on the data of the assumed damaged sensor, set it to 0, and at the same time use its real data as a label for the next step of training.
[0079] Step S5: Use the data during the period when all sensors are working properly for training.
[0080] Put the data of all sensors during the normal working period into the multi-source time series generative adversarial network for training to learn the relationship between the data of each sensor. First, the generator trains and reconstructs the masked data samples to obtain the reconstructed data of all assumed damaged sensors, and then compares it with the real data of all assumed damaged sensors to calculate the weighted loss function; the discriminator simultaneously discriminates the reconstructed data and the real data and calculates the composite loss function to backpropagate and optimize the network performance. In order to reduce the loss function, the generator improves its own performance during training and generates more real reconstructed data, and finally makes it basically consistent with the real data.
[0081] Step S6: Use the trained network for data reconstruction.
[0082] Use the multi-source time series generative adversarial network trained in S5 to reconstruct the data samples during the data missing period, and recover the data of the damaged sensors according to the data of the normally operating sensors.
[0083] In this embodiment, the data set can adopt the structural acceleration signals collected from the actual structure, or the simulation data generated by finite element simulation. When adopting the actual structure data, operations such as filtering and denoising can be omitted, simplifying the data analysis process and steps and improving the algorithm efficiency.
[0084] As shown in FIGS. 3(a) and 3(b), by using the method of the present invention for sensor data reconstruction, high-precision data reconstruction results can be obtained, and the reconstructed data is consistent with the real data both in the time domain and the frequency domain.
[0085] As Figure 4 shown, the method of the present invention (MTSGAN) is compared with the deep convolutional generative adversarial network (DCGAN) and the convolutional autoencoder (CAE). The accuracy is significantly higher than that of DCGAN and CAE in each round, demonstrating the superiority of the method of the present invention in data reconstruction.
[0086] The above description of the embodiments is to enable those of ordinary skill in the art to understand and apply the present invention. Obviously, those who are familiar with the technology in this field can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art according to the disclosure of the present invention should be within the protection scope of the present invention.
Claims
1. A sensor data reconstruction method based on a multi-source time-series generative adversarial network, comprising the following steps: (1) Obtain sensor data of the bridge structure under different load excitations at different times through sensors in the bridge health monitoring system, and slice the sensor data to obtain a large number of data samples; (2) Perform a masking operation on the data samples, and use the data samples before masking as labels, i.e., real data; (3) Build a multi-source time-series generative adversarial network, including: A generator for reconstructing the input data samples to generate reconstructed data; A discriminator for discriminating the authenticity of real data and reconstructed data; (4) Input the masked data samples and their labels into the multi-source time-series generative adversarial network for training; (5) Use the trained multi-source time-series generative adversarial network to reconstruct the sensor data with partial data missing to restore it to complete data.
2. The sensor data reconstruction method based on a multi-source temporal generative adversarial network according to claim 1, wherein: The sensor data is an n×l data matrix, where n is the number of sensors and l is the length of the data collected by the sensors, i.e., the number of sampling points; in step (1), a sliding window is used to slice the sensor data, the window length is m, and the sliding step is s, to obtain a plurality of data samples of size n×m.
3. The sensor data reconstruction method based on a multi-source temporal generative adversarial network according to claim 1, wherein: In step (2), a masking operation is performed on the data samples, that is, assuming some damaged sensors, the data of these sensors is set to 0.
4. The sensor data reconstruction method based on a multi-source temporal generative adversarial network according to claim 1, characterized in that: The multi-source time-series generative adversarial network adopts a 2D-DCGAN architecture.
5. The sensor data reconstruction method based on a multi-source temporal generative adversarial network according to claim 1, wherein: The generator adopts an encoder-decoder structure, where the encoder is used to extract and analyze the features of the data samples to achieve the encoding function. The encoder consists of five cascaded convolutional layers. Except for the first convolutional layer, the output of each other convolutional layer is activated by batch normalization operation and the LeakyReLU function; the decoder is used to perform deconvolution operations on the features generated by the encoder, decode the features and reconstruct them into time-series data. The decoder consists of five cascaded deconvolutional layers. Except for the last deconvolutional layer, the output of each other deconvolutional layer is activated by batch normalization operation and the ReLU function; the discriminator consists of five cascaded convolutional layers D1 to D5, where the output of D1 is only activated by the LeakyReLU function, the outputs of D2 to D4 are both activated by batch normalization operation and the LeakyReLU function, and the output of D5 is not subjected to batch normalization operation and function activation.
6. The sensor data reconstruction method based on a multi-source temporal generative adversarial network according to claim 1, characterized in that: During the training process of the multi-source time-series generative adversarial network in step (4), the loss function of the generator adopts an adaptive weighted loss function, and the loss function of the discriminator adopts a composite loss function.
7. The sensor data reconstruction method based on a multi-source temporal generative adversarial network according to claim 6, characterized in that: The calculation expression of the loss function of the generator is as follows: Where: Loss G is the loss function of the generator, W(p,q) is the Wasserstein distance between p and q, and p and q are the probability distributions of the real data and the reconstructed data respectively, λ i represents the weight coefficient of the i-th sensor, X i and X′ i are the real data and the reconstructed data of the i-th sensor respectively, MSE(X i ,X′ i ) is the mean square error between X i and X′ i , and n is the number of sensors.
8. The sensor data reconstruction method based on a multi-source temporal generative adversarial network according to claim 7, characterized in that: The weight coefficient λ i has the following calculation expression: Where: abs() represents taking the absolute value of each element in the vector, and mean() represents calculating the average value of all elements in the vector.
9. The method for reconstructing sensor data based on a multi-source temporal generative adversarial network according to claim 6, wherein: The calculation expression of the loss function of the discriminator is as follows: Loss D = W(p, q) + L MSE (X, X′) Among them: Loss D is the loss function of the discriminator, W(p, q) is the Wasserstein distance between p and q, and p and q are the probability distributions of the real data and the reconstructed data respectively, and L MSE (X, X′) is the mean square error between X and X′, and X and X′ are the real data and the reconstructed data respectively.
10. The sensor data reconstruction method based on a multi-source temporal generative adversarial network according to claim 1, characterized in that: The multi-source temporal generative adversarial network learns the data relationships between sensors using sensor data during normal working periods. Due to the masking operation, the generator cannot know the masked sensor data and can only reconstruct based on the existing sensor data that is working properly, generating reconstructed data. The discriminator obtains the real data of the assumed partially damaged sensors and the reconstructed data generated by the generator, and judges the authenticity of these two types of data. The generator continuously generates reconstructed data to try to deceive the discriminator into recognizing it as real, while the discriminator tries to correctly distinguish the authenticity of the reconstructed data. Through this adversarial process, the generator and the discriminator continuously improve their respective performances. Eventually, the generator can generate reconstructed data that is realistic enough to be almost identical to the real data.