A method for generating vehicle-bridge coupled vibration data based on image feature coding

By converting vibration signals into image features and generating new vibration signals using convolutional Transformer and diffusion probability model, the data scarcity problem is solved, the accuracy and generalization ability of signal analysis are improved, and applied to structural health monitoring and bridge anomaly detection.

CN120336790BActive Publication Date: 2025-08-19CHANGAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510817272.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-08-19
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

In the prior art, the data of axle coupled vibration signal is scarce, resulting in low accuracy and poor generalization capabilities of data drive models, which affects the accuracy and reliability of structural health monitoring and axle coupling dynamic analysis.

Method used

The vibration signal is converted into image feature representations, and a new vibration signal is generated using a convolutional Transformer encoder and diffusion probability model to improve signal quality through end-to-end training optimization model.

Benefits of technology

A large amount of vibration data with real statistical characteristics has been generated, which improves the accuracy and generalization ability of axle coupled vibration signal analysis, and supports structural health monitoring and bridge abnormality detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336790B_ABST
    Figure CN120336790B_ABST
Patent Text Reader

Abstract

This invention provides a method for generating vehicle-bridge coupled vibration data based on image feature coding. This method belongs to the field of image feature coding technology. A vibration data training set is constructed using simulated signals and actual engineering signals, and is divided into training subsets under different operating conditions. The original vibration signal is then preprocessed, and a time-frequency image is obtained through a short-time Fourier transform. Next, a convolutional Transformer encoder is used to extract image spatial features and generate an image latent space feature vector. Subsequently, the feature vector is enhanced and sampled using a diffusion probability model to generate a new feature vector. A convolutional Transformer decoder is then used to reconstruct and generate a realistic vibration signal. An end-to-end joint training optimization model is then used to further improve the quality of the generated signal. This method can generate a large amount of vibration data with realistic statistical characteristics, effectively improving the accuracy and generalization of vehicle-bridge coupled vibration signal analysis. It is expected to be applied to structural health monitoring, bridge anomaly detection, and intelligent defect diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image feature coding, and in particular to a method for generating vehicle-bridge coupling vibration data based on image feature coding. Background Art

[0002] Vehicle-bridge coupled dynamics analysis is crucial for a deep understanding of the dynamic response of bridge structures under actual traffic loads. It can not only optimize bridge design and improve structural safety, but also improve the driving comfort of autonomous vehicles, providing important technical support for future transportation infrastructure construction. At the same time, by collecting vibration signals generated during the interaction between vehicles and bridges, the real-time dynamic response characteristics of the bridge structure can be effectively reflected. These signals can reveal potential problems such as structural damage, stiffness degradation or abnormal vibration, and provide an accurate basis for bridge safety assessment, early damage warning and maintenance decisions. In addition, the monitoring method combined with vehicle-bridge coupled vibration signals also has the advantages of non-contact, high efficiency, and real-time, which can effectively improve the accuracy and reliability of monitoring.

[0003] In practical engineering monitoring, the effective collection of vehicle-bridge coupled vibration signals often faces the problem of data scarcity. This scarcity primarily manifests itself in the difficulty of acquiring large quantities of high-quality, reliable, field-measured vibration data, especially when the bridge is damaged or under special operating conditions. Furthermore, factors such as cost, sensor installation conditions, environmental interference, and traffic control restrictions make it difficult to collect authentic, continuous, and valid vibration signals. This data scarcity restricts the accuracy and robustness of data-driven models, further impacting the widespread application of vibration-based structural health monitoring methods and the accuracy and reliability of vehicle-bridge coupled dynamics analysis. Therefore, studying how to expand scarce vibration signal datasets through data augmentation methods to improve the accuracy and generalization performance of signal processing and analysis is of great engineering significance and practical value. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for generating vehicle-bridge coupled vibration data based on image feature coding, so as to solve the technical problem of data scarcity of vehicle-bridge coupled vibration signals in existing practical engineering projects.

[0005] First, the existing vibration signal is converted into image feature representation, and then the data is effectively augmented using the deep generation process of encoding-diffusion-decoding to generate a large number of new vibration signals with real statistical characteristics.

[0006] A vibration data training set is constructed using simulated and actual engineering signals and divided into training subsets for different operating conditions. The original vibration signal is then preprocessed and a short-time Fourier transform is used to obtain a time-frequency image. A convolutional Transformer encoder is then used to extract image spatial features and generate a latent feature vector. This feature vector is then upsampled using a diffusion probability model to generate a new feature vector. A convolutional Transformer decoder is then used to reconstruct and generate a realistic vibration signal. The model is optimized through end-to-end joint training to further improve the quality of the generated signal. This method can generate a large amount of vibration data with realistic statistical characteristics, effectively improving the accuracy and generalization of vehicle-bridge coupled vibration signal analysis. It is expected to be applied to structural health monitoring, bridge anomaly detection, and intelligent defect diagnosis.

[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] A method for generating vehicle-bridge coupled vibration data based on image feature coding, the method comprising the following steps:

[0009] S1. Signal preprocessing: Filter and de-noise the original vibration signal, remove DC current, normalize it, and segment it to construct a training set of vehicle-bridge coupled vibration signals.

[0010] S2, Convolutional Transformer Encoding: The vibration signal is transformed into time-frequency domain image data of uniform size through short-time Fourier transform (STFT). The spatial features are first extracted through multiple convolution modules. Each convolution module includes a convolution layer, a batch normalization layer, and a ReLU activation function. The output of the convolution module is then input into the Transformer Encoder module, and feature extraction is performed through a multi-head self-attention mechanism to obtain the image latent space feature vector. H ;

[0011] S3, diffusion probability model enhancement: the image latent space feature vector obtained in step S2 is H Input diffusion probability model (DDPM) for enhanced sampling to obtain the enhanced image latent space feature vector H e ;

[0012] S4, Convolutional Transformer decoding: The enhanced image latent space feature vector He obtained in step S3 is input into the Transformer Decoder module, and the features are accurately restored through the self-attention mechanism and the cross-attention mechanism. The spatial features are then reconstructed through multiple transposed convolution modules. Each transposed convolution module contains a transposed convolution layer, a batch normalization layer, and a ReLU activation function. Finally, it is reconstructed into vibration signal data through upsampling, Sigmoid activation function, and a fully connected layer (fc).Y r ;

[0013] S5. Model joint training optimization: performing end-to-end joint training optimization on the convolutional Transformer encoding-diffusion enhancement-convolutional Transformer decoding model until the model converges, and outputting a vibration signal generation model;

[0014] S6. Determine the required number of vibration signals, and output a corresponding number of axle-coupled vibration signals.

[0015] Furthermore, in step S1, the signal preprocessing step includes filtering and denoising, DC removal and normalization, and signal segmentation.

[0016] Furthermore, in step S1, the training set of the vehicle-bridge coupled vibration signal is divided into two parts, wherein the first part is the vehicle-bridge coupled vibration data collected under a flat road condition, which serves as normal data, and the second part is the vehicle-bridge coupled vibration data collected under a road wear condition, which serves as abnormal data; during the training process, corresponding image feature data are generated based on the normal data and the abnormal data, and further corresponding vibration signals are generated.

[0017] Furthermore, the Transformer Encoder module in step S2 contains multiple Transformer encoding layers, each of which includes a multi-head self-attention mechanism (MHA), a layer normalization layer, and a feedforward neural network (FFN); the specific architecture of each Transformer encoding layer is: input -> multi-head self-attention mechanism -> layer normalization layer -> feedforward neural network -> layer normalization layer.

[0018] Furthermore, the multi-head cross attention mechanism includes the following steps:

[0019] SS1: The features processed by the multi-head self-attention mechanism and normalized by the layer are used as the query (Q);

[0020] SS2, use the encoder output features as keys (K) and values (V);

[0021] SS3. Calculate the attention weight between the query (Q) and the key (K) in each subspace, and perform a weighted sum of the value (V) based on the attention weight;

[0022] SS4. Concatenate the weighted summation results in each subspace and use linear mapping to obtain the output features of the multi-head cross attention mechanism.

[0023] Furthermore, the feedforward neural network includes sequentially connected: a first linear layer for expanding the feature dimension; a ReLU nonlinear activation function layer for performing a nonlinear transformation on the expanded features; and a second linear layer for restoring the nonlinearly transformed features to the initial feature dimension.

[0024] Furthermore, the time-frequency domain image data in step S2 are uniformly obtained through short-time Fourier transform (STFT), and the window length, overlap rate and FFT points are fixed, so that all image sizes are uniform.

[0025] Furthermore, enhanced sampling is divided into a forward diffusion process and a reverse denoising process; the U-Net-based Transformer module contains a downsampling path and an upsampling path of the Transformer attention mechanism, and the two are connected through Transformer feature splicing to improve feature extraction efficiency.

[0026] Furthermore, the Transformer Decoder module in step S4 includes multiple Transformer decoding layers, each decoding layer consists of a multi-head self-attention mechanism, a multi-head cross-attention mechanism, a layer normalization layer and a feedforward neural network; the specific architecture of each Transformer decoding layer is: input->multi-head self-attention mechanism->layer normalization layer->multi-head cross-attention mechanism->layer normalization layer->feedforward neural network->layer normalization layer; the decoded features are further reconstructed into spatial features through multiple transposed convolution modules, each transposed convolution module includes a transposed convolution layer, a batch normalization layer and a ReLU activation function, and finally the vibration signal is reconstructed through an upsampling layer, a Sigmoid activation function and a fully connected layer (fc).

[0027] Furthermore, a generative adversarial network (GAN) discriminator is introduced in the end-to-end joint training stage for auxiliary optimization; the total loss function consists of three parts: time domain loss, frequency domain loss, and adversarial loss.

[0028] Furthermore, the loss function of the model includes three parts: time domain, frequency domain and adversarial loss, and the combined total loss function is shown in formula (1):

[0029] (1);

[0030] in, L is the total loss, is the time domain loss of the vibration signal, is the frequency domain loss after the vibration signal is converted into a spectrum diagram, To combat losses. 、 、 are the coefficients corresponding to time domain loss, frequency domain loss, and adversarial loss respectively;

[0031] The specific definition of time domain loss is shown in formula (2):

[0032] (2);

[0033] in, N represents the total number of samples; Representative A real vibration signal; Representative A vibration signal generated by a real vibration signal;

[0034] The specific definition of frequency domain loss is shown in formula (3):

[0035] (3);

[0036] in Represents the transformation function that converts the time domain signal into the frequency domain; : represents the square of the Euclidean distance;

[0037] The specific definition of adversarial loss is shown in formula (4):

[0038] (4);

[0039] in, It is a discriminator, which is used to determine whether the input signal is a real signal or a generated signal; Indicates taking the average of all samples; Represents the real signal Y t After passing through the discriminator, we hope that the probability of the discriminator output is as close to 1 as possible; Indicates the generation of a signal Y r After passing through the discriminator, we hope that the probability of the discriminator output is as close to 0 as possible.

[0040] Furthermore, after the vibration signal is generated, the signal-to-noise ratio can be used to preliminarily judge the quality of the generated signal. The higher the signal-to-noise ratio, the higher the quality of the generated signal.

[0041] Furthermore, the normal and abnormal vibration data generated by this method can be used as a dataset for training deep learning models to achieve tasks such as vibration signal denoising and anomaly detection.

[0042] The present invention has the following beneficial effects due to the adoption of the above technical solution:

[0043] The present invention can effectively solve the problems of low accuracy and poor generalization ability of data-driven models caused by insufficient data, and provide a reliable data basis for the promotion and application of vehicle-bridge coupling dynamics analysis and structural health monitoring methods based on vibration signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is an overall flow chart of the data generation method of the present invention;

[0045] Figure 2 It is a schematic diagram of the numerical simulation performed when the present invention prepares the training set;

[0046] Figure 3 This is a time domain schematic diagram of the vehicle-bridge coupled vibration signal of the present invention;

[0047] Figure 4 is a frequency domain schematic diagram of the vehicle-bridge coupled vibration signal after conversion according to the present invention;

[0048] Figure 5 : is a network structure diagram of the Transformer encoder of the present invention;

[0049] Figure 6 is a structural diagram of the DDPM network of the present invention;

[0050] Figure 7 : is the network structure diagram of the Transformer decoder of the present invention;

[0051] Figure 8 It is a frequency domain schematic diagram of the vehicle-bridge coupled vibration signal generated by the present invention. DETAILED DESCRIPTION

[0052] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and by way of preferred embodiments. However, it should be noted that many of the details listed in this specification are merely provided to help the reader gain a thorough understanding of one or more aspects of the present invention, and these aspects of the present invention can be practiced even without these specific details.

[0053] like Figure 1-7 As shown, the present invention provides a technical solution: a method for generating vehicle-bridge coupled vibration data based on image features, specifically performing the following steps S1 to S6 to obtain a generated vibration signal:

[0054] S1. Signal preprocessing: Filter, remove noise, remove DC current, and normalize the original vibration signal, then segment it to construct a training set of vehicle-bridge coupled vibration signals.

[0055] S2, Convolutional Transformer Encoding: The vibration signal is transformed into time-frequency domain image data of uniform size through short-time Fourier transform (STFT). The spatial features are first extracted through multiple convolution modules. Each convolution module includes a convolution layer, a batch normalization layer, and a ReLU activation function. The output of the convolution module is then input into the Transformer Encoder module, and feature extraction is performed through a multi-head self-attention mechanism to obtain the image latent space feature vector. H ;

[0056] S3, diffusion probability model enhancement: the image latent space feature vector obtained in step S2 is H Input diffusion probability model (DDPM) for enhanced sampling to obtain the enhanced image latent space feature vector H e ;

[0057] S4, Convolutional Transformer decoding: The enhanced image latent space feature vector He obtained in step S3 is input into the Transformer Decoder module, and the features are accurately restored through the self-attention mechanism and the cross-attention mechanism. The spatial features are then reconstructed through multiple transposed convolution modules. Each transposed convolution module contains a transposed convolution layer, a batch normalization layer, and a ReLU activation function. Finally, it is reconstructed into vibration signal data through upsampling, Sigmoid activation function, and a fully connected layer (fc). Y r ;

[0058] S5. Model joint training optimization: performing end-to-end joint training optimization on the convolutional Transformer encoding-diffusion enhancement-convolutional Transformer decoding model until the model converges, and outputting a vibration signal generation model;

[0059] S6. Determine the required number of vibration signals, and output a corresponding number of axle-coupled vibration signals.

[0060] The overall framework of the proposed method is as follows Figure 1 Specifically, the time-frequency domain image is obtained by performing short-time Fourier transform on the original vehicle-bridge coupled vibration signal, and the spatial feature vectors are extracted in sequence through the convolutional neural network and the Transformer encoder. H ; Then use the diffusion probability model (DDPM) to realize the diffusion transformation of the feature vector and obtain the new spatial feature vector H e , then the vibration signal is reconstructed by jointly decoding with the Transformer encoder and multiple convolution modules, and joint training and optimization are performed to finally generate vibration signals of different working conditions that meet the target standards.

[0061] In order to demonstrate the specific implementation of the present invention, Figure 2 The numerical simulation shown in the figure is combined with the signals obtained in the actual project to construct a training set. The vibration signal is as follows Figure 3 As shown, the short-time Fourier transform is performed to obtain the time-frequency domain image as shown in Figure 4 As shown in the figure, the DDPM branch network structure, convolution-Transformer encoder and convolution-Transformer decoder networks in S1 to S4 will be introduced in detail.

[0062] Convolutional-Transformer Encoder:

[0063] like Figure 3 As shown in the figure, the input vibration signal is first obtained by short-time Fourier transform (STFT) to obtain time-frequency domain image data of uniform size; then the time-frequency domain image data is input into the convolutional feature extraction module, which contains three convolutional modules connected in sequence, each of which is composed of a convolutional layer (Conv2d), a batch normalization layer (BatchNorm2d) and a ReLU activation function; then, the output image features of the third convolutional module are flattened into a sequence form and input into the Transformer encoding module, which contains multiple layers of Transformer Encoder layers, each of which is composed of a multi-head self-attention mechanism and a feedforward neural network, wherein the multi-head self-attention mechanism is responsible for extracting the global correlation information between different positions in the sequence, and the feedforward neural network further extracts deep-level features; finally, the sequence features output by the Transformer encoding module are reconstructed into a latent space feature vector corresponding to the input image size H , its size is 128×128×128, and the detailed parameters of the model are shown in Table 1 below.

[0064] Table 1 shows the network parameters of the convolution-Transformer model

[0065]

[0066] DDPM network structure:

[0067] The internal architecture of DDPM is as follows Figure 4As shown in the figure, its main structure adopts the typical U-Net framework. The network's convolutional layer consists of two layers of convolutional modules. The downsampling layer and upsampling layer are both composed of convolution and upsampling modules. After the downsampling or upsampling operation, an attention layer is introduced into the network structure to enhance the ability to capture spatial features. After downsampling, the feature maps at each level of the network are connected to the corresponding upsampling path through skip connections. Before the connection, the feature maps undergo a 1×1 convolution operation to adjust the number of channels and feature dimensions. During the upsampling process, the network concatenates the spatial features generated by each downsampling level with the current features to form a mixed feature map to support feature reconstruction.

[0068] Convolutional-Transformer Decoder:

[0069] like Figure 1 As shown, the enhanced feature vector H e The first step is to pass through the Transformer Decoder module to capture long-range dependencies in the sequence data. It then passes through three sequentially connected transposed convolution modules to gradually increase the size of the feature map, and uses upsampling to restore the spatial size. The feature map then undergoes nonlinear mapping using a Sigmoid activation function, and then passes through a fully connected (FC) layer to achieve feature compression and dimensionality matching, ultimately outputting the reconstructed vibration signal. Y r The detailed parameters of the model are shown in Table 2, Table 3, and Table 4.

[0070] Table 2 shows the parameters of the Transformer Decoder module.

[0071]

[0072] Table 3 shows the parameters of the transposed convolution module

[0073]

[0074] Table 4 Upsampling and fully connected (fc) layer parameters

[0075]

[0076] In the method proposed in the present invention, the network model is further optimized by a hybrid loss function. The loss function of the model includes three parts: time domain, frequency domain and adversarial loss. The combined total loss function is shown in formula (1):

[0077] (1);

[0078] in, L is the total loss, is the time domain loss of the vibration signal, is the frequency domain loss after the vibration signal is converted into a spectrum diagram, To combat losses. 、 、 are the coefficients corresponding to time domain loss, frequency domain loss, and adversarial loss respectively;

[0079] The specific definition of time domain loss is shown in formula (2):

[0080] (2);

[0081] in, N represents the total number of samples; Representative A real vibration signal; Representative A vibration signal generated by a real vibration signal;

[0082] The specific definition of frequency domain loss is shown in formula (3):

[0083] (3);

[0084] in Represents the transformation function that converts the time domain signal into the frequency domain; : represents the square of the Euclidean distance;

[0085] The specific definition of adversarial loss is shown in formula (4):

[0086] (4);

[0087] in, It is a discriminator, which is used to determine whether the input signal is a real signal or a generated signal; Indicates taking the average of all samples; Represents the real signal Y t After passing through the discriminator, we hope that the probability of the discriminator output is as close to 1 as possible; Indicates the generation of a signal Y rAfter passing through the discriminator, we hope that the probability of the discriminator output is as close to 0 as possible.

[0088] In this embodiment, the model reaches a fully converged state after 1000 training iterations. Based on the data generation method proposed in the present invention, a target signal is generated, and its frequency domain characteristics are as follows: Figure 8 shown.

[0089] Matters not covered by the present invention are known technologies.

[0090] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for generating vehicle-bridge coupled vibration data based on image feature coding, characterized by: The method comprises the following steps: S1. Signal preprocessing: Filter, remove noise, remove DC current, and normalize the original vibration signal, then segment it to construct a training set of vehicle-bridge coupled vibration signals. S2, Convolutional Transformer Encoding: The vibration signal is Fourier transformed to obtain time-frequency domain image data of uniform size. First, the spatial features are extracted through several convolution modules. Each convolution module includes a convolution layer, a batch normalization layer and a ReLU activation function. The output of the convolution module is then input into the Transformer Encoder module, and feature extraction is performed through the multi-head self-attention mechanism to obtain the image latent space feature vector H ; S3, diffusion probability model enhancement: the image latent space feature vector obtained in step S2 is H Input the diffusion probability model for enhanced sampling to obtain the enhanced image latent space feature vector H e ; S4, Convolutional Transformer decoding: The enhanced image latent space feature vector obtained in step S3 is converted to H e The input is the Transformer Decoder module, which accurately recovers features through self-attention mechanism and cross-attention mechanism, and then reconstructs spatial features through several transposed convolution modules. Each transposed convolution module contains a transposed convolution layer, a batch normalization layer and a ReLU activation function. Finally, it is reconstructed into vibration signal data through upsampling, Sigmoid activation function and full connection layer. Y r ; S5. Model joint training optimization: performing end-to-end joint training optimization on the convolutional Transformer encoding-diffusion enhancement-convolutional Transformer decoding model until the model converges, and outputting a vibration signal generation model; S6. Determine the required number of vibration signals and output a corresponding number of axle-coupled vibration signals.

2. The method for generating vehicle-bridge coupled vibration data based on image feature coding according to claim 1, characterized in that: In step S1, the training set of vehicle-bridge coupled vibration signals is divided into two parts, wherein the first part is the vehicle-bridge coupled vibration data collected under smooth road conditions, which is regarded as normal data, and the second part is the vehicle-bridge coupled vibration data collected under road wear conditions, which is regarded as abnormal data; During the training process, corresponding image feature data are generated based on normal data and abnormal data, and corresponding vibration signals are further generated.

3. The method for generating vehicle-bridge coupled vibration data based on image feature coding according to claim 1, characterized in that: In step 2, the Transformer Encoder module contains several Transformer encoding layers. Each layer includes a multi-head self-attention mechanism, a layer normalization layer, and a feedforward neural network. The specific architecture of each Transformer encoding layer is: input -> multi-head self-attention mechanism -> layer normalization layer -> feedforward neural network -> layer normalization layer.

4. The method for generating vehicle-bridge coupled vibration data based on image feature coding according to claim 3, characterized in that: The multi-head cross attention mechanism includes the following steps: SS1: The features processed by the multi-head self-attention mechanism and layer-normalized are used as queries; SS2, use encoder output features as keys and values; SS3, calculate the attention weights between the query and the key in each subspace, and perform weighted summation of the values based on the attention weights; SS4. Concatenate the weighted summation results in each subspace and use linear mapping to obtain the output features of the multi-head cross attention mechanism.

5. The method for generating vehicle-bridge coupled vibration data based on image feature coding according to claim 3, characterized in that: The feedforward neural network includes a first linear layer, a ReLU nonlinear activation function layer and a second linear layer connected sequentially. The first linear layer is used to expand the feature dimension; the ReLU nonlinear activation function layer is used to perform nonlinear transformation on the expanded features; and the second linear layer is used to restore the nonlinearly transformed features to the initial feature dimension.

6. The method for generating vehicle-bridge coupled vibration data based on image feature coding according to claim 1, characterized in that: In step S3, enhanced sampling is divided into a forward diffusion process and a reverse denoising process. The U-Net-based Transformer module contains a downsampling path and an upsampling path of the Transformer attention mechanism. The two are connected through Transformer feature splicing to improve feature extraction efficiency.

7. The method for generating vehicle-bridge coupled vibration data based on image feature coding according to claim 1, characterized in that: In step S3, the loss function of the diffusion probability model includes three parts: time domain, frequency domain and adversarial loss. The combined total loss function is shown in formula (1): (1); in, L is the total loss, is the time domain loss of the vibration signal, is the frequency domain loss after the vibration signal is converted into a spectrum diagram, To combat losses, a 、 b 、 c are the coefficients corresponding to time domain loss, frequency domain loss, and adversarial loss respectively; The specific definition of time domain loss is shown in formula (2): (2); in, N represents the total number of samples, Representative A real vibration signal, Representative A vibration signal generated by a real vibration signal; The specific definition of frequency domain loss is shown in formula (3): ; in represents the transformation function that converts the time domain signal to the frequency domain, : represents the square of the Euclidean distance; The specific definition of adversarial loss is shown in formula (4): ; in, is a discriminator, which is used to determine whether the input signal is a real signal or a generated signal. It means taking the average of all samples. Represents the real signal Y t The output probability after the discriminator is Indicates the generation of a signal Y r The output probability after the discriminator.

8. The method for generating vehicle-bridge coupled vibration data based on image feature coding according to claim 1, characterized in that: The Transformer Decoder module in step S4 contains several Transformer decoding layers. Each decoding layer consists of a multi-head self-attention mechanism, a multi-head cross-attention mechanism, a layer normalization layer, and a feedforward neural network. The specific architecture of each Transformer decoding layer is: input->multi-head self-attention mechanism->layer normalization layer->multi-head cross-attention mechanism->layer normalization layer->feedforward neural network->layer normalization layer. The decoded features are further reconstructed into spatial features through several transposed convolution modules. Each transposed convolution module includes a transposed convolution layer, a batch normalization layer, and a ReLU activation function. Finally, the vibration signal is reconstructed through an upsampling layer, a Sigmoid activation function, and a fully connected layer.

9. The method for generating vehicle-bridge coupled vibration data based on image feature coding according to claim 1, characterized in that: In step S4, after the vibration signal is generated, the quality of the generated signal is preliminarily judged using the signal-to-noise ratio. The higher the signal-to-noise ratio, the higher the quality of the generated signal.

10. The method for generating vehicle-bridge coupled vibration data based on image feature coding according to claim 1, characterized in that: In step S5, a generative adversarial network discriminator is introduced in the end-to-end joint training stage for auxiliary optimization. The total loss function consists of three parts: time domain loss, frequency domain loss, and adversarial loss.

Citation Information

Patent Citations

  • Axle safety health assessment system and method based on axle coupling analysis

    CN105825014A

  • Subway tunnel vehicle-induced vibration data set construction method

    CN117150291A