Axle coupling vibration data generation method based on image feature coding

By converting vibration signals into image features and using the generative model to generate new vibration signals, the data scarcity problem is solved, the accuracy and generalization ability of axle coupled vibration signal analysis is improved, and it is applied to structural health monitoring and bridge abnormality detection.

CN120336790AActive Publication Date: 2025-07-18CHANGAN UNIV

Patent Information

Application Number
CN202510817272.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

In the prior art, the data of axle coupled vibration signal is scarce, resulting in low accuracy and poor generalization capabilities of data drive models, which affects the accuracy and reliability of structural health monitoring and axle coupling dynamic analysis.

Method used

The vibration signal is converted into image feature representations, and a new vibration signal is generated using a convolutional Transformer encoder and diffusion probability model to improve signal quality through end-to-end training optimization model.

Benefits of technology

A large amount of vibration data with real statistical characteristics has been generated, which improves the accuracy and generalization ability of axle coupled vibration signal analysis, and supports structural health monitoring and bridge abnormality detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336790A_ABST
    Figure CN120336790A_ABST
Patent Text Reader

Abstract

The invention provides an axle coupling vibration data generation method based on image feature coding, and belongs to the technical field of image feature coding, and the method comprises the steps: building a vibration data training set through employing an analog signal and an actual engineering signal, and dividing the vibration data training set into training subsets under different working conditions; then, the original vibration signal is preprocessed, and a time-frequency image is obtained through short-time Fourier transform; thirdly, extracting image space features by using a convolution Transform encoder, and generating image hidden space feature vectors; then, enhanced sampling is carried out on the feature vector through a diffusion probability model, and a new feature vector is generated; then, a convolution Transform decoder is used for carrying out reconstruction to generate a simulated vibration signal; and the quality of the generated signal is further improved through an end-to-end joint training optimization model. According to the method, a large amount of vibration data with real statistical characteristics can be generated, the precision and generalization ability of axle coupling vibration signal analysis are effectively improved, and the method is expected to be applied to the fields of structural health monitoring, bridge anomaly detection and intelligent defect diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image feature encoding, and particularly to a method for generating vehicle-bridge coupling vibration data based on image feature encoding. Background Art

[0002] The dynamic analysis of vehicle-bridge coupling is crucial for deeply understanding the dynamic response of bridge structures under actual traffic loads. It can not only optimize bridge design and enhance structural safety, but also improve the driving comfort of autonomous vehicles, providing important technical support for future traffic infrastructure construction. At the same time, by collecting the vibration signals generated during the interaction between vehicles and bridges, the real-time dynamic response characteristics of bridge structures can be effectively reflected. These signals can reveal potential problems such as structural damage, stiffness degradation, or abnormal vibrations, providing accurate basis for bridge safety assessment, early damage warning, and maintenance decision-making. In addition, the monitoring method combining vehicle-bridge coupling vibration signals also has advantages such as non-contact, high efficiency, and real-time, which can effectively improve the accuracy and reliability of monitoring. In actual engineering monitoring, the effective acquisition of vehicle-bridge coupling vibration signals often faces the problem of data scarcity. This scarcity is mainly manifested in the difficulty of obtaining a large amount of high-quality and reliable on-site measured vibration data, especially the data under bridge damage or special working conditions is even rarer. In addition, affected by factors such as cost, sensor installation conditions, environmental interference, and traffic control restrictions, it is difficult to collect real, continuous, and effective vibration signals. This data scarcity restricts the accuracy and robustness of data-driven models, further affecting the popularization and application of structural health monitoring methods based on vibration signals, as well as the accuracy and reliability of vehicle-bridge coupling dynamic analysis. Therefore, studying how to expand the scarce vibration signal dataset through data augmentation methods to improve the accuracy and generalization performance of signal processing and analysis has important engineering significance and practical value. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for generating vehicle-bridge coupling vibration data based on image feature encoding, so as to solve the technical problem of data scarcity existing in vehicle-bridge coupling vibration signals in existing actual engineering.

[0004] Firstly, convert the existing vibration signals into image feature representations, and then use the deep generation process of encoding-diffusion-decoding to effectively augment the data, thereby generating a large number of new vibration signals with real statistical characteristics.

[0005] A vibration data training set is constructed by using analog signals and actual engineering signals, and is divided into training subsets under different working conditions. Then, the original vibration signal is preprocessed, and a time-frequency image is obtained through short-time Fourier transform. Next, a convolutional Transformer encoder is used to extract the spatial features of the image, generating an image latent space feature vector. Subsequently, the feature vector is enhanced and sampled through a diffusion probability model to generate a new feature vector. Then, a convolutional Transformer decoder is used to reconstruct and generate a realistic vibration signal. The model is optimized through end-to-end joint training to further improve the quality of the generated signal. This method can generate a large amount of vibration data with real statistical characteristics, effectively improving the accuracy and generalization ability of the analysis of vehicle-bridge coupling vibration signals, and is expected to be applied in the fields of structural health monitoring, bridge anomaly detection, and intelligent defect diagnosis.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for generating vehicle-bridge coupling vibration data based on image feature encoding, the method comprising the following steps: S1. Signal preprocessing: Filtering, denoising, removing DC bias, normalizing, and segmenting the original vibration signal to construct a vehicle-bridge coupling vibration signal training set; S2. Convolutional Transformer encoding: The vibration signal is transformed into time-frequency domain image data of a unified size through short-time Fourier transform (STFT). First, spatial features are extracted through multiple convolutional modules, and each convolutional module includes a convolutional layer, a batch normalization layer, and a ReLU activation function. Then, the output of the convolutional module is input into the Transformer Encoder module, and feature extraction is performed through a multi-head self-attention mechanism to obtain an image latent space feature vector H ; S3. Diffusion probability model enhancement: The image latent space feature vector obtained in step S2 H is input into a diffusion probability model (DDPM) for enhanced sampling to obtain an enhanced image latent space feature vector H e ; S4. Convolutional Transformer decoding: The enhanced image latent space feature vector He obtained in step S3 is input into the Transformer Decoder module, and the features are accurately restored through a self-attention mechanism and a cross-attention mechanism. Then, the spatial features are reconstructed through multiple transposed convolutional modules, and each transposed convolutional module includes a transposed convolutional layer, a batch normalization layer, and a ReLU activation function. Finally, it is reconstructed into vibration signal data through upsampling, a Sigmoid activation function, and a fully connected layer (fc) Y r ; S5. Model Joint Training Optimization: Perform end-to-end joint training optimization on the convolutional Transformer encoding-diffusion enhancement-convolutional Transformer decoding model until the model converges, and output the vibration signal generation model; S6. Determine the required number of vibration signals and output the corresponding number of vehicle-bridge coupling vibration signals.

[0007] Further, in step S1, the signal preprocessing steps include filtering and denoising, removing DC and normalizing, and signal segmentation.

[0008] Further, in step S1, the training set of vehicle-bridge coupling vibration signals is divided into two parts. The first part is the vehicle-bridge coupling vibration data collected under the condition of a flat road surface, which is used as normal data, and the second part is the vehicle-bridge coupling vibration data collected under the condition of road surface wear, which is used as abnormal data. During the training process, the corresponding image feature data are generated respectively based on the normal data and the abnormal data, and the corresponding vibration signals are further generated.

[0009] Further, the Transformer Encoder module in step S2 includes multiple Transformer encoding layers, and each layer includes a multi-head self-attention mechanism (MHA), a layer normalization layer, and a feed-forward neural network (FFN). The specific architecture of each Transformer encoding layer is: input -> multi-head self-attention mechanism -> layer normalization layer -> feed-forward neural network -> layer normalization layer.

[0010] Further, the multi-head cross-attention mechanism includes the following steps: SS1. Use the features processed by the multi-head self-attention mechanism and normalized by the layer normalization layer as the query (Q); SS2. Use the encoder output features as the key (K) and value (V); SS3. Calculate the attention weights between the query (Q) and the key (K) in each subspace, and perform weighted summation on the value (V) based on the attention weights; SS4. Concatenate the weighted summation results in each subspace and obtain the output features of the multi-head cross-attention mechanism through linear mapping.

[0011] Further, the feed-forward neural network includes, connected in sequence: a first linear layer for expanding the feature dimension; a ReLU non-linear activation function layer for performing non-linear transformation on the expanded features; a second linear layer for restoring the non-linearly transformed features to the initial feature dimension.

[0012] Further, the time-frequency domain image data in step S2 is uniformly obtained through the Short-Time Fourier Transform (STFT), with the window length, overlap rate, and number of FFT points fixed, so that all image sizes are unified.

[0013] Further, the enhanced sampling is divided into a forward diffusion process and a reverse denoising process; the Transformer module based on U-Net contains a downsampling path and an upsampling path with Transformer attention mechanisms, and the two are connected through Transformer feature splicing to improve the feature extraction efficiency.

[0014] Further, the Transformer Decoder module in step S4 includes multiple Transformer decoding layers, and each decoding layer consists of a multi-head self-attention mechanism, a multi-head cross-attention mechanism, a layer normalization layer, and a feed-forward neural network; the specific architecture of each Transformer decoding layer is: input -> multi-head self-attention mechanism -> layer normalization layer -> multi-head cross-attention mechanism -> layer normalization layer -> feed-forward neural network -> layer normalization layer; the decoded features are further reconstructed into spatial features through multiple transposed convolution modules, and each transposed convolution module includes a transposed convolution layer, a batch normalization layer, and a ReLU activation function, and finally the vibration signal is reconstructed through an upsampling layer, a Sigmoid activation function, and a fully connected layer (fc).

[0015] Further, a Generative Adversarial Network (GAN) discriminator is introduced for auxiliary optimization in the end-to-end joint training stage; the total loss function consists of three parts: time-domain loss, frequency-domain loss, and adversarial loss.

[0016] Further, the loss function of the model includes three parts: time-domain, frequency-domain, and adversarial losses, and its combined total loss function is shown in formula (1): (1); Where, L is the total loss, is the time-domain loss of the vibration signal, is the frequency-domain loss after the vibration signal is converted into a spectrogram, is the adversarial loss. , , are the coefficients corresponding to the time-domain loss, frequency-domain loss, and adversarial loss respectively; The specific definition of the time-domain loss is shown in formula (2): (2); Where, N represents the total number of samples; represents the th true vibration signal; The vibration signal generated on behalf of the th real vibration signal; The frequency-domain loss is specifically defined as shown in formula (3): (3); where represents the transformation function that converts the time-domain signal to the frequency domain; : represents the square of the Euclidean distance; The adversarial loss is specifically defined as shown in formula (4): (4); where, is the discriminator, which is used to determine whether the input signal is a real signal or a generated signal; represents taking the average of all samples; represents the real signal Y t After passing through the discriminator, it is hoped that the probability output by the discriminator is as close to 1 as possible; represents the generated signal Y r After passing through the discriminator, it is hoped that the probability output by the discriminator is as close to 0 as possible.

[0017] Furthermore, after the vibration signal is generated, the signal-to-noise ratio can be used to preliminarily judge the quality of the generated signal. The higher the signal-to-noise ratio, the higher the quality of the generated signal.

[0018] Furthermore, the normal and abnormal vibration data generated by this method can be used as a dataset for training a deep learning model, for tasks such as denoising processing and anomaly detection of vibration signals.

[0019] Due to the adoption of the above technical solutions, the present invention has the following beneficial effects: The present invention can effectively solve the problems of low accuracy and poor generalization ability of data-driven models caused by insufficient data, and provide a reliable data basis for the promotion and application of vehicle-bridge coupling dynamics analysis and structural health monitoring methods based on vibration signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is the overall flowchart of the data generation method of the present invention; Figure 2 is the schematic diagram of numerical simulation carried out when making the training set of the present invention; Figure 3 is the time-domain schematic diagram of the vehicle-bridge coupling vibration signal of the present invention; Figure 4 is the frequency-domain schematic diagram of the vehicle-bridge coupling vibration signal after conversion of the present invention; Figure 5It is the network structure diagram of the Transformer encoder of the present invention; Figure 6 It is the structure diagram of the DDPM network of the present invention; Figure 7 It is the network structure diagram of the Transformer decoder of the present invention; Figure 8 It is the frequency-domain schematic diagram of the vehicle-bridge coupling vibration signal generated by the present invention. Detailed implementation manners

[0021] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following preferred embodiments are cited with reference to the accompanying drawings for further detailed description of the present invention. However, it should be noted that many details listed in the specification are only for enabling the reader to have a thorough understanding of one or more aspects of the present invention, and these aspects of the present invention can be implemented even without these specific details.

[0022] As Figure 1-7 shown, the present invention provides a technical solution: a method for generating vehicle-bridge coupling vibration data based on image features, which specifically executes the following S1 to S6 to obtain the generated vibration signal: S1. Signal preprocessing: Filter and denoise, remove direct current and normalize the original vibration signal and segment it to construct a vehicle-bridge coupling vibration signal training set; S2. Convolutional Transformer encoding: Obtain time-frequency domain image data of a unified size by performing short-time Fourier transform (STFT) on the vibration signal. First, extract spatial features through multiple convolutional modules, and each convolutional module includes a convolutional layer, a batch normalization layer, and a ReLU activation function; then input the output of the convolutional module into the Transformer Encoder module to extract features through the multi-head self-attention mechanism to obtain an image latent space feature vector H ; S3. Diffusion probability model enhancement: Input the image latent space feature vector H obtained in step S2 into a diffusion probability model (DDPM) for enhanced sampling to obtain an enhanced image latent space feature vector H e ; S4. Convolutional Transformer decoding: Input the enhanced image latent space feature vector He obtained in step S3 into the Transformer Decoder module, accurately restore features through the self-attention mechanism and the cross-attention mechanism, and then reconstruct spatial features through multiple transposed convolutional modules. Each transposed convolutional module includes a transposed convolutional layer, a batch normalization layer, and a ReLU activation function. Finally, it is reconstructed into vibration signal data through upsampling, a Sigmoid activation function, and a fully connected layer (fc) Yr ; S5. Model Joint Training Optimization: End-to-end joint training optimization is performed on the convolutional Transformer encoding-diffusion enhancement-convolutional Transformer decoding model until the model converges, and a vibration signal generation model is output; S6. Determine the required number of vibration signals and output the corresponding number of vehicle-bridge coupling vibration signals.

[0023] The overall framework of the proposed method is as Figure 1 shown. Specifically, by performing short-time Fourier transform on the original vehicle-bridge coupling vibration signal, a time-frequency domain image is obtained, and spatial feature vectors are sequentially extracted by a convolutional neural network and a Transformer encoder H ; Then, the diffusion transformation of the feature vectors is realized by using the diffusion probability model (DDPM) to obtain new spatial feature vectors H e , and then the vibration signal is reconstructed by joint decoding of the Transformer encoder and multiple convolutional modules, and joint training optimization is carried out to finally generate vibration signals under different working conditions that meet the target standards.

[0024] To demonstrate the specific implementation cases of the present invention, numerical simulations as Figure 2 shown are carried out, and a training set is constructed by combining the signals obtained in actual engineering. The vibration signals are as Figure 3 shown, and the time-frequency domain image obtained by performing short-time Fourier transform is as Figure 4 shown. Next, the DDPM branch network structure, convolutional-Transformer encoder, and convolutional-Transformer decoder network in S1 to S4 will be introduced in detail.

[0025] Convolutional-Transformer Encoder: As Figure 3As shown, the input vibration signal first obtains time-frequency domain image data of a unified size through short-time Fourier transform (STFT); then the time-frequency domain image data is input into a convolutional feature extraction module, which consists of three sequentially connected convolutional modules, and each convolutional module is composed of a convolutional layer (Conv2d), a batch normalization layer (BatchNorm2d), and a ReLU activation function; next, the output image features of the third convolutional module are flattened into a sequence form and then input into a Transformer encoding module, which contains multiple layers of Transformer Encoder layers, and each Transformer Encoder layer consists of a multi-head self-attention mechanism and a feed-forward neural network. Among them, the multi-head self-attention mechanism is responsible for extracting the global correlation information between different positions in the sequence, and the feed-forward neural network further extracts deep features; finally, the sequence features output by the Transformer encoding module are reconstructed into a latent space feature vector corresponding to the input image size H , and its size is 128×128×128. The detailed parameters of the model are shown in Table 1 below

[0026] Table 1 shows the network parameters of the convolutional-Transformer model

[0027] DDPM network structure: The internal architecture of DDPM is as Figure 4 shown, and its main body adopts a typical U-Net framework. The convolutional layer of the network consists of two convolutional modules. The downsampling layer and the upsampling layer are both composed of convolutional and sampling modules. And after the downsampling or upsampling operation, an attention layer is introduced in the network structure to enhance the ability to capture spatial features. After each level of feature map in the network undergoes downsampling, it is connected to the corresponding upsampling path through a skip connection. And before the connection, the feature map undergoes a 1×1 convolutional operation to adjust the number of channels and feature dimensions. During the upsampling process, the network concatenates the spatial features generated by each level of downsampling with the current features to form a mixed feature map to support the feature reconstruction process

[0028] Convolutional-Transformer decoder: As Figure 1 shown, the enhanced feature vector H eFirst, it passes through the Transformer Decoder module to capture the long-range dependencies in the sequence data. Subsequently, it passes through three sequentially connected Transposed Convolution modules to gradually increase the size of the feature map and restore the spatial dimensions using the upsampling operation. After that, the feature map undergoes a non-linear mapping through the Sigmoid activation function and then passes through the fully connected (fc) layer to achieve feature compression and dimension matching, finally outputting the reconstructed vibration signal. Y r . The detailed parameters of the model are shown in Tables 2, 3, and 4 below.

[0029] Table 2 shows the parameters of the Transformer Decoder module.

[0030] Table 3 shows the parameters of the Transposed Convolution module.

[0031] Table 4 shows the parameters of the upsampling and fully connected (fc) layers.

[0032] In the method proposed in the present invention, the network model is further optimized through a mixed loss function. The loss function of the model includes three parts: time domain, frequency domain, and adversarial loss. The combined total loss function is shown in Equation (1): (1); Where L is the total loss, is the time domain loss of the vibration signal, is the frequency domain loss after the vibration signal is converted into a spectrogram, is the adversarial loss. , , are the coefficients corresponding to the time domain loss, frequency domain loss, and adversarial loss respectively; The specific definition of the time domain loss is shown in Equation (2): (2); Where N represents the total number of samples; represents the th real vibration signal; represents the th vibration signal generated from the real vibration signal; The specific definition of the frequency domain loss is shown in Equation (3): (3); Where A transformation function for converting a time-domain signal to a frequency domain; : Represents the square of the Euclidean distance; The adversarial loss is specifically defined as shown in formula (4): (4); Wherein, is a discriminator for determining whether the input signal is a real signal or a generated signal; Represents taking the average over all samples; Represents the real signal Y t After passing through the discriminator, it is desired that the probability output by the discriminator is as close to 1 as possible; Represents the generated signal Y r After passing through the discriminator, it is desired that the probability output by the discriminator is as close to 0 as possible.

[0033] In this embodiment, the model reaches a fully convergent state after 1000 training iterations. Based on the data generation method proposed by the present invention, a target signal is generated, and its frequency domain characteristics are as Figure 8 shown.

[0034] Matters not covered by the present invention are well-known techniques.

[0035] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for generating axle-bridge coupling vibration data based on image feature coding, characterized in that: The method includes the following steps: S1. Signal preprocessing: Filter and denoise the original vibration signal, remove the direct current and normalize it, and segment it to construct a training set of vehicle-bridge coupled vibration signals; S2. Convolutional Transformer Encoding: The vibration signal is subjected to Fourier transform to obtain time-frequency domain image data of a unified size. First, spatial features are extracted through several convolutional modules, each of which includes a convolutional layer, a batch normalization layer, and a ReLU activation function. Then, the output of the convolutional module is input into the Transformer Encoder module, and feature extraction is performed through the multi-head self-attention mechanism to obtain the image latent space feature vector H ; S3. Enhancement of the diffusion probability model: The image latent space feature vector obtained in step S2 H is input into the diffusion probability model for enhanced sampling to obtain an enhanced image latent space feature vector H e ; S4. Convolutional Transformer Decoding: Input the enhanced image latent space feature vector He obtained in step S3 into the Transformer Decoder module. Accurately restore the features through the self-attention mechanism and the cross-attention mechanism, and then reconstruct the spatial features through several transposed convolution modules. Each transposed convolution module contains a transposed convolution layer, a batch normalization layer, and a ReLU activation function. Finally, it is reconstructed into vibration signal data through upsampling, the Sigmoid activation function, and a fully connected layer. Y r ; S5. Model joint training and optimization: Perform end-to-end joint training and optimization on the convolutional Transformer encoder-diffusion enhancement-convolutional Transformer decoder model until the model converges, and output a vibration signal generation model; S6. Determine the number of required vibration signals and output the corresponding number of vehicle-bridge coupled vibration signals.

2. A method for generating axle coupling vibration data based on image feature coding according to claim 1, characterized in that: In step S1, the training set of vehicle-bridge coupled vibration signals is divided into two parts. The first part is the vehicle-bridge coupled vibration data collected under the condition of a flat road surface, which is used as normal data, and the second part is the vehicle-bridge coupled vibration data collected under the condition of road surface wear, which is used as abnormal data; During the training process, image feature data corresponding to the normal data and the abnormal data are generated respectively, and the corresponding vibration signals are further generated.

3. A method for generating axle-bridge coupling vibration data based on image feature coding according to claim 1, characterized in that: In step 2, the Transformer Encoder module includes several Transformer encoding layers, each layer including a multi-head self-attention mechanism, a layer normalization layer, and a feed-forward neural network. The specific architecture of each Transformer encoding layer is: input -> multi-head self-attention mechanism -> layer normalization layer -> feed-forward neural network -> layer normalization layer.

4. A method for generating axle coupling vibration data based on image feature coding according to claim 3, wherein: The multi-head cross-attention mechanism includes the following steps: SS1. Use the features processed by the multi-head self-attention mechanism and normalized by the layer normalization layer as queries; SS2. Use the encoder output features as keys and values; SS3. Calculate the attention weights between the queries and the keys in each subspace, and perform weighted summation on the values based on the attention weights; SS4. Concatenate the weighted summation results in each subspace and obtain the output features of the multi-head cross-attention mechanism through a linear mapping.

5. A method for generating axle coupling vibration data based on image feature encoding according to claim 3, characterized in that: The feed-forward neural network includes a first linear layer, a ReLU non-linear activation function layer, and a second linear layer connected in sequence. The first linear layer is used to expand the feature dimension; the ReLU non-linear activation function layer is used to perform non-linear transformation on the expanded features, and the second linear layer is used to restore the non-linearly transformed features to the initial feature dimension.

6. A method for generating axle coupling vibration data based on image feature encoding according to claim 1, characterized in that: In step S3, the enhanced sampling is divided into a forward diffusion process and a reverse denoising process. The Transformer module based on U-Net contains a downsampling path and an upsampling path with Transformer attention mechanisms, and the two are connected through Transformer feature splicing to improve the feature extraction efficiency.

7. A method for generating axle coupling vibration data based on image feature encoding according to claim 1, characterized in that: In step S3, the loss function of the diffusion probability model includes three parts: time domain, frequency domain, and adversarial loss. The combined total loss function is shown in formula (1): (1); Among them, L is the total loss, is the time-domain loss of the vibration signal, is the frequency-domain loss after the vibration signal is converted into a spectrogram, is the adversarial loss, 、 、 are the coefficients corresponding to the time-domain loss, frequency-domain loss, and adversarial loss, respectively; The time domain loss is specifically defined as shown in formula (2): (2); Among them, N represents the total number of samples, represents the th true vibration signal, represents the vibration signal generated by the th true vibration signal; The frequency domain loss is specifically defined as shown in formula (3): (3); Among them represents the transformation function that converts the time-domain signal to the frequency domain, : represents the square of the Euclidean distance; The adversarial loss is specifically defined as shown in formula (4): (4) Among them, is a discriminator for determining whether the input signal is a real signal or a generated signal, represents taking the average over all samples, represents the real signal Y t After passing through the discriminator, the probability output by the discriminator is close to 1, represents the generated signal Y r After passing through the discriminator, the probability output by the discriminator is close to 0.

8. A method for generating axle-coupled vibration data based on image feature coding according to claim 1, characterized in that: The Transformer Decoder module in step S4 contains a number of Transformer decoding layers. Each decoding layer consists of a multi-head self-attention mechanism, a multi-head cross-attention mechanism, a layer normalization layer, and a feed-forward neural network. The specific architecture of each Transformer decoding layer is: input -> multi-head self-attention mechanism -> layer normalization layer -> multi-head cross-attention mechanism -> layer normalization layer -> feed-forward neural network -> layer normalization layer. The decoded features are further reconstructed into spatial features through a number of transposed convolution modules. Each transposed convolution module includes a transposed convolution layer, a batch normalization layer, and a ReLU activation function. Finally, the vibration signal is reconstructed through an upsampling layer, a Sigmoid activation function, and a fully connected layer (fc).

9. A method for generating axle coupling vibration data based on image feature coding according to claim 1, characterized in that: In step S4, after the vibration signal is generated, the signal-to-noise ratio is used to preliminarily judge the quality of the generated signal. The higher the signal-to-noise ratio, the higher the quality of the generated signal.

10. A method for generating axle-coupled vibration data based on image feature coding according to claim 1, characterized in that: In step S5, a generative adversarial network discriminator is introduced for auxiliary optimization in the end-to-end joint training stage. The total loss function consists of three parts: time-domain loss, frequency-domain loss, and adversarial loss.

Citation Information

Patent Citations

  • Axle safety health assessment system and method based on axle coupling analysis

    CN105825014A

  • Subway tunnel vehicle-induced vibration data set construction method

    CN117150291A

  • Bridge indirect damage identification method based on axle coupling vibration and deep learning

    CN117435991A

  • Screening method and device of bridge monitoring data, equipment and medium

    CN118172673A

  • Coupled vibration data analysis method for axle

    CN118965222A

Cited By

  • Method for repairing sound quality, electronic equipment, storage medium and program product

    CN120510857A