Highway vehicle track monitoring method and system
By performing Hilbert transform and Gaussian kernel function enhancement processing on highway pavement vibration signals and combining them with generative adversarial network training, the real-time and noise control problems of traditional methods in vehicle monitoring in complex environments are solved, and efficient and accurate monitoring of vehicle trajectories on highways is achieved.
Patent Information
- Application Number
- CN202510696322.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-23
AI Technical Summary
Traditional highway vehicle monitoring methods based on distributed fiber optic sensing face challenges in real-time and noise control when processing complex spatiotemporal data. This is especially true when there is a lot of signal interference and noise caused by high-speed vehicles on highways, making it difficult to achieve accurate and real-time vehicle monitoring.
The Hilbert transform is used to preprocess the road vibration signal, the Gaussian kernel function is used for trajectory enhancement, and a denoising model is generated through adversarial training using a generative adversarial network to improve the robustness and stability of the model, thereby achieving real-time and efficient monitoring of vehicle trajectories in complex environments.
The robustness and stability of the denoising model have been improved, and it can efficiently and in real time extract vehicle trajectory information in complex and fast-flowing traffic scenarios on highways, reducing manual labeling costs and enhancing the diversity and versatility of the dataset.
Smart Images

Figure CN120687761A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a method and system for monitoring vehicle trajectories on a highway. Background Art
[0002] Traditional highway vehicle monitoring methods based on distributed fiber optic sensing face several key issues and challenges in practical applications, particularly in processing complex spatiotemporal data, real-time requirements, and noise control. This is because traditional processing methods primarily focus on the perception function of fiber optic sensors, namely, sensing environmental changes through Rayleigh scattered light generated in the optical fiber. While this approach has certain advantages in sensing environmental signals, it lacks sufficient real-time processing and classification analysis capabilities when dealing with large amounts of spatiotemporal data images.
[0003] Furthermore, these spatiotemporal data images are not only subject to the technical limitations of optical fiber itself, but also to interference from external environmental factors, such as uneven road surfaces, temperature fluctuations, and climate change. These factors often cause signal distortion and increase the complexity of signal processing. This is especially true in conditions with rapid dynamics and high levels of environmental interference, such as on highways. For example, with numerous vehicles traveling at high speeds, signals generated by different vehicles in open spaces interfere with each other, resulting in high noise levels in the spatiotemporal data images and unstable data quality. This affects the accuracy of the information and the effectiveness of the system's monitoring, making it difficult for traditional sensing methods to accurately and real-timely monitor vehicles.
[0004] Furthermore, while deep learning-based technologies have made significant progress in the field of target detection in recent years, they still face numerous challenges when processing complex, noisy spatiotemporal data and images. For example, deep learning models often struggle to fully exploit the effective information in the data when faced with tasks like detecting multiple small targets, such as highway vehicles. This is especially true when the data is significantly affected by noise and the external environment, significantly reducing target detection accuracy. Therefore, while traditional deep learning models can achieve good detection results under ideal conditions, in complex environments, especially when fiber optic sensor data is noisy, the robustness and stability of the models often fail to meet the requirements for real-time, efficient monitoring.
[0005] In related technologies, the paper "Research on Key Technologies for Highway Vehicle Detection and Trajectory Prediction Based on Distributed Fiber Optic Sensing, PhD thesis, Wang Maoning" uses signal denoising and trajectory enhancement to detect and extract vehicles from DAS vibration signals collected in real highway environments, enabling long-distance vehicle tracking. However, this approach directly enhances and denoises the dataset before performing trajectory extraction and prediction. The dataset used only covers traffic flow within a fixed time period on a single road section, lacking versatility. The paper "Research on Real-Time High-Precision Detection Methods for DAS Vibration Sources, Master's thesis, Liu Churui" proposes using wavelet packet threshold denoising and an ESBMV algorithm based on empirical mode decomposition and adaptive beamforming for DAS vibration source data denoising preprocessing. A YOLO-based real-time monitoring method for DAS vibration source signals is then used to detect multiple vibration sources in complex scenarios. However, this approach uses YOLO to extract a small amount of trajectory information and does not address a large number of overlapping, noisy trajectories. The paper "Research on Self-Supervised Learning and Prediction of Vehicle Trajectory Based on Generative Adversarial Science, Master's Thesis, Zhou Danyang" uses an improved GAN model to learn and predict vehicle behavior trajectories based on drone aerial image data, rather than extracting vehicle trajectories based on fiber optic vibration signals. Summary of the Invention
[0006] The technical problem to be solved by the present invention is how to improve the robustness and stability of the denoising model to meet the needs of real-time and efficient vehicle monitoring.
[0007] The present invention solves the above technical problems through the following technical means:
[0008] A method for monitoring vehicle trajectories on a highway is proposed, comprising the following steps:
[0009] The road vibration signal of a certain road section is preprocessed using Hilbert transform to obtain the original vehicle trajectory image with noise as the training data in the dataset;
[0010] The original vehicle trajectory image is enhanced using the Gaussian kernel function to obtain an enhanced vehicle trajectory image, and the label of the enhanced vehicle trajectory image is marked as the label data in the dataset;
[0011] Using vehicle trajectory images with different characteristics similar to the label data in the dataset as trajectory image conditions, and generating synthetic noisy images with real trajectory characteristics under the control of the trajectory image conditions;
[0012] The synthetic noisy image and trajectory image conditions are used as training data and label data in the dataset respectively to obtain the expanded dataset;
[0013] The denoising generative adversarial model is trained based on the expanded dataset, and the trained generator is used as a denoising model to denoise the road vibration signal of any road section to obtain vehicle trajectory information.
[0014] Furthermore, the method of preprocessing the road vibration signal of a certain road section by using Hilbert transform to obtain the original vehicle trajectory image with noise as the training data in the data set includes:
[0015] The road vibration signal collected from a certain road section is subjected to Fourier transformation and then processed through a low-pass filter to obtain a quasi-static signal. The road vibration signal is a time series signal containing vehicle driving signals and background noise signals at multiple locations.
[0016] The envelope of the quasi-static signal is calculated using Hilbert transform as the original vehicle trajectory image with noise, and the original vehicle trajectory image is used as the training data in the dataset. The original vehicle trajectory images of multiple position points within a time period constitute a spatiotemporal distribution waterfall diagram.
[0017] Furthermore, the use of the Gaussian kernel function to enhance the original vehicle trajectory image to obtain an enhanced vehicle trajectory image, and marking the label of the enhanced vehicle trajectory image as label data in the data set, includes:
[0018] Use the Gaussian kernel function to smooth the original vehicle trajectory image to enhance the trajectory and obtain an enhanced vehicle trajectory image;
[0019] Label the enhanced vehicle trajectory image to obtain the label data in the dataset;
[0020] Among them, the formula of Gaussian kernel function is expressed as:
[0021]
[0022] Where a is a space-time vector, which represents the space-time coordinate offset relative to the center of the Gaussian kernel; ∑ is a two-dimensional covariance matrix, which is used to control the shape and direction of the kernel function.
[0023] Furthermore, the method of using a vehicle trajectory image with different characteristics similar to the label data in the dataset as a trajectory image condition and controlling the generation of a synthetic noisy image with real trajectory characteristics under the trajectory image condition includes:
[0024] Mapping the trajectory image condition to a low-dimensional latent space through a first encoder to obtain trajectory image condition information;
[0025] Taking trajectory image condition information as input, embedding the trajectory image condition information into a semantic image controlled diffusion model to obtain prediction noise, wherein the semantic image controlled diffusion model adopts a U-Net framework and is pre-trained using the dataset;
[0026] Under the control of trajectory image condition information, reverse diffusion is performed based on the semantic image controlled diffusion model. Starting from pure Gaussian noise, prediction noise is gradually input, and a low-dimensional latent variable image is gradually constructed through variational inference.
[0027] The low-dimensional latent variable image is restored through a first decoder to obtain a synthetic noisy image, wherein the first encoder and the first decoder have symmetrical structures.
[0028] Furthermore, the processing process of the first encoder on the trajectory image condition or the initial noise image is expressed as:
[0029] E0=LeakyReLU(Conv 3×3 (x))
[0030] ResBlock(F)=ReLU(BN(Conv 3×3 (ReLU(Conv 3×3 (F)))))
[0031]
[0032] z=Conv 3×3 (E L )
[0033] Where: E0 is the initial feature extracted from the trajectory image condition or the initial noise image, ResBlock(F) represents the feature obtained by inputting the feature F into the residual block, and E l Represents the concatenation of the output features of the l-1th residual convolution block and the output features of the l-2th residual convolution block after convolution, l = 1, 2, ..., L, L is the total number of residual convolution blocks, Conv 3×3 represents a 3×3 convolution kernel, represents the Concat operation, LeakyReLU and ReLU represent activation functions, z is the low-dimensional latent variable obtained by the first encoder processing the initial noise image, Pic tr is the trajectory image condition information obtained by the first encoder processing the trajectory image condition, x is the trajectory image condition or the initial noise image, E VAE (x) = z or E VAE (x) = Pic tr The process representation for the first encoder.
[0034] Furthermore, the low-dimensional latent variable is used as input, and the predicted noise is obtained through the U-Net framework under the control of the trajectory image condition information. The formula is expressed as:
[0035]
[0036] Where: ∈ θ To predict noise, Conv 3×3 represents a 3×3 convolution kernel, represents the Concat operation, Swish is the activation function, and GP represents group normalization. represents the output features of the last residual block in the second decoder, represents the output feature of the l-1th residual block, represents the output feature of the lth residual block in the second encoder, Represents the input features of the first residual block in the second encoder, UpSample represents upsampling, DownSample represents downsampling, ResBlock ×L represents L consecutive residual blocks, Attention represents the attention mechanism, Pic tr is the trajectory image condition information, H mid It represents the feature obtained by the middle bottleneck layer using M consecutive residual blocks to operate on the output feature of the Lth residual block in the second encoder, PicEmb (Pic tr ) indicates that DownSample is used to convert the trajectory image condition information Pic tr The size of is reduced to embed into each downsampling path and upsampling path, and TimeEmb represents the use of Transformer-style sinusoidal position encoding as the embedding of time step t, where the second encoder and the second decoder are symmetrical in structure.
[0037] Furthermore, the back diffusion is performed based on the U-Net framework under the control of the trajectory image condition information, and the prediction noise is gradually input starting from pure Gaussian noise, and a low-dimensional latent variable image is gradually constructed through variational inference, including:
[0038] Under the control of trajectory image condition information, reverse diffusion is performed based on the U-Net framework, starting from pure Gaussian noise and gradually inputting prediction noise. The formula is expressed as:
[0039]
[0040] Where: α t is the noise scheduling coefficient at time step t, α t-1 is the noise scheduling coefficient at time step t-1, ∈ θ (x t ,t,c) is the prediction noise, σt is the parameter that controls randomness, x t is a pure Gaussian noise image, ∈ is a vector sampled from a standard Gaussian distribution;
[0041] Through variational inference on the pure Gaussian noise image x t Denoising is performed step by step to obtain a low-dimensional latent variable image x0.
[0042] Furthermore, the denoising generative adversarial model is trained based on the expanded data set to obtain a trained generator as a denoising model for denoising the road vibration signal of any road section to obtain vehicle trajectory information, including:
[0043] The denoising generative adversarial model is trained based on the expanded dataset. The denoising generative adversarial model includes a generator and a discriminator. The generator is used to denoise the training data in the expanded dataset to obtain a denoised trajectory image. The discriminator is used to distinguish whether the input denoised trajectory image is real data or generated data. The training loss function used includes:
[0044]
[0045] Where: Denotes the loss function of the discriminator, G(x) = x denoise , x label is the denoised trajectory image x denoise The trajectory label of , D represents the discriminator, and BCE represents the binary cross entropy loss; represents the loss function of the generator, z denoise is the denoised trajectory image x denoise The first encoder maps the latent variables into the latent feature space, z label is the potential feature vector of the label data, MAE is the mean absolute error loss, and MSE is the mean square error loss.
[0046] Furthermore, the process of obtaining vehicle trajectory information by denoising the road vibration signal of any road section by the denoising model is expressed as:
[0047] G enc1 =LeakyReLU(Dropout(Conv 4×4 (x ′ )))
[0048] G enc2 =LeakyReLU(Dropout(BN(Conv 4×4 (G enc1 ))))
[0049] G enc3=LeakyReLU(Dropout(BN(Conv 4×4 (G enc2 ))))
[0050] G mid =Conv 3×3 (G enc3 )
[0051] G dec1 =LeakyReLU(Dropout(BN(ConvTrans 4×4 (G mid ))))
[0052]
[0053] G out =Tanh(Dropout(Conv 3×3 (G dec3 )))
[0054] Where x ′ is the noisy vehicle trajectory image obtained by preprocessing the road vibration signal of any road section, Dropout is the loss layer, Conv 4×4 is a 4×4 convolution kernel, LeakyReLU is the activation function, BN is the normalization layer, Conv 3×3 Represents a 3×3 convolution kernel, ⊕ represents the Concat operation, ConvTrans 4×4 is the transposed convolution, Tanh is the activation function, G enc1 , G enc2 , G enc3 Represents the output features of the three-layer structure of the third encoder in the generator, G mid is the output feature of the middle bottleneck layer in the generator, G dec1 , G dec2 , G dec3 Represents the output features of the three-layer structure of the third decoder in the generator, G out To obtain the denoised vehicle trajectory information, the third encoder and the third decoder have symmetrical structures.
[0055] In addition, the present invention also proposes a highway vehicle trajectory monitoring system, the system comprising:
[0056] The preprocessing module is used to preprocess the road vibration signal of a certain road section using Hilbert transform to obtain the original vehicle trajectory image with noise as the training data in the dataset;
[0057] An enhancement processing module is used to enhance the original vehicle trajectory image using a Gaussian kernel function to obtain an enhanced vehicle trajectory image, and mark the label of the enhanced vehicle trajectory image as label data in the data set;
[0058] An image synthesis module is configured to use a vehicle trajectory image with different characteristics similar to the label data in the dataset as a trajectory image condition, and generate a synthetic noisy image with real trajectory characteristics under the control of the trajectory image condition;
[0059] The dataset expansion module is used to use the synthetic noisy image and trajectory image conditions as training data and label data in the dataset respectively to obtain the expanded dataset;
[0060] The adversarial training module is used to train the denoising generative adversarial model based on the expanded dataset, and the trained generator is used as a denoising model to denoise the road vibration signal of any road section to obtain vehicle trajectory information.
[0061] The advantages of the present invention are:
[0062] (1) The present invention first uses Hilbert transform to pre-process the road vibration signal of a certain road section to obtain a noisy original vehicle trajectory map, which can highlight the vehicle trajectory signal in a complex noise environment, so as to facilitate the subsequent label annotation and denoising model to perform denoising processing. Otherwise, the obtained vehicle trajectory image data will be unable to be labeled or processed due to a large amount of noise; then the Gaussian kernel function is used to perform trajectory enhancement processing on the original vehicle trajectory image to highlight the vehicle trajectory features partially covered by noise, which is convenient for the annotation of trajectory labels; then, based on the vehicle trajectory image similar to the label in the dataset set by the self-set one as the trajectory image condition, a high-quality synthetic noisy image that meets the vehicle trajectory image condition is generated to achieve the expansion of the dataset, which can greatly increase the diversity of the dataset, thereby improving the versatility and robustness of the subsequent denoising model and significantly reducing the manual annotation cost; finally, the denoising generative adversarial model is trained using the expanded dataset, so that the trained denoising model can extract vehicle trajectory information in the image under complex environments, especially when there is a lot of noise in the fiber optic sensor data. The model has stronger robustness and stability, so as to meet the real-time and efficient requirements of vehicle trajectory monitoring in complex and fast traffic scenes on highways.
[0063] (2) By mapping image data into latent variables in a low-dimensional latent feature space, the input spatial resolution is reduced, thereby reducing the computational complexity of model diffusion generation, giving the denoising generative adversarial model the ability to express data features and improving denoising robustness.
[0064] (3) The denoising generative adversarial model optimizes both pixel-level loss and latent feature space loss through adversarial training, achieving efficient and real-time denoising of noisy vehicle trajectory images and extracting vehicle trajectory information from the images, thus meeting the low-latency requirements of highway vehicle monitoring.
[0065] (4) The structural setting of the generator and discriminator in the denoising generative adversarial model is lightweight and easy to deploy.
[0066] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 This is a flow chart of a method for monitoring vehicle trajectories on a highway proposed in one embodiment of the present invention;
[0068] Figure 2 This is an example of a spatiotemporal waterfall diagram after preprocessing in one embodiment of the present invention;
[0069] Figure 3 This is an example of the denoising result obtained by training using the traditional U-Net structure in this invention;
[0070] Figure 4 This is an example of the denoising result obtained by training only the generator G in the LD-GAN model.
[0071] Figure 5 This is an example of the result image after denoising using the complete adversarially trained LD-GAN model in this invention;
[0072] Figure 6 This is a schematic structural diagram of a highway vehicle trajectory monitoring system proposed in one embodiment of the present invention;
[0073] Figure 7 2 is a schematic diagram of the structure of a latent diffusion model for semantic image conditional control in one embodiment of the present invention;
[0074] Figure 8 2 is a schematic diagram of the structure of a denoising generative adversarial model in one embodiment of the present invention. DETAILED DESCRIPTION
[0075] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0076] like Figure 1 As shown, an embodiment of the present invention provides a method for monitoring vehicle trajectories on a highway, the method comprising the following steps:
[0077] S10, using Hilbert transform to pre-process the road vibration signal of a certain road section to obtain the original vehicle trajectory image with noise as training data in the data set;
[0078] It should be noted that in practical applications, the vibration signals caused by vehicles traveling on a certain section of highway can be measured to obtain a time series of vibration signals at multiple locations over a period of time. The original vibration signals of the time series can be preprocessed to highlight the vehicle trajectory signals in the quasi-static signals in a complex noise environment.
[0079] S20, using a Gaussian kernel function to enhance the original vehicle trajectory image to obtain an enhanced vehicle trajectory image, and marking the label of the enhanced vehicle trajectory image as label data in the data set;
[0080] It should be noted that in order to efficiently construct the trajectory label data in the dataset, the Gaussian kernel function is used to smooth the original vehicle trajectory images in the dataset to enhance the trajectory, so as to highlight the vehicle trajectory features that are partially covered by noise, which is convenient for the annotation of trajectory labels.
[0081] S30, using a vehicle trajectory image with different characteristics similar to the label data in the dataset as a trajectory image condition, and generating a synthetic noisy image with real trajectory characteristics under the control of the trajectory image condition;
[0082] It should be noted that by using self-set vehicle trajectory images with labels similar to those in the dataset as trajectory image conditions, high-quality synthetic noisy images with real trajectory characteristics that conform to the characteristics of the dataset are generated. This can simulate and generate different traffic conditions (traffic density, speed, noise distribution, etc.), expand the dataset, increase the diversity and versatility of the data, and thus improve the robustness of the denoising model obtained through subsequent training. It can also significantly reduce the cost of manual labeling and reduce labor costs.
[0083] S40, using the synthesized noisy image and the trajectory image conditions as training data and label data in the data set, respectively, to obtain an expanded data set;
[0084] S50. The denoising generative adversarial model is trained based on the expanded data set to obtain a trained generator as a denoising model for denoising the road vibration signal of any road section to obtain vehicle trajectory information.
[0085] It should be noted that the denoising generative adversarial model is trained using the expanded dataset, so that the trained denoising model can denoise noisy vehicle trajectory images corresponding to any road section in complex environments, especially when there is a lot of noise in the fiber optic sensor data, and extract vehicle trajectory information from the image efficiently and in real time. The model is more robust and stable, which meets the needs of real-time and efficient vehicle trajectory monitoring on highways.
[0086] It should be understood that the method of this embodiment can also be used to monitor vehicle trajectories on other types of roads, by selecting road vibration signals of other types of roads as training data to perform the above processing and train the denoising generative adversarial model.
[0087] Furthermore, this embodiment collects road vibration signals based on existing distributed fiber optic sensing equipment. The specific implementation process is: a communication optical fiber is deployed under the surface of the central green belt of the highway as a test optical fiber for monitoring. The communication optical fiber is 16,000 meters long, and the data is divided into 32,000 channels along the length of the optical fiber. The spatial interval between adjacent channels is 0.5 meters, so as to perform real-time monitoring of the vibration signals caused by vehicles traveling on the highway.
[0088] The distributed fiber-optic sensing device in this embodiment utilizes Φ-OTDR (phase-sensitive optical time-domain reflectometry) technology and primarily consists of a high-coherence pulsed light source and an optical signal detector. The ends of the sensing optical fibers are connected to a pulse transmitter and a signal receiver via a fiber coupling device, forming a closed-loop optical path. As the pulsed light source propagates along the fiber, it is Rayleigh-scattered by the fiber material, generating backscattered light. This backscattered light interferes with the incident pulsed light, which is then received by the detector and converted into an electrical signal. This signal is then transmitted to a host computer via a data interface for processing.
[0089] When external vibrations induce micro-deformations in the optical fiber, the local refractive index and axial stress change, causing nonlinear shifts in the phase and intensity of the Rayleigh backscattered light. This characteristically distorts the original interference waveform captured by the device. The device, deployed in communication cabinets along highways, uses a high-frequency sampling rate of 2kHz.
[0090] It should be noted that this embodiment can perform real-time measurement of the ground vibration signal of the highway through distributed optical fiber sensing equipment to obtain original road vibration signals at multiple locations including vehicle driving signals and background noise signals.
[0091] As a further preferred technical solution, step S10: pre-processing the road vibration signal of a certain road section using Hilbert transform to obtain the noisy original vehicle trajectory image as training data in the data set, specifically includes the following steps:
[0092] S11, performing Fourier transform on a road vibration signal collected from a certain road section and then processing the signal through a low-pass filter to obtain a quasi-static signal, wherein the road vibration signal is a time series signal including vehicle driving signals and background noise signals at multiple locations;
[0093] Specifically, this embodiment performs Fourier transform on the road vibration signal of each time series and then filters it through a low-pass filter to highlight the quasi-static signal concentrated in the range of 0 to 2.5 Hz in the original road vibration signal of each time series, so as to preliminarily filter out some obvious noise.
[0094] S12. Calculate the envelope of the quasi-static signal using Hilbert transform as the original vehicle trajectory image with noise, and use the original vehicle trajectory image as training data in the dataset, wherein the original vehicle trajectory images of multiple position points within the time period constitute a spatiotemporal distribution waterfall diagram.
[0095] Specifically, this embodiment uses Hilbert transform to calculate the envelope of each time series quasi-static signal, and represents the quasi-static signal of the kth time series (k=1, 2, ..., 512) as X k (t), then the envelope E of the time series quasi-static signal k (t) can be expressed as:
[0096]
[0097] Where: H(·) is the Hilbert transform.
[0098] This embodiment completes data preprocessing by arranging the original vehicle trajectory image data at multiple locations within a period of time into a 512×512 spatiotemporal distribution waterfall diagram to form the training data in the original data set I.
[0099] It should be noted that this embodiment performs Hilbert transform processing on the quasi-static signal to highlight the vehicle trajectory signal in the quasi-static signal in a complex noise environment, which is convenient for manual labeling and subsequent denoising processing. Otherwise, the obtained data will be unable to be labeled or processed due to a large amount of noise.
[0100] As a further preferred technical solution, step S20: using a Gaussian kernel function to enhance the original vehicle trajectory image to obtain an enhanced vehicle trajectory image, and marking the label of the enhanced vehicle trajectory image as label data in the data set, specifically includes the following steps:
[0101] S21, using a Gaussian kernel function to perform trajectory smoothing on the original vehicle trajectory image to enhance the trajectory, thereby obtaining an enhanced vehicle trajectory image;
[0102] Among them, the formula of Gaussian kernel function is expressed as:
[0103]
[0104] Where: a=(s,t) T is a space-time vector, which represents the space-time coordinate offset relative to the center of the Gaussian kernel; ∑ is a two-dimensional covariance matrix, which is used to control the shape and direction of the kernel function.
[0105] Furthermore, The covariance matrix Σ follows Σ=VΛV T Decompose, σ s is the spatial standard deviation, σ st is the spatiotemporal covariance, σ s is the time standard deviation, where V is the eigenvector matrix and Λ is the eigenvalue matrix. The expressions are as follows:
[0106]
[0107] make The final expression of the two-dimensional covariance matrix Σ is:
[0108]
[0109] Where γ is the average direction angle in the velocity range (v1, v2), and λ1 and λ2 are characteristic values.
[0110] Furthermore, the speed range (v1, v2) and the spatial standard deviation σ are changed according to the actual situation of the highway section to be monitored. s To control the two-dimensional Gaussian kernel function, use it to convolve the waterfall map in the dataset to enhance the trajectory of the vehicle in the image to highlight the trajectory covered by the noise.
[0111] S22. Label the enhanced vehicle trajectory image to obtain label data in the data set.
[0112] It should be noted that the trajectory image data thus obtained is combined with some manual annotation to construct the training label data I in the dataset. label , so far the training data I and label data I in the vehicle trajectory dataset are completed label 's construction.
[0113] As a further preferred technical solution, step S30: using a vehicle trajectory image with different characteristics similar to the label data in the dataset as a trajectory image condition, and generating a synthetic noisy image with real trajectory characteristics under the control of the trajectory image condition, specifically includes the following steps:
[0114] S31, mapping the trajectory image condition to a low-dimensional latent space through a first encoder to obtain trajectory image condition information;
[0115] It should be noted that this embodiment maps the image data, namely the input trajectory image conditions and the initial noise image, to the latent feature vectors and low-dimensional latent variables in the low-dimensional latent feature space respectively, so as to reduce the input space resolution and thus reduce the computational complexity of the semantic image controlled diffusion model, while giving the denoising generative adversarial model the ability to learn the latent space data features, thereby improving the denoising robustness.
[0116] Among them, the first encoder is the encoding part in the residual VAE (Residual-Variational Autoencoder, RES-VAE) module, and the first decoder described below is the decoding part in the residual VAE module. The residual VAE module uses the residual block in the traditional variational autoencoder (Variational Autoencoder, VAE).
[0117] S32, taking the trajectory image condition information as input, embedding it into the semantic image controlled diffusion model via the semantic image condition embedding module to obtain predicted noise, wherein the semantic image controlled diffusion model adopts a U-Net framework and is pre-trained using the dataset for dataset expansion;
[0118] S33, performing reverse diffusion based on a semantic image controlled diffusion model under the control of trajectory image condition information, gradually inputting predicted noise starting from pure Gaussian noise, and gradually constructing a denoised image through variational inference;
[0119] Specifically, the semantic image control diffusion model is an improved semantic image control DDIM module (SIC-DDIM) based on the diffusion model (Denoising Diffusion Implicit Models, DDIM). The semantic image control DDIM module is used to perform a forward diffusion process of continuous noise addition and a backward diffusion process of gradual denoising under the control of trajectory image condition information to construct a denoised image.
[0120] S34. Restoring the denoised image through a first decoder to obtain a synthesized noisy image, wherein the first encoder and the first decoder have symmetrical structures.
[0121] Specifically, as mentioned above, the first decoder is the decoding part in the residual VAE module, which is symmetrical with the first encoder structure. The denoised image is restored by the first decoder to obtain the final required synthetic noisy image generated by the conditional trajectory image.
[0122] It should be noted that since the vehicle trajectories in the original road vibration signal data are completely submerged by noise, it may be impossible to obtain complete vehicle trajectories for some road vibration signal data using only the previous steps, and manual annotation is required. In order to reduce labor costs, increase data diversity, and improve the robustness of subsequent denoising models, this embodiment proposes a semantic image control-latent diffusion model (SIC-LDM), which includes the above-mentioned RES-VAE module, SIC-DDIM module, and semantic image condition embedding module PicEmb. The RES-VAE module maps the input image data into latent variables or latent feature vectors in a low-dimensional latent space to reduce the input spatial resolution and thus reduce the model calculation amount. The PicEmb module embeds any vehicle trajectory image with similar label data as a semantic image condition, i.e., a trajectory image condition, into the SIC-DDIM module. The SIC-DDIM module controls the generation of a synthetic noisy image with vehicle trajectories in the image that meets the characteristics of the training data according to the trajectory image condition, thereby expanding the dataset.
[0123] It should be noted that this embodiment pre-uses the training data I and trajectory label data I in the data set label The semantic image conditioned latent diffusion model (SIC-LDM) is trained, and then the trained semantic image conditioned latent diffusion model (SIC-LDM) is used to input self-set trajectory image conditions similar to the labels in the dataset and an initial noise image generated from a standard Gaussian distribution. Under the control of the trajectory image conditions, a high-quality synthetic noisy image with real trajectory features that conforms to the characteristics of the dataset is generated based on the initial noise image.
[0124] As a further preferred technical solution, before step S30: using a vehicle trajectory image with different characteristics similar to the label data in the dataset as a trajectory image condition and generating a synthetic noisy image with real trajectory characteristics under the control of the trajectory image condition, the method further includes:
[0125] Construct the semantic image conditional control latent diffusion model SIC-LDM and use the training data I and trajectory label data I in the dataset label The latent diffusion model of semantic image conditional control is trained, where the structure of the latent diffusion model of semantic image conditional control is as follows Figure 7 As shown in Figure 2, the training process of the latent diffusion model conditioned on semantic images is as follows:
[0126] (1) Structure and training of RES-VAE module
[0127] 1-1) The RES-VAE module includes a first encoder E VAE and a first decoder D VAE ,in:
[0128] The first encoder E VAE , including the initial feature extraction layer, the residual convolution block group and the potential space mapping layer, where the residual convolution block group adopts stride connection. Taking the residual convolution block group adopting the four-stage residual structure as an example, the encoder E VAE For the input image x (dimension 512×512×1), the following calculations are performed in sequence:
[0129] The initial feature extraction layer includes a 3×3 convolution kernel followed by a LeakyReLU activation function:
[0130] E0=LeakyReLU(Conv 3×3 (x))
[0131] The input image x is expanded to 32 channels using a 3×3 convolution kernel, and the initial feature extraction layer outputs a feature map E0 (dimension 512×512×32).
[0132] The residual convolution block group includes a four-stage residual structure. Each residual structure adopts stride connection. Each residual structure consists of two convolutional layers and a normalization layer connected in sequence. The feature extraction process of the residual convolution block group is as follows:
[0133] ResBlock(F)=ReLU(BN(Conv 3×3 (ReLU(Conv 3×3 (F)))))
[0134]
[0135] Among them, F represents the input of ResBlock, For the Concat operation, the encoder E passes through four residual blocks, and the last two residual blocks are followed by a convolutional layer Conv with a step size of 2. 3×3 , the spatial resolution is gradually halved to 128×128, and the number of channels is increased to 128.
[0136] The latent space mapping layer contains a 3×3 convolution, which proceeds as follows:
[0137] z=Conv 3×3 (E L )
[0138] This completes the encoder E VAE The encoding operation maps the input image to a low-dimensional latent space z = E VAE (x).
[0139] Among them, the first decoder D VAE , satisfying D VAE (z)≈x, the first decoder is symmetrical with the first encoder. The difference is that the convolution Conv in the residual block ResBlock is converted into 3×3 Replaced with transposed convolution ConvTrans 3×3 To reconstruct the image step by step, the image reconstruction process is as follows:
[0140] ResBlock ′ (F) = ReLU(BN(ConvTrans 3×3 (ReLU(ConvTrans 3×3 (F)))))
[0141]
[0142] D 0ut =Sigmoid(Conv 3×3 (D L ))
[0143] The first two residual blocks ResBlock ′ Gradually expand the spatial resolution back to 512×512, and finally restore the feature channel to 1 to obtain a synthetic noisy image that approximates the input x
[0144] 1-2) The training process of the RES-VAE module is as follows:
[0145] Use the RES-VAE module to pre-train the training data I and the trajectory label I label The first encoder maps the data to a low-dimensional latent space, and the first decoder reconstructs the original data from the latent space. In this process, the first decoder D VAE Reconstruct the latent variable z into an image that approximates the features of the input image x
[0146] The goal of the RES-VAE module is to minimize the sum of the reconstruction error and the KL divergence, which measures the difference between the latent space distribution q(z|x) and the prior distribution p(z), and adjusts the coefficient through the hyperparameter β. The total loss function during RES-VAE module training is for:
[0147]
[0148] Where: E Q(Z|x) [log p(x|z)] is the expected log-likelihood of the decoder generating data x under the encoder distribution q(z|x), DKL is the KL divergence, D KL [q(z|x)∥p(z)] measures the difference between the two distributions q(z|x) and p(z).
[0149] This embodiment minimizes the loss function Pre-training allows the RES-VAE module to effectively learn the latent representation of the input image data. After pre-training, the RES-VAE module can compress the image x in the dataset into a low-dimensional latent variable z of size 128×128. At this point, the RES-VAE module is trained. In the following model training steps, the parameters of the RES-VAE module are frozen.
[0150] (2) Semantic image condition embedding module PicEmb, including convolution layer, group normalization layer, etc. The PicEmb module includes the trajectory image condition information Pic tr Processing and embedding of time step TimeEmb:
[0151]
[0152] Where: DownSample is the downsampling operation, GP is the group normalization, For the Concat operation, TimeEmb uses Transformer-style sinusoidal position encoding as the embedding of time step t. During the training process, Pic tr By trajectory label I label The first encoder E of the previously pre-trained RES-VAE VAE Get, that is, Pic tr =E VAE (I label ).
[0153] (3) Structure and training of the SIC-DDIM module
[0154] The SIC-DDIM module includes a forward process with continuous noise addition and a backward diffusion process with gradual noise removal. Both processes use the same U-Net framework. In the forward process, the low-dimensional latent variable z obtained by the first encoder of the pre-trained RES-VAE module is used as input. Under the control of the trajectory image condition information embedded by the semantic image condition embedding module PicEmb, the predicted noise ∈ θ The back diffusion process starts with pure Gaussian noise and gradually inputs it into the U-Net framework to obtain the predicted noise ∈ θ , and gradually construct the denoised image through variational inference.
[0155] 3-1) The SIC-DDIM module includes a second encoder, a first intermediate bottleneck layer, and a second decoder. The second encoder contains a four-stage downsampling path. Each downsampling path contains four consecutive residual blocks and a self-attention submodule. During the downsampling path, the trajectory image condition information Pic obtained by the first encoder is embedded through the semantic image condition embedding module PicEmb. tr :
[0156]
[0157] In this process, each layer of downsampling DownSample reduces the image size by half, and the feature channels of each layer are doubled. The semantic image condition embedding module PicEmb uses DownSample to embed the trajectory image condition information PicEmb synchronously. tr The size of is reduced to embed into each downsampling path, and the number of feature channels is also doubled with the number of downsampling layers, where
[0158] The first intermediate bottleneck layer of the U-Net framework consists of two consecutive residual blocks, and the number of features remains unchanged, namely:
[0159]
[0160] The structure of the second decoder of the U-Net framework is symmetrical to the second encoder structure of the forward process, and additionally includes a stride connection part. Each upsampling path expands the image size by upsampling UpSample, and finally obtains the predicted noise ∈ θ ,as follows:
[0161]
[0162] Furthermore, the output of the U-Net framework can be expressed as: ∈ θ (x t ,t,c), where c is the trajectory image condition information, namely Pic tr .
[0163] 3-2) Training of the forward process of the SIC-DDIM module
[0164] The training data in the training set I is denoted as x, corresponding to the label data set I label The label data x in label , x and x label Input the RES-VAE module to obtain latent variables z and z labelWhen training the forward process of the SIC-DDIM module, Gaussian noise needs to be gradually added to z until the data becomes pure noise. In practice, the total time step is set to 1000. First, a time step t~U(1,T) is randomly sampled from the uniform distribution. The low-dimensional latent variable z is denoised after t steps to obtain a pure Gaussian noise image x. t , the noise is calculated by the following formula:
[0165]
[0166] Among them, ∈ is standard Gaussian noise, α t is the noise scheduling coefficient, and x′0 is the starting image of the forward process, that is, the low-dimensional latent variable z output by the first encoder.
[0167] Next, z label and time step t are embedded into the U-Net framework through the PicEmb module for conditional control, while x t Input U-Net framework and get prediction noise ∈ θ (x t ,t,c), where c is the semantic image condition, i.e. Pic tr . Use the following loss function To train this process:
[0168]
[0169] Where: MSE is the mean square error function.
[0170] 3-3) Backward diffusion process of SIC-DDIM module
[0171] After the forward process of the SIC-DDIM module is trained, the same U-Net framework can be used for back-diffusion to generate images. The back-diffusion process of the SIC-DDIM module is constructed through variational inference, and its denoising process is:
[0172]
[0173] Where: α t is the noise scheduling coefficient, ∈ θ (x t ,t,c) is the prediction noise, σ t To control the randomness parameters, σ can be adjusted according to the actual situation. t Whether to set to 0 controls whether the diffusion process is deterministic.
[0174] The back diffusion process starts from a pure Gaussian noise image x t To begin, combine the image condition information Pic trEmbedded by PicEmb module, the U-Net framework is used to estimate the current prediction noise ∈ θ (x t ,t,c), Since the SIC-DDIM module does not rely on the Markov chain and can perform step-by-step processing, the scheduling coefficient α is recursively calculated in advance through the above formula t Sequence, then simplify the original 1000 time steps into k steps as needed according to the formula, and gradually denoise the image by k-step calculation to obtain the potential variable x0 generated by diffusion, completing the reverse diffusion process.
[0175] Finally, x0 is restored through the previously pre-trained first decoder to obtain the final required synthetic noisy image generated by the conditional trajectory image:
[0176] x final =D VAE (x0).
[0177] As a further preferred technical solution, after the SIC-LDM model training is completed, the trained SIC-LDM model is used to train the self-set label data similar to the data set I label The vehicle trajectory images with different features are mapped into the image condition information Pic in the low-dimensional latent space by executing the above steps S31 to S34. tr =E VAE (trajectory), then generate a pure Gaussian noise image x from a standard Gaussian distribution y , through the reverse diffusion process of the SIC-DDIM module, the noise is gradually removed, and the synthetic noisy image x generated by the conditional trajectory image is obtained. gen , change x gen and trajectory respectively expand the training data and label data in the training set.
[0178] Furthermore, this embodiment randomly extracts 20% of the data from the expanded data set for data enhancement, including scaling, rotation, perspective transformation and flipping. The scaling amplitude does not exceed 2 times, and the rotation amplitude does not exceed 15 degrees to avoid abnormal data that does not conform to the actual situation. At this point, the generation and expansion of the data set is completed.
[0179] As a further preferred technical solution, Figure 8 As shown, the structure of the denoising generative adversarial model in step S50 includes a generator G and a discriminator D, wherein:
[0180] The structure of the generator G uses the U-Net architecture, including the third encoder, the second intermediate bottleneck layer and the third decoder part. The third encoder G enc The downsampling includes three layers:
[0181] G enc1 =LeakyReLU(Dropout(Conv 4×4 (x)))
[0182] G enc2 =LeakyReLU(Dropout(BN(Conv 4×4 (G enc1 ))))
[0183] G enc5 =LeakyReLU(Dropout(BN(Conv 4×4 (G enc2 ))))
[0184] Dropout is the loss layer, and the coefficient is set to 0.2. Conv 4×4 Using a stride factor of 2, in this process, the resolution of the 512×512×1 input image is reduced by half each time, the features are first expanded to 64, and then doubled each time, finally obtaining G enc3 The size is 64×64×256.
[0185] The second intermediate bottleneck layer G of the generator mid As follows, the number of features remains unchanged:
[0186] G mid =Conv 3×3 (G enc3 )
[0187] The third decoder G of the generator dec Symmetric to the third encoder structure, use transposed convolution and add stride connection, and finally add the output layer as follows:
[0188] G dec1 =LeakyReLU(Dropout(BN(ConvTrans 4×4 (G mid ))))
[0189]
[0190] G out =Tanh(Dropout(Conv 3×3 (G dec3 )))
[0191] Among them, ConvTrans 4×4 This is a transposed convolution. In this process, the image dimension is gradually restored from 64×64×256 to 512×512×1 to obtain the generated denoised image.
[0192] The discriminator D consists of four layers as follows:
[0193] D1=LeakyReLU(Conv 4×4 (x))
[0194] D2=LeakyReLU(BN(Conv 4×4 (D1)))
[0195] D3=LeakyReLU(BN(Conv 4×4 (D2)))
[0196] D4=Sigmoid(Conv 3×3 (D3)))
[0197] Among them, Conv 4×4 The stride coefficient is 2. In this process, the size of the input image of 512×512×1 is reduced by half in each of the first three layers, and the number of features is doubled from 64 in each layer, that is, the D3 dimension is 64×64×256, and finally the Conv 3×3 Convert the dimensions to 64×64×1.
[0198] It should be noted that in order to extract vehicle trajectories, this embodiment proposes a lightweight denoising generative adversarial model LD-GAN (Light Denoise-Generative Adversarial Networks) to perform denoising training on the previously constructed vehicle trajectory dataset. The LD-GAN model is obtained by improving the architecture of the generative adversarial model GAN, including a generator G and a discriminator D. In the denoising process, only the generator is used without the discriminator. The generator only uses a three-layer U-Net structure, which contains fewer parameters and has low computational overhead. Generator G is used to denoise noisy data. The goal is to generate high-quality denoised images to deceive the discriminator D. The discriminator D is used to distinguish whether the input data is real data or generated data. The two perform adversarial training together. The specific training process of the LD-GAN model is as follows:
[0199] First, the RES-VAE module is used to classify the labels I in the vehicle trajectory dataset. label Perform preprocessing to obtain the potential feature vector z of the label data label In the following training process, the parameters of the RES-VAE module are frozen first. Next, the LD-GAN model is trained. The training process is as follows:
[0200] (1) Training the discriminator D: Freeze the parameters of the generator G, and let the generator denoise the noisy image x in the original dataset I to obtain x denoise , that is, G(x)=x denoise, whose corresponding trajectory label is x label , the loss function is calculated by the following formula
[0201]
[0202] (2) Train the generator G, adversarial loss: freeze the discriminator D parameters and unfreeze the generator G. denoise The input RES-VAE encoder is mapped into the latent variable z in the latent feature space denoise , combined with the corresponding z label Calculate the loss function in the latent feature space
[0203]
[0204] Where: Denotes the loss function of the discriminator, G(x) = x denoise , x label is the denoised trajectory image x denoise The trajectory label of , D represents the discriminator, and BCE represents the binary cross entropy loss; represents the loss function of the generator, z denoise is the denoised trajectory image x denoise The first encoder maps the latent variables into the latent feature space, z label is the potential feature vector of the label data, MAE is the mean absolute error loss, and MSE is the mean square error loss.
[0205] Furthermore, this embodiment uses the generator G in the trained LD-GAN model, sets the dropout layer discard rate in the generator G to 0, and uses it as a denoising model.
[0206] Furthermore, during actual vehicle trajectory monitoring, the highway ground vibration signal of any road section detected by the distributed optical fiber sensing equipment is pre-processed by Hilbert transform and then input into the denoising model for denoising to obtain vehicle trajectory information.
[0207] The following experimental effect diagram is given to compare and illustrate the advantages of this solution. Figure 2 The spatiotemporal waterfall diagram obtained after Hilbert transform preprocessing is shown. Figure 3 The denoising results obtained by training using the traditional U-Net structure are shown in the figure. Figure 4 The denoising results shown are obtained by training only the generator G in the LD-GAN model and Figure 5Comparing the denoising results using a fully adversarially trained LD-GAN model, the LD-GAN architecture demonstrates superior denoising performance compared to the traditional U-Net architecture, resulting in clearer and more complete vehicle trajectories. Using the LD-GAN generator alone without the discriminator for adversarial training yields even poorer denoising results.
[0208] In addition, if Figure 6 As shown, another embodiment of the present invention provides a highway vehicle trajectory monitoring system, the system comprising:
[0209] The preprocessing module 10 is used to preprocess the road vibration signal of a certain road section using Hilbert transform to obtain the original vehicle trajectory image with noise as the training data in the data set;
[0210] An enhancement processing module 20 is used to enhance the original vehicle trajectory image using a Gaussian kernel function to obtain an enhanced vehicle trajectory image, and mark the label of the enhanced vehicle trajectory image as label data in the data set;
[0211] An image synthesis module 30 is configured to use a vehicle trajectory image with different characteristics similar to the label data in the dataset as a trajectory image condition, and to generate a synthetic noisy image with real trajectory characteristics under the control of the trajectory image condition;
[0212] The data set expansion module 40 is used to use the synthetic noisy image and the trajectory image conditions as training data and label data in the data set respectively to obtain an expanded data set;
[0213] The adversarial training module 50 is used to train the denoising generative adversarial model based on the expanded data set, and obtain the trained generator as a denoising model for denoising the road vibration signal of any road section to obtain vehicle trajectory information.
[0214] As a further preferred technical solution, the pre-processing module 10 includes:
[0215] A quasi-static signal extraction unit is used to perform Fourier transform on a road vibration signal collected from a certain road section and then process it through a low-pass filter to obtain a quasi-static signal, wherein the road vibration signal is a time series signal including vehicle driving signals and background noise signals at multiple locations;
[0216] The envelope extraction unit is used to calculate the envelope of the quasi-static signal using the Hilbert transform as the original vehicle trajectory image with noise, and use the original vehicle trajectory image as the training data in the data set, wherein the original vehicle trajectory images of multiple position points within a time period constitute a spatiotemporal distribution waterfall diagram.
[0217] As a further preferred technical solution, the enhanced processing module 20 specifically includes:
[0218] a trajectory enhancement unit, configured to perform trajectory smoothing on the original vehicle trajectory image using a Gaussian kernel function to enhance the trajectory, thereby obtaining an enhanced vehicle trajectory image;
[0219] The labeling unit is used to label the enhanced vehicle trajectory image and obtain the label data in the dataset
[0220] As a further preferred technical solution, the image synthesis module 30 is deployed with Figure 7 The SIC-LDM model shown in the figure can be used to generate synthetic noisy images after training the SIC-LDM model using the training data and label data in the dataset. The specific process is as follows:
[0221] Mapping the trajectory image condition to a low-dimensional latent space through a first encoder to obtain trajectory image condition information;
[0222] Taking trajectory image condition information as input, embedding the trajectory image condition information into a semantic image controlled diffusion model to obtain prediction noise, wherein the semantic image controlled diffusion model adopts a U-Net framework and is pre-trained using the dataset;
[0223] Under the control of trajectory image condition information, reverse diffusion is performed based on the semantic image controlled diffusion model. Starting from pure Gaussian noise, prediction noise is gradually input, and a low-dimensional latent variable image is gradually constructed through variational inference.
[0224] The low-dimensional latent variable image is restored through a first decoder to obtain a synthetic noisy image, wherein the first encoder and the first decoder have symmetrical structures.
[0225] Furthermore, if Figure 7 As shown in Figure 2, the SIC-LDM model includes the RES-VAE module, the SIC-DDIM module, and the semantic image conditional embedding module PicEmb, where:
[0226] The RES-VAE module includes a first encoder and a first decoder. The first encoder includes an initial feature extraction layer, a residual convolution block group 1, and a potential space mapping layer connected in sequence. The initial feature extraction layer includes a 3×3 convolution, and the 3×3 convolution is followed by a LeakyReLU activation function. The residual convolution block group 1 includes a first residual convolution block E1, a second residual convolution block E2, a third residual convolution block E3, and a fourth residual convolution block E4. The input of the previous residual convolution block and its output are concat-operated to obtain the feature as the input of the next residual convolution block. The output of the fourth residual convolution block E4 is connected to the potential The spatial mapping layer and the potential space mapping layer adopt a 3×3 convolution; the first residual convolution block E1 and the second residual convolution block E2 both include a residual block ResBlock and a 3×3 convolution connected in sequence; the third residual convolution block E3 and the fourth residual convolution block E4 both include a residual block ResBlock and a 3×3 convolution connected in sequence, and the 3×3 convolution is followed by a convolution layer with a stride of 2; the residual block ResBlock includes two convolution layers connected in sequence, wherein the first convolution layer is followed by an activation function ReLU, and the second convolution layer is followed by a normalization layer and an activation function ReLU in sequence.
[0227] Among them, the structure of the first decoder is symmetrical with that of the first encoder. The first decoder includes a residual convolution block group 2 and a feature output layer connected in sequence. The residual convolution block group 2 includes a first residual convolution block D1, a second residual convolution block D2, a third residual convolution block D3 and a fourth residual convolution block D4 connected in sequence. The feature obtained by the concat operation of the input and output of the previous residual convolution block serves as the input of the next residual convolution block; the output of the fourth residual convolution block D4 is connected to the feature extraction layer, which includes a 3×3 convolution and an activation function Sigmoid connected in sequence;
[0228] The first residual convolution block D1 and the second residual convolution block D2 both include a sequentially connected residual block ResBlock' and a 3×3 convolution, and the 3×3 convolution is followed by a transposed convolution ConvTrans 3×3 , the third residual convolution block D3 and the fourth residual convolution block D4 both include sequentially connected residual blocks ResBlock' and a 3×3 convolution; compared with the residual block ResBlock, the residual block ResBlock' replaces the two convolution layers in the residual block ResBlock with transposed convolution ConvTrans 3×3 .
[0229] The SIC-DDIM module includes a second encoder, a first intermediate bottleneck layer, and a second decoder. The second encoder includes a convolutional layer and four sequentially connected downsampling paths. Each downsampling path includes a sequentially connected self-attention submodule and four consecutive residual blocks ResBlock. In each downsampling path, the fourth residual block ResBlock is connected to the second encoder via the downsampling layer DownSample. The semantic image conditional embedding module PicEmb synchronously uses DownSample to embed the trajectory image condition information Pic tr The size of the image is reduced to embed it into each downsampling path; the output of the fourth downsampling path is connected to the second decoder through the first intermediate bottleneck layer. The second decoder includes four upsampling paths and a noise prediction layer connected in sequence. Each upsampling path includes a self-attention submodule and four consecutive residual blocks ResBlock connected in sequence. The output of the fourth residual block ResBlock is connected to the upsampling layer UpSample. The semantic image conditional embedding module PicEmb synchronously uses DownSample to embed the trajectory image conditional information Pic tr The size of is reduced to embed into each upsampling path, and the noise prediction layer includes a group normalization layer GP, an activation function Swish, and a 3×3 convolution connected in sequence.
[0230] The semantic image condition embedding module PicEmb includes a processing layer and an inclusion layer, tracking image condition information Pic tr The outputs of the processing layer and the embedding layer are concat-operated and used as the input of the convolution layer. The processing layer includes a convolution layer, a group normalization layer, and an activation function Swish connected in sequence. The inclusion layer uses a Transformer-style sinusoidal position encoding as the embedding of time step t.
[0231] As a further preferred technical solution, the adversarial training module 50 is specifically used to train the denoising generative adversarial model based on the expanded data set. The denoising generative adversarial model includes a generator and a discriminator. The generator is used to denoise the training data in the expanded data set to obtain a denoised trajectory image. The discriminator is used to distinguish whether the input denoised trajectory image is real data or generated data. The training loss function used includes:
[0232]
[0233] Where: Denotes the loss function of the discriminator, G(x) = x denoise , x label is the denoised trajectory image x denoise The trajectory label of , D represents the discriminator, and BCE represents the binary cross entropy loss; represents the loss function of the generator, zdenoise is the denoised trajectory image x denoise The first encoder maps the latent variables into the latent feature space, z label is the potential feature vector of the label data, MAE is the mean absolute error loss, and MSE is the mean square error loss.
[0234] Furthermore, the structure of the denoising generative adversarial model is as follows Figure 8 As shown, it specifically includes a generator and a discriminator. The generator includes a third encoder, a second intermediate bottleneck layer and a third decoder. The third encoder includes a three-layer encoding structure, wherein the first layer encoding structure includes a 4×4 convolution layer, which is followed by a dropout layer Dropout, and the dropout layer is followed by an activation function LeakyReLU; the second layer encoding structure is the same as the third layer encoding structure, including a 4×4 convolution layer, a normalization layer BN, and a dropout layer Dropout connected in sequence, and the dropout layer is followed by an activation function LeakyReLU; the second intermediate bottleneck layer uses a 3×3 convolution layer, and the third decoder It includes three layers of decoding structures and an output layer connected in sequence. The first layer decoding structure and the second layer decoding structure are symmetrical with the third layer encoding structure and the second layer encoding structure respectively, except that the 4×4 convolution layer is replaced by a 4×4 transposed convolution. Compared with the first decoding structure, the third encoding structure replaces the 4×4 convolution layer with a 4×4 transposed convolution. The output layer includes a 3×3 convolution layer and a loss layer connected in sequence, and the loss layer is followed by an activation function Tanh; the third layer encoding structure is connected to the first decoding structure via the second intermediate bottleneck layer, the second encoding structure is connected to the second decoding structure via a stride connection, and the first encoding structure is connected to the third decoding structure via a stride connection.
[0235] The discriminator includes four layers of structures D1, D2, D3 and D4 connected in sequence. The first layer structure D1 includes a 4×4 convolution layer followed by an activation function LeakyReLU; the second layer structure D2 and the third layer structure D3 both include a 4×4 convolution layer and a normalization layer BN connected in sequence, and the normalization layer is followed by an activation function LeakyReLU; the fourth layer structure D4 includes a 3×3 convolution layer followed by an activation function Sigmoid.
[0236] It should be noted that other embodiments or specific implementation methods of the highway vehicle trajectory monitoring system of the present invention can refer to the above-mentioned method embodiments and will not be repeated here.
[0237] It should be noted that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disk read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0238] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0239] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0240] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0241] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A method for monitoring vehicle trajectories on a highway, characterized in that: include: The road vibration signal of a certain road section is preprocessed using Hilbert transform to obtain the original vehicle trajectory image with noise as the training data in the dataset; The original vehicle trajectory image is enhanced using the Gaussian kernel function to obtain an enhanced vehicle trajectory image, and the label of the enhanced vehicle trajectory image is marked as the label data in the dataset; Using vehicle trajectory images with different characteristics similar to the label data in the dataset as trajectory image conditions, and generating synthetic noisy images with real trajectory characteristics under the control of the trajectory image conditions; The synthetic noisy image and trajectory image conditions are used as training data and label data in the dataset respectively to obtain the expanded dataset; The denoising generative adversarial model is trained based on the expanded dataset, and the trained generator is used as a denoising model to denoise the road vibration signal of any road section to obtain vehicle trajectory information.
2. The highway vehicle trajectory monitoring method according to claim 1, characterized in that: The method of preprocessing the road vibration signal of a certain road section by using Hilbert transform to obtain the original vehicle trajectory image with noise as the training data in the data set includes: The road vibration signal collected from a certain road section is subjected to Fourier transformation and then processed through a low-pass filter to obtain a quasi-static signal. The road vibration signal is a time series signal containing vehicle driving signals and background noise signals at multiple locations. The envelope of the quasi-static signal is calculated using Hilbert transform as the original vehicle trajectory image with noise, and the original vehicle trajectory image is used as the training data in the dataset. The original vehicle trajectory images of multiple position points within a time period constitute a spatiotemporal distribution waterfall diagram.
3. The highway vehicle trajectory monitoring method according to claim 1, characterized in that: The method of using the Gaussian kernel function to enhance the original vehicle trajectory image to obtain an enhanced vehicle trajectory image and marking the label of the enhanced vehicle trajectory image as label data in the data set includes: Use the Gaussian kernel function to smooth the original vehicle trajectory image to enhance the trajectory and obtain an enhanced vehicle trajectory image; Label the enhanced vehicle trajectory image to obtain the label data in the dataset; Among them, the formula of Gaussian kernel function is expressed as: Where a is a space-time vector, which represents the space-time coordinate offset relative to the center of the Gaussian kernel; ∑ is a two-dimensional covariance matrix, which is used to control the shape and direction of the kernel function.
4. The highway vehicle trajectory monitoring method according to claim 1, characterized in that: The method uses a vehicle trajectory image with different characteristics similar to the label data in the dataset as a trajectory image condition, and generates a synthetic noisy image with real trajectory characteristics under the control of the trajectory image condition, including: Mapping the trajectory image condition to a low-dimensional latent space through a first encoder to obtain trajectory image condition information; Taking trajectory image condition information as input, embedding the trajectory image condition information into a semantic image controlled diffusion model to obtain prediction noise, wherein the semantic image controlled diffusion model adopts a U-Net framework and is pre-trained using the dataset; Under the control of trajectory image condition information, reverse diffusion is performed based on the semantic image controlled diffusion model. Starting from pure Gaussian noise, prediction noise is gradually input, and a low-dimensional latent variable image is gradually constructed through variational inference. The low-dimensional latent variable image is restored through a first decoder to obtain a synthetic noisy image, wherein the first encoder and the first decoder have symmetrical structures.
5. The highway vehicle trajectory monitoring method according to claim 4, characterized in that: The processing process of the first encoder on the trajectory image condition or the initial noise image is expressed as: E0=LeakyReLU(Conv 3×3 (x)) ResBlock(F)=ReLU(BN(Conv 3×3 (ReLU(Conv 3×3 (F))))) z=Conv 3×3 (E L ) Where: E0 is the initial feature extracted from the trajectory image condition or the initial noise image, ResBlock(F) represents the feature obtained by inputting feature F into the residual block, and E l Represents the concatenation of the output features of the l-1th residual convolution block and the output features of the l-2th residual convolution block after convolution, l = 1, 2, ..., L, L is the total number of residual convolution blocks, Conv 3×3 represents a 3×3 convolution kernel, represents the Concat operation, LeakyReLU and ReLU represent activation functions, z is the low-dimensional latent variable obtained by the first encoder processing the initial noise image during training, Pic tr is the trajectory image condition information obtained by the first encoder processing the trajectory image condition, x is the trajectory image condition or the initial noise image, E VAE (x) = z or E VAE (x) = Pic tr The process representation for the first encoder.
6. The highway vehicle trajectory monitoring method according to claim 4, characterized in that: The low-dimensional latent variable is used as input, and the predicted noise is obtained through the U-Net framework under the control of the trajectory image condition information. The formula is expressed as: Where: ∈ θ To predict noise, Conv 3×3 represents a 3×3 convolution kernel, represents the Concat operation, Swish is the activation function, and GP represents group normalization. represents the output features of the last residual block in the second decoder, represents the output feature of the l-1th residual block, represents the output feature of the lth residual block in the second encoder, Represents the input features of the first residual block in the second encoder, UpSample represents upsampling, DownSample represents downsampling, ResBlock ×L represents L consecutive residual blocks, Attention represents the attention mechanism, Pic tr is the trajectory image condition information, H mid It represents the feature obtained by the middle bottleneck layer using M consecutive residual blocks to operate on the output feature of the Lth residual block in the second encoder, PicEmb (Pic tr ) indicates that DownSample is used to convert the trajectory image condition information Pic tr The size of is reduced to embed into each downsampling path and upsampling path, and TimeEmb represents the use of Transformer-style sinusoidal position encoding as the embedding of time step t, where the second encoder and the second decoder are symmetrical in structure.
7. The highway vehicle trajectory monitoring method according to claim 4, characterized in that: The method performs back diffusion based on the U-Net framework under the control of trajectory image condition information, gradually inputs prediction noise starting from pure Gaussian noise, and gradually constructs a low-dimensional latent variable image through variational inference, including: Under the control of trajectory image condition information, reverse diffusion is performed based on the U-Net framework, starting from pure Gaussian noise and gradually inputting prediction noise. The formula is expressed as: Where: α t is the noise scheduling coefficient at time step t, α t-1 is the noise scheduling coefficient at time step t-1, ∈ θ (x t ,t,c) is the prediction noise, σ t is the parameter that controls randomness, x t is a pure Gaussian noise image, and ∈ is a vector sampled from a standard Gaussian distribution. Through variational inference on the pure Gaussian noise image x t Denoising is performed step by step to obtain a low-dimensional latent variable image x0.
8. The highway vehicle trajectory monitoring method according to claim 1, characterized in that: The denoising generative adversarial model is trained based on the expanded data set to obtain a trained generator as a denoising model for denoising road vibration signals of any road section to obtain vehicle trajectory information, including: The denoising generative adversarial model is trained based on the expanded dataset. The denoising generative adversarial model includes a generator and a discriminator. The generator is used to denoise the training data in the expanded dataset to obtain a denoised trajectory image. The discriminator is used to distinguish whether the input denoised trajectory image is real data or generated data. The training loss function used includes: Where: Denotes the loss function of the discriminator, G(x) = x denoise , x label is the denoised trajectory image x denoise The trajectory label of , D represents the discriminator, and BCE represents the binary cross entropy loss; represents the loss function of the generator, z denoise is the denoised trajectory image x denoise The first encoder maps the latent variables into the latent feature space, z label is the potential feature vector of the label data, MAE is the mean absolute error loss, and MSE is the mean square error loss.
9. The highway vehicle trajectory monitoring method according to claim 8, characterized in that: The process of obtaining vehicle trajectory information by denoising the road vibration signal of any road section by the denoising model is expressed as follows: G enc1 =LeakyReLU(Dropout(Conv 4×4 (x ′ ))) G enc2 =LeakyReLU(Dropout(BN(Conv 4×4 (G enc1 )))) G enc3 =LeakyReLU(Dropout(BN(Conv 4×4 (G enc2 )))) G mid =Conv 3×3 (G enc3 ) G dec1 =LeakyReLU(Dropout(BN(ConvTrans 4×4 (G mid )))) G out =Tanh(Dropout(Conv 3×3 (G dec3 ))) Where x ′ is the noisy vehicle trajectory image obtained by preprocessing the road vibration signal of any road section, Dropout is the loss layer, Conv 4×4 is a 4×4 convolution kernel, LeakyReLU is the activation function, BN is the normalization layer, Conv 3×3 represents a 3×3 convolution kernel, Represents the Concat operation, ConvTrans 4×4 is the transposed convolution, Tanh is the activation function, G enc1 , G enc2 , G enc3 Represents the output features of the three-layer structure of the third encoder in the generator, G mid is the output feature of the middle bottleneck layer in the generator, G dec1 , G dec2 , G dec3 Represents the output features of the three-layer structure of the third decoder in the generator, G out For the denoised vehicle trajectory information, the third encoder and the third decoder have symmetrical structures.
10. A highway vehicle trajectory monitoring system, characterized in that: include: The preprocessing module is used to preprocess the road vibration signal of a certain road section using Hilbert transform to obtain the original vehicle trajectory image with noise as the training data in the data set; An enhancement processing module is used to enhance the original vehicle trajectory image using a Gaussian kernel function to obtain an enhanced vehicle trajectory image, and mark the label of the enhanced vehicle trajectory image as label data in the data set; An image synthesis module is configured to use a vehicle trajectory image with different characteristics similar to the label data in the dataset as a trajectory image condition, and generate a synthetic noisy image with real trajectory characteristics under the control of the trajectory image condition; The dataset expansion module is used to use the synthetic noisy image and trajectory image conditions as training data and label data in the dataset respectively to obtain the expanded dataset; The adversarial training module is used to train the denoising generative adversarial model based on the expanded dataset, and the trained generator is used as a denoising model to denoise the road vibration signal of any road section to obtain vehicle trajectory information.
Citation Information
Cited By
Student behavior data enhancement method based on improved DDIM model
CN121392481A
A student behavior data enhancement method based on an improved DDIM model
CN121392481B
Method and system for monitoring vehicles in tunnel based on linear frequency modulation pulse DAS
CN121545364A
Vehicle tracking method based on fusion of all-in-one machine and distributed optical fiber
CN121831760A