Road high-frequency data frequency resampling processing method and system based on neural network

Through the neural network-based method, convolutional neural network and Transformer model are used to automatically learn the deep features of multi-sensor data, and adaptive sampling frequency processing of high-frequency data on the road is realized, complex problems of data fusion and analysis are solved, and efficient data processing and storage are realized.

CN120010742APending Publication Date: 2025-05-16RES INST OF HIGHWAY MINIST OF TRANSPORT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510072275.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Due to the sampling frequency differences in road high-frequency data, data fusion and analysis are complex and time-consuming, and the prior art is difficult to effectively solve this problem.

Method used

Using a neural network-based method, the deep features related to sampling frequency selection in multi-sensor data are automatically learned through convolutional neural networks and Transformer models, and adaptive downsampling or upsampling processing of high-frequency data is realized.

Benefits of technology

The optimal downsampling or upsampling processing of multi-sensor data is realized, maintaining the characteristics and correlation of data, significantly saving storage space, and improving the efficiency of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010742A_ABST
    Figure CN120010742A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and particularly discloses a road high-frequency data frequency resampling processing method and system based on a neural network, and through adaptive sampling frequency processing based on the neural network and a Transform large model, the system can realize optimal down-sampling or up-sampling processing of multi-sensor data. The processed data can keep the characteristics of the original data of the sensor and the correlation with the data of other sensors, and meanwhile, the storage space can be remarkably saved under the condition that the data analysis and storage efficiency is not influenced. The method is suitable for high-frequency data analysis scenes such as intelligent transportation and automatic driving, and has good expansibility and application potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and in particular relates to a method and system for frequency resampling processing of high-frequency road data based on a neural network. Background Art

[0002] High-frequency road data usually has a very high sampling frequency, but the data sampling frequencies of various sensors often vary greatly, making data fusion and analysis complex and time-consuming. After completing the collection of these massive high-frequency data, the subsequent data timestamp alignment, multi-sensor data fusion, and subsequent data analysis often face huge technical challenges. These difficulties are not only due to the complexity and diversity of sensor data, but are also further exacerbated by the specific requirements of different application scenarios for sampling frequency.

[0003] In order to effectively address these challenges, this application proposes a method for adaptive sampling frequency processing of high-frequency road data based on a neural network. Summary of the invention

[0004] The present invention aims to provide a mechanism for adaptively downsampling or upsampling high-frequency data of different sampling frequencies, thereby achieving efficient processing of multiple sensor data and ultimately outputting fused data with an optimal sampling frequency.

[0005] In order to solve the above technical problems, the specific technical solutions of the present invention are as follows:

[0006] In some embodiments of the present application, a method and system for frequency resampling of high-frequency road data based on a neural network are provided, comprising the following steps:

[0007] Step 1) sensor data preprocessing;

[0008] Step 2) labeling the sampling frequency of the sensor data, and determining a reasonable sampling frequency for the data according to different sensor data characteristics and application requirements through sampling frequency labeling;

[0009] Step 3) Align the sampling frequencies of all sensor data, and process the data of all sensors to the same sampling frequency;

[0010] Step 4) using a convolutional neural network to predict the optimal downsampling or upsampling frequency of the data, by automatically learning deep features in the multi-sensor data related to the sampling frequency selection using a convolutional neural network (CNN);

[0011] Step 5) Use the Transformer model to process the data according to the predicted sampling frequency, and optimize the sampling frequency of the sensor data by using the Transformer model.

[0012] In some embodiments of the present application, the sensor data preprocessing in step 1 includes:

[0013] Step 1.1) Confirmation of sampling frequency;

[0014] Step 1.2) Data format conversion: The data conversion formula is as follows:

[0015]

[0016] Among them, scale_factor is set according to the range of the sensor;

[0017] Step 1.3) Data cleaning and missing value filling;

[0018] Step 1.3.1) Linear interpolation: Use the average value of adjacent data points to fill in the missing values, as follows:

[0019]

[0020] Step 1.3.2) Mean filling: For data with less non-time series impact, the mean of all data of the sensor can be used for filling:

[0021] x filled =mean(x non-missing );

[0022] Step 1.3.3) Filling based on statistical models: For missing data, Bayesian inference or Markov chain Monte Carlo (MCMC) methods are used to fill in the missing data; among them, mean filling is for data with less impact on non-time series, and the mean of all data of the sensor is used to fill in the missing data:

[0023] x filled =mean(x non-missing );

[0024] Step 1.3.4) Statistical model-based imputation: For missing data, Bayesian inference or Markov chain Monte Carlo method is used to impute the missing data to improve the accuracy of the imputed value;

[0025] Step 1.4) Outlier detection and processing;

[0026] Step 1.4.1) 3 times standard deviation method: When the difference between a data point and the mean exceeds 3 times the standard deviation, it is considered an outlier:

[0027] x outlier >μ+3σorx outlier <μ-3σ;

[0028] Step 1.4.2) Box plot method: The upper and lower quartile ranges of the data were identified through the box plot, and points exceeding 1.5 times the interquartile range were considered outliers;

[0029] Step 1.4.3) Processing method: For detected outliers, choose to replace them with the mean of the previous and next values, interpolation method, or directly eliminate them;

[0030] Step 1.5) Data normalization;

[0031] Step 1.5.1) Min-max normalization: Scale the data to the range [0,1] using the following formula:

[0032]

[0033] Step 1.5.2) Z-score normalization: For normally distributed data, scale the data by the standard deviation and center it around the mean:

[0034]

[0035] Among them, μ is the mean and σ is the standard deviation;

[0036] Step 1.5.3) Logarithmic normalization: For data with a large range, use logarithmic normalization:

[0037] x norm =log(x+1).

[0038] In some embodiments of the present application, the sampling frequency marking of the sensor data in step 2 includes:

[0039] Step 2.1) Determine the original sampling frequency and data characteristics of each sensor. Before labeling, the original sampling frequency f of each sensor is i and its data characteristics for analysis;

[0040] Step 2.2) Manual experience annotation of optimal frequency;

[0041] Step 2.3) Data analysis assists in determining the labeling, using data analysis methods to compare and analyze data at different sampling frequencies to determine the optimal downsampling or upsampling frequency;

[0042] Step 2.4) Establish a sampling frequency annotation table. After completing the data feature analysis and manual annotation, the optimal sampling frequency annotation results need to be recorded in the form of a structured table to form a sampling frequency annotation table;

[0043] Step 2.5) Optimization and verification of annotations.

[0044] In some embodiments of the present application, step 3: aligning the sampling frequencies of all sensor data includes:

[0045] Step 3.1) Determine a unified sampling frequency benchmark. During the alignment process, align the data sampling frequencies of all sensors to the highest actual sampling frequency in the data set, and select the highest frequency as the benchmark to retain the maximum detail information of the original data;

[0046] Step 3.2) Use data interpolation to up-convert low-frequency data. During the multi-sensor alignment process, sensor data with low sampling frequency needs to be up-converted to align its sampling points to the target frequency.

[0047] Step 3.3) Repeat sampling method is used to reduce the frequency of high-frequency data. For some sensors whose data sampling frequency is higher than the alignment frequency, the data volume can be reduced by reducing the frequency;

[0048] Step 3.4) Timestamp synchronization: After completing the up-conversion and down-conversion, it is necessary to ensure that the data of each sensor is completely consistent with the timestamp at the target sampling frequency;

[0049] Step 3.5) Sampling frequency alignment result verification: after completing the sampling frequency alignment, the accuracy of the alignment result is evaluated by a data verification method;

[0050] Step 3.6) Output the aligned multi-sensor data. After completing the sampling frequency alignment, interpolation, standardization and other processing, the multi-sensor data is output as a unified time series data.

[0051] In some embodiments of the present application, step 4: using a convolutional neural network to predict the optimal downsampling or upsampling frequency of data includes:

[0052] Step 4.1) Construction of input data matrix: all sensor data are integrated into a matrix, each row of which represents the time-aligned data from different sensors;

[0053] Step 4.2) Model output;

[0054] Step 4.3) Design of the structure of convolutional neural network;

[0055] Step 4.4) Use the model to make predictions.

[0056] In some embodiments of the present application, step 5: using the Transformer model to process the data according to the predicted sampling frequency includes:

[0057] Step 5.1) Input and structure of the Transformer model. The input of the Transformer model consists of some preprocessed sensor data and the optimal sampling frequency;

[0058] Step 5.2) Transformer encoding and decoding modules, including encoder and decoder modules;

[0059] Step 5.3) Use Transformer to make predictions;

[0060] Step 5.4) Save the prediction results and apply;

[0061] Step 5.5) Verification and feedback of prediction results.

[0062] In some embodiments of the present application, a neural network-based road high-frequency data frequency resampling processing system is disclosed, which adopts the above technical solution, including:

[0063] A data processing module, which is used to unify the format of the collected data, clean up missing data values ​​and adjust the numerical scale;

[0064] A data labeling module, which receives the data information processed by the data processing module and classifies and labels it according to different features;

[0065] A data alignment module, wherein the data alignment module performs alignment processing on the classified and annotated data information;

[0066] The data recognition module receives the aligned data information and uses a convolutional neural network to predict the optimal downsampling or upsampling frequency of the data, thereby identifying and distinguishing the frequency information of the data and its local pattern in the time series, and identifying the relationship and connection between different sensor data.

[0067] Compared with the prior art, the beneficial effect of the present invention is that, through adaptive sampling frequency processing based on neural networks and Transformer large models, the system can achieve optimal downsampling or upsampling processing of multi-sensor data. The processed data can not only maintain the characteristics of the original sensor data and the correlation with other sensor data, but also significantly save storage space without affecting data analysis and storage efficiency. This method is suitable for high-frequency data analysis scenarios such as intelligent transportation and autonomous driving, and has good scalability and application potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0069] Figure 1A schematic diagram of the working principle provided by an embodiment of the present invention; DETAILED DESCRIPTION

[0070] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0071] In order to better understand the purpose, structure and function of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings.

[0072] See attached Figure 1 As shown, according to some embodiments of the present application, the following steps are included:

[0073] Step 1) Sensor data preprocessing: Data preprocessing is crucial in the processing of multi-sensor data. Due to the complex environment of high-frequency road data collection and the variety of equipment, the sampling frequencies and data formats of different sensors are different. Therefore, the data preprocessing step aims to unify the data format, clean up missing data values, and adjust the numerical scale, laying a solid foundation for subsequent frequency alignment and data fusion. The following is a detailed process of data preprocessing.

[0074] Step 1.1) Confirmation of sampling frequency Since the data sampling frequencies of different sensors may vary greatly, for example, the accelerometer may be sampled at a high frequency of 2000Hz, while the gyroscope sensor is usually sampled at a frequency of 500Hz. Therefore, the data sampling frequency f of each sensor must be accurately recorded first. i For each sensor’s sampling frequency, they are summarized into a sampling frequency vector F = [f1, f2, ... f n ] to facilitate subsequent sampling alignment.

[0075] For example, if there are three sensors in the system and their sampling frequencies are 2000Hz, 500Hz and 200Hz respectively, then the sampling frequency vector is F = [2000, 500, 200].

[0076] Step 1.2) Data format conversion

[0077] The data formats collected by each sensor may be different. For example, some sensors output floating point numbers, while others may be integers or strings. In order to unify the data format, all data are usually converted to floating point format to facilitate subsequent processing. The data conversion formula is as follows:

[0078]

[0079] The scale_factor is set according to the range of the sensor to ensure that the data units are unified.

[0080] Step 1.3) Data cleaning (missing value filling)

[0081] In the process of high-frequency data collection, data may be missing due to signal loss, sensor failure or communication delay. These missing values ​​will affect the accuracy of data analysis, so they need to be filled.

[0082] Step 1.3.1) Linear interpolation: Use the average value of adjacent data points to fill in the missing values. For example, if the kth data point is missing, use the average value of the two adjacent points to fill it in:

[0083]

[0084] Step 1.3.2) Mean filling: For data with less non-time series impact, the mean of all data of the sensor can be used for filling:

[0085] x filled =mean(x non-missing )

[0086] Step 1.3.3) Statistical model-based imputation: For context-dependent missing data, methods such as Bayesian inference or Markov Chain Monte Carlo (MCMC) are used for imputation to improve the accuracy of the imputed values.

[0087] Step 1.4) Outlier detection and processing

[0088] Multi-sensor data may contain outliers, such as sensors suddenly detecting abnormally high / low temperatures or speeds, which may be due to interference or equipment failure. These outliers will mislead subsequent data analysis and must be identified and processed in the preprocessing stage.

[0089] Step 1.4.1) 3 times standard deviation method: If a data point xxx differs from the mean by more than 3 times the standard deviation, it is considered an outlier:

[0090] x outlier >μ+3σorx outlier <μ-3σ

[0091] Step 1.4.2) Box plot method: The upper and lower quartile ranges of the data were identified through the box plot, and points exceeding 1.5 times the interquartile range were considered outliers.

[0092] Step 1.4.3) Processing method: For detected outliers, you can choose to replace them with the mean of the previous and next values, interpolation method, or directly eliminate them.

[0093] Step 1.5) Data normalization

[0094] Since the dimensions of different sensor data are obviously different, such as temperature in degrees Celsius and acceleration in meters per square second, data normalization can eliminate these differences and avoid the effect of dimension effects in subsequent models. Common normalization methods are as follows:

[0095] Step 1.5.1) Minimum-maximum normalization: Scale the data to the range [0,1], which is suitable for cases where the data range is known:

[0096]

[0097] Step 1.5.2) Z-score normalization: For normally distributed data, scale the data by the standard deviation and center it around the mean:

[0098]

[0099] Among them, μ is the mean and σ is the standard deviation.

[0100] Step 1.5.3) Logarithmic normalization: For data with a large range, such as acceleration data, logarithmic normalization can be used:

[0101] x norm =log(x+1)

[0102] Through the above detailed data preprocessing steps, various types of multi-sensor data can be effectively standardized and integrated, laying a solid foundation for subsequent frequency alignment, downsampling or upsampling, and neural network model training.

[0103] Step 2) Label the sampling frequency of the sensor data. Through sampling frequency labeling, a reasonable sampling frequency is determined for the data according to different sensor data characteristics and application requirements. Sampling frequency labeling is one of the key steps in multi-sensor data processing. Its purpose is to determine the optimal downsampling or upsampling frequency for each sensor data to achieve more efficient sampling frequency distribution during data processing and storage. Considering that different sensors have different requirements for data sampling frequency, this step provides an important frequency alignment basis for subsequent model training through empirical analysis and labeling. The following is the detailed process of sampling frequency labeling:

[0104] Step 2.1) Determine the original sampling frequency and data characteristics of each sensor

[0105] Before labeling, the original sampling frequency f of each sensor needs to be i The data characteristics are analyzed. This includes but is not limited to the number of samples per second, the typical data change amplitude, the stability of the signal, etc. These characteristics help determine whether downsampling or upsampling is needed. The specific analysis content is as follows:

[0106] Step 2.1.1) Sampling frequency: The sampling frequency of each sensor is initially determined based on the rate of data change. For example, an accelerometer may require a higher frequency to capture subtle road vibrations, while a gyroscope sensor may require a relatively low sampling frequency.

[0107] Step 2.1.2) Signal characteristics: such as noise level, value range, mutation frequency, etc. These characteristics will affect the smoothness and volatility of the data. For example, data with high noise may require higher accuracy when downsampling.

[0108] Step 2.2) Manual experience to label the optimal frequency

[0109] Due to the unique characteristics of data collected by some sensors and the different sampling frequency requirements in different scenarios, manual experience is usually required to mark the optimal downsampling or upsampling frequency for each sensor data. This process is usually carried out in combination with the knowledge and experience of domain experts to ensure that the marked frequency can effectively retain key data features.

[0110] Step 2.2.1) Domain experts participate: Experts mark the optimal downsampling or upsampling frequency of sensors based on their data analysis experience. For example, for data that does not change much over a period of time, a smaller downsampling frequency can be marked.

[0111] Step 2.2.2) Labeling method: There are many ways to label, including scoring method (priority labeling) and interval method (frequency range labeling), so as to combine the labeling frequency with the data characteristics.

[0112] Step 2.3) Data analysis to assist in labeling

[0113] In order to improve the scientificity of the annotation, data analysis methods can be used to compare and analyze data with different sampling frequencies to determine the optimal downsampling or upsampling frequency. The analysis content usually includes:

[0114] Step 2.3.1) Frequency domain analysis: Use frequency domain analysis (such as Fourier transform) to understand the main frequency components of the signal and determine the main frequency bandwidth contained in the original sampling frequency to ensure that the downsampled data can still capture the main features. The Fourier transform formula is as follows:

[0115]

[0116] Step 2.3.2) Wavelet transform: For non-stationary signals, wavelet transform can be used to decompose the frequency components and identify the appropriate frequency for downsampling or upsampling by comparing the distribution of wavelet coefficients at different sampling frequencies.

[0117] Step 2.3.3) Signal-to-noise ratio (SNR) analysis: Calculate the signal-to-noise ratio of the data at different frequencies and select frequency points where downsampling will not seriously affect the data quality.

[0118] Step 2.4) Create a sampling frequency labeling table

[0119] After completing data feature analysis and manual labeling, the optimal sampling frequency labeling results need to be recorded in a structured table to form a sampling frequency labeling table. This table is used to prepare subsequent neural network training data to ensure that the model can accurately learn the requirements of different frequencies. The basic structure of the sampling frequency labeling table is as follows:

[0120] Step 2.4.1) Sensor ID: used to uniquely identify each sensor.

[0121] Step 2.4.2) Original sampling frequency f original : Record the original sampling frequency of the sensor.

[0122] Step 2.4.3) The sampling frequency f of the data processing marked suggested : The optimal sampling frequency obtained based on data feature analysis and manual experience annotation.

[0123] Step 2.4.4) Data characterization: record the main characteristics of the data such as stability and noise level.

[0124] Example:

[0125]

[0126] Step 2.5) Annotation optimization and verification

[0127] In order to ensure the accuracy and applicability of the labeling frequency, you can select some data from the data sample for testing to evaluate the performance of the labeling frequency. In this process, you can use different downsampling or upsampling frequencies to sample the data and compare it with the original data:

[0128] Step 2.5.1) Mean Square Error (MSE): Evaluate the mean square error of the data sampling results at different frequencies. The formula for the mean square error is as follows:

[0129]

[0130] Step 2.5.2) Cross-validation: Divide the data into training and validation sets, and test the performance of the sampling frequency on different data sets to optimize the selection of the sampling frequency.

[0131] Through the sampling frequency labeling process, it is possible to determine a reasonable sampling frequency for the data according to different sensor data characteristics and application requirements. This not only ensures that the key information of the data is effectively retained during the downsampling or upsampling process, but also significantly reduces the amount of data, improves storage and computing efficiency, and provides higher quality input for subsequent neural network models.

[0132] Step 3) Align the sampling frequencies of all sensor data, and process the data of all sensors to the same sampling frequency;

[0133] Sampling frequency alignment is the core step to achieve multi-sensor data fusion and analysis. Since different sensors often have different sampling frequencies when collecting data, the alignment process standardizes the data of all sensors to a common sampling frequency, so that subsequent data fusion and analysis can be performed on a unified time dimension. In this step, the data of all sensors will be processed to the same sampling frequency, thereby simplifying multi-source data analysis and improving the accuracy of the model in high-frequency data processing. The following is the detailed process of sampling frequency alignment.

[0134] Step 3.1) Determine a unified sampling frequency benchmark

[0135] During the alignment process, the data sampling frequencies of all sensors are usually aligned to the highest actual sampling frequency in the dataset. Selecting the highest frequency as the reference can retain the maximum detail information of the original data, thereby avoiding the loss of key features due to downsampling.

[0136] Step 3.1.1) Highest frequency alignment: Assuming that the sampling frequency of sensor 1 is 2000 Hz and the sampling frequency of sensor 2 is 500 Hz, 2000 Hz is used as the unified alignment frequency to ensure that all sensor data can be aligned at the time point of 2000 Hz.

[0137] Step 3.2) Use data interpolation to up-convert low-frequency data. During the multi-sensor alignment process, the sensor data with low sampling frequency needs to be up-converted to align its sampling points to the target frequency. To achieve this, data interpolation is usually used to generate missing sampling points so that the low-frequency data is aligned with the high-frequency data on the time axis. Common interpolation methods include linear interpolation, spline interpolation, and polynomial interpolation. The specific methods are as follows:

[0138] Step 3.2.1) Linear interpolation: Assume that the low-frequency sensor data at adjacent time points t i and t i+i If there are no sampling points between the two points, linear interpolation can be used to generate additional sampling points between the two points. The linear interpolation formula is:

[0139]

[0140] Where f(t) is the interpolated data value at time t, and t is in [t i ,t i+1 ] range.

[0141] Step 3.2.2) Spline interpolation: Spline interpolation can generate smooth interpolation results and is suitable for signals with stable data changes or less fluctuations. By constructing a cubic spline function between every two adjacent sampling points, the continuity and smoothness of the data can be guaranteed during the upscaling process.

[0142] Step 3.2.3) Polynomial interpolation: Polynomial interpolation is suitable for more complex signal patterns, but it is prone to overfitting and is generally used for upscaling operations in a smaller range.

[0143] Step 3.3) Repeated sampling method to reduce the frequency of high-frequency data If the data sampling frequency of some sensors is higher than the alignment frequency, the amount of data can be reduced by reducing the frequency. Frequency reduction is usually achieved by repeated sampling method, that is, discarding some sampling points while retaining important sampling points. There are mainly the following repeated sampling methods:

[0144] Step 3.3.1) Extraction method: directly extract sampling points from high-frequency data at fixed intervals. For example, if the sensor data sampling frequency is 4000Hz and the target frequency is 1000Hz, then retain one every 4 sampling points.

[0145] Step 3.3.2) Sliding average method: The downsampled data is obtained by applying a sliding window to the high-frequency data and calculating the average value of each window. The size of the sliding window is usually set to the ratio of the high-frequency data frequency to the target frequency. For example, if the window size is 4, a new sampling point is generated for every 4 sampling points.

[0146] Step 3.3.3) Low-pass filtering: Apply a low-pass filter to the high-frequency data to retain signal components below the target frequency. The low-pass filter can better retain the low-frequency characteristics of the signal by filtering out high-frequency noise.

[0147] Step 3.4) Timestamp synchronization

[0148] After completing the up-conversion and down-conversion, it is necessary to ensure that the data of each sensor is completely consistent with the timestamp on the target sampling frequency. This timestamp synchronization is usually achieved by assigning a uniform time interval. For example, all sensor data are rearranged to correspond to the time point of the target sampling frequency (such as every 0.5ms). This not only ensures the consistency of the time dimension, but also provides a basis for subsequent data fusion.

[0149] Step 3.4.1) Offset Correction: The sampling time of some sensors may have a slight offset. In this case, the timestamp offset correction method can be used to adjust the data to align with the unified time axis. For example, by adjusting the timestamp of each data point to the nearest standard time point.

[0150] Step 3.5) Verification of sampling frequency alignment results

[0151] After completing the sampling frequency alignment, the accuracy of the alignment result can be evaluated through data verification methods. The mean square error (MSE) and signal-to-noise ratio (SNR) of the aligned data are calculated to verify whether the data has feature loss or distortion during the alignment process.

[0152] Step 3.5.1) Mean square error (MSE): Calculate the error between the aligned data and the original data, and evaluate the alignment effect by the mean square error.

[0153] Step 3.5.2) Signal-to-noise ratio (SNR): Calculate the signal-to-noise ratio of the aligned data to assess signal fidelity.

[0154] Step 3.6) Output aligned multi-sensor data

[0155] After completing the sampling frequency alignment, interpolation, and standardization, the multi-sensor data is output as a unified time series data. The output data can be used for multi-sensor data fusion and subsequent neural network model training. In practical applications, this aligned multi-sensor data is not only easy to process, but also can improve the accuracy and real-time performance of subsequent models.

[0156] Through the above steps, sampling frequency alignment provides a unified time dimension basis for multi-sensor data, ensuring the consistency, reliability and applicability of the data. This alignment process significantly simplifies the analysis and processing of data with different frequencies, and provides high-quality input data for subsequent deep learning models.

[0157] Step 4) using a convolutional neural network to predict the optimal downsampling or upsampling frequency of the data, by automatically learning deep features in the multi-sensor data related to the sampling frequency selection using a convolutional neural network (CNN);

[0158] In this critical step, deep features related to sampling frequency selection in multi-sensor data can be automatically learned by using convolutional neural networks (CNN). Convolutional neural networks have significant advantages in processing multi-dimensional data with spatiotemporal characteristics. They can identify and distinguish the frequency information of data and its local patterns in time series, and can identify the relationship and connection between different sensor data. Therefore, they are very suitable for optimizing the sampling frequency of complex multi-sensor high-frequency data.

[0159] Step 4.1) Construction of input data matrix

[0160] First, all sensor data are integrated into a matrix, where each row represents time-aligned data from different sensors. The sampled data of each sensor is preprocessed, labeled, and frequency-aligned in the previous steps to form a unified sampling benchmark, providing a stable input source for the model. Specifically, the structure of the matrix X is as follows:

[0161] Step 4.1.1) Time step dimension: The columns of the matrix represent the aligned time steps, which record the sampling values ​​of each sensor at each moment with a uniform time granularity.

[0162] Step 4.1.2) Sensor dimension: The rows of the matrix represent the channels of different sensors, and each channel corresponds to the data sequence of a sensor.

[0163] For example, assuming the aligned frequency is 1000 Hz, there are 5 sensors, and each sample time window covers 1 second, the size of the input matrix is ​​(5, 1000), where 5 is the number of sensors and 1000 is the number of time steps after alignment.

[0164] Step 4.2) Model output

[0165] The output of the convolutional neural network is a sampling frequency prediction vector y for the optimal processing of all sensors. The detailed description of the output format is as follows:

[0166] Step 4.2.1) Output categories: The output is a vector where each value corresponds to the optimal sampling frequency of a sensor. The optimal sampling frequency of each sensor is processed as a classification task in the model, so the output may be a probability distribution of fixed categories (such as 500Hz, 1000Hz, etc.).

[0167] Step 4.2.2) Output shape: The output shape is (N,), which is a vector with length equal to the number of sensors. Each position corresponds to the predicted optimal downsampling / upsampling frequency for that sensor. Specifically, if the optimal sampling frequency for a sensor is predicted to be 500Hz, then the value of that position is 500.

[0168] Step 4.3) Convolutional Neural Network Architecture Design

[0169] The structure of this convolutional neural network is specially designed for multi-sensor data frequency learning tasks, aiming to extract deep frequency features in sensor data and predict the optimal downsampling or upsampling frequency. Its network structure includes the following main parts:

[0170] Step 4.3.1) Input layer: Receives multi-sensor data in matrix form. The shape of the input tensor is (number of sensors, number of time steps), which is regarded as a two-dimensional signal with multi-channel input in CNN.

[0171] Step 4.3.2) One-dimensional convolution layer: Through a one-dimensional convolution kernel (i.e., a convolution kernel that slides only in the time step dimension), the time series features are extracted for the data sequence of each sensor channel. Using multiple convolution kernels of different sizes (such as 3, 5, 7, etc.) can capture different frequency features.

[0172] Step 4.3.3) Pooling layer: Each convolutional layer is followed by a pooling layer (such as max pooling or average pooling) to reduce the dimension, reduce the amount of calculation and retain important features. This step reduces noise interference while retaining important frequency features.

[0173] Step 4.3.4) Feature channel merging: The features of different sensor data are merged and interactively learned in the deep convolutional layer, so that CNN can understand the association between different sensor frequency patterns.

[0174] Step 4.3.5) Fully connected layer: The convolution and pooling feature data are expanded in the fully connected layer and mapped to specific sampling frequency categories. The number of output nodes of the fully connected layer is equal to the number of optional sampling frequency categories. The constructed convolutional neural network model is input into the convolutional neural network, and the optimally processed sampling frequency prediction vector y of all sensors is output, which is expressed as:

[0175] y = CNN(X)

[0176] Among them, CNN represents the trained convolutional neural network model.

[0177] Step 4.4) Use the model to make predictions

[0178] The trained model can be used to predict the optimal sampling frequency of sensor data in real time. The prediction process is as follows:

[0179] Step 4.4.1) Input data preprocessing: The collected sensor data is preprocessed and frequency aligned in the previous steps and organized into an input matrix with the same structure as the training data.

[0180] Step 4.4.2) Model prediction: The input matrix is ​​passed into the model, and the model outputs the optimal sampling frequency category for each sensor based on the features extracted by the convolution kernel. The output is a vector representing the optimal sampling frequency for each sensor.

[0181] Example: Assume that the predicted output of the model is [500, 1000, 500], that is, the optimal sampling frequency of the first and third sensors is 500Hz, and the second sensor is 1000Hz.

[0182] Step 5) Use the Transformer model to process the data according to the predicted sampling frequency, and optimize the sampling frequency of the sensor data by using the Transformer model.

[0183] In this step, the sensor data is further processed through the Transformer model to achieve downsampling or upsampling after the sampling frequency is optimized. The purpose of using the Transformer is to take advantage of its advantages in time series and multi-dimensional feature processing to generate downsampled or upsampled data at the optimal frequency, thereby preserving the original characteristics of the sensor and the correlation between sensors. The specific steps are as follows:

[0184] Step 5.1) Input and structure of Transformer model

[0185] The input of the Transformer model consists of two main parts: some preprocessed sensor data and the optimal sampling frequency. The detailed description of each input part is as follows:

[0186] Step 5.1.1) Sensor raw data sequence: The data of a certain sensor is preprocessed, including removing missing values, normalization, etc. Assuming that the frequency after final alignment is 1000 Hz and the time window is T seconds, the input vector shape of the sensor data is T×1000, and the input vector length of the sensor data is N.

[0187] Step 5.1.2) Optimal sampling frequency label: In the output of the convolutional neural network in step 4, each sensor has been assigned an optimal downsampling or upsampling frequency. This frequency is encoded as an additional input to the model to guide the Transformer to generate data at the target frequency.

[0188] Step 5.2) Transformer encoding and decoding modules

[0189] Step 5.2.1) The Transformer model usually consists of two modules: encoder and decoder. It processes input data and sampling frequency labels through a hierarchical attention mechanism to achieve adaptive reconstruction of data frequency.

[0190] Step 5.2.2) Encoder: The task of the encoder is to extract the feature information of the input sensor data and integrate the temporal features of the sensor data into a unified representation. The encoder structure is as follows:

[0191] Step 5.2.3) Input Embedding: The sensor data sequence is first converted to a high-dimensional representation through a linear layer to form an input embedding matrix. The shape of the embedding matrix is ​​N×d model , where dmodel is the hidden layer dimension of the Transformer model.

[0192] Step 5.2.4) Position encoding: Since Transformer does not have temporal information, position encoding is used to mark the position of each time step, and the position encoding matrix is ​​added to the output of the embedding layer to ensure the order of the time series.

[0193] Step 5.2.5) Multi-head self-attention layer: The multi-head self-attention mechanism extracts the dependencies between each time step, captures the local and global patterns of temporal features, and enables the model to recognize the dynamic changes of sensor data.

[0194] Step 5.2.6) Feedforward network: After the multi-head attention layer, a feedforward network is added for nonlinear mapping to enrich the feature representation. The network output is a global representation of the sensor data.

[0195] Step 5.2.7) Decoder: The task of the decoder is to generate the target data after sampling frequency conversion based on the feature representation and optimal sampling frequency output by the encoder.

[0196] Step 5.2.7.1) Input and frequency combination: The decoder input is an information matrix with frequency annotations. The frequency information is embedded into the data generation process through position encoding and attention mechanism.

[0197] Step 5.2.7.2) Masking mechanism: When generating new data, a masking mechanism is used to mask future time steps to ensure that the model makes predictions based only on historical data to avoid information leakage.

[0198] Step 5.2.7.3) Output layer: The final linear layer maps the decoder output to a sequence form consistent with the target sampling frequency, thereby generating processed sensor data.

[0199] Step 5.3) Make predictions using Transformer

[0200] After training, the Transformer model can perform predictions on sensor data with different sampling frequencies to obtain adaptive downsampled or upsampled data. The prediction process is as follows:

[0201] Step 5.3.1) Input data preparation: Arrange the real-time collected sensor data and the optimal sampling frequency labels into a matrix that conforms to the input shape.

[0202] Step 5.3.2) Model prediction: The input matrix is ​​passed into the Transformer model, and the model will output the sequence data at the target frequency based on the feature extraction of the encoder and the frequency adjustment mechanism of the decoder. For example, if the input sensor data frequency is 500Hz and the optimal sampling frequency is 200Hz, the model will process the input sensor data into 200Hz and output it sequentially.

[0203] Step 5.3.3) Data format reconstruction: Reconstruct the sampled data sequence based on the model prediction results for subsequent storage and use.

[0204] Step 5.4) Save the prediction results and apply

[0205] The sampling frequency optimized data output by the Transformer model needs to be saved for use in future multi-sensor data fusion and analysis:

[0206] Step 5.4.1) Result format: The predicted sampled data is usually stored in CSV or JSON format for easy reading and subsequent processing. Each record contains a timestamp, sensor ID, and processed data value.

[0207] Step 5.4.2) Data storage: Store in a database or file, arranged in chronological order, to provide a basis for further data mining, analysis or modeling.

[0208] Step 5.4.3) Data review: Evaluate the accuracy and completeness of the generated data by comparing it with the original data to ensure that key features are not lost.

[0209] Step 5.5) Verification and feedback of prediction results

[0210] In practical applications, the frequency adjustment effect of the model is verified through the feedback mechanism. For example, regular sampling is performed to analyze the consistency between the data generated by the model and the actual situation, and the impact of different sampling frequencies on the model performance is evaluated. Based on the feedback results, the structure of the Transformer model can be further optimized or the sampling frequency annotation standard can be adjusted.

[0211] By using the Transformer model to optimize the sampling frequency of sensor data, the integrity of data features during downsampling and upsampling is ensured. This model can maintain the correlation and consistency between multi-sensor data while ensuring storage efficiency, providing a high-quality input basis for multi-sensor data analysis.

[0212] Ultimately, through this adaptive sampling frequency processing based on neural networks and Transformer large models, the system can achieve optimal downsampling or upsampling of multi-sensor data. The processed data can not only maintain the characteristics of the original sensor data and the correlation with other sensor data, but also significantly save storage space without affecting data analysis and storage efficiency. This method is suitable for high-frequency data analysis scenarios such as intelligent transportation and autonomous driving, and has good scalability and application potential.

[0213] Through the above technical solution, the technical effects produced in the embodiments of the present application are:

[0214] 1. Efficient data compression and storage optimization

[0215] This method achieves intelligent data compression by adaptively adjusting the sampling frequency of different sensor data, thereby retaining key data to the greatest extent possible under limited storage resources. Unlike simple fixed sampling methods, this adaptive downsampling and upsampling method can dynamically adjust the frequency according to the characteristics of the data, ensuring that effective information is not lost, and greatly optimizing the utilization of storage space. The storage and transmission requirements for high-frequency road data are therefore significantly reduced, saving bandwidth and costs, and is suitable for resource-constrained embedded systems and edge computing scenarios.

[0216] 2. Data Integrity and Information Retention

[0217] During the downsampling process, traditional methods may cause data loss or loss of key information, while this adaptive method uses a neural network model to accurately identify and retain the core features of the data. During the data alignment and annotation stage, the system analyzes the characteristics, frequency, and trends of various sensor data to ensure that the timing information and characteristic patterns of the data are still retained after downsampling. This can avoid interference with the data during the sampling frequency adjustment process, so that the data still has a high degree of analysis and application value after downsampling or upsampling.

[0218] 3. Convenience of multi-sensor data fusion

[0219] A major challenge in multi-sensor fusion processing is the difference in sampling frequencies of different data sources, and the sampling frequency alignment mechanism of this method can effectively solve this problem. By aligning data with different sampling frequencies to a unified time base, the method ensures the synchronization and consistency between multi-source data, facilitating subsequent data fusion, analysis, and decision-making. The alignment operation reduces the need for manual intervention through model-driven data processing, and improves the compatibility and applicability of multi-sensor systems through automated and intelligent sampling frequency management.

[0220] 4. Adaptive modeling and dynamic response capabilities

[0221] Traditional sampling methods usually use a predefined fixed frequency, while this method uses a neural network to achieve dynamic adaptive adjustment of the sampling frequency. This method allows the model to automatically select the optimal sampling frequency based on the actual content, characteristics and environmental changes of the input data. This means that when faced with real-time changing road data, the system can respond quickly and adjust the sampling strategy, thereby improving the adaptability to complex road conditions and meeting the needs of road data processing and analysis.

[0222] 5. Intelligent frequency labeling and processing

[0223] This method introduces a combination of expert knowledge-based empirical annotation and neural network training in the sampling frequency annotation stage, achieving a more refined sampling frequency selection. The optimal sampling frequency of different sensor data is more accurate through network training, which can mine implicit information from data features and improve the accuracy of data processing. This intelligent annotation method that combines prior knowledge with machine learning models can greatly reduce errors caused by human intervention and make the processing process more robust and scientific.

[0224] 6. Advantages of deep learning processing of high-frequency data

[0225] One of the core advantages of this method is the use of convolutional neural networks and Transformer models for data processing. Convolutional neural networks can capture temporal and local features in high-frequency data, and are suitable for identifying the core patterns of various sensor data, thereby maintaining the temporal consistency and local details of the data when downsampling or upsampling. The Transformer large model further provides global data association analysis capabilities, allowing the model to determine the best downsampling or upsampling processing effect on a global scale, improving the overall quality and consistency of data processing.

[0226] 7. Save bandwidth and processing resources

[0227] Through adaptive sampling frequency adjustment, this method can significantly reduce the amount of data while maintaining the key characteristics and quality of the data, thereby effectively reducing the data transmission bandwidth usage. Especially for remote data transmission and cloud computing scenarios, the reduction in data volume can significantly reduce network load and transmission delay, further improving the overall operating efficiency of the system. In addition, the optimization of data volume also reduces the computing pressure during downstream data processing and analysis, allowing the system to maintain efficient operation under limited computing resources.

[0228] 8. Improve system reliability and stability

[0229] Since this method effectively handles the complexity and diversity of multi-sensor data, it ensures data quality while also improving the robustness and reliability of the system. The automated sampling frequency adaptive processing method reduces manual intervention and the possibility of operational errors. Especially in high-frequency, high-load road environments, the system can maintain stable data processing capabilities and improve the reliability of the overall data acquisition system.

[0230] In summary, the adaptive sampling frequency processing method for high-frequency road data based on neural networks provides strong technical support for the efficient collection, analysis and application of high-frequency road data by virtue of its efficient storage optimization, intelligent data retention, multi-sensor data fusion, adaptive dynamic response and deep learning processing advantages. It can not only improve the accuracy and efficiency of data processing, but also provide significant optimization for storage, computing and transmission resources, making it have broad application prospects and practical value in application fields such as road condition analysis and road maintenance.

[0231] In some embodiments of the present application, a neural network-based road high-frequency data frequency resampling processing system is disclosed, which adopts the above technical solution, including:

[0232] A data processing module, which is used to unify the format of the collected data, clean up missing data values ​​and adjust the numerical scale;

[0233] A data labeling module, which receives the data information processed by the data processing module and classifies and labels it according to different features;

[0234] A data alignment module, wherein the data alignment module performs alignment processing on the classified and annotated data information;

[0235] The data recognition module receives the aligned data information and uses a convolutional neural network to predict the optimal downsampling or upsampling frequency of the data, thereby identifying and distinguishing the frequency information of the data and its local pattern in the time series, and identifying the relationship and connection between different sensor data.

[0236] In the description of the present application, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.

[0237] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.

[0238] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0239] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0240] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for frequency resampling of high-frequency road data based on a neural network, characterized in that: The following steps are involved: Step 1) sensor data preprocessing; Step 2) labeling the sampling frequency of the sensor data, and determining a reasonable sampling frequency for the data according to different sensor data characteristics and application requirements through sampling frequency labeling; Step 3) Align the sampling frequencies of all sensor data, and process the data of all sensors to the same sampling frequency; Step 4) using a convolutional neural network to predict the optimal downsampling or upsampling frequency of the data, by automatically learning deep features in the multi-sensor data related to the sampling frequency selection using a convolutional neural network (CNN); Step 5) Use the Transformer model to process the data according to the predicted sampling frequency, and optimize the sampling frequency of the sensor data by using the Transformer model.

2. The method for frequency resampling of high-frequency road data based on neural network according to claim 1, characterized in that: The sensor data preprocessing in step 1 includes: Step 1.1) Confirmation of sampling frequency; Step 1.2) Data format conversion: The data conversion formula is as follows: Among them, scale_factor is set according to the range of the sensor; Step 1.3) Data cleaning and missing value filling; Step 1.3.1) Linear interpolation: Use the average value of adjacent data points to fill in the missing values, as follows: Step 1.3.2) Mean filling: For data with less non-time series impact, the mean of all data of the sensor can be used for filling: x filled =mean(x non-missing ); Step 1.3.3) Filling based on statistical models: For missing data, Bayesian inference or Markov chain Monte Carlo (MCMC) methods are used to fill in the missing data; among them, mean filling is for data with less impact on non-time series, and the mean of all data of the sensor is used to fill in the missing data: x filled =mean(x non-missing ); Step 1.3.4) Statistical model-based imputation: For missing data, Bayesian inference or Markov chain Monte Carlo method is used to impute the missing data to improve the accuracy of the imputed value; Step 1.4) Outlier detection and processing; Step 1.4.1) 3 times standard deviation method: When the difference between a data point and the mean exceeds 3 times the standard deviation, it is considered an outlier: x outlier >μ+3σ or x outlier <μ-3σ; Step 1.4.2) Box plot method: The upper and lower quartile ranges of the data were identified through the box plot, and points exceeding 1.5 times the interquartile range were considered outliers; Step 1.4.3) Processing method: For detected outliers, choose to replace them with the mean of the previous and next values, interpolation method, or directly eliminate them; Step 1.5) Data normalization; Step 1.5.1) Min-max normalization: Scale the data to the range [0,1] using the following formula: Step 1.5.2) Z-score normalization: For normally distributed data, scale the data by the standard deviation and center it around the mean: Among them, μ is the mean and σ is the standard deviation; Step 1.5.3) Logarithmic normalization: For data with a large range, use logarithmic normalization: x norm =log(x+1)。 3. The method for frequency resampling of high-frequency road data based on neural network according to claim 1, characterized in that: The step 2 of marking the sampling frequency of the sensor data includes: Step 2.1) Determine the original sampling frequency and data characteristics of each sensor. Before labeling, the original sampling frequency f of each sensor is i and its data characteristics for analysis; Step 2.2) Manual experience annotation of optimal frequency; Step 2.3) Data analysis assists in determining the labeling, using data analysis methods to compare and analyze data at different sampling frequencies to determine the optimal downsampling or upsampling frequency; Step 2.4) Establish a sampling frequency annotation table. After completing the data feature analysis and manual annotation, the optimal sampling frequency annotation results need to be recorded in the form of a structured table to form a sampling frequency annotation table; Step 2.5) Optimization and verification of annotations.

4. The method for frequency resampling of high-frequency road data based on neural network according to claim 1, characterized in that: The step 3: aligning the sampling frequencies of all sensor data includes: Step 3.1) Determine a unified sampling frequency benchmark. During the alignment process, align the data sampling frequencies of all sensors to the highest actual sampling frequency in the data set, and select the highest frequency as the benchmark to retain the maximum detail information of the original data; Step 3.2) Use data interpolation to up-convert low-frequency data. During the multi-sensor alignment process, sensor data with low sampling frequency needs to be up-converted to align its sampling points to the target frequency. Step 3.3) Repeat sampling method is used to reduce the frequency of high-frequency data. For some sensors whose data sampling frequency is higher than the alignment frequency, the data volume can be reduced by reducing the frequency; Step 3.4) Timestamp synchronization: After completing the up-conversion and down-conversion, it is necessary to ensure that the data of each sensor is completely consistent with the timestamp at the target sampling frequency; Step 3.5) Sampling frequency alignment result verification: after completing the sampling frequency alignment, the accuracy of the alignment result is evaluated by a data verification method; Step 3.6) Output the aligned multi-sensor data. After completing the sampling frequency alignment, interpolation, standardization and other processing, the multi-sensor data is output as a unified time series data.

5. The method for frequency resampling of high-frequency road data based on neural network according to claim 1, characterized in that: The step 4: using a convolutional neural network to predict the optimal downsampling or upsampling frequency of the data includes: Step 4.1) Construction of input data matrix: all sensor data are integrated into a matrix, each row of which represents the time-aligned data from different sensors; Step 4.2) Model output; Step 4.3) Design of the structure of convolutional neural network; Step 4.4) Use the model to make predictions.

6. The method for frequency resampling of high-frequency road data based on neural network according to claim 1, characterized in that: The step 5: using the Transformer model to process the data according to the predicted sampling frequency includes: Step 5.1) Input and structure of the Transformer model. The input of the Transformer model consists of some preprocessed sensor data and the optimal sampling frequency; Step 5.2) Transformer encoding and decoding modules, including encoder and decoder modules; Step 5.3) Use Transformer to make predictions; Step 5.4) Save the prediction results and apply; Step 5.5) Verification and feedback of prediction results.

7. A neural network-based road high-frequency data frequency resampling processing system, using a neural network-based road high-frequency data frequency resampling processing method according to any one of claims 1 to 6, characterized in that: include: A data processing module, which is used to unify the format of the collected data, clean up missing data values ​​and adjust the numerical scale; A data labeling module, which receives the data information processed by the data processing module and classifies and labels it according to different features; A data alignment module, wherein the data alignment module performs alignment processing on the classified and annotated data information; The data recognition module receives the aligned data information and uses a convolutional neural network to predict the optimal downsampling or upsampling frequency of the data, thereby identifying and distinguishing the frequency information of the data and its local pattern in the time series, and identifying the relationship and connection between different sensor data.