Road high-frequency data redundancy detection method and system

Through the staged neural network model, the problem of large data volume and redundancy is solved, and efficient data compression and information retention are achieved.

CN119961839APending Publication Date: 2025-05-09RES INST OF HIGHWAY MINIST OF TRANSPORT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510059905.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The huge amount and redundancy of high-frequency data on the road make data processing and storage a difficult problem, and the prior art is difficult to effectively compress data and retain key changes information.

Method used

The data redundancy detection is performed using a staged neural network model, and the data is redundantly evaluated using convolutional neural networks and fully connected neural networks at the minute, hour and day levels, and the appropriate downsampling frequency is selected for data compression based on the detection results.

Benefits of technology

Through the application of staged redundant detection and neural network model, efficient compression of high-frequency data on the road is achieved, the core information of the data is retained to the maximum extent, and the effectiveness and representativeness of the data are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961839A_ABST
    Figure CN119961839A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent traffic, and particularly relates to a road high-frequency data redundancy detection method and system, the method comprises minute-level redundancy detection, hour-level redundancy detection and day-level redundancy detection, and the data redundancy degree in different time intervals can be detected through neural network models of the minute-level redundancy detection, the hour-level redundancy detection and the day-level redundancy detection. According to a detection result, the data can be compressed by selecting different downsampling frequencies, so that key change information is ensured to be reserved, and the validity and representativeness of the data are improved. According to the method, efficient compression of the data is achieved, meanwhile, core information of the road data is reserved to the maximum extent, and a more optimized data basis is provided for high-frequency data analysis in practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent transportation technology, and in particular relates to a method and system for detecting redundancy of high-frequency road data. Background Art

[0002] Since the sampling frequency of high-frequency road data is very high (usually 2000Hz), the amount of data is extremely large. Taking a single sensor as an example, the amount of data per minute reaches 120,000, the amount of data per hour is as high as 7.2 million, and the amount of data per day is even 172.8 million. Faced with such a huge amount of data, this application proposes a method for detecting redundancy of high-frequency road data, so as to effectively compress the amount of data while retaining key changes, thereby improving its representativeness and practical application value. Summary of the invention

[0003] The purpose of the present invention is to provide a method and system for detecting high-frequency road data redundancy, which is divided into three processing stages. Each stage is based on the degree of data redundancy in a specific time interval, and the redundancy is gradually detected and reduced through a neural network model.

[0004] In order to solve the above technical problems, the specific technical solutions of the present invention are as follows:

[0005] In some embodiments of the present application, a method and system for detecting high-frequency road data redundancy are provided, comprising the following steps:

[0006] Step 1) Minute-level redundancy detection: pre-process the collected road data and divide the data into time intervals per minute. The redundancy of the data in each minute is evaluated based on manual annotation, and then the annotated data is used to train the convolutional neural network (CNN) model;

[0007] Step 2) Hourly redundancy detection: Based on minute-level data redundancy detection, redundancy annotation is performed on hourly data intervals using manual experience as training data. Based on these annotated data, a fully connected neural network (FCNN) model is used for training.

[0008] Step 3) Daily redundancy detection: Based on the previous hourly redundancy detection results and manual labeling experience, the daily data redundancy level is evaluated. Then, these labeled data are used to train a fully connected neural network model, which enables it to identify the redundancy level of the daily interval and provide a basis for daily data compression and analysis.

[0009] Through the neural network model in the above steps, the degree of data redundancy in different time intervals is detected. According to the detection results, different downsampling frequencies are selected for compression processing to retain key change information and improve the effectiveness and representativeness of the data.

[0010] In some embodiments of the present application, step 1 includes:

[0011] Step 1.1) Data division and organization: divide the original high-frequency road data into one-minute time periods;

[0012] Step 1.2) Manually label the redundancy level, and manually label the redundancy level of each minute of data;

[0013] Step 1.3) Training of a minute-level data redundancy detection model based on a convolutional neural network: using a convolutional neural network to train the labeled minute-level data, thereby generating a model for identifying the redundancy of a one-minute data length interval, and training the model on a large number of samples of one-minute data intervals;

[0014] Step 1.4) Use the trained convolutional neural network model to predict the redundancy level of minute-length intervals.

[0015] In some embodiments of the present application, step 1.1 includes:

[0016] Step 1.1.1) Raw data collection;

[0017] Step 1.1.2) Data is divided into minute intervals, and the collected raw data is divided into an independent data interval per minute;

[0018] Step 1.1.3) Data standardization: standardize the divided data intervals so that the data distribution is concentrated within a fixed range. The standardization formula is as follows:

[0019] d' i,j =(d i,j -μ) / σ

[0020] in:

[0021] d i,j is the original data point, μ is the mean value in the interval, σ is the standard deviation in the interval, d' i,j is the standardized data;

[0022] After standardization, the original data is transformed into a standard normal distribution with a mean of 0 and a variance of 1.

[0023] In some embodiments of the present application, step 1.2 includes:

[0024] Step 1.2.1) Define the redundancy level annotation range, divide the redundancy level into continuous values ​​from 0 to 1, and define the meaning of redundancy features in different value intervals:

[0025] 0 means no redundancy at all, that is, the data changes significantly and contains a lot of valid information;

[0026] 1 means complete redundancy, that is, the data has almost no changes and is highly repetitive;

[0027] Intermediate values ​​indicate varying degrees of redundancy;

[0028] Step 1.2.2) Feature analysis and redundancy degree labeling. During the labeling process, manual experience is used to analyze the data interval of each minute and determine the redundancy degree based on the data change pattern, fluctuation amplitude, and periodic characteristics;

[0029] Step 1.2.3) Design of annotation tools and processes;

[0030] Step 1.2.4) Construction of the labeled data set. All manually labeled minute data will form a complete training data set for subsequent model training.

[0031] In some embodiments of the present application, step 1.3 includes:

[0032] Step 1.3.1) Design of convolutional neural network model: build a convolutional neural network model with multiple layers of convolution, pooling and fully connected layers;

[0033] Step 1.3.2) Model training.

[0034] In some embodiments of the present application, the design of the convolutional neural network model in step 1.3.1 includes:

[0035] Step 1.3.1.1) Input layer, the input layer receives the preprocessed data matrix X minutes , the shape is 60×2000×1, where "1" indicates a single channel;

[0036] Step 1.3.1.2) Convolutional layers and activation functions, use multiple convolutional layers to extract local features of the time series, the convolution kernel size is designed to be 3×3 or 5×5, the depth of the first convolutional layer output feature map is set to 32, and more feature maps are gradually expanded to 64 and 128 layers;

[0037] The activation function uses ReLU (Rectified Li near Un it), which is defined as:

[0038] f(x)=max(0,x)

[0039] The nonlinear characteristics of ReLU can enhance the expressiveness of the model.

[0040] Step 1.3.1.3) Pooling layer, use the maximum pooling layer for downsampling, and the pooling window size is set to 2×2 or 3×3;

[0041] Step 1.3.1.4) Fully connected layer, after the convolution feature extraction is completed, a fully connected layer is added to integrate all features and output the redundancy probability;

[0042] Step 1.3.1.5) Output layer and activation function. The output layer uses one node to represent the predicted value of the redundancy of the one-minute interval data. The activation function uses the Sigmoid function to limit the output value to [0,1] to facilitate comparison with the actual label. Sigmoid is defined as:

[0043]

[0044] In some embodiments of the present application, the model training in step 1.3.2 includes:

[0045] Step 1.3.2.1) Loss function, select mean square error as the loss function to measure the difference between the model prediction value and the actual label. The loss function formula is:

[0046]

[0047] Where: N represents the number of training samples; y" minute,j is the prediction redundancy of the jth one-minute data interval; y minute,j is the true redundant label of the jth sample.

[0048] Step 1.3.2.2) Optimize the algorithm and use the stochastic gradient descent or Adam optimization algorithm to adjust the parameters of the model so that the loss function gradually converges to the minimum value.

[0049] In some embodiments of the present application, step 1.4 includes:

[0050] Step 1.4.1) Preprocess all data. Process all collected high-frequency data according to the data preprocessing method in step 1. Divide each minute data interval into a 60×2000 matrix format, and normalize and standardize each matrix;

[0051] Step 1.4.2) Minute-level data segmentation

[0052] All collected high-frequency road data are divided into one-minute intervals, and the amount of raw data collected throughout the day is D total , then the number of minute-level intervals N minutes It can be calculated as:

[0053]

[0054] Step 1.4.3) Batch input and model inference to predict the degree of data redundancy per minute. Input the data matrix of each one-minute interval into the convolutional neural network model for batch inference. The matrix of a certain interval is represented as X minute,j , then the model predicts the redundancy level of this interval y" minute,j It can be expressed as:

[0055] y" minute,j =CNN(X minute,j ),

[0056] Among them, CNN represents the trained convolutional neural network model, which is constructed by traversing N minutes intervals, the model can predict the redundancy level of all one-minute intervals;

[0057] Step 1.4.4) Storage and recording: save the redundancy prediction results of each interval to a result file or database for subsequent processing and analysis.

[0058] In some embodiments of the present application, step 2 includes:

[0059] Step 2.1) Redundancy level labeling of hourly data

[0060] Step 2.2) Hourly redundancy detection model training based on fully connected neural network

[0061] Step 2.3) Use the trained fully connected neural network model to predict the redundancy level of all road high-frequency data in hours.

[0062] In some embodiments of the present application, step 3 includes:

[0063] Step 3.1) Redundancy level labeling of daily data

[0064] Step 3.2) Training of the day-level redundancy detection model based on a fully connected neural network

[0065] Step 3.3) Use the trained fully connected neural network model to predict the redundancy level of all road high-frequency data in terms of length.

[0066] In some embodiments of the present application, a system for detecting redundancy of high-frequency road data is disclosed, comprising:

[0067] A minute-level redundancy detection module processes the collected data, divides the data into time intervals in units of minutes, evaluates the redundancy of the data in each minute based on manual annotation, and uses the annotated data to train a convolutional neural network model so that the trained model can automatically identify the redundancy in the data in each minute;

[0068] An hourly redundancy detection module, which uses manual experience to mark the redundancy of each hour's data interval based on the minute-level redundancy detection module, and uses a fully connected neural network model for training to form a model that can identify redundancy in each hour's interval;

[0069] A daily redundancy detection module evaluates the daily data redundancy based on the hourly redundancy detection results and manual labeling experience, and uses these labeled data to train a fully connected neural network model to identify the redundancy level of the daily interval, providing a basis for daily data compression and analysis.

[0070] Compared with the prior art, the beneficial effect of the present invention is that it includes minute-level redundancy detection, hour-level redundancy detection, and day-level redundancy detection. Through the neural network model of the above three stages, the data redundancy degree in different time intervals can be detected. According to the detection results, the data can be compressed at different downsampling frequencies to ensure that key change information is retained and the validity and representativeness of the data are improved. This method achieves efficient data compression while retaining the core information of road data to the maximum extent, providing a more optimized data foundation for high-frequency data analysis in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0072] Figure 1 A schematic diagram of the overall phase flow provided for an embodiment of the present invention;

[0073] Figure 2 The present invention provides an overall flow chart of the embodiment of the present invention. DETAILED DESCRIPTION

[0074] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0075] In order to better understand the purpose, structure and function of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings.

[0076] See attached Figure 1-2 As shown, according to some embodiments of the present application, the method is divided into three processing stages as a whole, and each stage is based on the degree of data redundancy in a specific time interval, and the redundancy is gradually detected and reduced through a neural network model.

[0077] Step 1: Minute-level redundancy detection

[0078] First, the collected road data is preprocessed and divided into time intervals in units of minutes. The redundancy of the data in each minute is evaluated based on manual annotations, and then these annotated data are used to train a convolutional neural network (CNN) model to enable it to identify the redundancy of data at the minute level. The trained model can automatically identify the redundancy in each minute of data, laying the foundation for the subsequent processing stage.

[0079] It should be noted that in the high-frequency road data with high-speed sampling (2000Hz), the amount of data is extremely large. About 120,000 data are generated every minute, about 7.2 million data are generated in one hour, and the amount of data in one day is as high as 172.8 million. Such a large amount of data makes redundancy detection extremely challenging and cannot be processed using a single model. To this end, this method proposes a phased redundancy detection scheme, which realizes effective redundancy detection and data compression by hierarchical processing of data according to different time intervals (minutes, hours, days). The following are the specific implementation steps and method details:

[0080] Step 1.1: Data division and organization

[0081] In this step, we divide the original high-frequency road data into one-minute time periods so that the subsequent redundant detection operation is more targeted. The specific implementation details are as follows:

[0082] Step 1.1.1: Raw data collection

[0083] Since the sampling frequency of high-frequency road data is very high (usually 2000Hz), 2000 sampling points are generated per second. Assuming that each sampling point contains data of multiple dimensions, such as position signals, vibration signals, etc., the total amount of data generated in one minute can reach 120,000 sampling points. This amount of data is huge and continuous, so it needs to be pre-processed and reasonably divided before processing, so as to facilitate subsequent phased analysis and calculation.

[0084] Step 1.1.2: Divide the data into minute intervals

[0085] The collected raw data is divided into an independent data interval per minute. In this step, all data points within each minute are organized into a 60×2000 matrix structure, where the number of rows (60) represents the number of seconds per minute, and the number of columns (2000) represents the number of data points per second. This matrix structure can intuitively represent the data distribution per minute as a two-dimensional array, making subsequent feature extraction and analysis more convenient.

[0086] For example, within one minute, assuming that the data source is a vibration sensor with a sampling frequency of 2000 Hz, the data of this sensor for one minute can be represented as the following matrix:

[0087]

[0088] Where D minute Represents the data matrix for that minute, d i,j represents the jth sampling point at the i-th second.

[0089] Step 1.1.3: Data standardization

[0090] In order to ensure the consistency of data features and the training effect of the model, the divided data intervals need to be standardized. Standardization can eliminate the dimensional differences of the data and make the data distribution concentrated in a fixed range, which helps the model learn the characteristics of the data. The standardization formula is as follows:

[0091] d' i,j =(d i,j -μ) / σ

[0092] in:

[0093] d i,j is the original data point,

[0094] μ is the average value in this interval,

[0095] σ is the standard deviation within this interval,

[0096] d' i,j is the standardized data.

[0097] After standardization, the original data can be converted into a standard normal distribution with a mean of 0 and a variance of 1, thereby improving the sensitivity of the convolutional neural network model to data features.

[0098] Step 1.2: Manual annotation redundancy

[0099] In order to train a redundancy detection model with accurate judgment capabilities, the redundancy level of each minute of data needs to be manually labeled in the first stage. This step is the basis of the entire solution. Manual labeling provides the label data required for model training. The specific operations are as follows:

[0100] Step 1.2.1: Redundancy level annotation scope definition

[0101] The redundancy degree is divided into continuous values ​​from 0 to 1, and the meaning of the redundant features in different value intervals is defined:

[0102] 0 means no redundancy at all, that is, the data changes significantly and contains a lot of valid information;

[0103] 1 means complete redundancy, that is, the data has almost no changes and is highly repetitive;

[0104] The intermediate values ​​represent different degrees of redundancy. Depending on the degree of redundancy, the characteristics of the data can gradually transition from "significantly changing" to "stable and unchanged", covering a variety of different redundancy levels.

[0105] This definition provides more detailed labeled data for model training, enabling the model to accurately judge data with different levels of redundancy.

[0106] Step 1.2.2: Feature analysis and redundancy labeling

[0107] In the labeling process, human experience is used to analyze the data interval of each minute and determine the degree of redundancy based on the data change pattern, fluctuation amplitude, periodicity and other characteristics. For example:

[0108] If the data fluctuates frequently and shows obvious characteristics (such as periodic changes or mutations) within a minute, the redundancy level of that minute is marked as low;

[0109] If the data within a minute is stable, highly repetitive, and lacks significant features, the redundancy level of that minute is marked as high.

[0110] In order to improve the accuracy of annotation, the following formula can be used to quantify the degree of redundancy:

[0111]

[0112] in:

[0113] R represents the redundancy level of the minute.

[0114] d i,j is the data of the jth sampling point in the i-th second,

[0115] μ i is the mean value at the i-th second,

[0116] Var(D minute ) is the overall variance of the minute data matrix. This quantitative formula can be used to assist manual labeling and determine the degree of redundancy.

[0117] Step 1.2.3: Annotation tools and process design

[0118] In order to improve the efficiency and consistency of annotation, a dedicated annotation tool is designed to display the data every minute in a graphical manner and provide an interface for manually inputting the degree of redundancy. The operation process is as follows:

[0119] Step 1.2.3.1: Data display: Display the data every minute in the form of a waveform chart or heat map to facilitate observation of data change trends;

[0120] Step 1.2.3.2: Redundancy labeling: The operator can enter the corresponding redundancy value below each minute of data;

[0121] Step 1.2.3.3: Auxiliary calculation: The system automatically calculates the variance, peak change and other statistics of the minute data as an auxiliary reference.

[0122] The development of annotation tools can effectively reduce the workload of manual annotation and enhance the accuracy of annotation through visualization.

[0123] Step 1.2.4: Construction of labeled dataset

[0124] All manually annotated minute data will form a complete training data set for subsequent model training. Each data record contains:

[0125] Input data: 60×2000 data matrix of the minute;

[0126] Label data: the redundancy value annotated by humans.

[0127] The complete annotated dataset ensures the data quality of model training and can reflect the redundant characteristics of high-frequency road data in different scenarios and time periods, providing a guarantee for the generalization performance of the model.

[0128] Step 1.3: Minute-level data redundancy detection model training based on convolutional neural network

[0129] In order to accurately identify the redundancy of high-frequency data at the minute level, we use convolutional neural networks (CNNs) to train the labeled minute-level data in the first stage to generate a model that can identify the redundancy of one-minute data intervals. By training the model on a large number of samples of one-minute data intervals, the convolutional neural network can effectively learn the redundant features in the data, thereby achieving automated and efficient redundancy detection.

[0130] Step 1.3.1: Design of Convolutional Neural Network Model

[0131] In order to effectively detect the degree of data redundancy, a convolutional neural network model containing multiple layers of convolution, pooling and fully connected layers is constructed. The specific design is as follows:

[0132] Step 1.3.1.1: Input layer

[0133] The input layer receives the preprocessed data matrix X minutes, the shape is 60×2000×1, where "1" represents a single channel (i.e. grayscale feature).

[0134] Step 1.3.1.2: Convolutional layer and activation function

[0135] Multiple convolutional layers are used to extract local features of the time series, and the convolution kernel size is designed to be 3×3 or 5×5 to capture local correlation.

[0136] The depth of the output feature map of the first convolutional layer is set to 32, and more feature maps are gradually expanded to 64 and 128 layers to improve the model recognition ability.

[0137] The activation function uses ReLU (Rect ifi ed L i near Un it), which is defined as:

[0138] f(x)=max(0,x)

[0139] The nonlinear characteristics of ReLU can enhance the expressiveness of the model.

[0140] Step 1.3.1.3: Pooling layer

[0141] Use the maximum pooling layer (Max Pooling) for downsampling, and the pooling window size is usually set to 2×2 or 3×3. The role of the pooling layer is:

[0142] Reduce the dimension of the feature map, thereby reducing the amount of computation;

[0143] This makes the features more spatially invariant, thereby improving the generalization performance of the model.

[0144] Step 1.3.1.4: Stacking Deep Convolution and Pooling

[0145] By stacking multiple layers of convolution-pooling structures, higher-level features are gradually extracted, thereby enhancing the model's ability to understand data redundancy patterns. Each layer of convolution operation is followed by a pooling layer to reduce data dimensions and feature redundancy layer by layer.

[0146] Step 1.3.1.5: Fully connected layer

[0147] After the convolution feature extraction is completed, a fully connected layer is added to integrate all features and output the redundant probability. The function of the fully connected layer is to map high-dimensional features to low-dimensional space and aggregate them layer by layer into the final redundant label output.

[0148] Step 1.3.1.6: Output layer and activation function

[0149] The output layer uses one node to represent the predicted value of the redundancy of the one-minute interval data. The activation function uses the Sigma ID function to limit the output value to [0,1] for easy comparison with the actual label. Sigma ID is defined as:

[0150]

[0151] Step 1.3.2: Model training

[0152] Step 1.3.2.1: Loss Function

[0153] The mean squared error (MSE) is selected as the loss function to measure the difference between the model prediction value and the actual label. The loss function formula is:

[0154]

[0155] in:

[0156] N represents the number of training samples;

[0157] y" minute,j is the prediction redundancy of the j-th one-minute data interval;

[0158] y minute,j is the true redundant label of the jth sample.

[0159] Step 1.3.2.2: Optimize the algorithm

[0160] The stochastic gradient descent (SGD) or Adam optimization algorithm is used to adjust the parameters of the model so that the loss function gradually converges to the minimum value.

[0161] SGD: A single-sample update strategy based on gradient descent, suitable for large data sets;

[0162] Adam: An adaptive learning rate algorithm that can automatically adjust the learning rate, with stable convergence effect and fast speed.

[0163] Through the training and tuning of the above steps, the convolutional neural network will be able to accurately identify the redundancy level of a one-minute data length interval, providing an intelligent and automated solution for minute-level redundancy detection of high-frequency road data. The output of this model can be directly used for hour-level data redundancy detection in the next stage, realizing data layer-by-layer processing and reducing the redundancy rate of data storage and calculation.

[0164] Step 1.4: Use the trained convolutional neural network model to predict the redundancy of minute-length intervals

[0165] After the convolutional neural network model is successfully trained and optimized, it can be applied to minute-level redundancy detection of high-frequency data on actual roads. The goal of this step is to deploy the trained model on all collected high-frequency data and automatically predict the redundancy level of each one-minute data interval through the model's reasoning ability.

[0166] Step 1.4.1: Preprocess all data

[0167] According to the data preprocessing method in step 1.1, all collected high-frequency data are processed. Each minute data interval is divided into a 60×2000 matrix format, and each matrix is ​​normalized and standardized to ensure that the data is consistent with the format during training.

[0168] Step 1.4.2: Minute-level data segmentation

[0169] All collected high-frequency road data are divided into one-minute intervals. Assume that the amount of raw data collected throughout the day is D total , then the number of minute-level intervals N minutes It can be calculated as:

[0170]

[0171] Step 1.4.3: Batch input and model inference to predict the redundancy of data per minute

[0172] The data matrix of each one-minute interval is input into the convolutional neural network model for batch inference. Assume that the matrix of a certain interval is represented by X minute,j , then the model predicts the redundancy level of this interval y" minute,j It can be expressed as:

[0173] y" minute,j =CNN(X minute,j )

[0174] Among them, CNN represents the trained convolutional neural network model. By traversing N minutes intervals, the model can predict the redundancy level of all one-minute intervals.

[0175] Step 1.4.4: Storage and Recording

[0176] The redundancy prediction results of each interval are saved in the result file or database for subsequent processing and analysis. At the same time, the storage format should include the following fields:

[0177] Timestamp: identifies the start time of each minute interval;

[0178] Redundancy level: the probability of redundancy predicted by the model;

[0179] Interval marking: a label indicating whether it is redundant, used for subsequent data screening and downsampling.

[0180] By using the convolutional neural network model to comprehensively predict the minute-length data interval, the redundancy in the data can be effectively identified and marked. This process not only realizes the automation of minute-level redundancy detection, but also provides a reliable minute-level redundant data foundation for hour-level and day-level redundancy detection, which is conducive to further optimization and downsampling of data.

[0181] Step 2: Hourly Redundancy Detection

[0182] Based on the minute-level data redundancy detection, artificial experience is used to label the redundancy of the hourly data interval as the training data for the second stage. Based on these labeled data, a fully connected neural network (FCNN) model is trained to enable it to detect and identify data redundancy in the hourly interval. When the system needs to perform redundancy detection in the hourly time interval, the model at this stage can meet the needs and help identify redundant information in long-term data.

[0183] It should be noted that the redundancy of the minute-by-minute data in the first stage is used to further analyze and annotate the redundancy of the hourly data. The redundancy detection in this stage uses a fully connected neural network (FCNN) to process the redundancy of the hourly data length, and finally forms a model that can identify the redundancy of the hourly interval. The detailed steps are as follows:

[0184] Step 2.1: Redundancy labeling of hourly data

[0185] Step 2.1.1: Division and construction of hourly data intervals

[0186] According to the minute-level redundancy detection results of the first phase, the integrated hourly data interval is divided into a data set containing 60 one-minute intervals. The data matrix of the hourly interval is defined as:

[0187]

[0188] Among them, R i Represents the redundancy level label of the i-th minute, which has been generated by manual annotation or convolutional neural network model prediction in the first stage.

[0189] Step 2.1.2: Hourly redundant annotation strategy

[0190] The redundancy level of each hour is based on the aggregate analysis of the minute redundancy level and is labeled in combination with manual experience. The specific labeling rules are as follows:

[0191] If the redundancy level is low in most minutes within the hour and the data fluctuates significantly, the redundancy level of the hour interval is also low and marked as a value close to 0;

[0192] If the data of most minutes in the hour show a stable, unchanged or repeated pattern, the redundancy level of the hour interval is defined as high, marked as a value close to 1;

[0193] When the redundancy level is uneven and the minute data is highly volatile, it is necessary to consider the frequent changes in data and set a medium or above-medium redundancy label for the hourly interval based on professional judgment.

[0194] To ensure the accuracy and consistency of the annotations, the hourly redundant labels can be quantified and calculated as the comprehensive result of the mean and variance of the minute-level labels, as follows:

[0195]

[0196] in:

[0197] α and β are weight coefficients, representing the influence weight of the mean and variance on the redundancy degree. Var(R 1 ,R 2 ,...,R 60 ) represents the variance of the minute redundancy in the hour interval.

[0198] This formula can quantify the volatility and repeatability of hourly data, and combined with manual judgment, it can further determine the labels to form an hourly dataset suitable for training.

[0199] Step 2.1.3: Annotation aids

[0200] A dedicated annotation auxiliary tool is used to display the minute redundancy distribution diagram of hourly data for annotation personnel to observe in detail. The auxiliary tool will display the redundancy distribution and variance of 60 minutes within each hour, and provide an annotation interface to input the redundancy level.

[0201] Step 2.2: Hourly redundancy detection model training based on fully connected neural network

[0202] Step 2.2.1: Build a fully connected neural network model

[0203] In order to effectively identify and judge the redundancy of hourly data, a fully connected neural network (FCNN) is used for training. The specific model architecture includes the following parts:

[0204] Step 2.2.1.1: Input layer: The input data dimension is 1×60, which means that it contains 60 minute-level redundancy labels per hour;

[0205] Step 2.2.1.2: Hidden layer: The model uses multiple layers of fully connected hidden layers, each layer contains several nodes, and nonlinear activation functions (such as ReLU) are used to improve the model's ability to recognize complex data features;

[0206] Step 2.2.1.3: Output layer: The output is a redundancy value ranging from 0 to 1, which is used to represent the overall redundancy level of the hourly data interval.

[0207] Step 2.2.2: Training data preparation and training process

[0208] The model is trained using the hourly dataset labeled in step 2.1. The training data contains 60-dimensional input (minute-level redundancy labels) and 1-dimensional output (hour-level redundancy labels). The model calculates the error using the following loss function:

[0209]

[0210] in:

[0211] L represents the training error,

[0212] N is the number of training samples,

[0213] R" hour,j is the prediction redundancy of the jth sample,

[0214] R hour,j is the actual redundant label of the jth sample.

[0215] Use back-propagation and gradient descent algorithms to adjust model parameters, minimize training errors, and ensure that the model can accurately identify redundancy in hourly data.

[0216] Step 2.2.3: Model validation and evaluation

[0217] After model training, an independent validation set is used to evaluate model performance. The main evaluation indicators include:

[0218] Mean Square Error (MSE): used to measure how close the model prediction value is to the true label;

[0219] Accuracy: The accuracy of judging the redundancy level of the classification model.

[0220] Step 2.2.4: Model tuning and optimization

[0221] During the training process, you can try a variety of optimization methods, such as:

[0222] Adjust the number of hidden layer nodes and layers to optimize model complexity;

[0223] Regularization, to avoid model overfitting and improve generalization ability;

[0224] Learning rate optimization can speed up model convergence.

[0225] Through the above operations, a fully connected neural network model that can accurately identify hourly redundancy is trained. If the final redundancy detection only requires hourly intervals, the second stage can meet the requirements.

[0226] Step 2.3: Use the trained fully connected neural network model to predict the redundancy of the hour-length of all road high-frequency data

[0227] After completing minute-level redundancy detection, we can then apply the hour-level fully connected neural network model to predict the redundancy level of each hour’s data. The goal of this step is to obtain redundant information for a larger time interval by performing hour-by-hour detection.

[0228] Step 2.3.1: Collect minute-level redundant prediction results

[0229] From the minute-level redundant prediction result set obtained in step 1.4, extract the redundant prediction results of 60 one-minute intervals contained in each hour. Assume that the minute-level redundant result set contained in the hth hour is

[0230] Y" minutes,h ={y” minute,(h-1)×60+1 ,y” minute,(h-1)×60+2 ,...,y” minute,h×60}

[0231] Step 2.3.1.1: Construct hour-level model input

[0232] The redundant result set Y" at the minute level for each hour minutes,h Construct a 1×60 input vector X hour,j , used as input to the fully connected neural network model.

[0233] Step 2.3.1.2: Predict using hourly fully connected neural network model

[0234] The constructed hourly input vector X hour,h Input the fully connected neural network model and output the predicted value y" of the redundancy level of the hourly data interval hour,h , expressed as:

[0235] y" hour,h =FCNN 第二阶段 (X hour,h )

[0236] Among them, FCNN 第二阶段 Represents a trained fully connected neural network model.

[0237] Step 2.3.1.3: Save hourly redundant prediction results

[0238] The hourly redundancy prediction results are saved in the database, including the time interval, redundancy value and mark, for subsequent daily redundancy analysis.

[0239] Step 2.3.1.4: Analysis and screening

[0240] By analyzing the hourly redundancy prediction results, the data that needs further processing can be determined, and the hourly redundancy threshold can be set to filter out hourly intervals with high redundancy, which facilitates subsequent data downsampling and compression processing.

[0241] Step 3: Daily level redundancy detection

[0242] If the redundancy detection requirement for high-frequency road data is a daily time interval, the third stage is entered. In this stage, the daily data redundancy level is evaluated based on the previous hourly redundancy detection results and manual labeling experience. Then, the fully connected neural network model is trained using these labeled data to enable it to identify the redundancy level of the daily interval, providing a basis for daily data compression and analysis.

[0243] It should be noted that the hourly data results are further integrated to evaluate and annotate the redundancy of the daily data. The redundancy detection at this stage also relies on the fully connected neural network for processing.

[0244] Step 3.1: Redundancy labeling of daily data

[0245] Step 3.1.1: Division and construction of daily data

[0246] The high-frequency data of the whole day is divided into 24-hour intervals, and the redundancy level label of each hour is generated by the second-stage model or manual annotation. The daily data matrix is ​​represented as follows:

[0247]

[0248] Where R hour,i is the redundancy level label of the i-th hour.

[0249] Step 3.1.2 Day-level redundant annotation strategy

[0250] When marking the degree of daily redundancy, you can refer to the hourly redundancy distribution and variance:

[0251] If the intervals with low redundancy are the majority within 24 hours, the redundancy of the data for the whole day is marked as a low value (close to 0);

[0252] If the intervals with high redundancy are the majority within 24 hours, the redundancy of the data for the whole day is marked as a high value (close to 1);

[0253] If the redundancy distribution is uneven and fluctuates greatly, medium or high redundancy labels are marked based on actual application requirements and manual experience.

[0254] The quantification formula is:

[0255]

[0256] Among them, γ and δ are control parameters, which balance the influence of the average value and variance of the redundancy degree throughout the day on the overall redundancy.

[0257] Step 3.2: Training of a day-level redundancy detection model based on a fully connected neural network

[0258] Step 3.2.1: Model structure

[0259] The input layer contains 24 hour-level redundancy labels, the hidden layer includes several neuron layers, and the output layer provides a redundancy value (0 to 1).

[0260] Step 3.2.1: Training data preparation and training process

[0261] Use the labeled dataset in step 3.1 for training, minimize the loss function by the mean square error, and adjust the parameters to improve the prediction accuracy.

[0262] Step 3.2.1: Model evaluation and optimization

[0263] Use the daily validation set for evaluation and optimize based on error rate and accuracy.

[0264] Step 3.3: Use the trained fully connected neural network model to predict the redundancy of the day length of all road high-frequency data

[0265] After completing the hourly redundancy detection, the hourly data redundancy prediction results are used as the input for the daily redundancy detection to further detect and mark the data redundancy throughout the day.

[0266] Step 3.3.1: Collect hourly redundant prediction results

[0267] From the hourly redundancy detection results in step 2.3, extract the redundant prediction values ​​for the 24-hour interval of the whole day. Assume that the hourly redundancy result set for one day is

[0268] Y" hours,d =[y” hour,1 ,y” hour,2 ,...,y” hour,24}

[0269] Step 3.3.2: Constructing the day-level model input

[0270] The redundant hourly prediction results for the whole day are set Y" hours,d Organized as a 1×24 vector X day,j , and input it into the fully connected neural network model.

[0271] Step 3.3.3: Predict using a sky-level fully connected neural network model

[0272] Enter the day level into vector X day,j Input into the fully connected neural network model to obtain the redundancy prediction value y" of the whole day data interval day,j , expressed as:

[0273] y" day,j =FCNN 第三阶段 (X day,j )

[0274] Among them, FCNN 第三阶段 is a trained fully connected neural network model.

[0275] Step 3.3.4: Storage and analysis

[0276] The prediction results of the redundancy level for the whole day are saved, and redundant data analysis is performed at the whole day level. This result can be used for further decision-making, such as data downsampling and storage optimization for the whole day, or for marking key data intervals for the whole day.

[0277] Compression Processing and Application of Redundant Detection Results

[0278] Data compression strategy

[0279] According to the detected redundancy level, the data can be downsampled and compressed in a targeted manner.

[0280] The specific operations are as follows:

[0281] Low redundancy interval: maintain the original sampling frequency to preserve data integrity;

[0282] High redundancy interval: According to the redundancy level, the sampling frequency is appropriately reduced to reduce the amount of data. For example, for a data interval with a redundancy level of R, the downsampling frequency can be set to:

[0283] f new =f original ×(1-R)

[0284] f original is the original sampling frequency.

[0285] This downsampling method can not only effectively compress data, but also ensure that key data information is retained.

[0286] Practical Application of Redundancy Detection

[0287] Through the results of redundancy detection in different time intervals, data can be stored and accessed in a hierarchical manner according to application requirements. The data after redundancy detection is more representative in actual application scenarios such as intelligent transportation and road safety monitoring.

[0288] This method uses a hierarchical model architecture to implement multi-scale data redundancy detection and compression processing solutions based on minute, hour and day level redundancy detection. Combining the advantages of convolutional neural networks and fully connected neural networks, it not only improves detection accuracy, but also optimizes data storage efficiency, providing an effective solution for high-frequency road data analysis.

[0289] Through the above technical solution, the technical effects produced in the embodiments of the present application are:

[0290] Through the neural network model in the above three stages, the data redundancy degree in different time intervals can be detected. According to the detection results, the data can be compressed at different downsampling frequencies to ensure the retention of key change information and improve the validity and representativeness of the data. This method achieves efficient data compression while retaining the core information of road data to the maximum extent, providing a more optimized data foundation for high-frequency data analysis in practical applications.

[0291] This method of road high-frequency data redundancy detection based on phased and neural networks has the following technical advantages:

[0292] 1. High data processing efficiency

[0293] Since the sampling frequency of high-frequency road data is very high, a single sensor generates a large amount of data every minute. By adopting a layer-by-layer phased detection method, the redundancy level is gradually identified from minutes, hours to days, so that the computational burden of data redundancy detection is dispersed to different stages instead of being concentrated on one-time processing. This step-by-step screening process significantly improves the efficiency of data processing and avoids the processing bottleneck caused by excessive overall data volume.

[0294] 2. High accuracy of redundant detection

[0295] The use of convolutional neural network and fully connected neural network models in stages can capture the characteristics of each time scale. For example, the convolutional neural network extracts important features from minute-level data through convolutional layers and pooling layers to accurately detect redundancy within minutes; the fully connected neural network provides a more macro redundancy assessment on hourly and daily data by integrating the input minute and hourly redundant information. Therefore, the dedicated model at each stage ensures the accuracy of redundancy detection at each time scale, further improving the accuracy of the entire method.

[0296] 3. Strong scalability

[0297] The method is based on staged processing, so when new sensor types are added or longer detection time is required, new models can be simply added or the stages can be further divided. For example, redundant detection with longer periods (such as weeks or months) can be introduced to achieve analysis on a larger time scale. Such a method is highly scalable and can meet the needs of different application scenarios.

[0298] 4. Flexible data compression strategy

[0299] By detecting the degree of redundancy at different time scales, this method can achieve targeted data compression. For example, for minute-level or hour-level data with high redundancy, a lower sampling rate can be used for downsampling, thereby reducing data storage space. For key data fragments, a high sampling frequency can be maintained to ensure data details. This redundancy detection-guided compression method avoids blind compression and ensures that the data meets storage efficiency while retaining important information in practical applications.

[0300] 5. Reduce manual annotation costs

[0301] In the initial stage, only some minute-level and hour-level data redundancy needs to be manually annotated, and the rest of the redundant information can be automatically predicted and generated by the model. This redundancy detection model trained on a small amount of data can be expanded to a large amount of minute-level, hour-level and day-level data, significantly reducing the cost of manual annotation and making the system more efficient.

[0302] 6. Great potential for real-time applications

[0303] With the continuous advancement of computing hardware, this method can be further optimized into a real-time redundant detection solution. Since the phased model can be executed separately, the prediction time of each stage can be minimized. The system can perform minute-level redundant detection while collecting data, and then gradually expand to hourly and daily levels as needed. This design has great potential in real-time application scenarios (such as transportation systems, vehicle-mounted systems, etc.).

[0304] 7. High degree of automation and intelligence

[0305] With the help of convolutional neural networks and fully connected neural network models, the entire redundancy detection process greatly reduces human intervention. The system can automatically determine the degree of redundancy in different time intervals, significantly improving the automation and intelligence of the detection process. Automation and intelligence are particularly important in large-scale data processing, which can reduce the burden on data scientists and engineers and make the method more practical.

[0306] 8. Wide applicability

[0307] This method is not only applicable to high-frequency road data, but can also be applied to other high-frequency data scenarios (such as industrial control data, network traffic monitoring, etc.), through a similar multi-layer neural network structure and phased detection strategy to achieve redundancy identification. Therefore, this method has application potential and versatility in various high-frequency data analysis and processing scenarios.

[0308] In summary, this method combines phased detection, deep learning models and data compression strategies, which not only effectively improves data processing efficiency and accuracy, but also reduces costs and improves the level of intelligence, and has significant technical advantages in practical applications.

[0309] In some embodiments of the present application, the above technical solution is adopted, wherein a system for detecting redundancy of high-frequency road data is disclosed, comprising:

[0310] A minute-level redundancy detection module processes the collected data, divides the data into time intervals in units of minutes, evaluates the redundancy of the data in each minute based on manual annotation, and uses the annotated data to train a convolutional neural network model so that the trained model can automatically identify the redundancy in the data in each minute;

[0311] An hourly redundancy detection module, which uses manual experience to mark the redundancy of each hour's data interval based on the minute-level redundancy detection module, and uses a fully connected neural network model for training to form a model that can identify redundancy in each hour's interval;

[0312] A daily redundancy detection module evaluates the daily data redundancy based on the hourly redundancy detection results and manual labeling experience, and uses these labeled data to train a fully connected neural network model to identify the redundancy level of the daily interval, providing a basis for daily data compression and analysis.

[0313] In the description of the present application, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.

[0314] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.

[0315] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0316] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0317] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting redundancy of high-frequency road data, characterized in that: The following steps are involved: Step 1) Minute-level redundancy detection: pre-process the collected road data and divide the data into time intervals per minute. The redundancy of the data in each minute is evaluated based on manual annotation, and then the annotated data is used to train the convolutional neural network model; Step 2) Hourly redundancy detection: Based on minute-level data redundancy detection, redundancy annotation of hourly data intervals is performed using manual experience as training data. Based on these annotated data, a fully connected neural network model is used for training; Step 3) Daily redundancy detection: Based on the previous hourly redundancy detection results and manual labeling experience, the daily data redundancy level is evaluated. Then, these labeled data are used to train a fully connected neural network model, which enables it to identify the redundancy level of the daily interval and provide a basis for daily data compression and analysis. Through the neural network model in the above steps, the degree of data redundancy in different time intervals is detected. According to the detection results, different downsampling frequencies are selected for compression processing to retain key change information and improve the validity and representativeness of the data.

2. The method for detecting high-frequency road data redundancy according to claim 1, characterized in that: The step 1 includes: Step 1.1) Data division and organization: divide the original high-frequency road data into one-minute time periods; Step 1.2) Manually label the redundancy level, and manually label the redundancy level of each minute of data; Step 1.3) Training of a minute-level data redundancy detection model based on a convolutional neural network: using a convolutional neural network to train the labeled minute-level data, thereby generating a model for identifying the redundancy of a one-minute data length interval, and training the model on a large number of samples of one-minute data intervals; Step 1.4) Use the trained convolutional neural network model to predict the redundancy level of minute-length intervals.

3. The method for detecting high-frequency road data redundancy according to claim 2, characterized in that: The step 1.1 includes: Step 1.1.1) Raw data collection; Step 1.1.2) Data is divided into minute intervals, and the collected raw data is divided into an independent data interval per minute; Step 1.1.3) Data standardization: standardize the divided data intervals so that the data distribution is concentrated within a fixed range. The standardization formula is as follows: d' i,j =(d i,j -m) / s in: d i,j is the original data point, μ is the mean value in the interval, σ is the standard deviation in the interval, d' i,j is the standardized data; After standardization, the original data is transformed into a standard normal distribution with a mean of 0 and a variance of 1.

4. The method for detecting high-frequency road data redundancy according to claim 2, characterized in that: The step 1.2 includes: Step 1.2.1) Define the redundancy level annotation range, divide the redundancy level into continuous values ​​from 0 to 1, and define the meaning of redundancy features in different value intervals: 0 means no redundancy at all, that is, the data changes significantly and contains a lot of valid information; 1 means complete redundancy, that is, the data has almost no changes and is highly repetitive; Intermediate values ​​indicate varying degrees of redundancy; Step 1.2.2) Feature analysis and redundancy degree labeling. During the labeling process, manual experience is used to analyze the data interval of each minute and determine the redundancy degree based on the data change pattern, fluctuation amplitude, and periodic characteristics; Step 1.2.3) Design of annotation tools and processes; Step 1.2.4) Construction of the labeled data set. All manually labeled minute data will form a complete training data set for subsequent model training.

5. The method for detecting high-frequency road data redundancy according to claim 2, characterized in that: The step 1.3 includes: Step 1.3.1) Design of convolutional neural network model: build a convolutional neural network model with multiple layers of convolution, pooling and fully connected layers; Step 1.3.2) Model training.

6. A method for detecting high-frequency road data redundancy according to claim 5, characterized in that: The design of the convolutional neural network model in step 1.3.1 includes: Step 1.3.1.1) Input layer, the input layer receives the preprocessed data matrix X minutes , the shape is 60×2000×1, where "1" indicates a single channel; Step 1.3.1.2) Convolutional layers and activation functions, use multiple convolutional layers to extract local features of the time series, the convolution kernel size is designed to be 3×3 or 5×5, the depth of the first convolutional layer output feature map is set to 32, and more feature maps are gradually expanded to 64 and 128 layers; The activation function uses ReLU (Rectified Linear Unit), which is defined as: f(x)=max(0,x) The nonlinear characteristics of ReLU can enhance the expressiveness of the model; Step 1.3.1.3) Pooling layer, use the maximum pooling layer for downsampling, and the pooling window size is set to 2×2 or 3×3; Step 1.3.1.4) Fully connected layer, after the convolution feature extraction is completed, a fully connected layer is added to integrate all features and output the redundancy probability; Step 1.3.1.5) Output layer and activation function. The output layer uses one node to represent the predicted value of the redundancy of the one-minute interval data. The activation function uses the Sigmoid function to limit the output value to [0,1] to facilitate comparison with the actual label. Sigmoid is defined as:

7. The method for detecting high-frequency road data redundancy according to claim 5, characterized in that: The model training in step 1.3.2 includes: Step 1.3.2.1) Loss function, select mean square error as the loss function to measure the difference between the model prediction value and the actual label. The loss function formula is: Where: N represents the number of training samples; y" minute,j is the prediction redundancy of the jth one-minute data interval; y minute,j is the true redundant label of the jth sample; Step 1.3.2.2) Optimize the algorithm and use the stochastic gradient descent or Adam optimization algorithm to adjust the parameters of the model so that the loss function gradually converges to the minimum value.

8. The method for detecting high-frequency road data redundancy according to claim 1, characterized in that: The step 2 includes: Step 2.1) Redundancy level labeling of hourly data Step 2.2) Hourly redundancy detection model training based on fully connected neural network Step 2.3) Use the trained fully connected neural network model to predict the redundancy level of all road high-frequency data in hours.

9. The method for detecting high-frequency road data redundancy according to claim 1, characterized in that: The step 3 includes: Step 3.1) Redundancy level labeling of daily data Step 3.2) Training of the day-level redundancy detection model based on a fully connected neural network Step 3.3) Use the trained fully connected neural network model to predict the redundancy level of all road high-frequency data in terms of length.

10. A system for detecting redundancy of high-frequency road data, using a method for detecting redundancy of high-frequency road data as described in any one of claims 1 to 9, characterized in that: include: A minute-level redundancy detection module processes the collected data, divides the data into time intervals in units of minutes, evaluates the redundancy of the data in each minute based on manual annotation, and uses the annotated data to train a convolutional neural network model so that the trained model can automatically identify the redundancy in the data in each minute; An hourly redundancy detection module, which uses manual experience to mark the redundancy of hourly data intervals based on the minute-level redundancy detection module, and uses a fully connected neural network model for training to form a model that can identify redundancy in hourly intervals; A daily redundancy detection module evaluates the daily data redundancy based on the hourly redundancy detection results and manual labeling experience, and uses these labeled data to train a fully connected neural network model to identify the redundancy level of the daily interval, providing a basis for daily data compression and analysis.