Low-cost storage and data reconstruction method for traffic time series based on feature clustering

By using time-frequency decomposition and supervised clustering of traffic time series data, combined with road topology information, a feature semantic database is constructed, achieving low-cost storage and accurate data reconstruction. This solves the problems of high cost and data loss in traditional storage methods, and supports the development of intelligent traffic management systems and data prediction.

CN118965029BActive Publication Date: 2026-02-06XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411042753.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2026-02-06
Estimated Expiration
2044-07-31

AI Technical Summary

Technical Problem

Existing traffic time series data storage methods cannot meet the low-cost storage needs of large-scale data, and traditional compression methods are difficult to balance compression ratio and reconstruction accuracy, resulting in high storage costs and difficulty in accurate reconstruction when data is lost or damaged.

Method used

By performing time-frequency decomposition on segmented traffic time series, and inputting the road topology directed graph into a traffic sequence compression and reconstruction deep model, supervised clustering and feature extraction are performed to build a feature semantic library. The loss function is then used to adjust the classification parameters to achieve low-cost storage and data reconstruction.

Benefits of technology

It reduces the storage space requirements for large-scale traffic flow data, reduces storage costs, and enables accurate data reconstruction when data is damaged or lost, supporting the development of intelligent traffic management systems and data prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118965029B_ABST
    Figure CN118965029B_ABST
Patent Text Reader

Abstract

The application provides a traffic time series low-cost storage and data reconstruction method based on feature clustering, which comprises the following steps: obtaining n1-dimensional frequency sequences by performing time-frequency decomposition on segmented traffic time series, and inputting the n1-dimensional frequency sequences and a road topology directed graph into a constructed traffic sequence compression and reconstruction deep model to perform a training process; in the training process, features of the input objects are extracted first, and then supervised clustering is performed to obtain n2-dimensional feature class sets; a loss function is designed to repeatedly adjust the classification parameters of the supervised clustering until a training cutoff condition is reached, so as to obtain a traffic time series feature database; according to the time sequence time stamp and the corresponding road network in the reconstruction requirement, the initial value of the sequence, the n2-dimensional feature class set and the corresponding road topology directed graph of the road network are extracted from the traffic time series feature database to perform reconstruction, so as to obtain a reconstructed traffic flow time sequence. The application reduces the space required for storing large-scale traffic flow data, thereby reducing the storage cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of intelligent transportation and artificial intelligence, and specifically relates to a low-cost storage and data reconstruction method for traffic time series based on feature clustering. BACKGROUND

[0002] The intelligent transportation management system relies on massive traffic time series data, which provides rich and important basic data for traffic management, road planning, travel services, and related research and development work. Through in-depth analysis and application of these data, the intelligent transportation management system not only can realize efficient scheduling of road resources, improve road traffic efficiency, but also can enhance road safety, optimize people's travel experience, and ultimately promote the sustainable development of the entire intelligent transportation field. However, as the intelligent transportation system becomes more and more perfect, the scale of traffic time series data also presents an explosive growth, and the traditional data storage method has been unable to meet the big data storage demand, and brings high storage cost. For example, the DAIR-V2X car-road cooperation trajectory prediction data set provides more than 200,000 traffic segments, each segment is 10 seconds, the total time length is about 556 hours of traffic data, and the data set size is more than 30G; the highway network PeMS traffic data set collects data from about 40,000 sensors on the highway in real time, processes about 2GB of time series data every day, and stores more than 10 years of data in total. Therefore, it is necessary to study a low-cost data storage method, and to realize original data reconstruction when needed, so as to more efficiently and economically store and manage traffic time series data.

[0003] Traffic time series data is the basis for the realization of intelligent transportation. Its collection, form and storage method are as follows:

[0004] The collection of traffic time series usually relies on various sensors, monitoring devices and data sources. These devices and sources can capture vehicle and traffic environment related information, providing rich traffic data. The collection of traffic time series data includes, but is not limited to, traffic sensors, GPS devices, traffic signal control systems, weather sensors, vehicle-mounted sensors, and various traffic applications and platforms. Traffic sensors, such as inductive coils, infrared sensors and microwave sensors, are used to detect the presence, number and speed of vehicles; GPS devices provide vehicle location, speed and travel trajectory information for vehicle tracking, travel time measurement and real-time navigation; traffic signal control systems record traffic signal control signals; weather sensors provide meteorological information such as temperature, precipitation, visibility and wind speed for analyzing the impact of weather on traffic; vehicle-mounted sensors include speed sensors, ABS sensors, tire pressure sensors, etc., to provide vehicle status and performance data to support vehicle performance monitoring and health management; traffic applications and platforms, such as traffic navigation applications and intelligent traffic management systems, collect real-time traffic data through users' mobile phones or vehicle-mounted devices to provide real-time traffic information, navigation services and traffic management. These collection methods are usually used in combination to provide comprehensive traffic time series data to support the needs of traffic management, travel services and research applications.

[0005] Existing low-cost storage solutions for traffic data mainly include lossless compression and lossy compression. Lossless compression is a method that can compress data without losing any data, and the compressed data can be completely restored to the original state. Existing lossless compression methods include Huffman encoding, Run-Length encoding, LZW algorithm, etc. Lossy compression is a method that loses part of the data during compression, which can achieve higher compression ratio. Existing lossy compression methods include sequence approximation methods and sequence transformation methods. Sequence approximation methods include low-density sampling, multi-sequence aggregation, line segment representation, etc.; sequence transformation methods include Fourier transform, discrete cosine transform, wavelet transform, etc.

[0006] The above lossless compression methods have the disadvantage of relatively low compression ratio, which is not suitable for compressing a large amount of data. The above lossy compression sequence approximation methods, such as low-density sampling, multi-sequence aggregation, line segment representation, etc., can achieve higher compression ratio, but will lose a large amount of data, which is not conducive to data reuse and accurate analysis. The above lossy compression sequence transformation methods are affected by the time-frequency characteristics of the data itself, the selection of transformation parameters, etc., making it difficult to balance the compression ratio and reconstruction accuracy, and lacking universality and robustness. SUMMARY

[0007] In order to solve the above problems in the prior art, the present application provides a maintenance method and system for train equipment failure. The technical problem to be solved by the present application is solved by the following technical scheme:

[0008] In a first aspect, the application provides a traffic time series low-cost storage and data reconstruction method based on feature clustering, comprising:

[0009] S100, obtaining an n1-dimensional frequency sequence by performing time-frequency decomposition on the segmented traffic time sequence, and inputting the n1-dimensional frequency sequence and a road topology directed graph into a constructed traffic sequence compression and reconstruction deep model;

[0010] S200, performing a training process of the traffic sequence compression and reconstruction deep model, and in the training process, first extracting features of an input object, then performing supervised clustering to obtain an n2-dimensional feature class set, and composing a feature semantic library, repeatedly adjusting classification parameters of the supervised clustering by designing a loss function until a training cutoff condition is reached, and obtaining a traffic time sequence feature database when the training cutoff condition is reached;

[0011] S300, according to a time sequence timestamp in a reconstruction requirement and a corresponding road network, extracting a sequence initial value, the n2-dimensional feature class set, and a road topology directed graph corresponding to the road network from the traffic time sequence feature database, and using the three to reconstruct to obtain a reconstructed traffic flow time sequence.

[0012] In a first aspect, the application provides a traffic time series low-cost storage and data reconstruction device based on feature clustering, comprising:

[0013] An input module configured to obtain an n1-dimensional frequency sequence by performing time-frequency decomposition on the segmented traffic time sequence, and input the n1-dimensional frequency sequence and a road topology directed graph into a constructed traffic sequence compression and reconstruction deep model;

[0014] A training module configured to perform a training process of the traffic sequence compression and reconstruction deep model, and in the training process, first extract features of an input object, then perform supervised clustering to obtain an n2-dimensional feature class set, and compose a feature semantic library, repeatedly adjust classification parameters of the supervised clustering by designing a loss function until a training cutoff condition is reached, and obtain a traffic time sequence feature database when the training cutoff condition is reached;

[0015] A reconstruction module configured to, according to a time sequence timestamp in a reconstruction requirement and a corresponding road network, extract a sequence initial value, the n2-dimensional feature class set, and a road topology directed graph corresponding to the road network from the traffic time sequence feature database, and use the three to reconstruct to obtain a reconstructed traffic flow time sequence.

[0016] Beneficial effects:

[0017] (1) The application reduces the space required for storing large-scale traffic flow data, thereby reducing the storage cost.

[0018] (2) When data is damaged or lost, the method of the present application can be used to reconstruct the damaged or lost data, thereby ensuring the integrity and accuracy of the data.

[0019] (3) The method of the present application provides strong data support for intelligent traffic management systems, thereby promoting the development of intelligent traffic technology, and can also be used for traffic time series prediction, and has a wide application prospect.

[0020] The present application will be further described in detail below in combination with the drawings and examples. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is a flowchart of a traffic time series low-cost storage and data reconstruction method based on feature clustering provided by the present application;

[0022] Figure 2 is a traffic time series feature database input and output diagram provided by the present application;

[0023] Figure 3 is a traffic time series compression and reconstruction deep model training flowchart based on feature clustering provided by the present application;

[0024] Figure 4 is a traffic time series reconstruction flowchart provided by the present application;

[0025] Figure 5 is a prediction result diagram of a predicted traffic time series provided by the present application. DETAILED DESCRIPTION

[0026] The present application will be further described in detail below in combination with the drawings and examples.

[0027] In a first aspect, as shown in the drawings, the present application provides a traffic time series low-cost storage and data reconstruction method based on feature clustering, comprising: Figure 1

[0028] S100, by time-frequency decomposition of segmented traffic time series, n1-dimensional frequency sequence is obtained, and combined with road topology directed graph, input into the constructed traffic sequence compression and reconstruction deep model;

[0029] In this step, the traffic time series is represented by a, and all segmented traffic time series are represented by [a1, a2, … a i, ].

[0030] ​S200, a training process of the traffic sequence compression and reconstruction depth model is performed, and in the training process, features of an input object are first extracted, then supervised clustering is performed to obtain an n2-dimensional feature class set, and a feature semantic library is formed, a classification parameter of the supervised clustering is repeatedly adjusted by designing a loss function until a training stop condition is reached, and a traffic time sequence feature database when the training stop condition is reached is obtained;

[0031] Reference Figure 2 The traffic time sequence feature database includes a road topology directed graph database, a traffic section time sequence class library and a feature semantic library; the road topology directed graph is stored in the road topology directed graph database; the sequence initial value and the corresponding class are stored in the traffic section time sequence class library; n2-dimensional traffic sequence clustering features f x and the corresponding class are stored in the feature semantic library; the traffic section time sequence class library and the feature semantic library have a data index relationship.

[0032] The traffic time sequence class feature library stores a traffic time sequence class feature set wherein refers to a class feature of any dimension, and saves a k i class label and a feature value after clustering. The traffic section time sequence class library stores a timestamp of a section traffic time sequence, a corresponding road and n2-dimensional class label data.

[0033] S300, according to a time sequence timestamp in a reconstruction requirement and a corresponding road network, a sequence initial value, an n2-dimensional feature class set and a road topology directed graph corresponding to the road network are extracted from the traffic time sequence feature database, and the three are used for reconstruction to obtain a reconstructed traffic flow time sequence.

[0034] As an optional embodiment of the application, S100 includes:

[0035] S110, traffic time sequences are divided into equidistant section traffic time sequences, and a road topology directed graph is constructed by using a road grid distance between traffic time sequence collection devices; the section traffic time sequences and the road topology directed graph are both normalized;

[0036] S120, the normalized section traffic time sequences are decomposed into n1-dimensional frequency sequences by a time-frequency transform decomposition method, and the n1-dimensional frequency sequences and the normalized road topology directed graph are jointly input into a traffic sequence compression and reconstruction depth model that has been constructed.

[0037] The time-frequency transform decomposition method of the application includes but is not limited to Fourier transform, wavelet transform, Hilbert transform and various improved time-frequency transforms.

[0038] As an optional embodiment of the present application, S110 comprises:

[0039] S111, collecting traffic time series by using traffic time series collection devices and dividing the traffic time series into equidistant segmented traffic time series; wherein the traffic time series comprises motor vehicle flow, non-motor vehicle flow, traffic flow rate, traffic occupancy rate, meteorological information and vehicle state;

[0040] S112, calculating the road grid distance between the traffic time series collection devices and processing the road grid distance into an adjacency matrix data structure, thereby obtaining a road topology directed graph;

[0041] S113, normalizing the segmented traffic time series and the road topology directed graph by using a minimum-maximum normalization method, thereby obtaining normalized segmented traffic time series and normalized road topology directed graph.

[0042] wherein the length of the equidistant segmented traffic time series is Assuming that the total length of the traffic time series is l a , then any is obtained to achieve a higher compression ratio. The predetermined formula of the minimum-maximum normalization method is: wherein x nor is the normalized value, x is the original value of the sequence, x min is the minimum value of the sequence, and x max is the maximum value of the sequence.

[0043] In combination with Figure 3 and Figure 4 , as an optional embodiment of the present application, S200 comprises: performing a training process of the traffic sequence compression and reconstruction depth model, the training process comprising the following steps:

[0044] a. in the current training round, extracting an n2-dimensional traffic sequence feature vector from the n1-dimensional frequency sequence by using the traffic sequence compression and reconstruction depth model, and then clustering the n2-dimensional traffic sequence feature vector to obtain multi-class features of each dimension in the n2-dimensional traffic sequence feature vector; combining the multi-class features of each dimension in the n2-dimensional traffic sequence feature vector, the road topology directed graph and the sequence initial value to form a traffic time sequence feature database; performing feature fusion on the features in the traffic time sequence feature database and the deep features of the road topology directed graph and then reconstructing to obtain a reconstructed segmented traffic time series;

[0045] b. calculating a total loss function by using the reconstructed segmented traffic time series and the segmented traffic time series in S100, and adjusting the weight of the traffic sequence compression and reconstruction depth model in each training round by using the loss function; using the traffic sequence compression and reconstruction depth model with the adjusted weight as the traffic sequence compression and reconstruction depth model in the next training round;

[0046] e, repeating a to b until the traffic sequence compression and reconstruction deep model reaches the training stopping condition, to obtain the traffic time sequence feature database when the training stopping condition is reached.

[0047] The training stopping condition includes reaching the maximum training round or the loss function value no longer decreases.

[0048] Reference Figure 3 The traffic sequence compression and reconstruction deep model of the application adopts an encoder-decoder structure, including a traffic sequence feature extraction and encoding module, a road topology structure directed graph feature extraction module, a feature supervised clustering module, and a sequence reconstruction decoding module.

[0049] The traffic sequence feature extraction and encoding module, the road topology structure directed graph feature extraction module, the feature supervised clustering module, and the sequence reconstruction decoding module are composed of multiple neural network layers, including but not limited to a Transformer layer, a long short-term memory (LSTM) neural network layer, a convolutional neural network (CNN) layer, a fully connected layer (DNN), and various improved neural network layers. The activation functions used by the neural network layers include but are not limited to Rule, Tanh, Softmax, Sigmid, and various improved activation functions.

[0050] The input of the traffic sequence feature extraction and encoding module is the segmented traffic time sequence a i The decomposed n1-dimensional frequency sequence, and the output is an n2-dimensional traffic sequence feature vector f r , wherein n1≥n2;

[0051] The input of the road topology structure directed graph feature extraction module is the road topology directed graph, and the output is the deep feature f g of the road topology directed graph.

[0052] The input of the feature supervised clustering module is the n2-dimensional traffic sequence feature vector f r , and the output is an n2-dimensional traffic sequence clustering feature f x and the category to which it belongs.

[0053] The input of the sequence reconstruction decoding module is at least one sequence initial value of the segmented traffic time sequence a i , the n2-dimensional traffic sequence clustering feature f x , and the deep feature f g , and the output is the reconstructed traffic time sequence a i .

[0054] In combination with Figure 3 and Figure 4 , in an optional embodiment, a includes:

[0055] a1, in the current training round, the feature of the n1-dimensional frequency sequence is extracted by using the traffic sequence feature extraction coding module of the current training round, and an n2-dimensional traffic sequence feature vector f is output r to the feature supervised clustering module of the current training round, and the deep feature f of the normalized road topology directed graph is extracted by using the road topology directed graph feature extraction module of the current training round g , and output to the sequence reconstruction decoding module of the current training round;

[0056] a2, the n2-dimensional traffic sequence feature vector f is classified by using the feature supervised clustering module of the current training round r , so as to output the n2-dimensional traffic sequence clustering feature f x and the category to which it belongs, the n2-dimensional traffic sequence clustering feature f x and the category to which it belongs constitute the feature semantic library of the current training round, and the feature semantic library of the current training round, the sequence initial value in the segmented traffic time sequence a i and the normalized road topology directed graph constitute the traffic time sequence feature database of the current training round;

[0057] a3, the segmented traffic time sequence a i , the normalized road topology directed graph and the n2-dimensional traffic sequence clustering feature f x are input into the sequence reconstruction decoding module of the current training round, so as to output the reconstructed segmented traffic time sequence a

[0058] Reference Figure 3 , in an optional embodiment, b comprises:

[0059] b1, the reconstructed segmented traffic time sequence a and the segmented traffic time sequence a i are used to calculate the first loss value of the current training round;

[0060] , wherein the first loss value is a loss value calculated by the traffic sequence reconstruction loss loss1, the traffic sequence reconstruction loss loss1 is the error between the original traffic flow and the reconstructed traffic flow, and is used for training and evaluation of the traffic sequence feature extraction coding module, the road topology directed graph feature extraction module and the sequence reconstruction decoding module.

[0061] b2, the reconstructed segmented traffic time sequence a and segmented traffic time series a i Both are decomposed by n2-dimensional time-frequency transform, and the sum of errors between each one-dimensional frequency sequence of the decomposed two is calculated.

[0062] b3, taking the sum of errors as the third loss value of the current training round, and calculating the second loss value of the current training round in combination with the first loss value of the current training round;

[0063] Wherein, the second loss value is a loss value calculated by the sum of errors loss3, and loss3 is the sum of errors between each one-dimensional frequency signal after n2-dimensional time-frequency transform decomposition of the reconstructed traffic flow and the original traffic flow.

[0064] b4, adjusting the weights of the sequence reconstruction decoding module and the road topology directed graph feature extraction module of the current training round by using the first loss value of the current training round, and calculating the weights of the feature supervised clustering module of the current training round by using the second loss value of the current training round;

[0065] Wherein, the second loss value is a loss value calculated by the supervised clustering loss loss2, and the supervised clustering loss loss2 = loss1 + loss3, which is used for training and evaluation of the feature supervised clustering module.

[0066] b5, taking the adjusted sequence reconstruction decoding module, road topology directed graph feature extraction module and feature supervised clustering module as the sequence reconstruction decoding module, road topology directed graph feature extraction module and feature supervised clustering module of the next training round.

[0067] In an optional embodiment, the application can also update the road topology directed graph in the road topology directed graph database by adding new vertices, directed edges and distance weights; and set the update threshold of the traffic segmented time series category library, and when the threshold is met, update the sequence initial value and the category in the traffic segmented time series category library.

[0068] It is worth mentioning that: the running mode of the traffic time sequence feature database includes road topology directed graph updating, traffic segment time sequence category updating, and feature semantic data updating. The road topology directed graph updating updates the road topology network by adding vertices, directed edges, and distance weights. The feature semantic data updating refers to generating new traffic time sequence categories and features due to factors such as road structure, urban spatial structure, and traffic environment. A traffic time sequence category feature updating threshold is set, and when the threshold is met, the traffic time sequence category feature is updated. The traffic segment time sequence category data updating extracts features from any newly added real-time traffic segment time sequence and clusters them into n2-dimensional feature categories, and updates the segment time sequence category data.

[0069] In a second aspect, the application provides a traffic time sequence low-cost storage and data reconstruction device based on feature clustering, comprising:

[0070] An input module configured to obtain n1-dimensional frequency sequences by performing time-frequency decomposition on the segmented traffic time sequence, and inputting the n1-dimensional frequency sequences into the constructed traffic sequence compression and reconstruction deep model in combination with the road topology directed graph;

[0071] A training module configured to perform a training process of the traffic sequence compression and reconstruction deep model, and in the training process, first extract features of the input object, then perform supervised clustering to obtain an n2-dimensional feature class set, and form a feature semantic library. The classification parameters of the supervised clustering are repeatedly adjusted by designing a loss function until a training cutoff condition is reached, and a traffic time sequence feature database is obtained when the training cutoff condition is reached;

[0072] A reconstruction module configured to extract sequence initial values, an n2-dimensional feature class set, and a corresponding road topology directed graph of the road network from the traffic time sequence feature database according to the time sequence timestamp and the corresponding road network in the reconstruction requirement, and use the three to reconstruct to obtain a reconstructed traffic flow time sequence.

[0073] The effect of the application will be illustrated by simulation as follows.

[0074] For a 24-hour traffic time sequence, the traffic time sequence is segmented into 24 segments, each with a 1-hour traffic time sequence. The road topology structure directed graph is calculated. All data is standardized to the [0, 1] interval using the minimum maximum value normalization method. The traffic time sequence is one of motor vehicle flow, non-motor vehicle flow, traffic flow rate, traffic occupancy rate, meteorological information, and vehicle state.

[0075] The 1-hour segmented time series is decomposed into a 5-dimensional frequency sequence by a time-frequency transform decomposition method. The traffic sequence feature extraction and encoding module inputs the 5-dimensional frequency sequence decomposed from the 1-hour segmented traffic time series, and outputs a 5-dimensional traffic sequence feature vector f r . The feature supervised clustering module inputs the 5-dimensional traffic sequence feature, and outputs a 5-dimensional traffic sequence clustering feature f x . The sequence reconstruction and decoding module consists of a feature fusion layer, a feature extraction layer, and a multi-layer perception (MLP) output layer. The input is the initial value of the traffic segmented time series at two time steps, the 5-dimensional traffic sequence clustering feature value f x and the road topology directed graph depth feature f g , and the output is the reconstructed traffic sequence The feature semantic library stores a set of traffic time sequence category features X f =[f x1 ,f x2 ,f x3 ,f x4 ,f x5 ]. For the 5-dimensional feature sequence, each 1-dimensional feature sequence is clustered into 10 categories, and the label and its feature value are saved.

[0076] The segmented traffic time sequence category library stores the time stamp, the road to which the segmented time sequence belongs, and the 5-dimensional category label data. For 365 sets of 24-hour traffic time sequence data, the traffic time sequence category feature update threshold is set to 24 times, and when the number of similar features appears 24 times, the traffic time sequence category feature is updated. The segmented traffic time sequence category data is updated, and any newly added 1-hour segmented time sequence is extracted and clustered into a 5-dimensional feature category, and the segmented time sequence category data is updated.

[0077] According to the time sequence time stamp and the road network to which it belongs, the 2-time-step initial value, the 5-dimensional feature category set, and the road topology directed graph stored in the traffic time sequence feature database are extracted. The above data is input into the sequence reconstruction and decoding module, and the output is the reconstructed traffic time sequence.

[0078] The application can recover data for the lost traffic time sequence. For example, for a lost 10-minute traffic time sequence data, the 2-time-step initial value, the 5-dimensional feature category set, and the road topology directed graph within the previous 1 hour from the traffic time sequence feature database are extracted according to the time stamp and the road segment. The above data is input into the sequence reconstruction and decoding module, and the output is the recovered traffic time sequence.

[0079] The application can predict traffic flow time series, and the prediction process comprises the following steps: step 1, collecting traffic flow data in real time and performing standardization processing; step 2, extracting a 5-dimensional feature class set within 1 hour of the road and a road topology directed graph based on a traffic time series feature library; and step 3, inputting the above data into a sequence reconstruction decoding module to output a prediction value of the traffic flow data. Figure 5 The prediction error evaluation index is calculated from the prediction result of the traffic time series data, and the error evaluation index is MAE 1.66, RMSE 3.44, and MAPE 3.75%. Figure 5

[0080] The application provides a low-cost storage and data reconstruction method for traffic time series based on feature clustering by using time series frequency decomposition, automatic encoder, supervised clustering and other methods, and a traffic time series feature database is constructed, so that the original data can be accurately reconstructed while the traffic time series data is compressed and stored.

[0081] It should be noted that the terms "first" and "second" in the application are only used for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined as "first" and "second" can explicitly or implicitly include one or more of the features. In the description of the application, the meaning of "a plurality of" is two or more, unless otherwise specifically limited.

[0082] Although the application is described herein in conjunction with various embodiments, other variations of the disclosed embodiments can be understood and implemented by those skilled in the art by referring to the drawings, disclosure and appended claims in implementing the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality.

[0083] The above is a further detailed description of the application in conjunction with specific preferred embodiments, and the specific implementation of the application cannot be limited to these descriptions. For those skilled in the art to which the application belongs, without departing from the concept of the application, a number of simple deductions or substitutions can be made, which should be regarded as falling within the protection scope of the application.​

Claims

1. A low-cost storage and data reconstruction method for traffic time series based on feature clustering, characterized in that, Comprising: S100, obtaining by time-frequency decomposition on the segmented traffic time series the frequency sequence, and combining the road topology directed graph input into the constructed traffic sequence compression and reconstruction deep model; S100 comprises: S110, dividing the traffic time series into equidistant segmented traffic time series, and constructing a road topology directed graph using the road grid distance between traffic time series collection devices, both the segmented traffic time series and the road topology directed graph are normalized; S120, the normalized segmented traffic time series is decomposed into frequency sequence, and the normalized road topology directed graph is input into the constructed traffic sequence compression and reconstruction deep model. Wherein, the traffic sequence compression and reconstruction deep model comprises: traffic sequence feature extraction coding module, road topology structure directed graph feature extraction module, feature supervised clustering module, sequence reconstruction decoding module; The traffic sequence feature extraction coding module inputs segmented traffic time sequence Decomposed The frequency sequence is output The traffic sequence feature vector Wherein ≥ ; The road topology directed graph feature extraction module inputs the road topology directed graph and outputs a deep feature of the road topology directed graph ; The feature supervised clustering module has an input of dimensional traffic sequence feature vector and an output of dimensional traffic sequence clustering feature and the category to which it belongs; The sequence reconstruction decoding module inputs at least one sequence initial value of segmented traffic time sequence , dimensional traffic sequence clustering feature and deep feature , and outputs reconstructed traffic time sequence ; S200, performing a training process of a traffic sequence compression and reconstruction depth model, and in the training process, first extracting features of an input object, then performing supervised clustering to obtain a feature class set, and composing a feature semantic library, repeatedly adjusting classification parameters of the supervised clustering by designing a loss function until a training stop condition is reached, to obtain a traffic time sequence feature database when the training stop condition is reached; The traffic time series feature database comprises a road topology directed graph database, a traffic segmented time series category library and a feature semantic library; S300, according to the time sequence timestamp in the reconstruction demand and the corresponding road, extracting the sequence initial value from the traffic time sequence feature database, the road topology directed graph corresponding to the road of the feature class set, and inputting the three to the sequence reconstruction decoding module for reconstruction to obtain the reconstructed traffic flow time sequence.

2. The feature clustering based traffic time series low cost storage and data reconstruction method according to claim 1, characterized in that, S110 comprises: S111, collecting traffic time series using traffic time series collection devices, and dividing them into equidistant segmented traffic time series; wherein the traffic time series includes motor vehicle flow, non-motor vehicle flow, traffic flow, traffic occupancy, weather information and vehicle state; S112, calculate the road grid distance between the traffic time series collection devices, and process it into an adjacency matrix data structure, so as to obtain a road topology directed graph; S113, normalize the segmented traffic time series and the road topology directed graph using the minimum and maximum normalization method, to obtain the normalized segmented traffic time series and the normalized road topology directed graph.

3. The feature clustering based traffic time series low cost storage and data reconstruction method of claim 1, wherein, S200 comprises: performing a training process of the traffic sequence compression and reconstruction deep model, the training process comprises the following steps: a. In the current training round, the traffic sequence compression and reconstruction deep model is used to extract the data from the aforementioned... Extracting from 3D frequency sequences The feature vector of the traffic sequence is then clustered to obtain... Multi-class features for each dimension in the dimension; The multi-class features of each dimension in the road topology directed graph and the initial values ​​of the sequence constitute a traffic time series feature database; the features in the traffic time series feature database are fused with the depth features of the road topology directed graph and then reconstructed to obtain the reconstructed segmented traffic time series; b, calculate the total loss function using the reconstructed segmented traffic time series and the segmented traffic time series in S100, and adjust the weight of the traffic sequence compression and reconstruction deep model in each training round using the loss function; the traffic sequence compression and reconstruction deep model with adjusted weight is used as the traffic sequence compression and reconstruction deep model of the next training round; e, repeat a to b until the traffic sequence compression and reconstruction deep model reaches the training cutoff condition, and obtain the traffic time series feature database when the training cutoff condition is reached.

4. The feature clustering based traffic time series low cost storage and data reconstruction method of claim 3, wherein a Comprising: a1, in the current training round, using the traffic sequence feature extraction and coding module of the current training round to extract the traffic sequence feature of the current training round , and output the feature of the frequency sequence of the current training round , and output the traffic sequence feature vector of the current training round to the feature supervised clustering module of the current training round, and using the road topology directed graph feature extraction module of the current training round to extract the deep feature of the normalized road topology directed graph , and output to the sequence reconstruction decoding module of the current training round; a2, the feature of the current training round is supervised clustering module using the current training round of features dimensional traffic sequence feature vector each dimension is classified to output dimensional traffic sequence clustering features and the category to which it belongs, the dimensional traffic sequence clustering features and the category to which it belongs, the feature semantic library of the current training round is composed, and the feature semantic library of the current training round, the sequence initial value in the segmented traffic time sequence and the normalized road topology directed graph constitute the traffic time sequence feature database of the current training round; a3, selecting a segmented traffic time series from the traffic time series feature database of the current training round , a normalized road topology directed graph and a traffic sequence clustering feature input into a sequence reconstruction decoding module of the current training round to output a reconstructed segmented traffic time series under the current training round by using the sequence reconstruction decoding module to first perform feature fusion and then feature extraction .

5. The feature clustering based traffic time series low cost storage and data reconstruction method of claim 4, wherein b Comprising: b1, the segmented traffic time series reconstructed at the current training round and the segmented traffic time series , a first loss value of the current training round is calculated; b2, the segmented travel time series reconstructed under the current training round and the segmented travel time series are both subjected to a multidimensional frequency transform decomposition, and the sum of the errors between each of the decomposed one-dimensional frequency sequences is calculated; b3, taking the sum of the errors as the third loss value of the current training round, and calculating the second loss value of the current training round in combination with the first loss value of the current training round; b4, adjust the weights of the sequence reconstruction decoding module and the road topology structure directed graph feature extraction module of the current training round using the first loss value of the current training round, and calculate the weight of the feature supervised clustering module of the current training round using the second loss value of the current training round; b5, take the adjusted sequence reconstruction decoding module, road topology structure directed graph feature extraction module, and feature supervised clustering module as the sequence reconstruction decoding module, road topology structure directed graph feature extraction module, and feature supervised clustering module of the next training round.

6. The feature clustering based traffic time series low cost storage and data reconstruction method according to claim 4, characterized in that, The road topology directed graph is stored in the road topology directed graph database; and the sequence initial value and the corresponding category are stored in the traffic section time sequence category library. Traffic sequence clustering features The features are stored in the feature semantic library together with the categories to which the features belong. The traffic segmented time series category library and the feature semantic library have a data index relationship.

7. The feature clustering based traffic time series low cost storage and data reconstruction method of claim 6, wherein, The traffic time series low-cost storage and data reconstruction method based on feature clustering further comprises: The road topology directed graph in the road topology directed graph database is updated by adding vertices, directed edges and distance weights; An update threshold of the traffic segment time sequence category library is set, and when the threshold is met, the sequence initial value and the category in the traffic segment time sequence category library are updated.

8. A device for low-cost storage and data reconstruction of traffic time series based on feature clustering, characterized by Comprise: an input module configured to obtain a time-frequency sequence of segmented traffic time series by time-frequency decomposition a time-frequency sequence, and input into a constructed traffic sequence compression and reconstruction deep model in combination with a road topology directed graph; The input module is configured to: S110, the traffic time sequence is divided into equidistant segmented traffic time sequences, and a road topology grid distance between traffic time sequence acquisition devices is used to construct a road topology directed graph, and the segmented traffic time sequences and the road topology directed graph are normalized; S120, the normalized segmented traffic time series is decomposed into frequency sequence, and the normalized road topology directed graph is input into the constructed traffic sequence compression and reconstruction deep model. The traffic sequence compression and reconstruction deep model comprises a traffic sequence feature extraction and coding module, a road topology structure directed graph feature extraction module, a feature supervised clustering module and a sequence reconstruction decoding module. The traffic sequence feature extraction and coding module inputs segmented traffic time sequence decomposed dimensional frequency sequence, and outputs dimensional traffic sequence feature vector wherein ≥ ; The road topology directed graph feature extraction module inputs the road topology directed graph and outputs a deep feature of the road topology directed graph ; The features supervised clustering module input is dimensional traffic sequence feature vector The output is dimensional traffic sequence clustering features and the category to which it belongs; The sequence reconstruction decoding module inputs at least one sequence initial value of segmented traffic time sequence , dimensional traffic sequence clustering feature and deep feature , and outputs reconstructed traffic time sequence ; The traffic time sequence feature database comprises a road topology directed graph database, a traffic segmented time sequence category library and a feature semantic library. The training module is configured to perform a training process of a traffic sequence compression and reconstruction deep model, and first extracts features of an input object and then obtains a supervised clustering of the features to form a feature semantic library in the training process a feature class set, and forms a feature semantic library. A loss function is designed to repeatedly adjust a classification parameter of the supervised clustering until a training stop condition is reached, and a traffic time sequence feature database is obtained when the training stop condition is reached. a reconstruction module configured to extract sequence initial values from the traffic time series feature database according to time sequence timestamps in the reconstruction requirement and the corresponding road, a road topology directed graph corresponding to the road corresponding to the feature class set, and inputting the three to a sequence reconstruction decoding module for reconstruction to obtain the reconstructed traffic flow time series.

Citation Information

Patent Citations

  • Traffic prediction method based on enhanced space-time diagram neural network

    CN112241814A

  • Road network traffic state discrimination method based on clustering and graph convolutional network

    CN113450562A