Traffic data compensation model determination method, traffic data compensation method and related device

By combining the denoising stack autoencoder and the generative adversarial network, the problem of large-scale high-deletion traffic data is solved, efficient and accurate recovery of data is achieved, and the integrity and prediction performance of traffic data are improved.

CN120234535AActive Publication Date: 2025-07-01SHENYANG AEROSPACE UNIVERSITY
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510390738.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

When faced with large-scale and high missing rates, existing traffic data compensation technologies are difficult to accurately recover the original distribution of data, and are inefficient in computing, which cannot meet real-time or near-real-time traffic flow forecasting and decision support needs.

Method used

The method of combining denoising stack autoencoder and generative adversarial network is adopted to train the generative adversarial network by random missing processing and defect mapping of traffic data sets, and generate high-quality compensatory data, retaining the inherent laws and statistical characteristics of the data.

Benefits of technology

It improves the completeness and accuracy of the data, enhances the generalization ability of the model, and can accurately and efficiently compensate for traffic data at high missing rates, and supports traffic data analysis, prediction and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234535A_ABST
    Figure CN120234535A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic data compensation model determination method, a traffic data compensation method and a related device, and relates to the technical field of traffic data compensation, and the determination method comprises the steps: carrying out the random missing processing and defect mapping processing of a complete traffic data set, and obtaining a defect vector set; using the complete traffic data set and the defect vector set to train a de-noising stack type auto-encoder to obtain a trained de-noising stack type auto-encoder, and inputting the defect vector set into the trained de-noising stack type auto-encoder to carry out dimension raising processing to obtain first high-dimensional feature vector data; training a generative adversarial network by using the complete traffic data set and the first high-dimensional feature vector data to obtain a trained generative adversarial network; and sequentially connecting an encoder of the trained de-noising stack type auto-encoder, the generative adversarial network and a decoder of the de-noising stack type auto-encoder to obtain a traffic data compensation model. According to the invention, the missing part in the traffic data can be accurately and efficiently made up.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of traffic data compensation, and particularly to a method for determining a traffic data compensation model, a traffic data compensation method, and related devices. Background Art

[0002] With the rapid development of the Intelligent Transportation System (ITS), the collection, analysis, and application of traffic data have become the core of modern traffic management and planning. ITS obtains a large amount of information such as traffic flow, speed, and density in real time through sensors, cameras, and other data collection devices. These data provide important support for traffic flow prediction, congestion management, and route optimization. However, due to reasons such as sensor failures, data transmission errors, and environmental interference, the collected traffic data often has missing values. Data missing not only affects the integrity of the data but also poses a serious challenge to the accuracy of subsequent data analysis and model prediction. Therefore, how to effectively compensate for the missing data and restore the original distribution and statistical characteristics of the data has become a key problem to be solved urgently in the field of intelligent transportation systems.

[0003] Existing data compensation techniques, such as mean filling, linear interpolation, k-nearest neighbor (kNN), etc., although perform well when dealing with small-scale or low-missing-rate data sets, tend to be inadequate when faced with large-scale and high-missing-rate traffic data. Summary of the Invention

[0004] The purpose of the present application is to provide a method for determining a traffic data compensation model, a traffic data compensation method, and related devices, which can accurately and efficiently compensate for the missing parts in traffic data and provide strong support for traffic data analysis, prediction, and decision-making.

[0005] To achieve the above purpose, the present application provides the following solutions:

[0006] In a first aspect, the present application provides a method for determining a traffic data compensation model, including:

[0007] Performing random missing processing and defect mapping processing on a complete traffic data set to obtain a set of defect vectors; the complete traffic data set includes traffic flow data, spatio-temporal data, and environmental data;

[0008] Training a denoising stacked autoencoder using the complete traffic data set and the set of defect vectors to obtain a trained denoising stacked autoencoder;

[0009] Inputting the set of defect vectors into the encoder of the trained denoising stacked autoencoder for dimensionality elevation to obtain first high-dimensional feature vector data;

[0010] Train a generative adversarial network using the complete traffic data set and the first high-dimensional feature vector data to obtain a trained generative adversarial network;

[0011] Connect the encoder of the trained denoising stacked autoencoder, the trained generative adversarial network, and the decoder of the trained denoising stacked autoencoder in sequence to obtain a traffic data compensation model.

[0012] In a second aspect, the present application provides a traffic data compensation method, including:

[0013] Obtain original traffic data;

[0014] Input the original traffic data into the traffic data compensation model to obtain compensated traffic data; the traffic data compensation model is a model trained according to the traffic data compensation model determination method described above.

[0015] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the traffic data compensation model determination method or the traffic data compensation method described above.

[0016] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the traffic data compensation model determination method or the traffic data compensation method described above.

[0017] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the traffic data compensation model determination method or the traffic data compensation method described above.

[0018] According to the specific embodiments provided by the present application, the following technical effects are disclosed:

[0019] The present application provides a method for determining a traffic data compensation model, a traffic data compensation method and related devices. The present application can process a data set including multiple dimensions such as traffic flow, spatio-temporal information, and environmental data, comprehensively consider the influence of various factors on traffic conditions, and use a denoising stacked autoencoder for data processing. It can remove noise in the original data while processing missing data, avoid errors caused by data missing, and improve the prediction performance of subsequent models. Moreover, the generative adversarial network not only improves the quality of generated data through adversarial training, but also reduces overfitting during the recovery of missing data, enhances the generalization ability of the model, and combines the generative adversarial network with the denoising stacked autoencoder, enabling the model to better capture the internal laws of data when generating missing data, accurately and efficiently compensate for the missing parts in traffic data, improve the integrity and accuracy of data, and provide strong support for traffic data analysis, prediction and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 It is an application environment diagram of a method for determining a traffic data compensation model in an embodiment of the present application;

[0022] Figure 2 It is a flowchart of a method for determining a traffic data compensation model provided in an embodiment of the present application;

[0023] Figure 3 It is a structural diagram of a denoising autoencoder provided in an embodiment of the present application;

[0024] Figure 4 It is a structural diagram of a denoising stacked autoencoder provided in an embodiment of the present application;

[0025] Figure 5 It is a structural diagram of a traffic data compensation model provided in an embodiment of the present application;

[0026] Figure 6 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0028] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0029] The applicant's research found that: traditional methods, such as mean imputation, linear interpolation, k-nearest neighbor (kNN), etc., usually rely on simple statistics or neighboring data points and lack an in-depth understanding of the internal structure and patterns of the data. Traffic data has complex spatio-temporal correlations and highly non-linear characteristics. Traditional methods are difficult to effectively restore the original distribution of the data, resulting in significant differences between the compensated data and the real data. In addition, traditional methods have low computational efficiency and high time complexity when dealing with large-scale data sets, making it difficult to meet the requirements of real-time or near-real-time traffic flow prediction and decision support. Therefore, there is an urgent need for a new data compensation method that can efficiently process large-scale traffic data with a high missing rate and can maintain the original distribution and statistical characteristics of the data.

[0030] In view of this, the traffic data compensation model determination method provided in the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, placed on the cloud or other servers. The terminal 102 can send the complete traffic data set to the server 104. After the server 104 receives the complete traffic data set, the server 104 performs random missing processing and defect mapping processing on the complete traffic data set to obtain a defect vector set; uses the complete traffic data set and the defect vector set to train a denoising stacked autoencoder to obtain a trained denoising stacked autoencoder, and inputs the defect vector set for dimensionality increase processing to obtain the first high-dimensional feature vector data; uses the complete traffic data set and the first high-dimensional feature vector data to train a generative adversarial network to obtain a trained generative adversarial network; connects the encoder of the trained denoising stacked autoencoder, the generative adversarial network, and the decoder of the denoising stacked autoencoder in sequence to obtain a traffic data compensation model. The server 104 can feedback the obtained traffic data compensation model to the terminal 102. In addition, in some embodiments, the method for determining the traffic data compensation model can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly process the complete traffic data set, or the server 104 can obtain the complete traffic data set from the data storage system and process it.

[0031] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, and Internet of Things devices. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.

[0032] In an exemplary embodiment, as Figure 2 shown, a method for determining a traffic data compensation model is provided. This method is executed by a computer device, and specifically can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in

[0033] Step 201, perform random missing processing and defect mapping processing on the complete traffic data set to obtain a defect vector set; the complete traffic data set includes traffic flow data, spatio-temporal data, and environmental data.

[0034] Step 202, use the complete traffic data set and the defect vector set to train a denoising stacked autoencoder to obtain a trained denoising stacked autoencoder.

[0035] Step 203: Input the defective vector set into the encoder of the trained denoising stacked autoencoder for dimensionality increase to obtain the first high-dimensional feature vector data.

[0036] Step 204: Use the complete traffic data set and the first high-dimensional feature vector data to train a generative adversarial network to obtain a trained generative adversarial network.

[0037] Step 205: Connect the encoder of the trained denoising stacked autoencoder, the trained generative adversarial network, and the decoder of the trained denoising stacked autoencoder in sequence to obtain a traffic data compensation model.

[0038] Implementing the above steps 201 to 205 can obtain a traffic data compensation model. This model can better capture the internal laws of the data when generating missing data, accurately and efficiently compensate for the missing parts in the traffic data, improve the integrity and accuracy of the data, and provide strong support for traffic data analysis, prediction, and decision-making.

[0039] Furthermore, the denoising stacked autoencoder (DSAE) has two basic component modules: the autoencoder AE (autoencoder) and the denoising autoencoder DAE (denoising autoencoder). A single AE can extract features from the original input data, while stacked AEs can be stacked into a deep neural network and obtain an abstract representation of the input data in a way of gradually extracting features. A single DAE is a stochastic version of the AE and can capture the statistical correlations between inputs. To better compensate for missing traffic data, this application uses DSAE, which retains the functional advantages of both DAE and stacked AEs. Among them, DAE mainly realizes the function of denoising defective traffic data, while stacked AEs mainly help extract features from traffic data. DSAE realizes data recovery through feature extraction and statistical correlation learning. Next, the component structure and training algorithm of DSAE will be specifically introduced.

[0040] Specifically, DAE is a variant of AE, and its component structure is as Figure 3 shown. The input part of DAE sends the defective original input data into the input layer. The purpose of DAE training is to reconstruct the original data from a defective original input data, and the reconstructed original data is obtained from the output layer of the network. The defective vector of the original input data x is obtained through the corruption mapping:

[0041]

[0042] After this mapping is completed, is sent into the input layer of the network. The encoder f θMap the data of the input layer to a representation h of the hidden layer, that is f θ is a non-linear transformation function of the following form:

[0043]

[0044] where θ represents the parameters of the encoder in the denoising autoencoder, including W and b; W is the weight matrix between the hidden layer and the input layer, and b is the bias vector of the hidden layer. In the denoising autoencoder, the decoder g' θ maps the representation h of the hidden layer back to a reconstructed vector y of the input vector and outputs it at the output layer, that is y = g' θ (h). g' θ is also a non-linear transformation function, and its form is as follows:

[0045]

[0046] where θ' represents the parameters of the decoder, including W' and b'; W' is the weight matrix between the output layer and the hidden layer, and b' is the bias vector of the output layer; s is the activation function.

[0047] To reconstruct the clean original input data, the reconstruction error must be a measure of the error between the reconstructed vector y and the clean original vector x. The process of training the DAE is the process of minimizing the reconstruction error, which is achieved by solving the following optimization problem:

[0048] (θ,θ′)=argminL(X,Y);

[0049] where X is the set of input vectors x; Y is the set of corresponding reconstructed vectors y; L is the loss function, usually defined in the following form:

[0050]

[0051] where L(X,Y) is the value of the loss function; M is the number of reconstructed vectors; N is the dimension of the reconstructed vectors; x ij is the i-th j-dimensional input vector; y ij is the i-th j-dimensional reconstructed vector.

[0052] After training is completed, the hidden layer representation of the DAE is regarded as a useful representation of the input, because the clean input data corresponding to a defective data can be recovered from this hidden layer representation. In addition, for a new defective vector fed into the input layer, the trained DAE can recover its corresponding clean vector.

[0053] DSAE consists of a bottom - layer DAE and a stacked AE in the middle layer, and its composition structure is as Figure 4 shown. The purpose of the bottom - layer DAE is to recover the clean input vector from the defective vector, while the stacked AE in the middle layer extracts features from the input layer.

[0054] For the defective vector fed into the input layer, it is mapped to the reconstructed vector y through the following formula:

[0055]

[0056] where W1 is the weight matrix between the first hidden layer and the input layer; b1 is the bias vector of the first hidden layer; h l is the hidden - layer output representation of the l - th hidden layer; W l is the weight matrix between the l - th hidden layer and the (l - 1)-th hidden layer; b l is the bias vector of the l - th hidden layer; W l+1 is the weight matrix between the output layer and the (l + 1)-th hidden layer; b l+1 is the bias vector of the output layer.

[0057] The process of training DSAE is also a process of minimizing the reconstruction error, which is achieved by solving the following optimization problem:

[0058]

[0059] where θ represents the parameters including W l and b l (l = 1, 2,..., L + 1); X is the complete traffic data set, Y is the set of data vectors reconstructed by the model, and L is the loss function, which is defined as:

[0060]

[0061] where L(X,Y) is the value of the loss function; M is the number of reconstructed vectors; N is the dimension of the reconstructed vectors; x ij is the i - th j - th - dimensional input vector; y ij is the i - th j - th - dimensional reconstructed vector.

[0062] Different from training AE or DAE, this process includes two steps: pre - training and global fine - tuning. The training trains the parameters of each AE (i.e., hidden layer), while the global fine - tuning adjusts all the weight and bias parameters of DSAE. After training, DSAE can recover its corresponding clean input vector from the defective vector.

[0063] Furthermore, using the complete traffic data set and the defective vector set to train the denoising stacked auto - encoder, a trained denoising stacked auto - encoder is obtained, specifically including:

[0064] Determine whether the current hidden layer is the first hidden layer.

[0065] When the current hidden layer is the first hidden layer, input the defective vector set from the input layer, pass through the first hidden layer to obtain the reconstruction vector of the first hidden layer, and update the first hidden layer based on the complete traffic data set, the reconstruction vector of the first hidden layer, and the loss function to obtain the updated first hidden layer.

[0066] When the current hidden layer is not the first hidden layer, input the reconstruction vector of the previous hidden layer into the current hidden layer to obtain the reconstruction vector of the current hidden layer, and update the parameters of the current hidden layer using the backpropagation algorithm based on the complete traffic data set, the reconstruction vector of the current hidden layer, and the loss function to obtain the updated current hidden layer.

[0067] Take each hidden layer as the current hidden layer in turn until each hidden layer is updated to obtain the first denoising stacked autoencoder; the first denoising stacked autoencoder is obtained by training each hidden layer of the denoising stacked autoencoder one by one.

[0068] Input the defective vector set into the first denoising stacked autoencoder to obtain the output data of the first denoising stacked autoencoder.

[0069] Update all the parameters in the first denoising stacked autoencoder using the backpropagation algorithm based on the complete traffic data set, the output data of the first denoising stacked autoencoder, and the loss function, that is, obtain the trained denoising stacked autoencoder by globally fine-tuning the first denoising stacked autoencoder.

[0070] Furthermore, the generative adversarial network model of this embodiment is the core part of the algorithm architecture. Both the generator and the discriminator are deep neural network structures composed of multiple activation functions. The activation functions used include the PRELU function, the Dropout function, the Log Sigmoid function, and the Tanh function, etc. Specifically, the activation function used in the hidden layer of the generator is the PRELU function, the activation function used in the output layer of the generator is the Dropout function, the activation function used in the hidden layer of the decoder is the Tanh function, and the activation function used in the output layer of the decoder is the Log Sigmoid function. During the training process of the network, the generator G fits the distribution of the entire data and generates the sample z, while the discriminator D is responsible for judging whether the sample z comes from G or is sampled from the real data x. The goal of the generative adversarial network is to continuously train G to make the generated data closer to the real samples, and at the same time continuously optimize D to make it distinguish the source of the data as accurately as possible. The optimization functions for G and D can be expressed as follows:

[0071] minmaxV(D,G) = Ex~pdata(x)[logD(x)] + Ez~p(z)[log(1 - D(G(z)))];

[0072] Among them, V(D,G) is the optimization function; logD(x) is the logarithm of the output probability of the discriminator for the real data x; log(1 - D(G(z))) is the logarithm of the output probability of the discriminator for the data G(z) generated by the generator.

[0073] In the function V(D,G), the first term is the entropy of the data from the real distribution pdata(x) passing through the discriminator, and the discriminator attempts to maximize it to 1; the second term is the entropy of the data from the random input p(z) passing through the generator. The generator generates a fake sample, and the discriminator identifies the falsity. In this term, the discriminator attempts to minimize it to 0. Overall, the discriminator attempts to maximize the function V(D,G). On the other hand, the task of the generator is exactly the opposite. It attempts to minimize the function V(D,G) to minimize the difference between real data and fake data. When the classification accuracy rate of the discriminator D converges to 50%, it indicates that the generator G has enabled the discriminator D to be unable to distinguish the data source, and it is considered that the model training is completed. At this time, the high-dimensional feature vector after dimension elevation by the encoder of the denoising autoencoder can be used as the input, and the generator G is used to generate a new data set.

[0074] Furthermore, training a generative adversarial network using the complete traffic data set and the first high-dimensional feature vector data to obtain a trained generative adversarial network specifically includes:

[0075] Inputting the first high-dimensional feature vector data into the generator to obtain the generated data output by the generator.

[0076] Inputting the complete traffic data set and the generated data output by the generator into the discriminator to obtain a probability distribution;

[0077] Constructing an optimization function and updating the generator and the discriminator according to the probability distribution and the optimization function to obtain a trained generative adversarial network.

[0078] Combining the foregoing algorithm principles and certain preliminary hyperparameter experiments, this application combines the denoising stacked autoencoder with the generative adversarial network, and the traffic data compensates for the model structure as Figure 5As shown in the figure. The learning rate of the denoising stacked autoencoder is 0.005. The encoder includes an input layer, seven hidden layers, and an output layer connected in sequence, and the decoder includes an input layer, seven hidden layers, and an output layer connected in sequence. Considering that the highest dimension of the experimental data selected in this embodiment is an 18-dimensional dataset, weighing the computational cost of the network and the data representation ability of the feature vectors, it is set that the encoder raises the original traffic data to 200 - 300 dimensions through a deep neural network structure. At this time, its information volume already meets the sample dimension required for the training of the generative adversarial network. In subsequent dataset experiments, it is set to mutually convert the original traffic data and the 216-dimensional feature vectors. During the training process, the encoder performs 2% random erasure on the input signal and adds Gaussian white noise with a mean of 0.

[0079] After the denoising stacked autoencoder is trained, its encoder part is used to raise the dimension of the defective vectors, and the transformed first high-dimensional feature vectors are sent into the generative adversarial network for training. The learning rate of the generator G is 0.001, and it consists of an input layer, fourteen hidden layers, and an output layer. The learning rate of the discriminator D is 0.005, and it consists of an input layer, eighteen hidden layers, and an output layer. When designing the architecture, this application makes the discriminator slightly more powerful than the generator, making the discriminator more strict during training and the generated data closer to the real data. When the classification accuracy rate of the discriminator approaches 50%, the training is completed. The training time of the model on a single Tesla graphics card is about 60 to 500 minutes according to the scale of the dataset.

[0080] Compared with the prior art, the traffic data filling model generated by this application has the following advantages:

[0081] (1) This application can effectively handle a missing data rate of up to 80%. In contrast, traditional methods often cannot provide accurate data filling in the case of a high missing rate, resulting in a significant decline in the performance of the prediction model. By combining the use of a denoising stacked autoencoder (DSAE) and a generative adversarial network (GAN), the DSAE enhances the model's robustness to missing data by introducing random noise during the training process, and realizes the dimensionality increase of data through stacking multiple autoencoders to extract richer feature representations; the GAN then uses these high-dimensional feature vectors to generate samples similar to the original data, thereby generating high-quality filling data under high missing rate conditions.

[0082] (2) Since the discriminator of the generative adversarial network is responsible for distinguishing between the generated data and the real data, the quality of the generated data is improved through adversarial training, making the generated data closer to the real data in terms of statistical characteristics.

[0083] (3) Compared with traditional data compensation methods, the method of this application can better preserve the overall distribution of data and enhance the generalization ability of subsequent prediction models. By means of random masking and noise injection in the pre-training stage of DSAE, the actual situation of data loss is simulated, enabling the model to have better generalization ability when facing real missing data.

[0084] (4) The method of this application is through an optimized algorithm architecture design, especially the combined use of a denoising stacked autoencoder and a generative adversarial network. The pre-training stage of DSAE can process multiple data batches in parallel, while the training process of GAN reduces the model training time through optimized network structures and parameter settings.

[0085] (5) Through the combination of DSAE and GAN, the generated data not only has similar statistical characteristics to the original data, but also can better reflect the dynamic changes of traffic flow, providing higher-quality input data for the traffic flow prediction model, thereby improving the accuracy and reliability of the prediction.

[0086] This application also provides an application scenario that applies the above traffic data compensation model determination method. Specifically: The traffic data compensation model determination method provided in this embodiment can be applied in a traffic data compensation scenario. The traffic data compensation scenario includes an original traffic data acquisition link and an original traffic data compensation link; the original traffic data is input into the traffic data compensation model to obtain the compensated traffic data. The traffic data compensation model determination method provided in this embodiment belongs to the training link of the traffic data compensation model in the original traffic data compensation link.

[0087] In an exemplary embodiment, a method for compensating a traffic data compensation model is provided, including:

[0088] Obtain the original traffic data.

[0089] Input the original traffic data into the traffic data compensation model to obtain the compensated traffic data.

[0090] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store processed data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for determining a traffic data compensation model.

[0091] Those skilled in the art can understand that Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0092] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0093] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0094] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0095] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0096] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.

[0097] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0098] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for determining a traffic data compensation model, characterized in that: The traffic data compensation model determination method comprises: Performing random missing processing and defect mapping processing on a complete traffic data set to obtain a defect vector set; the complete traffic data set includes traffic flow data, spatiotemporal data, and environmental data; Using the complete traffic data set and the defect vector set to train a denoising stacked autoencoder to obtain a trained denoising stacked autoencoder; Inputting the defective vector set into the encoder of the trained denoising stacked autoencoder for dimension upgrading to obtain first high-dimensional feature vector data; Using the complete traffic data set and the first high-dimensional feature vector data to train a generative adversarial network to obtain a trained generative adversarial network; The trained encoder of the denoising stacked autoencoder, the trained generative adversarial network and the trained decoder of the denoising stacked autoencoder are connected in sequence to obtain a traffic data compensation model.

2. The method for determining a traffic data compensation model according to claim 1, characterized in that: The denoising stacked autoencoder comprises an input layer, a plurality of hidden layers and an output layer connected in sequence; the complete traffic data set and the defect vector set are input into the denoising stacked autoencoder, and the denoising stacked autoencoder is trained to obtain a trained denoising stacked autoencoder, specifically comprising: Obtain a first denoising stacked autoencoder; the first denoising stacked autoencoder is obtained by training each hidden layer of the denoising stacked autoencoder one by one; Inputting the defect vector set into the first denoising stacked autoencoder to obtain output data of the first denoising stacked autoencoder; Based on the complete traffic data set, the output data of the first denoising stacked autoencoder and the loss function, the first denoising stacked autoencoder is updated to obtain a trained denoising stacked autoencoder.

3. The method for determining a traffic data compensation model according to claim 2, characterized in that: Get the first denoising stacked autoencoder, including: Determine whether the current hidden layer is the first hidden layer; When the current hidden layer is the first hidden layer, the defect vector set is input from the input layer, passed through the first hidden layer, and a reconstruction vector of the first hidden layer is obtained, and based on the complete traffic data set, the reconstruction vector of the first hidden layer and the loss function, the first hidden layer is updated to obtain an updated first hidden layer; When the current hidden layer is not the first hidden layer, a reconstruction vector of a previous hidden layer is input into the current hidden layer to obtain a reconstruction vector of the current hidden layer, and the current hidden layer is updated based on the complete traffic data set, the reconstruction vector of the current hidden layer and the loss function to obtain an updated current hidden layer; Each hidden layer is used as the current hidden layer in turn until each hidden layer is updated to obtain the first denoising stacked autoencoder.

4. The method for determining a traffic data compensation model according to claim 1, characterized in that: The generative adversarial network includes a generator and a discriminator; the complete traffic data set and the first high-dimensional feature vector data are input into the generative adversarial network, and the generative adversarial network is trained to obtain a trained generative adversarial network, which specifically includes: Inputting the first high-dimensional feature vector data into the generator to obtain generated data output by the generator; Inputting the complete traffic data set and the generated data into the discriminator to obtain a probability distribution; An optimization function is constructed, and the generator and the discriminator are updated according to the probability distribution and the optimization function to obtain a trained generative adversarial network.

5. The method for determining a traffic data compensation model according to claim 1, characterized in that: The encoder of the trained denoising stacked autoencoder includes an input layer, seven hidden layers and an output layer connected in sequence; the decoder of the trained denoising stacked autoencoder includes an input layer, seven hidden layers and an output layer connected in sequence; the trained generative adversarial network includes a trained generator and a discriminator, wherein the trained generator includes an input layer, fourteen hidden layers and an output layer connected in sequence; the trained decoder includes an input layer, eighteen hidden layers and an output layer connected in sequence.

6. The method for determining a traffic data compensation model according to claim 5, characterized in that: The activation function used by the hidden layer of the trained generator is the PRELU function; the activation function used by the output layer of the trained generator is the Dropout function; the activation function used by the hidden layer of the trained decoder is the Tanh function; the activation function used by the output layer of the trained decoder is the Log Sigmoid function.

7. A method for compensating traffic data, characterized in that: The traffic data compensation method includes: Get raw traffic data; The original traffic data is input into a traffic data compensation model to obtain compensated traffic data; the traffic data compensation model is a model obtained according to the traffic data compensation model determination method according to any one of claims 1 to 6.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the traffic data compensation model determination method described in any one of claims 1 to 6 or the traffic data compensation method described in claim 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for determining a traffic data compensation model according to any one of claims 1 to 6 or the method for determining traffic data compensation according to claim 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for determining a traffic data compensation model according to any one of claims 1 to 6 or the method for determining traffic data compensation according to claim 7 is implemented.

Citation Information

Patent Citations

  • Significant object detection method based on stack-typed denoising self-coding machine

    CN103955936A

  • Traffic data make-up method

    CN104091081A

  • A data reduction method based on a stack noise reduction self-coding neural network

    CN109598336A

  • Improved data cleaning method for stack noise reduction auto-encoder

    CN109978079A

  • Road network traffic data restoration method based on SAE-GA-SAD

    CN110942624A