A traffic data prediction method and device

By adopting a multi-model layer structure prediction model in traffic data prediction, combining timing, static space and dynamic space sub-models, the problem that the existing technology is difficult to achieve effective traffic data prediction in different scenarios is solved, and the adaptability and generalization ability of prediction is improved.

CN119672961BActive Publication Date: 2025-05-30HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510171827.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-30
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

It is difficult to achieve effective generalization in different scenarios, especially in complex scenarios such as highways and urban roads. A single modeling method is difficult to adapt to changes in multiple scenarios.

Method used

A prediction model with a multi-model layer structure, including a time-series sub-model, a static spatial sub-model and a dynamic spatial sub-model, uses the routing layer to select the output results of each sub-model to realize traffic data prediction in multiple scenarios.

Benefits of technology

It improves the adaptability and generalization ability of traffic data prediction, avoids dependence on the prediction results of a single model, and enhances the timing generalization ability of the prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672961B_ABST
    Figure CN119672961B_ABST
Patent Text Reader

Abstract

The present application discloses a traffic data prediction method, which includes: using a trained prediction model for traffic data prediction, based on the historical traffic data at the historical time of the traffic position to be predicted, obtaining the predicted traffic data at the future time of the traffic position to be predicted, wherein the prediction model includes: a multi-model layer composed of at least two of a time series sub-model for obtaining the time series features of historical traffic data to obtain a first prediction result, a static space sub-model for obtaining the static space features of historical traffic data to obtain a second prediction result, and a dynamic space sub-model for obtaining the dynamic space features of historical traffic data to obtain a third prediction result, and a routing layer for selecting the prediction results output by each sub-model in the multi-model layer. The present application realizes the prediction combination of at least two of the time series features, static space features, and dynamic space features, and improves the adaptability of traffic data prediction in various application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation, and in particular, to a traffic data prediction method. Background Art

[0002] Accurate traffic data prediction in an intelligent transportation system is a prerequisite for precise analysis of traffic problems. Traffic data includes and is not limited to traffic road network indicators such as traffic flow, length of the waiting queue, and passing speed at a set traffic location. For example, traffic flow, length of the waiting queue, passing speed, etc. at a certain intersection. Traffic data prediction involves a variety of different scenarios, including but not limited to highways and urban roads. Among them, in addition to normal road connections, there are also some spatially isolated roads in the road network, such as some roads without connections to other roads widely existing at the urban fringe; there are also roads with complex spatio-temporal characteristics, such as highway ramps in the city, which have special road structures and are easily affected by specific events such as holidays.

[0003] Related traffic data prediction methods usually use a single modeling method to better adapt to traffic data prediction in a certain scenario, but it is difficult to generalize to other scenarios and achieve good results. Summary of the Invention

[0004] The present invention provides a traffic data prediction method to achieve the generalization ability of traffic data prediction in different scenarios.

[0005] The first aspect of the present application provides a traffic data prediction method, which includes:

[0006] Using a trained prediction model for traffic data prediction, based on the historical traffic data at the traffic location to be predicted at historical time, obtaining the predicted traffic data at the traffic location to be predicted at future time,

[0007] Wherein,

[0008] The prediction model includes: a multi-model layer composed of at least two of a time series sub-model for obtaining the time series characteristics of historical traffic data to obtain a first prediction result, a static space sub-model for obtaining the static space characteristics of historical traffic data to obtain a second prediction result, and a dynamic space sub-model for obtaining the dynamic space characteristics of historical traffic data to obtain a third prediction result, and a routing layer for selecting the prediction results output by each sub-model in the multi-model layer.

[0009] As a possible implementation manner, the prediction model is trained in the following manner:

[0010] The historical traffic data is segmented into time series segments with a set sample time length, and the time series segments are used as sample data and input into the prediction model to be trained,

[0011] Calculate the prediction loss function value of the prediction model based on the number of sub-models, the classification prediction loss function values of each sub-model, and the sample output probabilities of each sub-model. Among them, the classification prediction loss function values include: the first type of prediction loss function value used to discard poor sub-models, and the second type of prediction loss function value used to select better sub-models.

[0012] Perform backpropagation based on the prediction loss function value of the prediction model.

[0013] Train repeatedly until the training ends.

[0014] As a possible implementation, the prediction loss function value of the prediction model is calculated in the following way:

[0015] For any sub-model, calculate the product between the classification prediction loss function value of this sub-model and the logarithm of the sample output probability of this sub-model to obtain the prediction loss function value of this sub-model.

[0016] Calculate the average value of the prediction loss function values of all sub-models to obtain the prediction loss function value of the prediction model;

[0017] The first type of prediction loss function value is determined in the following way:

[0018] For any sub-model,

[0019] In the case where the difference between the predicted sample value output by this sub-model and the true value is less than the set first difference threshold and the sample output probability of this sub-model is greater than the set sample output probability threshold, the first type of prediction loss function value of this sub-model is set to 1.

[0020] In the case where the difference between the predicted sample value output by this sub-model and the true value is greater than the first difference threshold and the sample output probability of this sub-model is not greater than the sample output probability threshold, the first type of prediction loss function value of this sub-model is set to the reciprocal of the number of sub-models minus 1.

[0021] Otherwise, the first type of prediction loss function value of this sub-model is set to 0;

[0022] The second type of prediction loss function value is determined in the following way:

[0023] For any sub-model,

[0024] In the case where the difference between the predicted sample value output by this sub-model and the true value is less than the second difference threshold and the sample output probability of this sub-model is greater than the sample output probability threshold, the second type of prediction loss function value of this sub-model is set to 1.

[0025] In the case where the difference between the predicted sample value output by the sub-model and the true value is greater than the second difference threshold and the sample output probability of the sub-model is not greater than the sample output probability threshold, the value of the second type of prediction loss function of the sub-model is set to the reciprocal of the number of sub-models minus 1.

[0026] Otherwise, the value of the second type of prediction loss function of the sub-model is set to 0.

[0027] Among them,

[0028] The second difference threshold is the difference between 1 and the first difference threshold. The difference threshold is the quantile threshold for characterizing the equal-value points within the probability distribution range.

[0029] The output probability threshold is determined according to the number of types of the predicted results of the sub-models to be output.

[0030] As a possible implementation manner, using the trained prediction model for traffic data prediction to obtain the predicted traffic data at a future time for a traffic position to be predicted based on the historical traffic data at the historical time of the traffic position to be predicted includes:

[0031] Extracting the historical traffic data at the traffic position to be predicted according to a set historical time step length to obtain the historical traffic data at each historical time point within the set historical time step length.

[0032] Sequentially inputting the historical traffic data at each historical time point into the trained prediction model in chronological order, so that: each sub-model in the multi-model layer of the prediction model respectively obtains the prediction result of each sub-model from the currently input historical sequence data, and the routing layer selects the prediction results of the sub-models with output probabilities greater than the set output probability threshold according to the output probabilities of the sub-models.

[0033] Among them,

[0034] The output probability of the sub-model is used to characterize: the similarity between the prediction result output by the sub-model and the pattern features of the current historical sequence data, and the proportion it occupies in the sum of the similarities between the prediction results output by each sub-model and the pattern features of the current historical sequence data.

[0035] The pattern features are used to characterize the correlation between the current historical sequence data and the current time series pattern.

[0036] The current time series pattern is used to characterize each traffic data pattern with a relatively high recurrence probability found through time series search.

[0037] As a possible implementation manner, the prediction result is the predicted sequence data composed of the predicted traffic data at each future time point, where the time step length between two adjacent future time points is the set prediction time step length.

[0038] As a possible implementation, the output probability is determined in the following manner:

[0039] For any sub-model,

[0040] According to the linear weight and the linear offset, calculate the multi-model layer prediction sequence data used to characterize the prediction result of the multi-model layer for the current historical sequence data,

[0041] According to the multi-model layer prediction sequence data and the current time series pattern, calculate the correlation degree between the multi-model layer prediction sequence data and each traffic data pattern in the current time series pattern, and weight each traffic data pattern in the current time series pattern through the calculated correlation degrees to obtain the pattern feature of the current historical sequence data,

[0042] According to the pattern feature of the current historical sequence data and the prediction result output by this sub-model, calculate the similarity between the prediction result output by this sub-model and the pattern feature of the current historical sequence data,

[0043] According to the pattern feature of the current historical sequence data and the prediction results output by each sub-model, calculate the similarity between the prediction results output by each sub-model and the pattern feature of the current historical sequence data to obtain the similarities of each sub-model, and accumulate the similarities of each sub-model to obtain the sum of the similarities,

[0044] Calculate the ratio of the similarity between the prediction result output by this sub-model and the pattern feature of the current historical sequence data to the sum of the similarities to obtain the output probability of this sub-model.

[0045] As a possible implementation, the step of calculating the multi-model layer prediction sequence data used to characterize the prediction result of the multi-model layer for the current historical sequence data according to the linear weight and the linear offset includes:

[0046] Multiply the linear weight by the current historical sequence data, and accumulate the linear offset to obtain the multi-model layer prediction sequence data of the current historical sequence data;

[0047] As a possible implementation, the step of calculating the correlation degree between the multi-model layer prediction sequence data and each traffic data pattern in the current time series pattern according to the multi-model layer prediction sequence data and the current time series pattern includes:

[0048] For any traffic data pattern in the current time series pattern,

[0049] Calculate the exponential function value of the exponential function with the base of the natural constant and the exponent being the product of the multi-model layer prediction sequence data and the transpose of the row vector used to characterize the traffic data pattern, to obtain the exponential function value of the traffic data pattern, where the multi-model layer prediction sequence data is a row vector, the transpose of the row vector of the traffic data pattern is a column vector, the current time series pattern is a time series pattern matrix, each row in the time series pattern matrix corresponds to a traffic data pattern, and each column corresponds to a time series.

[0050] Calculate the exponential function value of the exponential function with the base of the natural constant and the exponent being the product of the multi-model layer prediction sequence data and the transpose of the row vector of each traffic data pattern, to obtain the exponential function value of each traffic data pattern, and accumulate the exponential function values of each traffic data pattern.

[0051] Calculate the ratio of the exponential function value of the traffic data pattern to the accumulated exponential function values of each traffic data pattern, to obtain the correlation degree between the multi-model layer prediction sequence data and the traffic data pattern.

[0052] As a possible implementation manner, the weighting of each traffic data pattern in the current time series pattern by the calculated correlation degrees includes:

[0053] For the correlation degree of each traffic data pattern, calculate the product of the correlation degree of the traffic data pattern and the row vector of the traffic data pattern, to obtain the weighted row vector of the traffic data pattern.

[0054] Add the weighted row vectors of each traffic data pattern to obtain a pattern feature vector.

[0055] As a possible implementation manner, the selection of the prediction results of each sub-model with an output probability greater than the set output probability threshold further includes:

[0056] Determine the fusion weights according to the output probabilities of the selected sub-models.

[0057] Use the fusion weights to perform weighted summation on the selected prediction results to obtain the predicted traffic data.

[0058] As a possible implementation manner, the time series sub-model is a time series large model, and this time series large model is used to model that the traffic time series feature of the traffic data at any traffic location at the next time point is only related to the traffic time series feature of the traffic data at the previous time point at this traffic location.

[0059] As a possible implementation, the static spatial sub-model is a graph neural network model, which is used to model the traffic static spatial features of traffic data at any traffic location at the next time point, related to the traffic static spatial features of traffic data at the same traffic location at the previous time point, and the static spatial features of traffic data at static traffic locations adjacent to the static traffic location.

[0060] As a possible implementation, the dynamic spatial sub-model is a graph neural network model and an attention mechanism model, which is used to model the spatial features of traffic data at any traffic location at the next time point, related to the spatial features of all traffic locations at the previous time point, and the spatial feature correlation is obtained through the attention mechanism.

[0061] The second aspect of the present application provides a prediction model structure for traffic data prediction, and the prediction model structure includes:

[0062] A multi-model layer composed of at least two of a temporal sub-model for obtaining the temporal features of historical traffic data to obtain a first prediction result, a static spatial sub-model for obtaining the static spatial features of historical traffic data to obtain a second prediction result, and a dynamic spatial sub-model for obtaining the dynamic spatial features of historical traffic data to obtain a third prediction result, and

[0063] A routing layer for selecting the prediction results output by each sub-model in the multi-model layer.

[0064] The third aspect of the present application provides a traffic data prediction device, and the device includes:

[0065] A prediction module, which uses the trained prediction model for traffic data prediction, and based on the historical traffic data of the traffic location to be predicted at the historical time, obtains the predicted traffic data of the traffic location to be predicted at the future time.

[0066] Wherein,

[0067] The prediction model includes: a multi-model layer composed of at least two of a temporal sub-model for obtaining the temporal features of historical traffic data to obtain a first prediction result, a static spatial sub-model for obtaining the static spatial features of historical traffic data to obtain a second prediction result, and a dynamic spatial sub-model for obtaining the dynamic spatial features of historical traffic data to obtain a third prediction result, and a routing layer for selecting the prediction results output by each sub-model in the multi-model layer.

[0068] A traffic data prediction method provided by the present application realizes the fusion of the prediction results of historical traffic data by multiple sub-models and a routing layer in a multi-model layer of a prediction model, avoids relying on the prediction result of a single model, realizes the prediction combination of at least two of temporal features, static spatial features, and dynamic spatial features, and improves the adaptability of traffic data prediction in various application scenarios; further, the temporal large model used by the temporal sub-model in multiple sub-models is conducive to combining the temporal large model with a spatial small model, enhancing the overall temporal generalization ability of the prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 FIG. is a schematic flowchart of a traffic data prediction method according to an embodiment of the present application.

[0070] Figure 2 FIG. is a schematic flowchart of a traffic data prediction method according to an embodiment of the present application.

[0071] Figure 3 FIG. is a schematic diagram of modeling a temporal sub-model in this embodiment.

[0072] Figure 4 FIG. is a schematic diagram of modeling a static spatial sub-model in this embodiment.

[0073] Figure 5 FIG. is a schematic diagram of modeling a static spatial sub-model in this embodiment.

[0074] Figure 6 FIG. is a schematic diagram of obtaining control data by a gated network layer in this embodiment.

[0075] Figure 7 FIG. is a schematic diagram of a trained prediction model predicting based on input historical traffic data in this embodiment.

[0076] Figure 8 FIG. is a schematic diagram of a traffic data prediction device according to an embodiment of the present application.

[0077] Figure 9 FIG. is another schematic diagram of a traffic data prediction device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0078] In order to make the objectives, technical means, and advantages of the present application clearer and more understandable, the following further describes the present application in detail with reference to the accompanying drawings.

[0079] An embodiment of the present application uses a trained prediction model having a multi-model layer composed of multiple sub-models to obtain predicted traffic data based on historical traffic data.

[0080] See Figure 1 as shown Figure 1It is a schematic flowchart of a traffic data prediction method according to an embodiment of the present application. The method includes:

[0081] Using the trained prediction model for traffic data prediction, based on the historical traffic data of the traffic location to be predicted at historical times, obtain the predicted traffic data of the traffic location to be predicted at future times.

[0082] Wherein,

[0083] The prediction model includes: a multi-model layer composed of at least two of a time series sub-model for obtaining the time series characteristics of historical traffic data to obtain a first prediction result, a static spatial sub-model for obtaining the static spatial characteristics of historical traffic data to obtain a second prediction result, and a dynamic spatial sub-model for obtaining the dynamic spatial characteristics of historical traffic data to obtain a third prediction result, and a routing layer for selecting the prediction results output by each sub-model in the multi-model layer.

[0084] As an example, the first prediction result represents the predicted traffic data of the traffic location to be predicted at each future time point, the second prediction result represents the predicted traffic data of the traffic location to be predicted at each future time point based on the traffic data at each fixed traffic location related to the traffic location to be predicted in the road network and the historical traffic data of the traffic location to be predicted, and the third prediction result represents the predicted traffic data of the traffic location to be predicted at each future time point based on the historical traffic data of all traffic locations in the road network.

[0085] As an example, according to the set historical time step, extract the historical traffic data of the traffic location to be predicted to obtain the historical traffic data at each historical time point within the set historical time step.

[0086] In chronological order, input the historical traffic data at each historical time point into the trained prediction model in turn, so that: each sub-model in the multi-model layer of the prediction model respectively obtains the prediction result of each sub-model from the currently input historical sequence data, and the routing layer selects the prediction results of each sub-model with an output probability greater than the set output probability threshold according to the output probabilities of each sub-model.

[0087] Wherein, the prediction result is the predicted sequence traffic data for the set prediction duration, and the total number of future time points included in the predicted sequence traffic data may be the same as or different from the total number of historical time points included in the currently input historical sequence data, and the set prediction time step is between adjacent two future time points.

[0088] For example, in ascending order of time, the historical traffic data at each historical time point is input into the trained prediction model in sequence. In this way, the first future time point in the predicted sequence data is the next time point adjacent to the last historical time point in the currently input historical sequence data, and the difference between the first future time point and the last historical time point is the set prediction time step.

[0089] The set historical time step can be determined according to the collection time interval of the historical traffic data. It should be understood that the set historical time step can be a fixed value or a dynamically changing value, and this application does not limit this. The set prediction time step can be the same as or different from the historical time step, and this application does not limit this.

[0090] The output probability of the sub-model is used to characterize: the similarity between the prediction result output by the sub-model and the pattern features of the current historical sequence data, and the proportion it occupies in the sum of the similarities between the prediction results output by each sub-model and the pattern features of the current historical sequence data.

[0091] The pattern features are used to characterize the correlation between the current historical sequence data and the current time series pattern.

[0092] The current time series pattern is used to characterize each traffic data pattern with a relatively high recurrence probability obtained through the time series.

[0093] It should be understood that the traffic location to be predicted can be a single traffic location to be predicted, for example, any intersection in the road network, or multiple traffic locations to be predicted, for example, multiple intersections with a certain distance in the road network. The multiple traffic locations to be predicted can have different forms, including but not limited to intersections, T-junctions, one-way intersections, tidal intersections, multi-way intersections, designated locations, etc., and this application does not limit this.

[0094] For multiple traffic locations to be predicted, the historical traffic data at each traffic location to be predicted can be input into the prediction model in parallel, serially, or alternately in parallel and serial, and this application does not limit this. As a possible implementation, when inputting the historical traffic data at each traffic location to be predicted in parallel in the same time order, the input historical traffic data of each road can have the same set duration, which is beneficial to improving the prediction accuracy and reducing the complexity of the prediction process.

[0095] The traffic data prediction method provided by the embodiments of this application enables the prediction model to have the functions of multiple expert models through a multi-model layer composed of multiple types of sub-models with independent prediction capabilities. Through the routing layer, the prediction results of each sub-model can be selected and fused, which is equivalent to implementing a mixture of expert models and is beneficial to improving the generalization ability of the prediction model.

[0096] To facilitate the understanding of the embodiments of the present application, the following uses a prediction model with a multi-model layer composed of a time series sub-model, a static spatial sub-model, and a dynamic spatial sub-model as an example for illustration. It should be understood that the embodiments of the present application are not limited to this, and any two of the sub-models can also be equally applicable.

[0097] See Figure 2 as shown Figure 2 This is a schematic diagram of the prediction model of this embodiment. The prediction model includes: a first fully connected layer, a multi-model layer, a second fully connected layer, and a routing layer that are connected in sequence. Among them, the routing layer includes a gating network layer and a routing selection layer. The input end of the gating network layer is respectively connected to the output end of the multi-model layer and the input end of the first fully connected layer, and the output end of the gating network (Gating Network) layer is connected to the selection input end of the routing selection layer.

[0098] The first fully connected layer is used to combine and transform the traffic feature data at each historical time point at each traffic location.

[0099] The multi-model layer is used to perform spatio-temporal prediction based on the traffic feature data output from the first fully connected layer. The multi-model layer includes a time series sub-model, a static spatial sub-model, and a dynamic spatial sub-model. Among them,

[0100] The time series sub-model is used to realize the prediction of traffic data at each traffic location at each future time.

[0101] The static spatial sub-model is used to realize the prediction of traffic data at static traffic locations.

[0102] The dynamic spatial sub-model is used to realize the prediction of traffic data at dynamic traffic locations;

[0103] The static traffic location refers to a static traffic location with fixed attributes, such as road intersections, crossroads, etc. The dynamic traffic location refers to a dynamic traffic location with non-fixed attributes, such as road intersections and crossroads under temporary control or traffic restrictions.

[0104] The above multi-model layer includes various types of sub-models. Each type of sub-model serves as an expert model, responsible for processing a specific part of the input data or a subset of tasks. Through the division of labor and cooperation among the expert models, the performance and efficiency of the model are improved to be applicable to processing complex and diverse traffic data.

[0105] Given that there is a lot of related work on time-series large models, but a single time-series large model ignores spatial information. In a complex large transportation network, spatial information plays a very important role in tasks such as traffic time-series prediction and completion. At the same time, there are also various spatial modeling methods, such as those based on graph neural networks and attention mechanisms. However, the time-series part adopted by these spatial modeling methods during modeling is often a small time-series model. Since the time-series modeling ability and generalization ability of a small time-series model are relatively poor compared to a time-series large model, that is to say, the time-series large model and the small spatial model each have their own advantages and disadvantages. In this embodiment, a variety of spatial modeling methods are combined with the time-series modeling ability of the time-series large model. The multi-model includes a time-series sub-model, a static spatial sub-model, and a dynamic spatial sub-model.

[0106] As an example, the time-series sub-model adopts a time-series large model, which is a deep learning model used to process and analyze time-series data to capture long-term dependencies and dynamic changes in the time series, facilitating future trend prediction, anomaly detection, etc. The time-series sub-model is used to extract the time-series features of traffic data without considering the relationships between traffic positions. That is, to model that the traffic time-series feature of traffic data at any traffic position at the next time point is only related to the traffic time-series feature of the traffic data at this traffic position at the previous time point. In other words, based on the historical time-series features at the previous time point at this traffic position, the future time-series features at the next time point at this traffic position are predicted. Mathematically expressed as:

[0107]

[0108] where, H i (t) represents the traffic time-series mapping of traffic data at traffic position i at the next time t, and H i (t-1) represents the traffic time-series mapping of traffic data at traffic position i at the previous time t - 1. The traffic time-series mapping is used to characterize the time-series feature information of traffic data. The traffic data includes but is not limited to road network index information such as traffic flow information, speed information, and the length of the waiting queue to pass.

[0109] See Figure 3 shown below, Figure 3 This is a schematic diagram of the time-series sub-model modeling in this embodiment. In the figure, circles of different colors represent different traffic positions, and different planes represent different time points.

[0110] The static spatial sub - model is used to extract the static spatial features of traffic data at each traffic location. As an example, the static spatial sub - model adopts a GCN model to extract the static correlations between traffic locations, that is, to model the static spatial features of traffic data at any traffic location at the next time point, which are related to the static spatial features of traffic data at this traffic location at the previous time point and the static spatial features of traffic data at static traffic locations that are statically adjacent to this traffic location. That is to say, based on the historical static spatial features at this traffic location at the previous time point and the static spatial features at its respective static traffic locations, the future static spatial features at this location at the next time point are predicted. Mathematically expressed as:

[0111]

[0112] Among them, \(S\) i (t) represents the static spatial features of traffic data at traffic location \(i\) at the next time \(t\).

[0113] \(S\) i (t-1) represents the static spatial features of traffic data at traffic location \(i\) at the previous time \(t - 1\), \(g()\) represents the static spatial feature mapping, and \(\theta\) j is the static spatial feature of traffic data at static traffic location \(j\) that is statically adjacent to this traffic location \(i\). For example, the 4 sections in an intersection are fixedly distributed in space, and the entrance and exit of a highway are fixed in space.

[0114] See Figure 4 as shown Figure 4 is a schematic diagram of the modeling of the static spatial sub - model in this embodiment. In the figure, circles of different colors represent different traffic locations, different planes represent different time points, black line segments represent the spatial location relationships between the positions represented by the circles at both ends of the line segment, and red line segments represent the associations between the spatial positions represented by the circles at both ends of the line segment at different time points.

[0115] The dynamic spatial sub - model is used to extract the dynamic spatial features of traffic data at each traffic location. As an example, the dynamic spatial sub - model is an attention mechanism model to extract the dynamic correlations between traffic locations, that is, to model the spatial features of traffic data at any traffic location in the road network at the next time point as related to the spatial features of all traffic locations in the road network at the previous time point, and the correlation is obtained through the attention mechanism. That is to say, based on the historical traffic dynamic spatial features at all traffic locations at the previous time, the future traffic dynamic spatial features at this traffic location at the next time are predicted. Mathematically expressed as:

[0116]

[0117] Among them, D i (t) represents the traffic dynamic spatial characteristics of the traffic data at traffic location i at the next time t.

[0118] D J (t-1) represents the traffic dynamic spatial characteristics of the traffic data at all traffic locations J at the previous time t - 1, and Attention() represents the attention mechanism operation.

[0119] See Figure 5 as shown. Figure 5 This is a schematic diagram of the static spatial sub - model modeling in this embodiment. In the figure, circles of different colors represent different traffic locations, different planes represent different time points, black line segments represent the spatial position relationships at the positions represented by the circles at both ends of the line segment, and red line segments represent the associations between the spatial positions at the positions represented by the circles at both ends of the line segment at different time points.

[0120] In the above - mentioned sub - model, the time interval between time t and time t - 1 can be the set prediction time step.

[0121] The second fully - connected layer is used to combine and transform the prediction results from the sub - model.

[0122] The routing layer is used to select the prediction results of each sub - model to output the predicted traffic data. As an example, through the gated network layer, it is used to calculate the probabilities of different sub - models, and the routing selection layer selects each sub - model according to the probability vector and fuses the outputs of each sub - model to achieve the hybrid output of multiple models.

[0123] The routing layer includes a gated network layer and a routing selection layer. The prediction results output by each sub - model in the multi - model layer are respectively input to the gated network layer, the control data output by the gated network layer is input to the routing selection layer, and the traffic feature data of each sub - model output by the second fully - connected layer is input to the routing selection layer as the data to be selected.

[0124] See Figure 6 as shown. Figure 6 This is a schematic diagram of the gated network layer in this embodiment for obtaining control data. By learning the current time - series pattern for characterizing each traffic data pattern with a relatively high recurrence probability searched through the time series, a time - series pattern matrix for characterizing the traffic data pattern is obtained. Each row of this matrix represents a traffic data pattern, a row vector is used to represent the pattern feature vector of a traffic data pattern, each column represents time - series information, that is, each time t, and any element in this matrix represents the eigenvalue of the traffic data pattern where the element is located and the time - series at the column where the element is located. The traffic data pattern can be a characteristic behavior pattern. For example, traffic peak, flat peak, low peak, congestion and other patterns.

[0125] The current input sequence traffic data input to the prediction model is input to a linear weight network to obtain the predicted sequence traffic data of multiple model layers, which can be expressed by the mathematical formula:

[0126]

[0127] where, X i (t) is the current input sequence traffic data at traffic location i at time t, that is, the input data input to the prediction model, W q is the weight of the linear weight network, b q is the offset of the linear weight network, Q i (t) is the predicted sequence traffic data of multiple model layers at traffic location i at time t output by the linear weight network.

[0128] According to the predicted sequence traffic data of multiple model layers and the row vectors of each traffic data pattern in the time series pattern matrix, the correlation between the predicted sequence traffic data of multiple model layers and each traffic data pattern is obtained. Specifically, for any traffic data pattern in the current time series pattern:

[0129] Calculate the exponential function value with the natural constant as the base and the product result of the predicted sequence data of multiple model layers and the transpose of the row vector representing the traffic data pattern as the exponent, to obtain the exponential function value of the traffic data pattern, where the predicted sequence data of multiple model layers is a row vector and the transpose of the row vector of the traffic data pattern is a column vector,

[0130] Calculate the exponential function values with the natural constant as the base and the product results of the predicted sequence data of multiple model layers and the transposes of the row vectors of each traffic data pattern as exponents, to obtain the exponential function values of each traffic data pattern, and accumulate the exponential function values of each traffic data pattern,

[0131] Calculate the ratio of the exponential function value of the traffic data pattern to the accumulated exponential function values of each traffic data pattern, to obtain the correlation degree between the predicted sequence data of multiple model layers and the traffic data pattern.

[0132] It is expressed by the mathematical formula:

[0133]

[0134] where, α m is the correlation degree between the predicted sequence traffic data of multiple model layers and traffic data pattern m, M[m] T is the transpose of the row vector M[m] corresponding to traffic data pattern m in the time series pattern matrix, which is a column vector, Q i (t)is a row vector, and M is the total number of traffic data patterns, i.e., the total number of rows of the time series pattern matrix.

[0135] Perform a weighted sum on each traffic data pattern to obtain a pattern feature representing the traffic data. Specifically, for the relevance of each traffic data pattern, calculate the product of the relevance of the traffic data pattern and the row vector of the traffic data pattern to obtain the weighted row vector of the traffic data pattern. Add up the weighted row vectors of all traffic data patterns to obtain a pattern feature vector. It is expressed by the mathematical formula as:

[0136]

[0137] where O i (t) is the pattern feature vector of the traffic data at traffic location i at time t. This vector is a row vector, and α m M[m] is the weighted row vector of traffic data pattern m.

[0138] Perform a correlation calculation on the pattern feature vector of the traffic data and the prediction results of each sub-model in the multi-model layer to obtain a sub-model similarity representing the similarity degree between the prediction result of each sub-model and the pattern feature vector. It is expressed by the mathematical formula as:

[0139]

[0140] where r e represents the similarity between sub-model e and the pattern feature vector, and z e represents the prediction result of sub-model e, which can be sequence data and can be expressed in vector form. R( ) represents a similarity function, which can be cosine similarity, Euclidean distance, etc.

[0141] Based on the similarities of each sub-model, calculate the output probabilities of each sub-model, and use the calculated output probabilities of each sub-model as the control data output by the gating network layer and input them into the routing selection layer. Specifically, accumulate the similarities of each sub-model to obtain the sum of all similarities; for any sub-model, calculate the ratio of the similarity of the sub-model to the sum of all similarities to obtain the output probability of the sub-model. It is expressed by the mathematical formula as:

[0142]

[0143] where p e is the output probability of sub-model e. In this embodiment, the value of e is 3, that is, the total number of sub-models included.

[0144] It should be understood that the prediction results of the above-mentioned sub-models can be the traffic data output by the hidden layer of the sub-model, or the traffic data of this traffic data output by the output layer of each sub-model in the multi-model layer. This application does not limit this.

[0145] The routing layer selects, at each time t, the prediction result output by the sub-model whose output probability is greater than the set output probability threshold as the predicted traffic data according to the output probabilities of the sub-models. For example, the prediction results output by the k sub-models with the highest output probabilities are selected as the predicted traffic data.

[0146] In this embodiment, the prediction model is trained in the following manner. The historical traffic data is segmented into time series segments of a set sample time length, and the time series segments are used as sample data and input into the prediction model to be trained. In order to find a suitable routing, two types of classification prediction loss function values are used during backpropagation: one is the first type of prediction loss function value for abandoning poor models, and the other is the second type of prediction loss function value for selecting better models, so as to select the output results of the k sub-models with relatively better prediction results.

[0147] As an example, the loss function of the prediction model is calculated in the following manner: The prediction loss function value of the prediction model is calculated according to the classification prediction loss function values of the sub-models, the number of sub-models, and the sample output probabilities of the sub-models. Specifically,

[0148] For any sub-model, the product of the classification prediction loss function value of the sub-model and the logarithm of the sample output probability of the sub-model is calculated to obtain the prediction loss function value of the sub-model, where the logarithm can be the logarithm with base 2.

[0149] The average value of the prediction loss function values of all sub-models is calculated to obtain the prediction loss function value of the prediction model. The mathematical expression is:

[0150]

[0151] where L(P′) represents the loss function value of the prediction model when the sample data within the set sample time step, that is, the sample time interval P', is input, E is the total number of sub-models, p e ′ is the sample output probability of sub-model e, l e is the classification prediction loss function value of sub-model e.

[0152] When using the first type of prediction loss function value, a quantile threshold qth for characterizing the equal-value point within the probability distribution range is set to be greater than 0.5. In this way, there will be the following 4 situations for the sample prediction value accuracy of any sub-model and the output probability of the sub-model.

[0153] (1) The predicted value of the sub-model sample is less than qth different from the true value of the sub-model, and the output probability of the sub-model sample is among the top k.

[0154] (2) The predicted value of the sub-model sample is less than or equal to qth different from the true value of the sub-model, and the output probability of the sub-model sample is not among the top k.

[0155] (3) The predicted value of the sub-model sample is greater than or equal to qth different from the true value of the sub-model, and the output probability of the sub-model sample is among the top k.

[0156] (4) The predicted value of the sub-model sample is greater than qth different from the true value of the sub-model, and the output probability of the sub-model sample is not among the top k.

[0157] Among them, the prediction loss function values of cases (2) and (3) are for the relatively poor models, so their prediction loss function values le are set to 0.

[0158] Based on this, the first type of prediction loss function value of any sub-model can be determined in the following way:

[0159] In the case where the difference between the predicted sample value output by the sub-model and the true value is less than the set first difference threshold, and the sample output probability of the sub-model is greater than the set sample output probability threshold, the first type of prediction loss function value of the sub-model is set to 1.

[0160] In the case where the difference between the predicted sample value output by the sub-model and the true value is greater than the first difference threshold, and the sample output probability of the sub-model is not greater than the sample output probability threshold, the first type of prediction loss function value of the sub-model is set to the reciprocal of the number of sub-models minus 1.

[0161] Otherwise, the first type of prediction loss function value of the sub-model is set to 0;

[0162] Expressed in a mathematical formula as:

[0163]

[0164] Among them, l e is the loss function value of sub-model e, represents the predicted value of the sample of sub-model e and the difference between it and the true value y of the model, is the first difference threshold, and its value is the quantile threshold. top k(P′) represents the k highest sample output probabilities when the sample data within the sample time step, that is, the sample time interval P', is input into the prediction model. p e ′ is the sample output probability of sub-model e.

[0165] When the second type of prediction loss function value is adopted, there are the following four situations for the accuracy of the predicted value of any sub-model sample and the probability of this sub-model.

[0166] (1) The difference between the predicted value of the sub-model sample and the true value of this sub-model is less than 1 - qth, and the sample output probability of this sub-model is among the top k.

[0167] (2) The difference between the predicted value of the sub-model sample and the true value of this sub-model is less than or equal to 1 - qth, and the sample output probability of this sub-model is not among the top k.

[0168] (3) The difference between the predicted value of the sub-model sample and the true value of this sub-model is greater than or equal to 1 - qth, and the sample output probability of this sub-model is among the top k.

[0169] (4) The difference between the predicted value of the sub-model sample and the true value of this sub-model is greater than 1 - qth, and the sample output probability of this sub-model is not among the top k.

[0170] Among them, the prediction loss function values in situations (1) and (4) are for better models, so their prediction loss function values le are not set to 0.

[0171] Based on this, the second type of prediction loss function value of any sub-model can be determined in the following way: For any sub-model,

[0172] When the difference between the predicted sample value output by this sub-model and the true value is less than the second difference threshold and the sample output probability of this sub-model is greater than the sample output probability threshold, the second type of prediction loss function value of this sub-model is set to 1.

[0173] When the difference between the predicted sample value output by this sub-model and the true value is greater than the second difference threshold and the sample output probability of this sub-model is not greater than the sample output probability threshold, the second type of prediction loss function value of this sub-model is set to the reciprocal of the number of sub-models minus 1.

[0174] Otherwise, the second type of prediction loss function value of this sub-model is set to 0. Mathematically expressed as:

[0175]

[0176] Among them, 1 - qth is the second difference threshold.

[0177] See Figure 7 as shown Figure 7 This is a schematic diagram of the prediction model after training in this embodiment for prediction based on the input historical traffic data.

[0178] In this embodiment, the number of traffic positions to be predicted includes N. According to the set historical time step P, the historical traffic data at each traffic position to be predicted is extracted to obtain the historical traffic data at each historical time point within the set historical time step. In the order of increasing time, the historical traffic data at each historical time point at each traffic position to be predicted is sequentially input into the trained prediction model to input the historical sequence traffic data at each traffic position to be predicted. For example, Figure 7 the currently input historical traffic data X P includes the historical traffic data of N traffic positions to be predicted at each time point t - P - 1, t - P, … t. The traffic data can be one-dimensional, such as flow, or multi-dimensional, such as flow and rate.

[0179] In the prediction model:

[0180] The historical sequence traffic data at each traffic position to be predicted is input into the gated network layer in the routing layer and simultaneously input into the first fully connected layer.

[0181] After being combined by the first fully connected layer, the historical sequence traffic data at each traffic position to be predicted is simultaneously input into each sub-model in the multi-model layer.

[0182] Each sub-model respectively outputs the prediction results of each traffic position to be predicted at its prediction time step S. Among them, the prediction results are prediction sequence traffic data. The prediction time steps S at each traffic position to be predicted can be the same or different, and this embodiment does not limit this.

[0183] The prediction results output by each sub-model are respectively input into the routing selection layer in the routing layer through the second fully connected layer and are respectively input into the gated network layer in the routing layer.

[0184] The gated network layer calculates the output probabilities of each sub-model according to the input prediction results and the input historical traffic data and inputs them into the routing selection layer in the routing layer.

[0185] The routing selection layer selects the prediction results output by the sub-models whose output probabilities are greater than the set output probability threshold from the prediction results output by each sub-model according to the output probabilities of each sub-model.

[0186] Further, the prediction model can also fuse the selected prediction results according to the output probabilities of the selected sub-models to obtain the predicted traffic data. For example, Figure 7 the prediction results include the predicted traffic data of N traffic positions to be predicted at each time point t + 1, t + S + 1, …. It should be understood that fusing the selected prediction results can be performed in the routing selection layer or outside the routing layer, and this application does not limit this.

[0187] For ease of understanding this embodiment, the following takes the example of predicting the traffic flow data of a road network with 200 intersections.

[0188] Assume that the time interval for collecting road network traffic flow is 5 minutes, and the prediction task is to predict the future 1-hour traffic flow data of the entire road network based on the past 1-hour traffic flow data of the entire road network. Then, taking the collection time interval as the historical time step and the prediction time step, using the trained prediction model, based on the traffic flow data of 12 historical time steps, predict the traffic flow data of the next 12 prediction time steps.

[0189] The prediction model can be trained in the following way:

[0190] Obtain the historical traffic flow data of the road network in the past several days, such as 7 days. The historical traffic flow data of each intersection is segmented into time series segments with a time length of 2 hours in a sliding window manner as the training sample data of the prediction model. As Figure 7 shown, assume that the number of historical time steps P of the historical sample data is 12, the number of prediction time steps S is 12, the number of intersections N = 200, and the traffic flow data dimension C = 1.

[0191] Input the sample data with a dimension of [12, 200, 1]. First, transform the data through the first fully connected (FCN) layer. The output data of the first FCN layer is respectively sent to each sub-model in the multi-model to obtain the outputs z 1 、z 2 、z 3 .

[0192] Based on the input sample data and the time series pattern matrix M, the routing selection layer can obtain the pattern feature vector O according to formulas (4), (5), and (6); the pattern feature vector O and the z 1~3 output by each sub-model are used to calculate the similarity based on formula (7) to obtain r e , and according to r e the sample output probability p' of each sub-model can be obtained e .

[0193] Based on the sample prediction value y output by each sub-model and the sample output probability p' e of this sub-model, the prediction loss function value can be obtained based on the loss function of the prediction model, and backpropagation is performed to optimize the network parameters.

[0194] For example: Select 2 sub-models, that is, k = 2, and the sample output probabilities p' e of the 3 sub-models are respectively p 1 = 0.5, p 2 = 0.4, p 3=0.1, the differences between the sample prediction values ​​and the actual values ​​of the three sub-models are:

[0195]

[0196] qth Set to 0.7;

[0197] If we abandon the loss function value of the worse model, we have:

[0198] Submodel 1 is case (1), le = 1;

[0199] Submodel 2 is case (3), le = 0;

[0200] Submodel 3 is case (2), le = 0;

[0201] Therefore, sub-models 2 and 3 are worse, and the prediction loss function value of the prediction model is as follows:

[0202]

[0203] If a better model loss function value is selected, then:

[0204] Submodel 1 is case (1), le = 1;

[0205] Submodel 2 is case (3), le = 0;

[0206] Submodel 3 is case (4), le = 1 / 2;

[0207] Therefore, sub-models 1 and 3 are better choices, and the prediction loss function value of the prediction model is as follows:

[0208]

[0209] According to the two prediction loss function values ​​of the prediction model calculated by the two types of prediction loss functions adopted, the model parameters of the prediction model are adjusted until the training is completed to obtain the trained prediction model.

[0210] Based on the trained prediction model, the input data is processed: time series and space are modeled through three sub-models; at each moment, the routing layer is used to select the prediction results output by each sub-model to obtain the best k models to achieve the prediction of the entire road network traffic.

[0211] Specifically, at time t, the traffic flow data of intersection A is predicted, and the predicted traffic flows output by each sub-model are v1, v2, and v3, respectively. The output probabilities of each sub-model are 0.6, 0.37, and 0.03, respectively. Then, the predicted traffic flow at intersection A at time t is:

[0212]

[0213] In this embodiment, a prediction model composed of multiple sub-models is adopted to achieve a better selection of prediction results, which is beneficial to traffic data prediction in different scenarios. Specifically, prediction is carried out through a time-series sub-model, which increases the time-series generalization ability of the prediction model. Through the static spatial sub-model, it is beneficial to obtain regular traffic characteristics. Through the dynamic spatial sub-model, it is beneficial to obtain irregular traffic characteristics. Through the selection of the prediction results of each sub-model by the routing layer, both the fusion of the prediction results of each sub-model is realized and the dependence on the prediction results of a single model is avoided. The time-series large model adopted by the time-series sub-model can capture the long-term dependencies and dynamic changes in the time series, which is beneficial to combining the time-series large model with the spatial small model and enhancing the overall time-series generalization ability of the prediction model. During the training process, when performing backpropagation, the prediction loss function value of the prediction model is calculated through the classification prediction function value, and learning is realized in the way of selecting a better model and avoiding a worse model.

[0214] See Figure 8 as shown in Figure 8 which is a schematic diagram of a traffic data prediction device according to an embodiment of the present application. The device includes:

[0215] A prediction module, configured to use a trained prediction model for traffic data prediction to obtain predicted traffic data at a future time for a traffic location to be predicted based on historical traffic data of the traffic location to be predicted at historical times.

[0216] Wherein,

[0217] The prediction model includes: a multi-model layer composed of at least two of a time-series sub-model for obtaining time-series features of historical traffic data to obtain a first prediction result, a static spatial sub-model for obtaining static spatial features of historical traffic data to obtain a second prediction result, and a dynamic spatial sub-model for obtaining dynamic spatial features of historical traffic data to obtain a third prediction result, and a routing layer for selecting the prediction results output by each sub-model in the multi-model layer.

[0218] See Figure 9 as shown in Figure 9 which is another schematic diagram of a traffic data prediction device or an electronic device for traffic data prediction according to an embodiment of the present application. The device includes: a memory and a processor, the memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of the traffic data prediction method according to an embodiment of the present application.

[0219] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0220] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0221] The embodiments of the present invention also provide a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the traffic data prediction method described in this embodiment are implemented.

[0222] For the embodiments of the device / network-side device / storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, refer to the partial description of the method embodiments.

[0223] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0224] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A traffic data prediction method, characterized in that: The method includes: Using the trained prediction model for traffic data prediction, based on the historical traffic data of the traffic location to be predicted at the historical time, the predicted traffic data of the traffic location to be predicted at the future time is obtained, in, The prediction model includes: a multi-model layer composed of at least two of a time series sub-model for obtaining time series features of historical traffic data to obtain a first prediction result, a static space sub-model for obtaining static space features of historical traffic data to obtain a second prediction result, and a dynamic space sub-model for obtaining dynamic space features of historical traffic data to obtain a third prediction result, and a routing layer for selecting the prediction results output by each sub-model in the multi-model layer; The method of using the trained prediction model for traffic data prediction to obtain predicted traffic data for the traffic location to be predicted at a future time based on historical traffic data of the traffic location to be predicted at a historical time includes: Each sub-model in the multi-model layer obtains the prediction results of each sub-model from the currently input historical sequence data. The routing layer selects the prediction results of each sub-model whose output probability is greater than the set output probability threshold according to the output probability of each sub-model. in, The output probability of the sub-model is used to represent: the proportion of the similarity between the prediction results output by the sub-model and the pattern features of the current historical sequence data in the sum of the similarities between the prediction results output by each sub-model and the pattern features of the current historical sequence data, Pattern features are used to characterize the correlation between the current historical sequence data and the current time series pattern. The current time series pattern is used to characterize each traffic data pattern with a high probability of recurrence that is searched through the time series.

2. The traffic data prediction method according to claim 1, characterized in that: The prediction model is trained as follows: The historical traffic data is divided into time series segments of set sample time length, and the time series segments are used as sample data and input into the prediction model to be trained. The prediction loss function value of the prediction model is calculated according to the number of sub-models, the classification prediction loss function value of each sub-model, and the sample output probability of each sub-model, wherein the classification prediction loss function value includes: a first type of prediction loss function value for abandoning a poor sub-model, and a second type of prediction loss function value for selecting a better sub-model. Back propagation is performed based on the predicted loss function value of the prediction model, Repeat the training until it is completed.

3. The traffic data prediction method according to claim 2, characterized in that: The prediction loss function value of the prediction model is calculated as follows: For any sub-model, calculate the product of the classification prediction loss function value of the sub-model and the logarithm of the sample output probability of the sub-model to obtain the prediction loss function value of the sub-model. Calculate the average of the prediction loss function values ​​of all sub-models to obtain the prediction loss function value of the prediction model; The first type of prediction loss function value is determined as follows: For any sub-model, When the difference between the predicted sample value output by the sub-model and the true value is less than the set first difference threshold, and the sample output probability of the sub-model is greater than the set sample output probability threshold, the first type of prediction loss function value of the sub-model is set to 1. When the difference between the predicted sample value output by the sub-model and the true value is greater than the first difference threshold, and the sample output probability of the sub-model is not greater than the sample output probability threshold, the first type of prediction loss function value of the sub-model is set to the reciprocal of the number of sub-models minus 1. Otherwise, the first-class prediction loss function value of this sub-model is set to 0; The second type of prediction loss function value is determined as follows: For any sub-model, When the difference between the predicted sample value output by the sub-model and the true value is less than the second difference threshold, and the sample output probability of the sub-model is greater than the sample output probability threshold, the second type of prediction loss function value of the sub-model is set to 1. When the difference between the predicted sample value output by the sub-model and the true value is greater than the second difference threshold, and the sample output probability of the sub-model is not greater than the sample output probability threshold, the second type prediction loss function value of the sub-model is set to the reciprocal of the number of sub-models minus 1. Otherwise, the second-category prediction loss function value of this sub-model is set to 0. in, The second difference threshold is the difference between 1 and the first difference threshold. The difference threshold is the quantile threshold used to characterize the equal-division numerical points within the probability distribution range. The output probability threshold is determined according to the number of types of prediction results of the required output sub-model.

4. The traffic data prediction method according to any one of claims 1 to 3, characterized in that: The currently input historical sequence data is input into the trained prediction model in the following manner: according to the set historical time step, the historical traffic data at the traffic location to be predicted is extracted to obtain the historical traffic data at each historical time point of the set historical time step, In chronological order, the historical traffic data at each historical time point is input into the trained prediction model in sequence.

5. The traffic data prediction method according to claim 1, characterized in that: The prediction result is a prediction sequence data composed of the predicted traffic data at each future time point, wherein the prediction time step is set between two adjacent future time points; The output probability is determined as follows: For any sub-model, According to the linear weight and linear offset, the multi-model layer prediction sequence data used to characterize the prediction results of the multi-model layer for the current historical sequence data is calculated. According to the multi-model layer prediction sequence data and the current time series mode, the correlation between the multi-model layer prediction sequence data and each traffic data mode in the current time series mode is calculated, and each traffic data mode in the current time series mode is weighted by the calculated correlation to obtain the pattern characteristics of the current historical sequence data. According to the pattern features of the current historical sequence data and the prediction results output by the sub-model, the similarity between the prediction results output by the sub-model and the pattern features of the current historical sequence data is calculated. According to the pattern characteristics of the current historical sequence data and the prediction results output by each sub-model, the similarity between the prediction results output by each sub-model and the pattern characteristics of the current historical sequence data is calculated to obtain the similarity of each sub-model, and the similarity of each sub-model is accumulated to obtain the sum of each similarity. The ratio of the similarity between the prediction result output by the sub-model and the pattern features of the current historical sequence data to the sum of the similarities is calculated to obtain the output probability of the sub-model.

6. The traffic data prediction method according to claim 5, characterized in that: The step of calculating the multi-model layer prediction sequence data for representing the prediction result of the multi-model layer for the current historical sequence data according to the linear weight and the linear offset includes: The product of the linear weight and the current historical sequence data is added with the linear offset to obtain the multi-model layer prediction sequence data of the current historical sequence data; The step of calculating the correlation between the multi-model layer prediction sequence data and each traffic data pattern in the current time series pattern according to the multi-model layer prediction sequence data and the current time series pattern includes: For any traffic data mode in the current time series mode, Calculate an exponential function value with a natural constant as the base and a product of the multi-model layer prediction sequence data and the transpose of the row vector used to characterize the traffic data pattern as the exponent to obtain the exponential function value of the traffic data pattern, wherein the multi-model layer prediction sequence data is a row vector, the transpose of the row vector of the traffic data pattern is a column vector, the current time series pattern is a time series pattern matrix, each row in the time series pattern matrix corresponds to a traffic data pattern, and each column corresponds to a time series, Calculate the exponential function value with the natural constant as the base and the product result of the transposition of the row vector of each traffic data mode and the multi-model layer prediction sequence data as the exponent, obtain the exponential function value of each traffic data mode, and accumulate the exponential function values ​​of each traffic data mode, The ratio of the exponential function value of the traffic data pattern to the accumulated exponential function values ​​of each traffic data pattern is calculated to obtain the correlation between the multi-model layer prediction sequence data and the traffic data pattern.

7. The traffic data prediction method according to claim 6, characterized in that: The step of weighting each traffic data mode in the current time series mode by using each calculated correlation degree includes: For each traffic data pattern relevance, the product of the traffic data pattern relevance and the row vector of the traffic data pattern is calculated to obtain a weighted row vector of the traffic data pattern. Add the weighted row vectors of each traffic data pattern to obtain a pattern feature vector; The selecting the prediction results of each sub-model whose output probability is greater than the set output probability threshold further includes: According to the output probability of each selected sub-model, the fusion weight is determined. The selected prediction results are weighted and summed using the fusion weights to obtain the predicted traffic data.

8. The traffic data prediction method according to claim 6, characterized in that: The time series sub-model is a time series large model, which is used to model the traffic time series characteristics of traffic data at any traffic location at the next time point, which are only related to the traffic time series characteristics of traffic data at the previous time point at the traffic location. The static space sub-model is a graph neural network model, which is used to model the static space characteristics of traffic data at any traffic location at the next time point, which are related to the static space characteristics of traffic data at the traffic location at the previous time point, and the static space characteristics of traffic data at a static traffic location statically adjacent to the traffic location. The dynamic spatial sub-model is a graph neural network model and an attention mechanism model, which is used to model that the spatial characteristics of traffic data at any traffic location at the next time point are related to the spatial characteristics of all traffic locations at the previous time point, and the spatial feature correlation is obtained through the attention mechanism.

9. A prediction model structure for traffic data prediction, characterized in that: The prediction model structure includes: a multi-model layer consisting of at least two of a temporal sub-model for obtaining temporal features of historical traffic data to obtain a first prediction result, a static spatial sub-model for obtaining static spatial features of historical traffic data to obtain a second prediction result, and a dynamic spatial sub-model for obtaining dynamic spatial features of historical traffic data to obtain a third prediction result, and A routing layer for selecting the prediction results output by each sub-model in the multi-model layer; in, Each sub-model in the multi-model layer obtains the prediction results of each sub-model from the currently input historical sequence data. The routing layer selects the prediction results of each sub-model whose output probability is greater than the set output probability threshold according to the output probability of each sub-model. in, The output probability of the sub-model is used to represent: the proportion of the similarity between the prediction results output by the sub-model and the pattern features of the current historical sequence data in the sum of the similarities between the prediction results output by each sub-model and the pattern features of the current historical sequence data, Pattern features are used to characterize the correlation between the current historical sequence data and the current time series pattern. The current time series pattern is used to characterize each traffic data pattern with a high probability of recurrence that is searched through the time series.

10. A traffic data prediction device, characterized in that: The device includes: The prediction module uses the trained prediction model for traffic data prediction to obtain the predicted traffic data of the traffic location to be predicted at a future time based on the historical traffic data of the traffic location to be predicted at a historical time. in, The prediction model includes: a multi-model layer composed of at least two of a time series sub-model for obtaining time series features of historical traffic data to obtain a first prediction result, a static space sub-model for obtaining static space features of historical traffic data to obtain a second prediction result, and a dynamic space sub-model for obtaining dynamic space features of historical traffic data to obtain a third prediction result, and a routing layer for selecting the prediction results output by each sub-model in the multi-model layer; The method of using the trained prediction model for traffic data prediction to obtain predicted traffic data for the traffic location to be predicted at a future time based on historical traffic data of the traffic location to be predicted at a historical time includes: Each sub-model in the multi-model layer obtains the prediction results of each sub-model from the currently input historical sequence data. The routing layer selects the prediction results of each sub-model whose output probability is greater than the set output probability threshold according to the output probability of each sub-model. in, The output probability of the sub-model is used to represent: the proportion of the similarity between the prediction results output by the sub-model and the pattern features of the current historical sequence data in the sum of the similarities between the prediction results output by each sub-model and the pattern features of the current historical sequence data, Pattern features are used to characterize the correlation between the current historical sequence data and the current time series pattern. The current time series pattern is used to characterize each traffic data pattern with a high probability of recurrence that is searched through the time series.

Citation Information

Patent Citations

  • Short-time traffic flow control method based on deep learning and spatio-temporal data fusion

    CN110827543A