Spatiotemporal data prediction method based on neural wavelet rough differential equation
Patent Information
- Application Number
- CN202211253795.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-13
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-10-13
AI Technical Summary
还存在的问题是这个任务输入的原始数据长度较短,而使用的特征提取建模神经网络过大
[0059]本发明继承了神经受控微分方程训练高效内存利用率、处理缺失观测值的鲁棒性,又展示了处理长时间突变抖动信号的特殊优势,对时序数据趋势特征尤为敏感,能够胜任于建模动态长时间交通流量数据预测。
Smart Images

Figure CN115905779B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the application of intelligent transportation systems in predicting traffic flow data, specifically a spatiotemporal data prediction method based on neural wavelet rough differential equations; it involves the fields of ordinary differential dynamical system modeling and rough path theory. Background Technology
[0002] Traffic flow data is a type of time-series data, continuously recorded by deployed sensors at fixed time intervals. Therefore, for early-stage traffic prediction problems, it is natural to use classic time-series analysis models, among which the autoregressive model is a representative example.
[0003] Integrated Moving Average (ARIMA) and Vector Auto-Regressive (VAR) models are used. These methods rely on the assumption of dynamic linear correlation based on time-series data. Given the complex and non-linear evolution of traffic data, it's unsurprising that they perform poorly in practice. To relax the linear assumption, traditional machine learning-based models such as Support Vector Machines (SVR) and K-Nearest Neighbors (KNN) have been employed. These models offer better results than linear models, but their performance heavily depends on the time-consuming and labor-intensive feature extraction process. To circumvent this problem, recent research has focused on developing traffic prediction methods based on deep neural networks. Many researchers combine time-series processing modules such as Recurrent Neural Networks (RNNs) and Graph Convolutional Neural Networks (GCNs) for spatiotemporal prediction. Compared to traditional machine learning methods, they achieve better performance without relying on human intervention. Despite these achievements, accurate, long-term traffic prediction remains challenging. Another issue is the relatively short length of the original input data for this task, while the feature extraction and modeling neural networks used are often too large. For input data, it contains spatial and temporal information. In most cases, time series modeling and analysis often play a crucial role. Therefore, modeling time series problems is the key to spatiotemporal data prediction.
[0004] In the field of machine learning (ML), the combination of dynamical systems and deep learning has become a topic of interest for the research community. Especially for neural differential equations (NDEs), it demonstrates that neural networks and differential equations are two sides of the same coin. Traditional parametric differential equations are a special case; many popular neural network architectures, such as residual networks and recurrent networks, are discretized. Neural differential equations can provide high-capacity function approximations, exhibit strong prior properties in the model space, handle irregular data, and are highly memory efficient.
[0005] Neural differential equations offer a win-win approach: similar in structure to neural networks, they provide high-capacity function approximation and easy training performance. Their differential equation-like structure provides strong prior knowledge regarding model space, memory efficiency, and theoretical understanding through easily understood and well-tested theoretical literature. Compared to classical differential equation theory, neural differential equations inherently possess unprecedented modeling capabilities. Neural controlled differential equations address the problem in dynamic modeling of time series problems using ordinary neural differential equations, where the solution is determined by its initial values and lacks adjustment based on subsequent observation trajectories. Similar to how ordinary neural differential equations are a continuous mathematical representation of residual networks, neural controlled differential equations are a continuous representation of recurrent neural networks (RNNs). Combined with the conjugate backpropagation computation method, they solve problems such as vanishing and exploding gradients in long sequence models of RNNs.
[0006] Rough path theory developed in the 1990s (Lyons 1998). It studies rough paths, where "rough" means that although the path is continuous, it fluctuates wildly at every point. For example, a path generated by Brownian motion is "rough"—continuous but not differentiable at every point. The core concept in rough path theory is the path signature. This "signature" is a mapping function that transforms the original path information into a set of real numbers. Each real number in the set is calculated from the data points in the original path in different ways, representing a certain geometric feature of the original path. Theoretically, a path signature is "infinite-dimensional." In practical applications, we only use signatures with a finite number of dimensions (i.e., a finite number of real numbers in the set); such signatures are called truncated signatures. Using truncated signatures to replace the data information of the original high-dimensional path is called dimensionality reduction. The amount of information contained in a high-order signature decays according to the factorial of its order. This means that the information contained in a higher-order signature is negligible compared to that in a lower-order signature. Therefore, even if a lower-order truncated signature is used, it can be expected to effectively preserve the information of the original path.
[0007] Wavelet transform (WT) is a novel transform analysis method that inherits and develops the localization idea of short-time Fourier transform while overcoming the shortcomings of window size not changing with frequency. It provides a frequency-varying "time-frequency" window, making it an ideal tool for time-frequency signal analysis and processing. Its main characteristics are its ability to fully highlight certain features of a problem through transformation, its capacity for localized analysis of time (space) and frequency, and its ability to progressively refine signals (functions) at multiple scales through scaling and translation operations. This ultimately achieves time subdivision at high frequencies and frequency subdivision at low frequencies, automatically adapting to the requirements of time-frequency signal analysis. This allows for focusing on arbitrary details of the signal, solving the difficulties of Fourier transform and representing a major breakthrough in scientific methodology since its development. Compared to Fourier transform, wavelet transform is a localized transformation in both time and frequency domains, thus effectively extracting information from signals. Through scaling and translation operations, it performs multi-scale analysis of functions or signals, solving many difficult problems that Fourier transform cannot address. Summary of the Invention
[0008] This invention introduces path signatures from coarse path theory to extend neural controlled differential equations. A path signature is essentially a mapping function that transforms the original path information into a set of real numbers. Each real number in the set is calculated from continuous data points in the original path in different ways, representing a certain geometric feature of the original path. This serves as statistical data describing how the signal drives the differential equation within a small time interval. Wavelet decomposition is used to expand the dimensionality of the path signature. The path signature, aided by wavelet decomposition, significantly improves the local and global representation capabilities of the data flow. Specifically, the low-frequency dominant wave obtained from wavelet decomposition can effectively reflect the trend characteristics of the time series. The calculation process of the path signature involves calculating the path integral between continuous paths. The signatures of the original path and the dominant wave path can represent richer path geometric features. The combination of wavelet transform and path signature improves the feature extraction capability of the original time series. Neural controlled differential equations model the time series as a continuously changing trajectory, better utilizing the timestamp information of the data. Moreover, predictions can be made for arbitrary time points. The vector field in the equation uses a simplified neural network as the data-driven tool, enabling efficient and accurate traffic flow prediction. Introducing an attention mechanism can extract key feature information from the hidden state channels of neurally controlled differential equations. This invention also combines graph neural networks to construct the spatial dependencies between nodes of each road segment in traffic spatiotemporal data. This invention uses ordinary differential equations as its basic framework and simple neural networks in deep learning as data-driven tools. Thus, the basic framework of this invention becomes an ordinary differential dynamical system, whose training and prediction are both reduced to the numerical solution problem of ordinary differential equations.
[0009] This invention is achieved through the following technical solution:
[0010] The spatiotemporal data prediction method based on neural wavelet coarse differential equations is characterized by including four steps: wavelet decomposition to obtain multi-frequency traffic data, signature transformation to calculate path signatures, construction of neural controlled differential equations, solution of neural controlled differential equations and output mapping.
[0011] The prediction method described herein comprises the following specific steps:
[0012] Step 1: Perform wavelet decomposition on the original historical traffic data stream to obtain multiple low-frequency main wave traffic data, thereby expanding the dimensions of the original data. Then, use interpolation to transform the multidimensional discrete data into a continuous data path. (The content after “:” explains the meaning of X. For example, here X is a function that can map any real number R between 0 and T to obtain data with dimension v. For example, if v = 3, X(2) = [2,3,4] and X(3) = [3,4,5], where 2 and 3 are in the range of 0 to T.) is used to provide to step 2.
[0013] Step 2: Calculate the path signature of the continuous data path obtained in Step 1. (The path signature Sig(X) is a mapping of [0,T′]. The path signature is obtained by cutting [0,T] into small time slices, calculating them segment by segment, shortening the original [0,T] to [0,T′]), and taking the Nth order truncated path signature Sig. N (X), for Sig N (X) Obtain the log signature by removing redundant data items. (The change in the log signature reduces the dimension obtained after mapping the values in [0,T′] from v′ to v″, which is then provided to step 3;
[0014] Step 3: Construct two neural controlled differential equations. By adjusting the depth of their respective control signal path signatures and the length of the window (a,b], extract the local and global feature information of the log signatures. Each neural controlled differential equation contains two sub-equations: processing time information (Equation (1)) and spatial information (Equation (2)). The formulas are shown below:
[0015]
[0016]
[0017] Where X(t): t∈[0,T] is the continuous data path of the input, and (a,b] is the range of the path log signature window. It is the path log signature within the window (a,b]. The hidden state is obtained by calculating the time series equation (formula (1)). It is by The hidden state is obtained as the control signal of the controlled differential equation. and The initial value for the hidden state, and These are the vector gradient fields for the temporal and spatial modules, respectively, with an internal structure containing trainable parameters. and The neural network models constitute the gradient modules of the controlled differential equations;
[0018] Combining formulas (1) and (2) above, we obtain the following neural-controlled differential equation, formula (3), where... As a result of ordinary differential equations All spatiotemporal features extracted as neural controlled differential equations are linearly mapped using the output layer:
[0019]
[0020] Step 4: Through the above steps, the model becomes an ordinary differential dynamic system. Its training and prediction can be reduced to the solution process of the ordinary differential equation in formula (3). Formula (4) is used as the input of the ordinary differential equation solver. The following ordinary differential equation structure is trained by the adjoint sensitivity method to obtain the final prediction result.
[0021]
[0022] The upper part of the right side of formula (4) together constitutes dZ(t), and the lower part constitutes dH(t). The ODE solver only needs the initial values y and dy / dx to solve the differential equation.
[0023] The prediction method, specifically the interpolation method in step 1, includes: piecewise cubic Hermite interpolation, Lagrange interpolation, or Newton interpolation.
[0024] The prediction method, step 2: calculating the path signature of the continuous data path obtained in step 1. Assume the input is X = (X 1 ,...,X v ),in Then the Nth order signature of the path in the interval (a, b] The calculation formula is as follows:
[0025]
[0026] in:
[0027]
[0028] And obtain the Nth-order truncated path signature Sig N (X), for Sig N (X) Remove redundant data items to obtain the log signature (LogSig). N (X):[0,T′]→R (v″) ;
[0029] path The set of first-order signatures for path X is:
[0030]
[0031] for The increment at time s, This refers to the increment of path X in the i-th dimension from time a to time t, where t∈[a,b]. It is also a real-valued path, and the second-order signature set of path X is:
[0032]
[0033] X i and X j The increments at time r and time s, That is, the first-order path signature of path X in the i-th dimension [a,s). Let X be the second-order path signature of path X in the i-th dimension [a,t), where t∈[a,b], and it is itself a real-valued path.
[0034] Similarly, the set of N-order paths is obtained by concatenating the set of k-order paths (k∈{1,...,N}) into a new set, resulting in the N-order truncated path signature Sig. N (X):
[0035]
[0036] Since path signatures themselves have a certain amount of data redundancy, log signatures are essentially about removing redundant terms from the calculated signature transformation to obtain a minimal subset that is not unique and has no data representation loss.
[0037] The prediction method, step 3:
[0038] Two neural controlled differential equations are constructed to extract local and global feature information of log signatures by controlling the signature depth and the length of the signature window (a,b]. Each neural controlled differential equation contains two sub-equations: processing time information (see formula (1)) and spatial information (see formula (2)). The formulas are as follows:
[0039]
[0040]
[0041] Where X(t): t∈[0,T] is the continuous data path of the input, and (a,b] is the range of the path log signature window. The hidden state is calculated by the time series equation (see formula (1)). It is by The hidden state is obtained as the control signal of the controlled differential equation. and The initial values of the hidden states are obtained from the initial states of the input sequence through a nonlinear mapping. and These are the vector gradient fields of the temporal module in formula (1) and the spatial module in formula (2), respectively. and The internal structure consists of neural network models, each containing trainable parameters. and and Together, they constitute the gradient module of the controlled differential equation;
[0042] in, The neural network model contains only 2 to 3 fully connected layers, using LeakyReLU as the activation function. The first fully connected layer stores the hidden state at the current time step. The process involves a dimension-up mapping, with an SE (Squeeze-and-Excitation) attention module in the middle that assigns weight parameters to the data features of the hidden layer for feature extraction. Finally, a fully connected layer maps and aligns the hidden layer to ensure that the output performs matrix multiplication with the control signal at the current time step.
[0043] and Similarly, the only difference is the addition of a graph operation layer in the middle, which performs adaptive node embedding representation on the nodes in the graph, obtains the corresponding adjacency matrix through matrix multiplication and transpose, and uses Chebyshev polynomials (ChebyNet) to approximate the graph convolution kernel;
[0044] in, accomplish arrive The nonlinear mapping is used, where v″ is the dimension of the path log signature and h(v) is the dimension of the hidden layer. The first two fully connected layers inside perform dimensionality increase and mapping operations on the input to obtain the hidden layer state A1. A1 is then processed by the SE attention module for local feature extraction to obtain the hidden layer state A2. Finally, the hidden layer A2 is subjected to feature mapping to obtain the output. The specific formula is as follows:
[0045]
[0046]
[0047] A2 = SE Attention (A1) (12)
[0048]
[0049] In the above formula (10), FC stands for full connection, and its subscripts a→b, where a is the input dimension and b is the output dimension. Formula (11) SE Attention The attention module, SE (Squeeze-and-Excitation), focuses on extracting data features from channel information. In the above formula, ψ and σ are activation functions used to implement nonlinear mapping.
[0050] in, accomplish arrive The nonlinear mapping, where v″ is the dimension of the path log signature, and h(v) is the dimension of the hidden layer. The output calculated at the current moment is used as the control signal. First, it goes through the first fully connected layer for dimensionality increase to obtain the hidden layer state B0. Then, graph convolution is performed, where W in formula (15) spatial Here, E represents the spatial correlation weight coefficients between nodes, and E is the adaptive node embedding feature. Finally, the hidden layer state B1 obtained from the graph convolution operation is mapped and aligned through a fully connected layer, and the output result is shown in the following structure:
[0051]
[0052] B1=α(I+φ(σ(E·E T )))B0W spatial (15)
[0053]
[0054] To facilitate training, the above formulas are combined to obtain the following neural-controlled differential equation, where... The result of the ordinary differential equation is output as the prediction result through the output layer:
[0055]
[0056] The prediction method, step 4:
[0057] Through step 3, the framework of the present invention becomes an ordinary differential dynamic system. Its training and prediction can be reduced to the process of solving ordinary differential equations. Formula (4) is used as the input of the ordinary differential equation solver. The following ordinary differential equation structure is trained by the adjoint sensitivity method. The final hidden state of the equation solution is nonlinearly mapped through the output layer to obtain the final prediction result.
[0058]
[0059] This invention inherits the high memory utilization and robustness of neural controlled differential equation training and handling missing observations, and also demonstrates the special advantages of handling long-term abrupt jitter signals. It is particularly sensitive to the trend characteristics of time series data and is capable of modeling dynamic long-term traffic flow data prediction. Attached Figure Description
[0060] Figure 1 This is a system flowchart of the spatiotemporal data prediction method based on neural wavelet rough differential equations of the present invention.
[0061] Figure 2 This is a schematic diagram of the neural wavelet rough differential equation model structure in this invention.
[0062] Figure 3 A geometric intuition for signing the first two paths.
[0063] Figure 4 This is a schematic diagram of discrete wavelet decomposition for traffic flow data.
[0064] Figure 5 This is a traffic flow prediction and evaluation map for node 2 in the PEMS04 dataset during February 16th and 17th. Detailed Implementation
[0065] Spatiotemporal data prediction methods based on neural wavelet rough differential equations, including such Figure 1 The four steps shown are: wavelet decomposition to obtain multi-frequency traffic data, signature transformation to calculate path signatures, construction of neural controlled differential equations, solution of neural controlled differential equations and output mapping.
[0066] Example
[0067] like Figure 2 As shown, this invention, using the traffic dataset PEMS04 as an example, elaborates on the neural wavelet rough differential equation model in this invention, specifically including the following steps:
[0068] Step 1: Perform wavelet decomposition on the original traffic flow data stream, such as... Figure 4 The data stream is decomposed into four low-frequency main waves and four high-frequency sub-waves using four-level wavelet decomposition. Figure 2 The dominant wave obtained from wavelet decomposition has a similar trend to the original data curve. Therefore, it is only necessary to select the low-frequency dominant wave that best reflects the traffic flow trend to expand the dimensions of the original data. Then, the multidimensional discrete data is transformed into a continuous data path through interpolation. A continuous traffic flow data path is constructed from the multidimensional discrete data obtained by piecewise cubic Hermite interpolation. The path X is guaranteed to be quadratically differentiable.
[0069] Step 2: Calculate the data path signature Sig(X) of the continuous traffic flow data path X, and take the 3rd-order truncated path signature Sig. 3 (X), for Sig 3 (X) is transformed to obtain the log signature (LogSig). 3 (X).
[0070] Step 2 principle Calculate the path signature Sig(X) of the continuous data path obtained in step 1, assuming the input is X = (X1 ,...,X v ),in Then the Nth order signature of the path in the interval (a, b] The calculation formula is as follows:
[0071]
[0072] in:
[0073]
[0074] And obtain the Nth-order truncated path signature Sig N (X) (generally N=3), for Sig N (X) Remove redundant data items to obtain the log signature (LogSig). N (X):[0,T′]→R (v″) ;
[0075] path The set of first-order path signatures for path X in the i-th dimension [a,t) is:
[0076]
[0077] for The increment at time s, This refers to the increment of path X in the i-th dimension from time a to time t, where t∈[a,b]. It is also a real-valued path, and the second-order signature set of path X is:
[0078]
[0079] X i and X j The increments at time r and time s, That is, the first-order path signature of path X in the i-th dimension [a,s). Let X be the second-order path signature of path X in the i-th dimension [a,t), where t∈[a,b], and it is itself a real-valued path.
[0080] The second-order signature corresponds to integration over a triangle, which is itself a real-valued path. This can reflect the geometric features of path X in high dimensions. We provide a geometric intuition for the first two layers of the logarithmic signature. Figure 3 These have a natural geometric interpretation. The depth term corresponds to the variation of each channel within the interval, i.e. Figure 3 In the equation, ΔX1 and ΔX2; the two depth terms correspond to the signed region between the chord connecting the endpoint and the path itself, i.e. Figure 3 A in + -A - Higher-order terms correspond to higher-order integrals and iterative regions in higher-dimensional space. Similarly, we can obtain the set of N-order paths. By concatenating the set of k (k∈{1,...,N})-order paths into a new set, we can obtain the N-order truncated path signature Sig. N (X):
[0081]
[0082] Since path signatures themselves have a certain amount of data redundancy, log signatures are essentially about removing redundant terms from the calculated signature transformation to obtain a minimal subset that is not unique and has no data representation loss.
[0083] Step 3: Construct two spatiotemporal graph neural controlled differential equations. Each spatiotemporal graph neural controlled differential equation contains two modules: processing temporal information and spatial information. The formulas are shown below:
[0084]
[0085]
[0086] Where X(t): t∈[0,12] is the continuous input data path, H(t): t∈[0,12] is the hidden state obtained by the timing module, Z(t): t∈[0,12] is the hidden state controlled by H(t), and f(H(t); θ f ) and g(Z(t); θ g These are the vector gradient fields for the temporal and spatial modules, respectively, with an internal structure containing trainable parameters θ. f With θ g The neural network models constitute the gradient modules of the neural controlled differential equations.
[0087] Note: In formulas (1) and (2), H and Z with superscripts mean that the control signal is the path signature of the data. In formulas (17) and (18), the control signal without superscripts is the original data path.
[0088] The original continuous input path of formula (17,18) is updated to a path signature. The updated formulas are shown below:
[0089]
[0090]
[0091] Combining the above formulas, we obtain the following neural-controlled differential equation, which is the main equation of the neural-controlled differential equation:
[0092]
[0093] Step 4: Use formula (3) as Figure 2 The input to the neural controlled differential equation solver is used to train the neural controlled differential equation using the adjoint sensitivity method. The internal structure of the entire neural controlled differential equation is shown below. Figure 2 As shown, the final hidden state obtained from solving the equation is nonlinearly mapped through the output layer to output the prediction result:
[0094]
[0095] In step 1 above, different low-frequency main waves are selected through wavelet discretization to construct a multidimensional path for the traffic flow data path; piecewise cubic Hermite interpolation is used to map the discrete data onto a continuous path, and this path is second-differentiable, allowing the use of any point on the path or its gradient for subsequent calculations. The specific steps are as follows:
[0096] Step 1.1: The traffic flow data path is decomposed using Discrete Wavelet (DWT). A fourth-order wavelet decomposition is performed using the Dobessi wavelet function. The low-frequency dominant wavelet that reflects the overall trend is selected and concatenated with the original data path to construct a multi-dimensional traffic flow data path, such as... Figure 4 As shown.
[0097] Step 1.2: Apply piecewise cubic Hermite interpolation to the original discrete traffic flow data. The interpolation function not only passes through all discrete points but also guarantees that the derivative is equal at each discrete point. Each interpolation polynomial contains four unknowns a, b, c, and d.
[0098] f(x) = a + bx + cx 2 +dx 3 (20)
[0099] In step 3, and These are the vector gradient fields for the temporal and spatial modules, respectively, and their internal structure contains trainable parameters. and The neural network models, which constitute the gradient modules of the controlled differential equations, are constructed as follows:
[0100] Step 3.1: accomplish arrive The nonlinear mapping is defined as follows: where v″ is the dimension of the path log signature and h(v) is the dimension of the hidden layer.
[0101]
[0102]
[0103] A2 = SE Attention (A1) (12)
[0104]
[0105] Step 3.2: accomplish arrive The nonlinear mapping is defined as follows: where v″ is the dimension of the path log signature and h(v) is the dimension of the hidden layer.
[0106]
[0107] B1=α(I+φ(σ(E·E T )))B0W spatial (15)
[0108]
[0109] Finally, the solution obtained from formula (3) The multilayer perceptron outputs 12 dimensions of traffic data.
[0110] This invention introduces path signatures from coarse path theory to extend neural controlled differential equations. The path signature transforms the original path information into a set of real-valued paths. Each real number in the set is calculated from continuous data points in the original path in different ways, representing a certain geometric feature of the original path. This serves as statistical data describing how the signal drives the differential equation within a small time interval. Wavelet decomposition is used as a tool to expand the dimensional information of the path signature. The low-frequency dominant wave obtained by wavelet decomposition can effectively reflect the trend characteristics of the time series, greatly improving the local and global representation capabilities of the data flow. An attention mechanism is introduced to extract key feature information from the hidden state channels of the neural controlled differential equations, while graph neural networks are used to extract features from the spatial information of traffic data. This invention uses ordinary differential equations as its basic framework and simple neural networks from deep learning as the data-driven tool. Thus, the basic architecture of this invention becomes an ordinary differential dynamical system, whose training and prediction are reduced to the numerical solution of ordinary differential equations. The combination of wavelet transform and path signature enhances the feature extraction capability of the original time series, while neural controlled differential equations model the time series as a continuously changing trajectory, making better use of the timestamp information of the data. Moreover, it can make predictions for arbitrary time points during prediction. The vector field in the equation uses a simplified neural network as the data-driven tool, thus achieving efficient and accurate traffic flow prediction. This invention inherits the high memory utilization and robustness to missing observations of neural controlled differential equations, and also demonstrates the special advantages of handling long-term abrupt jitter signals. It is particularly sensitive to the trend characteristics of time series data and is well-suited for modeling dynamic long-term traffic flow data prediction.
[0111] To demonstrate the generality of this example, experiments were conducted on the traffic dataset PEMS04, which contains traffic flow data for 307 road segment nodes over two months, with traffic flow data observed every 5 minutes. The traffic flow for the next hour (12 discrete points) was predicted by inputting the traffic flow data for the past hour (12 discrete points). This invention achieves improvements in mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE) compared to similar models. The prediction results are shown in the figure. Figure 5 Among them, the DCRNN model is a model for predicting traffic data based on a diffuse graph convolutional network and GRU; STG-ODE uses the continuous representation of GCN's ordinary differential equations to increase the depth of GCN, expand the spatial receptive field, and capture deeper spatiotemporal dependencies; the Graph-WaveNet model combines graph convolutional neural networks and diffuse causal convolution to extract spatiotemporal dependency information of traffic data; and the STG-NCDE model uses neural controlled differential equations as the core to build a model for spatiotemporal data prediction.
[0112] Table 1 Comparison of Algorithm Experimental Data:
[0113] DCRNN 23.65 35.02 14.36% STG-ODE 20.84 32.82 13.77% Graph-WaveNet 19.36 31.72 13.31% STG-NCDE 19.21 31.09 12.76% This invention 18.57 30.42 12.28%
[0114] The above content is a further detailed description of the present invention, and it should not be considered that the specific embodiments of the present invention are limited to this. For those skilled in the art, several simple deductions or substitutions can be made without departing from the concept of the present invention, and all of these should be considered to fall within the scope of protection of the invention as defined by the claims of the present invention.
Claims
1. A spatiotemporal data prediction method based on neural wavelet coarse differential equations, characterized in that, It includes four steps: wavelet decomposition to obtain multi-frequency traffic data, signature transformation to calculate path signatures, construction of neural controlled differential equations, solution of neural controlled differential equations and output mapping. The specific steps are as follows: Step 1: Perform wavelet decomposition on the original historical traffic data stream to obtain multiple low-frequency main wave traffic data, thereby expanding the dimensions of the original data. Then, use interpolation to transform the multidimensional discrete data into a continuous data path. This is used to provide information to step 2; Step 2: Calculate the path signature of the continuous data path obtained in Step 1. and take truncation path signature ,right Redundant data items are removed to obtain the log signature. This is used to provide information to step 3; Step 3: Construct two neural controlled differential equations by adjusting the depth and window of their respective control signal path signatures ( The length of the log signature is used to extract local and global feature information. Each neural controlled differential equation contains two sub-equations: one for processing time information and the other for processing spatial information. The formulas are shown below: (1) (2) in It is the continuous path of the input data, [ It is the range of the path log signature window. It is a window [ Internal path log signature, These are the hidden states obtained from time-series equations. It is by The hidden state is obtained as the control signal of the controlled differential equation. and The initial value for the hidden state, and These are the vector gradient fields of the temporal and spatial modules, respectively, with an internal structure containing trainable parameters. and The neural network models constitute the gradient modules of the controlled differential equations; Combining the above formulas (1) and (2), we obtain the following neural-controlled differential equation, formula (3), where As a result of ordinary differential equations All spatiotemporal features extracted from the neural controlled differential equation are linearly mapped using the output layer: (3) Step 4: Through the above steps, the model becomes an ordinary differential dynamic system. Its training and prediction can be reduced to the solution process of the ordinary differential equation in formula (3). Formula (4) is used as the input of the ordinary differential equation solver. The following ordinary differential equation structure is trained by the adjoint sensitivity method to obtain the final prediction result. (4)。 2. The prediction method as described in claim 1, wherein the interpolation method in step 1 is: piecewise cubic Hermitian interpolation, Lagrange interpolation, or Newton interpolation.
3. The prediction method as described in claim 1, wherein step 2: calculates the path signature of the continuous data path obtained in step 1. Assuming the input is ,in The path lies in the interval [ of Rank Signature The calculation formula is as follows: (5) in: (6) and take truncation path signature ,right Log signature is obtained by removing redundant data items. ; path ,path In the Dimensions The set of first-order path signatures is: (7) for exist Increment of time, That is, the path In the Dimensions arrive Time increment, among which It is itself a real-valued path, a path The set of second-order signatures is as follows: (8) , They are respectively and exist Time and Increment of time, That is, the path In the Dimensions First-order path signature, For path In the Dimensions The second-order path signature, where It is itself a real-valued path; The second-order signature corresponds to integration over a triangle, which is itself a real-valued path. This reflects the path Geometric features in high dimensions; Similarly, A set of path orders, including paths exist On The sets of path signatures are concatenated into a new set to obtain the path. exist[ On truncation path signature : (9) Since path signatures themselves have a certain amount of data redundancy, log signatures are essentially about removing redundant terms from the calculated signature transformation to obtain a minimal subset that is not unique and has no data representation loss.
4. The prediction method as described in claim 1, wherein step 3: Two neurally controlled differential equations are constructed to control the signature depth and signature window. The length of the log signature is used to extract local and global feature information. Each neural controlled differential equation contains two sub-equations: one for processing temporal information and the other for processing spatial information. The formulas are shown below: (1) (2) in It is the continuous path of the input data. It is the range of the path log signature window. These are the hidden states obtained from time-series equations. It is by The hidden state is obtained as the control signal of the controlled differential equation. and The initial values of the hidden states are obtained from the initial states of the input sequence through a nonlinear mapping. and These are the vector gradient fields of the temporal module in formula (1) and the spatial module in formula (2), respectively. and The internal structure consists of neural network models, each containing trainable parameters. and , and Together, they constitute the gradient module of the controlled differential equation; in, The neural network model contains only 2-3 fully connected layers, using LeakyReLU as the activation function. The first fully connected layer stores the hidden state at the current time step. The process involves a dimension-up mapping, with an SE attention module in the middle that assigns weight parameters to the data features of the hidden layer for feature extraction. Finally, a fully connected layer maps and aligns the hidden layer to ensure that the output performs matrix multiplication with the control signal at the current time step. and Similarly, the only difference is that a graph operation layer is added in the middle. This layer performs adaptive node embedding representation on the nodes in the graph, obtains the corresponding adjacency matrix through matrix multiplication and transpose, and uses Chebyshev polynomials to approximate the graph convolution kernel. in, accomplish arrive The nonlinear mapping, where It is a dimension of path log signature. It is the dimension of the hidden layer; the first two fully connected layers inside perform dimensionality-upgrading and mapping operations on the input. The hidden layer state is obtained through mapping of the fully connected layer. , After passing through the SE channel attention module, i.e. Local feature extraction is performed to obtain the hidden layer state. Finally, the hidden layer The output is obtained by feature mapping, and the specific formula is as follows: (10) (11) (12) (13) in Indicates the implementation of the fully connected layer arrive linear mapping, and For activation function, accomplish arrive The nonlinear mapping, where It is a dimension of path log signature. It is the dimension of the hidden layer. Using the calculated output at the current moment as the control signal, the signal first passes through the first fully connected layer for dimensionality increase, resulting in the hidden layer state. Then perform graph convolution operation, where in formula (15) The spatial correlation weighting coefficient between nodes. To adapt node embedding features, the hidden layer states obtained from graph convolution operations are finally processed. After mapping and alignment by a fully connected layer, the output result has the following structure: (14) (15) (16) Combining the above formulas, we obtain the following neural-controlled differential equation, where The result of the ordinary differential equation is output as the prediction result through the output layer: (3)。 5. The prediction method as described in claim 1, wherein step 4: Through step 3, it becomes an ordinary differential dynamic system. Its training and prediction can be reduced to the solution process of ordinary differential equations. Formula (4) is used as the input of the ordinary differential equation solver. The following ordinary differential equation structure is trained by the adjoint sensitivity method. The final hidden state of the equation solution is nonlinearly mapped through the output layer to obtain the final prediction result. (4)。
Citation Information
Patent Citations
Deep neural network robust traffic prediction method based on multi-modal spatio-temporal data
CN112289034A
Space-time differential equation network for urban flow prediction
CN115048852A