Traffic flow prediction method and system based on interpretable graph and multi-scale temporal backbone

By employing a traffic flow prediction method based on interpretable graphs and multi-scale time backbones, and utilizing graph neural networks and an improved PatchTST model, the problem of insufficient accuracy in existing traffic flow prediction technologies is solved, achieving more accurate and stable traffic flow prediction, which is applicable to traffic control and travel guidance.

CN121456830BActive Publication Date: 2026-04-07湖南工商大学
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing traffic flow prediction methods are insufficient in terms of accuracy and interpretability, and are difficult to effectively capture the dynamic topological relationships and multi-scale features of traffic flow. This makes it difficult to balance prediction accuracy and computational efficiency, especially in long-term extrapolation where there are problems of unstable results and bias.

Method used

A traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones is adopted. The spatial features of traffic flow are extracted by graph neural networks, and the temporal features of multiple time scales are captured by an improved PatchTST model. The Koopman operator is combined to handle long-term dependencies, and a composite prediction model is constructed to improve prediction accuracy.

Benefits of technology

It achieves more accurate and reliable prediction of traffic flow, suppresses false convergence and noise amplification, improves the interpretability of the model and the stability of long-term prediction, and is applicable to scenarios such as traffic control and travel guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456830B_ABST
    Figure CN121456830B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data analysis, and particularly relates to a traffic flow prediction method and system based on an interpretable graph and a multi-scale time backbone. The method comprises: processing historical traffic flow information of a target road network according to a prediction model to obtain traffic flow prediction information of the target road network in a target period; the prediction model is a composite model formed by fusing a spatial backbone network and a time backbone network, the spatial backbone network comprises a graph neural network, the time backbone network comprises an improved PatchTST model, and the improved PatchTST model comprises multi-time scale patches, the multi-time scale patches comprising: scale patches corresponding to short-term fluctuations, scale patches corresponding to intraday rhythms, and scale patches corresponding to cross-day trends. The present application can combine the real constraints of each node in the road network, fully mine the complex time sequence characteristics of the historical traffic flow information of the road network in each time dimension, and obtain more accurate and reliable traffic flow prediction information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data analysis, and in particular to a traffic flow prediction method and system based on an interpretable graph and a multi-scale time backbone. BACKGROUND

[0002] With the continuous improvement of urban motorization level and the diversification of travel modes, the urban road network traffic operation state is increasingly complex, showing the characteristics of "morning and evening peak, holiday tidal flow, and temporary event interference". Traffic state is influenced by weather, accidents, construction, activities and other factors, and has obvious dynamic and non-stationary characteristics, resulting in frequent changes and strong coupling of road network operation conditions.

[0003] In the prior art, short-term traffic flow prediction methods mainly include ARIMA-based statistical model methods, support vector machine or random forest methods based on machine learning, and recurrent neural network, spatio-temporal graph neural network and other methods based on deep learning. Among them, the adaptability of statistical and traditional machine learning methods to non-linear and time-varying relationships is insufficient; although deep learning methods can capture complex spatio-temporal dependencies, they generally rely on fixed or implicit learning of adjacency matrices, making it difficult to reflect the dynamic topological relationship of traffic flow changes over time and events. The prediction model based on graph neural network mostly represents the connection between nodes with "correlation weight", lacks the constraint of traffic flow directionality and conservation, and is prone to false flow convergence or diffusion; in time modeling, a single time scale or global attention mechanism is mostly used, which cannot balance the prediction accuracy and computational efficiency, and cannot take into account the multi-scale characteristics of short-term fluctuations, daily cycles and sudden events. At the same time, some models have the problem of unstable or biased results in long-term extrapolation, and lack of interpretability, which is not conducive to application in traffic control and travel induction scenarios.

[0004] That is, the prior art has low prediction accuracy for traffic flow. SUMMARY

[0005] The purpose of the present application is to provide a traffic flow prediction method and system based on an interpretable graph and a multi-scale time backbone, to solve the technical problem of low prediction accuracy for traffic flow in the prior art.

[0006] In a first aspect, an embodiment of the present application provides a traffic flow prediction method based on an interpretable graph and a multi-scale time backbone, the method comprising:

[0007] The historical traffic flow information of the target road network is processed according to a pre-trained prediction model to obtain the traffic flow prediction information of the target road network during the target time period. The prediction model is a composite model formed by fusing a spatial backbone network and a temporal backbone network. The spatial backbone network includes a graph neural network, and the temporal backbone network includes an improved PatchTST model. The improved PatchTST model includes multiple time-scale patches, which include: a first time-scale patch corresponding to short-term fluctuations, a second time-scale patch corresponding to intraday rhythms, and a third time-scale patch corresponding to cross-day trends.

[0008] The historical traffic flow information is used to indicate the traffic flow of each node in the target road network during a historical period; the traffic flow prediction information is used to indicate the traffic flow of each node in the target road network during a target period.

[0009] In one embodiment, the patch length of the first timescale patch is less than the patch length of the second timescale patch, and the patch length of the second timescale patch is less than the patch length of the third timescale patch.

[0010] The patch step size of the first timescale patch is smaller than the patch step size of the second timescale patch, and the patch step size of the second timescale patch is smaller than the patch step size of the third timescale patch.

[0011] In one embodiment, in the multi-timescale patch, the embedding vector corresponding to each timescale patch includes a first time-domain feature and a second time-domain feature, wherein the first time-domain feature is used to indicate the hour position of the corresponding time within a day, and the second time-domain feature is used to indicate the day position of the corresponding time within a week.

[0012] In one embodiment, in the multi-timescale patch, the self-attention corresponding to each timescale patch includes a collaborative mask, wherein the collaborative mask is jointly determined based on a causal mask indicating that access to the future is prohibited, a local mask indicating that the receptive field is within a set window width, and an anchor mask indicating that the connection of remote anchors satisfies a set time interval.

[0013] In one embodiment, the spatial backbone network further includes a Koopman operator, the input of which is connected to the output of the graph neural network.

[0014] In one embodiment, the output data of the prediction model is obtained by weighted calculation of the spatial prediction data corresponding to the Koopman operator and the temporal prediction data corresponding to the improved PatchTST model.

[0015] In one embodiment, the spatial backbone network further includes an edge weight scoring module, which is used to obtain the edge weights between different nodes. The edge weights are determined based on the differences in traffic flow between different nodes in the target road network. The edge weights are used to indicate the probability that the traffic flow of the source node is allocated to the target node. The source node and the target node are any two different nodes in the target road network, and the sum of all edge weights corresponding to the same source node is 1.

[0016] In one embodiment, the edge weights are determined based on the differences in traffic flow between different nodes in the target road network, and the loss function corresponding to the spatial backbone network includes an inflow-outflow conservation deviation term. The inflow-outflow conservation deviation term is determined based on the average value of the inflow-outflow intensity deviations of each node in the target road network. The inflow-outflow intensity deviation is the absolute difference between the inflow intensity and the outflow intensity of a node. The inflow intensity is the sum of all edge weights with the corresponding node as the target node, and the outflow intensity is the sum of all edge weights with the corresponding node as the source node.

[0017] In one embodiment, the edge weight is determined based on the traffic flow difference direction component and traffic flow difference magnitude component between the corresponding source node and target node;

[0018] Wherein, when the traffic flow of the corresponding source node is greater than the traffic flow of the corresponding target node, the traffic flow difference direction component is positively correlated with the corresponding edge weight; when the traffic flow of the corresponding source node is less than or equal to the traffic flow of the corresponding target node, the traffic flow difference direction component is not related to the calculation of the corresponding edge weight.

[0019] The magnitude component of the traffic flow difference is negatively correlated with the corresponding edge weight.

[0020] Secondly, another embodiment of the present invention also provides a traffic flow prediction system based on interpretable graphs and multi-scale temporal backbones, the system comprising:

[0021] The prediction module is used to process the historical traffic flow information of the target road network according to the pre-trained prediction model to obtain the traffic flow prediction information of the target road network in the target time period. The prediction model is a composite model formed by fusing a spatial backbone network and a temporal backbone network. The spatial backbone network includes a graph neural network, and the temporal backbone network includes an improved PatchTST model. The improved PatchTST model includes multiple time-scale patches, which include: a first time-scale patch corresponding to short-term fluctuations, a second time-scale patch corresponding to intraday rhythms, and a third time-scale patch corresponding to cross-day trends.

[0022] The historical traffic flow information is used to indicate the traffic flow of each node in the target road network during a historical period; the traffic flow prediction information is used to indicate the traffic flow of each node in the target road network during a target period.

[0023] Thirdly, in another embodiment of the present invention, an electronic device is provided, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described in the first aspect.

[0024] Fourthly, in another embodiment of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0025] The present invention has the following beneficial effects:

[0026] This invention utilizes a graph neural network as the core of the spatial backbone network to accurately extract the spatial characteristics of traffic flow at each node in the target road network over historical periods, and constructs corresponding real-world constraints accordingly. Based on this, an improved PatchTST model is used as the core of the temporal backbone network, incorporating multi-timescale patches corresponding to short-term fluctuations, intraday rhythms, and cross-day trends. This comprehensively captures the intertwined temporal features of historical traffic flow information across various time dimensions, thus summarizing the traffic flow temporal features of each node in the target road network over historical periods. By fusing the prediction results based on spatial traffic flow features with those based on temporal traffic flow features, a more accurate and reliable traffic flow prediction information is ultimately formed. Specifically, the graph neural network establishes real-world constraints that match the spatial dependencies between nodes in the target road network, suppressing spurious convergence and noise amplification. The multi-timescale patch configuration fully exploits the intertwined temporal features of historical traffic flow information across various time dimensions, suppressing error accumulation and numerical drift from multi-step extrapolation. Attached Figure Description

[0027] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a flowchart illustrating a traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones provided in an embodiment of the present invention.

[0029] Figure 2 This is a schematic diagram of the framework structure of a prediction model provided in an embodiment of the present invention;

[0030] Figure 3 This is a schematic diagram of a modeling structure for an interpretable road network map provided in an embodiment of the present invention;

[0031] Figure 4 This is a model framework diagram of an improved PatchTST model provided in an embodiment of the present invention;

[0032] Figure 5 This is a schematic diagram of the structure of a traffic flow prediction system based on interpretable graphs and multi-scale time backbone provided in an embodiment of the present invention;

[0033] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0034] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0036] The following description, in conjunction with the accompanying drawings, details a specific scheme for a traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones provided by this invention.

[0037] In one embodiment, the present invention provides a traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones, such as... Figure 1 As shown, the method includes:

[0038] Step S1: Process the historical traffic flow information of the target road network according to the pre-trained prediction model to obtain the traffic flow prediction information of the target road network in the target time period.

[0039] The prediction model is a composite model formed by integrating a spatial backbone network and a temporal backbone network. The spatial backbone network includes a graph neural network, and the temporal backbone network includes an improved PatchTST model. The improved PatchTST model includes multiple time-scale patches, which include a first time-scale patch corresponding to short-term fluctuations, a second time-scale patch corresponding to intraday rhythms, and a third time-scale patch corresponding to cross-day trends.

[0040] The historical traffic flow information is used to indicate the traffic flow of each node in the target road network during a historical period; the traffic flow prediction information is used to indicate the traffic flow of each node in the target road network during a target period.

[0041] It should be understood that, in this embodiment, the only difference between the improved PatchTST model and the standard PatchTST model is the setting of the scale patch. The improved PatchTST model uses multiple time-scale patches, while the standard PatchTST model uses a single time-scale patch.

[0042] In traffic flow prediction scenarios, traffic flow exhibits complex variations over time. For example, sudden changes in traffic flow can occur within short periods (e.g., within minutes), and significant differences can exist between different times of the day (e.g., traffic flow differences between morning / evening rush hours and early morning hours), as well as between different dates within a week (e.g., traffic flow differences between weekdays and weekends). Single-time-scale patches are insufficient to cover the temporal characteristics of traffic flow across different time dimensions. Therefore, this invention proposes multi-time-scale patches corresponding to short-term fluctuations, intraday rhythms, and cross-day trends to comprehensively cover the temporal characteristics of traffic flow across different time dimensions and improve prediction accuracy.

[0043] It should be noted that the number of first-timescale patches, second-timescale patches, and third-timescale patches can be one or more. Users can adaptively determine the number of patches for each timescale according to actual needs while balancing the computational complexity of patch processing with the comprehensiveness of time dimension coverage. It should also be noted that the patch length and / or patch step size of timescale patches belonging to different categories (first-timescale patches, second-timescale patches, and third-timescale patches) are different, and the patch length and / or patch step size of different timescale patches belonging to the same category are also different.

[0044] In this invention, both the spatial backbone network and the temporal backbone network generate corresponding prediction information. The traffic flow prediction information is obtained by fusing the prediction information output by the spatial backbone network and the prediction information output by the temporal backbone network.

[0045] This invention utilizes a graph neural network as the core of the spatial backbone network to accurately extract the spatial characteristics of traffic flow at each node in the target road network over historical periods, and constructs corresponding real-world constraints accordingly. Based on this, an improved PatchTST model is used as the core of the temporal backbone network, incorporating multi-timescale patches corresponding to short-term fluctuations, intraday rhythms, and cross-day trends. This comprehensively captures the intertwined temporal features of historical traffic flow information across various time dimensions, thus summarizing the traffic flow temporal features of each node in the target road network over historical periods. By fusing the prediction results based on spatial traffic flow features with those based on temporal traffic flow features, a more accurate and reliable traffic flow prediction information is ultimately formed. Specifically, the graph neural network establishes real-world constraints that match the spatial dependencies between nodes in the target road network, suppressing spurious convergence and noise amplification. The multi-timescale patch configuration fully exploits the intertwined temporal features of historical traffic flow information across various time dimensions, suppressing error accumulation and numerical drift from multi-step extrapolation.

[0046] In some embodiments, the patch length of the first timescale patch is less than the patch length of the second timescale patch, and the patch length of the second timescale patch is less than the patch length of the third timescale patch.

[0047] The patch step size of the first timescale patch is smaller than the patch step size of the second timescale patch, and the patch step size of the second timescale patch is smaller than the patch step size of the third timescale patch.

[0048] In this embodiment, based on the above settings, the confusion of the "field of vision" corresponding to patches of different time scales can be avoided, and the adaptive allocation of limited computing resources can be achieved. Under the premise of ensuring that the temporal feature information is fully extracted, unnecessary computing overhead is minimized, while enhancing the interpretability of multi-time scale patch settings.

[0049] In some embodiments, in the multi-timescale patch, the embedding vector corresponding to each timescale patch includes a first time-domain feature and a second time-domain feature, wherein the first time-domain feature is used to indicate the hour position of the corresponding time within a day, and the second time-domain feature is used to indicate the day position of the corresponding time within a week.

[0050] In this embodiment, based on the settings of the first time-domain features and the second time-domain features, time semantics (indicating the time-domain position of the data) can be added to the embedding vector corresponding to each time-scale patch to support data alignment across days and differentiation of intra-week patterns. This helps the time backbone network extract intra-day rhythms and cross-day trends from historical traffic flow information more quickly and accurately, thereby improving the accuracy and reliability of the prediction information output by the time backbone network.

[0051] In some embodiments, in the multi-timescale patch, the self-attention corresponding to each timescale patch includes a collaborative mask, wherein the collaborative mask is jointly determined based on a causal mask indicating that access to the future is prohibited, a local mask indicating that the receptive field is within a set window width, and an anchor mask indicating that the connection of remote anchors satisfies a set time interval.

[0052] In this embodiment, the causal mask is set to ensure that each patch can only focus on its own and previous patch information during prediction, and cannot "peek" at the information of subsequent patches. As for the local mask, it further restricts each patch to interact with other patches within a set window width during prediction, based on the causal mask. This ensures that each patch can focus on and extract the temporal features at the corresponding time scale to the maximum extent, while reducing the computational complexity during patch calculation. The anchor mask is set as a mechanism to supplement the local mask, allowing each patch to link to remote anchors at a set time interval. This allows for the conditional and sparse introduction of key long-term dependencies to fully extract intraday and intraweek rhythmic features.

[0053] In some embodiments, the spatial backbone network further includes a Koopman operator, the input of which is connected to the output of the graph neural network.

[0054] In this embodiment, to address the issue that graph neural networks can only capture the spatial local dependencies and short-term temporal dependencies of nodes in a target road network, the Koopman operator is introduced and connected to the output of the graph neural network. This maps the spatiotemporal features extracted by the graph neural network into the Koopman space, thereby achieving dimensionality enhancement of the spatiotemporal features output by the graph neural network. The exponentiation operation of linear operators is used to replace complex nonlinear iterative predictions, avoiding the problem of error accumulation caused by the lack of long-term temporal dependencies in long-term predictions of graph neural networks. This helps the graph neural network to effectively capture long-term dependency information and stabilize the modeling of long-term trends, based on capturing local structures and short-term dependencies, thereby improving the accuracy of the prediction information output by the spatial backbone network.

[0055] In the above setup, the graph neural network is used to process the nonlinear relationship of traffic flow in the spatial dimension of each node in the target road network, while the Koopman operator is used to process the nonlinear relationship of traffic flow in the temporal dimension of each node. The two work together to achieve accurate prediction of future traffic flow in the target road network.

[0056] Furthermore, the output data of the prediction model is obtained by weighted calculation of the spatial prediction data corresponding to the Koopman operator and the temporal prediction data corresponding to the improved PatchTST model.

[0057] In this context, the spatial prediction data corresponding to the Koopman operator can be understood as the prediction data output by the spatial backbone network, and the temporal prediction data corresponding to the improved PatchTST model can be understood as the prediction data output by the temporal backbone network.

[0058] For example, the output data of the prediction model It can be represented as:

[0059]

[0060] in, This represents the spatial prediction data corresponding to the Koopman operator. This represents the time prediction data corresponding to the improved PatchTST model. The gating coefficient is used to control the degree to which the final output depends on long-term trends and short-term changes, so that the prediction model can adaptively choose the degree to which it depends on the spatial backbone network and the degree to which it depends on the temporal backbone network according to the actual training situation.

[0061] In some embodiments, the spatial backbone network further includes an edge weight scoring module, which is used to obtain the edge weights between different nodes. The edge weights are determined based on the differences in traffic flow between different nodes in the target road network. The edge weights are used to indicate the probability that the traffic flow of the source node is allocated to the target node. The source node and the target node are any two different nodes in the target road network, and the sum of all edge weights corresponding to the same source node is 1.

[0062] In this embodiment, based on the setting of the edge weight scoring module, the probability of edge weights is further introduced on the basis of the spatial constraints indicated by the adjacency matrix of the target road network, so as to adapt to the coupling effect of traffic flow changes between different nodes in the target road network. This allows the final modeling of the spatial backbone network to more accurately simulate the actual traffic flow changes of the target road network, thereby improving the accuracy of the predicted data output by the spatial backbone network.

[0063] Furthermore, the edge weights are determined based on the differences in traffic flow between different nodes in the target road network. The loss function corresponding to the spatial backbone network includes an inflow-outflow conservation deviation term. The inflow-outflow conservation deviation term is determined based on the average value of the inflow-outflow intensity deviation of each node in the target road network. The inflow-outflow intensity deviation is the absolute difference between the inflow intensity and the outflow intensity of a node. The inflow intensity is the sum of all edge weights with the corresponding node as the target node, and the outflow intensity is the sum of all edge weights with the corresponding node as the source node.

[0064] During the actual traffic flow changes in the target road network, the vehicles merging into the target road network and the vehicles merging out of the target road network remain consistent. Based on this prior knowledge, an entry and exit conservation deviation term is constructed accordingly to provide realistic constraints for the entire road network. This suppresses the probability of the spatial backbone network learning false knowledge during the training phase and improves the prediction accuracy of the final trained spatial backbone network.

[0065] It should be further explained that, in real-world scenarios, the number of vehicles merging into and out of the target road network is usually nearly identical. However, since the target road network is not an ideal closed system but only a part of a complex real-world road network, it is affected by external factors (such as ramp entry and exit). Therefore, the entry / exit conservation deviation term can be set to have a smaller impact on the loss function of the spatial backbone network. This adapts to the situation where the number of vehicles merging into and out of the target road network is nearly identical in real-world scenarios, while avoiding overfitting of the spatial backbone network to meet real-world constraints oriented towards the entire road network. This further improves the prediction accuracy of the final trained spatial backbone network. Furthermore, the edge weights are determined based on the traffic flow difference direction component and traffic flow difference amplitude component between the corresponding source node and target node.

[0066] Wherein, when the traffic flow of the corresponding source node is greater than the traffic flow of the corresponding target node, the traffic flow difference direction component is positively correlated with the corresponding edge weight; when the traffic flow of the corresponding source node is less than or equal to the traffic flow of the corresponding target node, the traffic flow difference direction component is not related to the calculation of the corresponding edge weight.

[0067] The magnitude component of the traffic flow difference is negatively correlated with the corresponding edge weight.

[0068] The traffic flow difference direction component is used to indicate the degree of deviation between the traffic flow difference between the corresponding source node and the target node and zero. When the traffic flow difference between the corresponding source node and the target node is positive, the greater the deviation between the traffic flow difference between the corresponding source node and the target node and zero, the larger the traffic flow difference direction component is, and the more necessary it is to increase the value of the corresponding edge weight to strengthen the directionality of the edge weight content in combination with the actual traffic flow situation. When the traffic flow difference between the corresponding source node and the target node is negative, the corresponding traffic flow difference direction component is fixed to the minimum value (e.g., 0).

[0069] The traffic flow difference amplitude component is the absolute value of the traffic flow difference between the corresponding source node and the target node.

[0070] The above-mentioned traffic flow difference direction component, traffic flow difference amplitude component, and corresponding constraint conditions can effectively combine the actual traffic flow changes in the target road network to accurately evaluate the edge weights between different nodes.

[0071] Specifically, when the traffic flow of the corresponding source node is greater than the traffic flow of the corresponding target node, it can be assumed that the traffic flow of the source node in the historical period may overflow to the target node, thereby causing a corresponding change in the traffic flow of the target node. However, it should be noted that in the case of traffic flow overflow, the difference in traffic flow between the source node and the target node is usually small. Therefore, when the directional component of the traffic flow difference is larger, it can be assumed that the possibility of matching vehicle overflow between the source node and the target node is lower. Therefore, it is necessary to reduce the value of the corresponding edge weight to combine the actual characteristics of vehicle overflow to suppress the similarity of edge weight content and suppress the problem of false strong edges caused by noise or other factors.

[0072] If the traffic flow of the corresponding source node is less than or equal to the traffic flow of the corresponding target node, it can be considered that the possibility of the traffic flow of the source node overflowing to the target node in the historical period is low. In this case, the value of the traffic flow difference direction component cannot affect the probability of the traffic flow of the source node being allocated to the target node (the probability should be zeroed). As for the larger the traffic flow difference magnitude component, the greater the risk that the source node and the target node may have false strong connections due to noise or other factors. In combination with the situation of the traffic flow difference direction component, it is more necessary to reduce the edge weight from the source node to the target node.

[0073] For ease of understanding, the following example is provided:

[0074] In this example, traffic flow data is organized in 15-minute time granularities, defining time-stamp vectors. Indicates a time period of the day. Indicates the day of the week. Indicate whether it is a working day The time attribute at any given moment The input features are incorporated into the vector form.

[0075] To characterize the spatial dependencies of nodes in the target road network, a directed graph is constructed. :in Each road segment constitutes a node set. The set of edges is formed by the upstream and downstream connections and the convergence relationship. Adjacency matrix .

[0076] Directed definition is adopted (if road segment) For the section If there is an impact in the next moment, then (Otherwise, it is 0). When available, supplement with a distance or odometer matrix. Generate initial weights .

[0077] In this invention, road and geometric priors are used to filter out nodes whose nodes do not exist in the target road network, and the corresponding element values ​​in the adjacency matrix are set to 0.

[0078] To characterize the static properties of nodes, a space vector is defined. ( The vector represents the F-dimensional features of the i-th node, including but not limited to direction, number of lanes, road and corridor markings, mileage, and positional order. This vector, along with the temporal attributes, serves as the node-side prior input during training.

[0079] In this example, the prediction problem can be defined as: at the current time... , before Historical observations at each time step (This can be understood as the aforementioned historical traffic flow information.) Used to indicate the current state of each node in the target road network. Traffic flow), combined with time attributes Space vectors Adjacency matrix Predicting the traffic flow at each node in the target road network at future times. To simultaneously cover both single-step and multi-step prediction capabilities. To verify the model's temporal adaptability, this paper... The evaluation is conducted within minutes, corresponding to Use the most frequently used settings. Evaluation based on... The primary focus is on stability and portability, while ensuring synchronized output across the entire network.

[0080] This example uses the NV36 dataset, which contains minute-level data from 36 sensors along two highways in Washington, D.C., Northern Virginia, sampled every 15 minutes. The goal is to predict traffic volume at each sensor 15 minutes later, given the spatial connectivity of the road network.

[0081] The input features of a single sample are a 36×48 matrix: containing historical traffic flow, weekday, hour, road direction, number of lanes and road name of nearly 10 sampling points, totaling 48 dimensions; the training and testing time slices are 1261 and 840 consecutive 15-minute segments, respectively, and the corresponding output is the traffic flow labels of 36 locations in the corresponding time slice. The specific description of the original data is shown in Table 1.

[0082] Table 1 Original MAT Data Fields

[0083]

[0084] After the original data is read in chronological order, it is uniformly converted into a .npz package, and zero-mean unit variance standardization is performed on the input and labels. At the same time, the scaler is fixed so that the scale can be de-standardized to restore the dimensions during evaluation and visualization. During the training phase, the input is organized as a [B,T,N,F] tensor (N=X_train.shape[2], F=X_train.shape[3]), and the labels are [B,N] synchronized across the network. The adjacency matrix is ​​converted into a sparse edge_index after removing self-loops to adapt to graph propagation / attention calculation. A data loader is constructed according to the training-validation-test partition and the optimal validation weights are saved.

[0085] The key fields, model input features, and descriptions after processing are shown in Tables 2 and 3.

[0086] Table 2. Data fields after preprocessing

[0087]

[0088] Table 3 Model Input Feature Fields

[0089]

[0090] The prediction model uses the road network topology provided by the adjacency matrix as a structural prior. It incorporates three types of inputs: historical traffic, temporal attributes, and node attributes, which are then fed into two branches: a spatial backbone network and a temporal backbone network. The spatial backbone network first standardizes and embeds numerical and categorical features to form node vectors, followed by edge weight scoring. In the attention mechanism, monotonicity and directionality constraints are explicitly applied, along with similarity suppression, to obtain task-related attention weights. Subsequently, spatial representations are obtained through 1–3 layers of GNN message passing, and spatial dependencies are constrained and monitored using physical consistency measures such as ingress and egress strengths and approximate conservation. The temporal backbone network performs multi-windowing, stride slicing, and multi-scale patching on the sequences from the spatial representations. These are then input into the PatchTST attention backbone network via linear projection and time / location encoding (using causal masks and recent bias), supplemented with lightweight DW convolutions to enhance local smoothing. Finally, the temporal features are converged and shaped. The framework diagram of the prediction model is shown below. Figure 2 As shown.

[0091] The core innovation of the spatial backbone network lies in interpretable road network graph modeling, specifically manifested as follows: under the constraint of directed adjacency priors, based on the representation formed by standardized node features and category embeddings, it adaptively learns time-varying edge weights through an attention mechanism of "content scoring + explanatory item addition". Specifically, in The content vector explicitly superimposes a monotonic directional term driven by flow difference and a similarity suppression term, giving the congestion impact a clear physical meaning and controllability across upstream and downstream links; combined with neighbor information... With lightweight gating, robust effective weights are obtained. Subsequently, with To achieve stable spatial representations, 1-3 layers of sparse message passing are implemented for the coefficients. Inflow and outflow intensities, along with approximate conservation measures, are introduced as structural regularization to suppress spurious convergence and noise amplification. Physical consistency constraints, used as structural regularization, monitor and statistically analyze the inflow and outflow intensities of each road segment during training, minimizing their differences. This effectively suppresses spurious convergence / discontinuity and amplification effects caused by structural noise without altering the forward computation. This allows the spatial backbone network to dynamically identify key propagation links and output link-level explanations while maintaining computational scalability. This provides reliable spatial priors and visual evidence for subsequent multi-scale temporal modeling and decoding extrapolation, enabling interpretable road network modeling. Figure 3 As shown.

[0092] The input and symbol descriptions for interpretable road network diagrams are as follows: Let the road network... The number of nodes is Adjacency matrix To achieve directed and loop-free operation, during training, the edge_index is converted to sparse and only applied to... Calculate the edges.

[0093] Given a single-time input:

[0094]

[0095] Where F represents the feature dimension of each node, each node's features It is composed of standardized numerical features and category feature embeddings (weekday, time_slot, is_workday, direction, lanes, road_id). Node embeddings are obtained through linear projection. :

[0096]

[0097]

[0098] Where (d is a hyperparameter representing the dimension of the model's latent space) is related to the time window. The following spatial process can be repeated in parallel to obtain... The representation stack is used by the time backbone network.

[0099] The edge weights of the interpretable road network graph are explained as follows: In the content scoring, two terms, monotonic / directional reinforcement and similarity suppression, are explicitly superimposed and non-negative constraints are applied. Sparse masking is first performed based on priors such as distance, lane, and capacity. Then, source normalization is performed on the outgoing edges of the same source node, so that the weights have the physical semantics of "probability distribution of outflow among outgoing edges". At the same time, the sharpness of the weights and channel contribution are controlled by temperature coefficient and gating, which suppresses spurious correlation and false convergence and diffusion from the mechanism, and enhances conservation consistency, robustness and cross-domain generalization.

[0100] Spatiotemporal graph models that rely solely on data-driven learning often suffer from limitations in the learned edge weights. The process generates spurious correlations and false convergences / discontinuities: attention may misconnect distant or reverse road segments, overfit noise time steps, and scaling may render edge weights lacking physical meaning. This violates the geometric / directional priors and approximate conservation of road networks, and weakens interpretability and cross-domain generalization. To address this, this invention introduces Interpretable Edge Weight Scoring (MES) in the edge weight construction stage: decomposing the weight of each edge into a monotonic factor of the outflow intensity related to the outflow. Similarity inhibitors It integrates physical priors such as direction consistency, distance decay, and lane / capacity, first obtains a sparse mask and then performs source normalization to form a column random adjacency. And it is given explicit probabilistic semantics for outflow allocation. This design naturally complements the subsequent conservation and consistency regularization, improving structural interpretability and stability while reducing peak errors and instability propagation caused by structural noise.

[0101] In the neighborhood group Internal learning edge The weight.

[0102] First, define the flow difference:

[0103]

[0104] Decompose it into directional components (This can be understood as the directional component and amplitude component of traffic flow differences) (This can be understood as the component of traffic flow difference amplitude):

[0105]

[0106] Perform a linear transformation on the endpoint embedding And concatenate it with the difference vector:

[0107]

[0108] in, , Let i and j represent the latent vectors of nodes i and j after the input projection, respectively. It represents the combination of static features such as geometry, relationship, and rules between nodes i and j, such as the distance between roads, whether it is a ramp, the difference in the number of lanes, etc.

[0109] Will Input linear layer and use Function activation yields a content score:

[0110]

[0111] The edge weight scoring weight vector represents learnable parameters. Then it represents the vector dot product. It is a scalar, a real number that can be positive or negative.

[0112] In addition to content scoring, two types of interpretable items are explicitly injected additively:

[0113]

[0114] (Monotonic / Directional): When upstream High and downstream Low time ( ),Increase The impact; if the directions do not match The cutoff value is 0.

[0115] (Similarity Suppression): When the difference is too large, the edge score is penalized to avoid "false strong edges" caused by noise or distant associations.

[0116] pass Nonnegative constraints maintain physical meaning stability and ensure they are not "cancelled" by content layer weights.

[0117] For the same target node The neighbors are normalized, and the border rights are obtained. :

[0118]

[0119] The temperature coefficient is used to adjust the weight sharpness. Gating can be further added.

[0120]

[0121] Where g is the scoring function, which takes a vector of two endpoints as input and outputs a scalar score.

[0122] If gating is not enabled, then... .

[0123] In real road networks, the number of vehicles exhibits an approximately conserved characteristic within a short time window: the change in vehicle stock on a certain road segment equals the inflow minus the outflow plus external factors (such as ramp entry / exit, U-turns, parking lots, missed detections, etc.). In most arterial road sections, external factors are relatively small on a minute-level scale and can be considered as disturbances, thus the conservation relationship is basically valid. If the model part corresponding to the spatial backbone network lacks this constraint, purely data-driven approaches are prone to learning spurious dispersion structures that only allow inflows or outflows in the long term, weakening the semantics of outflow allocation of edge weights, leading to poor interpretability, decreased cross-temporal generalization, and instability in multi-step extrapolation. To address this, this invention, without changing the forward computation, defines inflow and outflow strengths for the learned column random adjacency, and uses the difference between the two to form a conserved residual as soft regularization and training monitoring, so that the topology gradually approximates a physically feasible flow pattern, while allowing actual external factors and noise to be flexibly absorbed.

[0124] Spatiotemporal graph models that rely solely on data-driven learning are prone to problems with the learned graphs. The appearance of "false convergence / discontinuity" (where a road segment experiences prolonged periods of only inbound or only outbound traffic) not only violates the approximate conservation of traffic flow but also harms interpretability and temporal generalization. Therefore, this invention, without altering the forward computation, randomly connects the learned columns. Ingress / exgress strengths and conserved residuals are introduced as physical consistency regularization and training monitoring.

[0125] Therefore, this invention assumes that the learned sparse column random matrix satisfy:

[0126]

[0127] This can be interpreted as the "conditional probability of outflow allocation": That is, from the source section The outflow (which can be understood as the source node) is allocated to the receiving segment with probability. (This can be understood as the target node).

[0128] For each node Define input strength and output strength:

[0129]

[0130] in It is "merging into" "The sum of probability weights" It is "from" "The sum of probability weights for the outflow to each receiving node." Because Source normalization of a column, theoretically for any source have ,therefore:

[0131]

[0132] That is, the global average output intensity is 1, but the imbalance may still occur at each node (which is the "pseudo-sink / dissipate" phenomenon that this invention hopes to suppress).

[0133] Then, the difference between the input and output intensities at each node is used to characterize the conserved residuals, and a lightweight regularization is added to the loss corresponding to the spatial backbone network:

[0134]

[0135] in, The smaller the value, the closer the "inflow = outflow" relationship is for each node, and the more the graph structure conforms to approximate conservation. It is added to the total loss with a small weight. Without changing any forward path, the structure tends to be physically reasonable when "pushing" it only during backward propagation.

[0136] Edge flow (i.e., the mass / vehicle flow allocated from source node i to target node k within the corresponding time period) is represented by factorization:

[0137]

[0138] in It is a node The interpretable “outflow intensity” (is an interpretable scaling factor that represents how much “mass / traffic” needs to be allocated outward within the corresponding time step i). This is a conditional probability, representing the probability of distribution to each outgoing neighbor k given that the source is i. The column randomness is guaranteed. ,make It becomes a identifiable scale factor. At the same time, prompt and Maintaining approximate conservation at the layer level, thus... The physical semantics corroborate each other, reducing the degree of freedom of "generating false convergence / dispersion based on attention weight".

[0139] Among them, time node The outflow intensity is represented using a nonnegative regression head:

[0140]

[0141] in, For the model to nodes The current representation; This is a recent statistic (mean, peak value); Capacity scaling factor (e.g., node) (Historical 95th percentile flow rate or engineering upper limit). , All are learnable parameters. To ensure non-negativity and numerical stability of the output, the temporal backbone network improves the PatchTST model by: normalizing single-sequence instances within a fixed time window, then embedding the original sequence fragments into trainable tokens through linear projection and overlaying positional / time-period semantics (hour, day of the week, etc.); and employing multi-scale patching (different scales) to simultaneously characterize short-term fluctuations, intraday rhythms, and cross-day trends. Parallel slices are used to form short, medium, and long-range receptive fields, and temporal-dependent encoding is completed using Transformer blocks (multi-head attention—feedforward network—residual connection + layer normalization) with causal masks and local-global hybrid attention. After encoding, lightweight calibration of each scale representation is performed through cross-scale interaction, and scale weights are adaptively allocated using gated fusion. To suppress scale conflicts and automatically favor the most effective time granularity based on the scenario; to characterize extrapolable long-term stationary components, a Koopman-Graph (a composite structure combining the Koopman operator and a graph neural network, specifically where the output of the graph neural network is concatenated with the Koopman operator) is introduced at the tail to improve performance in shared linear operators. The evolution is carried out in the improvement space, and its dynamics are constrained by self-consistency and spectral stability regularity, combined with gating coefficients. An adaptive division of labor is implemented between the Koopman branch and the multi-scale branch, ultimately providing node and overall network predictions. This temporal backbone network strikes a balance between computational scalability and interpretability, providing reliable temporal priors for decoding and extrapolation. The improved PatchTST model framework diagram is shown below. Figure 4 .

[0142] The improved PatchTST model (PatchTST++) structurally expands upon PatchTST in three key areas: First, it introduces multi-timescale patches, employing a multi-window set instead of a single fixed patch scale. Patch length and step size are set in pairs and can be asynchronously evaluated, thus simultaneously covering short-term fluctuations, intraday rhythms, and cross-day trends within the same computational graph. Second, it implements a collaborative mask, using causal local window attention and adding a small number of remote anchors to capture cross-period dependencies, making the complexity nearly linear and more favorable for long sequences. Third, it incorporates scale-level gating fusion and temperature control, adaptively weighting the contributions of different scales to avoid single-scale dominance or noise amplification. Overall, multi-scale patching maintains the efficient patch embedding approach of PatchTST while improving the discriminative power and inference stability for multi-timescale mixed patterns through multi-window, sparse remote dependencies, and gating fusion.

[0143] Traffic flow data exhibits temporal characteristics across multiple time scales. Short-term fluctuations, intraday rhythms, and long-term trends are often intertwined, making single-scale modeling insufficient to capture this multi-layered information. Therefore, PatchTST++ introduces multi-scale patching, enabling the model to learn different temporal features at multiple time scales. To simultaneously characterize short-term fluctuations, intraday rhythms, and cross-day trends within a single computational graph, a set of scales is defined:

[0144]

[0145] in and respectively scale The patch length and step size. For nodes normalized sequence structure:

[0146]

[0147]

[0148] in, This represents the k-th time patch for node i at the s-th scale, i.e., the length window; This represents the start time index of the k-th patch. This indicates the step size of the scale. This indicates the number of patches, where L is the total history length. This indicates rounding down to the nearest integer.

[0149] For example, in a 15-minute sampling scenario, to cover three time scales—short-term disturbances, semi-diurnal rhythms, and trans-diurnal cycles—the numerical multiplication factor between multiple consecutive patch parameters can be defined as 3. For instance, three sets of right-aligned patch parameters (windows) can be defined. Step length ):

[0150]

[0151] These correspond to approximately 2-hour windows (patching every hour), approximately 6-hour windows (patching every 3 hours), and 1-day windows (patching every 12 hours), respectively. In this example, based on the above settings, it is possible to maintain computational efficiency while taking into account intraday rhythms such as sudden changes, travel peaks and troughs, as well as stable cross-day cyclical relationships. The maximum scale is aligned with the daily cycle, and "the same time of the previous day" is retained as a distant anchor point, which helps to stabilize long-term dependencies when phase shifts occur on holidays or abnormal days.

[0152] It should be noted that in practical applications, the specific values ​​of the patch parameters can be adaptively adjusted according to actual needs, provided that the aforementioned constraints are met. This invention does not limit the specific values ​​of the patch parameters. Each patch is linearly embedded as follows:

[0153]

[0154] Represents the patch embedding vector. This represents the linear mapping matrix at the s-th scale. This means taking this length as... The scalar sequence is flattened into a column vector. This represents the bias vector at the s-th scale. The embedding dimension at this scale is a hyperparameter.

[0155] Based on this, token sequences at various scales can be obtained. Multi-scale sequences can be fed into their respective encoders in parallel or spliced ​​together to a common dimension first. Post-shared encoder. The complexity is approximately O(n log n). While ensuring multi-scale detection modes, scalability is maintained.

[0156] Traffic flow data exhibits strong temporal periodicity, including intraday peaks and troughs, and differences between weekdays and weekends. To enable the model to understand and distinguish traffic flow changes across different time periods, temporal semantic encoding is introduced to enhance the model's ability to perceive these time-period variations. To achieve cross-day alignment and intra-weekly pattern differentiation, temporal semantic embeddings are overlaid on tokens at each scale.

[0157]

[0158] in, This can be understood as the embedding vector corresponding to the s-th scale. For intra-scale location encoding, These are respectively encoded as "hour-of-day" (which can be understood as the first time-domain feature) and "day of the week" (which can be understood as the second time-domain feature), and can be taken as learnable vectors or triangular Fourier bases (such as...). (Form). This semantically satisfies: if If the events fall on different dates but within the same hour / day of the week, their temporal semantics are consistent, which is beneficial for model transfer and alignment.

[0159] Traffic flow data exhibits not only short-term local dependencies (such as traffic signal cycles and traffic chain reactions) but also longer-term periodic variations (such as intraday patterns). Standard Transformer models, through global attention, may overlook short-term fluctuations; therefore, a strategy combining causal constraints and a local-global hybrid approach is needed to enhance the model's ability to perceive both short-term and long-term dependencies. To satisfy the causal requirements of predictions while considering both near-field impacts and long-field rhythms, the sequence at each scale... Construct a masked self-attention mechanism. Let:

[0160]

[0161] Q is the query vector for each patch, K is the key vector for each patch, and V is the value vector for each patch. Let be the patch sequence matrix at the s-th scale. These are learnable matrices.

[0162] Attention is calculated as follows:

[0163]

[0164] It is the dimension of query and key, using To perform scaling, prevent the dot product from becoming too large and causing the softmax to become too sharp.

[0165] mask (This can be understood as the aforementioned co-mask) consists of the sum of three parts:

[0166]

[0167] in (This can be understood as a causal mask) Access to the future is prohibited. (This can be understood as a local mask) limiting the receptive field to the width of a local window. , (This can be understood as an anchor mask) for sparse long-range anchors (such as...) (corresponding to the same time the previous day) open connection. This maintains... It achieves near-linear complexity while retaining the necessary long-range rhythm information.

[0168] Traffic flow data exhibits significant long-term dependencies, such as the regularity of intraday peaks and troughs, seasonal variations, and weekly patterns. These dependencies are difficult to model using standard graph neural networks (GNNs). While GNNs can capture local structures and short-term dependencies, they often perform poorly when dealing with global dependencies spanning longer periods. To effectively capture these long-term trends, the Koopman operator is introduced as a linearization method for nonlinear systems. It can transform the nonlinear dynamics of traffic flow data into linear dynamics through a boosting process, thereby effectively capturing long-term dependencies and stabilizing the modeling of long-term trends.

[0169] In this invention, the temporal node features in the graph neural network are defined. ,in It refers to the batch size. It is the number of time steps. It is the number of nodes. This refers to the feature dimension of each node. To capture long-term dependencies, we consider the node features at each time step. Nonlinear enhancement is achieved through a nonlinear encoder. Mapping to a higher-dimensional space:

[0170]

[0171] Here, It is a non-linear mapping function, usually implemented using a neural network, and its output is... This is the improved representation. It is about enhancing the dimensions of space.

[0172] Using the Koopman operator We can map the evolution of time series from a nonlinear space to a linear space. In this space, the relationships between states become linearized, allowing us to use the Koopman operator to describe the long-term dependencies of time series data.

[0173]

[0174] in, It is a shared Koopman operator, representing the time step. Time to step A linear transformation matrix of 1.

[0175] To ensure the temporal consistency of the Koopman operator, we introduce a self-consistency loss to constrain the consistency between states at adjacent time steps. This loss function calculates the difference between the current state and the previous state, and then performs a linearized evolution using the Koopman operator to minimize this difference.

[0176]

[0177] This loss function ensures that the Koopman operator consistently models the long-term trends of time-series data, thereby capturing long-term dependencies and preventing instability of the operator over time.

[0178] To prevent the Koopman operator from becoming unstable during training, we introduce spectral stability regularization, which constrains the spectral norm (i.e., the largest eigenvalue) of the Koopman operator to be close to 1. This prevents the operator from growing or decaying too rapidly, thus avoiding runaway long-term trends in predictions.

[0179]

[0180] in, It is the spectral norm of the Koopman operator, and the regularization term ensures its stability during training.

[0181] In summary, this invention establishes a method for short-term traffic flow data processing and synchronous prediction for urban road networks: taking sensor sequences and road network topology as inputs, it forms an integrated process of "data standardization—feature construction—network-wide joint modeling—online inference—result presentation"; it provides stable and traceable network-wide prediction outputs under single-step and multi-step time intervals, and simultaneously provides contribution and uncertainty indicators for key road segments; through standardized model interfaces and deployment schemes, it achieves low-latency online services and rapid cross-regional migration, and is equipped with an acceptable evaluation and operation log mechanism, ultimately providing directly callable prediction capabilities and auxiliary decision-making basis for traffic command and control scenarios. For example, in expressway management, this solution can predict traffic flow on each segment of the corresponding road network within a future timeframe (e.g., 30 or 60 minutes). Based on this, it can predict impending congestion on certain segments (such as downstream segments). Knowing this information, the timing of traffic lights on the corresponding segments (such as upstream ramps) can be proactively adjusted to delay congestion at its source. Simultaneously, dynamic route guidance can be disseminated to navigation apps and roadside information boards based on the prediction results, providing drivers with globally optimal route choices and effectively balancing the road network load. In one embodiment, the present invention also provides a traffic flow prediction system 100 based on interpretable graphs and multi-scale temporal trunks, such as... Figure 5 As shown, the system 100 includes:

[0182] The prediction module 101 is used to process the historical traffic flow information of the target road network according to the pre-trained prediction model to obtain the traffic flow prediction information of the target road network in the target time period. The prediction model is a composite model formed by fusing a spatial backbone network and a temporal backbone network. The spatial backbone network includes a graph neural network, and the temporal backbone network includes an improved PatchTST model. The improved PatchTST model includes multiple time-scale patches, which include: a first time-scale patch corresponding to short-term fluctuations, a second time-scale patch corresponding to intraday rhythms, and a third time-scale patch corresponding to cross-day trends.

[0183] The historical traffic flow information is used to indicate the traffic flow of each node in the target road network during a historical period; the traffic flow prediction information is used to indicate the traffic flow of each node in the target road network during a target period.

[0184] It should be noted that the system provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the traffic flow prediction system based on interpretable graphs and multi-scale time backbones provided in the above embodiments and the traffic flow prediction method based on interpretable graphs and multi-scale time backbones are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0185] This invention also provides an electronic device. Please refer to [link to relevant documentation]. Figure 6 The electronic device may include a processor 201, a memory 202, and a program 2021 stored in the memory 202 and executable on the processor 201.

[0186] When program 2021 is executed by processor 201, it can achieve the following: Figure 1 Any steps in the corresponding method embodiments and the achievement of the same beneficial effects will not be repeated here.

[0187] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by hardware related to program instructions, and the program can be stored in a readable medium.

[0188] This invention also provides a readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described functions. Figure 1 Any step in the corresponding method embodiment can achieve the same technical effect, and will not be repeated here to avoid repetition.

[0189] The computer-readable storage medium of this invention can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0190] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0191] The program code contained on the storage medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0192] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or terminal. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0193] This invention also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the traffic flow prediction method based on interpretable graphs and multi-scale time backbones provided in the above embodiments.

[0194] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0195] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones, characterized in that, The method includes: The historical traffic flow information of the target road network is processed according to a pre-trained prediction model to obtain the traffic flow prediction information of the target road network during the target time period. The prediction model is a composite model formed by fusing a spatial backbone network and a temporal backbone network. The spatial backbone network includes a graph neural network, and the temporal backbone network includes an improved PatchTST model. The improved PatchTST model includes multiple time-scale patches, which include: a first time-scale patch corresponding to short-term fluctuations, a second time-scale patch corresponding to intraday rhythms, and a third time-scale patch corresponding to cross-day trends. The historical traffic flow information is used to indicate the traffic flow of each node in the target road network during a historical time period; the traffic flow prediction information is used to indicate the traffic flow of each node in the target road network during a target time period. In the multi-timescale patch, the self-attention corresponding to each timescale patch includes a collaborative mask, wherein the collaborative mask is jointly determined based on a causal mask indicating that access to the future is prohibited, a local mask indicating that the receptive field is within a set window width, and an anchor mask indicating that the connection of remote anchors satisfies a set time interval.

2. The traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones according to claim 1, characterized in that, The patch length of the first timescale patch is less than the patch length of the second timescale patch, and the patch length of the second timescale patch is less than the patch length of the third timescale patch. The patch step size of the first timescale patch is smaller than the patch step size of the second timescale patch, and the patch step size of the second timescale patch is smaller than the patch step size of the third timescale patch.

3. The traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones according to claim 1, characterized in that, In the multi-timescale patch, the embedding vector corresponding to each timescale patch includes a first time-domain feature and a second time-domain feature. The first time-domain feature is used to indicate the hour position of the corresponding time within a day, and the second time-domain feature is used to indicate the day position of the corresponding time within a week.

4. The traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones according to claim 1, characterized in that, The spatial backbone network also includes a Koopman operator, the input of which is connected to the output of the graph neural network.

5. The traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones according to claim 4, characterized in that, The output data of the prediction model is obtained by weighted calculation of the spatial prediction data corresponding to the Koopman operator and the temporal prediction data corresponding to the improved PatchTST model.

6. The traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones according to claim 1, characterized in that, The spatial backbone network also includes an edge weight scoring module, which is used to obtain the edge weights between different nodes. The edge weights are determined based on the differences in traffic flow between different nodes in the target road network. The edge weights are used to indicate the probability that the traffic flow of the source node is allocated to the target node. The source node and the target node are any two different nodes in the target road network, and the sum of all edge weights corresponding to the same source node is 1.

7. The traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones according to claim 6, characterized in that, The edge weights are determined based on the differences in traffic flow between different nodes in the target road network. The loss function corresponding to the spatial backbone network includes an inflow-outflow conservation deviation term. The inflow-outflow conservation deviation term is determined based on the average value of the inflow-outflow intensity deviation of each node in the target road network. The inflow-outflow intensity deviation is the absolute difference between the inflow intensity and the outflow intensity of a node. The inflow intensity is the sum of all edge weights with the corresponding node as the target node, and the outflow intensity is the sum of all edge weights with the corresponding node as the source node.

8. The traffic flow prediction method based on interpretable graphs and multi-scale temporal backbones according to claim 6, characterized in that, The edge weights are determined based on the traffic flow difference direction component and traffic flow difference magnitude component between the corresponding source node and target node. Wherein, when the traffic flow of the corresponding source node is greater than the traffic flow of the corresponding target node, the traffic flow difference direction component is positively correlated with the corresponding edge weight; when the traffic flow of the corresponding source node is less than or equal to the traffic flow of the corresponding target node, the traffic flow difference direction component is not related to the calculation of the corresponding edge weight. The magnitude component of the traffic flow difference is negatively correlated with the corresponding edge weight.

9. A traffic flow prediction system based on interpretable graphs and multi-scale temporal backbones, characterized in that, The system includes: The prediction module is used to process the historical traffic flow information of the target road network according to the pre-trained prediction model to obtain the traffic flow prediction information of the target road network in the target time period. The prediction model is a composite model formed by fusing a spatial backbone network and a temporal backbone network. The spatial backbone network includes a graph neural network, and the temporal backbone network includes an improved PatchTST model. The improved PatchTST model includes multiple time-scale patches, which include: a first time-scale patch corresponding to short-term fluctuations, a second time-scale patch corresponding to intraday rhythms, and a third time-scale patch corresponding to cross-day trends. The historical traffic flow information is used to indicate the traffic flow of each node in the target road network during a historical time period; the traffic flow prediction information is used to indicate the traffic flow of each node in the target road network during a target time period. In the multi-timescale patch, the self-attention corresponding to each timescale patch includes a collaborative mask, wherein the collaborative mask is jointly determined based on a causal mask indicating that access to the future is prohibited, a local mask indicating that the receptive field is within a set window width, and an anchor mask indicating that the connection of remote anchors satisfies a set time interval.

Citation Information

Patent Citations

  • Transform-based adaptive space-time diagram neural network traffic flow prediction method and system

    CN114492992A

  • Large-scale traffic prediction method and system fused with long-distance space-time correlation

    CN117727178A