Methods, systems, and computer program products for space-time map interlayer converters for traffic flow prediction

The processing of historical data of the traffic network through the space-time graph mezzanine transformer solves the problem that existing models fail to effectively utilize spatial correlation and achieves higher-precision traffic flow prediction.

CN120457666APending Publication Date: 2025-08-08VISA INTERNATIONAL SERVICE ASSOCIATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380079877.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-18
Filing Date
2023-11-17
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing traffic flow prediction model fails to effectively utilize the spatial correlation of the traffic network, resulting in insufficient prediction accuracy.

Method used

Using a space-time graph mezzanine transformer, including a top time transformer, a space transformer and a bottom time transformer, processes historical timing traffic data through a multi-head attention and feedforward network, generates predicted timing traffic data, and trains using Huber loss function.

Benefits of technology

It improves the accuracy and efficiency of traffic flow prediction, can better capture the spatial characteristics of the traffic network, and improves the prediction accuracy of future traffic conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120457666A_ABST
    Figure CN120457666A_ABST
Patent Text Reader

Abstract

A method, system, and computer program product for traffic flow prediction: obtaining a graph representing a traffic network; obtaining historical time sequence traffic data associated with historical traffic conditions at a plurality of historical time steps in the traffic network; processing the graph and the historical timing traffic data with each of at least one space-time graph interlayer transducer to generate an interlayer transducer output, where each space-time graph interlayer transducer includes a top time transducer, a space transducer, and a bottom time transducer, the space transformer receives the output of the top time transformer and the graph as input, and the bottom time transformer receives the output of the space transformer as input; and generate predicted timing traffic data associated with predicted traffic conditions at a plurality of subsequent time steps in the traffic network based on the interlayer transducer outputs from each space-time map interlayer transducer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 426,473, filed on November 18, 2022, the entire disclosure of which is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure relates generally to traffic flow prediction and, in some non-limiting embodiments or aspects, to methods, systems, and computer program products for a space-time graph sandwich transformer for traffic flow prediction. Background Art

[0004] Traffic forecasting plays a crucial role in modern intelligent transportation systems. Efficient and accurate traffic flow predictions enable better traffic management and planning. Generally speaking, traffic flow forecasting aims to predict future traffic conditions by leveraging historical time-series traffic inputs and the underlying transportation network. Classical statistical and sequence models primarily emphasize time-series inputs but ignore the spatial correlation of transportation networks, leaving significant room for improvement. Summary of the Invention

[0005] Thus, improved methods, systems, and computer program products for traffic flow prediction are provided.

[0006] According to some non-limiting embodiments or aspects, a method is provided, comprising: obtaining, with at least one processor, a graph representing a traffic network; obtaining, with the at least one processor, historical time-series traffic data associated with historical traffic conditions at a plurality of historical time steps in the traffic network; processing, with the at least one processor, the graph and the historical time-series traffic data using each of at least one space-time graph sandwich transformers to generate a sandwich transformer output, wherein each space-time graph sandwich transformer comprises a top time transformer, a space transformer, and a bottom time transformer, the space transformer receiving the output of the top time transformer and the graph as input, and the bottom time transformer receiving the output of the space transformer as input; and generating, with the at least one processor, predicted time-series traffic data associated with predicted traffic conditions at a plurality of next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer.

[0007] In some non-limiting embodiments or aspects, the at least one space-time pattern sandwich converter comprises a plurality of space-time pattern sandwich converters.

[0008] In some non-limiting embodiments or aspects, processing the graph and the historical time-series traffic data with the at least one processor using each space-time graph sandwich transformer to generate the sandwich transformer output is performed in parallel with the plurality of space-time graph sandwich transformers.

[0009] In some non-limiting embodiments or aspects, the method further includes: providing the historical time series traffic data as input to an input transformation layer including a fully connected layer using the at least one processor; and receiving a hidden time series embedding as output from the input transformation layer using the at least one processor, wherein the top time transformer of each space-time graph sandwich transformer receives the hidden time series embedding as input.

[0010] In some non-limiting embodiments or aspects, generating the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer using the at least one processor includes: concatenating the sandwich transformer output from each space-time graph sandwich transformer; providing the concatenated sandwich transformer output from each space-time graph sandwich transformer as input to a multi-step prediction layer including at least two fully connected layers; and receiving the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network as output from the multi-step prediction layer.

[0011] In some non-limiting embodiments or aspects, processing the graph and the historical temporal traffic data with each spatio-temporal graph sandwich transformer using the at least one processor to generate the sandwich transformer output includes: applying multi-head attention to the input of the top temporal transformer using the top temporal transformer, followed by applying a feedforward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feedforward network.

[0012] In some non-limiting embodiments or aspects, processing the graph and the historical time-series traffic data with each space-time graph sandwich transformer using the at least one processor to generate the sandwich transformer output includes: applying degree-based encoding and singular value decomposition (SVD)-based encoding to the input of the spatial transformer using the spatial transformer, the input including the output of the top time transformer and the graph; and applying spatial encoding as a bias term in the multi-head attention applied to the input of the spatial transformer using the spatial transformer, the input including the output of the top time transformer and the graph.

[0013] In some non-limiting embodiments or aspects, the input of the bottom temporal transformer including the output of the spatial transformer includes a plurality of time series sequences, and wherein the processing of the graph and the historical time series traffic data with the at least one processor using each spatial-temporal graph sandwich transformer to generate the sandwich transformer output includes: appending a special token to the beginning of each of the plurality of time series sequences with the bottom temporal transformer; and applying a multi-head attention with the bottom temporal transformer to the plurality of token-appended time series sequences, followed by a feed-forward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

[0014] In some non-limiting embodiments or aspects, the method further includes: training the at least one space-time graph sandwich transformer with the at least one processor according to a Huber loss function, wherein the Huber loss function depends on the predicted time-series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network and the actual time-series traffic data associated with the actual traffic conditions at the multiple next time steps in the traffic network.

[0015] In some non-limiting embodiments or aspects, the transportation network comprises a payment network, wherein the graph comprises a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of payment processing servers, and wherein the plurality of edges are associated with a plurality of connections between the plurality of payment processing servers.

[0016] According to some non-limiting embodiments or aspects, a system is provided, comprising: at least one processor programmed or configured to: obtain a graph representing a traffic network; obtain historical time-series traffic data associated with historical traffic conditions at multiple historical time steps in the traffic network; process the graph and the historical time-series traffic data with each of at least one space-time graph sandwich transformers to generate a sandwich transformer output, wherein each space-time graph sandwich transformer comprises a top time transformer, a space transformer, and a bottom time transformer, the space transformer receiving the output of the top time transformer and the graph as input, and the bottom time transformer receiving the output of the space transformer as input; and generate predicted time-series traffic data associated with predicted traffic conditions at multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer.

[0017] In some non-limiting embodiments or aspects, the at least one space-time pattern sandwich converter comprises a plurality of space-time pattern sandwich converters.

[0018] In some non-limiting embodiments or aspects, the at least one processor is programmed or configured to process the graph and the historical time-series traffic data with each space-time graph sandwich transformer in parallel with the plurality of space-time graph sandwich transformers to generate the sandwich transformer output.

[0019] In some non-limiting embodiments or aspects, the at least one processor is further programmed or configured to: provide the historical time series traffic data as input to an input transformation layer comprising a fully connected layer; and receive a hidden time series embedding as output from the input transformation layer, wherein the top time transformer of each space-time graph sandwich transformer receives the hidden time series embedding as input.

[0020] In some non-limiting embodiments or aspects, the at least one processor is programmed or configured to generate the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer by: concatenating the sandwich transformer output from each space-time graph sandwich transformer; providing the concatenated sandwich transformer output from each space-time graph sandwich transformer as input to a multi-step prediction layer comprising at least two fully connected layers; and receiving the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network as output from the multi-step prediction layer.

[0021] In some non-limiting embodiments or aspects, the at least one processor is programmed or configured to process the graph and the historical temporal traffic data with each spatio-temporal graph sandwich transformer to generate the sandwich transformer output by applying a multi-head attention with the top temporal transformer to the input of the top temporal transformer, followed by a feed-forward network with residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

[0022] In some non-limiting embodiments or aspects, the at least one processor is programmed or configured to process the graph and the historical time-series traffic data with each space-time graph sandwich transformer to generate the sandwich transformer output by: applying, with the spatial transformer, degree-based encoding and singular value decomposition (SVD)-based encoding to the input of the spatial transformer, the input comprising the output of the top time transformer and the graph; and applying, with the spatial transformer, spatial encoding as a bias term in a multi-head attention applied to the input of the spatial transformer, the input comprising the output of the top time transformer and the graph.

[0023] In some non-limiting embodiments or aspects, the input of the bottom temporal transformer including the output of the spatial transformer comprises a plurality of time series sequences, and wherein the at least one processor is programmed or configured to process the graph and the historical time series traffic data with each space-time graph sandwich transformer to generate the sandwich transformer output by: appending a special token to the beginning of each of the plurality of time series sequences with the bottom temporal transformer; and applying a multi-head attention with the bottom temporal transformer to the plurality of token-appended time series sequences, followed by a feed-forward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

[0024] In some non-limiting embodiments or aspects, the at least one processor is further programmed or configured to: train the at least one space-time graph sandwich transformer according to a Huber loss function, wherein the Huber loss function depends on the predicted time-series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network and the actual time-series traffic data associated with the actual traffic conditions at the multiple next time steps in the traffic network.

[0025] In some non-limiting embodiments or aspects, the transportation network comprises a payment network, wherein the graph comprises a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of payment processing servers, and wherein the plurality of edges are associated with a plurality of connections between the plurality of payment processing servers.

[0026] According to some non-limiting embodiments or aspects, a computer program product is provided, comprising a non-transitory computer-readable medium comprising program instructions that, when executed by at least one processor, cause the at least one processor to: obtain a graph representing a traffic network; obtain historical time-series traffic data associated with historical traffic conditions at multiple historical time steps in the traffic network; process the graph and the historical time-series traffic data with each of at least one space-time graph sandwich transformers to generate a sandwich transformer output, wherein each space-time graph sandwich transformer comprises a top time transformer, a space transformer, and a bottom time transformer, the space transformer receiving the output of the top time transformer and the graph as input, and the bottom time transformer receiving the output of the space transformer as input; and generate predicted time-series traffic data associated with predicted traffic conditions at multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer.

[0027] In some non-limiting embodiments or aspects, the at least one space-time pattern sandwich converter comprises a plurality of space-time pattern sandwich converters.

[0028] In some non-limiting embodiments or aspects, the program instructions, when executed by the at least one processor, cause the at least one processor to process the graph and the historical time-series traffic data with each space-time graph sandwich transformer in parallel with the plurality of space-time graph sandwich transformers to generate the sandwich transformer output.

[0029] In some non-limiting embodiments or aspects, the program instructions, when executed by the at least one processor, further cause the at least one processor to: provide the historical time series traffic data as input to an input transformation layer comprising a fully connected layer; and receive a hidden time series embedding as output from the input transformation layer, wherein the top time transformer of each space-time graph sandwich transformer receives the hidden time series embedding as input.

[0030] In some non-limiting embodiments or aspects, the program instructions, when executed by the at least one processor, cause the at least one processor to generate the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer by: concatenating the sandwich transformer output from each space-time graph sandwich transformer; providing the concatenated sandwich transformer output from each space-time graph sandwich transformer as input to a multi-step prediction layer comprising at least two fully connected layers; and receiving the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network as output from the multi-step prediction layer.

[0031] In some non-limiting embodiments or aspects, the program instructions, when executed by the at least one processor, cause the at least one processor to process the graph and the historical temporal traffic data with each spatio-temporal graph sandwich transformer to generate the sandwich transformer output by applying a multi-head attention with the top temporal transformer to the input of the top temporal transformer, followed by a feed-forward network, with residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

[0032] In some non-limiting embodiments or aspects, the program instructions, when executed by the at least one processor, cause the at least one processor to process the graph and the historical time-series traffic data with each space-time graph sandwich transformer to generate the sandwich transformer output by: applying, with the spatial transformer, degree-based encoding and singular value decomposition (SVD)-based encoding to the input of the spatial transformer, the input comprising the output of the top time transformer and the graph; and applying, with the spatial transformer, spatial encoding as a bias term in a multi-head attention applied to the input of the spatial transformer, the input comprising the output of the top time transformer and the graph.

[0033] In some non-limiting embodiments or aspects, the input to the bottom temporal transformer including the output of the spatial transformer comprises a plurality of time series sequences, and wherein the program instructions, when executed by the at least one processor, cause the at least one processor to process the graph and the historical time series traffic data with each space-time graph sandwich transformer to generate the sandwich transformer output by: appending a special token to the beginning of each of the plurality of time series sequences with the bottom temporal transformer; and applying a multi-head attention with the bottom temporal transformer to the plurality of token-appended time series sequences, followed by a feed-forward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

[0034] In some non-limiting embodiments or aspects, the program instructions, when executed by the at least one processor, further cause the at least one processor to: train the at least one space-time graph sandwich transformer according to a Huber loss function, wherein the Huber loss function depends on the predicted time-series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network and the actual time-series traffic data associated with the actual traffic conditions at the multiple next time steps in the traffic network.

[0035] In some non-limiting embodiments or aspects, the transportation network comprises a payment network, wherein the graph comprises a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of payment processing servers, and wherein the plurality of edges are associated with a plurality of connections between the plurality of payment processing servers.

[0036] Additional non-limiting embodiments or aspects are set forth in the following numbered clauses:

[0037] Clause 1: A method comprising: obtaining, with at least one processor, a graph representing a traffic network; obtaining, with the at least one processor, historical time-series traffic data associated with historical traffic conditions at a plurality of historical time steps in the traffic network; processing, with the at least one processor, the graph and the historical time-series traffic data using each of at least one space-time graph sandwich transformers to generate a sandwich transformer output, wherein each space-time graph sandwich transformer comprises a top time transformer, a space transformer, and a bottom time transformer, the space transformer receiving as input the output of the top time transformer and the graph, and the bottom time transformer receiving as input the output of the space transformer; and generating, with the at least one processor, predicted time-series traffic data associated with predicted traffic conditions at a plurality of next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer.

[0038] Clause 2: The method of clause 1, wherein the at least one space-time graph sandwich transformer comprises a plurality of space-time graph sandwich transformers.

[0039] Clause 3: A method according to clause 1 or 2, wherein said processing of said graph and said historical time-series traffic data with said at least one processor using each space-time graph sandwich transformer to generate said sandwich transformer output is performed in parallel with said plurality of space-time graph sandwich transformers.

[0040] Clause 4: The method according to any one of clauses 1 to 3, further comprising: providing the historical time series traffic data as input to an input transformation layer comprising a fully connected layer using the at least one processor; and receiving a hidden time series embedding as output from the input transformation layer using the at least one processor, wherein the top time transformer of each space-time graph sandwich transformer receives the hidden time series embedding as input.

[0041] Clause 5: A method according to any one of clauses 1 to 4, wherein the generating of the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer using the at least one processor includes: concatenating the sandwich transformer output from each space-time graph sandwich transformer; providing the cascaded sandwich transformer output from each space-time graph sandwich transformer as input to a multi-step prediction layer including at least two fully connected layers; and receiving the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network as output from the multi-step prediction layer.

[0042] Clause 6: A method according to any one of clauses 1 to 5, wherein processing the graph and the historical temporal traffic data with the at least one processor using each space-time graph sandwich transformer to generate the sandwich transformer output includes: applying multi-head attention to the input of the top time transformer with the top time transformer, followed by applying a feedforward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feedforward network.

[0043] Clause 7: A method according to any one of clauses 1 to 6, wherein processing the graph and the historical time-series traffic data using each space-time graph sandwich transformer with the at least one processor to generate the sandwich transformer output includes: applying degree-based encoding and singular value decomposition (SVD)-based encoding to the input of the spatial transformer with the spatial transformer, the input including the output of the top time transformer and the graph; and applying spatial encoding as a bias term in the multi-head attention applied to the input of the spatial transformer with the spatial transformer, the input including the output of the top time transformer and the graph.

[0044] Clause 8: A method according to any one of clauses 1 to 7, wherein the input to the bottom temporal transformer including the output of the spatial transformer comprises a plurality of time series sequences, and wherein processing the graph and the historical time series traffic data using each space-time graph sandwich transformer with the at least one processor to generate the sandwich transformer output comprises: appending a special token to the beginning of each of the plurality of time series sequences with the bottom temporal transformer; and applying a multi-head attention to the plurality of token-appended time series sequences with the bottom temporal transformer, followed by a feed-forward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

[0045] Clause 9: A method according to any one of clauses 1 to 8, further comprising: training the at least one space-time graph sandwich transformer with the at least one processor according to a Huber loss function, wherein the Huber loss function depends on the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network and the actual time series traffic data associated with the actual traffic conditions at the multiple next time steps in the traffic network.

[0046] Clause 10: A method according to any one of clauses 1 to 9, wherein the transportation network includes a payment network, wherein the graph includes a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of payment processing servers, and wherein the plurality of edges are associated with a plurality of connections between the plurality of payment processing servers.

[0047] Clause 11: A system comprising: at least one processor, the at least one processor being programmed or configured to: obtain a graph representing a traffic network; obtain historical time-series traffic data associated with historical traffic conditions at multiple historical time steps in the traffic network; process the graph and the historical time-series traffic data with each of at least one space-time graph sandwich transformers to generate a sandwich transformer output, wherein each space-time graph sandwich transformer includes a top time transformer, a space transformer, and a bottom time transformer, the space transformer receiving the output of the top time transformer and the graph as input, and the bottom time transformer receiving the output of the space transformer as input; and generate predicted time-series traffic data associated with predicted traffic conditions at multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer.

[0048] Clause 12: The system of clause 11, wherein the at least one space-time graph sandwich transformer comprises a plurality of space-time graph sandwich transformers.

[0049] Clause 13: The system of clause 11 or 12, wherein the at least one processor is programmed or configured to process the graph and the historical time-series traffic data with each space-time graph sandwich transformer in parallel with the plurality of space-time graph sandwich transformers to generate the sandwich transformer output.

[0050] Clause 14: A system according to any one of clauses 11 to 13, wherein the at least one processor is further programmed or configured to: provide the historical time series traffic data as input to an input transformation layer comprising a fully connected layer; and receive a hidden time series embedding as output from the input transformation layer, wherein the top time transformer of each space-time graph sandwich transformer receives the hidden time series embedding as input.

[0051] Clause 15: A system according to any one of clauses 11 to 14, wherein the at least one processor is programmed or configured to generate the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer by: cascading the sandwich transformer output from each space-time graph sandwich transformer; providing the cascaded sandwich transformer output from each space-time graph sandwich transformer as input to a multi-step prediction layer comprising at least two fully connected layers; and receiving the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network as output from the multi-step prediction layer.

[0052] Clause 16: A system according to any one of clauses 11 to 15, wherein the at least one processor is programmed or configured to process the graph and the historical temporal traffic data with each spatio-temporal graph sandwich transformer to generate the sandwich transformer output by applying a multi-head attention with the top temporal transformer to the input of the top temporal transformer, followed by applying a feedforward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feedforward network.

[0053] Clause 17: A system according to any one of clauses 11 to 16, wherein the at least one processor is programmed or configured to process the graph and the historical time-series traffic data with each space-time graph sandwich transformer to generate the sandwich transformer output by: applying degree-based encoding and singular value decomposition (SVD)-based encoding to the input of the spatial transformer with the spatial transformer, the input comprising the output of the top time transformer and the graph; and applying spatial encoding as a bias term in a multi-head attention applied to the input of the spatial transformer with the spatial transformer, the input comprising the output of the top time transformer and the graph.

[0054] Clause 18: A system according to any one of clauses 11 to 17, wherein the input to the bottom temporal transformer including the output of the spatial transformer comprises a plurality of time series sequences, and wherein the at least one processor is programmed or configured to process the graph and the historical time series traffic data with each space-time graph sandwich transformer to generate the sandwich transformer output by: appending a special token to the beginning of each of the plurality of time series sequences with the bottom temporal transformer; and applying a multi-head attention with the bottom temporal transformer to the plurality of token-appended time series sequences, followed by a feed-forward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

[0055] Clause 19: A system according to any one of clauses 11 to 18, wherein the at least one processor is further programmed or configured to: train the at least one space-time graph sandwich transformer according to a Huber loss function, wherein the Huber loss function depends on the predicted time-series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network and the actual time-series traffic data associated with the actual traffic conditions at the multiple next time steps in the traffic network.

[0056] Clause 20: The system of any one of clauses 11 to 19, wherein the transportation network comprises a payment network, wherein the graph comprises a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of payment processing servers, and wherein the plurality of edges are associated with a plurality of connections between the plurality of payment processing servers.

[0057] Clause 21: A computer program product comprising a non-transitory computer-readable medium, the non-transitory computer-readable medium comprising program instructions that, when executed by at least one processor, cause the at least one processor to: obtain a graph representing a traffic network; obtain historical time-series traffic data associated with historical traffic conditions at multiple historical time steps in the traffic network; process the graph and the historical time-series traffic data with each of at least one space-time graph sandwich transformers to generate a sandwich transformer output, wherein each space-time graph sandwich transformer comprises a top time transformer, a space transformer, and a bottom time transformer, the space transformer receiving the output of the top time transformer and the graph as input, and the bottom time transformer receiving the output of the space transformer as input; and generate predicted time-series traffic data associated with predicted traffic conditions at multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer.

[0058] Clause 22: The computer program product of clause 21, wherein the at least one space-time graph sandwich transformer comprises a plurality of space-time graph sandwich transformers.

[0059] Clause 23: A computer program product according to clause 21 or 22, wherein the program instructions, when executed by the at least one processor, cause the at least one processor to process the graph and the historical time-series traffic data with each space-time graph sandwich transformer in parallel with the plurality of space-time graph sandwich transformers to generate the sandwich transformer output.

[0060] Clause 24: A computer program product according to any one of clauses 21 to 23, wherein the program instructions, when executed by the at least one processor, further cause the at least one processor to: provide the historical time series traffic data as input to an input transformation layer comprising a fully connected layer; and receive a hidden time series embedding as output from the input transformation layer, wherein the top time transformer of each space-time graph sandwich transformer receives the hidden time series embedding as input.

[0061] Clause 25: A computer program product according to any one of clauses 21 to 24, wherein the program instructions, when executed by the at least one processor, cause the at least one processor to generate the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer by: concatenating the sandwich transformer output from each space-time graph sandwich transformer; providing the concatenated sandwich transformer output from each space-time graph sandwich transformer as input to a multi-step prediction layer comprising at least two fully connected layers; and receiving the predicted time series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network as output from the multi-step prediction layer.

[0062] Clause 26: A computer program product according to any one of clauses 21 to 25, wherein the program instructions, when executed by the at least one processor, cause the at least one processor to process the graph and the historical temporal traffic data with each space-time graph sandwich transformer to generate the sandwich transformer output by applying a multi-head attention with the top temporal transformer to the input of the top temporal transformer, followed by applying a feedforward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feedforward network.

[0063] Clause 27: A computer program product according to any one of clauses 21 to 26, wherein the program instructions, when executed by the at least one processor, cause the at least one processor to use each space-time graph sandwich transformer to process the graph and the historical time-series traffic data to generate the sandwich transformer output by: applying degree-based encoding and singular value decomposition (SVD)-based encoding to the input of the spatial transformer, the input comprising the output of the top time transformer and the graph; and applying spatial encoding as a bias term in a multi-head attention applied to the input of the spatial transformer, the input comprising the output of the top time transformer and the graph.

[0064] Clause 28: A computer program product according to any one of clauses 21 to 27, wherein the input to the bottom temporal transformer including the output of the spatial transformer comprises a plurality of time series sequences, and wherein the program instructions, when executed by the at least one processor, cause the at least one processor to process the graph and the historical time series traffic data with each space-time graph sandwich transformer to generate the sandwich transformer output by: appending a special token to the beginning of each of the plurality of time series sequences with the bottom temporal transformer; and applying a multi-head attention with the bottom temporal transformer to the plurality of token-appended time series sequences, followed by a feed-forward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

[0065] Clause 29: A computer program product according to any one of clauses 21 to 28, wherein the program instructions, when executed by the at least one processor, further cause the at least one processor to: train the at least one space-time graph sandwich transformer according to a Huber loss function, wherein the Huber loss function depends on the predicted time-series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network and the actual time-series traffic data associated with the actual traffic conditions at the multiple next time steps in the traffic network.

[0066] Clause 30: A computer program product according to any one of clauses 21 to 29, wherein the transportation network comprises a payment network, wherein the graph comprises a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of payment processing servers, and wherein the plurality of edges are associated with a plurality of connections between the plurality of payment processing servers.

[0067] These and other features and characteristics of the present disclosure, as well as the methods of operation and function of the related structural elements and combinations of parts, and the economies of manufacture, will become more apparent upon consideration of the following description and appended claims with reference to the accompanying drawings, all of which form a part of this specification, wherein like reference numerals designate corresponding parts throughout the several views. It is to be expressly understood, however, that the drawings are for purposes of illustration and description only and are not intended as a definition of the limits of the disclosed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Additional advantages and details are explained in more detail below with reference to exemplary embodiments shown in the schematic drawings, in which:

[0069] Figure 1 is an illustration of a non-limiting embodiment or aspect of an environment in which the systems, apparatus, products, devices, and / or methods described herein may be implemented;

[0070] Figure 2 yes Figure 1 illustrations of non-limiting embodiments or aspects of one or more devices and / or components of one or more systems;

[0071] Figure 3 is a flow chart of a method for a space-time graph sandwich transformer (STGST) according to a non-limiting embodiment or aspect;

[0072] Figure 4 The architecture of non-limiting embodiments or aspects of STGST is presented;

[0073] Figure 5 It is the dataset statistics table of the two traffic datasets selected for the experiment;

[0074] Figure 6 is a table including a performance comparison between non-limiting embodiments or aspects of STGST and a baseline;

[0075] Figure 7 is a graph comparing the performance between STGST variants; and

[0076] Figure 8 is a graph showing the training loss and validation loss of STGST on two traffic datasets selected for the experiment; and

[0077] Figure 9 It is a graph showing the impact of model depth and hidden dimensions on STGST. DETAILED DESCRIPTION

[0078] For purposes of the following description, the terms "end," "upper," "lower," "right," "left," "vertical," "horizontal," "top," "bottom," "transverse," "longitudinal," and their derivatives will be used relative to the orientation of the embodiments in the accompanying drawings. However, it will be understood that the embodiments may employ various alternative variations and step orders, except where expressly specified to the contrary. It will also be understood that the specific devices and processes illustrated in the accompanying drawings and described in the following specification are merely exemplary embodiments or aspects of the disclosed subject matter. Accordingly, specific dimensions and other physical characteristics related to the embodiments or aspects disclosed herein should not be considered limiting.

[0079] It should be understood that, unless expressly specified to the contrary, the present disclosure may employ various alternative variations and step sequences. It should also be understood that the specific devices and processes illustrated in the accompanying drawings and described in the following specification are merely exemplary and non-limiting embodiments or aspects. Therefore, specific dimensions and other physical characteristics associated with the embodiments or aspects disclosed herein should not be considered limiting.

[0080] Some non-limiting embodiments or aspects are described herein in conjunction with threshold values. As used herein, satisfying a threshold value may refer to a value greater than a threshold value, more than a threshold value, higher than a threshold value, greater than or equal to a threshold value, less than a threshold value, less than a threshold value, lower than a threshold value, less than or equal to a threshold value, equal to a threshold value, etc.

[0081] As used herein, the aspects, components, elements, structures, actions, steps, functions, instructions, etc. should not be understood as being critical or necessary unless explicitly described as such. Furthermore, as used herein, the article "one" is intended to include one or more items and can be used interchangeably with "one or more" and "at least one". Furthermore, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, a combination of related items and unrelated items, etc.), and can be used interchangeably with "one or more" or "at least one". In the case of wishing only one item, the term "one" or similar language is used. Furthermore, as used herein, the term "having" and / or its analogs are intended to be open terms. Additionally, unless explicitly stated otherwise, the phrase "based on" is intended to mean "at least partially based on". Additionally, reference to an action "based on" a condition may refer to the action being "in response to" the condition. For example, in some non-limiting embodiments or aspects, the phrases "based on" and "in response to" may refer to conditions that automatically trigger an action (e.g., specific operations of electronic devices such as computing devices and processors).

[0082] As used herein, the term "communication" may refer to the reception, acceptance, transmission, transfer, provision, etc. of data (e.g., information, signals, messages, instructions, commands, etc.). A unit (e.g., a device, a system, a component of a device or system, a combination thereof, etc.) communicating with another unit means that the unit is able to directly or indirectly receive information from the other unit and / or send information to the other unit. This may refer to a direct or indirect connection (e.g., a direct communication connection, an indirect communication connection, and / or the like) that is wired and / or wireless in nature. In addition, although the information sent may be modified, processed, relayed, and / or routed between a first unit and a second unit, the two units may also communicate with each other. For example, a first unit may communicate with a second unit even if the first unit passively receives information and does not actively send information to the second unit. As another example, a first unit may communicate with a second unit if at least one intermediate unit processes information received from the first unit and transmits the processed information to the second unit. In some non-limiting embodiments or aspects, a message may refer to a network packet (e.g., a data packet, etc.) comprising data. It should be understood that there may be many other arrangements.

[0083] As used herein, the term "computing device" may refer to one or more electronic devices configured to process data. In some examples, a computing device may include the necessary components to receive, process, and output data, such as a processor, a display, a memory, an input device, a network interface, and / or the like. A computing device may be a mobile device. As examples, a mobile device may include a cellular phone (e.g., a smartphone or a standard cellular phone), a portable computer, a wearable device (e.g., a watch, glasses, lenses, clothing, and / or the like), a personal digital assistant (PDA), and / or other similar devices. A computing device may also be a desktop computer or other form of non-mobile computer.

[0084] As used herein, the term "server" may refer to or include one or more computing devices that are operated by or facilitate communications and processing by multiple parties in a network environment such as the Internet, but it should be understood that communications may be facilitated through one or more public or private network environments, and that various other arrangements may be possible. In addition, multiple computing devices (e.g., servers, point-of-sale (POS) devices, mobile devices, etc.) that communicate directly or indirectly in a network environment may constitute a "system." As used herein, references to a "server" or "processor" may refer to a previously described server and / or processor, a different server and / or processor, and / or a combination of servers and / or processors that are stated to perform a previous step or function. For example, as used in the specification and claims, a first server and / or first processor stated to perform a first step or function may refer to the same or different server and / or processor stated to perform a second step or function.

[0085] As used herein, the term "system" may refer to one or more computing devices or combinations of computing devices (e.g., processors, servers, client devices, software applications, components of such computing devices, etc.). As used herein, references to "device," "server," "processor," and the like may refer to a previously stated device, server, or processor stated as performing a previous step or function, a different server or processor, and / or a combination of servers and / or processors. For example, as used in the specification and claims, a first server or first processor stated as performing a first step or first function may refer to the same or a different server or the same or a different processor stated as performing a second step or second function.

[0086] As used herein, the term "transaction service provider" may refer to an entity that receives transaction authorization requests from merchants or other entities and, in some cases, provides payment assurance through an agreement between the transaction service provider and the issuer organization. For example, a transaction service provider may include, for example The term "transaction processing system" may refer to one or more computing devices operated by or on behalf of a transaction service provider, such as a transaction processing server executing one or more software applications. A transaction processing system may include one or more processors and, in some non-limiting embodiments, may be operated by or on behalf of a transaction service provider.

[0087] As used herein, the term "account identifier" may include one or more primary account numbers (PANs), tokens, or other identifiers associated with a customer account. The term "token" may refer to an identifier that is used as a replacement or substitute identifier for an original account identifier such as a PAN. An account identifier may be alphanumeric or any combination of characters and / or symbols. A token may be associated with a PAN or other original account identifier in one or more data structures (e.g., one or more databases, etc.) such that the token can be used to conduct transactions without directly using the original account identifier. In some instances, an original account identifier such as a PAN may be associated with multiple tokens used for different individuals or purposes.

[0088] As used herein, the terms "issuer institution," "portable financial device issuer," "issuer," or "issuer bank" may refer to one or more entities that provide one or more accounts to a user (e.g., a client, consumer, organization, etc.) to conduct transactions (e.g., payment transactions), such as initiating a credit card payment transaction and / or a debit card payment transaction. For example, an issuer institution may provide an account identifier, such as a PAN, to a user that uniquely identifies one or more accounts associated with the user. The account identifier may be included on a portable financial device, such as a physical financial instrument (e.g., a payment card), and / or may be electronic and used for electronic payments. In some non-limiting embodiments or aspects, the issuer institution may be associated with a bank identification number (BIN) that uniquely identifies the issuer institution. As used herein, the term "issuer institution system" may refer to one or more computer systems operated by or on behalf of the issuer institution, such as a server computer that executes one or more software applications. For example, an issuer institution system may include one or more authorization servers for authorizing payment transactions.

[0089] As used herein, the term "merchant" may refer to an individual or entity that provides goods and / or services or access to goods and / or services to a user (e.g., a customer) based on a transaction (e.g., a payment transaction). As used herein, the term "merchant" or "merchant system" may also refer to one or more computer systems, computing devices, and / or software applications operated by or on behalf of a merchant, such as a server computer that executes one or more software applications. As used herein, a "point of sale (POS) system" may refer to one or more computers and / or peripheral devices used by a merchant to conduct payment transactions with a user, including one or more card readers, near field communication (NFC) receivers, radio frequency identification (RFID) receivers, and / or other contactless transceivers or receivers, contact-based receivers, payment terminals, computers, servers, input devices, and / or other similar devices that can be used to initiate payment transactions. A POS system may be part of a merchant system. A merchant system may also include a merchant plug-in for facilitating online internet-based transactions through a merchant webpage or software application. A merchant plug-in may include software that runs on a merchant server or is hosted by a third party to facilitate such online transactions.

[0090] As used herein, the term "payment device" may refer to a portable financial device, an electronic payment device, a payment card (e.g., a credit or debit card), a gift card, a smart card, smart media, a payroll card, a healthcare card, a wristband, a machine-readable medium containing account information, a keychain device or pendant, an RFID transponder, a retailer discount or loyalty card, a cellular phone, an electronic wallet mobile application, a PDA, a pager, a security card, a computer, an access card, a wireless terminal, a transponder, etc. In some non-limiting embodiments or aspects, a payment device may include volatile or non-volatile memory to store information (e.g., an account identifier, an account holder's name, etc.).

[0091] As used herein, the term "acquirer" may refer to an entity that is licensed and / or approved by a transaction service provider to initiate a transaction using the transaction service provider's portable financial device. An acquirer may also refer to one or more computer systems operated by or on behalf of an acquirer, such as a server computer that executes one or more software applications (e.g., an "acquirer server"). An "acquirer" may be a merchant's bank, or in some cases, a merchant system may be an acquirer. The transactions may include original credit transactions (OCTs) and account fund transactions (AFTs). A transaction service provider may authorize an acquirer to sign a merchant of the service provider to initiate a transaction using the transaction service provider's portable financial device. An acquirer may sign a contract with a payment service provider to enable the service provider to sponsor merchants. An acquirer may monitor the compliance of a payment service provider in accordance with the transaction service provider's regulations. An acquirer may conduct due diligence on a payment service provider and ensure that appropriate due diligence is conducted before signing a sponsored merchant. An acquirer may be responsible for all transaction service provider programs they operate or sponsor. An acquirer may be responsible for the actions of its payment service provider and the merchants it or its payment service provider sponsors.

[0092] As used herein, the term "payment gateway" may refer to an entity and / or a payment processing system operated by or on behalf of such an entity that provides payment services (e.g., transaction service provider payment services, payment processing services, etc.) to one or more merchants (e.g., transaction service provider payment services, payment processing services, etc.). The payment services may be associated with the use of a portable financial device managed by a transaction service provider. As used herein, the term "payment gateway system" may refer to one or more computer systems, computer devices, servers, server groups, etc. operated by or on behalf of a payment gateway.

[0093] By modeling traffic networks as graphs with nodes and edges representing the spatial connectivity of traffic sensors and traffic sensors, spatial-temporal graph models have been intensively studied and have achieved state-of-the-art performance in traffic flow prediction. For example, existing work has explored graph neural networks (GNNs), such as spatial / spectral graph convolutional networks (GCNs) and graph attention networks (GATs), to characterize spatial dependencies by incorporating the inherent structural information of traffic networks. However, in the temporal dimension, existing work can apply recurrent neural networks (RNNs), such as gated recurrent units (GRUs) and long short-term memory networks (LSTMs), or convolution-based sequence learning models, such as temporal convolutional networks (TCNs), to describe the temporal dependencies of time-series traffic data. By integrating outputs from the spatial and temporal domains, existing work can jointly capture spatial-temporal correlations. Although existing spatial-temporal models have demonstrated superior performance over sequence-based models, these existing spatial-temporal models are still limited in at least three aspects: limited-range temporal dependencies, shallow spatial dependencies, and weak spatial-temporal interactions.

[0094] For example, RNNs used in the temporal domain can process sequential inputs incrementally while maintaining past information in a hidden state. However, RNNs face the problem of long-range dependencies. When processing long sequences, the probability of maintaining context distant from the currently processed word decreases exponentially with the distance from the word. The same issues with RNNs that arise in the natural language processing (NLP) domain often occur in traffic prediction scenarios. These limitations lead to the described limited-range temporal dependencies, limiting their use when long-range historical time series data are required. In the spatial domain, GNNs used in existing work typically follow a message-passing scheme that iteratively aggregates neighbor information. However, GNNs have been shown to suffer from oversmoothing problems due to repeated local aggregation (e.g., node representations become indistinguishable and model depth increases), resulting in poor performance in practice. These inherent shortcomings limit the ability of GNNs to learn deep and global spatial features, resulting in the shallow spatial dependencies captured in existing work. Furthermore, most existing methods characterize temporal and spatial dependencies separately and combine them serially or in parallel. These designs can weaken the connection between the spatial and temporal domains and produce weak spatial-temporal interactions.

[0095] Non-limiting embodiments or aspects of the present disclosure may provide methods, systems, and / or computer program products that obtain a graph representing a traffic network; obtain historical time-series traffic data associated with historical traffic conditions at multiple historical time steps in the traffic network; process the graph and the historical time-series traffic data using each of at least one space-time graph sandwich transformer to generate a sandwich transformer output, wherein each space-time graph sandwich transformer includes a top time transformer, a space transformer, and a bottom time transformer, the space transformer receiving the output of the top time transformer and the graph as input, and the bottom time transformer receiving the output of the space transformer as input; and generate predicted time-series traffic data associated with predicted traffic conditions at multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer. For example, non-limiting embodiments or aspects of the present disclosure may provide two time transformers and one space transformer to characterize long-range temporal dependencies and deep spatial dependencies, respectively; and construct the time and space transformers in a sandwich manner to capture prosperous space-time interactions.

[0096] As an example, a non-limiting embodiment or aspect of the present disclosure may provide an STGST for traffic flow prediction, which includes a sandwich transformer cluster, wherein the sandwich transformer cluster includes a group of STGSTs. Each sandwich transformer may include a top temporal transformer and a bottom temporal transformer as a "bun" and a spatial transformer as a "meat". The two temporal transformers can alleviate limited-range temporal dependencies. The temporal input can be processed as a whole, rather than processed step by step in the temporal transformer, making the non-limiting embodiment or aspect less likely to forget information and able to capture long-range temporal dependencies. By equipping the spatial transformer with structural encoding and spatial encoding, the spatial transformer can adapt to shallow spatial dependencies to incorporate graph structure information into the transformer architecture. Applying the designed spatial transformer enables the non-limiting embodiment or aspect to capture deep and global spatial dependencies. To cope with weak space-time interactions, the time and space transformers can be constructed in a "sandwich" manner, which can capture prosperous space-time interactions. Additionally, assembling several sandwich transformers in a cluster can further enhance space-time correlations. Comprehensive experiments on public transportation benchmarks are described below with promising results demonstrating the superior performance of STGST according to non-limiting embodiments or aspects by comparison with ten prior art baselines.

[0097] In this way, non-limiting embodiments or aspects of the present disclosure can enable (i) processing of temporal inputs as a whole, rather than processing them step-by-step in a temporal transformer, making it less likely to forget information and enabling the capture of long-range temporal dependencies; (ii) enabling the capture of deep and global spatial dependencies; and (iii) enabling the capture of flourishing space-time interactions. Furthermore, assembling several sandwich transformers in a cluster can further enhance space-time correlations.

[0098] Now refer to Figure 1 , Figure 1 is an illustration of an exemplary environment 100 in which the apparatus, systems, methods, and / or products described herein may be implemented. Figure 1 As shown in FIG, an environment 100 includes a transaction processing network 101, a user device 112, and / or a communication network 116. The transaction processing network may include a merchant system 102, a payment gateway system 104, an acquirer system 106, a transaction service provider system 108, and an issuer system 110. The transaction processing network 101, the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, the issuer system 110, and / or the user device 112 may be interconnected (e.g., connected to communicate, etc.) via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.

[0099] The merchant system 102 may include one or more devices that are capable of receiving information and / or data from the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, the issuer system 110, and / or the user device 112 (e.g., via a communication network 116, etc.) and / or transmitting information and / or data to the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, the issuer system 110, and / or the user device 112 (e.g., via a communication network 116, etc.). The merchant system 102 may include one or more devices that are capable of communicating with the user device 112 via a communication connection (e.g., an NFC communication connection, an RFID communication connection, The merchant system 102 is a device that receives information and / or data from the user device 112 and / or transmits information and / or data to the user device 112 via the communication connection (e.g., a communication connection). For example, the merchant system 102 may include a computing device, such as a server, a server group, a client device, a client device group, and / or other similar devices. In some non-limiting embodiments or aspects, the merchant system 102 may be associated with a merchant as described herein. In some non-limiting embodiments or aspects, the merchant system 102 may include one or more devices, such as computers, computer systems, and / or peripheral devices, that can be used by a merchant to conduct payment transactions with a user. For example, the merchant system 102 may include a POS device and / or a POS system.

[0100] The payment gateway system 104 may include one or more devices capable of receiving information and / or data from the merchant system 102, the acquirer system 106, the transaction service provider system 108, the issuer system 110, and / or the user device 112 (e.g., via a communication network 116, etc.) and / or transmitting information and / or data to the merchant system 102, the acquirer system 106, the transaction service provider system 108, the issuer system 110, and / or the user device 112 (e.g., via a communication network 116, etc.). For example, the payment gateway system 104 may include a computing device, such as a server, a server group, and / or other similar devices. In some non-limiting embodiments or aspects, the payment gateway system 104 is associated with the payment gateway described herein.

[0101] The acquirer system 106 may include one or more devices capable of receiving information and / or data from the merchant system 102, the payment gateway system 104, the transaction service provider system 108, the issuer system 110, and / or the user device 112 (e.g., via a communication network 116, etc.) and / or transmitting information and / or data to the merchant system 102, the payment gateway system 104, the transaction service provider system 108, the issuer system 110, and / or the user device 112 (e.g., via a communication network 116, etc.). For example, the acquirer system 106 may include a computing device, such as a server, a server group, and / or other similar devices. In some non-limiting embodiments or aspects, the acquirer system 106 may be associated with an acquirer as described herein.

[0102] The transaction service provider system 108 may include one or more devices capable of receiving information and / or data from the merchant system 102, the payment gateway system 104, the acquirer system 106, the issuer system 110, and / or the user device 112 (e.g., via a communication network 116, etc.) and / or transmitting information and / or data to the merchant system 102, the payment gateway system 104, the acquirer system 106, the issuer system 110, and / or the user device 112 (e.g., via a communication network 116, etc.). For example, the transaction service provider system 108 may include computing devices, such as servers (e.g., transaction processing servers, etc.), server farms, and / or other similar devices. In some non-limiting embodiments or aspects, the transaction service provider system 108 may be associated with a transaction service provider as described herein. In some non-limiting embodiments or aspects, the transaction service provider system 108 may include and / or access one or more internal and / or external databases containing transaction data.

[0103] The issuer system 110 may include one or more devices capable of receiving information and / or data from the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the user device 112 (e.g., via a communication network 116, etc.) and / or transmitting information and / or data to the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the user device 112 (e.g., via a communication network 116, etc.). For example, the issuer system 110 may include a computing device, such as a server, a server cluster, and / or other similar devices. In some non-limiting embodiments or aspects, the issuer system 110 may be associated with an issuer organization as described herein. For example, the issuer system 110 may be associated with an issuer organization that issues payment accounts or instruments (e.g., credit accounts, debit accounts, credit cards, debit cards, etc.) to users (e.g., users associated with the user device 112, etc.).

[0104] In some non-limiting embodiments or aspects, the transaction processing network 101 includes multiple systems in a communication path for processing transactions. For example, the transaction processing network 101 may include a merchant system 102, a payment gateway system 104, an acquirer system 106, a transaction service provider system 108, and / or an issuer system 110 in a communication path (e.g., a communication path, a communication channel, a communication network, etc.) for processing electronic payment transactions. For example, the transaction processing network 101 may process (e.g., initiate, conduct, authorize, etc.) electronic payment transactions via the communication paths between the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the issuer system 110.

[0105] The user device 112 may include one or more devices that are capable of receiving information and / or data from the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the issuer system 110 (e.g., via the communication network 116, etc.) and / or transmitting information and / or data to the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the issuer system 110 (e.g., via the communication network 116, etc.). For example, the user device 112 may include a client device, etc. In some non-limiting embodiments or aspects, the user device 112 is capable of communicating with the merchant system 102, the payment gateway system 104, the acquirer system 106, the transaction service provider system 108, and / or the issuer system 110 via a short-range wireless communication connection (e.g., an NFC communication connection, an RFID communication connection, The user device 112 may receive information via a short-range wireless communication connection (e.g., from the merchant system 102, etc.) and / or transmit information via a short-range wireless communication connection (e.g., to the merchant system 102). In some non-limiting embodiments or aspects, the user device 112 may include an application associated with the user device 112, such as an application stored on the user device 112, a mobile application stored and / or executed on the user device 112 (e.g., a mobile device application, a native application for the mobile device, a mobile cloud application for the mobile device, an electronic wallet application, an issuer bank application, etc.). In some non-limiting embodiments or aspects, the user device 112 may be associated with a sender account and / or a recipient account in the payment network for one or more transactions in the payment network.

[0106] The communication network 116 may include one or more wired and / or wireless networks. For example, the communication network 116 may include a cellular network (e.g., Long Term Evolution networks, third generation (3G) networks, fourth generation (4G) networks, fifth generation (5G) networks, code division multiple access (CDMA) networks, etc.), public land mobile networks (PLMNs), local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), telephone networks (e.g., public switched telephone networks (PSTNs), private networks, ad hoc networks, intranets, the Internet, fiber-optic-based networks, cloud computing networks, etc., and / or combinations of these or other types of networks.

[0107] supply Figure 1 The number and arrangement of devices and systems shown are examples. Figure 1 Additional devices and / or systems, fewer devices and / or systems, different devices and / or systems, and / or differently arranged devices and / or systems than those shown. Furthermore, a single device and / or system may be implemented. Figure 1 two or more devices and / or systems shown in , or Figure 1 The single device and / or system shown in the environment 100 may be implemented as multiple distributed devices and / or systems. Additionally or alternatively, one or more devices and / or systems of the environment 100 (e.g., one or more devices or systems) may perform one or more functions described as being performed by another group of devices and / or systems of the environment 100.

[0108] Now refer to Figure 2, a diagram of exemplary components of an apparatus 200 according to a non-limiting embodiment is shown. As an example, apparatus 200 may correspond to merchant system 102, payment gateway system 104, acquirer system 106, transaction service provider system 108, issuer system 110 and / or user device 112. In some non-limiting embodiments, such systems or apparatuses may include at least one apparatus 200 and / or at least one component of apparatus 200. The number and arrangement of components shown are provided as examples. In some non-limiting embodiments, apparatus 200 may include additional components, fewer components, different components, or differently arranged components compared to those shown. Additionally or alternatively, one or more components of apparatus 200 (e.g., one or more components) may perform one or more functions described as being performed by another group of components of apparatus 200.

[0109] like Figure 2 As shown, the device 200 may include a bus 202, a processor 204, a memory 206, a storage component 208, an input component 210, an output component 212, and a communication interface 214. The bus 202 may include components that permit communication between the components of the device 200. In some non-limiting embodiments, the processor 204 may be implemented in hardware, firmware, or a combination of hardware and software. For example, the processor 204 may include a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), etc.), a microprocessor, a digital signal processor (DSP), and / or any processing component that can be programmed to perform a function (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.). The memory 206 may include a random access memory (RAM), a read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, optical memory, etc.) that stores information and / or instructions for use by the processor 204.

[0110] Continue to refer Figure 2, storage component 208 can store information and / or software related to the operation and use of device 200. For example, storage component 208 can include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, a solid-state disk, etc.) and / or another type of computer-readable medium. Input component 210 can include components that allow device 200 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, buttons, switches, a microphone, etc.). Additionally or alternatively, input component 210 can include sensors for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, an actuator, etc.). Output component 212 can include components that provide output information from device 200 (e.g., a display, a speaker, one or more light-emitting diodes (LEDs), etc.). Communication interface 214 can include transceiver-type components (e.g., a transceiver, a separate receiver and transmitter, etc.) that enable device 200 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. The communication interface 214 may permit the device 200 to receive information from another device and / or provide information to another device. For example, the communication interface 214 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, interface, cellular network interface, etc.

[0111] Device 200 can perform one or more processes described herein. Device 200 can perform these processes based on processor 204 executing software instructions stored by a computer-readable medium such as memory 206 and / or storage component 208. The computer-readable medium may include any non-transitory memory device. The memory device includes a memory space located within a single physical storage device or a memory space extended across multiple physical storage devices. The software instructions can be read from another computer-readable medium or from another device into memory 206 and / or storage component 208 via communication interface 214. When executed, the software instructions stored in memory 206 and / or storage component 208 can cause processor 204 to perform one or more processes described herein. Additionally or alternatively, hard-wired circuitry can be used in place of or in combination with software instructions to perform one or more processes described herein. Therefore, the embodiments described herein are not limited to any specific combination of hardware circuitry and software. The term "programmed or configured" as used herein refers to the arrangement of software, hardware circuitry, or any combination thereof on one or more devices.

[0112] Now refer to Figure 3 , a flow chart illustrating a method 300 for a space-time graph sandwich converter according to some non-limiting embodiments or aspects. Figure 3The steps shown are for example purposes only. It will be appreciated that in some non-limiting embodiments or aspects, additional, fewer, different, and / or different orders of steps may be used. In some non-limiting embodiments or aspects, steps may be automatically performed in response to the execution and / or completion of previous steps.

[0113] like Figure 3 As shown, at step 302, method 300 includes obtaining a graph representing a transportation network. For example, transaction service provider system 108 may obtain a graph representing a transportation network.

[0114] The transportation network can be represented as a graph in represents a set of N nodes (e.g., sensors, etc.), and ε is a set of edges indicating the connectivity between nodes. The adjacency matrix derived from the graph can be given by Indicates that if (v i ,v j ), then A ij =1.

[0115] In some non-limiting embodiments or aspects, the transportation network includes a payment network, wherein the graph includes a plurality of edges and a plurality of nodes N for the plurality of edges, the plurality of nodes N being associated with a plurality of payment processing servers, and / or the plurality of edges being associated with a plurality of connections between the plurality of payment processing servers.

[0116] like Figure 3 As shown, at step 304, the method 300 includes obtaining historical time-series traffic data associated with historical traffic conditions at a plurality of historical time steps in the traffic network. For example, the transaction service provider system 108 may obtain historical time-series traffic data associated with historical traffic conditions at a plurality of historical time steps in the traffic network.

[0117] The traffic situation at time step t can be formulated as Where D indicates the number of traffic measurements (e.g., volume, speed, etc.). Given the S-step historical traffic conditions [X (t-S+1) ,…,X (t) ], the prediction model F can be learned to predict the traffic conditions [X' in the future T steps according to the following equation (1) (t+1) ,…,X' (t+T) ]:

[0118]

[0119] in and Represent input and output respectively.

[0120] In some non-limiting embodiments or aspects, the traffic conditions may correspond to the processing of one or more payments at one or more payment processing servers in a payment network.

[0121] like Figure 3 As shown, at step 306, the method 300 includes providing the historical time series traffic data as input to the input transformation layer including the fully connected layer. For example, the transaction service provider system 108 may provide the historical time series traffic data as input to the input transformation layer including the fully connected layer. As an example, and also referring to Figure 4 , Figure 4 The architecture of non-limiting embodiments or aspects of STGST is shown. Figure 4 As shown, STGST can use S-step historical time series traffic data and the traffic network input, and output for the next T steps The STGST can include the following modules: input transformation, sandwich transformer cluster, and multi-step prediction. The input transformation can include a fully connected layer to project low-dimensional traffic data into a high-dimensional information space.

[0122] like Figure 3 As shown, at step 308, the method 300 includes receiving the hidden temporal embedding as an output from the input transformation layer. For example, the transaction service provider system 108 may receive the hidden temporal embedding as an output from the input transformation layer. As an example, the input transformation layer may generate the hidden temporal embedding where S, N, and d represent the sequence length, number of nodes, and hidden dimension, respectively.

[0123] like Figure 3 As shown, at step 310, method 300 includes processing the graph and the historical time-series traffic data with each of the at least one space-time graph sandwich transformer to generate a sandwich transformer output. For example, the transaction service provider system 108 may process the graph and the historical time-series traffic data with each of the at least one space-time graph sandwich transformer to generate the sandwich transformer output. As an example, the transaction service provider system 108 may process the hidden time-series embedding received as an output from the input transformation layer with each of the at least one space-time graph sandwich transformer.

[0124] The transformer architecture can be composed of a set of transformer layers as described in the following paper: Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, AN, Kaiser, L., and Polosukhin, I., “Attention is all you need”, Conference on Neural Information Processing Systems (NIPS) 30 (2017), the entire disclosure of which is hereby incorporated by reference in its entirety. Each transformer layer may include a multi-head attention (MHA) block followed by a point-wise feed-forward (FFN) block, each surrounded by residual connections and layer normalization (LN). Assume is the input of the transformer layer, then the MHA block can be calculated according to the following equations (2) and (3):

[0125]

[0126] in and and is the query, key, and value weight matrix that linearly projects the input X into the h-th attention head.

[0127] Residual connections and layer normalization can be further applied to the output of MHA, expressed as X^=LN(MHA(X)+X). The FFN block can be calculated according to the following equation (4):

[0128]

[0129] Where σ represents the activation function, W1, W2, b1, and b2 are the weight matrices and bias. The final output of the transformer layer is X~=LN(FFN(X^)+X^).

[0130] Now refer to Figure 4 , which shows the architecture of a non-limiting embodiment or aspect of STGST, which can use S-step historical time series traffic data and the traffic network input, and output for the next T steps For predictions, the sandwich transformer cluster may include at least one space-time graph sandwich transformer and / or a group of space-time graph sandwich transformers (e.g., multiple space-time graph sandwich transformers, etc.). Each space-time graph sandwich transformer may include a top time transformer, a spatial transformer, and a bottom time transformer, wherein the spatial transformer receives the output of the top time transformer and the graph as input, and the bottom time transformer receives the output of the spatial transformer as input. For example, each space-time graph sandwich transformer may include two time transformers as "bread" and a spatial transformer as "meat". As an example, each space-time graph sandwich transformer in the sandwich transformer cluster may include a combination of a top time transformer, a spatial transformer, and a bottom time transformer in a "sandwich" manner. The time transformer and the spatial transformer may individually characterize long-range temporal dependencies and deep spatial dependencies, and / or the sandwich combination may capture prosperous space-time interactions. The cluster of sandwich transformers may further strengthen such connections and / or the space-time graph sandwich transformers may be executed in parallel to accelerate the training process. For example, processing the graph and historical time-series traffic data with each spatio-temporal graph sandwich transformer to generate sandwich transformer outputs can be performed in parallel with multiple spatio-temporal graph sandwich transformers. In multi-step prediction, the sandwich transformer cluster outputs can be first concatenated and then fed into two fully connected layers to generate prediction results.

[0131] In some non-limiting embodiments or aspects, processing the graph and historical temporal traffic data with each spatio-temporal graph sandwich transformer to generate a sandwich transformer output includes applying a multi-head attention to the input of the top temporal transformer with a top temporal transformer, followed by applying a feed-forward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feed-forward network. For example, the top temporal transformer can be executed on each node in the traffic network to characterize the long-range temporal dependencies of the temporal input in the temporal dimension. Given the hidden temporal embedding generated by the input transformation module Where S, N and d represent the sequence length, number of nodes and hidden dimension respectively, which can be expressed as XFMR according to the following equation (5): 顶部 The top time transformer module of (·) is formulated as:

[0132]

[0133] in is with have the same size output, and Θ 顶部 is a trainable parameter. In this module, the dimension N can be regarded as the batch size.

[0134] The design of XFMR(·) can be built on top of the transformer structure described in the paper: Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, AN, Kaiser, L., and Polosukhin, I. “Attention is all you need”, Conference on Neural Information Processing Systems (NIPS) 30 (2017), the entire disclosure of which is hereby incorporated by reference in its entirety. It is from The nodes obtained in the calculation of MHA can be time-coded Added to the hidden temporal embedding H to incorporate the time-dependent factors of the temporal input. 时间 Each element p(t,i) in (1≤t≤S and 1≤i≤d) can be derived from a frequency encoding function that characterizes a time-dependent sinusoid. For example, p(t,i) = sin(t / 10000 2i / / d )(if i is even) or cos(t / 10000 2i / d )(if i is odd).

[0135] By adding temporal encoding to the latent embeddings at different time steps, the latent embeddings can become temporally discriminative and / or can follow the Transformer Encoder architecture to learn temporal dependencies, which can be formulated according to the following equation (6):

[0136]

[0137] The output of each node is collected to form

[0138] In some non-limiting embodiments or aspects, processing the graph and the historical time-series traffic data with each space-time graph sandwich transformer to generate the sandwich transformer output includes: applying degree-based encoding and singular value decomposition (SVD)-based encoding to the input of the spatial transformer with the spatial transformer, the input including the output of the top time transformer and the graph; and applying spatial encoding as a bias term in the multi-head attention applied to the input of the spatial transformer with the spatial transformer, the input including the output of the top time transformer and the graph. For example, unlike the time transformer, the spatial transformer can be performed at each time step with the goal of describing the global spatial dependencies across all nodes in the traffic network at this time step. Given the traffic network and the last component from Output of the spatial transformer XFMR 空间(·) can be defined according to the following equation (7):

[0139]

[0140] The output With the same size, and Θ 空间 is a learnable parameter. When describing the spatial dependencies of N nodes, the dimension S can be regarded as the batch size in the spatial transformer.

[0141] Based on the observations described in the following paper: Min, E., Chen, R., Bian, Y., Xu, T., Zhao, K., Huang, W., Zhao, P., Huang, J., Ananiadou, S. and Rong, Y., entitled "Transformer for graphs: An overview from architecture perspective", arXiv preprint arXiv:2202.08455 (2022), the entire disclosure of which is hereby incorporated by reference in its entirety, non-limiting embodiments or aspects may introduce two strategies to encode graph structure information into the transformer architecture. For example, non-limiting embodiments or aspects may add degree-based encoding and singular value decomposition (SVD)-based encoding to the input, and add spatial encoding as a bias term in the MHA module. By performing these two structure-preserving strategies, the transformer layer can adaptively adjust the attention coefficient according to the graph structure information. For example, the position encoding E added to the input can be made according to the following equation (8) 结构 formulation:

[0142] E 结构 =E 度 +E svd , (8)

[0143] in Denotes degree-based coding and SVD-based coding. 度 It can be expressed as E 入 and E 出 The sum of E, which can be a learnable matrix identified by in-degree and out-degree respectively. svd can be added to further distinguish two nodes with the same degree and can be formed using the largest r singular value and the corresponding left and right singular vectors according to the following equation (9):

[0144]

[0145] in Include r left and right singular vectors, respectively, corresponding to the diagonal matrix The top r singular values in , representing the concatenation operator along the columns.

[0146] The spatial encoding E can be defined based on the concept of the shortest path spd . E spd Each element in can represent the shortest path distance (SPD) between two corresponding nodes. According to the following equation (10), the spatial encoding can act as a bias term in the MHA module:

[0147]

[0148] Before model training, the in-degree and out-degree of each node, the SVD vector, and the SPD between two connected nodes can be pre-computed, which may not have much impact on the training time and resources used. The spatial dependence can be formulated according to the following equation (11):

[0149]

[0150] The output of stacking S time steps can be obtained

[0151] In some non-limiting embodiments or aspects, the input to the bottom temporal transformer including the output of the spatial transformer includes a plurality of time series sequences, and processing the graph and the historical time series traffic data with each space-time graph sandwich transformer to generate the sandwich transformer output includes: appending a special token to the beginning of each of the plurality of time series sequences with the bottom temporal transformer; and applying a multi-head attention with the bottom temporal transformer to the plurality of token-appended time series sequences, followed by a feed-forward network, wherein there are residual connections and layer normalization around each of the multi-head attention and the feed-forward network. For example, another temporal transformer can be connected to the spatial transformer to capture space-time interactions. As input, the bottom time transformer XFMR 底部 (·) can be defined according to the following equation (12):

[0152]

[0153] in And Θ 底部is a trainable parameter of the bottom temporal transformer module. The overall design of the bottom temporal transformer module can be the same or similar to that of the top temporal transformer, including temporal encoding, MHA, and FFN; however, in the bottom temporal transformer module, a special token [SEQ] can be appended at the beginning of each temporal sequence to summarize the sequence-level representation.

[0154] like Figure 3 As shown, at step 312, the method 300 includes generating predicted time-series traffic data associated with predicted traffic conditions at a plurality of next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer. For example, the transaction service provider system 108 may generate predicted time-series traffic data associated with predicted traffic conditions at a plurality of next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer.

[0155] In some non-limiting embodiments or aspects, generating the predicted time-series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer includes: concatenating the sandwich transformer output from each space-time graph sandwich transformer; providing the concatenated sandwich transformer output from each space-time graph sandwich transformer as input to a multi-step prediction layer including at least two fully connected layers; and receiving the predicted time-series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network as output from the multi-step prediction layer. For example, and again with reference to Figure 4 , can be obtained from Get [SEQ] token H 序列 The embedding of is taken as the output of each sandwich transformer. By cascading the outputs of the sandwich transformer cluster consisting of K space-time graph sandwich transformers, in A two-layer MLP can be provided as the output layer for final prediction.

[0156] like Figure 3As shown, at step 314, the method 300 includes training at least one space-time graph sandwich transformer according to a Huber loss function. For example, the transaction service provider system 108 can train at least one space-time graph sandwich transformer (e.g., train STGST, etc.) according to the Huber loss function, wherein the Huber loss function depends on the predicted time series traffic data associated with the predicted traffic conditions at multiple next time steps in the traffic network and the actual time series traffic data associated with the actual traffic conditions at multiple next time steps in the traffic network. As an example, the Huber loss defined according to the following equation (13) can be selected as the loss function because the Huber loss is less sensitive to outliers than the squared error loss to outliers:

[0157]

[0158] where у is the actual traffic volume, and δ controls the sensitivity to outliers.

[0159] like Figure 3 As shown, at step 316, the method 300 includes providing at least one trained space-time graph sandwich transformer. For example, the transaction service provider system 108 may provide at least one trained space-time graph sandwich transformer (e.g., provide a trained STGST, etc.). As an example, the transaction service provider system 108 may store (e.g., in a memory, in a database, etc.) the trained at least one space-time graph sandwich transformer (e.g., the trained STGST, etc.). In such an example, the transaction service provider system 108 may apply the trained STGST to the S-step historical time series traffic data. and traffic network input, and receives traffic data or traffic conditions from STGST for the next T steps The output is a prediction of (e.g., upcoming traffic conditions in a payment network, etc.).

[0160] experiment

[0161] In this section, the performance of non-limiting embodiments or aspects of STGST is compared with the state-of-the-art baseline on a public transportation network benchmark, with the goal of answering the following research questions: RQ1. How does STGST perform on the traffic flow forecasting task?; RQ2. How does each component of STGST contribute to the prediction?; and RQ3. How stable is the learning of STGST? And how do certain hyperparameters affect the model performance?

[0162] Non-limiting embodiments or aspects of STGST can be validated in real time every 30 seconds on real-world freeway traffic data originally collected by the California Department of Transportation (Caltrans) Performance Measurement System (PeMS), as described in the following paper: Chen, C., Petty, K., Skabardonis, A., Varaiya, P. and Jia, Z. entitled "Freeway performance measurement system: mining loop detector data", Transportation Research Record 1748(1), 96–102 (2001), the entire disclosure of which is hereby incorporated by reference in its entirety. Figure 5 This is a table of dataset statistics for two traffic datasets selected for the experiment, PeMS04 and PeMS08, which have been widely studied in the field of traffic flow forecasting and are described and published in the following paper: Song, C., Lin, Y., Guo, S. and Wan, H. entitled "Spatial-temporal synchronous graph convolutional networks: A new framework for spatial-temporal network data forecasting", Association for the Advancement of Artificial Intelligence (AAAI). Vol. 34, pp. 914-921 (2020), the entire disclosure of which is hereby incorporated by reference in its entirety. Traffic data (including flow, speed and occupancy) are aggregated every 5 minutes, and Z-score normalization is used to standardize the data input. Spatial relationships are constructed based on the actual road network.

[0163] Non-limiting embodiments or aspects of STGST are compared with three types of prior art baselines, including RNN-based models, convolution-based models, and spatio-temporal graph-based models. LSTM and TCN are selected for the first two categories, and STG2Seq, STGCN, DCRNN, GraphWaveNet, STSGCN, ASTGCN, MSTGCN, and STGODE are selected for the third category.

[0164] Following the standard benchmark setup in this field, each dataset was split chronologically into training, validation, and test sets with a ratio of 6:2:2. Using one hour of historical data to predict the next hour's data means using the past 12 consecutive time steps to predict the next 12 consecutive time steps. Mean absolute error (MAE), mean absolute percentage error (MAPE), and root mean square error (RMSE) were used to measure performance.

[0165] The non-limiting embodiments or aspects of STGST can be implemented using Python 3.9.13, Py-Torch 1.12.1 and DGL (Deep Graph Library (Deep Graph Library)) 0.9.1. All experiments are carried out on a Linux server equipped with Intel (R) Xeon (R) Gold 6130CPU and NVIDIA Tesla V100 GPU. According to the figure ratio in each data set, for PEMS04 and PEMS08, hidden dimension is set to 128, the number of attention heads is set to 8, and the loss rate is set to 0.1, and the number of interlayer transformers is set to 2, and the number of time transformer layer and space transformer layer is set to 4 and 3 respectively. STGST is trained using Adam optimizer, for 200 periods (epoch), batch size is 32, learning rate is 3e-4, and weight decay is 5e-4. For reproducibility, random seed is set to 0.

[0166] Figure 6 is a table including performance comparisons between non-limiting embodiments or aspects of STGST and a baseline. Figure 6 The prediction errors of STGST and ten baselines on the PEMS04 and PEMS08 datasets are shown, with the best and second-best results for each dataset highlighted in bold and underlined, respectively. Comparing the results of LSTM and TCN with the other baselines shows that space-time graph models that consider both temporal and spatial dependencies perform better than sequence models that only capture temporal information. This further confirms the benefits of spatial dependencies in traffic flow prediction. The results also show that STGST consistently outperforms all baselines on both datasets in terms of the three evaluation metrics. Figure 6 The last row of the table in

[15] shows the improvement of STGST over the most competitive baseline (i.e., STGODE). On average, STGST outperforms STGODE by 3.28%, 1.61%, and 2.01% on MAE, MAPE, and RMSE, respectively. The superior performance comes from a well-designed space-time graph intercalation transformer that is able to capture long-range temporal dependencies and deep spatial dependencies and describe prosperous space-time interactions.

[0167] To verify the contribution of each module of STGST in the prediction task, three STGST variants were prepared and ablation studies were performed on two datasets. The "no time" variant removed the two temporal transformers and added an average pooling layer to the output of the spatial transformer. The "no space" variant removed the spatial transformer module and directly connected the two temporal transformers. The "no encoding" variant removed the spatial position encoding module in the spatial transformer, such as the degree encoder and SVD encoder.

[0168] Figure 7 is a graph comparing the performance of STGST variants. Figure 7 As shown, STGST equipped with all designed components outperforms its three variants, demonstrating that each component contributes to the final performance. By comparing STGST with the timeless and spaceless variants, it can be seen that ignoring the spatial or temporal domains significantly degrades the performance. In addition, the spaceless variant performs slightly better than the timeless variant, indicating that temporal dependencies are more advantageous than spatial dependencies in traffic flow prediction. In addition, spatial position encoding helps capture spatial dependencies, which can be observed from the results of the codeless variant and STGST.

[0169] We now investigate the learning stability of STGST. Figure 8 The training and validation loss curves for STGST are plotted on two traffic datasets selected for experimentation. Analyzing the curves for both datasets, we note that STGST converges quickly at epoch 8 and epoch 5 for PEMS04 and PEMS08, respectively, and that the loss decreases steadily with increasing epochs. This finding indicates that STGST's learning procedure is stable. Furthermore, the validation loss curve is slightly higher than the training loss curve, indicating that the trained model generalizes well to the validation set and that learning does not suffer from overfitting or underfitting.

[0170] We now investigate how different parameter choices affect model performance. We vary the model depth (e.g., the number of transformer encoders, etc.) from 1 to 5 to investigate the performance of STGST on the PEMS04 dataset. The experimental results are shown in Figure 9 (Left) Middle. As can be seen from the curve, as the model depth increases, the performance first improves and then begins to gradually decline. The best results are obtained when the model depth is set to 4 on PEMS04. The embedding dimension can be varied from 16 to 256 to study its impact. Figure 9 (Right) shows that the impact of the latent dimension has a similar pattern to the model depth, and the best performance is achieved when setting the latent dimension to 128.

[0171] Therefore, non-limiting embodiments or aspects of the present disclosure may provide an STGST comprising an input transformation module that projects a temporal input into a high-dimensional feature space, a sandwich transformer cluster that processes temporal dependencies and spatial dependencies and captures space-time interactions, and a multi-step prediction module that maps hidden representations to an output space. The sandwich transformer cluster may be a plurality of space-time graph sandwich transformers, each of which contains a top time transformer and a bottom time transformer as "bread" and a spatial transformer as "meat," which enables comprehensive capture of space-time interactions, thereby providing superior performance over prior art baselines.

[0172] Although the embodiments have been described in detail for purposes of illustration, it should be understood that such detail is solely for that purpose and that the present disclosure is not limited to the disclosed embodiments or aspects, but, on the contrary, is intended to cover modifications and equivalent arrangements within the spirit and scope of the appended claims. For example, it should be understood that the present disclosure contemplates that, to the extent possible, one or more features of any embodiment or aspect can be combined with one or more features of any other embodiment or aspect.

Claims

1. A method comprising: obtaining, with at least one processor, a graph representing a transportation network; obtaining, with the at least one processor, historical time-series traffic data associated with historical traffic conditions at a plurality of historical time steps in the traffic network; processing the graph and the historical time-series traffic data with the at least one processor using each of at least one space-time graph sandwich transformer to generate a sandwich transformer output, wherein each space-time graph sandwich transformer includes a top time transformer, a space transformer, and a bottom time transformer, the space transformer receiving as input the output of the top time transformer and the graph, and the bottom time transformer receiving as input the output of the space transformer; as well as Generate, with the at least one processor, predicted time-series traffic data associated with predicted traffic conditions at a plurality of next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer. 2 . The method of claim 1 , wherein the at least one space-time pattern sandwich converter comprises a plurality of space-time pattern sandwich converters.

3. The method of claim 2, wherein said processing the graph and the historical time-series traffic data with said at least one processor using each space-time graph sandwich transformer to generate said sandwich transformer output is performed in parallel with said plurality of space-time graph sandwich transformers.

4. The method according to claim 2, further comprising: providing, using the at least one processor, the historical time-series traffic data as input to an input transformation layer comprising a fully connected layer; as well as The at least one processor receives as output a hidden temporal embedding from the input transformation layer, wherein the top temporal transformer of each spatio-temporal graph sandwich transformer receives as input the hidden temporal embedding.

5. The method of claim 2 , wherein generating, with the at least one processor, the predicted time-series traffic data associated with the predicted traffic conditions at the plurality of next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer comprises: cascading the sandwich transformer output from each space-time graph sandwich transformer; providing as input a multi-step prediction layer comprising at least two fully connected layers the concatenated sandwich transformer outputs from each of the spatio-temporal graph sandwich transformers; as well as The predicted time-series traffic data associated with the predicted traffic conditions at the plurality of next time steps in the traffic network is received as output from the multi-step prediction layer.

6. The method of claim 1 , wherein processing the graph and the historical time-series traffic data with each space-time graph sandwich transformer with the at least one processor to generate the sandwich transformer output comprises: A multi-head attention is applied to the input of the top temporal transformer with the top temporal transformer, followed by a feed-forward network, with residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

7. The method of claim 1 , wherein processing the graph and the historical time-series traffic data with each space-time graph sandwich transformer with the at least one processor to generate the sandwich transformer output comprises: applying, with the spatial transformer, degree-based coding and singular value decomposition (SVD)-based coding to the input of the spatial transformer, the input comprising the output of the top temporal transformer and the map; and The spatial transformer is used to apply spatial encoding as a bias term in a multi-head attention applied to the input of the spatial transformer, the input comprising the output of the top temporal transformer and the map.

8. The method of claim 1 , wherein the input to the bottom temporal transformer, including the output of the spatial transformer, comprises a plurality of time-series sequences, and wherein processing the graph and the historical time-series traffic data with each space-time graph sandwich transformer with the at least one processor to generate the sandwich transformer output comprises: appending a special token to the beginning of each of the plurality of time sequence sequences using the bottom time transformer; as well as Multi-head attention is applied to the temporal sequence of multiple additional tokens using the bottom temporal transformer, followed by a feed-forward network with residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

9. The method according to claim 1, further comprising: The at least one space-time graph sandwich transformer is trained with the at least one processor according to a Huber loss function, wherein the Huber loss function depends on the predicted time-series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network and the actual time-series traffic data associated with the actual traffic conditions at the multiple next time steps in the traffic network.

10. The method of claim 1 , wherein the transportation network comprises a payment network, wherein the graph comprises a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of payment processing servers, and wherein the plurality of edges are associated with a plurality of connections between the plurality of payment processing servers.

11. A system comprising: at least one processor coupled to the memory and programmed or configured to: Obtaining a graph representing a transportation network; Obtaining historical time-series traffic data associated with historical traffic conditions at a plurality of historical time steps in the traffic network; processing the graph and the historical time-series traffic data with each of at least one space-time graph sandwich transformer to generate a sandwich transformer output, wherein each space-time graph sandwich transformer comprises a top time transformer, a space transformer, and a bottom time transformer, the space transformer receiving as input the output of the top time transformer and the graph, and the bottom time transformer receiving as input the output of the space transformer; and Predicted time-series traffic data associated with predicted traffic conditions at a plurality of next time steps in the traffic network is generated based on the sandwich transformer output from each space-time graph sandwich transformer.

12. The system of claim 11 , wherein the at least one space-time graph sandwich transformer comprises a plurality of space-time graph sandwich transformers, wherein the at least one processor is programmed or configured to process the graph and the historical time-series traffic data with each space-time graph sandwich transformer in parallel with the plurality of space-time graph sandwich transformers to generate the sandwich transformer output.

13. The system of claim 12, wherein the at least one processor is further programmed or configured to: Providing the historical time-series traffic data as input to an input transformation layer including a fully connected layer; and A hidden temporal embedding is received as output from the input transformation layer, wherein the top temporal transformer of each spatio-temporal graph sandwich transformer receives the hidden temporal embedding as input.

14. The system of claim 12 , wherein the at least one processor is programmed or configured to generate the predicted time-series traffic data associated with the predicted traffic conditions at the plurality of next time steps in the traffic network based on the sandwich transformer output from each space-time graph sandwich transformer by: cascading the sandwich transformer output from each space-time graph sandwich transformer; providing as input a multi-step prediction layer comprising at least two fully connected layers the concatenated sandwich transformer outputs from each spatio-temporal graph sandwich transformer; and The predicted time-series traffic data associated with the predicted traffic conditions at the plurality of next time steps in the traffic network is received as output from the multi-step prediction layer.

15. The system of claim 11 , wherein the at least one processor is programmed or configured to, with each space-time graph sandwich transformer, process the graph and the historical time-series traffic data to generate the sandwich transformer output by: A multi-head attention is applied to the input of the top temporal transformer with the top temporal transformer, followed by a feed-forward network, with residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

16. The system of claim 11 , wherein the at least one processor is programmed or configured to, with each space-time graph sandwich transformer, process the graph and the historical time-series traffic data to generate the sandwich transformer output by: applying, with the spatial transformer, degree-based coding and singular value decomposition (SVD)-based coding to the input of the spatial transformer, the input comprising the output of the top temporal transformer and the map; and The spatial transformer is used to apply spatial encoding as a bias term in a multi-head attention applied to the input of the spatial transformer, the input comprising the output of the top temporal transformer and the map.

17. The system of claim 11 , wherein the inputs to the bottom temporal transformer, including the outputs of the spatial transformers, comprise a plurality of time-series sequences, and wherein the at least one processor is programmed or configured to, with each space-time graph sandwich transformer, process the graph and the historical time-series traffic data to generate the sandwich transformer outputs by: appending a special token to the beginning of each of the plurality of time sequence sequences using the bottom time transformer; and Multi-head attention is applied to the temporal sequence of multiple additional tokens using the bottom temporal transformer, followed by a feed-forward network with residual connections and layer normalization around each of the multi-head attention and the feed-forward network.

18. The system of claim 11, wherein the at least one processor is further programmed or configured to: The at least one space-time graph sandwich transformer is trained according to a Huber loss function, which depends on the predicted time-series traffic data associated with the predicted traffic conditions at the multiple next time steps in the traffic network and the actual time-series traffic data associated with the actual traffic conditions at the multiple next time steps in the traffic network.

19. The system of claim 11, wherein the transportation network comprises a payment network, wherein the graph comprises a plurality of edges and a plurality of nodes for the plurality of edges, wherein the plurality of nodes are associated with a plurality of payment processing servers, and wherein the plurality of edges are associated with a plurality of connections between the plurality of payment processing servers.

20. A computer program product comprising a non-transitory computer-readable medium comprising program instructions that, when executed by at least one processor, cause the at least one processor to: Obtaining a graph representing a transportation network; Obtaining historical time-series traffic data associated with historical traffic conditions at a plurality of historical time steps in the traffic network; processing the graph and the historical time-series traffic data with each of at least one space-time graph sandwich transformer to generate a sandwich transformer output, wherein each space-time graph sandwich transformer comprises a top time transformer, a space transformer, and a bottom time transformer, the space transformer receiving as input the output of the top time transformer and the graph, and the bottom time transformer receiving as input the output of the space transformer; and Predicted time-series traffic data associated with predicted traffic conditions at a plurality of next time steps in the traffic network is generated based on the sandwich transformer output from each space-time graph sandwich transformer.