Training method and device, spatiotemporal data prediction method and device
By using the cross-scale Koopman operator to extract and fuse multi-dimensional features in spatiotemporal systems, the problem of inaccurate prediction under a single perspective is solved, and more accurate prediction of the future state of spatiotemporal systems is achieved.
Patent Information
- Application Number
- CN202510947249.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Existing technologies typically employ a single perspective when predicting the future state of spatiotemporal systems, resulting in inaccurate predictions and difficulty in capturing causal emergence and dynamic characteristics at different spatial scales and levels.
We employ cross-scale Koopman operators to map observation data of spatiotemporal systems into a representation space. By using globally shared, community-specific, and node-specific Koopman operators, we extract and fuse predictive feature variables to construct a stable representation space and improve prediction accuracy.
By extracting and fusing multi-dimensional features, the state of a spatiotemporal system can be described more comprehensively, improving the accuracy of predicting the future state of the spatiotemporal system.
Smart Images

Figure CN120449972B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data prediction, in particular to a training method and device, and a spatio-temporal data prediction method and device. BACKGROUND
[0002] Real-world spatio-temporal systems, such as traffic networks, atmospheric systems, and epidemiology, are intuitive reflections of the real world, implying the basic laws of natural behavior, and the state of the spatio-temporal system will change continuously with time and space.
[0003] Most current researches analyze spatio-temporal systems from a single perspective and predict the future state of the spatio-temporal system, but if only limited to a specific observation dimension, it may lead to inaccurate prediction of the future state of the spatio-temporal system. SUMMARY
[0004] The purpose of the present application is to provide a training method and device, and a spatio-temporal data prediction method and device, which can improve the accuracy of predicting the future state of the spatio-temporal system.
[0005] To achieve the above-mentioned purpose, the present application provides the following solutions:
[0006] In a first aspect, the present application provides a training method, comprising:
[0007] obtaining samples; each sample includes a spatio-temporal observation data sequence in an observation space; inputting the samples into a spatio-temporal prediction network for training to obtain a prediction model, wherein the spatio-temporal prediction network includes a mapping function, an inverse mapping function, and a cross-scale Koopman operator; the spatio-temporal prediction network and the prediction model are used to perform spatio-temporal data prediction according to input data; the spatio-temporal data prediction includes: mapping the input data to a first feature space by the mapping function to obtain at least a feature variable of a current time step; predicting the feature variable of the current time step by a global shared operator, a community-specific operator, and a node-specific operator in the cross-scale Koopman operator respectively to obtain a global predicted feature variable, a community predicted feature variable, and a node predicted feature variable of a next time step, and fusing them to obtain a predicted feature variable of the next time step; mapping the predicted feature variable of the next time step to the observation space by the inverse mapping function to obtain a prediction result; the prediction result includes predicted spatio-temporal observation data of the next time step.
[0008] In a second aspect, the present application provides a spatio-temporal data prediction method, comprising:
[0009] obtaining target spatio-temporal observation data; inputting the target spatio-temporal observation data into a prediction model trained by the training method according to any one of the first aspect to obtain predicted spatio-temporal observation data within a target time span.
[0010] According to specific embodiments provided in the present application, the present application discloses the following technical effects:
[0011] In the present disclosure, the global shared operator, the community-specific operator and the node-specific operator in the cross-scale Koopman operator are used to respectively predict the characteristic variables of the current time step, to obtain predicted characteristic variables at different levels, and then to fuse them. This multi-dimensional feature extraction and fusion method can comprehensively consider the characteristics of the spatio-temporal system at different scales and levels, capture more comprehensive information, and thus more accurately describe the state of the spatio-temporal system. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is a schematic diagram of observation sequences of the same spatio-temporal system at different perspectives in an embodiment of the present application;
[0013] Figure 2 is a flowchart of a training method according to an exemplary embodiment;
[0014] Figure 3 is an architecture diagram of a spatio-temporal prediction network in an embodiment of the present application;
[0015] Figure 4 is a schematic diagram of mapping by a mapping function in an embodiment of the present application;
[0016] Figure 5 is a schematic diagram of prediction and fusion by a cross-scale Koopman operator in an embodiment of the present application;
[0017] Figure 6 is a schematic diagram of mapping by an inverse mapping function in an embodiment of the present application;
[0018] Figure 7 is a schematic diagram of linear decomposition of a Koopman operator in an embodiment of the present application;
[0019] Figure 8 is a distribution of eigenvalues in a Koopman operator at different scales and a corresponding causal graph in an embodiment of the present application;
[0020] Figure 9 is a schematic diagram of reconstructing an observation data sequence according to an invariant causal pattern in an embodiment of the present application;
[0021] Figure 10 is a schematic diagram of a third training phase in an embodiment of the present application;
[0022] Figure 11 is a functional module schematic diagram of a training device provided in an embodiment of the present application;
[0023] Figure 12 is a flow chart of a spatio-temporal data prediction method according to an exemplary embodiment;
[0024] Figure 13 a functional module schematic diagram of a null data prediction device provided for an embodiment of the present application;
[0025] Figure 14 a comparison result schematic diagram provided for an embodiment of the present application. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.
[0027] Real-world spatio-temporal systems, such as traffic networks, atmospheric systems, and epidemiological systems, are intuitive reflections of the real world, implying the basic laws of natural behavior. Exploring the internal causal relationships and dynamic characteristics of spatio-temporal systems is crucial for predicting future states and achieving reliable control. Spatio-temporal representation learning, as the core of understanding spatio-temporal systems, aims to map observable and redundant spatio-temporal data into a low-rank and robust latent representation space, i.e., to map data in the observation space to the representation space and vectorize data in the representation space, and may also involve causal relationship construction. In the prediction of spatio-temporal data through a prediction model, spatio-temporal representation learning is a core technical means for constructing a stable representation space, and it has a close dependence relationship with the prediction model. Constructing a stable representation space not only breaks through the shackles of the curse of dimensionality, but also enhances the model generalization ability, solves the problem of data sparsity, and significantly improves the prediction accuracy. Currently, this technology has been widely studied in many tasks such as spatio-temporal sequence prediction and anomaly detection of urban traffic flow and road speed.
[0028] A spatio-temporal system can be represented as a spatio-temporal network, in which nodes represent certain regions, and the attributes of nodes represent observable spatio-temporal observation data sequences of regions over time, and the connections between nodes represent the interactions between regions. Due to the limitations of the macro level, real-world spatio-temporal observation data sequences (such as traffic flow and road speed) usually do not have explicit causal relationships, but are generated by latent causal hidden variables or causally related confounding factors. Changes in data distribution are due to changes in causal hidden variables behind the data, so the observed data depend on the nonlinear mixing of causal hidden variables. Causal hidden variables describe phenomena that cannot be directly observed but actually affect the travel of people in the region, such as regional attributes, economic conditions, education levels, and job-to-residence ratios.
[0029] In the context of spatiotemporal representation learning, accurately modeling and inferring causal latent variables and causal relationships can generate a more robust representation space, particularly demonstrating better generalization to out-of-distribution samples. This indicates that the model overcomes selection bias and statistical shortcuts caused by correlation. However, methods that solely track causal latent variables are insufficient to explain complex urban phenomena. Further exploration of the dynamic interactions between causal latent variables is needed to form nonlinear spatiotemporal dynamics among variables, which describe the causal structure that changes over time.
[0030] Meanwhile, due to the non-steady, nonlinear, and cross-scale emergent characteristics of the causal structure of spatiotemporal dynamic systems, the causal structure between latent causal variables is difficult to fully describe. Compared to single time-series data, the spatial dimension of spatiotemporal data exhibits emergent properties, meaning that self-organization at the microscopic level may manifest stronger causal relationships at the macroscopic level. Furthermore, the observed phenomena of spatiotemporal systems at different spatial scales also show significant differences; that is, observational sequences within a spatiotemporal system can differ considerably at the micro, meso, and macroscopic levels. Figure 1 As shown, taking the micro-level as the node perspective, the meso-level as the community perspective, and the macro-level as the global perspective as an example, from... Figure 1 As can be seen, the observation sequences of the same spatiotemporal system on the left exhibit significant differences from the global perspective, the community perspective, and the node perspective. This makes it difficult for traditional methods that capture distribution changes at different spatial scales under the same perspective to uncover causal emergence at different spatial scales, which may lead to inaccurate predictions of the future state of the spatiotemporal system.
[0031] To address the aforementioned technical problems, this disclosure proposes a training method, a training device, a spatiotemporal data prediction method, and a spatiotemporal data prediction device. The training method and device are used to train a prediction model. The functions of the spatiotemporal data prediction method and device are implemented based on the prediction model.
[0032] Figure 2 This is a flowchart illustrating a training method according to an exemplary embodiment, such as... Figure 2 As shown, the method includes the following steps S101-S102:
[0033] In step S101, samples are acquired; each sample includes a sequence of spatiotemporal observation data in the observation space.
[0034] The content of the samples may vary in different application areas. For example, in applications that predict urban traffic flow data (such as average road speed) and detect anomalies in urban traffic flow data, the samples may include spatiotemporal observation data sequences of urban spatiotemporal systems (e.g., one sample corresponds to one time step, and each sample includes the average road speed of different roads at each sampling time at the same time step).
[0035] Exemplarily, the spatio-temporal data of the region in the city that needs to be predicted can be collected, and the spatio-temporal data is quantified to obtain the time period data set of each region, that is, the spatio-temporal observation data sequence in the observation space.
[0036] In step S102, the sample is input into the spatio-temporal prediction network to obtain a prediction model.
[0037] As shown in the figure, the spatio-temporal prediction network includes a spatio-temporal synchronization network and a dynamics layer in architecture, the spatio-temporal synchronization network includes an encoder and a decoder, wherein the encoder fits a mapping function, the dynamics layer fits a cross-scale Koopman operator, and the decoder fits an inverse mapping function, and each data flow in the figure will be described in detail below. Figure 3 Figure 3
[0038] The spatio-temporal prediction network and the prediction model are used to perform spatio-temporal data prediction according to input data. The spatio-temporal prediction network is used to perform spatio-temporal data prediction according to input data in the training stage, and the prediction model is used to perform spatio-temporal data prediction according to input data in the inference stage, that is, the application of the trained spatio-temporal prediction network.
[0039] In one example, performing spatio-temporal data prediction includes the following sub-steps A1-A3:
[0040] A1, mapping the input data into a first feature space by a mapping function to obtain at least a feature variable of a current time step.
[0041] It should be noted that when the spatio-temporal prediction network performs spatio-temporal data prediction according to input data, the input data is the sample in S101, and when the prediction model performs spatio-temporal data prediction according to input data, the input data is the spatio-temporal observation data sequence in the observation space to be predicted.
[0042] Consider a spatio-temporal system affected by a controlled input , , denotes the time span; denotes the control input or control variable with a time span of , denotes the control input or control variable at any time step t ; N denotes the number of spatial nodes, T +1 denotes the total length of time, m denotes the dimension of the control information.
[0043] Taking the application field of predicting urban traffic flow data (such as road average speed) as an example, This can include variables such as the inherent attributes of each region (spatial node), weather conditions, and regional gathering activities. The above-mentioned variables are defined as follows: The spatiotemporal observation data sequences in the affected spatiotemporal systems are , ; Indicates the dimension of the observation sequence.
[0044] Following the previous example, assuming the sampling period for traffic flow data in the aforementioned urban spatiotemporal system is 1 minute, and each time step is 5 minutes, then the traffic flow data for a total duration of 1 hour can be divided into 12 time steps, i.e. T =11, Spatiotemporal observation data containing 12 time steps can be used express any time step t Spatiotemporal observation data, Including a certain time step t The average road speed in different areas at each sampling time.
[0045] The following section introduces mapping functions, inverse mapping functions, and the Koopman operator. The relationship between the first representation space, etc.:
[0046] Establish the Koopman operator The key is to discover the observation space and representation space Mapping function for an invertible mapping relationship between them and inverse mapping function ,in, It represents the dimension of the character space.
[0047] The representation space is the subspace spanned by the mapping function. If for any mapping function All meet That is, in the Koopman operator Under the influence of , the representation space possesses completeness, and is therefore called the representation space. It is a Koopman invariant subspace. At this point, the mapping function... Nonlinear dynamics in spatiotemporal systems can be encapsulated into linear Koopman operators. In the middle, also the space for observation Mapping to representation space In, and satisfies the following dynamic evolution equation:
[0048] ;
[0049] in, Representation space One of the characterization elements, also known as the characteristic variable at any time step t, is a... D dimensional vector, Indicates in t At time 1, the feature variables of node 1 t [0:] T Any one of the following, Indicates in t At time t, the characteristic variables of node N, Indicates the next time step t +1 predictive feature variable, Indicates in t At time +1, the predicted feature variables of node 1, Indicates in t The predicted feature variables of node N at time +1. The first representation space in this disclosure can also be called the representation space.
[0050] Therefore, the first representation space in this application is the Koopman invariant subspace.
[0051] This disclosure uses an arbitrary spatiotemporal synchronization network to obtain the mapping function. and inverse mapping function Any neural network capable of synchronously extracting spatial and temporal features from a spatiotemporal system can serve as a spatiotemporal synchronization network, without specific requirements. Therefore, this disclosure offers high flexibility, allowing users to design suitable spatiotemporal synchronization networks for specific tasks. For example, a conventional GraphGRU can be used as a spatiotemporal synchronization network.
[0052] like Figure 4 As shown, the spatiotemporal observation data sequence at the current time step Through mapping function Mapping to the first representation space yields the feature variables of the current time step. Here, for , , , that is, ,for , .
[0053] A2. The feature variables of the current time step are predicted by the global shared operator, community-specific operator and node-specific operator in the cross-scale Koopman operator to obtain the global predicted feature variables, community predicted feature variables and node predicted feature variables of the next time step, and then fused to obtain the predicted feature variables of the next time step.
[0054] like Figure 1As shown, due to causal emergence, the observation sequences on different spatial scales are different, and the evolution law inside the phenomenon also follows different dynamics. As one of the main methods for analyzing complex dynamics, Koopman theory can model any nonlinear dynamics in the feature space through a linear Koopman operator. Similar to the spatiotemporal causal structure, the core of the Koopman operator also encapsulates the dynamics of the input sequence into a linear matrix, and encodes the temporal transition process between hidden variables using the matrix. Therefore, the method based on the Koopman operator is suitable for learning nonlinear spatiotemporal dynamics.
[0055] In order to adapt to the nonlinearity and emergence of the spatiotemporal system, the present disclosure proposes a cross-scale Koopman operator to encapsulate the dynamics at different scales into respective linear matrices, so as to limit the feature space to the Koopman invariant subspace.
[0056] For the non-stationarity and emergence of the spatiotemporal system, the present disclosure splits the Koopman operator into three scales, and establishes the Koopman operator from the macro, meso and micro three scales, wherein the globally shared operator represents the Koopman operator at the macro scale, the community-specific operator represents the Koopman operator at the meso scale, and the node-specific operator represents the Koopman operator at the micro scale. The globally shared operator can also be referred to as a globally shared Koopman operator, the community-specific operator can also be referred to as a community-specific Koopman operator, and the node-specific operator can also be referred to as a node-specific Koopman operator. Each kind of Koopman operator will be described in detail as follows:
[0057] 1. Globally shared Koopman operator:
[0058] The globally shared Koopman operator is a learnable matrix that is shared in both spatial and temporal dimensions, and is used to depict the trend and behavior of global dynamics. The update equation is as follows:
[0059] ;
[0060] wherein, represents the global prediction feature variable at the next time step t+1 calculated using the globally shared Koopman operator .
[0061] 2. Community-specific Koopman operator:
[0062] The emergence of the spatio-temporal system mainly shows that the spatial nodes have community characteristics, i.e., the nodes in the same community have similar dynamics. Since the prior knowledge of the nodes is unknown, the present disclosure uses the concept of prototype learning to construct a community pattern.
[0063] The present disclosure initializes a learnable prototype embedding vector as a hidden factor representing the community pattern, C is a hyperparameter artificially given according to experience or statistical methods. Different values can be set according to different application scenarios. The original embedding vector is randomly initialized and updated during the training process, and becomes a model parameter after the training is completed. The total dimension is , and each D dimensional vector represents the implicit law inside the community c any one of the C communities. The inference process is directly used and will not be updated again. This shows that the prototype embedding vector and the learned law information in the data set during the training process.
[0064] Based on the prototype embedding vector, in the step of predicting the community prediction feature variable of the next time step by using the Koopman operator specific to the community, the following operations can be performed:
[0065] Operation 1: Calculate the similarity between the node embedding vector and the prototype embedding vector, and divide the nodes into communities based on the similarity.
[0066] First, use an arbitrary multi-layer perceptron (MLP) to compress the time dimension of the feature variable to obtain a node embedding matrix as shown in equation (1): The node embedding matrix includes node embedding vectors of N nodes:
[0067] (1)
[0068] Then, the similarity between the node embedding vector and the prototype embedding vector is calculated, and the Gumbel-Softmax trick is used to obtain an indicator variable to divide the nodes into corresponding community classes. It can be expressed as follows:
[0069]
[0070] where ; let c ∈ [1, C ], represents any one vector in , and is an indicator variable indicating the similarity, indicating which node belongs to the i-th community. a community; denotes a node embedding matrix; denotes a prototype embedding vector denotes the transpose of denotes a temperature coefficient, used to control the smoothness of the softmax function output. denotes a matrix inner product, the value after the inner product represents the similarity, specifically the similarity between each node embedding vector in the node embedding matrix and each of the transpose of the C prototype embedding vectors . denotes noise randomly sampled from .
[0071] Operation 2, based on the community-specific Koopman operator, the community prediction feature variables of each community at the next time step are calculated.
[0072] The present disclosure defines a community-specific Koopman operator is a learnable matrix of , which is shared across the time dimension and is independent between different communities, and the update formula is as follows:
[0073] ;
[0074] wherein denotes the community prediction feature variables at the next time step calculated using the community-specific Koopman operator , which includes the community prediction feature variables at the next time step corresponding to each community; since N nodes may belong to different communities, therefore denotes the community prediction feature variables at the next time step corresponding to all nodes belonging to community c , indicating that the variable will filter out the nodes belonging to community c , denotes the group-specific Koopman operator, denotes the Hadamard product.
[0075] In order to ensure the effectiveness of prototype learning, the present disclosure designs a prototype embedding orthogonal loss function, which is defined as follows:
[0076] ;
[0077] In the formula, denotes an identity matrix, denotes C prototype embedding vectors Combined Multiplying the matrix by its transpose, the loss function is used for training. C prototype embedding vectors This is to ensure that the cluster centers are sufficiently distinctive.
[0078] It should be noted that the initialization C There are three prototype embedding vectors, which are orthogonal to each other. Orthogonal vectors have zero similarity to each other, and the inner product of any two orthogonal vectors is zero. Ideally, if the learned vectors... C If the prototype embedding vectors are pairwise orthogonal, then The resulting matrix is a diagonal matrix (the elements on the main diagonal are not 0, and the other elements are 0).
[0079] If the inner product of any prototype embedding vector and its transpose is equal to 1, then in the learned... C When the prototype embedding vectors are pairwise orthogonal The resulting matrix is an identity matrix.
[0080] Therefore, the above prototype embedding orthogonal loss function is intended to make the learned... C The prototype embedding vectors are as orthogonal as possible to each other.
[0081] In one example, conditions can be set for learning prototype embedding vectors, restricting the inner product of any learned prototype embedding vector and its transpose to be equal to 1.
[0082] 3. Node-specific Koopman operators:
[0083] To capture the local dynamics of each node and adapt to changes in local distribution, this disclosure proposes a node-specific Koopman operator to capture the node-specific dynamics. Indicates the first The node-specific Koopman operator of each node, then The calculation method is as follows:
[0084] ;
[0085] in, This indicates that the first term defined in formula (1) The node embedding vector of each node. and Indicates learnable parameters, It represents the number of nodes.
[0086] The aforementioned "predicting the feature variables of the current time step using node-specific operators to obtain the node prediction feature variables of the next time step" can be obtained using the following formula:
[0087]
[0088] wherein, represents the node prediction feature variable of the i-th node at the time step t+1 calculated under the action of the Koopman operator unique to the i-th node. represents the node prediction feature variable of the i-th node at the time step t+1 calculated under the action of the Koopman operator unique to the i-th node. t represents the node prediction feature variable of the i-th node at the time step t+1 calculated under the action of the Koopman operator unique to the i-th node. represents the transpose of . represents the node prediction feature variable of the i-th node at the time step t+1 calculated under the action of the Koopman operator unique to the i-th node.
[0089] The fusion of the following:
[0090] As shown in FIG. 1, in the current time step t, the feature variable of the current time step t is mapped into the first feature space by the mapping function to obtain the feature variable of the current time step t. Figure 5 Figure 4 For different nodes, the action ability of Koopman operators of different scales is different, and the present disclosure uses a learnable scalar weight to fuse the calculation results under the action of the cross-scale Koopman operator to obtain the prediction feature variable of the next time step t+1.
[0091]
[0092] A3, the prediction feature variable of the next time step is mapped into the observation space by the inverse mapping function to obtain the prediction result; the prediction result includes the predicted spatiotemporal observation data of the next time step.
[0093] As shown in FIG. 1, after obtaining the prediction feature variable of the next time step, each is mapped into the observation space by the inverse mapping function to obtain the prediction result. Figure 6
[0094] In the present disclosure, since the Koopman operator used for predicting the spatiotemporal observation data is refined into a globally shared operator, a community-specific operator and a node-specific operator, when predicting the spatiotemporal observation data, the prediction is performed from three different scales, so that the predicted spatiotemporal data is more accurate.
[0095] The following describes how to train.
[0096] In an embodiment, the prediction model is obtained by using three-stage training, and the training process of each stage is described in detail as follows:
[0097] The first training stage:
[0098] The spatiotemporal data prediction is performed by the spatiotemporal prediction network to obtain a prediction result.
[0099] How to perform the spatiotemporal data prediction is described above and will not be repeated here.
[0100] Perform the first loss calculation: use the first loss function to obtain the first loss according to the prediction result obtained in the current training stage and the sample.
[0101] Update the spatiotemporal prediction network according to the first loss.
[0102] In one embodiment, the first loss function includes at least a spatial loss function , a reconstruction loss function and a prediction loss function .
[0103] In one example, the first loss function includes a spatial loss function, a reconstruction loss function, a prediction loss function and a prototype embedding orthogonal loss function, i.e. . The first loss function is denoted as L1.
[0104] The above use of the first loss function to obtain the first loss according to the prediction result obtained in the current training stage and the sample includes the following sub-steps B1-B3:
[0105] B1, calculate the spatial loss from the spatial loss function according to the first predicted feature variable and the first actual feature variable.
[0106] In one example, in order to enable the Koopman operator to have the ability to represent nonlinear dynamics, the present disclosure designs a spatial loss function to constrain the first representation space into a Koopman invariant subspace, and the definition of the spatial loss function is as follows:
[0107] ;
[0108] wherein, denotes the mean square error, denotes the time step t the actual feature variable at time step denotes the first predicted feature variable, denotes the time step t the actual feature variable at time step
[0109] The following describes how to determine the first or second actual feature variable from the posterior distribution.
[0110] In the first training phase, in order to increase randomness when determining the first or second actual feature variable, the posterior distribution of the first feature space is fitted. The specific definition is as follows:
[0111]
[0112] Specifically, in the present disclosure, the posterior distribution is fitted with a learnable Gaussian distribution, is a Gaussian distribution, including mean and variance two learnable parameters; The calculated results are the mean and variance of the Gaussian distribution, so the posterior distribution is equal to this mapping function, that is, By sampling the posterior distribution , and and can be obtained.
[0113] B2, the reconstruction loss is calculated from the spatiotemporal observation data sequence in the sample and the above-mentioned second actual feature variable according to the reconstruction loss function.
[0114] In order to ensure the reversibility of the mapping function and the inverse mapping function, the present disclosure designs the reconstruction loss function :
[0115] ;
[0116] wherein, denotes the reconstruction loss function, denotes the spatiotemporal observation data sequence in the sample at time step t ; denotes the actual feature variable at time step t (the second actual feature variable), sampled from the posterior distribution fitted by the mapping function , represents the cross-scale Koopman operator in the current training stage; represents the inverse mapping function in the current training stage.
[0117] B3、calculating the prediction loss from the prediction loss function according to the spatio-temporal observation data sequence in the sample and the second actual feature variable mentioned above, to obtain the prediction loss; the first loss includes the spatial loss, the reconstruction loss and the prediction loss.
[0118] To ensure the cross-scale Koopman operator can fit the dynamics of the spatio-temporal system, the prediction loss function is designed in the present disclosure :
[0119] ;
[0120] wherein, represents the prediction loss function, represents the spatio-temporal observation data sequence in the sample at the time step t+ 1; represents the actual feature variable (which can be referred to as the second actual feature variable) at the time step t ; represents the cross-scale Koopman operator in the current training stage (the cross-scale Koopman operator updated in the last training stage or the last iteration).
[0121] After obtaining the first loss, the mapping function , the cross-scale Koopman operator and the inverse mapping function in the spatio-temporal prediction network can be updated according to the first loss.
[0122] After the first training stage ends, the mapping function, the cross-scale Koopman operator and the inverse mapping function obtained by training can be referred to as the first mapping function , the first cross-scale Koopman operator and the first inverse mapping function .
[0123] The second training stage includes C1-C3:
[0124] C1, performing the above spatio-temporal data prediction by the target spatio-temporal prediction network to obtain a prediction result; the target spatio-temporal prediction network is the spatio-temporal prediction network updated in the last training stage.
[0125] C2, performing second loss calculation to obtain a second loss. C2 includes C21-C23:
[0126] C21, performing the above first loss calculation to obtain the first loss in the current training stage.
[0127] For instructions on how to perform the first loss calculation, please refer to the aforementioned records, which will not be repeated here.
[0128] In step C21, , and Replace B1-B3 respectively , and By recalculating the first loss, we can obtain the first loss for the second training phase.
[0129] C22. Using the prediction results of the current training phase, perform spectral decomposition on the cross-scale Koopman operator in the above target spatiotemporal prediction network to obtain multiple eigenvalues and corresponding eigenvectors.
[0130] C23. Calculate the spectral loss based on the above eigenvalues; the second loss includes the first loss and the spectral loss in the current training phase.
[0131] The spectral loss can be obtained from the identity matrix, the diagonal matrix composed of the aforementioned multiple eigenvalues, and the complex conjugate matrix of the aforementioned diagonal matrix.
[0132] C3. Update the above-mentioned target spatiotemporal prediction network based on the second loss mentioned above.
[0133] Since the Koopman operator is a linear operator, its eigenvalues can be analyzed using spectral decomposition techniques to understand the dynamic evolution process. Spectral decomposition can then be used to find stable and unstable states. Spectral decomposition is a fundamental concept in matrix theory, enabling the decomposition of matrix eigenvalues.
[0134] The following explains the principles for finding stable and unstable states:
[0135] By the first cross-scale Koopman operator The eigenvalues obtained from the decomposition. and eigenvectors The following relationship exists between them:
[0136]
[0137] Among them, the first representation space in the first representation space i The node at the th Feature variables at each time step All of these can be calculated from the aforementioned eigenvectors and eigenvalues.
[0138] For each feature variable Spectral decomposition can be performed to obtain the first... i Transpose of the initial feature variables of each node and the i th node at the th time step:
[0139]
[0140]
[0141] where, denotes the th power of the first cross-scale Koopman operator, is the initial feature variable of the i th node is the projection of the eigen vector , also called the amplitude, is obtained by the inner product of and , the superscript T in the above formula denotes the transpose.
[0142] As shown in FIG. 1, based on the Koopman mode decomposition technique, the Koopman operator is linearly decomposed into eigen vectors and eigen values. The eigen vectors are the basis of the feature space, which contains the basic information about the underlying dynamics, describing the temporal transition rules between hidden variables. While the eigen values contain the evolution process of the corresponding dynamic mode. Figure 7 It is worth noting that the first cross-scale Koopman operator
[0143] in the above formula is a linear combination of the globally shared Koopman operator Figure 7 trained after the first training stage, the community-specific Koopman operator and the node-specific Koopman operator , that is, . In order to more clearly understand the oscillation and frequency of the dynamic evolution process, the eigen value can be converted to continuous time: where
[0144] is the infinitesimal, representing a very small numerical difference, the difference between two consecutive time points), where and (θ represents the amplitude angle, representing the angle formed by the corresponding vector of the eigen value in the complex plane counterclockwise from the positive half of the x-axis). Since the eigen value belongs to the complex space , the real part of the time-continuous eigen value represents the evolution growth rate of the corresponding dynamic mode, and the imaginary part represents the evolution frequency of the corresponding dynamic mode. The dynamics evolution process can be regarded as the contribution rate of the corresponding dynamics model on the time-space system, The dynamics model is regarded as. Or The first model can be determined to be attenuated, static, or divergent. The change frequency of the corresponding model is determined. The real part of the eigenvalue in continuous time represents the growth rate of the corresponding model, and the imaginary part represents the model frequency. The relative size of the eigenvalue basically determines the proportion of the corresponding dynamics model to the system.
[0145] Where: Or The contribution of the corresponding dynamics model to the system does not change over time, so the dynamics model captures stable information, and the present disclosure defines it as an invariant causal model, as shown in Figure 7 The sequence recovered based on the invariant causal model (the waveform indicated by the invariant causal model in Figure 7 ) can accurately reflect the stable trend of the observed sequence; And Both are close to 0 (for example, between -0.1 and 0.1), indicating that the contribution of the corresponding dynamics model has a relatively low change frequency and follows a certain change rule, so the model captures the changing causal relationship and is a supplement to static information. The present disclosure defines it as a time-varying causal model. As shown in Figure 7 The sequence recovered based on the time-varying causal model (the waveform indicated by the time-varying causal model in Figure 7 ) can reflect the fluctuation degree of the observed sequence, depicting the local change degree, for example, the position with smaller value of the observed sequence has larger fluctuation, and the position with larger value of the observed sequence has smaller fluctuation; Smaller (for example, between -0.1) or larger (for example, between 0.1 and ), indicating that the eigenvalue is far from the unit circle in the complex space, indicating that the contribution evolution frequency of the corresponding dynamics model is large and irregular, and the present disclosure defines it as a noise model. As shown in Figure 7 The sequence recovered based on the noise model (the waveform indicated by the noise model in Figure 7 ) is random noise without rules.
[0146] According to the above analysis, the relatively stable dynamics model and the causal relationship have a close relationship, so the present disclosure regards the relatively stable dynamics evolution model (invariant causal model and time-varying causal model) as a causal model, and the linear matrix reconstructed based thereon is regarded as a causal structure.
[0147] The relatively stable dynamic mode can be separated and a linear matrix is reconstructed therefrom. It can be found in the sequence results recovered under different dynamic matrices that the stable dynamic mode can effectively recover the stable trend in the observation sequence, which is referred to as a causal component in the present application. Therefore, the present application takes the stability of the dynamic evolution process reflected by the eigenvalue as a link between dynamics and causality, and regards the relatively stable dynamic evolution mode as a causal mode, and the linear matrix reconstructed therefrom as a causal structure.
[0148] In the context of deep learning, the present embodiment explores the dynamic process and causal relationship between latent variables, and the present disclosure can find interpretability from eigenvalues at different scales. The core of the present embodiment is to explore the link between linear dynamics and causal structure, so as to reconstruct the causal structure from the perspective of dynamics.
[0149] The following explains the link between the spatiotemporal dynamics captured by the corresponding Koopman operator at different scales and the spatiotemporal causal relationship, which embodies the interpretability and effectiveness of the present disclosure at the bottom principle:
[0150] Figure 8 The distribution of eigenvalues in the Koopman operator at different scales and the corresponding causal graph is shown. When performing interpretability experiments, that is, Figure 8 , the Koopman operator at different scales is respectively subjected to spectral decomposition, and the causal graph obtained is only a specific scale causal graph obtained after decomposition of a Koopman operator at a certain scale. It is worth noting that this is only to illustrate the relationship between the Koopman operator at each scale and the causal graph, and what each scale can capture in the spatiotemporal system, and is only for interpretability illustration.
[0151] The globally shared Koopman operator uses a learnable matrix to capture the dynamics shared by spatial nodes. It is stable in the spatiotemporal dimension, depicting global behavior and trends. As Figure 8 shown, this stability ensures that all eigenvalues of the globally shared Koopman operator are close to the unit circle, keeping the contribution of the corresponding mode to the system dynamics unchanged (i.e., stable mode or causal mode). This property also conforms to the intuition of causal relationship, which is the law of stability.
[0152] The community-specific Koopman operator captures several typical modes, which follow the assumption that nodes in the same community have similar dynamic systems. In addition, in order to prevent over-smoothing, the community-specific Koopman operator is often used as a compensation for the globally shared Koopman operator to improve overall performance. As Figure 8 shown, the discrete eigenvalues of the community-specific Koopman matrix are mostly concentrated in the region with positive real part. The eigenvalues or will cause the modes to decay to zero over time, while or will cause the system to be unstable and diverge. This fact is the key principle that separates causality (relative stability) from dynamics, i.e. or . Therefore, the community-specific Koopman operator mainly extracts low-frequency running dynamic modes. The matrix recovered from these modes is defined as the invariant causal relationship in the present disclosure, which is more important in the self-mapping.
[0153] The node-specific Koopman operator is more flexible, aiming to capture the specific properties caused by the local dynamic of each node and adapt to the distribution changes. As shown in Figure 8 , the eigenvalues of the node-specific Koopman operator are mainly distributed in the left half of the unit circle, where is larger, so the corresponding dynamic mode has a higher frequency of change. This also proves that the node-specific Koopman operator pays more attention to local changes and can effectively capture the non-stationarity of the spatiotemporal sequence. The matrix recovered from these modes is defined as the time-varying causal relationship in the present disclosure.
[0154] The darker points in the causal graph represent causal relationships, and it can be seen from Figure 8 that the globally shared Koopman operator and the community-specific Koopman operator have causal relationships in the causal graphs corresponding to them, and the node-specific Koopman operator has few darker points in the causal graph corresponding to it, i.e. the causal relationship is relatively sparse. In contrast, the causal graph induced by the node-specific Koopman operator is very sparse, indicating that the local dynamic with high-frequency changes has fewer causal relationships.
[0155] The present disclosure extracts all invariant causal modes in the complete cross-scale Koopman operator (i.e. eigenvalues or ), and reconstructs the observation data sequence according to the invariant causal modes, and the result is shown in Figure 9 . In Figure 9 , the components are the sequences predicted by the new Koopman operator composed of the eigenvectors constrained by the eigenvalues, i.e. the decomposition from the observation sequence, which is predicted by the new Koopman operator. It can be seen from Figure 9 that the sequence predicted by the invariant causal component can effectively depict the core mode of the observation sequence while ignoring local oscillations. This is also the main task of causal representation learning, which involves mining core rules from redundant and noisy data and has certain migration and expansion capabilities, i.e. causal structure.
[0156] In addition, the disclosure extracts time-varying causal patterns from node-specific Koopman operators and reconstructs the observed data sequence based on them (here, reconstruction is also equivalent to prediction, and the sequence predicted by the new Koopman operator composed of the eigenvector after the eigenvalue constraint is also referred to as the thing decomposed from the observed sequence, which is predicted by the new Koopman operator), and the result is shown in FIG. 9. It can be observed that the reconstructed result (time-varying causal component in Figure 9 has a trend opposite to the invariant causal component. This is because the node pays more attention to local oscillation. When the value of the sequence is large, a small change will not have a significant impact on the entire sequence, thereby masking the local oscillation. When the value of the sequence is very small, oscillation becomes the main factor, and a small change will also cause fluctuations. The main task of the time-varying causal component is to extract causal patterns from these fluctuations. While Figure 9 the noise component in is random noise without rules.
[0157] Based on the above principle, in order to ensure the effectiveness of decoupling stable causal patterns, the disclosure proposes a spectral loss function to regularize the eigenvalues:
[0158] ;
[0159] wherein, represents a complex conjugate matrix of a matrix, represents a unit matrix, represents a diagonal matrix composed of a plurality of eigenvalues.
[0160] The purpose of the spectral loss function is to make the modulus of the eigenvalue close to 1. In the second training stage, spectral normalization is added to make the Koopman eigenvalue close to the unit circle in the complex space, thereby enhancing the stability of the dynamic pattern. The loss function in this stage is .
[0161] It is mentioned above that the target spatiotemporal prediction network can be updated according to the second loss. Specifically, in the second training stage, the first mapping function , the first cross-scale Koopman operator , and the first inverse mapping function in the spatiotemporal prediction network can be updated according to the second loss. After the second training stage ends, the second mapping function , the second cross-scale Koopman operator , and the second inverse mapping function are obtained.
[0162] The third training stage includes D1-D7:
[0163] D1, performing the above spatiotemporal data prediction by the target spatiotemporal prediction network to obtain a prediction result.
[0164] The target spatiotemporal prediction network here can refer to a spatiotemporal prediction network trained through the second training stage.
[0165] D2, performing the second loss calculation described above.
[0166] In this step, the , and in C1-C3 are replaced with , and respectively, and the second loss is recalculated to obtain the second loss of the third training stage.
[0167] D3, from the plurality of eigenvalues obtained in the current training stage, obtaining a target eigenvalue satisfying a preset condition; the eigenvector corresponding to the target eigenvalue is a target eigenvector.
[0168] It should be noted that the eigenvalues and eigenvectors of the current training stage are obtained by performing spectral decomposition on the second cross-scale Koopman operator .
[0169] The preset condition will be introduced later.
[0170] D4, constructing a causal graph through the target eigenvalue and the target eigenvector.
[0171] D5, based on the causal graph, mapping the feature variable of the current time step into a second representation space with causal constraints through a causal flow network; the second representation space is a subspace of the first representation space.
[0172] D6, calculating the KL divergence between the posterior distribution of the current training stage and the prior distribution of the second representation space.
[0173] D7, updating the current target spatiotemporal prediction network according to the third loss; the third loss includes the second loss obtained in the current training stage and the KL divergence.
[0174] The following will be introduced in detail.
[0175] In the third training stage, as Figure 10As shown, involving the causal structure extractor and the causal flow network, the causal structure extractor takes the relatively stable dynamic mode as the causal mode, takes the linear matrix reconstructed by the stable mode as the causal graph, decouples the evolution of the relatively stable dynamic mode, and takes the reconstructed linear matrix as the causal structure based on this. The causal flow network fits the structural causal equation based on the extracted causal graph to ensure the identifiability of the causal variable and model the subspace of the causal structure constraint. The above-mentioned representation space constraint method narrows the distribution distance between the Koopman subspace (the first representation space in the above embodiment) and the causal constraint subspace (the second representation space in the above embodiment), and obtains a robust representation space.
[0176] Second cross-scale Koopman operator Perform spectral decomposition:
[0177]
[0178] (Specific methods please refer to the relevant description of the second training stage), wherein, , is the left eigenvector, is the right eigenvector, and the eigenvalue is .
[0179] Where the matrix composed of the eigenvector is , and the matrix composed of the eigenvalue is . The present disclosure extracts the corresponding causal mode by constraining the eigenvalue range, and the calculation process is as follows:
[0180]
[0181] Wherein, and represent the matrix composed of all invariant causal modes and time-varying causal modes, is the relaxation factor, which is used to control the strictness of the causal mode, and is an empirical value. is the reconstructed causal graph, and the causal graph is also a matrix, because the reconstructed matrix A is the stable part of the Koopman operator matrix K, and the matrix established by this stable part is regarded as the causal graph in the present application. The constraint indicates that and is the matrix composed of the eigenvector whose modulus satisfies 1±δ, is the matrix composed of the eigenvalue whose modulus satisfies 1±δ.
[0182] The aforementioned preset condition includes the above-mentioned constraint on the eigenvalue range.
[0183] In Figure 10 , in the causal structure extractor part, a matrix composed of eigenvectors satisfying the requirements is found, and the dimension is ,so Dimensional transformation is After dimensional transformation The causal graph is formed, and the left eigenvectors form the time-varying causal pattern and the invariant causal pattern. The right eigenvectors also form time-varying causal patterns and invariant causal patterns. .
[0184] Obtaining a causal graph Then, a causal structure equation needs to be established based on the causal edges. To clarify the generation mechanism of latent variables, among which Represents the characteristic variables in the representation space; express In the middle, according to the cause-and-effect diagram Received The set of parent variables, representing the set of parent variables in right The causal effect, among which, It is obtained through sampling from the posterior distribution, and also The direct cause of its generation; Represents an exogenous variable that follows a standard Gaussian distribution. It is obtained by random sampling from a standard Gaussian distribution, which is The generated external causes are usually random noise and have no practical significance. This disclosure proposes to use an affine transformation modulated by a causal parent variable as the structural causal equation, defined as follows:
[0185]
[0186] Among them, such as Figure 10 The causal flow network shown is an SNet, which is a multilayer perceptron, and its purpose is to obtain the time steps in the affine transformation. Corresponding scaling factor TNet is a multilayer perceptron, and its purpose is to obtain time steps. Corresponding translation factor , It is a vector alternating between 0 and 1, the purpose of which is to represent the time steps. Corresponding exogenous variables Divided into two groups and , Used to calculate scaling and translation factors, representing ride The remaining portion is 0; the other portion... Used to calculate the result of affine transformation. This indicates that modulation parameters are generated using a multilayer perceptron based on the parent variable. Modulation parameters for modifying scaling factor and translation factor , modulation parameter for modifying translation factor , and thereby achieve causal effect passing. The last formula in the structural causal equation aims to convert the exogenous variable into a causal variable based on the influence of the causal parent variable . The causal structural equation designed by the present disclosure is a reversible function, which can effectively improve the identifiability of the causal variable. For convenience, the present disclosure records .
[0187] The causal effect contained in the causal graph can be embedded into the causal variable through the structural causal equation. Since the parent variable of the causal variable comes from the second representation space, the generated causal variable also belongs to the second representation space. Therefore, the present disclosure uses the causal flow network to obtain the prior distribution of the second representation space, thereby forming a representation space with structural causal constraints and enhancing the causality of the feature variables therein.
[0188] The present disclosure uses the variable theorem to decompose the prior distribution into the product of the exogenous variable distribution , which is defined as follows:
[0189] ;
[0190] wherein, represents the exogenous variable distribution; is the th element in k , which represents the process of converting from the causal variable to the exogenous variable ; represents the value of the i th node, the t th time step, and the k th exogenous variable; represents the Jacobian matrix of the function f from to , wherein, , , represents a dimensional identity matrix; represents the determinant of this Jacobian matrix, represents the causal variable of the i th node, the t th time step, and the k th exogenous variable, represents a diagonal matrix. Through the above formula, the prior distribution can be solved.
[0191] In the third training stage, the target spatio-temporal prediction network represents the evolution process of the spatio-temporal system dynamics based on the posterior distribution of the second cross-scale Koopman operator fitting, while the prior distribution of the causal structure equation fitting contains the natural causal relationship in the system. By reducing the distribution distance between the prior distribution and the posterior distribution, the spatio-temporal system can be driven in a causal dynamic manner in the representation space. Specifically, the KL divergence of them is calculated as a constraint. Since the prior distribution has no explicit analytical expression, it is calculated by sampling :
[0192] ; wherein, is used to calculate the expectation.
[0193] In the third training stage, the structural causal constraint is added, aiming to propagate the causal relationship in the entire representation space, and the loss function of this stage is , and the obtained loss is the third loss.
[0194] In this training stage, the second mapping function , the second cross-scale operator and the second inverse mapping function are updated according to the above third loss. Then, after the third training stage, the mapping function, the cross-scale operator and the mapping function can be represented as , and , that is, the parameters in the prediction model are , and .
[0195] Of course, in other embodiments, the training can also be performed without dividing the training stages, using the loss function .
[0196] It is worth noting that, in the training, the spatio-temporal observation data at each time step in the spatio-temporal observation data sequence can be used to predict the spatio-temporal data at the next time step, and then the training of the model is performed based on the predicted spatio-temporal data and the corresponding spatio-temporal observation data, for example: T +1 is 3 minutes, and each time step is 1 minute, so , the spatio-temporal observation data at the first minute is used to predict the spatio-temporal data at the second minute , the spatio-temporal observation data at the second minute is used to predict the spatio-temporal data at the third minute , at least and are used, and and Training of the model is performed.
[0197] This embodiment describes the equivalence between spatiotemporal dynamics and causal structure from another perspective, and understands the evolution mechanism of causal structure from the perspective of dynamic evolution process. The spatiotemporal causal structure describes the causal dependence relationship between the hidden variables in space over time, and the spatiotemporal dynamics describes the evolution law of the state of the space system over time, which naturally has a certain relationship in essence. The present disclosure linearizes the representation space using the Koopman operator, and then uses spectral analysis techniques to explain the relationship between the dynamics and the causal structure of the spatiotemporal system.
[0198] It should be noted that the existing spatiotemporal representation learning method assumes that the causal structure is stable and globally shared, and usually uses a learnable causal matrix to learn the causal relationship of the causal hidden variables in an unsupervised manner. However, the pursuit of global causal relationship between causal hidden variables will lead to over-smoothing effect of the learned causal structure, that is, only a small number of edges exist in the learned causal matrix and most of them are 0 values, which will lead to the loss of variability of the learned causal structure, making it difficult to generalize to a specific scenario.
[0199] To solve the above problems, the present disclosure provides a solution based on a structural causal constraint Koopman operator model. The present disclosure starts from the perspective of spatiotemporal dynamics, explores the stable modes in the evolution process, thereby revealing the generation principle of the causal structure emerging across scales, and imposes a constraint on the representation space. The present application regards the relatively stable dynamic evolution mode as a causal mode, and the reconstructed linear matrix based thereon as a causal structure. For this purpose, the present disclosure proposes a structural causal constraint Koopman neural operator model (Structural Causality-Constrained Koopman Neural Operator Model, SCKNO) to learn a robust representation space in a spatiotemporal system. The core idea is to reconstruct the causal structure by decomposing the stable dynamic mode, and impose a causal constraint in the representation space. First, in order to adapt to the nonlinearity and emergence of the spatiotemporal system, a cross-scale Koopman neural operator is proposed to encapsulate the dynamics at different scales into respective linear matrices, thereby limiting the representation space to the Koopman invariant subspace. Second, due to the linear nature of the space, the model can perform spectral analysis and check its eigenvalues, thereby better understanding the oscillation frequency. For this purpose, the present disclosure decouples the relatively stable dynamic mode based on Koopman mode decomposition, and extracts the causal structure therefrom. Third, the present disclosure proposes a causal flow network to fit the causal structure equation, thereby generating a subspace with structural causal constraints. Finally, the present disclosure narrows the distribution of the Koopman subspace and the structural causal constraint subspace, thereby generating a robust representation space.
[0200] In summary, the present disclosure has the following advantages over the prior art:
[0201] Unlike traditional methods that calculate feature variables through complex encoders, the present disclosure starts from the perspective of spatiotemporal dynamics, explores stable patterns in evolutionary processes, thereby revealing the generation principle of causal structures emerging across scales, and constrains the representation space. Traditional methods only focus on the generation of observation sequences depending on feature variables, while ignoring the generation principle of feature variables. The present application further explores the generation principle of feature variables, and defines that the generation of feature variables is derived from the causal structure between variables, while the generation of causal structure depends on the stable patterns of spatiotemporal dynamics.
[0202] The present disclosure proposes to encapsulate non-stationary, nonlinear, and cross-scale emergent dynamics into a linear matrix using cross-scale Koopman operators, and to analyze the relationship between stable dynamic patterns and causal relationships using spectral decomposition techniques, thereby explaining the generation mechanism of causal structures.
[0203] The present disclosure uses prototype learning techniques to divide spatial nodes into communities, with similar dynamic processes within the same community, and different dynamic processes between different communities, which helps to capture spatial emergence in spatiotemporal systems and improve the cross-scale feature extraction capability of the model.
[0204] The present disclosure uses linear dynamics and causal structures to jointly constrain the representation space, thereby obtaining robust feature variables. In this space, nonlinear dynamics are described by linear operators, and feature vectors with structural causal constraints are generated according to the causal direction, thereby increasing the robustness of the representation space by narrowing the distance between prior and posterior distributions.
[0205] In summary, the present disclosure uses global shared operators, community-specific operators, and node-specific operators in the cross-scale Koopman operator to predict feature variables at the current time step, obtaining prediction feature variables at different levels, and then fusing them. This multi-dimensional feature extraction and fusion method can consider the features of the spatiotemporal system at different scales and levels, capture more comprehensive information, and more accurately describe the state of the spatiotemporal system. For example, global shared operators can capture macro features of the entire system, community-specific operators can consider specific features of different communities or sub-regions in the system, and node-specific operators can focus on the unique properties of each node. By fusing these feature variables at different levels, the model can better grasp the complex characteristics of the spatiotemporal system.
[0206] The use of the mapping function and the inverse mapping function in the present disclosure enables the model to perform nonlinear mapping on spatiotemporal data, better adapting to the complex dynamic changes of the spatiotemporal system. The state changes of the spatiotemporal system are often nonlinear, and traditional single-perspective analysis methods may be difficult to accurately describe such nonlinear relationships. The present scheme maps the data through the mapping function to a feature space more suitable for model processing, in which space the nonlinear features of the data can be more effectively captured, and then maps the prediction results back to the observation space through the inverse mapping function, so that the prediction results can correspond to the actual spatiotemporal observation data. At the same time, the cross-scale Koopman operator can model the dynamic changes of the spatiotemporal system, processing different scale features through different operators, which can more accurately capture the evolution law of the spatiotemporal system over time and space, thereby improving the prediction accuracy of the future state.
[0207] The present disclosure trains the spatiotemporal prediction network by obtaining a large amount of sample data, and the model can automatically learn the potential patterns and laws in the spatiotemporal data. During the training process, the model will continuously adjust its parameters according to the input sample data to minimize the error between the prediction results and the actual observation data. This data-driven adaptive learning approach enables the model to adapt to different spatiotemporal systems and various complex situations, unlike existing technologies which are limited to specific observation dimensions or fixed analysis methods. The model can automatically extract appropriate features and establish a corresponding prediction model according to the specific data set and problem, thereby improving the prediction accuracy and generalization ability.
[0208] Based on the same inventive concept, the present application also provides a training device for implementing the above-mentioned training method. The implementation scheme of the problem solving provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more training device embodiments provided below can refer to the limitations of the training method in the foregoing, which will not be repeated here.
[0209] In one exemplary embodiment, as shown in Figure 11 a training device is provided, comprising:
[0210] The sample acquisition module 11 is configured to acquire samples, each sample including a time-space observation data sequence in an observation space; the training module 12 is configured to input the samples into a time-space prediction network to obtain a prediction model through training, wherein the time-space prediction network includes a mapping function, an inverse mapping function, and a cross-scale Koopman operator; the time-space prediction network and the prediction model are configured to perform time-space data prediction according to input data; wherein the time-space data prediction includes: mapping the input data into a first feature space through the mapping function to obtain at least a feature variable of a current time step; performing prediction on the feature variable of the current time step through a global shared operator, a community-specific operator, and a node-specific operator in the cross-scale Koopman operator respectively to obtain a global predicted feature variable, a community predicted feature variable, and a node predicted feature variable of a next time step, and performing fusion to obtain a predicted feature variable of the next time step; and mapping the predicted feature variable of the next time step into the observation space through the inverse mapping function to obtain a prediction result; the prediction result includes predicted time-space observation data of the next time step.
[0211] In one embodiment, the training includes a first training phase, and the training module 12 is specifically configured to: perform the time-space data prediction by the time-space prediction network to obtain the prediction result; perform first loss calculation: obtain a first loss by using a first loss function according to the prediction result obtained in the current training phase and the samples; and update the time-space prediction network according to the first loss.
[0212] In one embodiment, the first loss function includes a spatial loss function, a reconstruction loss function, and a prediction loss function; in the aspect of obtaining the first loss by using the first loss function according to the prediction result obtained in the current training phase and the samples, the training module 12 is specifically configured to: calculate a spatial loss by the spatial loss function according to a first predicted feature variable and a first actual feature variable; wherein the first predicted feature variable is determined by a second actual feature variable and a current cross-scale Koopman operator; the first actual feature variable and the second actual feature variable are both determined by the samples and a posterior distribution corresponding to the current training phase of the first feature space; calculate a reconstruction loss by the reconstruction loss function according to the time-space observation data sequence in the samples and the second actual feature variable; calculate a prediction loss by the prediction loss function according to the time-space observation data sequence in the samples and the second actual feature variable; and the first loss includes the spatial loss, the reconstruction loss, and the prediction loss.
[0213] In one embodiment, the training further includes a second training phase, and the training module 12 is specifically configured to: perform the spatiotemporal data prediction by the target spatiotemporal prediction network to obtain a prediction result; the target spatiotemporal prediction network is the spatiotemporal prediction network updated in the previous training phase; perform the second loss calculation to obtain a second loss; perform the first loss calculation to obtain a first loss of the current training phase; perform spectral decomposition on the cross-scale Koopman operator in the target spatiotemporal prediction network by using the prediction result of the current training phase to obtain a plurality of eigenvalues and corresponding eigenvectors; calculate a spectral loss according to the plurality of eigenvalues; the second loss includes the first loss of the current training phase and the spectral loss; and update the target spatiotemporal prediction network according to the second loss.
[0214] In the aspect of calculating the spectral loss according to the plurality of eigenvalues, the training module 12 is specifically configured to: obtain the spectral loss according to a unit matrix, a diagonal matrix composed of the plurality of eigenvalues, and a complex conjugate matrix of the diagonal matrix.
[0215] In one embodiment, the training further includes a third training phase, and the training module 12 is specifically configured to: perform the spatiotemporal data prediction by the target spatiotemporal prediction network to obtain a prediction result; perform the second loss calculation; obtain a target eigenvalue satisfying a preset condition from the plurality of eigenvalues obtained in the current training phase; an eigenvector corresponding to the target eigenvalue is a target eigenvector; construct a causal graph by using the target eigenvalue and the target eigenvector; map the feature variable of the current time step to a second representation space with causal constraints by using a causal flow network based on the causal graph; the second representation space is a subspace of the first representation space; calculate the KL divergence between the posterior distribution of the current training phase and the prior distribution of the second representation space; update the current target spatiotemporal prediction network according to a third loss; and the third loss includes the second loss obtained in the current training phase and the KL divergence.
[0216] Figure 12 is a flowchart of a spatiotemporal data prediction method according to an exemplary embodiment, as shown in Figure 12 the method includes the following steps S201-S202:
[0217] In step S201, target spatiotemporal observation data is obtained. In step S202, the target spatiotemporal observation data is input into a prediction model to obtain predicted spatiotemporal observation data within a target time span.
[0218] The prediction model is trained and obtained, and the related introduction is described in the foregoing description, which is not repeated here.
[0219] In one example, the aforementioned target prediction spatiotemporal observation data includes prediction spatiotemporal observation data for at least one time step; obtaining the target time span prediction spatiotemporal observation data in step S202 includes the following sub-steps S2021-S2022:
[0220] S2021. Input the input data into the above prediction model to obtain the predicted spatiotemporal observation data for the next time step.
[0221] S2022. Update the above input data with the predicted spatiotemporal observation data of the next time step until the predicted spatiotemporal observation data of all time steps in the target time span are obtained.
[0222] After obtaining the prediction model through the above embodiments, when it is necessary to use the prediction model for spatiotemporal data prediction, the target spatiotemporal observation data is first obtained during prediction. The target spatiotemporal observation data is then used as input data and processed through a mapping function. The spatiotemporal observation data of the target are mapped to the first representation space to obtain the feature variables of the current time step. These feature variables of the current time step are then processed by the cross-scale Koopman operator. Global shared operators, community-specific operators, and node-specific operators are used for prediction to obtain global, community, and node predicted feature variables for the next time step. These are then fused to obtain the predicted feature variables for the next time step. This is achieved through an inverse mapping function. The predicted feature variables for the next time step are mapped to the observation space to obtain the predicted spatiotemporal observation data for the next time step. After obtaining the predicted spatiotemporal observation data for the next time step, the input data is updated using this data, and the prediction model is used to obtain the predicted spatiotemporal observation data for the next time step based on the new input data. This process is repeated until the predicted spatiotemporal observation data for all time steps within the target time span are obtained. This is based on the aforementioned spatiotemporal observation data sequence. For example, the predicted spatiotemporal observation data for the next time step would be: ;use Update the input data to obtain a new spatiotemporal observation data sequence. As new input data, it is used by the prediction model based on Obtain the predicted spatiotemporal observation data for the next time step ; then use Update the input data to obtain a new spatiotemporal observation data sequence. As input data, the same principle applies, without further explanation.
[0223] Based on the same inventive concept, the embodiments of the present application also provide a spatio-temporal data prediction device for implementing the above-mentioned spatio-temporal data prediction method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more spatio-temporal data prediction device embodiments provided below can refer to the limitations of the spatio-temporal data prediction method described above, which will not be repeated here. In one exemplary embodiment, as shown in Figure 13 FIG. 16, a spatio-temporal data prediction device is provided, which includes an acquisition module 21 configured to acquire target spatio-temporal observation data. A prediction module 22 is configured to input the target spatio-temporal observation data into a prediction model trained by any one of the training methods in the above embodiments to obtain predicted spatio-temporal observation data in a target time span.
[0224] In one embodiment, the predicted spatio-temporal observation data in the target time span includes predicted spatio-temporal observation data of at least one time step; in terms of obtaining the predicted spatio-temporal observation data in the target time span, the prediction module 22 is specifically configured to: input the input data into the prediction model to obtain predicted spatio-temporal observation data of the next time step; and update the input data using the predicted spatio-temporal observation data of the next time step until predicted spatio-temporal observation data of all time steps in the target time span is obtained.
[0225] The above model is further described below in relation to the urban highway speed spatio-temporal system specifically involved in the present embodiment. The globally shared Koopman operator is used to model the common information within all regions, the community-specific Koopman operator is used to model the information specific to the same community, wherein the dynamic evolution process of the same community is similar, and the dynamic evolution process of different communities is different, and the node-specific Koopman operator is used to model the private information within a single region, which aims to capture the speciality and local details of the region.
[0226] For multi-region urban highway speed spatio-temporal data prediction, the multi-region urban highway speed itself forms a spatio-temporal system, and the spatio-temporal system is mapped into a high-dimensional space, called a representation space, through deep learning technology. The feature variables in the representation space are unobservable hidden variables that affect the regional highway speed. The present application uses deep learning technology to mine the hidden variables to model and explain the generation process of urban road speed. In the training process, the feature variables are constrained by the causal flow network, causing causal effects between the feature variables, and then converting the feature variables into causal variables. The hidden variables obtained by spatio-temporal dynamics are defined as feature variables, and the feature variables are converted into causal variables after passing through the causal flow network. In order to make the representation space cover both spatio-temporal dynamics and causality, the present application uses a structural causal constraint method to narrow the distribution distance between the feature variable distribution (located in the Koopman invariant subspace) and the causal variable distribution (located in the subspace subject to causal constraints), obtaining a robust representation space. Finally, the spatio-temporal data is predicted from the robust representation space using an inverse mapping function.
[0227] Characteristics of regional highway speed in cities may include, but are not limited to, the following cases (population density, building density, economic activity intensity, traffic conditions, etc.): Population density: Population density is an important spatial feature, because areas with high population density often have more traffic demand to meet people's living and business needs. For example, the highway speed in large cities is slower than that in ordinary urban areas, because cities have more people and business activities. Building density: Building density is also an important spatial feature, because areas with high building density require more traffic demand to meet people's work, entertainment and other living needs. For example, the traffic speed in commercial areas during the evening peak period is usually slower than that in residential areas, because commercial areas have higher building density and people travel more frequently. Economic activity intensity: Economic activity intensity is also an important factor affecting traffic speed. Developed areas often require more traffic demand to meet higher work and trade needs. For example, the traffic speed in an industrialized area may be slower than that in an agricultural area. The purpose of the present application is to infer the characteristic variables that may affect the regional highway speed from the observable regional highway speed data in cities, and to establish the causal relationship between the various characteristic variables, and then to model and explain the generation rule and causal structure of the regional highway speed in the complex system of cities. In addition, the application takes the variables that affect the operation of the spatio-temporal system, such as regional properties, weather conditions and regional gathering activities, as system control variables, and summarizes the time series pattern of regional traffic speed as a node embedding vector, which describes the operation mode, for example, the time series pattern of regional traffic speed has periodicity, trend, seasonality and suddenness, when the external variables such as weather change, the time series pattern of regional traffic speed will also change accordingly, at the same time, the sudden regional gathering activities will also cause the dynamic change of regional traffic speed pattern, therefore, the dynamics of characteristic variables and causal structure will be affected by control variables.
[0228] The following examples of the present application are used to compare the performance of the SCKNO of the present application with the baseline model. The characteristic space dimension is set for all SCKNO components . For hyperparameters, the number of communities . For fairness, the hidden layers of all baseline models are adjusted to 64 layers, while ensuring consistent input information, and all models are optimized for a maximum of 300 epochs by Adam optimizer. The best parameters of all deep learning models are selected through a careful parameter adjustment process on the validation set.
[0229] Example 1: The predictive performance of SCKNO was evaluated on a real-world, publicly available spatiotemporal dataset collected by the Transportation Bureau Performance Measurement System (PeMS) of Region A, which consists of the average speeds of Region B. The research nodes are... The road is observed at 5-minute intervals (each 5-minute interval is a time step), with the observation dimension being... Using data from the past hour ( To predict the data for the next hour ( This dataset records the average speed of 170 roads in Zone B from July to August 2016, with a sampling period of 1 minute. To ensure the stability of the time series, regions with missing values were removed, and the dataset was split at 5-minute intervals to obtain 17,833 samples from the PeMS08 dataset. Each sample includes the average speed of each of the 170 roads over one time step (5 minutes) (if the sampling period is 1 minute, the average speed includes the average of five 1-minute road speed samples). This application uses data from the past hour. To predict the data for the next hour ( 60% of the data is used for training, 20% for validation, and the remainder for testing.
[0230] This application uses root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE) to evaluate model performance. This application compares SCKNO with state-of-the-art spatiotemporal representation learning methods / models (DGCRN, DMSTGCN, PDFormer, and STNSCM) to verify SCKNO's predictive performance. The final average results are shown in Table 1.
[0231] Table 1. Quantitative analysis results of the prediction model of this application and other baseline models on the PeMS08 dataset.
[0232]
[0233] Figure 14 This paper shows the changes in the predicted MAE and MAPE indices of the model of this disclosure and the baseline model as time steps increase in a multi-step prediction task on the PeMS08 dataset. Thanks to its ability to extract causal structure from stable dynamic patterns, the model of this disclosure performs best across all time periods, reflecting its stability. It is worth noting that... Figure 14 The results of quantitative analysis of other baseline models besides those in Table 1, including GCIM, ST-SSL, DSTAGNN, RGSL, D2STGNN, and KNF, are also shown.
[0234] Example 2: The prediction performance of SCKNO was evaluated on a real-world public spatio-temporal dataset, i.e., the BJ2022 dataset. The dataset records the traffic flow data of multiple transportation means, including the inflow and outflow of bicycles, taxis, and buses, and the average speed of regional roads, in 235 regions of B district from January 2022 to December 2022, with a sampling period of 30 minutes. The dataset was segmented at 30-minute intervals to obtain 17,518 samples of the BJ2022 dataset, with each sample including the traffic flow data of each region in the 235 regions at a time step (30 minutes). The prediction of 90 minutes of data was performed using 4 hours of historical data . 60% of the data was used for training, 20% for validation, and the rest for testing.
[0235] The root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE) methods were used to evaluate the performance of the model. The prediction performance of SCKNO was compared with the current advanced spatio-temporal representation learning methods / models to verify the prediction performance of SCKNO, and the final average results are shown in Table 2:
[0236] Table 2 Quantitative analysis results of the prediction model of the present application and other models on the BJ2022 dataset
[0237]
[0238] Example 3: The prediction performance of SCKNO was evaluated on a real-world public spatio-temporal dataset, i.e., the NYC2016 dataset. The dataset records the traffic flow data of multiple transportation means, including the inflow and outflow of bicycles and taxis, in 51 regions of D district of C city from April 2016 to June 2016, with a sampling period of 30 minutes. The dataset was segmented at 30-minute intervals to obtain 4,358 samples of the NYC2016 dataset, with each sample including the traffic flow data of each region in the 51 regions at a time step (30 minutes). The prediction of 90 minutes of data was performed using 4 hours of historical data . 60% of the data was used for training, 20% for validation, and the rest for testing.
[0239] The root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE) methods were used to evaluate the performance of the model. The prediction performance of SCKNO was compared with the current advanced spatio-temporal representation learning methods / models to verify the prediction performance of SCKNO, and the final average results are shown in Table 3:
[0240] Table 3 Quantitative analysis results of the prediction model of the present application and other models on the NYC2016 dataset
[0241]
[0242] For spatio-temporal sequences, the higher the dimension of observation, the more complex the causal relationship within the system, and MAPE can effectively reflect the ability of the model to resist random fluctuations. From the corresponding MAPE of each model, it can be seen that the SCKNO of the present disclosure always outperforms the baseline model with overwhelming advantage. In particular, on the two multi-modal traffic datasets of BJ2022 and NYC2016 (Tables 2 and 3), SCKNO shows significant improvement in MAPE. At the same time, in terms of multi-step spatio-temporal sequences on the PeMS08 dataset (Table 1), the Koopman operator method based on nonlinear dynamics has a natural advantage, which makes the method of the present disclosure more advantageous on BJ2022.
[0243] For correlation-based representation learning methods, DGCRN and DMSTGCN embed the spatio-temporal correlation of training data into a learnable implicit graph. The difference is that DMSTGCN learns an implicit graph structure at each time step. Although the data information can be maximally utilized, it is not conducive to the generalization of the model, leading to a significant decline in performance when dealing with complex situations (such as multi-modal data). PDFormer focuses on using complex attention to extract long-term spatio-temporal correlation. Although it has a certain competitiveness, it still has deficiencies when dealing with complex non-stationary environments due to the lack of stable causal guidance. In contrast, causal representation learning-based methods have obvious advantages in complex scenarios (i.e., multi-modal datasets such as BJ or NYC). STNSCM eliminates confounding factors in spatio-temporal representation based on causal intervention, but does not simulate the causal relationship between latent variables. STNSCM considers the causal graph to be globally shared and stable, which may lead to over-stationary effects and poor performance in multi-step prediction tasks (i.e., PeMS dataset).
[0244] In an example embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above training method embodiments when executing the computer program. The computer device can be a server or a terminal. In an example, the computer device specifically includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement a training method. In an example embodiment, a computer readable storage medium is provided, storing a computer program, which is executed by the processor to implement the steps in the above training method embodiments. In an example embodiment, a computer program product is provided, including a computer program, which is executed by the processor to implement the steps in the above training method embodiments. In an example embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above spatio-temporal data prediction method embodiments when executing the computer program. The computer device can be a server or a terminal. In an example, the computer device specifically includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement a spatio-temporal data prediction method. In an example embodiment, a computer readable storage medium is provided, storing a computer program, which is executed by the processor to implement the steps in the above spatio-temporal data prediction method embodiments.In an example embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps of the above-mentioned spatio-temporal data prediction method embodiments. Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. The database involved in each embodiment provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, and the like, without being limited thereto. The processor involved in each embodiment provided by the present application can be a general processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, and the like, without being limited thereto. Each technical feature of the above embodiments can be combined arbitrarily. In order to make the description simple, each technical feature of the above embodiments is not described in all possible combinations, however, as long as the combination of these technical features does not exist contradictory, it should be considered as the scope of the present application. The principles and implementation manners of the present application are described by applying specific examples herein. The above embodiments are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application range can be changed. In conclusion, the content of the present application should not be understood as a limitation.
Claims
1. A training method characterized by, The training method is applied to the field of spatiotemporal sequence prediction of urban traffic flow and road speed, and the method comprises: obtaining samples; each sample comprises a spatiotemporal observation data sequence in an observation space; each sample comprises at least one of the following information: road average speed of different roads at each sampling time at the same time step: inputting the samples into a spatiotemporal prediction network to obtain a prediction model through training, wherein the spatiotemporal prediction network comprises a mapping function, an inverse mapping function and a cross-scale Koopman operator; the spatiotemporal prediction network and the prediction model are used to perform spatiotemporal data prediction according to input data; the spatiotemporal data prediction comprises: mapping the input data into a first feature space through a mapping function to obtain at least a feature variable of a current time step; the input data is road average speed of different roads at each sampling time at the current time step, and the feature variable of the current time step is a feature variable corresponding to the road average speed of different roads at each sampling time at the current time step; The feature variables corresponding to the average road speed of different roads at each sampling time step in the current time step are predicted by the global shared operator, community-specific operator, and node-specific operator in the cross-scale Koopman operator, respectively, to obtain the global predicted feature variables, community predicted feature variables, and node predicted feature variables for the next time step. These are then fused to obtain the predicted feature variables corresponding to the average road speed of different roads at each sampling time step in the next time step. The global shared operator is a... The learnable matrix, shared across the spatiotemporal dimensions, is used to characterize the trends and behaviors of global dynamics. The update equation is as follows: ,in, Indicates the use of globally shared operators Calculate the global predicted feature variables for the next time step t+1. The characteristic variable representing the current time step t; the community-specific operator is a The learnable matrix, which is shared across time dimensions and independent between different communities, is updated using the following formula: ; in, Indicates the use of community-specific operators The community prediction feature variables for the next time step are calculated, which include the community prediction feature variables for the next time step corresponding to each community. This represents the community prediction feature variable for all nodes belonging to community c at the next time step; the indicator variable is... It will filter out the nodes that belong to community c. Indicates the first Each group has its own unique Koopman operator. Let Hadamard product be represented. The following formula is used to predict the node-specific feature variables of the current time step using node-specific operators, thus obtaining the node-predicted feature variables for the next time step: in, Indicates the first Each node in its unique Koopman operator The node prediction feature variables calculated under the influence of the action; express Transpose of; This represents the node prediction feature variable for the next time step, which includes the node prediction feature variables for the next time step of each node. mapping the predicted feature variable corresponding to the road average speed of different roads at each sampling time at the next time step into the observation space through an inverse mapping function to obtain a prediction result; the prediction result comprises the road average speed of different roads at each sampling time at the next time step.
2. The training method of claim 1, wherein the training comprises a first training phase: performing the spatiotemporal data prediction to obtain a prediction result; performing first loss calculation: using a first loss function to obtain a first loss according to the prediction result obtained in the current training phase and the sample; updating the spatiotemporal prediction network according to the first loss.
3. The training method of claim 2, wherein the first loss function comprises a spatial loss function, a reconstruction loss function and a prediction loss function; the first loss is obtained by using the first loss function according to the prediction result obtained in the current training phase and the sample, comprising: calculating a spatial loss by a spatial loss function according to a first predicted feature variable and a first actual feature variable; wherein the first predicted feature variable is determined by a second actual feature variable and a current cross-scale Koopman operator; the first actual feature variable and the second actual feature variable are both determined by the sample and a posterior distribution corresponding to the first feature space in the current training phase; calculating a reconstruction loss by a reconstruction loss function according to the spatiotemporal observation data sequence in the sample and the second actual feature variable; calculating a prediction loss by a prediction loss function according to the spatiotemporal observation data sequence in the sample and the second actual feature variable; the first loss comprises the spatial loss, the reconstruction loss and the prediction loss.
4. The training method of claim 3, wherein the training further comprises a second training phase: performing the spatiotemporal data prediction by a target spatiotemporal prediction network to obtain a prediction result; the target spatiotemporal prediction network is the spatiotemporal prediction network updated in the last training phase; performing second loss calculation to obtain a second loss: performing the first loss calculation to obtain the first loss in the current training phase; perform spectral decomposition on the cross-scale Koopman operator in the target spatio-temporal prediction network by using the prediction result of the current training stage, to obtain a plurality of eigenvalues and corresponding eigenvectors; calculate a spectral loss according to the plurality of eigenvalues; the second loss includes the first loss of the current training stage and the spectral loss; update the target spatio-temporal prediction network according to the second loss.
5. The training method of claim 4, wherein, The method comprises: obtaining the spectral loss according to a unit matrix, a diagonal matrix composed of the plurality of eigenvalues, and a complex conjugate matrix of the diagonal matrix.
6. The training method of claim 4, wherein the training further comprises a third training stage: performing the spatio-temporal data prediction by the target spatio-temporal prediction network to obtain a prediction result; performing the second loss calculation; from the plurality of eigenvalues obtained from the current training stage, obtaining a target eigenvalue satisfying a preset condition; the eigenvector corresponding to the target eigenvalue is a target eigenvector; constructing a causal graph through the target eigenvalue and the target eigenvector; mapping the feature variable of the current time step into a second representation space with causal constraints through a causal flow network based on the causal graph; the second representation space is a subspace of the first representation space; calculating the KL divergence between the posterior distribution of the current training stage and the prior distribution of the second representation space; updating the current target spatio-temporal prediction network according to a third loss; the third loss includes the second loss obtained in the current training stage and the KL divergence.
7. A training device, characterized by The device is applied to the field of spatio-temporal sequence prediction of urban traffic flow and road speed, and the device comprises: a sample acquisition module configured to acquire samples; each sample comprises a spatio-temporal observation data sequence in an observation space; each sample comprises at least one of the following information: road average speed of different roads at each sampling time at the same time step: a training module configured to input the samples into a spatio-temporal prediction network to obtain a prediction model through training, wherein the spatio-temporal prediction network comprises a mapping function, an inverse mapping function, and a cross-scale Koopman operator; the spatio-temporal prediction network and the prediction model are configured to perform spatio-temporal data prediction according to input data; wherein the spatio-temporal data prediction comprises: mapping the input data into a first representation space through the mapping function to obtain at least a feature variable of a current time step; the input data is road average speed of different roads at each sampling time at the current time step, and the feature variable of the current time step is a feature variable corresponding to the road average speed of different roads at each sampling time at the current time step; The feature variables corresponding to the road average speeds of different roads at each sampling time at the next time step are obtained by predicting the feature variables corresponding to the road average speeds of different roads at each sampling time at the current time step through the global shared operator, the community-specific operator and the node-specific operator in the cross-scale Koopman operator respectively, and fusing the global predicted feature variables, the community predicted feature variables and the node predicted feature variables at the next time step, to obtain the predicted feature variables corresponding to the road average speeds of different roads at each sampling time at the next time step; the global shared operator is a learnable matrix, which is shared in the space-time dimension, and is used to describe the trend and behavior of global dynamics, and the update equation is as follows: , wherein, represents the global predicted feature variables at the next time step t+1 calculated using the global shared operator , represents the feature variables at the current time step t; the community-specific operator is a learnable matrix, which is shared across the time dimension and is independent between different communities, and the update formula is as follows: ; , wherein, represents the community predicted feature variables at the next time step calculated using the community-specific operator , which includes the community predicted feature variables at the next time step corresponding to each community; represents the community predicted feature variables at the next time step corresponding to all nodes belonging to the community c, and the indicator variable screens out the nodes belonging to the community c, represents the i-th community-specific Koopman operator, represents the Hadamard product; the feature variables at the current time step are predicted through the node-specific operator to obtain the node predicted feature variables at the next time step by using the following formula: , wherein, represents the node predicted feature variables at the next time step calculated under the i-th node-specific Koopman operator for the i-th node; represents the transpose of ; represents ; and represents the node predicted feature variables at the next time step, which includes the node predicted feature variables at the next time step of each node. mapping the predicted feature variable corresponding to the road average speed of different roads at each sampling time at the next time step into the observation space through the inverse mapping function to obtain a prediction result; the prediction result comprises road average speed of different roads at each sampling time at the next time step.
8. A spatiotemporal data prediction method, characterized by, The method comprises: obtaining target spatio-temporal observation data corresponding to the road average speed; inputting the target spatiotemporal observation data corresponding to the road average speed into the prediction model trained by the training method in any one of claims 1-6 to obtain predicted spatiotemporal observation data corresponding to the road average speed in the target time span.
9. The method of claim 8, wherein, The predicted spatiotemporal observation data corresponding to the road average speed in the target time span includes predicted spatiotemporal observation data corresponding to the road average speed of at least one time step. The obtaining of the predicted spatiotemporal observation data corresponding to the road average speed in the target time span includes: inputting the input data into the prediction model to obtain predicted spatiotemporal observation data corresponding to the road average speed of the next time step; updating the input data using the predicted spatiotemporal observation data corresponding to the road average speed of the next time step until predicted spatiotemporal observation data corresponding to the road average speed of all time steps in the target time span is obtained.
10. A spatio-temporal data prediction apparatus characterized by comprising: The device includes: an acquisition module configured to acquire target spatiotemporal observation data corresponding to a road average speed; a prediction module configured to input the target spatiotemporal observation data corresponding to the road average speed into a prediction model trained by the training method in any one of claims 1-6 to obtain predicted spatiotemporal observation data corresponding to the road average speed in a target time span.
Citation Information
Patent Citations
System and method for monitoring and training a manufacturing system
CN118331181A
Non-stationary time sequence flow prediction method and device based on delay space-time dependence, and medium
CN120257237A