Space-time Mangbar diffusion hybrid network for multivariable traffic sequence completion
By introducing a causal prior and Mamba diffusion hybrid model in traffic data processing, the problem that existing methods are difficult to capture spatiotemporal causal relationships is solved, and more efficient traffic data interpolation and real-time monitoring of intelligent traffic systems are achieved.
Patent Information
- Application Number
- CN202510312433.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
AI Technical Summary
Existing traffic data processing methods are difficult to effectively capture the complex causal relationships and local consistency in spatiotemporal data, resulting in insufficient accuracy of missing value interpolation, limiting the real-time monitoring and prediction capabilities of intelligent traffic systems.
A mamba diffusion mixed model based on causal prior guidance is proposed. By introducing the mamba structure and attention mechanism, we can learn the causal relationship in time and as conditional knowledge of the diffusion model, we can enhance the network's feature learning ability.
It effectively captures the causal relationship and fine-grained characteristics of time and space, improves the interpolation accuracy of traffic data, and enhances the real-time monitoring and prediction capabilities of intelligent traffic systems.
Smart Images

Figure CN120164324A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet of Things artificial intelligence, and more specifically, to a Mamba diffusion hybrid network for multivariable spatio-temporal sequence completion. Background Art
[0002] The rapid development of intelligent transportation systems has made the collection and analysis of traffic data crucial for effective traffic management and planning. The progress of the Internet of Things has further promoted the seamless integration of various sensors and monitoring devices. However, due to problems such as sensor failures, communication interruptions, and data sparsity, missing values often occur in traffic datasets, which weakens their usefulness and effectiveness and seriously hinders real-time traffic monitoring. Therefore, implementing effective methods to handle missing values is crucial for improving the accuracy of downstream application tasks, such as traffic prediction and route planning. Recently, spatio-temporal imputation techniques for multivariate time series missing data have attracted wide attention. These methods utilize temporal and spatial dependencies to effectively recover lost data. Therefore, finding efficient spatio-temporal imputation methods has become a key area of traffic data research.
[0003] Spatio-temporal traffic data is usually represented as a tensor, which makes tensor imputation a key method for solving missing data. However, many existing tensor imputation models overly rely on the algebraic properties of data. They often fail to fully capture local spatio-temporal consistency and network correlations. These models ignore the potential interactions between unobserved segments, thus limiting their imputation capabilities. In addition, although smoothing priors and low-rank structures are introduced to improve performance, in the context of large-scale networks, the global consistency of the low-rank assumption remains challenging.
[0004] With the progress of deep learning techniques, new models that utilize neural networks to extract data features are emerging. These methods can effectively solve the problems of temporal entanglement and inter-sequence correlations in time series data. Self-supervised generative models have shown remarkable performance in this regard. They significantly improve imputation accuracy by leveraging global and context sequence information. Although self-supervised generative models improve imputation accuracy, they are difficult to capture the complex spatio-temporal correlations in traffic data. Graph neural networks (GNNs) can effectively encode complex spatio-temporal patterns by aggregating domain signals, but they still face challenges in dealing with data heterogeneity and dynamic changes. In addition, the model's dependence on graph structure selection and construction may weaken its generalization ability. More importantly, current methods often neglect the inherent causal relationships in the data generation process. This oversight complicates the accurate modeling of complex spatio-temporal dynamic systems. In intelligent transportation systems, there are complex causal chains that link various events. Most existing models tend to emphasize correlations rather than causal relationships, which limits their ability to understand and predict system behavior. Summary of the Invention
[0005] The present invention proposes a Mamba diffusion hybrid model guided by causal priors for multivariate spatio-temporal sequence interpolation. We adopt a diffusion model as the core of the interpolation network, taking into account the causal relationships in both the temporal and spatial dimensions. In addition, the Mamba structure is introduced to enhance the learning of high-level spatio-temporal features. Specifically, variational ideas are used to learn spatio-temporal causal relationships, and the learned causal features are incorporated as prior conditions into the input of the diffusion model. We introduce a novel Mamba-attention dual structure integrated into the diffusion network. It effectively captures spatio-temporal causal relationships and various fine-grained features, thus enhancing the overall learning ability of the network.
[0006] To achieve the above objectives, the technical solution of the present invention is as follows:
[0007] A spatio-temporal Mamba diffusion hybrid network for multivariate traffic sequence completion, as Figure 1 shown, includes the following steps:
[0008] S1: Use sensors to collect spatio-temporal sequence data with missing values;
[0009] S2: Design a mask matrix according to whether there are missing values at each sampling point. The mask matrix is a 0 / 1 matrix, where data missing is 0 and successful sampling is 1. The matrix dimension is the same as the sampling data;
[0010] S3: Input the processed data into the temporal and spatial causal discovery networks respectively to obtain spatio-temporal causal prior matrices;
[0011] S4: Use the temporal and spatial causal prior matrices as weight matrices representing causal relationships, and multiply them with the collected original data through a sliding window to obtain the temporal and spatial causal prior features required for training;
[0012] S5: Design an innovative Mamba-attention dual structure, which includes a spatio-temporal attention structure, a spatio-temporal dual-channel Mamba module, a causal dual-channel Mamba module, and a multi-scale attention mechanism;
[0013] S6: Embed the Mamba-attention dual structure into the diffusion network to obtain a Mamba diffusion hybrid network with stronger feature learning ability;
[0014] S7: Use the causal prior features as conditional knowledge of the diffusion model, and input them together with the collected sequence and the mask matrix into the Mamba diffusion network;
[0015] S8: When accurate true values cannot be obtained for the missing data, mask some of the observed data as missing points as well. The mean square error between the values of this part of the observed data and their imputed values will be used as the loss function of the network;
[0016] S9: Start training the diffusion network using Adam. After the training is completed, the complete data restored by the generation network is obtained, and the data interpolation task is completed.
[0017] Preferably, the nth sequence among the N multivariate time series data sampled in step S1 can be expressed as:
[0018]
[0019] where 1 ≤ t ≤ T is the time step, and D represents the number of sensors used for sampling. represents the data sampled from D sensors at time t, where 1 ≤ d ≤ D.
[0020] Preferably, the masking matrix in step S2 is M (n) , and specifically can be expressed as:
[0021]
[0022] Preferably, the causal discovery method in step S3 is based on the functional causal model and uses Granger causality to characterize, representing the time series as a nonlinear autoregressive function. Through the transformation of the nonlinear autoregressive equation and the replacement of noise data, the approximate distribution of the causal matrix can be obtained by maximizing the evidence lower bound. The calculation methods in the spatial and temporal dimensions are the same, and only the reference dimension of the sequence needs to be converted.
[0023] Preferably, in step S4, the time dimension is taken as an example for detailed description. Given that the inherent causal relationship of the data is transitive, we only need to learn the causal associations of three consecutive time steps.
[0024] Specifically, we specify the parent nodes of the features at time τ as the eigenvalues of the data at τ - 2, τ - 1, and τ, considering the instantaneous causal relationship (the data at time τ). Therefore, each multivariate time series generates a fixed three-dimensional causal prior tensor G T , with dimensions (3, D, D). Here, "3" represents the three time steps τ - 2, τ - 1, τ. and correspond to the causal matrix between the data features at three time points, with dimensions D × D.
[0025] We apply a time sliding window of scale three, moving one unit at a time within the range of (2, T), processing the data within each window in turn, and using a simple fully connected layer to assign weights to the three causal matrices. Finally, the obtained causal prior feature representation is:
[0026]
[0027] Similarly, the parent nodes of the d-th sensor feature are also limited to three. According to the same calculation logic, the spatio-causal attention feature is calculated as follows:
[0028]
[0029] Preferably, the specific composition of the Mamba-Attention dual structure in step S5 is as Figure 2 shown:
[0030] The first module is a spatio-temporal attention structure. It is composed of a temporal attention structure and a spatial attention structure in series. Among them, each attention structure consists of a multi-head attention (MHA), a feed-forward network (FFN), and a layer normalization (LayerNorm):
[0031]
[0032] The output u of the spatio-temporal attention structure spatial , is fed into the spatio-temporal dual Mamba module. It consists of two Mamba layers (ML) and a skip connection with a one-dimensional convolution (ConvId):
[0033] v sdm = ML(u temporal ) + ML(u spatial ) + ConvId(u spatial )
[0034] In parallel, there is a causal dual-channel Mamba structure, which uses the previously learned spatio-causal prior features (h, l) to pass through the Mamba layers respectively, and at the same time preprocesses the temporal information (sideinfo) through a one-dimensional convolutional layer, and adds them together:
[0035] v cdm = ML(h) + ML(l) + ConvId(sideinfo)
[0036] The last part is a multi-scale attention structure (MSA). First, it uses a local attention mechanism to learn fine-grained local features, and then extracts global attention features based on these local representations. It integrates the output v of the spatio-temporal dual Mamba sdm and the result v of the causal dual-channel Mamba cdm :
[0037] v madm = v sdm + v cdm + MSA(v sdm + v cdm )
[0038] The above is the complete composition of the Mamba-Attention dual structure, and finally the result v madm ;
[0039] Preferably, in step S6, the complete Mamba-attention dual structure is embedded inside the diffusion network to obtain a Mamba diffusion hybrid network with stronger feature learning ability;
[0040] Preferably, step S7 will determine the specific input of the Mamba diffusion network, including the causal prior features H (n) , L (n) , the collected sequence, and the mask matrix;
[0041] Preferably, the specific diffusion training steps of steps S8-S9 are as follows:
[0042] The overall diffusion model consists of two processes: the forward diffusion process, which systematically introduces noise into the initial sequence and gradually changes it to generate a distorted data sequence. The reverse sampling process, which starts from the completely distorted data, systematically samples to remove the noise and finally restores the original data distribution. To restore the complete data, I will train the network to fit the reverse sampling process:
[0043]
[0044]
[0045] Among them, the diffusion process steps are represented as 1 ≤ s ≤ S, and obeys the standard Gaussian distribution. θ is the trainable parameter of the sampling network, then it means that the sampled data follows a Gaussian distribution with a mean of and a variance of . By performing S-step chained reverse sampling, finally the complete sampled data to be restored can be obtained.
[0046] The loss function is calculated specifically as follows:
[0047]
[0048] Among them, Ω (n) represents a set of indices corresponding to the missing terms to be calculated. We use the masking matrix M (n) to distinguish between observed values and missing values. When , it means that the data set point is included in Ω (n) and will participate in the calculation of the loss function. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is the overall model structure diagram of the present invention.
[0050] Figure 2 is the Mamba-attention dual structure diagram proposed by the present invention.
[0051] Figure 3 Visualization of the complete data distribution curve restored by CMDSI and the target true value of the missing data for the method of the present invention provided for the embodiment on the PEMS-BAY dataset at a 10% missing rate.
[0052] Figure 4 Visualization of the complete data distribution curve restored by CMDSI and the target true value of the missing data for the method of the present invention provided for the embodiment on the PEMS-BAY dataset at a 50% missing rate.
[0053] Figure 5 Visualization of the complete data distribution curve restored by CMDSI and the target true value of the missing data for the method of the present invention provided for the embodiment on the Guangzhou dataset at a 50% missing rate.
[0054] Figure 6 Visualization of the complete data distribution curve restored by CMDSI and the target true value of the missing data for the method of the present invention provided for the embodiment on the Guangzhou dataset at a 90% missing rate. Detailed implementation manners
[0055] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the present patent;
[0056] To better illustrate this embodiment, some components in the accompanying drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product;
[0057] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0058] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0059] Embodiment 1
[0060] This example provides a spatio-temporal mamba diffusion hybrid network for multivariate traffic sequence completion, as Figure 1 shown, including the following steps:
[0061] S1: Collect spatio-temporal sequence data with missing values using sensors;
[0062] S2: Design a mask matrix according to whether there are missing values at each sampling point. The mask matrix is a 0 / 1 matrix, where 0 indicates data missing and 1 indicates successful sampling. The matrix dimension is the same as that of the sampling data;
[0063] S3: Input the processed data into the time and space causal discovery networks respectively to obtain the spatio-temporal causal prior matrix;
[0064] S4: The time and space causal prior matrix, as the weight matrix representing the causal relationship, is multiplied by the collected original data through a sliding window to obtain the time and space causal prior features required for training;
[0065] S5: Design an innovative Mamba-Attention dual structure, which includes a spatio-temporal attention structure, a spatio-temporal dual-channel Mamba module, a causal dual-channel Mamba module, and a multi-scale attention mechanism;
[0066] S6: Embed the Mamba-Attention dual structure into the diffusion network to obtain a Mamba diffusion hybrid network with stronger feature learning ability;
[0067] S7: Use the causal prior features as the conditional knowledge of the diffusion model, and input them together with the collected sequence and the mask matrix into the Mamba diffusion network;
[0068] S8: When the accurate true value of the missing data cannot be obtained, part of the observed data is also masked as missing points, and the mean square error between the observed data value and its imputed value will be used as the loss function of the network;
[0069] S9: Use Adam to start training the diffusion network. After the training is completed, the complete data restored by the generative network is obtained, and the data interpolation task is completed.
[0070] Example 2
[0071] The nth sequence in the N multivariate time series data sampled in step S1 can be expressed as:
[0072]
[0073] where 1 ≤ t ≤ T is the time step, and D represents the number of sensors used for sampling. represents the data sampled from D sensors at time t, where 1 ≤ d ≤ D.
[0074] The mask matrix in step S2 is M (n) , and can be specifically expressed as:
[0075]
[0076] The causal discovery method in step S3 is based on the functional causal model and uses Granger causality to represent the time series as a non-linear autoregressive function. By converting the non-linear autoregressive equation and replacing the noise data, the approximate distribution of the causal matrix can be obtained by maximizing the evidence lower bound. The calculation methods in the two spatio-temporal dimensions are the same, and only the reference dimension of the sequence needs to be converted.
[0077] In step S4, a detailed explanation is given taking the time dimension as an example. Given that the inherent causal relationship of data is transitive, we only need to learn the causal associations of three consecutive time steps.
[0078] Specifically, we specify the parent nodes of the features at time τ as the feature values of the data at times τ - 2, τ - 1, and τ, taking into account the instantaneous causal relationship (the data at time τ). Thus, each multivariate time series generates a fixed three-dimensional causal prior tensor G T , with dimensions (3, D, D). Here, "3" represents the three time steps τ - 2, τ - 1, τ. and the causal matrix corresponding to the data features at three time points, with dimensions D × D.
[0079] We apply a time sliding window of scale three, moving one unit at a time within the range (2, T), processing the data within each window in turn, and using a simple fully connected layer to assign weights to the three causal matrices. Finally, the resulting causal prior feature representation is:
[0080]
[0081] Similarly, the parent nodes of the d-th sensor feature are also restricted to three. According to the same calculation logic, the spatial causal attention feature is calculated as follows:
[0082]
[0083] Preferably, the specific composition of the Mamba-Attention dual structure in step S5 is as Figure 2 shown:
[0084] The first module is the spatio-temporal attention structure. It is composed of a time attention structure and a spatial attention structure in series. Among them, each attention structure consists of multi-head attention (MHA), a feed-forward network (FFN), and layer normalization (LayerNorm):
[0085] u temporal = LayerNorm(FFN(x + LayerNorm(MHA(x) + x)))
[0086]
[0087] The output u spatial of the spatio-temporal attention structure is fed into the spatio-temporal dual Mamba module. It consists of two Mamba layers (ML) and a skip connection with a one-dimensional convolution (ConvId):
[0088] v sdm = ML(u temporal ) + ML(uspatial ) + ConvId(u spatial )
[0089] In parallel is the causal dual-channel Mamba structure, which uses the previously learned spatio-temporal causal prior features (h, l) to pass through the Mamba layer respectively, and at the same time preprocesses the temporal information (sideinfo) through a one-dimensional convolutional layer, and adds them together:
[0090] v cdm = ML(h) + ML(l) + ConvId(sideinfo)
[0091] The last part is the multi-scale attention structure (MSA). First, use the local attention mechanism to learn fine-grained local features, and then extract global attention features based on these local representations. It integrates the output v of the spatio-temporal dual Mamba sdm and the result v of the causal dual-channel Mamba cdm :
[0092] v madm = v sdm + v cdm + MSA(v sdm + v cdm )
[0093] The above is the complete composition of the Mamba-attention dual structure, and finally the result v will be obtained madm ;
[0094] In step S6, the complete Mamba-attention dual structure is embedded inside the diffusion network to obtain a Mamba diffusion hybrid network with stronger feature learning ability;
[0095] In step S7, the specific input of the Mamba diffusion network will be determined, including the causal prior features H (n) , L (n) , the collected sequence and the mask matrix;
[0096] The specific diffusion training steps of steps S8 - S9 are described as follows:
[0097] The overall diffusion model consists of two processes: the forward diffusion process, which systematically introduces noise into the initial sequence and gradually changes it to generate a distorted data sequence. The reverse sampling process, which starts from the completely distorted data, systematically samples to remove the noise and finally restores the original data distribution. To restore the complete data, I will train the network to fit the reverse sampling process:
[0098]
[0099]
[0100] Among them, the diffusion process step is represented as 1 ≤ s ≤ S, and obeys the standard Gaussian distribution. θ is the trainable parameter of the sampling network, then it means that the sampled data follows a Gaussian distribution with a mean of and a variance of . By performing S-step chained inverse sampling, finally the complete sampled data to be restored can be obtained.
[0101] The loss function is calculated specifically as follows:
[0102]
[0103] where Ω (n) represents a set of indices corresponding to the missing terms to be calculated. We use the masking matrix M (n) to distinguish between observed values and missing values. When , it means that the data set point is included in Ω (n) and will participate in the calculation of the loss function.
[0104] Example 3
[0105] Based on Example 1 and Example 2, this example provides the following specific examples:
[0106] We verified on three Internet of Things data sets, namely: PEMS-BAY, METR-LA, and Guangzhou data set.
[0107] PEMS-BAY was collected by the Performance Measurement System (PEMS) of the California Department of Transportation. It consists of highway traffic speed data (in km / h) recorded by 325 section detectors located in the San Francisco Bay Area of California. The data set covers six months from January 1, 2017 to June 30, 2017, with a resolution of 5 minutes. Overall, the data set includes 16,937,179 observed traffic data points. For verification and testing, we randomly deleted 10% to 90% of the observed values and used the remaining data as the ground truth to evaluate the imputation performance.
[0108] The METR-LA data set consists of traffic time series data collected by 207 loop detectors on the highways of Los Angeles County. Each time series was recorded over a four-month period from March 1, 2012 to June 30, 2012, with a sampling rate of 5 minutes. The data set includes a total of 6,519,002 observed traffic data points. For verification and testing, we randomly deleted 10% to 90% of the observed values and used the remaining data as the ground truth to evaluate the imputation ability.
[0109] The traffic speed record in Guangzhou. The data is collected by sensors on 214 anonymous road segments and is collected every 10 minutes over a two-month period (from August 1, 2016 to September 30, 2016, for 61 days). The original missing rate of this dataset is 1.29%. Compared with the other two datasets, it is considered to be somewhat representative due to its smaller sample size. We divide the data into a training set, a validation set, and a test set in the ratio of 0.7:0.2:0.1 respectively.
[0110] In our experiment, we set the batch size of the three datasets to 16. The learning rate of the Guangzhou dataset is set to 5e-4, and the other two are set to 1e-4. At the same time, we implemented an early stopping strategy to terminate the training after 100 epochs. In addition, the model is constructed using PyTorch, and the Adam optimizer is used in the training process.
[0111] The present invention compares the proposed CMDSI with seven advanced imputation methods, namely LRTC-TNN, IGNNK, LATC, GMAN, CSDI, SAITS, and MATCN. The results are shown in Table 1, which shows the performance of the method CMDSI of the present invention in completing imputation. For clarity, the best performance of each dataset is shown in bold.
[0112] Table 1: Performance comparison of imputation methods on three datasets. The reported metrics are MAE / RMSE.
[0113]
[0114] As shown in Table 1, the CMDSI of the present invention has achieved breakthrough results on the three datasets. It shows near-optimal performance in both MAE and RMSE, highlighting its robustness and universality. CMDSI performs better than other models, such as PEMS-BAY and METR-LA, on large-scale datasets. On smaller datasets, such as Guangzhou, although its RMSE performance is slightly behind MATCN, its MAE performance remains optimal and is still competitive overall.
[0115] To further illustrate the imputation performance, we visualize the imputation examples of PEMS-BAY and Guangzhou at different missing rates (10%, 50%, 90%). Figures 3-4 It shows the data distribution curve restored by CMSDI and the target ground truth of the missing data when the missing rate of the dataset PEMS-BAY is 10% and 50%. Figures 5-6It shows the reconstructed data distribution curves and the target true values of the missing data when the missing rates are 50% and 90% in Guangzhou. CMDSI effectively reconstructs the missing values in both datasets, accurately capturing the trends and outliers. Even in the presence of a large amount of missing data, CMDSI can successfully recover the data fluctuation trends and missing values, further demonstrating its ability to complete the interpolation task in data with a high missing rate. In summary, as the loss rate increases, CMDSI maintains a significant advantage, proving its excellent robustness and insensitivity to missing data. This highlights the effectiveness and stability of CMDSI in handling challenging time series interpolation tasks.
[0116] The above embodiments show that the proposed metric can achieve better processing results when the sampling rate is not high enough, thus achieving better detection performance.
[0117] The same or similar reference numerals correspond to the same or similar components;
[0118] The descriptions of the positional relationships in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;
[0119] Obviously, the above examples of the present invention are merely illustrations for clearly explaining the present invention and are not limitations on the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A spatiotemporal Mamba diffusion hybrid network for multivariate traffic sequence completion, characterized in that: The following steps are involved: S1: Use sensors to collect spatiotemporal series data with missing data; S2: Design a mask matrix based on whether each sampling point is missing. The mask matrix is a 0 / 1 matrix. If the data is missing, it is 0, and if it is successfully sampled, it is 1. The matrix dimension is the same as the sampled data. S3: Input the processed data into the temporal and spatial causal discovery networks respectively to obtain the spatiotemporal causal prior matrix; S4: The temporal and spatial causal prior matrices are used as weight matrices to represent causal relationships, and are multiplied by sliding windows with the collected raw data to obtain the temporal and spatial causal prior features required for training; S5: Design an innovative Mamba-attention dual structure, which includes spatiotemporal attention structure, spatiotemporal dual-channel Mamba module, causal dual-channel Mamba module and multi-scale attention mechanism; S6: Embed the Mamba-Attention dual structure into the diffusion network to obtain a Mamba-Diffusion hybrid network with stronger feature learning ability; S7: The causal prior features are used as the conditional knowledge of the diffusion model and input into the Mamba diffusion network together with the collected sequence and mask matrix; S8: When the missing data cannot obtain the accurate true value, part of the observed data will also be masked as missing points, and the mean square error between the observed data value and its interpolated value will be used as the loss function of the network; S9: Use Adam to start training the diffusion network. After the training is completed, the complete data restored by the generated network is obtained, completing the data interpolation task.
2. The spatiotemporal Mamba diffusion hybrid network for multivariate traffic sequence completion according to claim 1, characterized in that: In step S1, the nth sequence of the N multivariate time series data sampled can be expressed as: Among them, 1≤t≤T is the time step, and D represents the number of sensors used for sampling. Represents the data sampled from D sensors at time t, where 1≤d≤D.
3. The spatiotemporal Mamba diffusion hybrid network for multivariate traffic sequence completion according to claim 2, characterized in that: The mask matrix in step S2 is M (n) , which can be specifically expressed as:
4. The spatiotemporal Mamba diffusion hybrid network for multivariate traffic sequence completion according to claim 3, characterized in that: The causal discovery method in step S3 is based on the functional causal model and uses Granger causality representation to represent the time series as a nonlinear autoregressive function. By transforming the nonlinear autoregressive equation and replacing the noise data, the approximate distribution of the causal matrix can be obtained by maximizing the lower bound of evidence. The calculation methods for the two dimensions of time and space are the same, and only the base dimension of the sequence needs to be transformed.
5. The spatiotemporal Mamba diffusion hybrid network for multivariate traffic sequence completion according to claim 4, characterized in that: In step S4, the time dimension is used as an example for detailed explanation. Given that the inherent causal relationship of the data is transitive, we only need to learn the causal relationship of three consecutive time steps. Specifically, we assign the parent node of the feature at time τ to the feature value of the data at time τ-2, τ-1, and τ, taking into account the instantaneous causal relationship (the data at time τ). Therefore, each multivariate time series will produce a fixed three-dimensional causal prior tensor G T , the dimension is (3, D, D). Here, "3" represents three time steps τ-2, τ-1, τ. and The causal matrix between the data features corresponding to the three time points has a dimension of D×D. We apply a time sliding window of scale three, moving one unit at a time in the range of (2, T), process the data in each window in turn, and use a simple fully connected layer to assign weights to the three causal matrices. The final causal prior feature is expressed as: Similarly, the parent node of the dth sensor feature is also limited to three. According to the same calculation logic, the spatial causal attention feature is calculated as follows:
6. The spatiotemporal Mamba diffusion hybrid network for multivariate traffic sequence completion according to claim 5, characterized in that: The specific composition of the Mamba-Attention dual structure of step S5 is as follows: The first module is the spatiotemporal attention structure. It is composed of the temporal attention structure and the spatial attention structure in series. Each attention structure consists of multi-head attention (MHA), feedforward network (FFN) and layer normalization (LayerNorm): u temporal =LayerNorm(FFN(x+LayerNorm(MHA(x)+x))) u spatial =LayerNorm(FFN(u temporal +LayerNorm(MHA(u temporal )+u temporal ))) The output u of the spatiotemporal attention structure spatial , passed into the spatiotemporal dual Mamba module. It consists of two Mamba layers (ML) and a skip connection with a one-dimensional convolution (ConvId): v sdm =ML(u temporal )+ML(u spatial )+ConvId(u spatial ) In parallel, the causal dual-channel Mamba structure uses the previously learned spatiotemporal causal prior features (h, l) to pass through the Mamba layer respectively, and pre-processes the temporal information (sideinfo) through a one-dimensional convolutional layer, and then adds them together: v cdm =ML(h)+ML(l)+ConvId(sideinfo) The last part is the multi-scale attention structure (MSA), which first uses the local attention mechanism to learn fine-grained local features, and then extracts global attention features based on these local representations. It integrates the output v of the spatiotemporal dual Mamba sdm And the result of causal double pass mamba v cdm : v madm =v sdm +v cdm +MSA(v sdm +v cdm ) The above is the complete composition of the Mamba-Attention dual structure, and the final result is v madm .
7. The spatiotemporal Mamba diffusion hybrid network for multivariate traffic sequence completion according to claim 6, characterized in that: In step S6, the complete Mamba-attention dual structure is embedded in the diffusion network to obtain a Mamba-diffusion hybrid network with stronger feature learning ability.
8. The spatiotemporal Mamba diffusion hybrid network for multivariate traffic sequence completion according to claim 7, characterized in that: Step S7 determines the specific input of the Mamba diffusion network, including the causal prior features H learned by the causal discovery network. (n) , L (n) , with the acquired sequence and mask matrix.
9. The spatiotemporal Mamba diffusion hybrid network for multivariate traffic sequence completion according to claim 8, characterized in that: The specific diffusion training steps of steps S8-S9 are as follows: The overall diffusion model consists of two processes: the forward diffusion process and the reverse sampling process. To recover the complete data, I will train the network to fit the directional sampling process: where the diffusion process steps are represented as 1≤s≤S, and Obeys a standard Gaussian distribution. θ is a trainable parameter of the sampling network, This means that the sampled data follows the mean The variance is Gaussian distribution. By chaining inverse sampling from S steps, we finally get The complete sampling data to be restored can be obtained. The loss function is calculated as follows: Where Ω (n) represents a set of indices corresponding to the missing items that need to be calculated. We use the masking matrix M (n) to distinguish observed values from missing values. When , it means that the data set point is contained in Ω (n) , will participate in the calculation of the loss function.