Traffic data completion method and device based on time autoregression low-rank tensor

By introducing a time autoregressive regularization term and a truncated weighted nuclear norm into the low-rank tensor completion model, and combining it with the ADMM framework, we solve the problem of neglecting the time dimension in traditional methods, and achieve more accurate data recovery and imputation, especially in the efficient processing of complex missing scenarios.

CN121256205APending Publication Date: 2026-01-02CHONGQING SHUAIBANG MACHINERY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511370685.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing low-rank tensor completion methods neglect the crucial role of the time dimension when dealing with missing data, resulting in inaccurate data recovery and an inability to effectively utilize historical information.

Method used

By introducing a time autoregressive regularization term and a truncated weighted nuclear norm, and combining it with the ADMM framework, a low-rank tensor completion model is constructed by iteratively updating the tensor to be estimated, the multivariate time series, and the autoregressive coefficient matrix. This ensures that the model can learn historical information and maintain low rank during data recovery.

Benefits of technology

It significantly improves the accuracy and adaptability of data interpolation, especially in complex missing scenarios, and can better handle the local and global consistency of spatiotemporal data, providing more reliable interpolation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121256205A_ABST
    Figure CN121256205A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a traffic data completion method and device based on a time autoregression low-rank tensor, and the method comprises the following steps: obtaining incomplete traffic data; constructing a three-dimensional tensor to represent incomplete traffic data, and obtaining an observation index set and a tensor matrix; taking a low-rank tensor completion model based on rank minimization as a framework, and introducing time autoregressive regularization and truncation weighted nuclear norms to obtain a low-rank tensor completion model; performing problem decomposition solving on the low-rank tensor completion model on the basis of an ADMM framework to obtain four variable parameters, namely, a variable parameter, a variable parameter and a variable parameter; and iteratively updating the variables according to the sequence until a preset convergence standard is reached, and outputting the complemented complete traffic data. According to the method, the time change is introduced into the completion of the three-order tensor as a new regularization item, so that the historical information can be learned during data recovery, the low-rank performance of the completed tensor can be ensured, and a complex missing scene can be better processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for traffic data completion based on time autoregressive low-rank tensors. Background Technology

[0002] With the rapid development of computer technology, the internet, and the Internet of Things (IoT), human society is rapidly generating massive amounts of high-dimensional data, including but not limited to ultra-high-definition images, videos, traffic information, and remote sensing data. This growth and complexity presents unprecedented challenges to data analysis. However, in the real world, data gaps are a frequent problem, becoming a common challenge for many datasets. During data acquisition, data gaps are almost unavoidable due to various reasons such as sensor malfunctions, equipment errors, or communication problems. This results in incomplete datasets, limiting data integrity and usability. Data gaps have a profound negative impact on subsequent data analysis and applications. They can not only reduce sample size but also cause biases in feature distributions and even affect model accuracy, all of which hinder accurate data understanding and consequently impact decision-making. Therefore, effectively addressing data gaps is crucial for ensuring data quality and enhancing the value of data applications.

[0003] In academia, tensor representation and computation have been widely applied across multiple disciplines, including computer vision, signal processing, data mining, and machine learning. However, in practical data processing, data gaps are common, with diverse causes including, but not limited to, data conversion or communication failures, noise interference from signal acquisition equipment, and occlusion caused by environmental obstacles. Therefore, effectively recovering incomplete tensors has become a key research focus and challenge in many disciplines. Accurate recovery and supplementation of missing data are crucial in numerous applications such as image and video rendering, facial recognition, audio classification, motion capture, and traffic data completion. Therefore, researching and developing effective tensor completion techniques to address data gaps not only contributes to advancements in related disciplines but also has significant practical implications for improving the efficiency and accuracy of data processing.

[0004] To address the problem of missing data, researchers have proposed various data completion methods, including matrix factorization, tensor kernel norm, tensor sparsity, and tensor approximation. In recent years, tensor-based completion methods have attracted widespread attention in the field of data completion and have yielded certain research results. The emergence of tensor completion methods provides a new approach to solving the problem of missing values ​​in multidimensional data. Compared with traditional matrix completion methods, tensor completion methods can better utilize the high-dimensional structure and multimodal characteristics of data, thus achieving more accurate and robust data completion. These methods can not only handle different types of missing patterns but also effectively process large-scale high-dimensional datasets. In practical applications, tensors often have an approximately low-rank structure, therefore, low-rank structures can be used for tensor completion. The low-rank tensor completion problem is an extension of the low-rank matrix completion problem. Typically, traditional matrix kernel norm minimization methods are used to solve the LRTC, a convex relaxation method for matrix rank minimization. However, the solution obtained by this method may not be optimal, because minimizing the matrix nuclear norm of all singular values ​​at the same time may not approximate the matrix rank well, and it cannot guarantee that historical information can be learned during data recovery. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a traffic data completion method and apparatus based on time autoregressive low-rank tensors. The method introduces time variation as a new regularization term into the completion of third-order tensors, which ensures that historical information can be learned during data recovery and that the completed tensor has low rank, thus better handling complex missing scenarios.

[0006] The present invention solves the above-mentioned technical problems through the following technical means:

[0007] In a first aspect, embodiments of this application provide a traffic data completion method based on time autoregressive low-rank tensors, comprising the following steps:

[0008] Obtaining incomplete traffic data;

[0009] A three-dimensional tensor is constructed to represent the incomplete traffic data, and an observation index set and tensor matrix are obtained.

[0010] Based on the low-rank tensor completion model with rank minimization as the framework, time autoregressive regularization and truncated weighted nuclear norm are introduced to obtain the low-rank tensor completion model.

[0011] The low-rank tensor completion model is decomposed and solved using the ADMM framework to obtain the tensor to be estimated. Multivariate time series Dual variables Autoregressive coefficient matrix Four variable parameters;

[0012] Treating the estimation tensor Multivariate time series Dual variables Autoregressive coefficient matrix The system iterates and updates sequentially until the predetermined convergence criteria are met, then outputs the completed traffic data.

[0013] Secondly, embodiments of this application provide a traffic data completion device based on time autoregressive low-rank tensors, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in the first aspect above.

[0014] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0015] This application proposes a low-rank tensor completion model, TARLRTC-TWNN, based on temporal autoregressive regularization and truncated weighted nuclear norm. Temporal variation is introduced as a new regularization term into the completion of third-order tensors, ensuring that historical information is learned during data recovery while maintaining the low-rank nature of the completed tensor. The low-rank structure of the tensor better represents the global consistency of the data, such as inherent periodicity and similarity. By applying an autoregressive model to each time series to design temporal variation and using the autoregressive coefficients as learnable parameters, minimizing temporal variation better represents the correlation of data in local time and can better handle complex missing data scenarios. Within the ADMM framework, the low-rank tensor and autoregressive coefficient matrix are iteratively estimated, and the missing spatiotemporal data is estimated and imputed by minimizing the tensor truncated weighted nuclear norm and the temporal autoregressive regularization constraint term.

[0016] Traditional low-rank tensor completion techniques primarily focus on capturing spatial correlation and low-rank properties, sometimes neglecting the crucial role of the temporal dimension. This application significantly enhances the model's ability to capture the temporal characteristics and local variations of data by incorporating a temporal autoregressive regularization term. Leveraging the low-rank characteristics of spatiotemporal data and its continuity in the temporal dimension, this application combines temporal autoregressive regularization with low-rank tensor completion. This combination not only ensures the restoration of global data consistency by truncating the weighted nuclear norm but also guarantees the coherence of local temporal structures through temporal autoregressive regularization, thereby more accurately imputing missing temporal data.

[0017] The core of temporal autoregressive regularization lies in uncovering the autocorrelation in time series, that is, considering the correlation between the current moment and past moments. This mechanism greatly improves the accuracy of the interpolation process. This method can effectively utilize the temporal patterns and trends in historical data, providing more reliable interpolation results for spatiotemporal data. By combining correlation analysis of time series data and the introduction of temporal autoregressive regularization terms to model local temporal neighborhoods, this application provides a basis for understanding the temporal local variation patterns of spatiotemporal data.

[0018] The experimental results of this application verify the significant advantages of the time autoregressive low-rank tensor completion method in time-series data recovery, especially its superior performance in handling continuous missing data scenarios, which significantly outperforms other comparative methods. This indicates that the method has excellent adaptability and efficiency in handling complex data missing patterns, providing an effective technical means for accurate data imputation. Attached Figure Description

[0019] Figure 1 This is a flowchart of a traffic data completion method based on time autoregressive low-rank tensors;

[0020] Figure 2 This is a heatmap of correlation coefficients for traffic speed data from selected sensors in Guangzhou's urban traffic over a 60-day period.

[0021] Figure 3 It is a graph of traffic speed data from the same sensor over a week;

[0022] Figure 4 This is the algorithm flowchart for the TARLRTC-TWNN model;

[0023] Figure 5 These are parameters from the Guangzhou traffic dataset. and Heatmap of the impact of the values ​​on model performance;

[0024] Figure 6 Parameters under the Seattle dataset and Heatmap of the impact of the values ​​on model performance;

[0025] Figure 7 The parameters are from the Portland traffic dataset. and Heatmap of the impact of the values ​​on model performance;

[0026] Figure 8 This is the interpolation result of TARLRTC-TWNN on the Guangzhou traffic speed dataset;

[0027] Figure 9 This is the interpolation result of TARLRTC-TWNN on the Seattle traffic speed dataset;

[0028] Figure 10 This is the interpolation result of TARLRTC-TWNN on the Portland traffic flow dataset. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. In the description of embodiments of this application, unless otherwise stated, "a plurality of" means two or more, for example, "a plurality of processing units" means two or more processing units, "a plurality of elements" means two or more elements, etc.

[0031] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0032] This application selects spatiotemporal data with obvious periodicity and human characteristics (such as traffic data) as the experimental object. The characteristics of spatiotemporal data are introduced below, and the spatiotemporal data are analyzed using the Pearson correlation coefficient.

[0033] Spatiotemporal data is data that varies in time and space, and it has the following characteristics:

[0034] (1) Temporal sequence: Spatiotemporal data has a temporal sequence, that is, the data is arranged in chronological order. This means that the data may exhibit different patterns, trends and periodicities at different points in time.

[0035] (2) Spatial correlation: Spatiotemporal data are spatially correlated, that is, there may be spatial dependence between data in adjacent regions or locations. This correlation can be reflected as spatial autocorrelation, that is, there is a correlation between data values ​​of adjacent locations.

[0036] (3) Spatiotemporal interaction: The characteristics of spatiotemporal data interact with each other in time and space, which means that changes in data at a certain point in time may be affected by other locations in space, or changes at a certain location in space may change over time.

[0037] (4) Scale effect: Spatiotemporal data may exhibit scale effect, that is, data may show different patterns and characteristics at different scales. For example, some phenomena may appear random at a smaller spatial scale, while they may exhibit obvious spatial patterns at a larger spatial scale.

[0038] (5) Heterogeneity: Spatiotemporal data may be heterogeneous, that is, data from different regions or locations have different characteristics and change patterns. This heterogeneity may be affected by geographical, environmental, socioeconomic and other factors.

[0039] (6) Multivariability: Spatiotemporal data usually contains multiple variables or indicators, not just a single time series or spatial series. There may be complex correlations and interactions between these variables.

[0040] The Pearson correlation coefficient is used to obtain a correlation analysis index for spatiotemporal data. It is a statistic used to measure the degree of linear correlation between two variables, and its value ranges from -1 to 1. Given two variables X and Y, the Pearson correlation coefficient can be calculated using the following formula:

[0041] (1)

[0042] In equation (1), It is a variable and Pearson correlation coefficient, yes and covariance, and They are and The standard deviation.

[0043] when A positive correlation coefficient indicates a perfectly positive linear relationship between two variables. This means that as one variable increases, the other variable will also increase by a fixed proportion. A positive correlation coefficient indicates that the two variables are positively correlated.

[0044] when When the correlation coefficient is negative, it indicates that there is a completely negative linear correlation between the two variables. This means that as one variable increases, the other variable will decrease by a fixed proportion. A negative correlation coefficient indicates that there is a negative correlation between the two variables.

[0045] when When the correlation coefficient is close to zero, it indicates that there is no linear correlation between the two variables. This means that there is no linear relationship between the change of one variable and the change of the other variable. A correlation coefficient close to zero indicates that the linear relationship between the two variables is weak or close to non-existent.

[0046] In summary, The larger the value, the stronger the correlation between sample variables X and Y.

[0047] Traffic data, as a typical spatiotemporal data format, records traffic conditions (such as traffic flow and vehicle speed information) collected from sensors at different locations, exhibiting a continuous temporal dimension. Traffic data possesses strong temporal characteristics; this time dependence imbues spatiotemporal data with dynamism and historical context, enabling it to reflect trends and changes over time. The following analysis uses traffic data as an example to examine the correlation of spatiotemporal data:

[0048] (1) Temporal correlation of traffic data

[0049] In exploring the correlation characteristics of traffic data, this study selected 60 days of traffic data from a monitoring point in Guangzhou to conduct an in-depth analysis of the interrelationships over time. Figure 2 A heatmap showing the time correlation of Guangzhou's urban traffic speed data over a week is presented. Figure 2 In the image, variations in color intensity visually reveal the degree of correlation between traffic data from different days: a redder color indicates a strong correlation between traffic data from two days, while a lighter color indicates a weaker correlation. This graphical presentation allows for the rapid identification of correlation patterns in traffic speed data over time, providing an intuitive reference for further data analysis.

[0050] from Figure 2 As can be seen, the Pearson correlation coefficient of data collected by the same sensor over 60 days is above 0.5, indicating a high correlation and low dispersion among data from different days. This suggests that the periodic correlation of data from the same sensor at different times can be used for data recovery. Table 1 shows the correlation between days within a week. By calculating the correlation of traffic speed data on different days, it can be seen that the traffic speed data from the same sensor within a week has a strong correlation, especially the correlation coefficient between two adjacent days, which can reach above 0.8, indicating a high degree of autocorrelation. Figure 3 The graphs showing traffic speed changes from the same sensor for each day of the week provide a more intuitive view of the correlation between the same time on different days within the week.

[0051]

[0052] Table 1

[0053] Correlation analysis based on traffic data reveals patterns and relationships in traffic flow across different locations, and also demonstrates the high correlation of the same sensor across different days or weeks. However, traditional low-rank tensor completion methods neglect similarities within local temporal neighborhoods. To better model the local temporal neighborhoods of traffic data, this application proposes introducing a temporal autoregressive regularization term into low-rank tensor completion based on minimizing the tensor truncation weighted norm.

[0054] This application reveals the interdependencies among traffic flow data over consecutive days through in-depth correlation analysis of time-series data. To improve the imputation accuracy of time-series data and address the issue of missing data, this application proposes introducing a time autoregressive regularization term into the low-rank tensor completion framework. Figure 1 As shown, the traffic data completion method based on time autoregressive low-rank tensors in this application includes the following steps:

[0055] Step S100: Obtain incomplete traffic data;

[0056] Step S200: Construct a three-dimensional tensor to represent the incomplete traffic data and obtain the observation index set and tensor matrix;

[0057] Step S300: Based on the low-rank tensor completion model with rank minimization as the framework, time autoregressive regularization and truncated weighted nuclear norm are introduced to obtain the low-rank tensor completion model.

[0058] Step S400: Based on the ADMM framework, the low-rank tensor completion model is decomposed and solved to obtain the tensor to be estimated. Multivariate time series Dual variables Autoregressive coefficient matrix Four variable parameters;

[0059] Step S500, the tensor to be estimated Multivariate time series Dual variables Autoregressive coefficient matrix The system iterates and updates sequentially until the predetermined convergence criteria are met, then outputs the completed traffic data.

[0060] The traffic data completion method based on time autoregressive low-rank tensors proposed in this application will be described in detail below:

[0061] In step S100, incomplete traffic data is acquired. Specifically, Y represents the matrix of real data collected by M sensors, as shown in equation (2), where each row of data corresponds to a different time series collected by the sensor, and the columns correspond to time points.

[0062] (2)

[0063] In equation (2), I represents the number of blocks required for tensor quantization, and J represents the size of the tensor.

[0064] In step S200, a three-dimensional tensor is constructed to represent the incomplete traffic data, obtaining an observation index set and a tensor matrix. Specifically, in tensor form, the size of the tensor is... In matrix form, the data has a total of The column determines the observation period and performs observations; part of the observation matrix can be written as... The elements in the observation index set must satisfy formula (3):

[0065] (3)

[0066] In equation (3), Ω represents the observation set, m represents the row dimension of the matrix, and n represents the column dimension of the matrix.

[0067] Next, we introduce a positive tensor quantization operation. Tensor quantization can convert a multivariate time series matrix into a third-order tensor, dividing the time dimension of the time series into two dimensions, from which the third-order tensor can be obtained. Conversely, the inverse operation... It can be converted into a tensor matrix to obtain .

[0068] In step S300, based on the low-rank tensor completion model with rank minimization as the framework, time autoregressive regularization and truncated weighted nuclear norm are introduced to obtain the low-rank tensor completion model. Specifically, the time autoregressive regularization is as follows:

[0069] Given a time series matrix and an autoregressive coefficient matrix And a time lag set ,matrix The time autoregressive model can be expressed as equation (4):

[0070] (4)

[0071] In equation (4), autoregressive regularization By The coefficient matrix quantifies the total squared error of the autoregressive model fitting each time series, Z. m,t Let (m,t) be the (m,t)th data point in Z. For Z m,t Lag h i Data, To and The corresponding coefficient, l, represents the number of sequences.

[0072] Given an estimated coefficient matrix Minimizing the time variation makes each time series exhibit better consistency in local time, and can also better interpret and characterize multivariate time series matrices.

[0073] The low-rank tensor completion model based on time autoregressive regularization and stage-weighted nuclear norm is constructed as follows:

[0074] By employing tensor quantization, the traditional matrix completion problem is extended to the level of low-rank tensor completion, a transformation that significantly enhances the preservation of global data consistency. To achieve this, this application uses a truncated weighted nuclear norm minimization strategy as a tool to approximate and simplify tensor rank. With this method, the completed data not only maintains long-term global consistency but also demonstrates the ability to capture short-term local consistency. Global consistency typically refers to periodic patterns that may originate from human daily habits or inherent characteristics of certain mechanical operations, forming significant peaks at specific frequencies. Short-term trends, on the other hand, reveal temporary fluctuations or disturbances that deviate from the norm, such as brief changes caused by sudden events. Although these short-term changes are random in nature, they are prevalent over a short period.

[0075] To fully explore and utilize the global and local features inherent in multi-dimensional data, the TARLRTC-TWNN model employs a joint representation of matrices and tensors. This joint representation strategy endows the TARLRTC-TWNN model with a more comprehensive ability to capture the global and local features of the data, thereby achieving higher accuracy when performing traffic data completion.

[0076] The time-autoregressive low-rank tensor completion model obtained by adding the time-autoregressive regularization term to the objective function is shown in Equation (5):

[0077] (5)

[0078] In equation (5), It is a time series matrix composed of partial observation time series. Indicates To determine the degree of truncation, in For the tensor of the weighted vector, truncate the weighted nuclear norm, satisfying χ represents the tensor to be estimated. The observation matrix of matrix Z, This represents the observation matrix of matrix Y.

[0079] In step S400, the low-rank tensor completion model is decomposed and solved based on the ADMM framework to obtain the tensor to be estimated. Multivariate time series Dual variables Autoregressive coefficient matrix Four variable parameters. Details are as follows:

[0080] The model in this application integrates strategies for minimizing the truncated nuclear norm and minimizing local temporal variations, thereby ensuring both global and local consistency of the data. When constructing the objective function, weight parameters are introduced to balance the impact of the truncated nuclear norm and autoregressive regularization on the completion accuracy, achieving effective learning of the model both globally and locally. This balancing mechanism can be finely adjusted according to task requirements and data characteristics when facing different application scenarios, giving the model high adaptability and flexibility.

[0081] While many tensor completion models based on the nuclear norm tend to employ the Alternating Directional Multiplier (ADMM) algorithm to solve related optimization problems, the traditional ADMM algorithm is not directly applicable due to the introduction of an autoregressive coefficient matrix in this model. To address this issue, the original optimization problem is decomposed into two independent subproblems, and an alternating minimization method is used for solving them. This strategy not only successfully solves the optimization problem but also achieves a good balance between computational efficiency and model performance, resulting in superior performance in practical applications. The specific derivation steps of the model are described below:

[0082] First, by fixing the autoregressive coefficient matrix Update and The constructed subproblems are shown in formula (6):

[0083] (6)

[0084] In equation (6), This represents the number of iterations required for the ADMM algorithm to alternately minimize.

[0085] Then, Z is fixed to update the coefficient matrix A, and the constructed optimization problem is as shown in equation (7):

[0086] (7)

[0087] Once the coefficient matrix A is fixed, the first subproblem becomes a general low-rank tensor completion problem. The augmented Lagrangian function can be constructed using the ADMM algorithm, and by combining the objective function with the constraints, an unconstrained problem is formed. The augmented Lagrangian function of this optimization problem is expressed as equation (8):

[0088] (8)

[0089] In equation (8), These are the weighting parameters for the Frobenius norm. It is a dual variable. Represents time. In each iteration, to maintain consistency of observations, each time a time period is specified... The operation ensures the preservation of information from each observation. Within the ADMM framework, the augmented Lagrangian function can be iteratively transformed into the following three subproblems:

[0090] (9)

[0091] (10)

[0092] (11)

[0093] in, This represents the number of iterations in ADMM. Represents time.

[0094] In step S500, the four variable parameters are iteratively updated until a predetermined convergence criterion is reached, and the completed traffic data is output. Specifically:

[0095] (1) Update variable χ

[0096] The optimization of the tensor χ to be estimated is a problem of minimizing the truncated nuclear norm of a tensor. The truncated weighted nuclear norm of any given tensor is the sum of the truncated nuclear norms on the tensor expansion matrix minus the sum of the weight vectors, and its form is:

[0097] (12)

[0098] In equation (12), represents the tensor parameter, and k represents the tensor coefficient.

[0099] By introducing three copy tensors Used to store the information after each modality update, it can obtain each The optimal solution is:

[0100] (13)

[0101] In equation (13), fold k (·) represents the matrix being folded into a tensor modulo k, and D(·) represents the wide-area singular value thresholding algorithm related to minimizing the truncated nuclear norm. The final tensor to be completed can be obtained by the following equation (14):

[0102] (14)

[0103] (2) Update variables

[0104] By using the tensor transformation formula Substituting into formula (10), regarding The optimization problem becomes:

[0105] (15)

[0106] The optimal solution to this subproblem is given by the following derivation:

[0107] For any multivariate time series consisting of M time series at T consecutive time points... , any The autoregressive process of each element is as follows:

[0108] (16)

[0109] Based on the autoregressive coefficient matrix Time lag set The autoregressive process can also be represented by the following formula (17):

[0110] (17)

[0111] Among them, Z T Ψ represents the tensor parameter, Ψ0 represents the tensor parameter, Ψ i Represents the tensor parameter, a i T Denotes tensor parameters, A T represents the tensor parameter, and d represents the tensor parameter.

[0112] For each time series vector , The Khatri-Rao product, Ψ, is given by the following equation (18):

[0113] (18)

[0114] In equation (18), express The zero matrix, express The identity matrix, where T represents time, h d Represents tensor parameters.

[0115] When the autoregressive coefficient matrix With time lag set When known, update There are two options: one is to minimize the error in matrix form, as shown in Equation (16); the other is to minimize the error in vector form, as shown in Equation (17). The first solution involves complex operations and high computational costs. Therefore, this application chooses the second method, using vector form for optimization. The following proof is derived. The optimal solution in closed form:

[0116] Assumption And the autoregressive coefficient matrix For any The optimization problem is given by equation (19):

[0117] (19)

[0118] The optimal solution to the above optimization problem formula (19) is:

[0119] (20)

[0120] In equation (20), x m B represents the tensor parameter. m T B represents the tensor parameter. m α represents the tensor parameter, and a represents the tensor parameter. m,i Represents tensor parameters.

[0121] (3) Update the autoregressive coefficient matrix

[0122] It is the coefficient matrix in the autoregressive problem, which is updated by solving the problem of formula (7). :

[0123] (twenty one)

[0124] in, It is by The result obtained after the ADMM iteration is completed. The generated data includes all subsequent observations. This is the maximum number of iterations in ADMM. Clearly, the above optimization problem can be solved optimally using the least squares method. To prevent... For non-full rank, use generalized inverse to replace it. The inverse of the equation yields the least squares solution, and the optimal solution is given by equation (22):

[0125] (twenty two)

[0126] In formula (22), This indicates the Moore-Penrose pseudo-inverse. Represents tensor parameters.

[0127] Through the above steps, the variables , , , The algorithm iterates sequentially, continuing this process until it reaches a predetermined convergence criterion. In this application, the convergence criterion is the estimated tensor obtained from two consecutive iterations. The relative error between them is less than a certain predetermined threshold. The specific mathematical expression of the convergence judgment is shown in formula (23). The algorithm flow of the TARLRTC-TWNN model is as follows: Figure 4 As shown, it contains three key parameters: , and truncated rank Among them, parameters Used to control the penalty strength of the Frobenius norm and the degree of truncation of the truncated nuclear norm in ADMM. Parameters This is used to balance the trade-off between the truncated weighted nuclear norm of the tensor and the time autoregressive regularization term. In short, these parameters collectively affect the algorithm's accuracy in data recovery and its computational efficiency.

[0128] (twenty three)

[0129] in, Represents tensor parameters.

[0130] To better understand the traffic data completion method based on time autoregressive low-rank tensors proposed in this application, the following examples illustrate the method using the Guangzhou urban traffic dataset, the Seattle highway traffic speed dataset, and the Portland highway traffic volume dataset. The overview and data construction methods of each dataset are as follows:

[0131] The Guangzhou City Traffic Dataset contains traffic speed data for 214 road sections in Guangzhou, China, over a two-month period (August 1 to September 30, 2016). Data was collected at one sampling point every ten minutes (i.e., one sample per day). (number of data points), the size of which is... In the experiment, the data matrix was constructed with a slice length of one day (i.e., a slice length of 144) to a size of [missing information]. The tensor.

[0132] The Seattle Highway Traffic Speed ​​Dataset contains highway traffic speed data from 323 loop detectors in Seattle, USA, for the first four weeks of January 2015, at a resolution of 5 minutes (i.e., 288 time intervals per day). The data size is [data size missing]. In the experiment, the data matrix was constructed with a slice length of one day (i.e., a slice length of 288) to a size of [missing information]. The tensor.

[0133] The Portland Highway Traffic Dataset, collected from highways in the Portland-Vancouver metropolitan area, contains traffic data from 1156 loop detectors in January 2021, sampled at a frequency of 15 minutes (96 time intervals per day). The collected data size is [data missing]. In the experiment, the data matrix was constructed with a slice length of one day (i.e., a slice length of 96) to a size of [missing information]. The tensor.

[0134] To comprehensively evaluate the performance of the algorithm in imputing missing traffic data, this application constructs three different missing scenarios to simulate data missing situations in real-world environments, specifically including: random missing (RM), non-random missing (NM), and continuous missing (CM). In the random missing scenario, data loss is not constrained by specific factors but exhibits randomness; non-random missing is associated with certain specific factors and may have some discernible pattern or regularity; continuous missing involves the continuous loss of observations or variables in the dataset, which is usually caused by prolonged equipment failure or data acquisition interruption. To evaluate the imputation performance of the algorithm, RMSE and MAPE of the actual value and the true value at the missing position are used as evaluation indicators. The formulas for these two evaluation indicators are shown in equations (24) and (25), respectively:

[0135] (twenty four)

[0136] (25)

[0137] in, and These represent the actual value and the estimated value, respectively. This represents the data length.

[0138] The TARLRTC-TWNN model has several key parameters, including the learning rate. Weight parameters truncated rank and time lag set For all datasets, the learning rate... Pick Time lag collection It is used to learn local temporal correlation information. The weight ratio between the autoregressive regularization term and the Frobinus norm was controlled. The impact of different parameters on the model was verified by setting multiple sets of parameters, with the two sets of parameter values ​​being as follows: , .

[0139] Figure 5 The parameters are given in the Guangzhou traffic speed dataset. With truncated rank Heatmap of the impact of different values ​​on the model's RMSE. Figure 5 Figure (a) shows a scenario with 30% RM missing, Figure (b) shows a scenario with 30% NM missing, and Figure (c) shows a scenario with 30% CM missing. From... Figure 5 It can be seen that, in the case of RM, the optimal performance of the model is obtained when both values ​​are relatively large, but the truncated rank... The value of has little impact on the interpolation result. , To achieve optimal performance and improve model accuracy. In the NM case, a smaller truncated rank is desirable. It can improve the performance of the model, in , To achieve optimal performance and accuracy for the model. In the case of CM, , To achieve optimal model accuracy and performance. Table 2 shows the RMSE values ​​of the model for two parameter values ​​under random missing values. Numerical results indicate that for random missing values, the truncation level r has a relatively small impact on the model's completion results. The model's completion performance is best when the time is right.

[0140]

[0141] Table 2

[0142] Figure 6 A heatmap of the interpolated RMSE values ​​obtained from the model on Seattle traffic data is presented. Figure 6 Figure (a) shows the 30% RM deletion scenario, (b) shows the 30% NM deletion scenario, and (c) shows the 30% CM deletion scenario. The results indicate that for both RM and NM deletion scenarios, smaller... Ratio and larger cutoff This can significantly improve the interpolation accuracy of the model. However, for scenarios where CM is missing, when... , The model achieves optimal performance at this point. This indicates that for complex CM missing scenarios, it is necessary to enhance the performance of low-rank representations to achieve the best model performance.

[0143] Figure 7 A heatmap of the interpolated RMSE values ​​obtained from the model on Portland traffic data is presented. Figure 7 Figure (a) shows the scenario with 30% RM deletion, Figure (b) shows the scenario with 30% NM deletion, and Figure (c) shows the scenario with 30% CM deletion. The results indicate that in the RM deletion scenario, a small degree of truncation... With smaller The ratio can improve model performance, and changes in the truncation rank have a significant impact on interpolation accuracy. , To achieve optimal performance and model accuracy. In NM missing data scenarios, smaller Ratio and smaller cutoff It can significantly improve the accuracy of the model, when , The model achieves its highest accuracy during this time. However, in scenarios where CM is missing, the accuracy is significantly lower. Ratio and smaller cutoff It can significantly improve the accuracy of the model. The choice of both parameters has a great impact on the results. , The accuracy of the model reaches its highest level at that time.

[0144] Analysis of experimental results with various parameter combinations shows that model performance is affected to varying degrees by different parameter combinations under different datasets and missing data scenarios. Specifically, for continuous missing data scenarios, the introduction of temporal autoregressive regularization significantly improves model performance. This means that the model can better utilize information about time changes, thus more effectively filling in continuously missing data.

[0145] Table 3 shows the completion errors of different models under the RM model in the Guangzhou traffic dataset:

[0146]

[0147] Table 3

[0148] Table 4 shows the completion errors of different models under the NM model in the Guangzhou traffic dataset:

[0149]

[0150] Table 4

[0151] Table 5 shows the completion errors of different models under the central CM of the Guangzhou traffic dataset:

[0152]

[0153] Table 5

[0154] As shown in Tables 3, 4, and 5, the autoregressive low-rank tensor completion model outperforms other models in most cases. Comparing the autoregressive low-rank tensor completion model with the low-rank autoregressive matrix completion (LAMC) model reveals the advantage of tensor structures over matrix structures; that is, the introduction of autoregressive terms in tensor structures performs better than that in matrix structures. Compared to low-rank tensor completion models based on a single truncated weighted nuclear norm, the autoregressive low-rank tensor completion model demonstrates the introduction of a temporal autoregressive regularization term, which effectively improves imputation performance. Further comparison shows that, at the same missing rate, the overall imputation accuracy in the CM case is significantly higher than that in the RM and NM cases, indicating that continuous missing values ​​are more challenging for the completion model, but VARLRTC-TWNN still maintains imputation accuracy.

[0155] Table 6 shows the completion errors of different models under the RM model in the Seattle traffic dataset:

[0156]

[0157] Table 6

[0158] Table 7 shows the completion errors of different models under NM in the Seattle traffic dataset:

[0159]

[0160] Table 7

[0161] Table 8 shows the completion errors of different models in the Seattle traffic dataset CM:

[0162]

[0163] Table 8

[0164] Table 9 shows the completion errors of different models in the RM mode of the Portland highway traffic dataset:

[0165]

[0166] Table 9

[0167] Table 10 shows the completion errors of different models in the Portland Highway Traffic Data Set under the NM model:

[0168]

[0169] Table 10

[0170] Table 11 shows the completion errors of different models in the CM mode of the Portland highway traffic dataset:

[0171]

[0172] Table 11

[0173] Tables 6 to 11 detail the completion error results for the Seattle traffic speed dataset and the Portland traffic flow dataset under different missing data scenarios using different models. The experimental results on these two datasets also reflect the inherent temporal dependency characteristics of time-series data. The experimental results show that the TARLRTC-TWNN model performs excellently in handling various missing data patterns, especially demonstrating superior performance in dealing with the challenge of continuous data loss. For the Seattle traffic speed dataset, where temporal dependency is not very significant, the LRTC-TWNN-AW model, utilizing only low-rank characteristics, also provides excellent completion results. This phenomenon indicates that it is not always necessary to model local temporal variations in data; sometimes, low-rank characteristics are sufficient to address the challenges of data completion.

[0174] Even in complex data loss scenarios, such as the loss of data over multiple days or time periods, the TARLRTC-TWNN model can still effectively capture local smooth trends in the data. This advantage stems from the model's integration of autoregressive mechanisms and low-rank tensor completion techniques, enabling it to not only learn historical trends but also maintain the global structure of the data. In this way, even when faced with incomplete and continuously missing data, the model can accurately predict and complete the missing information, thus ensuring the continuity and reliability of traffic data analysis.

[0175] Furthermore, these experimental results further validate the model's robustness on traffic datasets of different types and sizes. Whether applied to the Seattle traffic speed dataset or the Portland traffic flow dataset, the TARLRTC-TWNN model demonstrates its powerful data imputation capabilities, providing an effective tool for processing real-world time-series data.

[0176] Figure 8 The images show the interpolation results of TARLRTC-TWNN on the Guangzhou traffic speed dataset. (a) shows the scene with 30% missing RM, (b) shows the scene with 30% missing NM, and (c) shows the scene with 30% missing CM. Figure 9 The images show the interpolation results of TARLRTC-TWNN on the Seattle traffic speed dataset. (a) shows the scene with 30% missing RM, (b) shows the scene with 30% missing NM, and (c) shows the scene with 30% missing CM. Figure 10The images show the imputation results of TARLRTC-TWNN on the Portland traffic flow dataset. In the image, (a) shows the scenario with 30% missing RM, (b) shows the scenario with 30% missing NM, and (c) shows the scenario with 30% missing CM.

[0177] Figures 8 to 10 The graph shows a comparison between the sampled data from the first two sensors and the imputed data during the first week. The blue curve represents the original data, while the thin red line depicts the imputed data. Blue markers at 0 values ​​clearly indicate the specific locations of missing data. The graph demonstrates that even in complex missing data patterns like NM and CM, the TARLRTC-TWNN model achieves high-precision data imputation by learning from observations in neighboring time series.

[0178] This application reveals the interdependencies among traffic flow data over consecutive days through in-depth correlation analysis of time-series data. To improve the imputation accuracy of time-series data and address the issue of missing data, this application proposes a low-rank tensor completion model that incorporates a time autoregressive regularization term into the low-rank tensor completion framework.

[0179] Traditional low-rank tensor completion techniques primarily focus on capturing spatial correlation and low-rank properties, sometimes neglecting the crucial role of the temporal dimension. This application significantly enhances the model's ability to capture the temporal characteristics and local variations of data by incorporating a temporal autoregressive regularization term. Leveraging the low-rank characteristics of spatiotemporal data and its continuity in the temporal dimension, this application combines temporal autoregressive regularization with low-rank tensor completion. This combination not only ensures the restoration of global data consistency by truncating the weighted nuclear norm but also guarantees the coherence of local temporal structures through temporal autoregressive regularization, thereby more accurately imputing missing temporal data.

[0180] The core of temporal autoregressive regularization lies in uncovering the autocorrelation in time series, that is, considering the correlation between the current moment and past moments. This mechanism greatly improves the accuracy of the interpolation process. This method can effectively utilize the temporal patterns and trends in historical data, providing more reliable interpolation results for spatiotemporal data. By combining correlation analysis of time series data with the introduction of temporal autoregressive regularization terms to model local temporal neighborhoods, this application provides a basis for understanding the local temporal variation patterns of spatiotemporal data.

[0181] The experimental results of this application verify the significant advantages of the time autoregressive low-rank tensor completion method in time-series data recovery, especially its superior performance in handling continuous missing data scenarios, which significantly outperforms other comparative methods. This indicates that the method has excellent adaptability and efficiency in handling complex data missing patterns, providing an effective technical means for accurate data imputation.

[0182] Another embodiment of the present invention provides a traffic data completion device based on a time-autoregressive low-rank tensor, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a traffic data completion method program based on a time-autoregressive low-rank tensor. When the processor executes the computer program, it implements the steps in the above-described embodiments of the traffic data completion method based on a time-autoregressive low-rank tensor, for example... Figure 1 The steps.

[0183] For example, the above-mentioned computer program can be divided into one or more modules / units, which are stored in memory and executed by a processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions. These instruction segments describe the execution process of the computer program in a traffic data completion device based on time-autoregressive low-rank tensors. For example, the computer program can be divided into a data acquisition module, a data representation module, a model building module, a variable acquisition module, and a variable update module. The specific functions of each module are as follows:

[0184] The data acquisition module is used to acquire incomplete traffic data; the data representation module is used to construct a three-dimensional tensor to represent the incomplete traffic data, obtaining the observation index set and tensor matrix; the model construction module is used to obtain a low-rank tensor completion model based on a rank minimization low-rank tensor completion model framework, introducing time autoregressive regularization and truncated weighted nuclear norm; the variable acquisition module is used to decompose and solve the low-rank tensor completion model based on the ADMM framework, obtaining... , , , Four variable parameters; the variable update module is used to update the variables. , , , The system iterates and updates sequentially until the predetermined convergence criteria are met, then outputs the completed traffic data.

[0185] Traffic data completion devices based on time autoregressive low-rank tensors can be computing devices such as desktop computers, laptops, handheld computers, and cloud servers. These devices may include, but are not limited to, processors and memory; for example, they may also include output devices, network access devices, and buses.

[0186] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the time-autoregressive low-rank tensor-based traffic data completion device, connecting all parts of the device via various interfaces and lines.

[0187] The memory can be used to store computer programs and / or modules. The processor implements various functions of the traffic data completion device based on time autoregressive low-rank tensors by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory.

[0188] If the integrated modules / units of the traffic data completion device based on time autoregressive low-rank tensors are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various embodiments of the traffic data completion method based on time autoregressive low-rank tensors described above.

[0189] Computer programs include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0190] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications and substitutions should be covered within the scope of the claims of the present invention. Technical aspects, shapes, and structures not described in detail in this invention are all well-known technologies.

Claims

1. A traffic data completion method based on time autoregressive low-rank tensors, characterized in that, Includes the following steps: Obtaining incomplete traffic data; A three-dimensional tensor is constructed to represent the incomplete traffic data, and an observation index set and tensor matrix are obtained. Based on the low-rank tensor completion model with rank minimization as the framework, time autoregressive regularization and truncated weighted nuclear norm are introduced to obtain the low-rank tensor completion model. The low-rank tensor completion model is decomposed and solved using the ADMM framework to obtain the tensor to be estimated. Multivariate time series Dual variables Autoregressive coefficient matrix Four variable parameters; Treating the estimation tensor Multivariate time series Dual variables Autoregressive coefficient matrix The system iterates and updates sequentially until the predetermined convergence criteria are met, then outputs the completed traffic data.

2. The traffic data completion method based on time autoregressive low-rank tensor according to claim 1, wherein, The steps of constructing a three-dimensional tensor to represent incomplete traffic data and obtaining an observation index set and tensor matrix include: The actual data matrix Y collected by M sensors is obtained as follows: ; In the real data matrix Y, each row of data corresponds to a different time series collected by the sensor, the column corresponds to a time point, I represents the number of blocks required for tensor quantization, and J represents the size of the tensor; Determine the observation period and conduct observations to obtain the observation matrix. The elements in the observation index set must satisfy: ; Where Ω represents the observation set, m represents the row dimension of the matrix, and n represents the column dimension of the matrix; We introduce positive tensor quantization to perform tensor quantization operations, transforming the multivariate time series matrix Y into a third-order tensor, as follows: ; The third-order tensor is matrixed to obtain: 。 3. The traffic data completion method based on time autoregressive low-rank tensor according to claim 1, wherein, The time autoregressive regularization is as follows: using a given time series matrix... and an autoregressive coefficient matrix And a time lag set ,matrix The time autoregressive model can be expressed as: ; in, Z represents the total squared error of the quantized time series. m,t Let (m,t) be the (m,t)th data point in Z. For Z m,t Lag h i Data, To and The corresponding coefficient, l, represents the number of sequences.

4. The traffic data completion method based on time autoregressive low-rank tensor according to claim 3, wherein, The low-rank tensor completion model is represented as follows: ; in, It is a time series matrix composed of partial observation time series. Indicates To determine the degree of truncation, in For the tensor of the weighted vector, truncate the weighted nuclear norm, satisfying χ represents the tensor to be estimated. The observation matrix of matrix Z, This represents the observation matrix of matrix Y.

5. The traffic data completion method based on time autoregressive low-rank tensor according to claim 4, wherein, The steps for decomposing and solving the low-rank tensor completion model based on the ADMM framework to obtain four variable parameters are as follows: Update by fixing the autoregressive coefficient matrix A and The constructed subproblems are as follows: ; Where ℓ represents the number of iterations for alternating minimization in the ADMM algorithm; fixed To update the coefficient matrix The optimization problem for the construction is as follows: ; The augmented Lagrangian function is constructed using the ADMM algorithm as follows: ; in, These are the weighting parameters for the Frobenius norm. It is a dual variable. Represents time; Within the ADMM framework, the augmented Lagrangian function is iteratively transformed into three subproblems as follows: ; ; ; in, This represents the number of iterations in ADMM. Represents time.

6. The traffic data completion method based on time autoregressive low-rank tensor according to claim 5, wherein, The variable χ is updated as follows: the truncated weighted nuclear norm of any given tensor is the sum of the truncated nuclear norms on the tensor expansion matrix minus the sum of the weight vectors, as follows: ; in, represents the tensor parameter, and k represents the tensor coefficient; Introducing three copy tensors Used to store the information after each modality update, to obtain each The optimal solution is as follows: ; The tensors that need to be completed are as follows: ; Among them, fold k (·) represents the matrix being folded into a tensor modulo k, and D(·) represents the wide-area singular value thresholding algorithm related to minimizing the truncated nuclear norm.

7. The traffic data completion method based on time autoregressive low-rank tensor according to claim 5, wherein, The variable The update is as follows: Tensor transform formula Substitute In China, regarding The optimization problem is as follows: ; Based on the autoregressive coefficient matrix Time lag set , The autoregressive process is as follows: ; Among them, Z T Ψ represents the tensor parameter, Ψ0 represents the tensor parameter, Ψ i Represents the tensor parameter, a i T Denotes tensor parameters, A T represents the tensor parameter, and d represents the tensor parameter; For each time series vector , The Khatri-Rao product, Ψ, is given by the following formula: ; In the above formula, express The zero matrix, express The identity matrix, where T represents time, h d Represents tensor parameters; Assumption And the autoregressive coefficient matrix For any The optimization problem is as follows: ; The optimal solution is as follows: ; in, x m B represents the tensor parameter. m T B represents the tensor parameter. m α represents the tensor parameter, and a represents the tensor parameter. m,i Represents tensor parameters.

8. The traffic data completion method based on time autoregressive low-rank tensor according to claim 5, wherein, The autoregressive coefficient matrix The update is as follows: The coefficient matrix in the autoregressive problem is updated as follows: ; in, It is by The result obtained after the ADMM iteration is completed. The generated data includes all subsequent observations. This is the maximum number of iterations in ADMM; Using generalized reverse substitution The inverse of is used to obtain the least squares solution. The optimal solution is as follows: ; in, This indicates the Moore-Penrose pseudo-inverse. Represents tensor parameters.

9. The traffic data completion method based on time autoregressive low-rank tensor according to claim 1, wherein, The predetermined convergence criterion is the estimated tensor obtained from two consecutive iterations. The relative error between them is less than a certain predetermined threshold. ,as follows: ; in, Represents tensor parameters.

10. A traffic data completion device based on time autoregressive low-rank tensors, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-8.

Citation Information

Cited By

  • Method and device for repairing low-rank single-sensor building energy consumption data

    CN121880705A

  • A low-rank single-sensor building energy consumption data repairing method and device

    CN121880705B