A pedestrian flow walking data completion method based on tensor decomposition and reconstruction

By improving the tensor decomposition and reconstruction algorithm, and combining the genetic algorithm and the adaptive gradient algorithm, the problem of high-dimensional sparse pedestrian flow data completion was solved, achieving higher data completion accuracy and efficiency, which is suitable for pedestrian dynamics research.

CN116881644BActive Publication Date: 2026-04-07BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively complete high-dimensional and sparse pedestrian flow data, resulting in poor data completion, especially in large-scale, real-world scenarios where it is difficult to obtain complete pedestrian flow data.

Method used

We employ a tensor decomposition and reconstruction approach, constructing an improved CP decomposition weighted optimization tensor decomposition and reconstruction algorithm. By combining genetic algorithm and adaptive gradient algorithm, we optimize the rank, regularization parameter and learning rate of the tensor, and complete the missing values ​​in the sparse pedestrian velocity tensor through tensor decomposition and reconstruction.

Benefits of technology

It improves the completion accuracy of sparse pedestrian flow data, is suitable for high-dimensional sparse data scenarios, and provides higher data completion accuracy and efficiency, making it suitable for pedestrian dynamics research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116881644B_ABST
    Figure CN116881644B_ABST
Patent Text Reader

Abstract

The application provides a pedestrian flow walking data completion method based on tensor decomposition and reconstruction ideas. The method comprises the following steps: constructing a tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization; using a genetic algorithm to optimize parameters of tensor rank R, regularization parameter and learning rate r in the algorithm η Three hyperparameters are optimized, and after the data completion effect of the improved algorithm is verified to be qualified, a trained tensor decomposition and reconstruction algorithm based on CP decomposition weighted optimization is obtained; high-dimensional sparse pedestrian walking speed tensors in different scenes are constructed, and the trained tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization is used for tensor decomposition and reconstruction to complete missing values in the sparse pedestrian speed tensor. The method can effectively predict missing values in high-dimensional and sparse data through a data-driven model for scene expansion, and is suitable for completion of high-dimensional and sparse data in pedestrian dynamics research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer application technology, and in particular to a method for completing pedestrian flow walking data based on the idea of ​​tensor decomposition and reconstruction. Background Technology

[0002] Complete, real-world pedestrian flow data is fundamental to the study of pedestrian dynamics. Currently, there are four main methods for obtaining pedestrian flow data:

[0003] (1) Manual data collection mainly includes manual counting, tracking and recording, intention survey (Stated Preference, SP), and behavior survey (Revealed Preference, RP), which involves manually collecting real pedestrian walking data on site.

[0004] (2) Indoor positioning technology refers to the method of positioning a terminal indoors by receiving signal data transmitted from a signal base station, uploading it to a server, and then calculating the coordinates of the terminal based on a positioning algorithm. This mainly includes positioning technologies such as UWB, WiFi, GPS, and Bluetooth. Automatic video analysis uses software such as PeTrack and Tracker to process real videos of pedestrian traffic, automatically extracting accurate pedestrian trajectories and providing data such as the location, speed, density, and flow of pedestrians at various time steps. Pedestrian detection and tracking algorithms use deep learning algorithms to process pedestrian videos and acquire pedestrian traffic data. Due to limitations in pedestrian walking environments and concerns about pedestrian safety and privacy, it is difficult to obtain large-scale, realistic pedestrian traffic data in many scenarios. Completing missing pedestrian traffic data has become a primary problem to be solved in the study of pedestrian traffic and pedestrian dynamics.

[0005] Currently, the main methods for pedestrian flow data completion in existing technologies are:

[0006] (1) Data completion algorithms based on statistics, including time series method, linear regression model, Kalman filter, etc., to complete missing values ​​by statistically analyzing the changing patterns of pedestrian flow parameters;

[0007] (2) Data completion algorithms based on machine learning, including algorithms based on probabilistic graphical models such as Bayesian theory and Markov, as well as nonparametric estimation methods such as decision trees and support vector machines;

[0008] (3) Data completion algorithms based on deep learning. Commonly used network structures are convolutional neural networks (CNN) and recurrent neural networks (RNN). These algorithms capture the spatiotemporal variation characteristics of pedestrian walking parameters and learn them to complete missing data.

[0009] The shortcomings of the existing pedestrian flow data completion methods mentioned above include: most data completion algorithms rely on traditional statistical methods such as averaging and interpolation, which are not suitable for highly sparsity data, resulting in poor completion performance. Deep learning and intelligent algorithms require large amounts of real-world data for training and learning, but large-scale, complete pedestrian flow data is difficult to obtain in real-world scenarios. In other words, pedestrian flow data obtained in real-world scenarios is high-dimensional and sparse, making it difficult for the aforementioned data completion algorithms to achieve high completion accuracy for sparse pedestrian flow data. Summary of the Invention

[0010] The embodiments of the present invention provide a pedestrian flow data completion method based on the idea of ​​tensor decomposition and reconstruction, so as to effectively utilize the high-dimensional sparsity of pedestrian flow data.

[0011] To achieve the above objectives, the present invention adopts the following technical solution.

[0012] A method for pedestrian flow data completion based on tensor decomposition and reconstruction includes:

[0013] Construct a tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization;

[0014] The rank R and regularization parameter of the tensor in the tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization are analyzed using a genetic algorithm. and learning rate r η The three hyperparameters were tuned, and after verifying that the data completion effect of the tensor decomposition and reconstruction algorithm based on the improved CP decomposition and weighted optimization was qualified, the trained tensor decomposition and reconstruction algorithm based on the improved CP decomposition and weighted optimization was obtained.

[0015] We construct high-dimensional sparse pedestrian walking speed tensors for different scenarios, and use a trained tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization to decompose and reconstruct the tensors, and fill in the missing values ​​in the sparse pedestrian speed tensors.

[0016] Preferably, the construction of the tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization includes:

[0017] Based on the CP decomposition weighted optimization algorithm, from the perspective of optimizing the weight objective function value and variable gradient value, the CPWOPT algorithm is improved using an adaptive gradient algorithm, resulting in an improved CPWOPT* tensor decomposition and reconstruction algorithm. The processing flow of the improved CPWOPT* tensor decomposition and reconstruction algorithm includes:

[0018] The objective function in the CPWOPT algorithm is regularized, and a third-order pedestrian walking parameter tensor model is constructed. Its factor matrix after decomposition is in, The objective function for optimizing the weights in tensor decomposition is:

[0019]

[0020] Add a Tikhonov regularization term to the objective function (29):

[0021]

[0022] in, The factor matrix weight decay parameter reflects the factor matrix S (n) Throughout the objective function The importance of it;

[0023] The optimized weight objective function is rewritten in the following form:

[0024]

[0025] in, Pedestrian velocity tensor It is certain that throughout the entire iteration process, the tensor It remains unchanged and can be calculated in advance, while the regularization term... The sum of squares of the Frobenius norms of each factor matrix is ​​given. At the beginning of each iteration, the factor matrix is ​​fixed, and the regularization term is calculated directly.

[0026] sparse tensors The index of a known value in the binary indicator tensor Positions where the element is 1 are added to the set in order. Given a three-dimensional vector Q∈{1,2,…,Q}, convert the tensor... The known values ​​are stored in a vector y of length Q, that is:

[0027]

[0028] for tensor The missing value is located in the corresponding tensor. The value in the middle must also be 0; only calculation is needed. corresponding tensor The elements in the vector z represent the tensor. The known values ​​in;

[0029]

[0030] During computer operations, let:

[0031]

[0032]

[0033] The method for calculating vector u is as follows:

[0034]

[0035] The method for calculating vector z is as follows:

[0036]

[0037] One vector is computed at a time in each iteration. The optimized weight objective function (31) is rewritten as:

[0038]

[0039] Where γ is a constant,

[0040] The partial derivatives of the optimized weight objective function with respect to each factor matrix are:

[0041]

[0042] Among them, Z (n) Y (n) For tensor n-order matrix transformation, S -(n) The calculation formula is as follows:

[0043] S (-1) =S (3) ⊙S (2) S (-2) =S (3) ⊙S (1) S (-3) =S (2) ⊙S (2) (40)

[0044] The learning rate of the tensor decomposition and reconstruction algorithm based on the improved CP decomposition weighted optimization is adaptively adjusted using an adaptive gradient algorithm.

[0045] Preferably, the rank R and regularization parameter of the tensor in the tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization using a genetic algorithm are... and learning rate r η Three hyperparameters were tuned, including:

[0046] The genetic algorithm population is initialized by determining the values ​​of initialization parameters, including population size, maximum number of iterations, crossover rate, mutation rate, and global optimal fitness initialization value. Chromosomes in the population are encoded using binary encoding, with a gene length of 10. The first 4 bits of the gene represent the rank R of the tensor, and the second 3 bits represent the regularization parameter. The third part, consisting of 3 digits, represents the learning rate r. η ;

[0047] The first generation population is randomly generated, and the fitness value of chromosomes is calculated using the fitness function. The fitness function is used to determine the fitness value of a certain chromosome to the environment. Based on the fitness function, chromosomes in the population are selected, and individuals with higher fitness values ​​are retained for crossover and mutation. The objective function f, expressed by formula (30), is used. w Fitness is represented by the reciprocal of the value.

[0048] Determine if the iteration termination condition has been met. If not, perform selection, crossover, and mutation operations on the population to obtain the next generation population, and continue to evaluate the population using the fitness function. If the termination condition has been met, output the globally optimal chromosome and fitness value.

[0049] Using a single-point crossover operator, a crossover point is randomly selected, and the two selected chromosomes are split. The binary values ​​at corresponding positions after the crossover point are swapped, resulting in two distinct chromosomes. Using a basic bit mutation operator, a mutation point is randomly determined on the selected chromosome, and the gene value is replaced with other alleles of that gene to form a new individual. A genetic algorithm is used for iterative optimization to determine the optimal value in each generation of the population. When the termination condition is met, the optimal chromosome value is the rank R and the regularization parameter. and learning rate r η The value of is converted to binary to obtain the rank R and the regularization parameter. and learning rate r η The specific values ​​of the three hyperparameters.

[0050] Preferably, the step of obtaining the trained tensor decomposition and reconstruction algorithm based on improved CP decomposition and weighted optimization after verifying that the data completion effect of the tensor decomposition and reconstruction algorithm based on improved CP decomposition and weighted optimization is qualified includes:

[0051] The pedestrian velocity tensor constructed using the tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization. The known values ​​in the data were randomly missing, with the missing rate gradually increasing from 10% to 60%. Evaluation indicators of the completion results under different methods were calculated.

[0052] Mean Absolute Error (MAE) represents the average distance between the estimated value and the true value. MAE is a non-negative value; the closer it is to 0, the more accurate the algorithm's prediction. MAE is calculated as follows:

[0053]

[0054] Where, x i Let be the true value of the walking parameters of the i-th pedestrian. The corresponding estimated value;

[0055] The root mean square error (RMSE) measures the deviation between the estimated value and the true value. The smaller the RMSE value, the more accurate the algorithm's prediction. The RMSE is calculated as follows:

[0056]

[0057] By setting different missing rates, the completion effects of the tensor decomposition and reconstruction algorithm based on CP decomposition and weighted optimization are compared with those of existing data completion algorithms. After verifying that the data completion effect of the tensor decomposition and reconstruction algorithm based on CP decomposition and weighted optimization is qualified, the trained tensor decomposition and reconstruction algorithm based on CP decomposition and weighted optimization is obtained.

[0058] Preferably, the construction of high-dimensional sparse pedestrian walking speed tensors under different scenarios involves decomposing and reconstructing the tensors using a trained tensor decomposition and reconstruction algorithm based on CP decomposition and weighted optimization, and filling in missing values ​​in the sparse pedestrian speed tensors, including:

[0059] ① Construct rank R for different scenarios s Sparse pedestrian walking speed tensor and the same-dimensional binary indicator tensor And calculate constants

[0060] ②Tenometer The known value index is placed in And store the known values ​​in the vector y;

[0061] ③ Use a genetic algorithm to determine the hyperparameter regularization parameter. The rank of the tensor R S Learning rate r η The value of is determined using the CP-ALS algorithm on the tensor. Decomposition yields the initial factor matrix S. (1) ,S (2) ,S (3) ;

[0062] ④ Construct a regularized objective function

[0063] ⑤ Order And construct vector z,

[0064] ⑥ Calculate the objective optimization function If the objective function value reaches the loss threshold or the maximum number of iterations J max If yes, proceed to step ⑧; otherwise, proceed to step ⑦.

[0065] ⑦ Calculate the partial derivatives of each factor matrix. And the gradient is updated using the AdaGread algorithm:

[0066] v = v + G (n) ⊙G (n)

[0067]

[0068] S (n) =S(n)-η·V -1 ·G(n)

[0069] Then proceed to step ⑤;

[0070] ⑧ Output the optimal factor matrix S (1) ,S (2) ,S (3) And use the following formula for tensors Fill in the missing values ​​in the tensor:

[0071]

[0072] Complete the evaluation of pedestrian walking parameters based on tensor decomposition and reconstruction.

[0073] As can be seen from the technical solutions provided by the embodiments of the present invention above, the present invention provides a method for pedestrian flow data completion based on the idea of ​​tensor decomposition and reconstruction. By expanding the scene through a data-driven model, it can effectively predict missing values ​​in high-dimensional and sparse data. It is suitable for completing high-dimensional sparse data in pedestrian dynamics research and provides a new and advantageous tool for the effective utilization of pedestrian flow data with typical high-dimensional sparsity.

[0074] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of the invention. Attached Figure Description

[0075] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0076] Figure 1 A flowchart illustrating a pedestrian flow data completion method based on tensor decomposition and reconstruction, provided for an embodiment of the present invention;

[0077] Figure 2 A flowchart illustrating the specific process of calculating tensor decomposition using an improved CPWOPT algorithm provided in this embodiment of the invention;

[0078] Figure 3 The following is an iterative flowchart of the AdaGrad algorithm provided in an embodiment of the present invention;

[0079] Figure 4 This is a schematic diagram illustrating the crossover and mutation process of the genetic algorithm provided in an embodiment of the present invention;

[0080] Figure 5 This is a flowchart of the genetic algorithm for obtaining model hyperparameters provided in an embodiment of the present invention;

[0081] Figure 6 This is a diagram illustrating the iterative optimization process of the hyperparameters of the velocity tensor model based on a genetic algorithm, provided in an embodiment of the present invention.

[0082] Figure 7 The mean absolute error of each method under different data missingness levels provided in the embodiments of the present invention;

[0083] Figure 8 The root mean square error of each method under different data missingness levels provided in the embodiments of the present invention. Detailed Implementation

[0084] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0085] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or couplings. The term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.

[0086] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.

[0087] To facilitate understanding of the embodiments of the present invention, the following will provide further explanation and description with reference to the accompanying drawings and several specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.

[0088] The flowchart of a pedestrian flow data completion method based on tensor decomposition and reconstruction provided in this embodiment of the invention is as follows: Figure 1 As shown, the processing steps include the following:

[0089] Step S10: Construct the improved CPWOPT* tensor decomposition and reconstruction algorithm.

[0090] Step S20: Use a genetic algorithm to refine the rank R and regularization parameter of the tensors in the improved CPWOPT* tensor decomposition and reconstruction algorithm. Learning rate r η After the three hyperparameters were fine-tuned and the data completion effect of the improved CPWOPT* tensor decomposition and reconstruction algorithm was verified to be satisfactory, the trained improved CPWOPT* tensor decomposition and reconstruction algorithm was obtained.

[0091] Step S30: Construct high-dimensional sparse pedestrian walking speed tensors for different scenarios, decompose and reconstruct the tensors using the trained and improved CPWOPT* tensor decomposition and reconstruction algorithm, and fill in the missing values ​​in the sparse pedestrian speed tensors.

[0092] This invention first presents an improved CPWOPT* tensor decomposition and reconstruction algorithm, which is proposed for the first time and is the biggest innovation of this invention. It is mainly based on the existing CP decomposition weighted optimization algorithm (CPWOPT algorithm), and improves the CPWOPT algorithm from the perspective of optimizing the weight objective function value and variable gradient value by using an adaptive gradient algorithm.

[0093] The second step is to use a genetic algorithm to fine-tune the hyperparameters of the proposed algorithm and determine the hyperparameter values ​​for the CPWOPT* algorithm. This step is essential for the algorithm's feasibility and is also an innovation of this invention. Other algorithms may also use similar evolutionary algorithms (including but not limited to genetic algorithms). Therefore, this technology is an application of existing technology to a completely new algorithm and applicable scenario.

[0094] The third step is to verify the proposed algorithm, proving its rationality and accuracy. These three steps are innovative proposals of this invention, while the technical means used in the second step are existing achievements.

[0095] The fourth step is algorithm application. The proposed tensor decomposition and reconstruction algorithm is used to process high-dimensional, sparse pedestrian data within urban rail stations.

[0096] Specifically, step S10 includes, based on the CP (CANDECOMP / PARAFAC) decomposition weighted optimization CPWOPT (CPWeighted OPTimization, CPWOPT) algorithm, improving the CPWOPT algorithm from the perspective of optimizing the weight objective function value and variable gradient value, and proposing a tensor decomposition and reconstruction algorithm based on CPWOPT*.

[0097] The CPWOPT algorithm improves performance by weighting the objective function, optimizing the estimation error of known data, and ignoring missing data. The specific calculation process of the CPWOPT algorithm is as follows.

[0098] First, we give a K-order tensor with missing data. And assume that its CP rank is known. For tensor Same-dimensional binary indicator tensor, i.e., tensor The index of a known element in a tensor The element at the corresponding position is 1, and the missing element is 0, as in formula (16).

[0099]

[0100] Where: m k ∈{1,2,…,M K}, k∈{1,2,…,K}. M1×M2×…×M K

[0101] This represents the K-order tensor mentioned above.

[0102] Given a series of factor matrices Then the factor matrix It can be represented as M1×M2×…×M K tensor, A represents the factor vector of the factor matrix. (1) A (2) ,…,A (K) Represents the factor matrix.

[0103] Tensor elements can be represented as:

[0104]

[0105] The optimization objective of the CPWOPT algorithm is to minimize the error between the original tensor and the reconstructed tensor. Its weight objective function is:

[0106]

[0107] make The weighted objective function (18) can then be written in the following form:

[0108]

[0109] Among them, tensor The data is known within the original tensor and does not change during the iteration process.

[0110] The objective function for weights can be written in tensor element form as follows:

[0111]

[0112] Next, the process of calculating the objective function and its gradient is given, which allows the gradient-based optimization algorithm to solve the optimization problem. Objective function f w The partial derivatives with respect to each element of the factor matrix are calculated as follows:

[0113]

[0114] Where, m k =1,2,…,M k , k=1,2,…,K, r=1,2,…,R.

[0115] From formula (16), we know that For a binary indicator tensor, then:

[0116]

[0117]

[0118] Therefore, formula (21) can be written in matrix form as follows:

[0119]

[0120] A -(k) =A (K) ⊙…⊙A (k+1) ⊙A (k-1) ⊙…⊙A (1) (25) The factor matrix of the tensor is updated during the iteration process as follows:

[0121] A (k) =A (k) -η·G (k) (26)

[0122] The algorithm terminates when the objective function error reaches a threshold or the maximum number of iterations is reached, yielding the optimal factor matrix A. (1) A (2) ,…,A (K) Then the missing data in the original tensor can be calculated using the following formula.

[0123] Where 1 represents the tensor Tensors of the same dimension with all elements being 1. Original tensor The completed tensor can be represented as Right now:

[0124]

[0125] (2) Figure 2 A flowchart of an improved CPWOPT algorithm for calculating tensor decomposition is provided in this embodiment of the invention. The specific processing steps include:

[0126] ① Objective function optimization

[0127] To improve the robustness and accuracy of the algorithm and prevent overfitting of the objective function, regularization is applied to the objective function in the CPWOPT algorithm. A third-order pedestrian walking parameter tensor model is constructed. Its factor matrix after decomposition is in, The objective function for optimizing the weights in tensor decomposition is:

[0128]

[0129] Add a Tikhonov regularization term to the objective function (29):

[0130]

[0131] in, The factor matrix weight decay parameter reflects the factor matrix S (n) Throughout the objective function The importance of it.

[0132] The optimized weight objective function is rewritten in the following form:

[0133]

[0134] in, Due to pedestrian velocity tensor It is deterministic; therefore, throughout the entire iteration process, the tensor... The variable remains unchanged and can be calculated in advance. The regularization term, however, remains unchanged. This is the sum of squares of the Frobenius norms of each factor matrix. At the start of each iteration, the factor matrices are fixed, and the regularization term can be calculated directly.

[0135] For sparse tensors To improve computational efficiency and reduce storage space, an efficient method for computing tensors is presented. The method. Tensor The index of a known value in the binary indicator tensor Positions where the element is 1 are added to the set in order. Given a three-dimensional vector, q∈{1,2,…,Q}. This operation transforms the tensor... The known values ​​are stored in a vector y of length Q, that is:

[0136]

[0137] therefore, for tensor The missing value is located in the corresponding tensor. The value in the middle must also be 0. Therefore, it is only necessary to calculate corresponding tensor The elements are sufficient. Given a vector z, represent the tensor. The known values ​​in.

[0138]

[0139] During computer operations, let:

[0140]

[0141]

[0142] Vector u can be calculated as:

[0143]

[0144] Then the vector z can be calculated as:

[0145]

[0146] Only one vector needs to be computed at a time in each iteration. This significantly reduces storage costs and improves algorithm efficiency. At this point, the optimized weight objective function (31) can be written as:

[0147]

[0148] Where γ is a constant, The above fast algorithm can effectively calculate the optimized weight objective function value, greatly improving computational efficiency.

[0149] The partial derivatives of the optimized weight objective function with respect to each factor matrix are:

[0150]

[0151] Among them, Z (n) Y (n) For tensor n-order matrix transformation, S -(n) The calculation formula is as follows:

[0152] S (-1) =S (3) ⊙S (2) S (-2) =S (3) ⊙S (1) S (-3) =S (2) ⊙S (2) (40)

[0153] ② The AdaGrad algorithm optimizes the learning rate.

[0154] To prevent the stochastic gradient algorithm from encountering "valleys" and "saddle points" that prevent convergence or cause extremely slow convergence, an adaptive gradient algorithm is used to adaptively adjust the learning rate. The iterative process of the AdaGrad algorithm is as follows:

[0155]

[0156]

[0157]

[0158] in, I represents the diagonal matrix of the element-wise product of the gradients taken in the first j steps; I is a unit symmetric matrix; ε is the smoothing coefficient, usually taken as 1e-8; the learning rate r η The hyperparameters are fixed. The AdaGrad algorithm uses an adaptive matrix. It replaces the fixed learning rate in the SGD algorithm. Because... The diagonal elements of the matrix represent the update weights for each dimension's parameters, allowing the learning rate to incorporate historically accumulated squared gradient information and reflecting the differences between dimensions when updating the step size. The specific steps of the AdaGrad algorithm are as follows: Figure 3 .

[0159] 2. A genetic algorithm is proposed to optimize the rank R and regularization parameter of the tensor in the CPWOPT* algorithm. Learning rate r η Three hyperparameters were tuned, including:

[0160] First, the population is initialized by determining the values ​​of initialization parameters, including population size, maximum number of iterations, crossover rate, mutation rate, and initial global optimal fitness. Specifically, chromosomes in the population are encoded using binary encoding, and the gene length is set to 10. The first part of the gene represents the rank R of the tensor, and the second part represents the regularization parameter. The third part is the learning rate r. η .

[0161] Secondly, a first-generation population is randomly generated, and the fitness value of chromosomes is calculated using a fitness function. The fitness function determines the fitness value of a particular chromosome to the environment; chromosome selection within the population is based on this fitness function. Individuals with higher fitness values ​​are retained for crossover and mutation, thus accelerating the search for the optimal value. Since the genetic algorithm restricts fitness values ​​to be non-negative, and the tensor model requires finding the minimum value of the objective function, the objective function f is used... w Fitness is represented by the reciprocal of the number.

[0162] Next, it is determined whether the iteration termination condition has been met. If the termination condition has not been met, selection, crossover, and mutation operations are performed on the population to obtain the next generation, and the fitness function is used to evaluate the population again. If the termination condition is met, the globally optimal chromosome and its fitness value are output. Specifically, a single-point crossover operator is used to randomly select a crossover point, and the two selected chromosomes are split, exchanging the binary values ​​at the corresponding positions after the crossover point to obtain two different chromosomes. A basic bit mutation operator is used to randomly determine the mutation point on the selected chromosome, replacing the gene value with other alleles of that gene to form a new individual. A schematic diagram of the crossover and mutation process in the genetic algorithm is shown below. Figure 4 .

[0163] Finally, iterative optimization is performed using a genetic algorithm. The process of obtaining the model hyperparameters using a genetic algorithm is as follows: Figure 5 The optimal set of hyperparameters is obtained, and the iterative optimization process is as follows: Figure 6 .

[0164] 3. Compare the data completion effects of the algorithm of this invention with those of classic data completion algorithms, including:

[0165] For the constructed pedestrian velocity tensor The known values ​​in the model are randomly missing, with the missing value rate gradually increasing from 10% to 60%. By comparing the completion effects of commonly used data completion methods and the CPWOPT* algorithm, the evaluation index of the completion results under different methods is calculated to verify the effectiveness of the model.

[0166] (1) Algorithm evaluation metrics

[0167] ①MAE

[0168] Mean Absolute Error (MAE) represents the average distance between the estimated value and the true value. MAE is a non-negative value; the closer it is to 0, the more accurate the algorithm's prediction. MAE is calculated as follows:

[0169]

[0170] Where, x i Let be the true value of the walking parameters of the i-th pedestrian. This is the corresponding estimated value.

[0171] ②RMSE

[0172] The root mean square error (RMSE) measures the deviation between the estimated value and the true value. A smaller RMSE value indicates a more accurate prediction. RMSE is calculated as follows:

[0173]

[0174] (2) Data completion effect analysis

[0175] By setting different missing percentages, the completion effects of the CPWOPT* algorithm were compared with those of classic data completion algorithms, namely: CPWOPT algorithm, CP-ALS algorithm, and Linear Interpolation (LI). The MAE (Maximum Achievement Value) after completion for each algorithm was also calculated (see [link to relevant documentation]). Figure 7 ) and RMSE values ​​(see Figure 8The results show that the filling effect of various data completion algorithms decreases with the increase of missing data in the tensor, especially when the missing data rate reaches 50% or above, the filling error increases and the filling effect is poor. Comparing various data completion algorithms, the data filling effect based on tensor decomposition and reconstruction is better than that of the classic linear interpolation method. For velocity tensors with a missing rate of 60%, the average absolute error of the CPWOPT* algorithm is only 0.260 m / s, the CPWOPT algorithm is 0.276 m / s, the CP-ALS algorithm is 0.322 m / s, while the linear interpolation method is 0.417 m / s. Because the linear interpolation method only considers local information, even when the missing data is small, the filling error is still large; for a velocity tensor with a missing data rate of 30%, the average absolute error after filling reaches 0.195 m / s. Furthermore, when the missing data rate of tensor data is 30% or less, the accuracy of the improved CPWOPT* algorithm and the CPWOPT algorithm are not much different, and both have good filling effect. However, when the missing data rate is large, the CPWOPT* algorithm has a better filling effect. In other words, the improved CPWOPT* algorithm is more suitable for processing sparse data and has a higher filling accuracy.

[0176] 4. Construct high-dimensional sparse pedestrian walking velocity tensors for different scenarios, and use the CPWOPT* algorithm to decompose and reconstruct the tensors, filling in missing values ​​in the sparse pedestrian velocity tensors, including:

[0177] The main steps are as follows:

[0178] (1) Construct high-dimensional sparse pedestrian velocity tensors for different scenarios;

[0179] (2) The CPWOPT* algorithm is used to decompose and reconstruct the tensor, and the missing values ​​in the sparse pedestrian velocity tensor are filled in. The specific steps are as follows:

[0180] ① Construct a rank R S tensor and the same-dimensional binary indicator tensor And calculate constants

[0181] ②Tenometer The known value index is placed in And store the known values ​​in the vector y;

[0182] ③ Use a genetic algorithm to determine the hyperparameter regularization parameter. The rank of the tensor R S Learning rate r η The value of is determined using the CP-ALS algorithm on the tensor. Decomposition yields the initial factor matrix S. (1) ,S(2) ,S (3) ;

[0183] ④ Construct a regularized objective function

[0184] ⑤ Order And construct vector z,

[0185] ⑥ Calculate the objective optimization function If the objective function value reaches the loss threshold or the maximum number of iterations J max If yes, proceed to step ⑧; otherwise, proceed to step ⑦.

[0186] ⑦ Calculate the partial derivatives of each factor matrix. And the gradient is updated using the AdaGread algorithm:

[0187] v = v + G (n) ⊙G (n)

[0188]

[0189] S (n) =S(n)-η·V -1 ·G(n)

[0190] Then proceed to step ⑤;

[0191] ⑧ Output the optimal factor matrix S (1) ,S (2) ,S (3) And use the following formula for tensors Fill in the missing values ​​in the tensor:

[0192]

[0193] Complete the evaluation of pedestrian walking parameters based on tensor decomposition and reconstruction.

[0194] (3) Complete the missing values ​​of high-dimensional, sparse pedestrian flow data.

[0195] In summary, this invention constructs a pedestrian walking speed evaluation model based on the CPWOPT* algorithm, based on the theoretical ideas of tensor decomposition and reconstruction. The model applies regularization constraints to the objective function to ensure the algorithm's learning and generalization capabilities. Considering the sparsity of tensor data, the objective function is iteratively solved using the adaptive gradient algorithm (AdaGrad). The model utilizes a genetic algorithm to determine the values ​​of hyperparameters and randomly handles missing tensor values. Comparison with classic data completion algorithms verifies the feasibility and accuracy of the proposed CPWOPT* algorithm.

[0196] This invention can expand the pedestrian walking scenarios in pedestrian traffic and pedestrian dynamics research, effectively predict missing values ​​in high-dimensional data, and provide a new and useful tool for the effective utilization of passenger flow video data with typical high-dimensional sparsity.

[0197] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.

[0198] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0199] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatus or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0200] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for completing pedestrian flow data based on tensor decomposition and reconstruction, characterized in that, include: Construct a tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization; The rank R and regularization parameter of the tensor in the tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization are analyzed using a genetic algorithm. and learning rate The three hyperparameters were tuned, and after verifying that the data completion effect of the tensor decomposition and reconstruction algorithm based on the improved CP decomposition and weighted optimization was qualified, the trained tensor decomposition and reconstruction algorithm based on the improved CP decomposition and weighted optimization was obtained. We construct high-dimensional sparse pedestrian walking speed tensors for different scenarios, and use a trained tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization to decompose and reconstruct the tensors, and fill in the missing values ​​in the sparse pedestrian speed tensors. The aforementioned construction of a tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization includes: Based on the CP decomposition weighted optimization algorithm, this paper improves the CPWOPT algorithm from the perspective of optimizing the objective function value of the weights and the gradient values ​​of the variables, using an adaptive gradient algorithm to obtain the improved CPWOPT. Tensor decomposition and reconstruction algorithm, the improved CPWOPT The processing flow of the tensor decomposition and reconstruction algorithm includes: The objective function in the CPWOPT algorithm is regularized, and a third-order pedestrian walking parameter tensor model is constructed. , Its factor matrix after decomposition is ,in, , The objective function for optimizing the weights in tensor decomposition is: (29) Add a Tikhonov regularization term to the objective function (29): (30) in, The factor matrix weight decay parameter reflects the factor matrix Throughout the objective function The importance of it; The optimized weight objective function is rewritten in the following form: (31) in, , pedestrian velocity tensor It is certain that throughout the entire iteration process, the tensor It remains unchanged and can be calculated in advance, while the regularization term... The sum of squares of the Frobenius norms of each factor matrix is ​​given. At the beginning of each iteration, the factor matrix is ​​fixed, and the regularization term is calculated directly. sparse tensors The index of a known value in the binary indicator tensor Positions where the element is 1 are added to the set in order. , It is a three-dimensional vector. , tensor The known values ​​are stored in a vector of length Q. In, that is: (32) ,for tensor Missing value location In the corresponding tensor The value in the middle must also be 0; only calculation is needed. corresponding tensor The elements in the vector are given. Tensor The known values ​​in; (33) During computer operations, let: (34) (35) vector The calculation method is as follows: (36) Then vector The calculation method is as follows: (37) One vector is computed at a time in each iteration. The optimized weight objective function (31) is rewritten as: (38) in, It is a constant. ; The partial derivatives of the optimized weight objective function with respect to each factor matrix are: (39) in, , For tensor , of n Matrix transformation, The calculation formula is as follows: (40) The learning rate of the tensor decomposition and reconstruction algorithm based on the improved CP decomposition weighted optimization is adaptively adjusted using an adaptive gradient algorithm.

2. The method according to claim 1, characterized in that, The method described above utilizes a genetic algorithm to evaluate the rank R and regularization parameter of the tensor in the tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization. and learning rate Three hyperparameters were used for parameter tuning, including: The genetic algorithm population is initialized by determining the values ​​of initialization parameters, including population size, maximum number of iterations, crossover rate, mutation rate, and global optimal fitness initialization value. Chromosomes in the population are encoded using binary encoding, with a gene length of 10. The first 4 bits of the gene represent the rank R of the tensor, and the second 3 bits represent the regularization parameter. The third part, consisting of 3 digits, represents the learning rate. ; The first generation population is randomly generated, and the fitness value of chromosomes is calculated using the fitness function. The fitness function is used to determine the fitness value of a certain chromosome to the environment. Based on the fitness function, chromosomes in the population are selected, and individuals with higher fitness values ​​are retained for crossover and mutation. The objective function expressed by formula (30) is used. Fitness is represented by the reciprocal of the value. Determine if the iteration termination condition has been met. If not, perform selection, crossover, and mutation operations on the population to obtain the next generation population, and continue to evaluate the population using the fitness function. If the termination condition has been met, output the globally optimal chromosome and fitness value. Using a single-point crossover operator, a crossover point is randomly selected, and the two selected chromosomes are split. The binary values ​​at corresponding positions after the crossover point are swapped, resulting in two distinct chromosomes. Using a basic bit mutation operator, a mutation point is randomly determined on the selected chromosome, and the gene value is replaced with other alleles of that gene to form a new individual. A genetic algorithm is used for iterative optimization to determine the optimal value in each generation of the population. When the termination condition is met, the optimal chromosome value is the rank R and the regularization parameter. and learning rate The value of is converted to binary to obtain the rank R and the regularization parameter. and learning rate The specific values ​​of the three hyperparameters.

3. The method according to claim 2, characterized in that, After verifying that the data completion effect of the tensor decomposition and reconstruction algorithm based on the improved CP decomposition weighted optimization is satisfactory, the trained tensor decomposition and reconstruction algorithm based on the improved CP decomposition weighted optimization is obtained, including: The pedestrian velocity tensor constructed using the tensor decomposition and reconstruction algorithm based on improved CP decomposition weighted optimization. The known values ​​in the data were randomly missing, with the missing rate gradually increasing from 10% to 60%. Evaluation metrics for the completion results under different methods were calculated. Mean Absolute Error (MAE) represents the average distance between the estimated value and the true value. MAE is a non-negative value; the closer it is to 0, the more accurate the algorithm's prediction. MAE is calculated as follows: (44) in, For the first i The true values ​​of pedestrian walking parameters The corresponding estimated value; The root mean square error (RMSE) measures the deviation between the estimated value and the true value. The smaller the RMSE value, the more accurate the algorithm's prediction. The RMSE is calculated as follows: (45) By setting different missing rates, the completion effects of the tensor decomposition and reconstruction algorithm based on CP decomposition and weighted optimization are compared with those of existing data completion algorithms. After verifying that the data completion effect of the tensor decomposition and reconstruction algorithm based on CP decomposition and weighted optimization is qualified, the trained tensor decomposition and reconstruction algorithm based on CP decomposition and weighted optimization is obtained.

4. The method according to any one of claims 1 to 3, characterized in that, The construction of high-dimensional sparse pedestrian walking velocity tensors under different scenarios involves decomposing and reconstructing the tensors using a trained tensor decomposition and weighted optimization algorithm based on CP decomposition, and filling in missing values ​​in the sparse pedestrian velocity tensors, including: ① Constructing rank in different scenarios Sparse pedestrian walking speed tensor and the same-dimensional binary indicator tensor And calculate constants ; ② Tensor The known value index is placed in And store the known values ​​in a vector. middle; ③ Use a genetic algorithm to determine the hyperparameter regularization parameter. rank of tensors Learning rate The value of is determined using the CP-ALS algorithm on the tensor. Decompose to obtain the initial factor matrix. ; ④ Construct a regularized objective function ; ⑤ Order and construct vectors , ; ⑥ Calculate the objective function If the objective function value reaches the loss threshold or the maximum number of iterations is reached... If yes, proceed to step ⑧; otherwise, proceed to step ⑦. ⑦ Calculate the partial derivatives of each factor matrix. And update the gradient using the AdaGrad algorithm: in, For the front j Step gradient element-wise product matrix, For the front j Partial derivatives of each factor matrix at each step, Indicates taking the first j The diagonal matrix of the step gradient element-wise product matrix; It is a unit symmetric matrix; The smoothing coefficient is usually taken as... Learning rate These are fixed hyperparameters; For the front j The step factor matrix, the AdaGrad algorithm uses an adaptive matrix Replaces the fixed learning rate in the SGD algorithm. The diagonal elements of the matrix are the update weights for each dimension parameter, so that the learning rate includes the historical accumulated squared gradient information, which reflects the differences of each dimension when updating the step size; and then jump to step ⑤; ⑧ Output the optimal factor matrix And use the following formula for tensors Fill in the missing values ​​in the tensor: Complete the evaluation of pedestrian walking parameters based on tensor decomposition and reconstruction.

Citation Information

Patent Citations

  • Data collection method for adaptive crowd sensing system based on tensor filling

    CN108830930A

  • Tensor structural missing filling method based on joint low rank and sparse representation

    CN109241491A