Pedestrian incomplete track filling method
By combining the convolutional network and self-attention mechanism in the trajectory filling method, the time domain and spatial characteristics of pedestrian trajectory data are solved, and the accuracy and robustness of pedestrian incomplete trajectory filling in complex environments is achieved, and the trajectory filling effect with high accuracy and high robustness is achieved.
Patent Information
- Application Number
- CN202411861352.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-23
AI Technical Summary
The existing trajectory filling methods are difficult to effectively process the time domain and spatial characteristics in pedestrian trajectory data, resulting in insufficient accuracy and robustness of pedestrian incomplete trajectory filling in complex environments with more occlusions.
The time domain input embedding module based on the convolutional network and the trajectory filling module of the self-attention mechanism are used to combine the time domain and spatial domain characteristics to identify the missing information location through the neural network input trajectory mask and perform data completion.
It realizes high accuracy and high robustness of pedestrian incomplete trajectory filling in complex environments with more occlusions, and is suitable for complex scenes of time domain and spatial characteristics of pedestrian trajectory data.
Smart Images

Figure CN120032089A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of trajectory filling, and in particular to a method for filling an incomplete trajectory of a pedestrian. Background Art
[0002] After years of development, trajectory filling methods have made great progress and have been tested and used in many fields. As an important part of data processing, trajectory filling can enhance the availability of trajectory data, improve the resolution of trajectory data, and provide strong data support for subsequent steps such as trajectory prediction, intention recognition, and situation awareness.
[0003] At present, a lot of scientific research and applications have been carried out in the field of trajectory filling, and fruitful results have been achieved. Du W et al. (Du W, D, Liu Y. Saits: Self-attention-based imputation for time series [J]. Expert Systems with Applications, 2023, 219: 119619.) designed a non-autoregressive trajectory filling model based on the self-attention mechanism, and used a large number of quantitative and qualitative experiments to show that SAITS effectively outperforms other methods in the time series filling task. Xue Yapu et al. proposed a method for filling missing vehicle trajectories based on the Floyd pathfinding algorithm in the literature (Xue Yapu. Research on vehicle trajectory reconstruction method based on discrete monitoring data [D]. Beijing University of Technology, 2020.), which realized the filling of missing trajectories during vehicle driving and improved the integrity and usable value of discrete data. The patent (Xiao Zhu, Zeng Fanzi, Sun Wenyuan, et al. A vehicle trajectory filling method based on intelligent identification of road environment [P]. Hunan Province: CN202010094936.0, 2022-02-25.) proposes a vehicle trajectory filling method based on intelligent identification of road environment, which divides vehicle trajectories into three categories: "straight line", "right-angle turn" and "ramp", and uses the GRU network to fill the vehicle trajectory.
[0004] The above trajectory filling methods consider scenarios such as statistical data such as reports, or vehicle trajectory data filling in road traffic, but do not consider pedestrian trajectory data filling, and rarely consider both the temporal and spatial characteristics of trajectory data. Pedestrian trajectory data usually has stronger nonlinearity and randomness. Conventional trajectory data filling methods will have problems such as low accuracy and low robustness when applied to pedestrian trajectory filling. Summary of the invention
[0005] Technical problem: The purpose of the present invention is to provide a method for filling in the incomplete trajectory of pedestrians. The method has high filling accuracy and strong robustness, and is suitable for filling in the incomplete trajectory of pedestrians in complex environments with many obstructions.
[0006] Technical solution: The purpose of the invention is to provide a method for filling incomplete pedestrian trajectories, which is suitable for filling incomplete pedestrian trajectories in complex environments with many obstructions.
[0007] The bird's-eye view camera of the block acquires pedestrian videos in a certain area, and extracts the trajectory data of pedestrians within a certain period of time from the video. The image acquisition frequency of the camera is not less than 25 frames per second, the sampling period of the extracted trajectory data is not less than 0.2 seconds, and the sampling frame number is not less than 8 frames.
[0008] Considering that the input of the neural network in this method is incomplete trajectory information, and trajectory filling requires targeted data completion of the missing information position, the neural network inputs a trajectory mask while inputting the incomplete pedestrian trajectory to represent the position of the missing information in the trajectory. Given an input sequence X with a length of k and a dimension of 2, that is:
[0009]
[0010] where the element x 1t and x 2t Respectively represent the horizontal and vertical coordinates of the position in the input trajectory at time t. The neural network requires a trajectory mask matrix Φ of size k×2 to indicate the location of missing information in the input trajectory X, that is:
[0011]
[0012] The trajectory mask matrix Φ is a zero-one matrix, whose elements represent whether the elements at the corresponding positions in the input trajectory X are missing values, that is:
[0013]
[0014] Where 1≤i≤k, 1≤j≤2. When φ ij =1, the corresponding position in the input sequence X is the observed value; when φ ij =1, the corresponding position information in the input sequence X is missing, that is, NaN.
[0015] The neural network used in the incomplete pedestrian trajectory filling method consists of two parts. The first part is a time domain input embedding module based on a convolutional network, which extracts time domain features and embeds time domain and spatial domain features into the input sequence; the second part is a trajectory filling module based on a self-attention mechanism, which is used to extract spatial and time domain features from the input sequence and fill the sequence.
[0016] The temporal input embedding module is divided into two parts: the temporal convolution module and the position embedding. The temporal convolution module perceives the temporal feature information in the trajectory sequence, and the position embedding provides additional position feature information for the embedding vector. The temporal convolution module uses a one-dimensional fully convolutional network with dilated causal convolution to extract temporal features. Causal convolution enables each convolution layer to obtain only the information of the current time step and the past time step in the previous convolution layer, thereby eliminating the leakage of future information when processing trajectory sequences; dilated convolution is used to expand the receptive field of the convolution operation, so that the multi-layer convolution network can extract feature information from a larger time range while reducing the amount of network computation. Each convolution layer in the temporal convolution module is composed of a dilated causal convolution layer, batch normalization, and ReLU activation function in series, and a residual structure is used to alleviate the negative impact caused by the common gradient vanishing problem in convolutional networks. After multiple layers of dilated causal convolution, the temporal convolution module adjusts the output dimension of the network through a fully connected network. Position embedding adds the relative and absolute distances between different positions in the input trajectory sequence to the embedding vector, thereby providing additional spatial domain features for the embedding vector. Each element in the embedding vector output by the time domain input embedding module is a time domain segment with a small variable receptive field containing position information, that is, each element contains spatial domain information and time domain feature information of a certain length.
[0017] Based on the self-attention mechanism, the present invention designs a trajectory filling module to extract the temporal and spatial domain features in the trajectory sequence and fill in the missing information in the trajectory sequence. The self-attention mechanism is implemented through three linear layers W Q , W K and W V The input embedding vector E is converted into query, key, and value vectors, namely Q, K, and V vectors. Then the Q, K, and V vectors are multiplied and normalized to obtain the weight matrix Z. The relationship between different positions in the input embedding vector is converted into mathematical expectations through weights. This process is also called proportional dot product attention. Q, K, V, and Z are usually operated in multiple heads. Different heads focus on the feature information of different dimensions of the embedding vector and calculate the weight matrix in parallel. Finally, the weight matrix Z is adjusted through an output fully connected layer to obtain the output vector O. The above process can be expressed by the formula O=SA(E), which can be specifically expressed as:
[0018]
[0019] Where d is the dimension of the embedding matrix.
[0020] The trajectory filling module in the present invention uses an encoder structure to obtain information from the past and future of the missing information position to fill the trajectory. Each layer in the encoder consists of a self-attention sublayer and a feedforward network sublayer. Both sublayers have serial batch normalization and ReLU activation functions, and the problem of gradient disappearance is alleviated by the residual structure. At the end of the encoder is a fully connected layer network, which converts the output of the trajectory filling network into a trajectory sequence that meets the requirements according to the output mode of the trajectory filling method.
[0021] Specifically:
[0022] a) Design a time domain input embedding module, and design a trajectory filling network structure based on the time domain input embedding module;
[0023] b) Collect pedestrian trajectory data and pre-process the data;
[0024] c) The preprocessed pedestrian trajectory is passed through the trajectory filling network to obtain the completed pedestrian trajectory, and then the loss function of the trajectory filling network is designed;
[0025] d) Repeat step c according to the back propagation algorithm to update the trajectory filling network until the network accuracy reaches the expected requirement.
[0026] This method has high filling accuracy and strong robustness, and is suitable for filling incomplete pedestrian trajectories in complex environments with many obstructions.
[0027] Wherein said step a specifically comprises:
[0028] a1) Based on the convolutional network, design the time domain input embedding module, and design the number of network layers, the number of neurons in each layer, the activation function and the position embedding;
[0029] a2) Based on the time domain input embedding module and self-attention mechanism, design the network structure, number of network layers, number of neurons in each layer, activation function and output format of the pedestrian trajectory filling network;
[0030] Wherein said step b specifically comprises:
[0031] b1) Obtain pedestrian trajectory data from the bird’s-eye view camera;
[0032] b2) Use a fixed-step sampling window to normalize the data, and give a mask matrix of the trajectory data based on the situation where the pedestrian trajectory is blocked.
[0033] Wherein said step c specifically comprises:
[0034] c1) Input the incomplete pedestrian history trajectory into the trajectory filling network to obtain the completed pedestrian trajectory;
[0035] c2) Design a loss function for the trajectory filling network according to the different output formats of the trajectory filling network.
[0036] Wherein said step d specifically comprises:
[0037] d1) Update the weights of the trajectory filling network according to the back propagation algorithm until the expected requirements are met.
[0038] The beneficial effects of the present invention are as follows: the method has high filling accuracy and strong robustness, and is suitable for filling incomplete trajectories of pedestrians in complex environments with many obstructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of the overall framework of the trajectory filling method;
[0040] Figure 2 This is a schematic diagram of the time domain input embedding module;
[0041] Figure 3 Schematic diagram of dilated convolution and causal convolution;
[0042] Figure 4 Fill the module structure diagram for the trajectory;
[0043] Figure 5 This is a schematic diagram of camera coordinate transformation;
[0044] Figure 6 This is a schematic diagram of the neural network learning rate curve;
[0045] Figure 7 Flowchart of the method for filling pedestrian trajectories. DETAILED DESCRIPTION
[0046] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0047] like Figure 1 and 7 As shown, the present invention provides a method for filling an incomplete trajectory of a pedestrian, and the specific steps are composed of P1 to P4, and each step is described as follows:
[0048] 1) Step P1
[0049] Step P1 designs a time domain input embedding module, and designs a trajectory filling network structure based on the time domain input embedding module. The specific steps are as follows:
[0050] Step 1: Design the time domain input embedding module. The time domain input embedding module is divided into the time domain convolution module ( Figure 3) and position embedding, input a k×2 trajectory sequence X and a mask matrix, and output a padded sequence of the same size as X The trajectory sequence length of the present invention is set to 8, that is, k = 8. Position embedding vector P = [p (pos,i) ] can be calculated by the following formula:
[0051]
[0052] where p (pos,i) is the element of the i-th dimension at the position pos in the position embedding vector P; d is the dimension of the embedding vector, and in the present invention, d=1024. The structure of the time domain input embedding module is as follows Figure 2 The parameters are shown in Table 1.
[0053] Table 1 Time domain input embedding module structure
[0054]
[0055] Step 2: Based on the time domain input embedding module and self-attention mechanism, design the network structure of the pedestrian trajectory filling network.
[0056] The specific steps are as follows:
[0057] The pedestrian trajectory filling network of the present invention includes two parts: a time domain input embedding module and a trajectory filling module based on a self-attention mechanism. The time domain input embedding module extracts the time domain features in the sequence and adds additional spatial domain features. The design of the time domain input embedding module is shown in the first step. Figure 4 As shown in Figure 2, the trajectory filling module is an 8-layer encoder structure, which is used to extract spatial and temporal feature information in the trajectory and fill the incomplete input trajectory. The parameters of the encoder layer and the output fully connected layer after the encoder are shown in Table 2, where the parameters of all encoder layers are the same and the self-attention mechanism sublayer has 8 heads.
[0058] Table 2 Trajectory filling module structure
[0059]
[0060]
[0061] The trajectory filling method of the present invention has two output modes, namely generative output and filling output. The generative output generates a new complete trajectory based on the input incomplete trajectory; the filling output generates the value of the missing position of the incomplete trajectory and fills it into the incomplete trajectory to form a complete pedestrian trajectory.
[0062] 2) Step P2
[0063] Step P2 collects pedestrian trajectory data and preprocesses the data. The specific steps are as follows:
[0064] Step 1: Obtain pedestrian trajectory data from the bird's-eye view camera. The camera's image acquisition speed is no less than 25 frames per second. The valid pedestrian data in the extracted video data requires that the pedestrian is in the camera's field of view for at least 4 seconds, and the number of valid pedestrians is no less than 20.
[0065] With the center of the video as the coordinate origin, the video width as the x-axis, and the video height as the y-axis, the trajectory data of pedestrians appearing in the video are identified and annotated. Figure 5 As shown, due to the perspective phenomenon of the bird's-eye view camera, the present invention uses the homography matrix H to transform the trajectory coordinates X marked in the video v Convert to real world coordinate X, that is:
[0066] X=HX
[0067] X=HX v
[0068]
[0069] Step 2: Use a fixed-step sampling window to normalize the data, and give a mask matrix of the trajectory data according to the situation where the pedestrian trajectory is blocked. Use a fixed-step sampling window to sample the pedestrian trajectory data preprocessed in the first step, with a sampling window length of 8 and a sampling interval of 0.4 seconds. If there are missing values in the sampled pedestrian trajectory data, the corresponding element in the mask matrix is set to 0, indicating that this position in the sampled trajectory sequence is a missing value, otherwise it is set to 1. Finally, a set of incomplete trajectory data with a step size of 8 and its mask matrix are obtained, and the incomplete trajectory data is batched with a batch size of 256.
[0070] 3) Step P3
[0071] The preprocessed pedestrian trajectory is passed through the trajectory filling network to obtain the completed pedestrian trajectory, and then the loss function of the trajectory filling network is designed. The specific contents are as follows:
[0072] Step 1: Input the incomplete pedestrian history trajectory into the trajectory filling network to obtain the complete pedestrian trajectory. The design is implemented in the following steps:
[0073] Multiply the trajectory sequence data X in step P1 and the mask matrix Φ bit by bit to obtain the trajectory sequence X that masks the missing values. m , all missing values in X are covered with 0. The process is as follows:
[0074] X m =X⊙Φ
[0075] X mInput to the time domain input embedding module to convert the input trajectory sequence into an embedding vector. The process is as follows:
[0076] E=TIE(X m ,P)=TCB(X m )+P=X'+P
[0077] Where TIE is the time domain input embedding module, TCB is the time domain convolution module, X' is the output of the TCB module, and P is the position embedding. Then the embedding vector E is input into the trajectory filling module, and the process is as follows:
[0078]
[0079] Among them, W imp is the output fully connected layer, is the completed trajectory sequence, It is an 8-layer encoder structure;
[0080] The encoder layer process is expressed as follows:
[0081]
[0082] Among them, SA is the self-attention mechanism, FFN is the feedforward network, That is Eight layers of superposition.
[0083] Step 2: Design a loss function for the trajectory filling network based on the different output formats of the trajectory filling network. The details are as follows:
[0084] The present invention uses Euclidean distance as a loss function for the neural network design of the trajectory filling method. However, since the trajectory filling method of the present invention has two comfortable modes, it is necessary to design loss functions for the two output modes respectively.
[0085] The loss function L of the padding output If for:
[0086]
[0087] where ω If =0.9 is the weight parameter, which is used to adjust the convergence speed of the neural network; n m is the total number of missing information in the input trajectory sequence; k = 8 is the length of the input trajectory sequence; and is the horizontal and vertical coordinates of the complete sequence at time t; x t and t is the horizontal and vertical coordinates at time t in the input trajectory; the loss function of the filled output only focuses on the output error of the missing information position;
[0088] The loss function L of the generative output Ig for:
[0089]
[0090] where ω Im =0.8 is the weight parameter of the missing information position in the trajectory sequence, ω Is = 0.5 is the weight parameter of other positions containing complete information; I is the unit matrix of the same size as the mask matrix Φ; the loss function L Ig At the same time, we pay attention to the output errors at the missing information location and other locations, and usually ω Im >ω Is , L Ig More attention will be paid to the output errors at the locations where information is missing.
[0091] 4) Step P4
[0092] Step P4 is based on the back propagation algorithm and uses the loss function obtained in P3 to update the weights in the neural network of the trajectory filling method until the desired requirements are met.
[0093] Step 1: Use the Adaptive Moment Estimation (Adam) optimizer to calculate the gradient of the neural network, and use the back-propagation algorithm to simultaneously update the weights of the network in the time domain input embedding module and the trajectory filling module in the trajectory filling method. Repeat step P3 of the above steps for the neural network after updating the weights to the maximum number of rounds, and finally obtain a result that converges approximately to the optimal value.
[0094] The initial learning rate of the network is 0.1, and the learning rate is increased linearly in the first 10% of the training rounds, and the learning rate is decayed according to the inverse square root of the number of rounds in the last 90% of the training rounds. The maximum number of training rounds of the present invention is set to 100, that is, the learning rate increases linearly in the first 10 rounds, and the learning rate decays by the inverse square root of the last 90 rounds, such as Figure 6 shown.
Claims
1. A method for filling in an incomplete trajectory of a pedestrian, characterized in that: It is used to fill in the incomplete trajectory of pedestrians in complex environments with many obstructions, including the following steps: a) Design a time domain input embedding module, and design a trajectory filling network structure based on the time domain input embedding module; b) Collect pedestrian trajectory data and pre-process the data; c) The preprocessed pedestrian trajectory is passed through the trajectory filling network to obtain the completed pedestrian trajectory, and then the loss function of the trajectory filling network is designed; d) Repeat step c according to the back propagation algorithm to update the trajectory filling network until the network accuracy reaches the expected requirement.
2. A pedestrian incomplete trajectory filling method according to claim 1, characterized in that: The step a) specifically comprises: a1) Based on the convolutional network, design the time domain input embedding module, and design the number of network layers, the number of neurons in each layer, the activation function and the position embedding; a2) Based on the trajectory filling module of the time domain input embedding module and the self-attention mechanism, the network structure, number of network layers, number of neurons in each layer of the network, activation function and output format of the pedestrian trajectory filling network are designed.
3. A pedestrian incomplete trajectory filling method according to claim 2, characterized in that: The time domain input embedding module includes a time domain convolution module and a position embedding module, the time domain convolution module perceives the time domain feature information in the trajectory sequence, and the position embedding provides additional position feature information for the embedding vector; By inputting a k×2 trajectory sequence X and a mask matrix Φ, the padded trajectory sequence with the same size as X is output. Among them, the element x 1t and x 2t They represent the horizontal and vertical coordinates of the position in the input trajectory at time t respectively.
4. A pedestrian incomplete trajectory filling method according to claim 3, characterized in that: The mask matrix Φ is a zero-one matrix, whose elements represent whether the elements at the corresponding positions in the input trajectory sequence X are missing values; Among them, 1≤i≤k, 1≤j≤2; when φ ij = 1, the corresponding position in the input trajectory sequence X is the observed value; when φ ij When =0, the corresponding position information in the input trajectory sequence X is missing, that is, NaN.
5. A method for filling in an incomplete pedestrian trajectory according to claim 4, characterized in that: The step b) specifically comprises: b1) Obtain pedestrian trajectory data from the bird’s-eye view camera; b2) Use a fixed-step sampling window to normalize the data, and give a mask matrix of the trajectory data based on the situation where the pedestrian trajectory is blocked.
6. A pedestrian incomplete trajectory filling method according to claim 5, characterized in that: The step c) specifically comprises: c1) Input the incomplete pedestrian history trajectory into the trajectory filling network to obtain the completed pedestrian trajectory; c2) Design a loss function for the trajectory filling network according to the different output formats of the trajectory filling network.
7. A method for filling in an incomplete pedestrian trajectory according to claim 6, characterized in that: In step c1), the input trajectory sequence X and the mask matrix Φ are multiplied bit by bit to obtain the trajectory sequence X that masks the missing values. m , all missing values in X are covered with 0, the formula is as follows: X m =X⊙Φ X m Input into the time domain input embedding module to transform the input trajectory sequence X into an embedding vector E. The formula is as follows: E=TIE(X m ,P)=TCB(X m )+P=X'+P Among them, TIE is the time domain input embedding module, TCB is the time domain convolution module, X' is the output of the TCB module, and P is the position embedding; then the embedding vector E is input into the trajectory filling module, and the formula is as follows: Among them, W imp is the output fully connected layer, is the filled trajectory sequence, It is an 8-layer encoder structure; The encoder layer process is expressed as follows: Among them, SA is the self-attention mechanism, FFN is the feedforward network, That is Eight layers of superposition; In step c2), the output formats are generative output and filling output. The generative output generates a new complete trajectory based on the input incomplete trajectory; the filling output generates the value of the missing position in the incomplete trajectory and fills it into the incomplete trajectory to form a complete pedestrian trajectory.
8. A method for filling in an incomplete pedestrian trajectory according to claim 7, characterized in that: The loss function L of the padding output If for: where ω If =0.9 is the weight parameter, which is used to adjust the convergence speed of the neural network; n m is the total number of missing information in the input trajectory sequence; k = 8 is the length of the input trajectory sequence; and is the horizontal and vertical coordinates of the complete sequence at time t; x t and t is the horizontal and vertical coordinates at time t in the input trajectory; the loss function of the filled output only focuses on the output error of the missing information position; The loss function L of the generative output Ig for: where ω Im =0.8 is the weight parameter of the missing information position in the trajectory sequence, ω Is = 0.5 is the weight parameter of other positions containing complete information; I is the unit matrix of the same size as the mask matrix Φ; the loss function L Ig At the same time, we pay attention to the output errors at the missing information location and other locations, and usually ω Im >ω Is , L Ig More attention will be paid to the output errors at the locations where information is missing.
9. A pedestrian incomplete trajectory filling method according to claim 8, characterized in that: The step d) specifically comprises: d1) Update the weights of the trajectory filling network according to the back propagation algorithm until the expected requirements are met.