A pedestrian group detection method and system based on pedestrian trajectory retrieval

By constructing a pedestrian space-time trajectory model and generating pedestrian cross-camera trajectory spatiotemporal information, the problem of difficulty in constructing a cross-camera space-time model in the existing technology is solved, and higher quality pedestrian video retrieval and group detection results are achieved.

CN114694093BActive Publication Date: 2025-06-27SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210256807.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-16
Publication Date
2025-06-27
Estimated Expiration
2042-03-16

AI Technical Summary

Technical Problem

In the construction of cross-camera space-time models, the problems of large amount of data calculation, the inability to calculate the space-time probability between camera pairs, and the lack of effective probability models to describe the space-time probability distribution.

Method used

By constructing a pedestrian space-time trajectory model, the spatio-time and peer information relationships between nodes and nodes are used to generate pedestrian cross-camera trajectory space-time information, and combining pedestrian trajectory reordering methods and joint distance measurement methods, the quality of pedestrian video retrieval results and pedestrian group detection results are improved.

Benefits of technology

The problem of node duplication in pedestrian trajectory was solved, better pedestrian trajectory information was obtained, and the quality of pedestrian video retrieval results and pedestrian group detection results were improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694093B_ABST
    Figure CN114694093B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian group detection method and system based on pedestrian trajectory retrieval. The method includes: obtaining an annotated data set and constructing a training data set; training an appearance feature model and a spatio-temporal model based on the training data set, and integrating them to obtain a pedestrian spatio-temporal trajectory model; generating pedestrian cross-camera trajectory spatio-temporal information based on the pedestrian spatio-temporal trajectory model; detecting a to-be-detected pedestrian image based on the pedestrian cross-camera trajectory spatio-temporal information, and outputting a pedestrian group detection result. The system includes: an obtaining module, a training module, a generating module, and an output module. By constructing a pedestrian spatio-temporal trajectory model, the present invention can improve the quality of pedestrian video retrieval results and the quality of pedestrian group detection results. As a pedestrian group detection method and system based on pedestrian trajectory retrieval, the present invention can be widely applied to the field of computer vision technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a pedestrian group detection method and system based on pedestrian trajectory retrieval. Background Art

[0002] The information used in previous surveillance camera applications is all image information, while ignoring the spatial topological information of the camera itself and the spatio-temporal information existing in the process of image acquisition. Existing cross-camera spatio-temporal models are all based on statistical methods or methods based on prior probability models. The defect of the statistical method is that a large amount of data is required to calculate the spatio-temporal probability between camera pairs. If there is no data between a certain camera pair, then the spatio-temporal probability of this camera pair cannot be calculated. The defect of the method based on the prior probability model is that a lot of data is required to estimate the parameters of the prior probability, and at the same time, there is no good probability model to describe the spatio-temporal probability distribution of different camera pairs. Summary of the Invention

[0003] In order to solve the above technical problems, the purpose of the present invention is to provide a pedestrian group detection method and system based on pedestrian trajectory retrieval, which can improve the quality of pedestrian video retrieval results and the quality of pedestrian group detection results.

[0004] The first technical solution adopted by the present invention is: A pedestrian group detection method based on pedestrian trajectory retrieval, including the following steps:

[0005] Obtain a labeled data set and construct a training data set;

[0006] Train the appearance feature model and the spatio-temporal model based on the training data set, and integrate them to obtain a pedestrian spatio-temporal trajectory model;

[0007] Generate pedestrian cross-camera trajectory spatio-temporal information based on the pedestrian spatio-temporal trajectory model;

[0008] Detect the pedestrian image to be measured based on the pedestrian cross-camera trajectory spatio-temporal information, and output the pedestrian group detection result.

[0009] Further, the step of obtaining a labeled data set and constructing a training data set specifically includes:

[0010] Obtain a data set through a network camera, perform annotation processing on the data set to obtain a labeled data set;

[0011] Construct a training data set based on the labeled data set, and divide the training data set into labeled single-camera pedestrian image information and labeled single-camera pedestrian trajectory information.

[0012] Further, the step of training the appearance feature model and the spatio-temporal model based on the training dataset and integrating them to obtain the pedestrian spatio-temporal trajectory model specifically includes:

[0013] Training the appearance feature model based on the annotated single-camera pedestrian image information and training the spatio-temporal model based on the annotated single-camera pedestrian trajectory information to obtain an optimized appearance feature model and an optimized spatio-temporal model;

[0014] Integrating the optimized appearance feature model and the optimized spatio-temporal model to obtain the pedestrian spatio-temporal trajectory model.

[0015] Further, the step of training the appearance feature model based on the annotated single-camera pedestrian image information and training the spatio-temporal model based on the annotated single-camera pedestrian trajectory information to obtain an optimized appearance feature model and an optimized spatio-temporal model specifically includes:

[0016] Based on the appearance feature model, performing feature extraction processing on the annotated single-camera pedestrian image information to obtain the annotated single-camera pedestrian image feature information, and the appearance feature model is composed of a convolutional neural network;

[0017] Calculating the annotated single-camera pedestrian image feature information respectively through the triplet loss function and the gradient descent algorithm to obtain the annotated single-camera pedestrian image optimization information;

[0018] Optimizing the appearance feature model based on the annotated single-camera pedestrian image optimization information to obtain an optimized appearance feature model;

[0019] Based on the spatio-temporal model, calculating the annotated single-camera pedestrian trajectory information respectively through the cross-entropy loss function and the gradient descent algorithm to obtain the annotated single-camera pedestrian trajectory optimization information;

[0020] Optimizing the spatio-temporal model based on the annotated single-camera pedestrian trajectory optimization information to obtain an optimized spatio-temporal model.

[0021] Further, the step of generating the pedestrian cross-camera trajectory spatio-temporal information based on the pedestrian spatio-temporal trajectory model specifically includes:

[0022] Based on the pedestrian spatio-temporal trajectory model, clustering the single-camera pedestrian trajectory spatio-temporal information to obtain a pedestrian trajectory association dataset and constructing a trajectory graph model;

[0023] The single-camera pedestrian trajectory spatio-temporal information includes the annotated single-camera pedestrian image optimization information and the annotated single-camera pedestrian trajectory optimization information;

[0024] The trajectory graph model is associated and calculated through an association algorithm to obtain the spatio-temporal information of pedestrian cross-camera trajectories.

[0025] Furthermore, the step of clustering the spatio-temporal information of single-camera pedestrian trajectories based on the pedestrian spatio-temporal trajectory model to obtain a pedestrian trajectory association data set and constructing a trajectory graph model specifically includes:

[0026] Based on the pedestrian spatio-temporal trajectory model, the spatio-temporal information of single-camera pedestrian trajectories is extracted to obtain the spatio-temporal extraction information of single-camera pedestrian trajectories;

[0027] The spatio-temporal extraction information of single-camera pedestrian trajectories is clustered and calculated through a hierarchical clustering algorithm to obtain the number of categories of the spatio-temporal extraction information of single-camera pedestrian trajectories;

[0028] The number of categories of the spatio-temporal extraction information of single-camera pedestrian trajectories is spliced through a companion clustering algorithm to obtain a pedestrian trajectory association data set;

[0029] A trajectory graph model is established according to the pedestrian trajectory association data set;

[0030] The trajectory graph model is composed of an adjacency matrix, and the adjacency matrix includes a directed edge adjacency matrix and an undirected edge adjacency matrix.

[0031] Furthermore, the step of associating and calculating the trajectory graph model through an association algorithm to obtain the spatio-temporal information of pedestrian cross-camera trajectories specifically includes:

[0032] The adjacency matrix in the trajectory graph model is sequentially subjected to threshold truncation processing and search processing to obtain pedestrian cross-camera trajectories;

[0033] The cosine value of the directed edge adjacency matrix in the trajectory graph model is calculated to obtain a trajectory energy matrix;

[0034] The undirected edge adjacency matrices in the trajectory graph model are multiplied by companions to obtain a companion energy matrix;

[0035] The trajectory energy matrix and the companion energy matrix are integrated to obtain an energy matrix;

[0036] The energy matrix is iteratively updated through an average field algorithm to obtain an updated energy matrix;

[0037] Based on the pedestrian cross-camera trajectories, the updated energy matrix is decomposed and calculated through a non-negative matrix factorization method to obtain the spatio-temporal information of pedestrian cross-camera trajectories.

[0038] Furthermore, the step of detecting a to-be-detected pedestrian image based on the spatio-temporal information of pedestrian cross-camera trajectories and outputting a crowd detection result specifically includes:

[0039] Train the image of the pedestrian to be measured based on the apparent feature model to obtain the trajectory information of the pedestrian to be measured;

[0040] Based on the pedestrian trajectory reordering method, combine the spatio-temporal information of the pedestrian's cross-camera trajectory to retrieve the trajectory information of the pedestrian to be measured, and obtain the pedestrian video retrieval result;

[0041] Based on the pedestrian group detection framework, combine the spatio-temporal information of the pedestrian's cross-camera trajectory to retrieve the trajectory information of the pedestrian to be measured, and output the pedestrian group detection result.

[0042] Further, the step of retrieving the trajectory information of the pedestrian to be measured based on the pedestrian group detection framework, combining the spatio-temporal information of the pedestrian's cross-camera trajectory, and outputting the pedestrian group detection result specifically includes:

[0043] Retrieve the trajectory information of the pedestrian to be measured to obtain a set of cross-camera trajectories of a single pedestrian;

[0044] Calculate the set of cross-camera trajectories of a single pedestrian through the pedestrian correlation matrix to obtain the cross-camera trajectory distance of a single pedestrian;

[0045] Calculate the correlation distances between multiple cross-camera trajectories of a single pedestrian through a metric algorithm to obtain the correlation distances between pedestrians;

[0046] Filter the correlation distances between pedestrians to obtain the pedestrian group detection result.

[0047] The second technical solution adopted by the present invention is: a pedestrian group detection system based on pedestrian trajectory retrieval, including:

[0048] An acquisition module for acquiring an annotated data set and constructing a training data set;

[0049] A training module for training an apparent feature model and a spatio-temporal model based on the training data set, and integrating them to obtain a pedestrian spatio-temporal trajectory model;

[0050] A generation module for generating spatio-temporal information of a pedestrian's cross-camera trajectory based on the pedestrian spatio-temporal trajectory model;

[0051] An output module for detecting the image of the pedestrian to be measured based on the spatio-temporal information of the pedestrian's cross-camera trajectory, and outputting the pedestrian group detection result.

[0052] The beneficial effects of the method and system of the present invention are: by constructing a pedestrian spatio-temporal trajectory model, on the one hand, the problem of repeated nodes in the pedestrian trajectory is solved, and on the other hand, better pedestrian trajectory information is obtained by using the spatio-temporal relationship and companion information relationship between nodes and nodes, and the quality of the pedestrian video retrieval result and the quality of the pedestrian group detection result are improved through the pedestrian trajectory reordering method and the joint distance metric method. Description of the Drawings

[0053] Figure 1 is a flowchart of the steps of a method for detecting a crowd of pedestrians based on pedestrian trajectory retrieval according to the present invention;

[0054] Figure 2 is a block diagram of the structure of a system for detecting a crowd of pedestrians based on pedestrian trajectory retrieval according to the present invention;

[0055] Figure 3 is a specific schematic diagram of a non - negative matrix factorization node according to the present invention;

[0056] Figure 4 is a schematic diagram of the principle of a companion - associated camera according to the present invention;

[0057] Figure 5 is a retrieval diagram of a surveillance camera for an application of a simulation experiment of the present invention. Detailed Description of the Invention

[0058] The present invention will be further described in detail below with reference to the drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0059] Referring to Figure 1 , the present invention provides a method for detecting a crowd of pedestrians based on pedestrian trajectory retrieval, and the method includes the following steps:

[0060] S1. Obtain an annotated data set and construct a training data set;

[0061] S11. Obtain a data set through a network camera, perform annotation processing on the data set to obtain an annotated data set;

[0062] S12. Construct a training data set based on the annotated data set;

[0063] S13. The training data set includes annotated single - camera pedestrian image information and annotated single - camera pedestrian trajectory information.

[0064] S2. Train an appearance feature model and a spatio - temporal model based on the training data set, and integrate them to obtain a pedestrian spatio - temporal trajectory model;

[0065] S21. Train the appearance feature model based on the annotated single - camera pedestrian image information to obtain an optimized appearance feature model;

[0066] S211. Based on the appearance feature model, perform feature extraction processing on the annotated single-camera pedestrian image information to obtain the annotated single-camera pedestrian image feature information, where the appearance feature model is composed of a convolutional neural network;

[0067] S212. Calculate the loss of the annotated single-camera pedestrian image feature information through the triplet loss function to obtain the annotated single-camera pedestrian image loss information;

[0068] S213. Perform optimization calculation on the annotated single-camera pedestrian image loss information through the gradient descent algorithm to obtain the annotated single-camera pedestrian image optimization information;

[0069] Specifically, use the Resnet50 convolutional neural network as the appearance feature extraction model. The input annotated single-camera pedestrian images for training are color images with a resolution of 256×128. In each iteration, first sample 4 pedestrians, with 8 images for each pedestrian, a total of 32 images. Then send these annotated single-camera pedestrian images into the feature extraction model for feature extraction to obtain the annotated single-camera pedestrian image feature information. Calculate the loss of the annotated single-camera pedestrian image feature information through the triplet loss function to obtain the annotated single-camera pedestrian image loss information. Finally, use the gradient descent algorithm to perform optimization calculation on the annotated single-camera pedestrian image loss information to obtain the annotated single-camera pedestrian image optimization information.

[0070] S214. Optimize the appearance feature model based on the annotated single-camera pedestrian image optimization information to obtain the optimized appearance feature model.

[0071] S22. Train the spatio-temporal model based on the annotated single-camera pedestrian trajectory information to obtain the optimized spatio-temporal model;

[0072] S221. Based on the spatio-temporal model, calculate the loss of the annotated single-camera pedestrian trajectory information through the cross-entropy loss function to obtain the annotated single-camera pedestrian trajectory loss information;

[0073] S222. Perform optimization calculation on the annotated single-camera pedestrian trajectory loss information through the gradient descent algorithm to obtain the annotated single-camera pedestrian trajectory optimization information;

[0074] Specifically, the spatio-temporal model includes an input layer, a hidden layer, and an output layer. When optimizing the spatio-temporal network model of the camera, this patent uses a multi-layer perceptron to obtain the spatio-temporal model. A multi-layer perceptron is also a type of neural network, which consists of three layers. The first layer is the input layer, which includes two inputs, namely the distance d between camera pairs and the time difference t passing through this camera pair. The second layer is the hidden layer, which consists of 100 nodes. The third layer is the output layer, and the value it outputs is between 0 and 1, which represents the probability that a person passes through them at time t under the condition of the distance d between the camera pairs. During optimization, this method only considers collecting positive and negative samples from the samples of camera pairs with sufficient positive and negative samples, while ignoring the data of camera pairs with very little data. Among them, positive samples refer to the distance and time between these camera pairs passed by in a person's cross-camera trajectory, and negative samples refer to the data pairs composed of camera pairs that are not nodes of the same cross-camera trajectory and their time differences. During the optimization process, 16 positive samples and 16 negative samples are randomly sampled for training, the cross-entropy loss function is used to calculate the loss of the single-camera pedestrian trajectory information with annotations, and the gradient descent algorithm is used to optimize the loss information of the single-camera pedestrian trajectory with annotations to obtain the spatio-temporal information of the single-camera pedestrian trajectory with annotations.

[0075] S223. Optimize the spatio-temporal model based on the optimized information of the single-camera pedestrian trajectory with annotations to obtain an optimized spatio-temporal model.

[0076] Specifically, the spatio-temporal model under the camera network is obtained through optimization, and its mathematical representation is d represents the shortest path distance between cameras, t represents time, and its output represents the probability of taking time t to pass through them under the condition that the shortest distance between cameras is d.

[0077] S23. Integrate the optimized appearance feature model and the optimized spatio-temporal model to obtain a pedestrian spatio-temporal trajectory model.

[0078] S3. Generate spatio-temporal information of a pedestrian's cross-camera trajectory based on the pedestrian spatio-temporal trajectory model;

[0079] S31. Based on the pedestrian spatio-temporal trajectory model, perform clustering processing on the spatio-temporal information of the single-camera pedestrian trajectory to obtain a pedestrian trajectory association dataset and construct a trajectory graph model;

[0080] S311. Based on the pedestrian spatio-temporal trajectory model, perform extraction processing on the spatio-temporal information of the single-camera pedestrian trajectory to obtain spatio-temporal extraction information of the single-camera pedestrian trajectory;

[0081] Specifically, define the input spatio-temporal information of the single-camera pedestrian trajectory as where S iDenote the set of tracking trajectories under the \(i\)-th camera, \(N\) c Denote the number of camera networks. Use the feature extraction model to extract the appearance features for each single-camera pedestrian trajectory in the set \(S\). All the extracted appearance features form the set where \(A\) i denotes a single-camera trajectory. \(n_0\) represents the number of all single-camera trajectories in \(S\).

[0082] S312. Perform clustering calculation on the spatio-temporal extraction information of single-camera pedestrian trajectories through the hierarchical clustering algorithm to obtain the number of categories of the spatio-temporal extraction information of single-camera pedestrian trajectories;

[0083] Specifically, use the hierarchical clustering method to perform hierarchical clustering on the set \(A\) to obtain the set of single-camera trajectories of pedestrians with similar appearances \(B = \{B_1, B_2, \ldots, B\) |B| \(\}\), where \(|B|\) represents the number of categories obtained by clustering.

[0084] S313. Perform splicing processing on the number of categories of the spatio-temporal extraction information of single-camera pedestrian trajectories through the companion clustering algorithm to obtain the pedestrian trajectory association data set;

[0085] S314. Establish a trajectory graph model according to the pedestrian trajectory association data set;

[0086] Specifically, referring to Figure 4 and using the hierarchical clustering method to perform clustering on the single-camera pedestrian trajectories under each camera in the set \(S\). The feature used for clustering is the time when each single-camera trajectory appears under the camera. If the time difference between two single-camera trajectories appearing under the same camera is less than a certain threshold, they are called companions and grouped into one category. Let the clustering result under the \(i\)-th camera be where \(|C\) i | represents the number of categories obtained by clustering under the \(i\)-th camera, and \(D\) i· represents the specific category. Concatenate the clustering results under all cameras to obtain which is the final companion clustering result, where \(n\) c represents the number of cameras.

[0087] S315. The trajectory graph model consists of an adjacency matrix, and the adjacency matrix includes a directed edge adjacency matrix and an undirected edge adjacency matrix.

[0088] Specifically, assume that the graph \(G\) consists of \(n_0\) nodes, and each node represents a single-camera trajectory. If two nodes belong to the same \(B\) iIf so, there is a directed edge between them, and the direction of the edge is from the trajectory that appears earlier in time to the one that appears later among the two single-camera tracks. The weight of the edge is obtained in the next step. If two nodes belong to the same category obtained from the spatio-temporal peer clustering, there is an undirected edge between them. Traverse the set B. For each B i Perform the following steps: Let B i There are m single-camera tracks of pedestrians with similar appearances, which can be represented by B i ={K1, K2, …, K m}, where K j represents the j-th single-camera track in B i . Define the adjacency matrix to represent the spatio-temporal relationship between every two in B i . Its size is m×m, and the value in the p-th row and q-th column is the spatio-temporal probability between K p and K p , which is also the weight of the edge between nodes K p and K q in the graph G. Let K p and K q appear under cameras a and b respectively, and the distance between cameras a and b is d ab . The times when K m and K n appear are t a and t b respectively, and the absolute value of the time difference is t ab =|t a -t b |. The corresponding spatio-temporal probability is Thus, |B| adjacency matrices can be obtained. These |B| adjacency matrices are the adjacency matrix representations of the directed edges in the graph G

[0089] S32. The spatio-temporal information of the single-camera pedestrian tracks includes the optimized information of the single-camera pedestrian images with annotations and the optimized information of the single-camera pedestrian tracks with annotations;

[0090] S33. Perform correlation calculations on the trajectory graph model through a correlation algorithm to obtain the spatio-temporal information of the pedestrian cross-camera tracks.

[0091] S331. Perform threshold truncation processing and search processing on the adjacency matrix in the trajectory graph model in sequence to obtain the pedestrian cross-camera tracks;

[0092] Specifically, perform threshold truncation processing on the adjacency matrix in the trajectory graph model. Traverse all the adjacency matrices in graph G. Without loss of generality, assume that the currently traversed adjacency matrix is R. Define the hyperparameter ξ, set the values in R that are less than ξ to 0, and at the same time delete the corresponding edges in graph G. Further, perform search processing on the adjacency matrix in the trajectory graph model. Traverse all the adjacency matrices in graph G, ignoring undirected edges. Find all the nodes with in-degree 0 as the starting points and all the nodes with out-degree 0 as the ending points in the adjacency matrix. Starting from the starting point, along the directed edges, and ending at the ending point, the path passed through in the middle is a cross-camera trajectory. Use the exhaustive method to obtain all the paths and get all the cross-camera trajectories of this adjacency matrix.

[0093] S332. Calculate the cosine value of the directed edge adjacency matrix in the trajectory graph model to obtain the trajectory energy matrix;

[0094] Specifically, the directed edge adjacency matrix describes the correlation degree between different single-camera trajectories in a cross-camera trajectory. Assume that B h has v nodes p1, p2, …, p v , and the corresponding adjacency matrix is R. The i-th row of the matrix can be expressed as R i: , and E1(h) represents the trajectory energy matrix of B h , which has the same size as R. E1(v i , v j ) represents the value of the i-th row and j-th column of E1, and its meaning is the trajectory energy between nodes v i and v j . Therefore, E1(v i , v j ) can be expressed by the following formula:

[0095] E1(v i , v j ) = u1 × R ij + u2 × sim(R i: , R j: )

[0096] In the above formula, E1(v i , v j ) represents the value of the i-th row and j-th column of E1, sim(·, ·) represents calculating the cosine similarity of two vectors, R i: represents the i-th row of the adjacency matrix, and R j: represents the j-th column of the adjacency matrix. u1 and u2 represent hyperparameters;

[0097] Furthermore, sim(·, ·) can be obtained by the following formula:

[0098]

[0099] In the above formula, R i: ·R j: represents the inner product between two row vectors of R i: and R j: . |R i: | represents the norm of the R i: vector itself, |R i: ·R j: | represents the vector inner product of R i: and R j: . |R j: | represents the norm of the R j: vector itself.

[0100] S333. Perform companion multiplication between the undirected edge adjacency matrices in the trajectory graph model to obtain a companion energy matrix;

[0101] Specifically, the undirected edge adjacency matrix describes the correlation degree between the trajectories belonging to companion nodes under the same camera. Since two nodes with an undirected edge adjacency matrix must belong to two different apparent clustering categories, it may be assumed that these two clustering categories are B h and B g . There are v nodes in B h , denoted as p1, p2, …, p v , and the corresponding adjacency matrices are R respectively. There are f nodes in B g , denoted as q1, q2, …, q f , and the corresponding adjacency matrices are T respectively. The x-th row in R can be expressed as R x: , and the y-th row in T can be expressed as T y: . Assume that the h in B and the g in B nodes are in one-to-one correspondence as companions. represents α1, α2, …, α w columns of the x-th row in R, represents β1, β2, …, β w columns of the y-th row in T. Then the companion energy vector of nodes p x and q y can be expressed as:

[0102]

[0103] In the above formula, ⊙ represents element-wise multiplication, |·| represents the norm, E2 represents a vector with w dimensions, represents a mapping of E2(p x , q y ), and its dimension is v. Denote α1, α2, …, α of x rows in R w columns, and denote β1, β2, …, β of y rows in T w columns;

[0104] The value of the α x -th dimension in E2(p y , q i ) is the value of the i-th dimension in , where i ranges from 1 to w and the values of the remaining dimensions are 0. Let M(p x ) denote the set of companions of p x , then represents the companion energy vector of node p x . Concatenating the energy vectors of all nodes in B h row by row from top to bottom gives the companion matrix E2(h) of B h .

[0105] S334. Integrate the trajectory energy matrix and the companion energy matrix to obtain the energy matrix;

[0106] Specifically, perform a summation operation on the trajectory energy matrix and the companion energy matrix. The summation formula is as follows:

[0107] E(h) = E1(h) + E2(h)

[0108] In the above formula, E(h) represents the comprehensive energy matrix, E1(h) represents the self-energy matrix, and E2(h) represents the companion matrix.

[0109] S335. Iteratively update the energy matrix through the mean field algorithm to obtain the updated energy matrix;

[0110] Specifically, use the mean field algorithm to iteratively update the energy matrix. First, perform preprocessing by grouping all the nodes in graph G. The grouping rule is: if there is an edge connecting two nodes, they belong to the same group; otherwise, they do not. Then, traverse all the groups in turn and update the adjacency matrices of the nodes within each group. Without loss of generality, assume that the group to be updated now consists of B1, B2, …, B k apparent clustering categories, and their corresponding adjacency matrices are R1, R2, …, R k . The corresponding spatio-temporal trajectory energies E(1), E(2), …, E(k) can be obtained through the summation formula. Next, use the mean field algorithm to update all the adjacency matrices. Define the matrix W(k) as the exponential representation of the spatio-temporal trajectory energy E(k), which has the same size as E(k). The value of its i-th row and j-th column can be obtained from the following formula:

[0111]

[0112] In the above formula, E ij (k) represents the element in the i-th row and j-th column of E(k), e represents the natural exponential, λ represents the hyperparameter, and W ij (k) represents the element in the i-th row and j-th column of W(k);

[0113] After that, perform row normalization on W(k). Let W i: (k) represent the i-th row of W(k), then the normalized row vector can be expressed as:

[0114] W i: (k) = W i: (k) ÷ W ii (k)

[0115] In the above formula, W i: (k) represents the i-th row of W(k), and W ii (k) represents the element in the i-th row and i-th column of W(k);

[0116] Traverse all rows of W(k) to perform row normalization. For all adjacency matrices in this group, perform the above steps to obtain the row-normalized W(1), W(2), …, W(k), and then assign them to R1, R2, …, R k , this round of iteration ends. Repeat the above process for t rounds. t is the hyperparameter. After the adjacency matrices of this group are updated, traverse all groups to update the adjacency matrices and then this step ends. Use a threshold to perform truncation processing on all adjacency matrices in graph G. Define the hyperparameter θ, traverse all adjacency matrices, and set all element values less than θ in the adjacency matrices to 0.

[0117] S336. Based on the pedestrian cross-camera trajectory, decompose and calculate the updated energy matrix through the non-negative matrix factorization method to obtain the spatio-temporal information of the pedestrian cross-camera trajectory.

[0118] Specifically, referring to Figure 3 , use the non-negative matrix factorization method to obtain the cross-camera trajectory of the pedestrian. Traverse all adjacency matrices in graph G in turn. Without loss of generality, assume that the currently traversed adjacency matrix is R h , with a size of |B h | × |B h |, where |B h | is the number of elements in the clustering category B h . Define a new matrix U with a size of |B h | × k, where k is the number of possible trajectories to be obtained, and initialize the matrix U with random values. Update the values in U according to the following method and iterate L times:

[0119]

[0120] In the above formula, U represents the decomposition matrix, represents the symbol of matrix element multiplication, sqrt(·) represents the square root of matrix elements, represents matrix element division, and R h represents the adjacency matrix, I2 represents a column vector of size N×1, and τ represents a hyperparameter. represents a row vector of size 1×k, and UU T represents the matrix multiplication of U and U T .

[0121] The value of k can be obtained by the following method. Calculate the h eigenvalues of R and take the absolute values. Define the hyperparameter σ. The number of eigenvalues of R h whose absolute values are greater than σ is the value of k. Finally, to obtain the decomposition matrix U, first traverse each row of U, set the value of the column with the largest value in each row to 1 and the other columns to 0. Then traverse each column, find the rows with a value of 1 in each column, and group the corresponding nodes into one trajectory. Finally, the pedestrian cross-camera trajectory information is obtained.

[0122] S4. Detect the pedestrian image to be measured based on the spatio-temporal information of the pedestrian cross-camera trajectory, and output the detection result of the pedestrian group.

[0123] S41. Train the pedestrian image to be measured based on the appearance feature model to obtain the trajectory information of the pedestrian to be measured;

[0124] S42. Based on the pedestrian trajectory reordering method, combine the spatio-temporal information of the pedestrian cross-camera trajectory to retrieve the trajectory information of the pedestrian to be measured, and obtain the pedestrian video retrieval result;

[0125] Specifically, assume that the obtained pedestrian cross-camera trajectory is the set T = {O1, O2,..., O |T|}, where |T| represents the number of all pedestrian cross-camera trajectories obtained by the pedestrian trajectory generation module, and O i represents a specific cross-camera trajectory, which consists of one or more single-camera trajectories. Without loss of generality, assume that the appearance features of these single-camera trajectories are where |O i | represents the number of single-camera trajectories in O i . Then the feature P i of the corresponding pedestrian cross-camera trajectory can be obtained by the following formula:

[0126]

[0127] In the above formula, P i represents the feature of the corresponding pedestrian cross-camera trajectory, represents the appearance feature of the single-camera trajectory, and |Oi | represents O i The number of single-camera trajectories.

[0128] Traverse all pedestrian cross-camera trajectories to obtain the set of all pedestrian cross-camera trajectory features P = {P1, P2, …, P |T|};

[0129] Use the appearance feature extraction model to extract the features of the pedestrian image sequence to be retrieved, and take the average as the feature vector to be retrieved. Compare and sort the image features of the pedestrian to be retrieved with the cross-camera trajectory features to obtain the pedestrian trajectory retrieval result. Calculate the cosine distance between the feature vector of the pedestrian to be retrieved and the cross-camera trajectory features of the pedestrian, and sort them from smallest to largest. The final sorting result is the pedestrian trajectory retrieval result;

[0130] Further explain the pedestrian trajectory re-ranking method. The pedestrian trajectory re-ranking method can convert the result of pedestrian trajectory retrieval into the retrieval result format of the pedestrian image sequence. Perform pedestrian trajectory retrieval to obtain the result of the first sorting. This step is the same as the above pedestrian trajectory retrieval step. Perform the second sorting, calculate the cosine distance by comparing the single-camera features of the pedestrians in the trajectory with the features of the pedestrian to be retrieved, and sort them from smallest to largest according to the distance. Perform re-ranking to obtain the final retrieval result. Define an empty set Y = {}, traverse all cross-camera trajectories according to the result of the first sorting, and traverse the single-camera trajectories of each cross-camera trajectory according to the result of the second sorting. If the current single-camera trajectory is not in Y, add it to Y, otherwise skip it. Finally, the order in which each single-camera trajectory is added to Y is the final pedestrian trajectory re-ranking result.

[0131] S43. Based on the pedestrian group detection framework, combine the spatio-temporal information of pedestrian cross-camera trajectories to retrieve the information of the pedestrian trajectory to be measured, and output the pedestrian group detection result.

[0132] S431. Retrieve the information of the pedestrian trajectory to be measured to obtain the set of single-pedestrian cross-camera trajectories;

[0133] Specifically, traverse all pedestrian image sequences to be grouped, and use the appearance feature model to extract their features. Here, it is assumed that the extracted feature set is Q = {q1, q2, …, q |Q|}, where |Q| represents the number of pedestrians to be grouped, and q iRepresent specific pedestrian features, determine the possible cross-camera trajectories of the pedestrian to be retrieved under the camera network, further define the hyperparameter δ, traverse all elements in Q. Without loss of generality, assume that the current pedestrian feature vector to be retrieved is q1, use the pedestrian trajectory retrieval module to retrieve the pedestrian trajectory, and then use the threshold screening method to filter the retrieval sorting results, deleting the cross-camera trajectories with a cosine distance greater than δ. The remaining trajectories are used as the possible cross-camera trajectory set corresponding to q1, which is defined as where |H1| represents the number of obtained cross-camera trajectories.

[0134] S432. Calculate the single-pedestrian cross-camera trajectory set through the pedestrian correlation matrix to obtain the single-pedestrian cross-camera trajectory distance;

[0135] Construct the pedestrian correlation matrix F. The size of the pedestrian correlation matrix F is |Q|×|Q|, initialized to 0. Its i-th row and j-th column describe the distance between the i-th pedestrian and the j-th pedestrian. If they belong to the same group, their distance approaches 0, otherwise it approaches infinity. Then calculate the correlation between the pedestrians to be detected. The pedestrians to be detected are paired pairwise to find the difference between their cross-camera trajectories. Without loss of generality, assume that the current correlation between q1 and q2 is being calculated, and their corresponding pedestrian cross-camera trajectories are sets H1 and H2. Therefore, finding the correlation distance d1(q1, q2) between the two people is equivalent to the distance between the two cross-camera trajectory sets:

[0136] d1(q1, q2) = d2(H1, H2)

[0137] In the above formula, d1(q1, q2) represents the correlation distance between the two people, and d2(H1, H2) represents the two pedestrian cross-camera trajectory sets;

[0138] S433. Calculate the correlation distances between pedestrians by calculating multiple single-pedestrian cross-camera trajectory distances through a metric algorithm;

[0139] The values of the two pedestrian cross-camera trajectory sets can be obtained through the following steps:

[0140] Define the method for measuring the distance between single-camera trajectories. The distance between single-camera trajectories consists of three parts: the spatial distance d4, the temporal overlap distance d5, and the appearance time distance d6. First, introduce how to calculate the spatial distance: Suppose there are two single-camera trajectories J1 and J2 under the same camera. Each single-camera trajectory includes an image sequence and the positions of the detection boxes of the corresponding images on the original frame. At the same time, the external and internal parameters of the current camera should also be known. First, match the two image sequences frame by frame. If there are frames with missing detection boxes in one image sequence, they are completed by linear interpolation. Then, traverse each frame in the order of frame numbers. Define the hyperparameter ε as the penalty term for non-matching. Define err to represent the spatial difference between the two detection boxes in the current frame. If there is only one trajectory with a detection box in the current frame, it means there is no match, and err = ε lost , otherwise there is a match, and err can be calculated through the following steps: Define the external parameter translation vector of the current camera as T, the rotation matrix as R, the internal parameter focal length as f, the image scaling factor as s, the image center coordinates as (c1, c2), and the pixel coordinates of the bottom center points of the two detection boxes as (x1, y1) and (x2, y2) respectively. Then, the coordinates of (x1, y1) and (x2, y2) in the top-down plane coordinate system can be obtained through the method of homography matrix mapping, and then the sum of the squared differences of the two coordinates is calculated as the err of the current frame. Finally, the sum of the err of all frames is divided by the total number of frames to obtain the average spatial distance d4 between the two trajectories. The second distance is the temporal overlap distance. Suppose the start timestamp of J1 is t 11 , the end timestamp is t 12 , and the duration of continuous existence is from t 11 to t 12 . The start timestamp of J2 is t 21 , the end timestamp is t 22 , and the duration of continuous existence is from t 21 to t 22 . Suppose there is an overlapping time of Δt between t 11 to t 12 and t 21 to t 22 , and the maximum time difference is max(t 22 , t 12 ) - min(t 11 , t 21 ), where max represents finding the maximum value and min represents finding the minimum value. Then, the second temporal coverage distance can be obtained from the following formula:

[0141]

[0142] In the above formula, max(t 22 , t12 ) - min(t 11 , t 21 ) represents the maximum time difference, Δt represents the overlapping time in the middle, and d5(J1, J2) represents the time coverage distance;

[0143] The third term is the minimum time difference distance, and its calculation method can be obtained by the following formula:

[0144] d6(J1, J2) = |t 11 - t 21 |

[0145] In the above formula, d6(J1, J2) represents the minimum time difference distance between two single - camera trajectories J1 and J2, and |t 11 - t 21 | represents the time difference;

[0146] Finally, the distance between single - camera trajectories is:

[0147] d3(J1, J2) = u4 × d4(J1, J2) + u5 × d5(J1, J2) + u6 × d6(J1, J2)

[0148] In the above formula, d3(J1, J2) represents the comprehensive distance between two single - camera trajectories J1 and J2, u4 represents a hyperparameter, d4(J1, J2) represents the spatial distance, u5 represents a hyperparameter, d5(J1, J2) represents the time coverage distance, u6 represents the hyperparameter of the minimum time difference distance, and d6(J1, J2) represents the appearance time distance;

[0149] In the first part, assume that the two cross - camera trajectories for which the distance is required are O1 and O2 respectively. Define the matrix d6(O1, O2) to represent the distance between the two cross - camera trajectories. It consists of two parts, namely the edit distance and the cross - camera trajectory matching distance. Define d7(O1, O2) to represent the ratio edit distance of two cross - camera trajectories. It can be obtained through the following steps: First, define the mapping from camera numbers to letters. For example, number 1 represents a, number 2 represents b, and so on. Convert the cross - camera trajectories in O1 and O2 into camera letter sequences. Define represents the edit distance between the first i characters of the letter sequence of O1 and the first j characters of the letter sequence of O2, and it can be recursively obtained by the following formula:

[0150]

[0151] In the above formula, represents an indicator function, Denote the edit distance between the first i characters of the alphabetic sequence of O1 and the first j characters of the alphabetic sequence of O2. min(i - 1, j - 1) represents the smaller one of i - 1 and j - 1, min(i, j) represents the smaller one of i and j, min(i - 1, j)+1 represents the smaller one of i - 1 and j, and min(i, j - 1)+1 represents the smaller one of i and j - 1;

[0152] If the i-th character of O1 is different from the j-th character of O2, it is 1, otherwise it is 0. Thus, d7(O1, O2) can be obtained by the following formula:

[0153]

[0154] In the above formula, d7(O1, O2) represents the ratio edit distance between two cross-camera trajectories, Denote the edit distance between sequences O1 and O2, and max(|O1|, |O2|) represents the length of the longest sequence among O1 and O2;

[0155] The second part is the cross-camera trajectory matching distance, denoted by d8(O1, O2). To obtain this distance, single-camera trajectory matching needs to be performed first. Define V as the matching distance matrix of single-camera trajectories of O1 and O2, with size |O1|×|O2|. The element in the i-th row and j-th column of it represents the distance between the i-th single-camera trajectory of O1 and the j-th single-camera trajectory of O2. If the i-th single-camera trajectory of O1 and the j-th single-camera trajectory of O2 are not from the same camera, then the element in the i-th row and j-th column is infinity. Otherwise, use the distance metric method defined in a) to calculate their distance and assign it to the element in the i-th row and j-th column. Then use the Hungarian matching method to match the single-camera trajectories in the two cross-camera trajectories. Input the matrix V and output the matching result. Define two hyperparameters τ and μ. Traverse all single-camera matching pairs. If the value of a single-camera trajectory pair in the V matrix is less than τ, then keep this pair of matches; otherwise, delete this pair of matches. Sum up the distances between the single-camera trajectories of all the matched single-camera trajectory pairs and take the average, and then add the number of unmatched single-camera trajectories multiplied by μ as the cross-camera trajectory matching distance between O1 and O2, that is:

[0156] d8(O1, O2) = u7×d7(O1, O2)+u8×N notmatch ×μ

[0157] In the above formula, d8(O1, O2) represents the ratio edit distance between two cross-camera trajectories, u7 and u8 represent weight coefficients, N notmatch Represents the number of unmatched single-camera trajectories, and μ represents the hyperparameter;

[0158] Further calculate the relevant distance between pedestrians, define the pedestrian correlation matrix Z, with a size of |H1|×|H2|, where Z ij represents the distance between the i-th cross-camera trajectory in H1 and the j-th cross-camera trajectory in H2. Then, use the Hungarian matching method to match the cross-camera trajectories in H1 and H2, and define hyperparameters. Define two hyperparameters τ and μ. Traverse all cross-camera matching pairs. If the value of the single-camera trajectory pair in the Z matrix is less than τ for a pair, then retain this pair of matches; otherwise, delete this pair of matches. Sum and average the distances between the single-camera trajectories of all the cross-camera trajectory pairs that are matched, and then add the number of cross-camera trajectories that are not matched multiplied by μ to represent the cross-camera trajectory matching distance between H1 and H2 as follows:

[0159] d2(H1,H2)=min(Z)

[0160] In the above formula, d2(H1,H2) represents the set of cross-camera trajectories of two pedestrians, and min(Z) represents the minimum value in Z.

[0161] S434. Screen the relevant distances between pedestrians to obtain the pedestrian group detection result;

[0162] Through the above steps, the relevant distances between two pedestrians are calculated. Traverse all pairs of pedestrians and assign the corresponding distances to the matrix F. Screen the group detection results through a threshold. Define the hyperparameter η. Traverse each row of F to obtain the columns with values less than η. Then, the pedestrians corresponding to these columns and the pedestrians corresponding to the current row form a pedestrian group.

[0163] Furthermore, the simulation experiment of the present invention is as follows:

[0164] Refer to Figure 5 , taking the video data of a camera network in a community collected by the present invention itself as an example. The data used comes from the camera video data under 9 cameras in a certain community. The spatio-temporal distribution of this camera network is as Figure 5As shown below. This dataset includes these 9 cameras. This patent divides this dataset into two datasets, a pedestrian trajectory retrieval dataset and a pedestrian group detection dataset, in two ways. The pedestrian trajectory retrieval dataset divides the data in the dataset according to the number of cross-camera pairs in the pedestrian cross-camera trajectory, such that the number of pedestrians in the training set and the number of pedestrians in the test set are close to 3 times, and at the same time, the number of cross-camera pairs of the cross-camera trajectory in the training set is also 3 times that of the test set. Among them, the trajectory data of this data is used for the training of the spatio-temporal model. Finally, there are 465 pedestrians and 12,993 images in the training set of the pedestrian trajectory retrieval dataset, and 197 pedestrians and 5,003 images in the test set. The pedestrian group detection dataset divides most of the single-person data into the training set, and all the pedestrian groups and some single-person data are divided into the test set. There are a total of 494 pedestrians in the training set of the pedestrian group detection dataset and 168 pedestrians in the test set. Among them, in the test set, 60 pedestrians belong to 29 groups, and the remaining 45 pedestrians do not belong to any group. Table 1 shows the pedestrian video retrieval data, and Table 2 shows the pedestrian group detection data. The specific tables are as follows:

[0165]

[0166] Table 1

[0167] Method Accuracy (%) Recall (%) F1 Score (%) Graph Search Model 72.8 76.2 74.4 Pedestrian Spatiotemporal Trajectory Model 74.8 81.0 77.8

[0168] Table 2

[0169] Referring to Figure 2 , a pedestrian group detection system based on pedestrian trajectory retrieval, includes:

[0170] An acquisition module, configured to acquire an annotated dataset and construct a training dataset;

[0171] A training module, which trains an appearance feature model and a spatio-temporal model based on the training dataset, and integrates them to obtain a pedestrian spatio-temporal trajectory model;

[0172] A generation module, which generates spatio-temporal information of pedestrian cross-camera trajectories based on the pedestrian spatio-temporal trajectory model;

[0173] An output module, which detects the pedestrian image to be measured based on the spatio-temporal information of the pedestrian cross-camera trajectory, and outputs the pedestrian group detection result.

[0174] The content in the above method embodiments is applicable to the system embodiments. The functions specifically implemented in the system embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0175] The above is a specific description of the preferred embodiment of the present invention. However, the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A pedestrian group detection method based on pedestrian trajectory retrieval, characterized in that, It includes the following steps: Obtain a labeled dataset and construct a training dataset; Train the appearance feature model and the spatio-temporal model based on the training dataset, and integrate them to obtain a pedestrian spatio-temporal trajectory model; Generate pedestrian cross-camera trajectory spatio-temporal information based on the pedestrian spatio-temporal trajectory model; Detect the pedestrian image to be measured based on the pedestrian cross-camera trajectory spatio-temporal information, and output the pedestrian group detection result; The step of generating pedestrian cross-camera trajectory spatio-temporal information based on the pedestrian spatio-temporal trajectory model specifically includes: Based on the pedestrian spatio-temporal trajectory model, perform clustering processing on the single-camera pedestrian trajectory spatio-temporal information to obtain a pedestrian trajectory association dataset and construct a trajectory graph model; The single-camera pedestrian trajectory spatio-temporal information includes the optimized information of the labeled single-camera pedestrian image and the optimized information of the labeled single-camera pedestrian trajectory; Perform association calculation on the trajectory graph model through an association algorithm to obtain pedestrian cross-camera trajectory spatio-temporal information.

2. The method for detecting a pedestrian group based on pedestrian trajectory retrieval according to claim 1, wherein, The step of obtaining a labeled dataset and constructing a training dataset specifically includes: Obtain a dataset through a network camera, perform annotation processing on the dataset to obtain a labeled dataset; Construct a training dataset based on the labeled dataset, and divide the training dataset into labeled single-camera pedestrian image information and labeled single-camera pedestrian trajectory information.

3. The method for detecting a pedestrian group based on pedestrian trajectory retrieval according to claim 2, wherein The step of training the appearance feature model and the spatio-temporal model based on the training dataset and integrating them to obtain a pedestrian spatio-temporal trajectory model specifically includes: Train the appearance feature model based on the labeled single-camera pedestrian image information and train the spatio-temporal model based on the labeled single-camera pedestrian trajectory information to obtain an optimized appearance feature model and an optimized spatio-temporal model; Integrate the optimized appearance feature model and the optimized spatio-temporal model to obtain a pedestrian spatio-temporal trajectory model.

4. The method for detecting a pedestrian group based on pedestrian trajectory retrieval according to claim 3, wherein The step of training the appearance feature model based on the labeled single-camera pedestrian image information and training the spatio-temporal model based on the labeled single-camera pedestrian trajectory information to obtain an optimized appearance feature model and an optimized spatio-temporal model specifically includes: Based on the appearance feature model, perform feature extraction processing on the labeled single-camera pedestrian image information to obtain labeled single-camera pedestrian image feature information, and the appearance feature model is composed of a convolutional neural network; Calculate the labeled single-camera pedestrian image feature information through a triplet loss function and a gradient descent algorithm respectively to obtain labeled single-camera pedestrian image optimized information; Perform optimization processing on the appearance feature model based on the labeled single-camera pedestrian image optimized information to obtain an optimized appearance feature model; Based on the spatio-temporal model, calculate the labeled single-camera pedestrian trajectory information through a cross-entropy loss function and a gradient descent algorithm respectively to obtain labeled single-camera pedestrian trajectory optimized information; Perform optimization processing on the spatio-temporal model based on the labeled single-camera pedestrian trajectory optimized information to obtain an optimized spatio-temporal model.

5. The method for detecting a pedestrian group based on pedestrian trajectory retrieval according to claim 4, characterized in that, The step of performing clustering processing on the single-camera pedestrian trajectory spatio-temporal information based on the pedestrian spatio-temporal trajectory model to obtain a pedestrian trajectory association dataset and construct a trajectory graph model specifically includes: Based on the pedestrian spatiotemporal trajectory model, the spatiotemporal information of the pedestrian trajectory of a single camera is extracted and processed to obtain the spatiotemporal extracted information of the pedestrian trajectory of a single camera; The hierarchical clustering algorithm is used to perform clustering calculation on the spatiotemporal information extracted from the single camera pedestrian trajectory, and the number of categories of the spatiotemporal information extracted from the single camera pedestrian trajectory is obtained; The number of categories of spatiotemporal information extracted from single-camera pedestrian trajectories is spliced ​​through a peer clustering algorithm to obtain a pedestrian trajectory association dataset. Establish a trajectory graph model based on the pedestrian trajectory association dataset; The trajectory graph model is composed of an adjacency matrix, which includes a directed edge adjacency matrix and an undirected edge adjacency matrix.

6. The method for detecting a pedestrian group based on pedestrian trajectory retrieval according to claim 5, wherein The step of performing association calculation on the trajectory graph model through an association algorithm to obtain the spatiotemporal information of the pedestrian's cross-camera trajectory specifically includes: The adjacency matrix in the trajectory graph model is subjected to threshold truncation and search processing in turn to obtain the pedestrian's cross-camera trajectory; The cosine value of the directed edge adjacency matrix in the trajectory graph model is calculated to obtain the trajectory energy matrix; Perform peer multiplication between the undirected edge adjacency matrices in the trajectory graph model to obtain the peer energy matrix; Integrate the trajectory energy matrix and the companion energy matrix to obtain the energy matrix; The energy matrix is ​​iteratively updated by the mean field algorithm to obtain an updated energy matrix; Based on the cross-camera trajectory of pedestrians, the updated energy matrix is ​​decomposed and calculated using the non-negative matrix decomposition method to obtain the spatiotemporal information of the cross-camera trajectory of pedestrians.

7. The method for detecting a pedestrian group based on pedestrian trajectory retrieval according to claim 6, wherein, The step of detecting the pedestrian image to be detected based on the spatiotemporal information of the pedestrian's cross-camera trajectory and outputting the pedestrian group detection result specifically includes: Based on the appearance feature model, the trajectory features of the image of the pedestrian to be tested are extracted to obtain the trajectory information of the pedestrian to be tested; Based on the pedestrian trajectory re-ranking method, the pedestrian trajectory information to be measured is retrieved by combining the spatiotemporal information of pedestrian trajectories across cameras to obtain the pedestrian video retrieval results; Based on the pedestrian group detection framework, the trajectory information of the pedestrians to be tested is retrieved by combining the spatiotemporal information of the pedestrian trajectories across cameras, and the pedestrian group detection results are output.

8. The method for detecting a pedestrian group based on pedestrian trajectory retrieval according to claim 7, wherein The step of searching the trajectory information of pedestrians to be detected based on the pedestrian group detection framework and combining the spatiotemporal information of pedestrian cross-camera trajectories to output the pedestrian group detection results specifically includes: Retrieve the trajectory information of the pedestrian to be measured and obtain the trajectory set of a single pedestrian across cameras; The trajectory set of a single pedestrian across cameras is calculated through the pedestrian correlation matrix to obtain the trajectory distance of a single pedestrian across cameras; The distances between multiple pedestrians across camera trajectories are calculated using a metric algorithm to obtain the relevant distances between pedestrians. The relevant distances between pedestrians are screened to obtain the pedestrian group detection results.

9. A pedestrian group detection system based on pedestrian trajectory retrieval, characterized in that, The method for detecting a group of pedestrians based on pedestrian trajectory retrieval according to claim 1 comprises the following modules: The acquisition module is used to obtain the labeled data set and construct the training data set; The training module trains the appearance feature model and spatiotemporal model based on the training data set, and integrates them to obtain the pedestrian spatiotemporal trajectory model; The generation module generates the spatiotemporal information of pedestrian trajectories across cameras based on the pedestrian spatiotemporal trajectory model; An output module that detects a pedestrian image to be measured based on the spatio-temporal information of pedestrian cross-camera trajectories and outputs the detection result of a pedestrian group.

Citation Information

Patent Citations

  • Pedestrian target movement track acquisition method and system based on multiple cameras

    CN110378931A

  • Deep learning-based cross-camera pedestrian multi-target tracking method and device

    CN112270310A