A method and system for three-dimensional tracking of multiple pedestrians based on monocular vision
An end-to-end neural network is used to solve the problem of three-dimensional tracking of multiple pedestrians using monocular vision. Combined with the Gaussian distribution assumption of height and a parameter-free positioning method, accurate detection and tracking of the 3D positions of pedestrians are achieved, improving the operating efficiency and accuracy of the system.
Patent Information
- Application Number
- CN202111148568.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-28
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-09-28
AI Technical Summary
Existing intelligent monitoring systems based on monocular vision lack an efficient, end-to-end method for real-time calculation and tracking of the 3D spatial positions of multiple pedestrians. In addition, monocular cameras have difficulty obtaining depth of field information, resulting in insufficient three-dimensional positioning and tracking.
A multi-pedestrian 3D tracking method based on monocular vision is adopted. Through an end-to-end neural network of feature extraction, pedestrian detection and ID embedding learning, 3D positioning, re-identification and 3D tracking, the Gaussian distribution assumption of height is combined to resolve scale ambiguity. The intermediate results of the pedestrian detection module are used to realize 3D position calculation without additional parameters. Module parallel learning and parameter-free positioning methods are adopted.
It achieves accurate detection and association of identical pedestrians, accurate prediction of pedestrian 3D positions, and drawing of pedestrian 3D activity trajectories. It also improves operational efficiency while ensuring tracking accuracy, thereby increasing computational efficiency and tracking accuracy.
Smart Images

Figure CN113888594B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-pedestrian three-dimensional tracking method based on monocular vision, belonging to the field of artificial intelligence monitoring technology, in particular to the field of pedestrian detection, three-dimensional positioning and three-dimensional tracking technology in artificial intelligence monitoring. Background Art
[0002] With the rapid development of the Internet of Things and artificial intelligence (AI) technologies, camera-based security monitoring, ecological monitoring, and smart home applications have been widely used. In these applications, pedestrians are the primary targets of surveillance. Due to the low cost and ease of use of monocular cameras, multi-pedestrian detection, re-identification, localization, and tracking based on monocular vision have become a research hotspot in this field. However, existing monocular vision-based intelligent monitoring systems suffer from at least two problems: 1. Due to the difficulty in obtaining depth of field information, existing monocular cameras are primarily used for 2D target detection and localization, lacking effective means for three-dimensional spatial positioning and tracking; 2. Existing intelligent monitoring systems are primarily designed to address the problems of multi-pedestrian detection, re-identification, localization, and tracking, lacking an efficient, end-to-end method for calculating the 3D spatial positions and tracking of multiple pedestrians in real time based on video streams captured by monocular cameras.
[0003] Qin et al. (Monogrnet: A geometric reasoning network for monocular 3D object localization, Zengyi Qin, et al., Proceedings of the AAAI Conference on Artificial Intelligence, 33, 8851-8858, 2019) designed a neural network to predict instance-level depth values of objects in a spatial coordinate system. They then used geometric relationships to calculate the horizontal and vertical offsets of the objects, and combined these three values to form the 3D position. However, their depth prediction involved learning a large number of parameters. Bertoni et al. (Monoloco: Monocular 3D pedestrian localization and uncertainty estimation, Bertoni et al., Proceedings of IEEE International Conference on Computer Vision, 6860-6870, 2019) predicted the spatial position of pedestrians using pose keypoints and provided an interval range for the pedestrian position using uncertainty estimation. However, this method required the use of another independent neural network to first predict the pedestrian pose keypoints from the monocular image. Summary of the Invention
[0004] In response to the above problems, the purpose of the present invention is to provide a multi-pedestrian three-dimensional tracking method and system based on monocular vision, which can accurately detect and associate the same pedestrians, accurately predict the 3D positions of pedestrians, draw the 3D activity trajectories of pedestrians, and achieve real-time operation efficiency. On the basis of ensuring competitive tracking accuracy, the operation efficiency of the entire model is improved.
[0005] To achieve the above objectives, the present invention adopts the following technical solutions: a method for three-dimensional tracking of multiple pedestrians based on monocular vision, comprising the following steps: S1 extracts features from a training image set obtained by a monocular camera; S2 simultaneously detects pedestrians and learns the ID embedding of the corresponding pedestrians based on the extracted features; S3 performs 3D positioning of pedestrians based on the pedestrian detection results; S4 re-identifies pedestrians based on the results of ID embedding learning; S5 determines the motion trajectory of each pedestrian based on the re-identification results and the 3D position of the pedestrian, thereby tracking it in three dimensions.
[0006] Furthermore, the method for extracting features from the training image set in step S1 is as follows: the category of object j is marked as k in the training image set, and the pixel coordinates of the upper left corner of the bounding box of object j are Pixel coordinates of the lower right corner The fixed-size image is input into the convolutional neural network, and the features are extracted through the convolution operation to obtain a feature map F of size W / S×H / S, where W is the width of object j, H is the height of object j, and S is the downsampling step size.
[0007] Furthermore, the method for pedestrian detection in step S2 is as follows: the pixel points in the input image are matched with the pixel points in the feature map F, and the center point offset of the feature map F is calculated by the loss function; the pixel coordinates of the upper left corner of the bounding box of the object j are calculated based on the pixel coordinates of the upper left corner of the bounding box of the object j. Pixel coordinates of the lower right corner Obtain the bounding box size in the input image; obtain a feature point (x', y') on the feature map F through convolution transformation, and determine the probability that other points (x, y) in the feature map F are the center points of the pedestrian.
[0008] Furthermore, the probability calculation formula is:
[0009]
[0010] Among them, p is the probability that other points (x, y) in the feature map are the pedestrian center points, σ is the standard deviation that can be adjusted according to the object size, and u and v are the physical width and height of the camera photosensitive element respectively.
[0011] Furthermore, the pedestrian center point probability prediction is optimized through the loss function with focal loss, and its calculation formula is:
[0012]
[0013] Among them, L cls is the pedestrian center probability prediction loss function, N is the total number of pedestrians in the input image, α and β are hyperparameters in focal loss, and p j is the probability that the jth point is the pedestrian center point, is the estimated probability that the jth point is the pedestrian center.
[0014] Furthermore, the ID embedding learning method corresponding to the pedestrian in step S2 is: use an embedding vector to represent the pedestrian's features, and the number of embedding vectors is the same as the number of pedestrians with different identities. For pedestrians with the same identity but appearing in different pictures, their embedding vectors are the same or similar, and the embedding vectors of pedestrians with different identities are different.
[0015] Furthermore, the method for 3D positioning of pedestrians in step S3 is as follows: assuming that the height of pedestrians is Gaussian distributed, the bounding box of the pedestrian in the image is projected into the three-dimensional coordinate system of the camera as the 3D position of the pedestrian. The 3D position is calculated by the following formula:
[0016]
[0017] Among them, loc is the position of the pedestrian in the 3D camera coordinate system; x c 、y c and z c are the x-axis, y-axis, and z-axis coordinate values of the pedestrian’s center point in the 3D camera coordinate system; f x and f y are the x- and y-direction coordinates of the camera principal point, H is the physical height of the pedestrian, and h is the pixel height of the pedestrian in the image. (x i ,y i ) is the coordinate of the pixel center point i in the image plane coordinate system, (x j ,y j ) is the coordinate of the pixel center point i in the image plane coordinate system, and the height H of the pedestrian in the image is obtained by taking the Gaussian distribution N(μ,σ 2 ) is the mean μ.
[0018] Furthermore, the re-identification method in step S4 is as follows: S4.1 If the ID embedding vector of pedestrian i at time t is represented by emb i , the ID embedding vector of pedestrian j at time t+1 is expressed as emb j , then the cosine similarity between the two is cos <embi ,emb j >;S4.2 If the bounding box of pedestrian i at time t is box i , the bounding box of pedestrian j at time t+1 is box j , then the intersection-over-union ratio of the two is IoU <box i ,box j >;S4.3 judge cos <emb i ,emb j > and IoU <box i ,box j >Whether the preset threshold conditions are met. If both values meet the preset conditions, proceed to the next step; S4.4 Based on the result at time t, the Kalman filter is used to predict the bounding box position of pedestrian k at time t+1. Detect the bounding box position of pedestrian j at time t+1 And calculate the distance between the two If the distance also meets the threshold condition, pedestrian i at time t and pedestrian j at time t+1 are considered to match.
[0019] Furthermore, the method for three-dimensional tracking of pedestrians in step S5 is: based on the pedestrian re-identification module, the same pedestrian is continuously associated in the time interval interval = {t, t+1, ..., t+n}, and the position of pedestrian i = 1, 2, ..., N at each time t is Thus, the movement trajectory of pedestrian i in the time period interval is obtained
[0020] The present invention also discloses a multi-pedestrian three-dimensional tracking system based on monocular vision, including: a feature extraction module, used to extract features from a training image set obtained by a monocular camera; a pedestrian detection and ID learning module, used to simultaneously perform pedestrian detection and corresponding pedestrian ID embedding learning based on the extracted features; a 3D positioning module, used to perform 3D positioning of pedestrians based on the pedestrian detection results; a re-identification module, used to re-identify pedestrians based on the pedestrian detection and ID embedding learning results; a pedestrian three-dimensional tracking module, used to determine the movement trajectory of each pedestrian based on the re-identification results and the pedestrian's position, thereby tracking it.
[0021] The present invention has the following advantages due to the adoption of the above technical solution:
[0022] 1. The present invention can accurately detect and associate identical pedestrians, accurately predict their 3D positions, and plot their 3D activity trajectories. This can achieve real-time operational efficiency, while ensuring competitive tracking accuracy and improving the operational efficiency of the entire model.
[0023] 2. This paper proposes a method to resolve the scale ambiguity of monocular vision by assuming pedestrian height distribution; a parameter-free pedestrian 3D positioning method to improve computational efficiency; and a pedestrian matching method based on the collaboration of target box IoU consistency, identity embedding consistency, and position consistency to improve the accuracy of pedestrian matching and 3D tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 1 is a schematic structural diagram of a method for three-dimensional tracking of multiple pedestrians based on monocular vision in one embodiment of the present invention;
[0025] Figure 2 is a schematic diagram of a pedestrian detection method according to an embodiment of the present invention;
[0026] Figure 3 FIG. 1 is a schematic diagram of a re-identification method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the technical direction of the present invention, the present invention will be described in detail through specific embodiments. However, it should be understood that the provision of specific embodiments is only for a better understanding of the present invention and they should not be construed as limitations of the present invention. In the description of the present invention, it should be understood that the terms used are for descriptive purposes only and are not to be construed as indicating or implying relative importance.
[0028] This invention discloses a method and system for three-dimensional tracking of multiple pedestrians based on monocular vision. Multiple pedestrian detection involves accurately detecting a tightly coupled rectangular pixel box around each pedestrian from a 2D image captured by a monocular camera. Re-identification involves identifying the same pedestrian from different frames. 3D localization involves calculating the pedestrian's three-dimensional (x, y, z) position. And 3D tracking involves tracking the changing positions of the same pedestrian in three dimensions. This invention integrates multiple pedestrian detection, embedded representation learning of user features, re-identification, 3D localization, and tracking into an end-to-end neural network. A method for 3D localization of pedestrians based on the spatial position of a target box is proposed. First, the ambiguity of pedestrian scale acquired by a monocular camera is addressed by assuming a Gaussian distribution of height. Second, the intermediate results of the pedestrian detection module are fully utilized to implement a parameter-free method for calculating pedestrian 3D positions. This method utilizes parallel module learning and a parameter-free 3D localization method, and demonstrates superior tracking accuracy and efficiency compared to current state-of-the-art methods. The technical solution of this invention is described in detail below through two embodiments, combined with the accompanying drawings.
[0029] Example 1
[0030] This embodiment discloses a method for three-dimensional tracking of multiple pedestrians based on monocular vision. Figure 1 、 2, 3, including the following steps:
[0031] S1 performs feature extraction on the training image set obtained by the monocular camera.
[0032] The method for feature extraction of training images is as follows: mark the category k of object j in the training image set, the pixel coordinates of the upper left corner of the bounding box of object j Pixel coordinates of the lower right corner Establish object category labels. Input a fixed-size image (e.g., W × H × 3) into a convolutional neural network. After convolution, extract features and obtain a feature map F of size W / S × H / S, where W is the width of object j, H is the height of object j, and S is the downsampling step size.
[0033] S2 simultaneously performs pedestrian detection and corresponding pedestrian ID embedding learning based on the extracted features.
[0034] A pixel point (x', y') on the input image corresponds to the pixel coordinates on the feature map F Since pixel coordinates must be integers, It is not necessarily strictly equal to x' / S, so the coordinate offset of the center point will be generated on the feature map F. The predicted center point offset is calculated using a loss function based on a norm. The loss function L offset The formula is:
[0035]
[0036] Among them, j is the real coordinate offset of object j, is the predicted coordinate offset of object j, and N is the number of pedestrians in the image. According to the pixel coordinates of the upper left corner of the bounding box of object j Pixel coordinates of the lower right corner Get the bounding box size s in the input image j , Combined with the center point offset, we can get the bounding box of object j predicted in feature map F. And optimize the feature map F, the optimization goal of object size prediction is:
[0037]
[0038] The true center point of object j on the input image After a series of convolution transformations, it will correspond to a feature point (u, v) on the optimized feature map F. For other points (x", y") on the feature map F, the probability of judging other points (x", y") in the feature map is the center point of the pedestrian
[0039] The probability is calculated as:
[0040]
[0041] in, is the probability that other points (x, y) in the feature map are pedestrian centers, and σ is the standard deviation that can be adjusted based on object size. The closer a point on the feature map F is to (u, v), the greater its probability of being a pedestrian center, and vice versa. The optimization objective for pedestrian probability prediction can be expressed as a logistic regression function with focal loss:
[0042]
[0043] Among them, L cls is the pedestrian center point prediction loss function, N is the total number of pedestrians in the input image, α and β are the hyperparameters of the focalloss function, and p j It is the calculated probability of the pedestrian center point. Therefore, a point (x", y") on the pedestrian detection module feature map F will generate a 1+2+2 dimensional vector, where the first dimension represents the probability that the object is a pedestrian, the middle two dimensions represent the horizontal and vertical offsets of the object center point, and the last two dimensions represent the width and height of the object. These prediction values are generated from a convolutional network with shared structure and parameters. Since there is no need to preset anchors before detection and the IoU of the bounding box is not calculated, the prediction generation only requires the probability value of the heatmap corresponding to the feature map F to perform non-maximum suppression, which improves the running efficiency.
[0044] The goal of learning ID embeddings for pedestrians is to associate the same pedestrian across different images with the same identity. This association is based on the similarity of pedestrian features, so it is necessary to learn the features of different pedestrians. Specifically, a single embedding vector represents pedestrian features. The number of embedding vectors is the same as the number of pedestrians with different identities, and the two correspond one-to-one. This ensures that the similarity between pedestrians with the same identity in different images is as high as possible, while the similarity with other pedestrians is relatively low.
[0045] To improve efficiency, such as Figure 1 As shown, the pedestrian embedding vector learning module and the pedestrian detection module share the same backbone network, and the input image is convolved to obtain a resolution of For each pedestrian center point on the feature map F, 512 3x3 convolution kernels are used to extract a 512-dimensional vector emb. The learning of the feature vector can be regarded as a classification problem: each pedestrian corresponds to a category k = 1, 2, ···, so that each pedestrian j with different identities generates a K-dimensional one-hot encoding h j(k) as its category distribution. Pedestrians j with different numbers may correspond to the same category k, indicating that they are the same person appearing in different pictures. The category distribution of pedestrian j predicted by the model is Pedestrian feature learning optimization vector emb j , try to classify it into the correct category k. The loss function of the learning model can be expressed as:
[0046]
[0047] S3 performs 3D positioning of pedestrians based on the pedestrian detection results.
[0048] After the camera tracks a pedestrian, it is necessary to accurately determine the pedestrian's position in space. Spatial position is described by a 3D vector, so a spatial rectangular coordinate system needs to be established. Since the camera has its own 3D spatial coordinate system, to simplify the problem, this embodiment directly calculates the pedestrian's position in the camera coordinate system. Assume that the center point of the pedestrian's bounding box A is projected onto the image plane at point A'. Based on the geometric relationship in the monocular camera imaging model, the following equation can be obtained:
[0049]
[0050]
[0051]
[0052] The height H of the pedestrian in the image is obtained by taking the Gaussian distribution N(μ,σ 2 ) is the mean μ. For a given monocular camera, its internal parameter matrix (f x 0 c x ,0 f y c y ,0 0 1) is a known quantity. Among them, f x 、f y Satisfies the following relationship:
[0053]
[0054]
[0055] The pedestrian is located in 3D by projecting the bounding box of the pedestrian in the image into the 3D coordinates of the camera. The positioning result is calculated by the following formula:
[0056]
[0057] Among them, loc is the position of the pedestrian in the three-dimensional coordinates of the camera; x c 、y c and z care the x-axis, y-axis, and z-axis coordinate values of the pedestrian’s center point in the camera’s three-dimensional coordinates; f x and f y are the focal lengths of the camera in the x and y directions, respectively, H is the height of the pedestrian in the image, and h is the height of the pedestrian in the camera.
[0058] S4 combines pedestrian detection and ID embedding learning results to re-identify pedestrians.
[0059] The re-identification method is:
[0060] S4.1 If the pedestrian’s ID embedding vector at time t is represented by emb i , the ID embedding vector of the pedestrian after learning at time t+1 is represented as emb j , then the cosine similarity between the two is cos <emb i ,emb j >
[0061] S4.2 If the bounding box of the pedestrian at time t is box i , the bounding box of the pedestrian at time t+1 is box j , then the intersection-over-union ratio of the two is IoU <box i ,box j >
[0062] S4.3 Determine cos <emb i ,emb j > and IoU <box i ,box j >Whether the preset conditions are met, if both values meet the preset conditions, proceed to the next step;
[0063] S4.4 Based on the result at time t, the Kalman filter is used to predict the bounding box position of the pedestrian at time t+1. Get the bounding box position of the pedestrian at time t+1 And calculate the distance between the two If the distance also meets the preset conditions, the pedestrians appearing in the images at time t and time t+1 are considered to match.
[0064] S5 determines the motion trajectory of each pedestrian based on the re-identification results and the pedestrian's 3D position, and thus tracks it in three dimensions.
[0065] The method for tracking pedestrians is as follows: Mono3DMOT is used to obtain the 3D tracking results of pedestrians based on pedestrian re-identification. Pedestrian re-identification is used to associate the same pedestrians in two consecutive images within a time interval. The position of pedestrian i=1,2,···,N at each time t is Thus, the movement trajectory of pedestrian i in the time period interval is obtained Usually the pedestrian trajectory is continuous. To prevent large local oscillations, a Kalman filter is used to filter the trace. i Forming a smooth 3D motion trajectory.
[0066] The parameters in this embodiment are shown in Table 1.
[0067] Table 1 Meaning and status of each variable in the method of this embodiment
[0068]
[0069]
[0070] Example 2
[0071] Based on the same inventive concept, this embodiment discloses a multi-pedestrian three-dimensional tracking system based on monocular vision, comprising:
[0072] A feature extraction module is used to extract features from a training image set obtained by a monocular camera;
[0073] The pedestrian detection and ID learning module is used to simultaneously detect pedestrians and embed the corresponding pedestrian IDs based on the extracted features; the 3D positioning module is used to locate pedestrians in 3D based on the pedestrian detection results;
[0074] The re-identification module is used to re-identify pedestrian detection results based on the ID learning results;
[0075] The pedestrian 3D tracking module is used to determine the motion trajectory of each pedestrian based on the re-identification results and the pedestrian's position, and thus track them.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific embodiments of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be included in the scope of protection of the claims of the present invention. The above content is only a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in the present application, which should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A three-dimensional tracking method for multiple pedestrians based on monocular vision, characterized in that: include: S1 performs feature extraction on the training image set obtained by the monocular camera; S2 simultaneously performs pedestrian detection and ID embedding learning of the corresponding pedestrian based on the extracted features; S3 performs 3D positioning of pedestrians based on pedestrian detection results; S4 re-identifies the pedestrian based on the result of the ID embedding learning; S5 determines the movement trajectory of each pedestrian based on the re-identification results and the pedestrian's 3D position, and thus performs three-dimensional tracking; The method for extracting features from the training image set in step S1 is as follows: in the training image set, the category of the object j is marked as k, and the pixel coordinates of the upper left corner of the bounding box of the object j are ( , ), pixel coordinates of the lower right corner ( , ), input the fixed-size image into the convolutional neural network, extract features through the convolution operation, and obtain a feature map F of size W / S×H / S, where W is the width of object j, H is the height of object j, and S is the downsampling step size; The method for pedestrian detection in step S2 is as follows: matching the pixel points in the input image with the pixel points in the feature map F, calculating the center point offset of the feature map F by using the loss function; and calculating the pixel coordinates of the upper left corner of the bounding box of the object j according to the pixel coordinates of the upper left corner of the bounding box of the object j ( , ), pixel coordinates of the lower right corner ( , ) obtain the bounding box size in the input image; obtain a feature point (x', y') on the feature map F through convolution transformation, and determine the probability that other points (x, y) in the feature map F are the center points of the pedestrian; The ID embedding learning method corresponding to the pedestrian in step S2 is: using an embedding vector to represent the pedestrian's features, the number of the embedding vectors is the same as the number of pedestrians with different identities. For pedestrians with the same identity but appearing in different pictures, their embedding vectors are the same or similar, and the embedding vectors of pedestrians with different identities are different.
2. The method for three-dimensional tracking of multiple pedestrians based on monocular vision according to claim 1, wherein: The probability is calculated as follows: in, p is the probability that other points (x, y) in the feature map are the center points of pedestrians, is the standard deviation that can be adjusted according to the size of the object, and They are the physical width and height of the camera's photosensitive element respectively.
3. The method for three-dimensional tracking of multiple pedestrians based on monocular vision according to claim 2, wherein: The pedestrian center point probability prediction is optimized by using a loss function with focal loss, and its calculation formula is: in, is the pedestrian center point probability prediction loss function, N is the total number of pedestrians in the input image, α and β are hyperparameters in focalloss, is the probability that the jth point is the pedestrian center.
4. The method for three-dimensional tracking of multiple pedestrians based on monocular vision according to claim 2 or 3, characterized in that: The method for 3D positioning of pedestrians in step S3 is as follows: assuming that the height of pedestrians is Gaussian distributed, the bounding box of the pedestrian in the image is projected into the three-dimensional coordinate system of the camera as the 3D position of the pedestrian, and the 3D position is calculated by the following formula: in, loc is the position of the pedestrian in the 3D camera coordinate system; 、 and are the x-axis, y-axis, and z-axis coordinate values of the pedestrian's center point in the three-dimensional camera coordinate system; are the x- and y-direction coordinates of the camera principal point, H is the physical height of the pedestrian, and h is the pixel height of the pedestrian in the image. , is the coordinate of the pixel center point i in the image plane coordinate system, , is the coordinate of the pixel center point i in the image plane coordinate system, and the height H of the pedestrian in the image is obtained by taking the Gaussian distribution N(µ,σ 2 ) is the mean µ.
5. The method for three-dimensional tracking of multiple pedestrians based on monocular vision according to claim 2 or 3, wherein: The re-identification method in step S4 is: S4.1 If the ID embedding vector of pedestrian i at time t is expressed as emb i , the ID embedding vector of pedestrian j at time t+1 is expressed as emb j , then the cosine similarity between the two is cos <emb i ,emb j > S4.2 If the bounding box of pedestrian i at time t is box i , the bounding box of pedestrian j at time t+1 is box j , then the intersection-over-union ratio of the two is IoU <box i ,box j > S4.3 Determine cos <emb i ,emb j > and IoU <box i ,box j >Whether the preset threshold conditions are met, if both values meet the preset conditions, proceed to the next step; S4.4 Based on the result at time t, use the Kalman filter to predict the bounding box position of pedestrian k at time t+1 , detect the bounding box position of pedestrian j at time t+1 , and calculate the distance between the two If the distance also meets the threshold condition, it is considered that pedestrian i at time t and pedestrian j at time t+1 match.
6. The method for three-dimensional tracking of multiple pedestrians based on monocular vision according to claim 5, wherein: The method for three-dimensionally tracking pedestrians in step S5 is: based on the pedestrian re-identification module, continuously associate the same pedestrian in the time interval interval={t, t+1, …, t+n}, and the position of pedestrian i=1,2,···,N at each time t is , thereby obtaining the movement trajectory of pedestrian i in the time interval .
7. A three-dimensional multi-pedestrian tracking system based on monocular vision, characterized in that: include: A feature extraction module is used to extract features from a training image set obtained by a monocular camera; Pedestrian detection and ID embedding learning module, which is used to simultaneously detect pedestrians and learn the corresponding pedestrian ID embedding based on the extracted features; The 3D positioning module uses the Gaussian distribution assumption of pedestrian height and combines the intermediate results of pedestrian detection to calculate the three-dimensional position of the pedestrian in the camera coordinate system; The re-identification module is used to re-identify pedestrian detection results based on the ID learning results; The 3D pedestrian tracking module determines the motion trajectory of each pedestrian in continuous time based on the pedestrian ID embedding representation and 3D position, thereby achieving 3D tracking; The method for extracting features from the training image set in the feature extraction module is as follows: in the training image set, the category of the object j is marked as k, the pixel coordinates of the upper left corner of the bounding box of the object j are ( , ), pixel coordinates of the lower right corner ( , ), input the fixed-size image into the convolutional neural network, extract features through the convolution operation, and obtain a feature map F of size W / S×H / S, where W is the width of object j, H is the height of object j, and S is the downsampling step size; The method for pedestrian detection in the pedestrian detection and ID embedding learning module is as follows: the pixel points in the input image are matched with the pixel points in the feature map F, and the center point offset of the feature map F is calculated by the loss function; the pixel coordinates of the upper left corner point of the bounding box of the object j are calculated according to the pixel coordinates of the upper left corner point ( , ), pixel coordinates of the lower right corner ( , ) obtain the bounding box size in the input image; obtain a feature point (x', y') on the feature map F through convolution transformation, and determine the probability that other points (x, y) in the feature map F are the center points of the pedestrian; The ID embedding learning method corresponding to the pedestrian in the pedestrian detection and ID embedding learning module is: use an embedding vector to represent the pedestrian features, and the number of the embedding vectors is the same as the number of pedestrians with different identities. For pedestrians with the same identity but appearing in different pictures, their embedding vectors are the same or similar, and the embedding vectors of pedestrians with different identities are different.
Citation Information
Patent Citations
Deep learning-based cross-camera pedestrian multi-target tracking method and device
CN112270310A
Vehicle-mounted terminal multi-target identification tracking prediction method
CN112307921A