A target tracking method applicable to vehicle-mounted environments

By simulating the opposing geometric model of binocular camera imaging and extracting feature information using convolutional neural networks, the problem of poor multi-objective tracking in the on-board environment is solved, and high-precision target matching is achieved.

CN114596333BActive Publication Date: 2025-07-22BEIJING HUAHANG RADIO MEASUREMENT & RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011396578.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-03
Publication Date
2025-07-22
Estimated Expiration
2040-12-03

AI Technical Summary

Technical Problem

The existing multi-objective tracking methods have poor results in on-board environments, especially when camera movement is not effective.

Method used

The two-frame data of the front and rear frames of monocular camera imaging are used to simulate the opposing geometric model of binocular camera imaging, combined with the convolutional neural network to extract the appearance feature information and motion information of the target, predict the position of the target in the current frame by calculating the essence matrix, and use the Marshallow distance and the appearance feature cosine distance for target matching.

Benefits of technology

It improves the multi-object tracking accuracy in the on-board environment, reduces the impact of camera motion on tracking, and achieves effective multi-object matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596333B_ABST
    Figure CN114596333B_ABST
Patent Text Reader

Abstract

The present invention discloses a target tracking method applicable to vehicle-mounted environments. The method uses a target detection algorithm to obtain the target detection results Pre_Detections and Cur_Detections in the previous frame and the current frame images, and extracts appearance feature vectors; calculates the essential matrix F of the epipolar geometry model for the two consecutive frame images; predicts its position G_Tracks in the current frame image according to Pre_Detections by using the essential matrix F; calculates the Mahalanobis distance correlation metric matrix corresponding to all targets in G_Tracks and Cur_Detections, and at the same time calculates the appearance feature cosine distance correlation metric matrix corresponding to all targets in Pre_Features and Cur_Features, and uses the metric matrix to make a decision on the correlation information between the targets in the front and rear frames, so as to achieve multi-target tracking. The design of the present invention can extract the appearance feature information of the target through a convolutional neural network, and at the same time calculate the motion information of the target in the video, and combine the two to achieve multi-target association between frames in the video, so as to achieve the effect of target tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an object tracking method applicable to a vehicle-mounted environment. Background Art

[0002] In recent years, the technical field of computer vision has developed rapidly. Among them, the multi-object tracking direction has gradually become a major research hotspot. Currently, the method with higher attention in the industrial field is DeepSort. The solution of this method is mainly divided into the following steps: 1. Use object detection methods such as YOLO series, SSD series, Faster-RCNN, etc. to detect the initial frame of the video and obtain the object detection boxes; 2. Extract the corresponding target regions in the image frame and then extract the feature information (appearance feature information, motion feature information); 3. Calculate the similarity between the objects in the front and back image frames; 4. Associate the objects matched in the two frames and assign IDs.

[0003] The above method is widely used in the field of security monitoring. In this application scenario, the camera is fixed in position, and most objects move linearly and uniformly, so the multi-object tracking effect is better. However, in an environment where the camera moves (such as a vehicle-mounted environment), its tracking and association effect is relatively poor. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an object tracking method applicable to a vehicle-mounted environment.

[0005] To solve the above technical problem, the technical solution adopted by the present invention is as follows:

[0006] Step S1: Use an object detection algorithm to obtain the object detection results Pre_Detections in the previous frame image and the object detection results Cur_Detections in the current frame image, and extract the appearance feature vectors Pre_Features and Cur_Features of the object detection results;

[0007] Step S2: Calculate the essential matrix F of the epipolar geometry model of the front and back two-frame images, and initialize the Kalman filter;

[0008] Step S3: According to Pre_Detections, predict its position G_Tracks in the current frame image by using the essential matrix F;

[0009] Step S4: Calculate the Mahalanobis distance association metric matrix corresponding to all objects in G_Tracks and Cur_Detections, and at the same time calculate the appearance feature cosine distance association metric matrix corresponding to all objects in Pre_Features and Cur_Features;

[0010] Step S5: Filter out the target pairs with too large distance values in the Mahalanobis distance correlation metric matrix, and then take the target pair with the smallest cosine distance among the remaining target pairs as the target matching pair, so as to achieve multi-target tracking.

[0011] Further, step S2 for calculating the essential matrix F of the epipolar geometry model of two consecutive frames of images includes the following steps:

[0012] Step S201: Extract the ORB or SIFT key point features of two consecutive frames of images;

[0013] Step S202: Use the Brute Force algorithm for key point matching to obtain matching point pairs;

[0014] Step S203: Use the findFundamentalMat function in opencv to calculate the essential matrix F between the above-mentioned matching point pairs.

[0015] Further, step S3 includes the following steps:

[0016] Step S301: Use the four-corner coordinates of Pre_Detections obtained in step S1 as the input. According to the coordinate transformation relationship of epipolar geometry in the binocular camera coordinate system: R T FL = 0, assuming GL = 0, calculate G = R T F,

[0017] where, x1, y1 are the upper left coordinates of Pre_Detections in the image; w, h are the width and height of the target; L is the three-dimensional vector coordinate of a certain point in the previous frame of image; R is the three-dimensional vector coordinate of a certain point corresponding to the current frame of image; F is the essential matrix between the two images; G is a 4x3 matrix;

[0018] Step S302: According to the G matrix calculated in step S301, assuming that the target scale remains unchanged, that is, the width and height (w, h) remain unchanged, expand and deduce GL = 0 as:

[0019]

[0020] where x, y are the upper left coordinates of the prediction result G_Tracks in the current frame image, and the coordinates of G_Tracks are calculated therefrom.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0022] 1. Extract the apparent feature information of the target through the convolutional neural network, and at the same time calculate the motion information of the target in the video. Combine the two to achieve multi-target association between frames in the video, so as to achieve the effect of target tracking;

[0023] 2. A method is proposed to simulate the epipolar geometry model in binocular camera imaging by using the data of two consecutive frames of monocular camera imaging, which reduces the influence of camera movement on target prediction and tracking in the vehicle environment and improves the multi-target tracking accuracy.

[0024] 3. According to the target detection results of the previous frame image, the position of the target in the current frame image is innovatively deduced by using the essential matrix, thus effectively ensuring the matching and tracking of multiple targets. Brief Description of the Drawings

[0025] Figure 1 It is a flowchart of the multi-target tracking method provided by the embodiment of the present invention. Detailed Embodiment

[0026] In the vehicle environment, the camera is also moving during the imaging process. Therefore, when using the DeepSort algorithm for multi-target tracking, there will be a problem of poor tracking effect when the vehicle speed is slightly faster. Based on this method, the present invention proposes a multi-target tracking method applicable to the vehicle environment, which uses the data of two consecutive frames of monocular camera imaging to simulate the epipolar geometry model in binocular camera imaging, reduces the interference effect of camera movement on the motion model, and thus improves the multi-target tracking effect in the vehicle environment.

[0027] The present invention will be further described below with reference to the drawings and embodiments.

[0028] As Figure 1 shown, the embodiment of the present invention provides a multi-target tracking method applicable to the vehicle environment, including the following steps:

[0029] Step S1: Use the YOLOv3 target detection algorithm to obtain the target detection results (Pre_Detections and Cur_Detections) in the previous frame and the current frame images, and extract the appearance feature vectors (Pre_Features and Cur_Features) of the target detection results.

[0030] Step S101: Use the YOLO series, SSD series or Faster-RCNN target detection algorithm to perform target detection on the previous frame and the current frame images to obtain the Pre_Detections and Cur_Detections results.

[0031] Step S102: Corresponding to the positions of Pre_Detections and Cur_Detections in the front and rear frame images, intercept the target slice images and normalize them to the same size.

[0032] Step S103: Input the normalized sliced images into the cosine metric learning network for appearance feature extraction to obtain appearance feature vectors Pre_Features and Cur_Features.

[0033] Step S2: Calculate the essential matrix F of the epipolar geometry model for the front and rear frame images;

[0034] Step S201: Extract the ORB or SIFT key point features of the front and rear frame images;

[0035] Step S202: Use the Brute Force algorithm for key point matching to obtain matching point pairs;

[0036] Step S203: Use the findFundamentalMat function in opencv to calculate the essential matrix F between the above matching point pairs.

[0037] Step S3: According to Pre_Detections, predict its position (G_Tracks) in the current frame image using the essential matrix F.

[0038] Step S301: Use the four corner coordinates of Pre_Detections obtained in Step S1 as the input. Utilize the principle that the front and rear frame images in a monocular camera can approximately simulate the left and right images in a binocular camera. According to the coordinate transformation relationship of epipolar geometry in the binocular camera coordinate system: R T FL = 0 (where L is the three-dimensional vector coordinate of a point in the previous frame image, R is the three-dimensional vector coordinate of the corresponding point in the current frame image, and F is the essential matrix between the two images). Assume GL = 0 (G is a 4x3 matrix), and calculate G = R T F, where (x1, y1 are the upper left coordinates of Pre_Detections in the image, and w, h are the width and height of the target).

[0039] Step S302: According to the G matrix calculated in the previous step, expand GL = 0 to obtain the following formula:

[0040]

[0041] where x, y are the upper left coordinates of the prediction result G_Tracks in the current frame image. Assuming the target scale remains unchanged, i.e., the width and height (w, h) remain unchanged, the above formula can be deduced as:

[0042]

[0043] The coordinates of G_Tracks can be calculated.

[0044] Step S4: Calculate the Mahalanobis distance correlation metric matrix for all targets corresponding to G_Tracks and Cur_Detections, and at the same time calculate the cosine distance correlation metric matrix for all targets corresponding to Pre_Features and Cur_Features.

[0045] Step S401: Take all the targets in G_Tracks and Cur_Detections in (where C x , C y is the center point of the target, a is the aspect ratio of the target, h is the height of the target, and the others are the change rates of the corresponding variables) to calculate the covariance matrix S between the tracking target and the detection target.

[0046] Step S402: Use the covariance matrix to perform normalization calculation on the Mahalanobis distance correlation matrix:

[0047]

[0048] Obtain the Mahalanobis distance metric value between the i-th G_Tracks target and the j-th Cur_Detections target.

[0049] Step S403: Calculate the cosine distance correlation metric matrix for the appearance features between all targets of Pre_Features and Cur_Features:

[0050] d cosine (i, j) = Cur_Features j T ·Pre_Features i

[0051] Step S5: Filter out the target pairs with too large distance values in the Mahalanobis distance metric matrix, and then take the pair with the smallest cosine distance among the remaining target pairs as the target matching pair.

[0052] Step S501: Judge whether the Mahalanobis distance metric value is greater than the threshold Th ma . If it is greater, then judge that the i-th G_Tracks target and the j-th Cur_Detections target are unpaired; if it is less, then go to the next step.

[0053] Step S502: For the target pairs that meet the above Mahalanobis distance metric conditions, compare the cosine distances between all Cur_Detections targets corresponding to the same G_Tracks target, find the smallest cosine metric value, and judge whether it is less than Th cosine . If it is less, then judge it as a matching pair, otherwise it is a non-matching pair.

Claims

1. A multi-object tracking method applicable to vehicle-mounted environments, characterized in that, It includes the following steps: Step S1: Use an object detection algorithm to obtain the object detection results Pre_Detections in the previous frame image and the object detection results Cur_Detections in the current frame image, and extract the appearance feature vectors Pre_Features and Cur_Features of the object detection results; Step S2: Calculate the essential matrix F of the epipolar geometry model for the two consecutive frame images, and initialize the Kalman filter; Step S3: According to Pre_Detections, use the essential matrix F to predict its position G_Tracks in the current frame image; Step S4: Calculate the Mahalanobis distance association metric matrix corresponding to all objects in G_Tracks and Cur_Detections, and at the same time calculate the appearance feature cosine distance association metric matrix corresponding to all objects in Pre_Features and Cur_Features; Step S5: Filter out the object pairs with too large distance values in the Mahalanobis distance association metric matrix, and then take the object pair with the smallest cosine distance among the remaining object pairs as the object matching pair, so as to achieve multi-object tracking. Step S3 includes the following steps: Step S301: Using the four-corner coordinates of Pre_Detections obtained in Step S1 as input, according to the coordinate transformation relationship of the epipolar geometry in the binocular camera coordinate system: R T FL = 0, assuming GL = 0, calculate G = R T F, Among them, x1 and y1 are the upper left corner coordinates of Pre_Detections in the image; w and h are the width and height of the target; L is the three-dimensional vector coordinate of a certain point in the previous frame image; R is the three-dimensional vector coordinate of the corresponding point in the current frame image; F is the essential matrix between the two images; G is a 4x3 matrix; Step S302: According to the G matrix calculated in Step S301, assuming that the object scale remains unchanged, that is, the width and height (w, h) remain unchanged, expand GL = 0 and deduce it as: where x and y are the upper left coordinates of the prediction result G_Tracks in the current frame image, and the coordinates of G_Tracks are calculated therefrom.

2. The multi-target tracking method applicable to a vehicle-mounted environment according to claim 1, wherein Step S2 for calculating the essential matrix F of the epipolar geometry model for the two consecutive frame images includes the following steps: Step S201: Extract the ORB or SIFT key point features of the two consecutive frame images; Step S202: Use the Brute Force algorithm to perform key point matching to obtain the matching point pairs; Step S203: Use the findFundamentalMat function in opencv to calculate the essential matrix F between the above-mentioned matching point pairs.

Citation Information

Patent Citations

  • Object recognizing apparatus and object recognizing method using epipolar geometry

    WO2006090735A1

  • Method and apparatus for detecting moving target, and electronic device and storage medium

    WO2020156341A1