A multi-level matching video racing car tracking method and system

Through the multi-level matching video racing tracking method, combining sports features and appearance features, the problem that the existing technology cannot take into account both sports features and appearance features is solved, and the accurate tracking of special lenses in video racing and the complete acquisition of vehicle trajectory is achieved.

CN114972410BActive Publication Date: 2025-06-10HUNAN XINGLAN ZHIYUAN NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210682158.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-16
Publication Date
2025-06-10
Estimated Expiration
2042-06-16

AI Technical Summary

Technical Problem

The existing video racing target tracking methods cannot take into account both the sports and appearance characteristics, resulting in the complete trajectory of the vehicle being unable to accurately obtain the situation of certain special lenses such as long-term occlusion and excessive target movement position.

Method used

The multi-level matching video racing tracking method is adopted to detect the target vehicle frame by frame through the pre-trained target detection model, combine the vehicle's motion characteristics and appearance characteristics, and perform multi-level matching based on the confidence of the target detection frame to complete vehicle tracking.

Benefits of technology

It realizes accurate tracking of special lenses in video racing, taking into account both sports and appearance characteristics, and can effectively deal with complex situations such as long-term occlusion and large target moving position to obtain the complete trajectory of the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972410B_ABST
    Figure CN114972410B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-level matching video racing car tracking method and system, which uses a pre-trained object detection model to perform object vehicle detection on each frame of a racing car video to obtain a detection result, and the result includes an object detection box and a detection result confidence level; inputs the image of the object detection box area extracted into a secondary network to extract the vehicle appearance feature; combines the motion feature and the appearance feature of the vehicle, and performs multi-level matching according to the object detection box confidence level to complete vehicle tracking, and obtains the association result between the objects in each frame of the video. By taking into account both the motion feature and the appearance feature, the tracking of some special shots can be realized, such as long-term occlusion, too large position movement of the object between two frames, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target tracking, and particularly to a multi-level matching video racing car tracking method and system. Background Art

[0002] With the development of artificial intelligence technology, more and more technologies are applied to video processing. In previous racing competitions, we sometimes need to extract exciting shots of different shot types for highlights. The usual method is to manually obtain fragments of certain vehicles, which will take a lot of time. We hope to use technical means to obtain the trajectories of each vehicle in the video and achieve automatic extraction. Most of the existing technical paths are to detect vehicles through object detection methods and determine the positions of a certain vehicle in the video segment in combination with the results of OCR. However, due to the rapid movement of racing cars, the switching of camera angles, and the mutual occlusion of vehicles, many license plates are blocked, resulting in the inability to accurately obtain the complete trajectory of vehicles within a shot. Existing object tracking methods usually cannot take into account both motion features and appearance features to achieve tracking of certain special shots, such as long-term occlusion, too large a position change of the target between two frames, etc. These are all problems that need to be solved in video racing. Summary of the Invention

[0003] Therefore, the present invention provides a multi-level matching video racing car tracking method and system to solve the problem that the existing video racing car target tracking method cannot take into account both motion features and appearance features, and for certain special shots such as long-term occlusion, too large a position change of the target between two frames, etc., it is impossible to accurately obtain the complete trajectory of vehicles within a shot.

[0004] To achieve the above object, the present invention provides the following technical solutions:

[0005] According to the first aspect of the embodiments of the present invention, a multi-level matching video racing car tracking method is proposed, characterized in that the method includes:

[0006] Using a pre-trained object detection model to detect target vehicles frame by frame in a racing car video to obtain detection results, the results including target detection boxes and detection result confidence levels;

[0007] Inputting the extracted target detection box region images into a secondary network to extract vehicle appearance features;

[0008] Combining the motion features and appearance features of the vehicle, and performing multi-level matching according to the target detection box confidence level to complete vehicle tracking, and obtaining the association results between targets in each frame of the video.

[0009] Further, by combining the motion characteristics and appearance characteristics of the vehicle, and performing multi-level matching according to the confidence of the target detection box to complete vehicle tracking, the association results between the targets in each frame of the video are obtained, specifically including:

[0010] The detection results are divided into high-score boxes and low-score boxes by setting a confidence threshold; for the created target tracking trajectories, first match among the high-score boxes. If no match is found, then use the low-score boxes and the tracking trajectories that have not matched the high-score boxes for matching.

[0011] Further, by combining the motion characteristics and appearance characteristics of the vehicle, and performing multi-level matching according to the confidence of the target detection box to complete vehicle tracking, the association results between the targets in each frame of the video are obtained, specifically including:

[0012] For the high-score boxes that have not matched the tracking trajectories and have a high enough score, a new tracking trajectory is created for them.

[0013] Further, by combining the motion characteristics and appearance characteristics of the vehicle, and performing multi-level matching according to the confidence of the target detection box to complete vehicle tracking, the association results between the targets in each frame of the video are obtained, specifically including:

[0014] For the tracking trajectories that have not matched the detection boxes, retain them for multiple consecutive frames until the target appears again for matching.

[0015] Further, by combining the motion characteristics and appearance characteristics of the vehicle, and performing multi-level matching according to the confidence of the target detection box to complete vehicle tracking, the association results between the targets in each frame of the video are obtained, specifically including:

[0016] For the target detection results of the current frame, a detection box for an adjacent frame is predicted through Kalman filtering;

[0017] According to the detection results and prediction box results of the target, the Mahalanobis distance is calculated based on the motion characteristics to obtain the spatial position difference; and the cosine distance is calculated based on the appearance characteristics of the targets in different frames to obtain the appearance similarity;

[0018] The calculated Mahalanobis distance and cosine distance are weighted and summed to obtain a cost matrix, which is matched through the Hungarian algorithm. The matching items that do not meet the Mahalanobis distance threshold are set to infinity and then removed, and multi-object cascade matching is performed on the results of each frame. Finally, the association results between the targets in each frame of the video are obtained.

[0019] Further, the method further includes training the target detection model, specifically:

[0020] Select video segments containing different racing car models for equidistant frame extraction, label the minimum bounding rectangle of each racing car for each extracted frame, construct a training set, and use the training set to train the model.

[0021] Further, the object detection model uses the YOLOX network model.

[0022] Further, the method further includes: adding a secondary network to the output head of the YOLOX backbone network, and extracting appearance features of the obtained object detection region through the secondary network.

[0023] According to the second aspect of the embodiments of the present invention, a multi-level matching video racing car tracking system is proposed. The system includes:

[0024] An object detection module, configured to use a pre-trained object detection model to perform object vehicle detection on each frame of a racing car video to obtain a detection result, where the result includes an object detection box and a detection result confidence level;

[0025] An appearance feature extraction module, configured to input the image of the extracted object detection box region into a secondary network to extract vehicle appearance features;

[0026] A vehicle tracking module, configured to combine the motion features and appearance features of the vehicle, and perform multi-level matching according to the object detection box confidence level to complete vehicle tracking, and obtain the association result between the objects in each frame of the video.

[0027] The present invention has the following advantages:

[0028] A multi-level matching video racing car tracking method and system proposed by the present invention uses a pre-trained object detection model to perform object vehicle detection on each frame of a racing car video to obtain a detection result, where the result includes an object detection box and a detection result confidence level; inputs the image of the extracted object detection box region into a secondary network to extract vehicle appearance features; combines the motion features and appearance features of the vehicle, and performs multi-level matching according to the object detection box confidence level to complete vehicle tracking, and obtains the association result between the objects in each frame of the video. Taking into account both motion features and appearance features, it can achieve tracking of some special shots, such as long-term occlusion, too large a position movement of the object between two frames, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained according to the provided drawings.

[0030] Figure 1 It is a schematic flowchart of a multi-level matching video racing car tracking method provided in Embodiment 1 of the present invention;

[0031] Figure 2 Schematic diagram of the specific implementation process of a multi-level matching video racing car tracking method provided in Embodiment 1 of the present invention;

[0032] Figure 3 Schematic diagram of the vehicle appearance extraction network in a multi-level matching video racing car tracking method provided in Embodiment 1 of the present invention;

[0033] Figure 4 Schematic diagram of the steps of cascade matching in a multi-level matching video racing car tracking method provided in Embodiment 1 of the present invention. Detailed implementation mode

[0034] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0035] Embodiment 1

[0036] As Figure 1 shown, this embodiment proposes a multi-level matching video racing car tracking method, and the method includes:

[0037] S100. Use a pre-trained object detection model to perform frame-by-frame detection on the racing car video to obtain detection results, and the results include target detection frames and detection result confidence levels.

[0038] S200. Input the extracted target detection frame area image into a secondary network to extract vehicle appearance features.

[0039] S300. Combine the motion features and appearance features of the vehicle, and perform multi-level matching according to the target detection frame confidence level to complete vehicle tracking, and obtain the association results between targets in each frame of the video.

[0040] The specific implementation process is as follows. Refer to Figure 2 :

[0041] 1. Vehicle detection

[0042] 1) Construction of the racing car data set

[0043] Select video segments containing different racing car models for equidistant frame extraction, and label the minimum bounding rectangle of each racing car in each extracted frame. The category is uniformly set to one category, and the number of labels is about 2000 frames. Construct a training set and use the training set to train the model.

[0044] 2) Model training and inference

[0045] Select yolox as the detection model, and use the labeled data to train a racing car detection model. Use this model to detect each frame in the video and output the detection results. For each frame, obtain the four parameters (x0, y0, w, h) of the vehicle position and the confidence conf of the detection result, and record the frame number frame of each vehicle. After obtaining the results of consecutive frames, the next step is to track the detection results.

[0046] 2. Vehicle tracking:

[0047] 1) Initialization: According to the detection results of the first frame, create an initialized tracker (tracks), and predict the detection boxes of adjacent frames through Kalman filtering. And determine the state of tracks.

[0048] 2) Calculation of Mahalanobis distance: The Mahalanobis distance uses motion features, that is, the spatial position information of the target between different frames. The Mahalanobis distance takes into account the uncertainty of state measurement by calculating the standard deviation between the detection position and the average tracking position, and reflects the difference in spatial position through the Mahalanobis distance. The formula for calculating the Mahalanobis distance similarity metric is as follows:

[0049]

[0050] d j represents the position of the j-th detection box; y i represents the predicted position of the target by the i-th tracker, and S i represents the covariance matrix between the detection box and the predicted box.

[0051] 3) Extraction of appearance information and calculation of similarity: Considering the fast movement of the racing cars in the video, the movement gap between two consecutive frames is often large. Relying solely on the basis of motion distance matching often fails to achieve ideal results. Especially when different vehicles cross each other, pure motion features often fail to achieve reasonable matching. By extracting the target detection area and using a lightweight secondary network to extract appearance features, the reid features are obtained.

[0052] In this embodiment, a secondary network is added to the output head of the yolox backbone network, and the secondary network is used to extract appearance features from the obtained target detection area. The structure of the feature extraction network is as Figure 3 shown. The input of the network is the target detection result area, and the output is a 1×512 feature vector. The formula for calculating the cosine distance of the apparent feature is as follows

[0053]

[0054] Among them, r j corresponds to the feature vector of the j-th detection, and the feature vector for tracking. The minimum cosine distance between all the feature vectors of the i-th object tracking and the j-th object is calculated by this formula. This distance represents the appearance similarity of the target between different frames.

[0055] 4) Target confidence level classification: Considering that the confidence levels within the same range in the detection results are more highly correlated, the detection results are divided into high-score boxes and low-score boxes by setting a confidence threshold. First, matching is performed among the high-score boxes. Second, the low-score boxes are used to match the tracking trajectories that did not match the high-score boxes in the first time (for example, objects whose scores have dropped due to severe occlusion in the current frame). For the detection boxes that did not match the tracking trajectories but have a high enough score, we create a new tracking trajectory for them. For the tracking trajectories that did not match the detection boxes, we will retain them for 30 frames and perform matching again when they appear again.

[0056] 5) Cascade matching:

[0057] Calculate the Mahalanobis distance of the motion features. Through the gating matrix, set the matching items that do not meet the Mahalanobis distance threshold to infinity to obtain result B;

[0058] The cosine distance and Mahalanobis distance of reid are used to obtain the cost matrix, denoted as C, and its calculation formula is as follows:

[0059] c i,j =λd (1)(i,j) +(1 - λ)d (2) (i,j)

[0060] According to the update status of the prediction box (here, the update status means the time since this prediction box was successfully matched last time), the newer the prediction box (that is, the shorter the number of frames since it was last matched), the more preferentially perform matching according to the result of C using the Hungarian algorithm. Finally, divide the set of matched and unmatched sets according to the result in B. Perform multi-object cascade matching on the results of each frame. Through matching, the numbers of the target vehicles in each frame can be obtained. Vehicles with the same number are classified into the same tracking trajectory, and thus the association results between the targets in each frame of the entire video series can be obtained. The specific steps of cascade matching are as Figure 4 shown.

[0061] Embodiment 2

[0062] Corresponding to the above Embodiment 1, this embodiment proposes a multi-level matching video racing car tracking system, and the system includes:

[0063] A target detection module, configured to perform frame-by-frame target vehicle detection on a racing car video using a pre-trained target detection model, and obtain detection results, where the results include target detection frames and detection result confidence levels;

[0064] An appearance feature extraction module, configured to input the extracted target detection frame area image into a secondary network to extract vehicle appearance features;

[0065] A vehicle tracking module, configured to combine the motion features and appearance features of a vehicle, and perform multi-level matching according to the target detection frame confidence level to complete vehicle tracking, and obtain the association results between targets in each frame of the video.

[0066] The functions performed by each component in a multi-level matching video racing car tracking system provided by an embodiment of the present invention have been described in detail in the above-mentioned Embodiment 1, and thus will not be elaborated here.

[0067] Although the present invention has been described in detail with general descriptions and specific embodiments above, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection of the present invention.

Claims

1. A multi-level matching video racing car tracking method, characterized in that, the method includes: Using a pre-trained object detection model to perform object vehicle detection frame by frame on the racing car video to obtain detection results, and the results include object detection frames and detection result confidence levels; Inputting the image of the object detection frame area extracted into a secondary network to extract the vehicle appearance features; Combining the motion features and appearance features of the vehicle, and performing multi-level matching according to the object detection frame confidence level to complete vehicle tracking, and obtaining the association results between the targets of each frame of the video; Combining the motion features and appearance features of the vehicle, and performing multi-level matching according to the object detection frame confidence level to complete vehicle tracking, and obtaining the association results between the targets of each frame of the video, specifically including: For the object detection results of the current frame, predicting through Kalman filtering to obtain a detection frame of an adjacent frame; According to the object detection results and the predicted frame results, calculating the Mahalanobis distance based on the motion features to obtain the spatial position difference; and calculating the cosine distance according to the appearance features of the objects in different frames to obtain the appearance similarity; Performing weighted summation on the calculated Mahalanobis distance and cosine distance to obtain a cost matrix, performing matching through the Hungarian algorithm, setting the matching items that do not meet the Mahalanobis distance threshold to infinity and then removing them, and performing multi-object cascade matching on the results of each frame to obtain the association results between the targets of each frame of the video.

2. A multi-level matching video racing car tracking method according to claim 1, characterized in that, Combining the motion features and appearance features of the vehicle, and performing multi-level matching according to the object detection frame confidence level to complete vehicle tracking, and obtaining the association results between the targets of each frame of the video, specifically including: Dividing the detection results into high-score frames and low-score frames by setting a confidence level threshold; starting to match the created object tracking trajectories among the high-score frames first, and if not matched, then using the low-score frames and the tracking trajectories that have not been matched with the high-score frames for matching.

3. A multi-level matching video racing car tracking method according to claim 2, characterized in that, Combining the motion features and appearance features of the vehicle, and performing multi-level matching according to the object detection frame confidence level to complete vehicle tracking, and obtaining the association results between the targets of each frame of the video, specifically further including: For the high-score frames that have not been matched with the tracking trajectories and have a high enough score, creating a new tracking trajectory for them.

4. A multi-level matching video racing car tracking method according to claim 3, characterized in that, Combining the motion features and appearance features of the vehicle, and performing multi-level matching according to the object detection frame confidence level to complete vehicle tracking, and obtaining the association results between the targets of each frame of the video, specifically further including: For the tracking trajectories that have not been matched with the detection frames, retaining them for multiple consecutive frames until the target appears again for matching.

5. A multi-level matching video racing car tracking method according to claim 1, characterized in that, the method further includes training the object detection model, specifically: Selecting video segments containing different racing car models for equal-interval frame extraction, annotating the minimum bounding rectangle of each racing car for each extracted frame, constructing a training set, and using the training set to train the model.

6. A multi-level matching video racing car tracking method according to claim 1, characterized in that, the target detection model adopts the yolox network model.

7. A multi-level matching video racing car tracking method according to claim 6, characterized in that, the method further includes: adding a secondary network to the output head of the yolox backbone network, and extracting appearance features of the obtained target detection area through the secondary network.

8. A multi-level matching video racing car tracking system, characterized in that, the system includes: a target detection module, configured to use a pre-trained target detection model to detect target vehicles frame by frame in a racing car video to obtain detection results, and the results include target detection frames and detection result confidence levels; an appearance feature extraction module, configured to input the extracted target detection frame area image into a secondary network to extract vehicle appearance features; a vehicle tracking module, configured to combine the motion features and appearance features of the vehicle, and perform multi-level matching according to the target detection frame confidence level to complete vehicle tracking, and obtain the association result between targets in each frame of the video; Combining the motion features and appearance features of the vehicle, and performing multi-level matching according to the target detection frame confidence level to complete vehicle tracking, and obtaining the association result between targets in each frame of the video, specifically including: For the target detection result of the current frame, a detection frame of an adjacent frame is predicted through Kalman filtering; According to the target detection result and the predicted frame result, the Mahalanobis distance is calculated based on the motion features to obtain the spatial position difference; and the cosine distance is calculated according to the appearance features of targets in different frames to obtain the appearance similarity; The calculated Mahalanobis distance and cosine distance are weighted and summed to obtain a cost matrix, and the Hungarian algorithm is used for matching. The matching items that do not meet the Mahalanobis distance threshold are set to infinity and then removed, and multi-target cascaded matching is performed on the results of each frame to obtain the association result between targets in each frame of the video.

Citation Information

Patent Citations

  • Vehicle detecting and tracking method and device

    CN113658222A

  • Data processing method and device, electronic equipment and storage medium

    CN113694528A