Target Tracking Method, Device, Electronic Device and Storage Medium

By introducing a second-order derivative of position change in the state vector of the Kalman filter, the target tracking accuracy is improved, and the problem of low target tracking accuracy in the prior art is solved.

CN119810153BActive Publication Date: 2025-06-20HANGZHOU HUAXI INTELLIGENT TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510252322.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

In the existing target tracking technology, the problem of low target tracking accuracy.

Method used

By introducing a second-order derivative of position change, the state vector of the Kalman filter is improved, and the accuracy of the prediction target box is improved, making it closer to the real target.

Benefits of technology

Improves target tracking accuracy, allowing the predicted target box to track mobile targets more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810153B_ABST
    Figure CN119810153B_ABST
Patent Text Reader

Abstract

The present application discloses a target tracking method, device, electronic device and storage medium. The present disclosure relates to the field of computer vision technology. The method includes: obtaining a detection target box of a target to be tracked in a current image frame; obtaining a predicted target box of the target to be tracked in the current image frame, where the predicted target box is predicted based on a Kalman filter, and the state vector of the Kalman filter includes the second derivative of the position change; performing a generalized intersection over union (GIoU) association matching on the detection target box of the target to be tracked in the current image frame and the predicted target box of the target to be tracked in the current image frame to obtain a matching result; and performing target tracking on the next image frame when the matching result meets the target tracking success condition. By introducing the second derivative of the position change, the above technical solution realizes the utilization of the change rate of the target motion speed, can improve the accuracy of the predicted target box, make the predicted target box closer to the real target, and thus improve the target tracking accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular, to an object tracking method, apparatus, electronic device, and storage medium. Background Art

[0002] Object tracking technology can continuously identify and locate one or more moving objects over a period of time, and is widely used in fields such as video surveillance, robot navigation, and autonomous driving.

[0003] Existing object tracking technologies often adopt object tracking algorithms based on the Kalman Filter, such as the SORT multi-object tracking algorithm or the DeepSORT algorithm, etc.

[0004] In the process of implementing the present disclosure, it is found that there are at least the following technical problems in the existing technology: the above-mentioned existing object tracking solutions have the problem of low object tracking accuracy. Summary of the Invention

[0005] The present disclosure provides an object tracking method, apparatus, electronic device, and storage medium. By introducing the second derivative of the position change, the utilization of the change rate of the object motion speed is realized, which can improve the accuracy of predicting the object bounding box, make the predicted object bounding box closer to the real object, and thus improve the object tracking accuracy.

[0006] According to one aspect of the present disclosure, an object tracking method is provided, including:

[0007] Obtain a detection object bounding box of an object to be tracked in a current image frame, where the detection object bounding box is detected based on an object detection algorithm;

[0008] Obtain a predicted object bounding box of the object to be tracked in the current image frame, where the predicted object bounding box is predicted based on a Kalman Filter, and the state vector of the Kalman Filter includes the second derivative of the position change;

[0009] Perform a generalized intersection over union (GIoU) association matching on the detection object bounding box of the object to be tracked in the current image frame and the predicted object bounding box of the object to be tracked in the current image frame to obtain a matching result between the detection object bounding box and the predicted object bounding box;

[0010] In the case where the matching result between the detection object bounding box and the predicted object bounding box meets the object tracking success condition, perform object tracking on the next image frame.

[0011] According to another aspect of the present disclosure, an object tracking apparatus is provided, including:

[0012] The detection target box acquisition module is used to acquire the detection target box of the target to be tracked in the current image frame, where the detection target box is detected based on a target detection algorithm;

[0013] The predicted target box acquisition module is used to acquire the predicted target box of the target to be tracked in the current image frame, where the predicted target box is predicted based on a Kalman filter, and the state vector of the Kalman filter includes the second derivative of the position change;

[0014] The generalized intersection over union correlation matching module is used to perform generalized intersection over union correlation matching on the detection target box of the target to be tracked in the current image frame and the predicted target box of the target to be tracked in the current image frame, and obtain the matching result between the detection target box and the predicted target box;

[0015] The target tracking success determination module is used to perform target tracking on the next image frame when the matching result between the detection target box and the predicted target box meets the target tracking success condition.

[0016] According to another aspect of the present disclosure, there is provided an electronic device, where the electronic device includes:

[0017] At least one processor;

[0018] And a memory communicatively connected to the at least one processor;

[0019] Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the target tracking method according to any embodiment of the present disclosure.

[0020] According to another aspect of the present disclosure, there is provided a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the target tracking method according to any embodiment of the present disclosure when executed by a processor.

[0021] In the technical solution of the embodiments of the present disclosure, a detection target box of a target to be tracked in the current image frame is obtained, where the detection target box is detected based on a target detection algorithm, and then a predicted target box of the target to be tracked in the current image frame is obtained, where the predicted target box is predicted based on a Kalman filter, and the state vector of the Kalman filter includes the second derivative of the position change. Furthermore, a general intersection over union (IoU) correlation matching is performed on the detection target box of the target to be tracked in the current image frame and the predicted target box of the target to be tracked in the current image frame to obtain a matching result between the detection target box and the predicted target box; when the matching result between the detection target box and the predicted target box meets the target tracking success condition, target tracking is performed on the next image frame. In the above technical solution, by introducing the second derivative of the position change, the utilization of the change rate of the target movement speed is realized, the accuracy of the predicted target box can be improved, the predicted target box is closer to the real target, and thus the target tracking accuracy is improved.

[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 is a flowchart of a target tracking method provided in Embodiment 1 of the present disclosure;

[0025] Figure 2 is a flowchart of a target tracking method provided in Embodiment 2 of the present disclosure;

[0026] Figure 3 is a flowchart of a target tracking method provided in Embodiment 3 of the present disclosure;

[0027] Figure 4 is a flowchart of a target tracking method provided in Embodiment 4 of the present disclosure;

[0028] Figure 5 is a schematic structural diagram of a target tracking device provided in Embodiment 5 of the present disclosure;

[0029] Figure 6 is a schematic structural diagram of an electronic device for implementing the target tracking method of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of the data in the technical solution of this application all comply with the relevant regulations of national laws and regulations.

[0032] The standard Kalman filter assumes that both the dynamic model and the observation model of the system are linear, which requires that the target motion is all linear motion. However, in the real world, the motion patterns of many targets and the observation processes of sensors are non-linear. For these situations, the standard Kalman filter may not be able to provide accurate state estimation, resulting in a decrease in target tracking accuracy. To this end, the embodiments of the present disclosure provide a target tracking method, device, electronic device and storage medium, which can effectively solve this problem. The target tracking method, device, electronic device and storage medium provided by the embodiments of the present disclosure will be further described in detail below.

[0033] Embodiment 1

[0034] Figure 1 It is a flowchart of a target tracking method provided in Embodiment 1 of the present disclosure. This embodiment is applicable to the situation where a robot tracks a moving target. This method can be executed by a target tracking device, which can be implemented in the form of hardware and / or software, and the target tracking device can be configured in an electronic device such as a robot. As Figure 1 shown, the method includes:

[0035] S110. Obtain a detection target box of the target to be tracked in the current image frame, where the detection target box is detected based on a target detection algorithm.

[0036] Among them, the current image frame refers to the image frame being processed by the electronic device at the current moment, which can be obtained by real-time acquisition through the camera of the robot or read from the preset storage path of the electronic device. The target to be tracked can be a moving target such as a person or an animal. The target detection algorithm can be a single-stage detection algorithm or a two-stage detection algorithm, etc., and no specific limitation is made here.

[0037] Exemplarily, the image frame real-time acquired by the robot camera can be subjected to target detection to identify and lock the target to be tracked, and the position information of the target to be tracked can be obtained. The position information of the target to be tracked can include a detected target box, that is, the detected target box is the target box obtained through the target detection algorithm.

[0038] S120. Obtain a predicted target box of the target to be tracked in the current image frame, where the predicted target box is predicted based on a Kalman filter, and the state vector of the Kalman filter includes the second derivative of the position change.

[0039] In the embodiment of the present disclosure, the predicted target box is a target box predicted by a Kalman filter. It should be noted that the state vector of the Kalman filter designed in this embodiment includes not only the target position and the first derivative of the position change, but also the second derivative of the position change. The second derivative of the position change refers to the change rate of the target motion speed, that is, the acceleration. By introducing the second derivative of the position change, the utilization of the change rate of the target motion speed is realized, so that the Kalman filter can adapt to non-linear variable-speed motion, thereby improving the accuracy of the predicted target box, making the predicted target box closer to the real target, and thus improving the target tracking accuracy.

[0040] Among them, the target position can include the abscissa of the center point of the target box, the ordinate of the center point of the target box, the area of the target box, and the aspect ratio of the target box. The first derivative of the position change can include the first derivative of the abscissa of the center point of the target box with respect to time, the first derivative of the ordinate of the center point of the target box with respect to time, and the first derivative of the area of the target box with respect to time. The second derivative of the position change can include the second derivative of the abscissa of the center point of the target box with respect to time, the second derivative of the ordinate of the center point of the target box with respect to time, and the second derivative of the area of the target box with respect to time.

[0041] Exemplarily, the state vector of the Kalman filter designed in the embodiment of the present disclosure can be expressed as follows:

[0042] ;

[0043] Among them, represents the abscissa of the center point of the target box, represents the ordinate of the center point of the target box, represents the area of the target box, Represents the aspect ratio of the target bounding box. Represents the first-order derivative of the abscissa of the center point of the target bounding box with respect to time. Represents the first-order derivative of the ordinate of the center point of the target bounding box with respect to time. Represents the first-order derivative of the area of the target bounding box with respect to time. Represents the second-order derivative of the abscissa of the center point of the target bounding box with respect to time. Represents the second-order derivative of the ordinate of the center point of the target bounding box with respect to time. Represents the second-order derivative of the area of the target bounding box with respect to time.

[0044] S130. Perform Generalized Intersection over Union (GIoU) correlation matching on the detected target bounding box of the target to be tracked in the current image frame and the predicted target bounding box of the target to be tracked in the current image frame, to obtain the matching result between the detected target bounding box and the predicted target bounding box.

[0045] S140. When the matching result between the detected target bounding box and the predicted target bounding box meets the target tracking success condition, perform target tracking on the next image frame.

[0046] Among them, the matching result between the detected target bounding box and the predicted target bounding box can be their Generalized IoU (GIoU). The Generalized IoU can be used to measure the overlap degree between the detected target bounding box and the predicted target bounding box. The larger the value, the higher the coincidence degree between the two, and vice versa.

[0047] Specifically, the target tracking success condition can be that the matching result between the detected target bounding box and the predicted target bounding box is greater than a preset threshold. Exemplarily, the preset threshold can be 0.4. If the matching result between the detected target bounding box and the predicted target bounding box is 0.8, it indicates that the target tracking is successful and the target tracking can continue for the next image frame; if the matching result between the detected target bounding box and the predicted target bounding box is 0.2, it indicates that the target tracking fails.

[0048] In the technical solution of the embodiment of the present disclosure, the detection target box of the target to be tracked in the current image frame is obtained, where the detection target box is detected based on a target detection algorithm. Furthermore, the predicted target box of the target to be tracked in the current image frame is obtained, where the predicted target box is predicted based on a Kalman filter, and the state vector of the Kalman filter includes the second derivative of the position change. Then, the general intersection over union (GIoU) association matching is performed on the detection target box of the target to be tracked in the current image frame and the predicted target box of the target to be tracked in the current image frame to obtain the matching result between the detection target box and the predicted target box. When the matching result between the detection target box and the predicted target box meets the target tracking success condition, target tracking is performed on the next image frame. In the above technical solution, by introducing the second derivative of the position change, the utilization of the change rate of the target motion speed is realized, the accuracy of the predicted target box can be improved, the predicted target box is closer to the real target, and thus the target tracking accuracy is improved.

[0049] Embodiment 2

[0050] Figure 2 The flowchart of a target tracking method provided by Embodiment 2 of the present disclosure is shown. The method of this embodiment can be combined with each optional solution in the target tracking method provided in the above embodiment. The target tracking method provided in this embodiment is further optimized. Optionally, the obtaining of the predicted target box of the target to be tracked in the current image frame includes: obtaining the estimated state vector of the previous image frame of the current image frame; performing Kalman filter prediction based on the estimated state vector of the previous image frame of the current image frame to obtain the predicted state vector of the current image frame; and determining the predicted target box of the target to be tracked in the current image frame based on the predicted state vector of the current image frame.

[0051] As Figure 2 shown, the method includes:

[0052] S210. Obtain the detection target box of the target to be tracked in the current image frame, where the detection target box is detected based on a target detection algorithm.

[0053] S220. Obtain the estimated state vector of the previous image frame of the current image frame.

[0054] Here, the previous image frame refers to the image frame processed at the previous moment of the current moment. The estimated state vector refers to the state vector predicted by the Kalman filter.

[0055] S230. Perform Kalman filter prediction based on the estimated state vector of the previous image frame of the current image frame to obtain the predicted state vector of the current image frame.

[0056] Specifically, the estimated state vector of the previous image frame of the current image frame can be used for state prediction through a first preset state transition function to obtain the predicted state vector of the current image frame.

[0057] Optionally, the first preset state transition function is:

[0058] ;

[0059] ;

[0060] ;

[0061] ;

[0062] Among them, represents the actual abscissa of the center point of the target box in the previous image frame, represents the actual ordinate of the center point of the target box in the previous image frame, represents the actual area of the target box in the previous image frame, represents the actual aspect ratio of the width and height of the target box in the previous image frame, represents the first-order derivative of the actual abscissa of the center point of the target box in the previous image frame with respect to time, represents the first-order derivative of the actual ordinate of the center point of the target box in the previous image frame with respect to time, represents the first-order derivative of the actual area of the target box in the previous image frame with respect to time, represents the second-order derivative of the actual abscissa of the center point of the target box in the previous image frame with respect to time, represents the second-order derivative of the actual ordinate of the center point of the target box in the previous image frame with respect to time, represents the second-order derivative of the actual area of the target box with respect to time, represents the time step, 、 、 and are random noises; represents the predicted abscissa of the center point of the target box in the current image frame, represents the predicted ordinate of the center point of the target box in the current image frame, represents the predicted area of the target box in the current image frame, represents the predicted aspect ratio of the width and height of the target box in the current image frame.

[0063] S240. Determine the predicted target box of the target to be tracked in the current image frame based on the predicted state vector of the current image frame.

[0064] In the embodiments of the present disclosure, the predicted state vector includes the horizontal and vertical coordinates of the center point of the target box, the area, and the aspect ratio. Furthermore, the predicted target box can be calculated based on the horizontal and vertical coordinates of the center point of the target box, the area, and the aspect ratio.

[0065] S250. Perform a generalized intersection over union (GIoU) association matching on the detected target box of the target to be tracked in the current image frame and the predicted target box of the target to be tracked in the current image frame, to obtain the matching result between the detected target box and the predicted target box.

[0066] S260. When the matching result between the detected target box and the predicted target box meets the target tracking success condition, perform target tracking on the next image frame.

[0067] Based on the above embodiments, optionally, after performing Kalman filter prediction on the state vector of the previous image frame of the current image frame, it further includes: obtaining the measurement position information of the target to be tracked in the current image frame; based on the measurement position information of the target to be tracked in the current image frame, performing Kalman filter update on the predicted state vector of the current image frame to obtain the estimated state vector of the current image frame, and updating the predicted state vector of the current image frame through a second preset state transition function.

[0068] Among them, the measurement position information is the position information of the target obtained through a target detection algorithm.

[0069] Exemplarily, the Kalman gain can be calculated:

[0070] ;

[0071] Among them, represents the Kalman gain of the current image frame, represents the observation matrix, represents the measurement noise covariance matrix, represents the predicted error covariance of the current image frame.

[0072] Furthermore, use the Kalman gain to perform Kalman filter update on the predicted state vector of the current image frame to obtain the estimated state vector of the current image frame. The formula is as follows:

[0073] ;

[0074] Among them, represents the estimated state vector of the current image frame, represents the predicted state vector of the current image frame, represents the measurement position information of the target to be tracked in the current image frame.

[0075] Furthermore, the error covariance can be updated. The formula is as follows:

[0076] ;

[0077] Among them, represents the estimated error covariance of the current image frame, represents the identity matrix.

[0078] In the embodiments of the present disclosure, the second preset state transition function is:

[0079] ;

[0080] ;

[0081] ;

[0082] ;

[0083] ;

[0084] ;

[0085] Among them, represents the actual abscissa of the center point of the target box in the current image frame, represents the actual ordinate of the center point of the target box in the current image frame, the actual area of the target box in the current image frame, represents the first-order derivative of the predicted abscissa of the center point of the target box in the current image frame with respect to time, represents the first-order derivative of the predicted ordinate of the center point of the target box in the current image frame with respect to time, represents the first-order derivative of the predicted area of the target box in the current image frame with respect to time; represents the second-order derivative of the predicted abscissa of the center point of the target box in the current image frame with respect to time, represents the second-order derivative of the predicted ordinate of the center point of the target box in the current image frame with respect to time, represents the second-order derivative of the predicted area of the target box in the current image frame with respect to time; , , , , and are random noises.

[0086] It should be noted that the first preset state transition function is used to update . The second preset state transition function is used to update . In other words, the values of all elements in the predicted state vector of the present disclosure are completed in two stages, namely the Kalman filter prediction stage and the Kalman filter update stage.

[0087] In the technical solution of the embodiment of the present disclosure, by obtaining the estimated state vector of the previous image frame of the current image frame, and then performing Kalman filter prediction based on the estimated state vector of the previous image frame of the current image frame to obtain the predicted state vector of the current image frame, and then determining the predicted target box of the target to be tracked in the current image frame based on the predicted state vector of the current image frame, the automatic construction of the predicted target box is realized.

[0088] Embodiment III

[0089] Figure 3 As shown in the figure, the flowchart of a target tracking method provided in Embodiment III of the present disclosure. The method in this embodiment can be combined with each optional solution in the target tracking method provided in the above embodiment. The target tracking method provided in this embodiment is further optimized. Optionally, after obtaining the matching result between the detected target box and the predicted target box, it further includes: when the matching result between the detected target box and the predicted target box meets the target tracking failure condition, extracting the appearance features of the current image frame to obtain the appearance features of all targets in the current image frame; performing similarity matching on the appearance features of all targets in the historical appearance feature library to obtain the tracking target, where the historical appearance feature library includes the appearance features of the tracking targets in the historical target tracking process.

[0090] As Figure 3 shown, the method includes:

[0091] S310. Obtain the detected target box of the target to be tracked in the current image frame, where the detected target box is detected based on a target detection algorithm.

[0092] S320. Obtain the predicted target box of the target to be tracked in the current image frame, where the predicted target box is predicted based on a Kalman filter, and the state vector of the Kalman filter includes the second derivative of the position change.

[0093] S330. Perform general intersection over union (GIoU) association matching on the detected target box of the target to be tracked in the current image frame and the predicted target box of the target to be tracked in the current image frame to obtain the matching result between the detected target box and the predicted target box.

[0094] S340. When the matching result between the detected target box and the predicted target box meets the target tracking success condition, perform target tracking on the next image frame.

[0095] S350. When the matching result between the detected target box and the predicted target box meets the target tracking failure condition, extract the appearance features of the current image frame to obtain the appearance features of all targets in the current image frame.

[0096] S360. Perform similarity matching on the appearance features of all the said targets in the historical appearance feature library to obtain the tracked targets, where the historical appearance feature library includes the appearance features of the tracked targets in the historical target tracking process.

[0097] When a target is temporarily occluded by other objects, there is no new measurement position information to update the estimated state vector, and in the actual scenario, the noise accumulates over time, resulting in a huge deviation in the state vector prediction after the target reappears, often leading to problems such as failure or loss of the tracked target.

[0098] Therefore, in the case of target tracking failure, the embodiments of the present disclosure extract the appearance features of the current image frame to obtain the appearance features of all targets in the current image frame, and then perform similarity matching on the appearance features of all targets in the historical appearance feature library to obtain the tracked targets, that is, retrieve the target and continue tracking, effectively improving the tracking effect when the target is occluded.

[0099] Among them, the appearance feature refers to the externally visible features of the target. For example, it can be physical appearance features or clothing features, etc. In the embodiments of the present disclosure, the appearance features can be extracted by an appearance feature extractor, and the appearance feature extractor can be a deep neural network. For example, the appearance feature extractor can be MobileNet or other trained convolutional neural networks, etc., which is not limited herein.

[0100] Exemplarily, the target tracking failure condition can be that the matching result between the detected target box and the predicted target box is less than a preset threshold. Specifically, in the case of target tracking failure, multiple targets in the current image frame are obtained through a target detection algorithm, and the appearance features of each target are obtained through an appearance feature extractor. The appearance features of all targets are subjected to similarity matching with the appearance features of the tracked targets in the historical target tracking process to retrieve the tracked targets. The similarity matching can be achieved by calculating the Euclidean distance or cosine similarity between the appearance features, etc., which is not specifically limited herein. After retrieving the tracked targets, the previous existing trajectories can be cleared, and the state vector of the Kalman filter can be initialized to re-perform target tracking.

[0101] The technical solution of the embodiments of the present disclosure, by performing appearance feature extraction on the current image frame when the matching result between the detected target box and the predicted target box meets the target tracking failure condition, obtaining the appearance features of all targets in the current image frame, and then performing similarity matching on the appearance features of all targets in the historical appearance feature library to retrieve the target and continue tracking, effectively improves the tracking effect when the target is occluded.

[0102] Embodiment 4

[0103] Figure 4The flowchart of a target tracking method provided in the fourth embodiment of the present disclosure. The method of this embodiment is a preferred example of the above embodiment. As Figure 4 shown, the method includes:

[0104] First step, initialize the state vector of the Kalman filter.

[0105] Specifically, the state vector of the Kalman filter can be initialized according to the measured position information of the target. Exemplarily, the . Initial values of can all be 0.

[0106] Second step, perform target detection on the image frame.

[0107] Specifically, a robot in an indoor scenario can collect an image frame through a camera and perform target detection on the image frame to obtain a detection target box of the target to be tracked. Exemplarily, the robot in the indoor scenario can be a robot for guiding services or a security monitoring robot, etc., which is not limited herein.

[0108] Third step, crop the image frame and extract the appearance features of the target.

[0109] Specifically, the image frame can be cropped according to the detection target box of the target to be tracked, and the cropped image is input into an appearance feature extractor to obtain the appearance features of the target and save them to a historical appearance feature library.

[0110] Fourth step, Kalman filter prediction.

[0111] Specifically, perform Kalman filter prediction based on the estimated state vector of the previous image frame of the current image frame to obtain the predicted state vector of the current image frame, and then determine the predicted target box of the target to be tracked in the current image frame based on the predicted state vector of the current image frame.

[0112] Fifth step, GIoU association matching.

[0113] Specifically, perform generalized intersection over union (GIoU) association matching on the detection target box of the target to be tracked in the current image frame and the predicted target box of the target to be tracked in the current image frame to obtain the matching result between the detection target box and the predicted target box.

[0114] Sixth step, Kalman filter update.

[0115] Specifically, obtain the measured position information of the target to be tracked in the current image frame; based on the measured position information of the target to be tracked in the current image frame, perform Kalman filter update on the predicted state vector of the current image frame to obtain the estimated state vector of the current image frame.

[0116] In the seventh step, in the case of target tracking failure, execute the eighth step, and in the case of successful target tracking, return to the second step.

[0117] In the eighth step, perform feature similarity matching.

[0118] Specifically, in the case of target tracking failure, obtain multiple targets in the current image frame through the target detection algorithm, and obtain the appearance features of each target through the appearance feature extractor. Perform similarity matching between the appearance features of all targets and the appearance features of the tracking target during the historical target tracking process to retrieve the tracking target again.

[0119] In the ninth step, clear the existing trajectories and return to the first step.

[0120] Specifically, after retrieving the tracking target again, the existing trajectories before can be cleared, the state vector of the Kalman filter can be initialized, and target tracking can be performed again.

[0121] The technical solution of the embodiment of the present disclosure realizes the utilization of the change rate of the target motion speed by introducing the second derivative of the position change, can improve the accuracy of predicting the target box, make the predicted target box closer to the real target, and thus improve the target tracking accuracy. In addition, in the case of target tracking failure, use appearance feature similarity matching to retrieve the tracking target again and automatically clear the tracking record, and start a new tracking, which improves the continuity and accuracy of the tracking.

[0122] Embodiment Five

[0123] Figure 5 It is a schematic structural diagram of a target tracking device provided in Embodiment Five of the present disclosure. As Figure 5 shown, the device includes:

[0124] A detection target box acquisition module 510, configured to acquire a detection target box of a target to be tracked in the current image frame, where the detection target box is detected based on a target detection algorithm;

[0125] A predicted target box acquisition module 520, configured to acquire a predicted target box of a target to be tracked in the current image frame, where the predicted target box is predicted based on a Kalman filter, and the state vector of the Kalman filter includes the second derivative of the position change;

[0126] A generalized intersection over union correlation matching module 530, configured to perform generalized intersection over union correlation matching on the detection target box of the target to be tracked in the current image frame and the predicted target box of the target to be tracked in the current image frame to obtain a matching result between the detection target box and the predicted target box;

[0127] The target tracking success judgment module 540 is used to perform target tracking on the next image frame when the matching result between the detected target box and the predicted target box meets the target tracking success condition.

[0128] In the technical solution of the embodiment of the present disclosure, by obtaining the detected target box of the target to be tracked in the current image frame, where the detected target box is detected based on the target detection algorithm, and then obtaining the predicted target box of the target to be tracked in the current image frame, where the predicted target box is predicted based on the Kalman filter, and the state vector of the Kalman filter includes the second derivative of the position change, and then performing a generalized intersection over union (GIoU) association matching on the detected target box of the target to be tracked in the current image frame and the predicted target box of the target to be tracked in the current image frame to obtain the matching result between the detected target box and the predicted target box; when the matching result between the detected target box and the predicted target box meets the target tracking success condition, perform target tracking on the next image frame. In the above technical solution, by introducing the second derivative of the position change, the utilization of the change rate of the target motion speed is realized, the accuracy of the predicted target box can be improved, the predicted target box is closer to the real target, and thus the target tracking accuracy is improved.

[0129] In some alternative embodiments, the second derivative of the position change includes the second derivative of the abscissa of the center point of the target box with respect to time, the second derivative of the ordinate of the center point of the target box with respect to time, and the second derivative of the area of the target box with respect to time.

[0130] In some alternative embodiments, the predicted target box acquisition module 520 includes:

[0131] The estimated state vector acquisition unit of the previous image frame is used to obtain the estimated state vector of the previous image frame of the current image frame;

[0132] The Kalman filter prediction unit is used to perform Kalman filter prediction based on the estimated state vector of the previous image frame of the current image frame to obtain the predicted state vector of the current image frame;

[0133] The predicted target box determination unit is used to determine the predicted target box of the target to be tracked in the current image frame based on the predicted state vector of the current image frame.

[0134] In some alternative embodiments, the Kalman filter prediction unit may specifically be used to:

[0135] Perform state prediction on the estimated state vector of the previous image frame of the current image frame through a first preset state transition function to obtain the predicted state vector of the current image frame.

[0136] In some alternative embodiments, the first preset state transition function is:

[0137] ;

[0138] ;

[0139] ;

[0140] ;

[0141] wherein, represents the actual abscissa of the center point of the target box in the previous image frame, represents the actual ordinate of the center point of the target box in the previous image frame, represents the actual area of the target box in the previous image frame, represents the actual aspect ratio of the target box in the previous image frame, represents the first-order derivative of the actual abscissa of the center point of the target box in the previous image frame with respect to time, represents the first-order derivative of the actual ordinate of the center point of the target box in the previous image frame with respect to time, represents the first-order derivative of the actual area of the target box in the previous image frame with respect to time, represents the second-order derivative of the actual abscissa of the center point of the target box in the previous image frame with respect to time, represents the second-order derivative of the actual ordinate of the center point of the target box in the previous image frame with respect to time, represents the second-order derivative of the actual area of the target box with respect to time, represents the time step, , , and are random noises; represents the predicted abscissa of the center point of the target box in the current image frame, represents the predicted ordinate of the center point of the target box in the current image frame, represents the predicted area of the target box in the current image frame, represents the predicted aspect ratio of the target box in the current image frame.

[0142] In some alternative embodiments, the target tracking device further includes:

[0143] a Kalman filter update module, configured to obtain the measurement position information of the target to be tracked in the current image frame; based on the measurement position information of the target to be tracked in the current image frame, perform Kalman filter update on the predicted state vector of the current image frame to obtain the estimated state vector of the current image frame; update the predicted state vector of the current image frame through a second preset state transition function, and the second preset state transition function is:

[0144] ;

[0145] ;

[0146] ;

[0147] ;

[0148] ;

[0149] ;

[0150] wherein, represents the actual abscissa of the center point of the target box in the current image frame, represents the actual ordinate of the center point of the target box in the current image frame, the actual area of the target box in the current image frame, represents the first-order derivative with respect to time of the predicted abscissa of the center point of the target box in the current image frame, represents the first-order derivative with respect to time of the predicted ordinate of the center point of the target box in the current image frame, represents the first-order derivative with respect to time of the predicted area of the target box in the current image frame; represents the second-order derivative with respect to time of the predicted abscissa of the center point of the target box in the current image frame, represents the second-order derivative with respect to time of the predicted ordinate of the center point of the target box in the current image frame, represents the second-order derivative with respect to time of the predicted area of the target box in the current image frame; , , , , and are random noises.

[0151] In some alternative embodiments, the target tracking device further includes:

[0152] a tracking target matching module, configured to, when the matching result between the detected target box and the predicted target box meets the target tracking failure condition, extract appearance features of the current image frame to obtain appearance features of all targets in the current image frame; perform similarity matching on the appearance features of all the targets in a historical appearance feature library to obtain a tracking target, wherein the historical appearance feature library includes appearance features of tracking targets in the historical target tracking process.

[0153] The target tracking device provided by the embodiments of the present disclosure can execute the target tracking method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0154] Embodiment Six

[0155] Figure 6 FIG. 1 shows a schematic structural diagram of an electronic device 10 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0156] As Figure 6 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The I / O interface 15 is also connected to the bus 14.

[0157] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0158] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the target tracking method, which includes:

[0159] Obtaining a detection target box of a target to be tracked in the current image frame, where the detection target box is detected based on a target detection algorithm;

[0160] Obtain the predicted target bounding box of the target to be tracked in the current image frame, where the predicted target bounding box is obtained based on a Kalman filter, and the state vector of the Kalman filter includes the second derivative of the position change;

[0161] Perform a Generalized Intersection over Union (GIoU) association matching on the detected target bounding box of the target to be tracked in the current image frame and the predicted target bounding box of the target to be tracked in the current image frame to obtain the matching result between the detected target bounding box and the predicted target bounding box;

[0162] When the matching result between the detected target bounding box and the predicted target bounding box meets the target tracking success condition, perform target tracking on the next image frame.

[0163] In some embodiments, the target tracking method can be implemented as a computer program, which is tangibly included in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the target tracking method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the target tracking method by any other suitable means (e.g., by means of firmware).

[0164] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), System-on-Chip (SOCs), Complex Programmable Logic Devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, and the programmable processor can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0165] A computer program for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer program may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0166] In the context of the present disclosure, a computer-readable storage medium may be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0167] In order to provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, speech input, or tactile input).

[0168] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0169] A computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0170] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this disclosure can be achieved, and no limitation is imposed herein.

[0171] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A target tracking method, characterized in that: include: Obtaining a detection target frame of a target to be tracked in a current image frame, wherein the detection target frame is detected based on a target detection algorithm; Obtaining an estimated state vector of the previous image frame of the current image frame; Performing state prediction on the estimated state vector of the previous image frame of the current image frame by using a first preset state transfer function to obtain a predicted state vector of the current image frame; Determine a predicted target frame of a target to be tracked in the current image frame based on a predicted state vector of the current image frame, wherein the predicted target frame is obtained based on a Kalman filter prediction, and the state vector of the Kalman filter includes a second-order derivative of position change, which refers to a rate of change of a target motion speed, and includes a second-order derivative of a horizontal coordinate of a center point of the target frame with respect to time, a second-order derivative of a vertical coordinate of a center point of the target frame with respect to time, and a second-order derivative of an area of ​​the target frame with respect to time; Performing generalized intersection-union-correlation matching on a detection target frame of the target to be tracked in the current image frame and a prediction target frame of the target to be tracked in the current image frame to obtain a matching result between the detection target frame and the prediction target frame; Acquire the measured position information of the target to be tracked in the current image frame; based on the measured position information of the target to be tracked in the current image frame, perform Kalman filtering update on the predicted state vector of the current image frame to obtain the estimated state vector of the current image frame; update the predicted state vector of the current image frame by a second preset state transfer function; When the matching result between the detected target frame and the predicted target frame satisfies the target tracking success condition, performing target tracking on the next image frame; The first preset state transfer function is: ; ; ; ; in, Indicates the actual horizontal coordinate of the center point of the target frame in the previous image frame. Indicates the actual vertical coordinate of the center point of the target frame in the previous image frame. Indicates the actual area of ​​the target box in the previous image frame. Indicates the actual aspect ratio of the target frame of the previous image frame. Represents the first-order derivative of the horizontal coordinate of the actual center point of the target frame of the previous image frame with respect to time, Represents the first-order derivative of the ordinate of the actual center point of the target frame of the previous image frame with respect to time, Represents the first-order derivative of the actual area of ​​the target box in the previous image frame with respect to time, Represents the second-order derivative of the horizontal coordinate of the actual center point of the target frame in the previous image frame with respect to time, Represents the second-order derivative of the ordinate of the actual center point of the target frame in the previous image frame with respect to time, Represents the second-order derivative of the actual area of ​​the target box with respect to time, represents the time step, , , and is random noise; Indicates the horizontal coordinate of the predicted center point of the target box in the current image frame, Indicates the ordinate of the predicted center point of the target box in the current image frame. Represents the predicted area of ​​the target box in the current image frame, Indicates the predicted aspect ratio of the target box in the current image frame; The second preset state transfer function is: ; ; ; ; ; ; in, Indicates the actual horizontal coordinate of the center point of the target frame of the current image frame. Indicates the actual center point ordinate of the target frame of the current image frame. The actual area of ​​the target box of the current image frame, Represents the first-order derivative of the horizontal coordinate of the predicted center point of the target box in the current image frame with respect to time, Represents the first-order derivative of the ordinate of the predicted center point of the target frame of the current image frame with respect to time, Represents the first-order derivative of the predicted area of ​​the target box in the current image frame with respect to time; Represents the second-order derivative of the horizontal coordinate of the predicted center point of the target frame of the current image frame with respect to time, Represents the second-order derivative of the ordinate of the predicted center point of the target frame of the current image frame with respect to time, Represents the second-order derivative of the predicted area of ​​the target box in the current image frame with respect to time; , , , , and is random noise.

2. The target tracking method according to claim 1, characterized in that: After obtaining the matching result between the detected target frame and the predicted target frame, it also includes: When the matching result between the detected target frame and the predicted target frame satisfies the target tracking failure condition, performing appearance feature extraction on the current image frame to obtain appearance features of all targets in the current image frame; The appearance features of all the targets are matched for similarity in a historical appearance feature library to obtain a tracking target, wherein the historical appearance feature library includes the appearance features of the tracking target in the historical target tracking process.

3. A target tracking device, characterized in that: include: A detection target frame acquisition module is used to acquire a detection target frame of a target to be tracked in a current image frame, wherein the detection target frame is detected based on a target detection algorithm; A predicted target frame acquisition module is used to acquire an estimated state vector of an image frame previous to a current image frame; perform state prediction on the estimated state vector of the image frame previous to the current image frame through a first preset state transfer function to obtain a predicted state vector of the current image frame; determine a predicted target frame of a target to be tracked in the current image frame based on the predicted state vector of the current image frame, wherein the predicted target frame is obtained based on a Kalman filter prediction, the state vector of the Kalman filter includes a second-order derivative of position change, the second-order derivative of position change refers to a rate of change of target motion speed, and the second-order derivative of position change includes a second-order derivative of a horizontal coordinate of a center point of the target frame with respect to time, a second-order derivative of a vertical coordinate of a center point of the target frame with respect to time, and a second-order derivative of an area of ​​the target frame with respect to time; A generalized intersection-and-union ratio association matching module is used to perform generalized intersection-and-union ratio association matching on a detection target frame of a target to be tracked in the current image frame and a predicted target frame of a target to be tracked in the current image frame to obtain a matching result between the detection target frame and the predicted target frame; A Kalman filter update module, configured to obtain the measured position information of the target to be tracked in the current image frame; perform Kalman filter update on the predicted state vector of the current image frame based on the measured position information of the target to be tracked in the current image frame to obtain an estimated state vector of the current image frame; and update the predicted state vector of the current image frame by a second preset state transfer function; A target tracking success judgment module is used to perform target tracking on the next image frame if the matching result between the detected target frame and the predicted target frame meets the target tracking success condition; The first preset state transfer function is: ; ; ; ; in, Indicates the actual horizontal coordinate of the center point of the target frame in the previous image frame. Indicates the actual vertical coordinate of the center point of the target frame in the previous image frame. Indicates the actual area of ​​the target box in the previous image frame. Indicates the actual aspect ratio of the target frame of the previous image frame. Represents the first-order derivative of the horizontal coordinate of the actual center point of the target frame of the previous image frame with respect to time, Represents the first-order derivative of the ordinate of the actual center point of the target frame of the previous image frame with respect to time, Represents the first-order derivative of the actual area of ​​the target box in the previous image frame with respect to time, Represents the second-order derivative of the horizontal coordinate of the actual center point of the target frame in the previous image frame with respect to time, Represents the second-order derivative of the ordinate of the actual center point of the target frame in the previous image frame with respect to time, Represents the second-order derivative of the actual area of ​​the target box with respect to time, represents the time step, , , and is random noise; Indicates the horizontal coordinate of the predicted center point of the target box in the current image frame, Indicates the ordinate of the predicted center point of the target box in the current image frame. Represents the predicted area of ​​the target box in the current image frame, Indicates the predicted aspect ratio of the target box in the current image frame; The second preset state transfer function is: ; ; ; ; ; ; in, Indicates the actual horizontal coordinate of the center point of the target frame of the current image frame. Indicates the actual center point ordinate of the target frame of the current image frame. The actual area of ​​the target box of the current image frame, Represents the first-order derivative of the horizontal coordinate of the predicted center point of the target box in the current image frame with respect to time, Represents the first-order derivative of the ordinate of the predicted center point of the target frame of the current image frame with respect to time, Represents the first-order derivative of the predicted area of ​​the target box in the current image frame with respect to time; Represents the second-order derivative of the horizontal coordinate of the predicted center point of the target frame of the current image frame with respect to time, Represents the second-order derivative of the ordinate of the predicted center point of the target frame of the current image frame with respect to time, Represents the second-order derivative of the predicted area of ​​the target box in the current image frame with respect to time; , , , , and is random noise.

4. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the target tracking method according to any one of claims 1 to 2.

5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the target tracking method according to any one of claims 1 to 2 when executed.

Citation Information

Patent Citations

  • Target tracking method and device and storage medium

    CN109816701A

  • Visual tracking and positioning method based on target detection

    CN116403139A

  • Vehicle tracking method suitable for roadside sensing scene

    CN117974710A