Target tracking method, device, electronic device and storage medium

By adaptively adjusting the detection features in the target video stream, the weight of each sub-feature is determined and the target feature information is generated, the poor tracking effect caused by target occlusion is solved, and more efficient and accurate target tracking is achieved.

CN113850843BActive Publication Date: 2025-08-22LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111134770.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-27
Publication Date
2025-08-22
Estimated Expiration
2041-09-27

AI Technical Summary

Technical Problem

The existing target tracking method has poor tracking effect when the target object is blocked, reducing tracking accuracy.

Method used

By adaptively adjusting the detection features in the target video stream, the weight of each sub-feature is determined, and the change information in the multi-frame image is used to generate target feature information to achieve adaptive tracking of the target.

Benefits of technology

Improve the efficiency and accuracy of target tracking, especially in complex scenarios such as the target occlusion, the target can be tracked more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113850843B_ABST
    Figure CN113850843B_ABST
Patent Text Reader

Abstract

This application discloses a target tracking method, device, electronic device, and storage medium. The method detects a target in a target video stream and obtains a detection feature of the target to be tracked. The detection feature includes at least two sub-features. The weight of each sub-feature is determined based on the change information of each sub-feature in each frame of the image. Each sub-feature is processed according to the weight of each sub-feature to obtain target feature information. The target to be tracked in the target video stream is tracked based on the target feature information to obtain a tracking trajectory of the target to be tracked. The method implements adaptive adjustment of the detection feature based on the weight, so that the obtained target feature information is more suitable for complex tracking scenarios, such as scenarios where the target is obscured, thereby improving the efficiency and accuracy of target tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information processing technology, and more specifically to a target tracking method, device, electronic device and storage medium. Background Art

[0002] Object tracking generally refers to tracking a target object of interest in a video and identifying the location of the target object from each image frame of the video.

[0003] At present, most target tracking methods are based on the principle of correlation filtering tracking. However, these methods rely on the apparent characteristics of the target. When the target object is occluded, the tracking effect is poor, which reduces the tracking accuracy. Summary of the Invention

[0004] In view of this, this application provides the following technical solutions:

[0005] A target tracking method, comprising:

[0006] Detecting a target in a target video stream to obtain a detection feature of the target to be tracked, wherein the target video stream includes multiple frames of images, and the detection feature includes at least two sub-features, each of the sub-features being different;

[0007] Determining a weight of each sub-feature based on change information of each sub-feature in the detection feature in each image;

[0008] Processing each sub-feature according to a weight of each sub-feature to obtain target feature information, wherein the weight of each sub-feature is used to control the influence of each sub-feature on generating the target feature information;

[0009] The target to be tracked in the target video stream is tracked based on the target feature information to obtain a tracking trajectory of the target to be tracked.

[0010] Optionally, determining the weight of each sub-feature based on change information of each sub-feature in the detection feature in each frame of image includes:

[0011] Get the size and position information of each sub-feature in each frame image, and

[0012] The corresponding relationship information of each sub-feature in each frame of the image;

[0013] Determining change information of each sub-feature based on the size and position information of each sub-feature in each frame of image and the corresponding relationship information of each sub-feature in each frame of image;

[0014] Get the initial weight of each sub-feature;

[0015] The initial weight is adjusted based on the change information of each sub-feature to determine the weight of each sub-feature.

[0016] Optionally, the sub-features include feature points and target detection frames, and obtaining corresponding relationship information of each sub-feature in each frame of the image includes:

[0017] Get the number and size of feature points in the target detection frame in each frame of the image.

[0018] Optionally, the processing of each sub-feature according to the weight of each sub-feature to obtain target feature information includes:

[0019] Obtain the feature matrix corresponding to each sub-feature;

[0020] A target feature matrix is ​​determined based on the weight of each sub-feature and the feature matrix corresponding to each sub-feature.

[0021] Optionally, performing tracking processing on the target to be tracked in the target video stream based on the target feature information to obtain a tracking trajectory of the target to be tracked includes:

[0022] Matching each target in the target video stream based on the target feature information to determine a matching result for each target;

[0023] Based on the matching results of the respective targets, obtaining tracking chain information of the target to be tracked;

[0024] According to the tracking chain information of the target to be tracked, the trajectory prediction result of the target to be tracked is corrected to obtain the tracking trajectory of the target to be tracked.

[0025] Optionally, the target feature information includes target detection frame information, and matching each target in the target video stream based on the target feature information to determine a matching result of each target includes:

[0026] Determine whether the target detection frame information corresponding to the (n+1)th frame image and the target detection frame information corresponding to the (n)th frame image meet the matching conditions;

[0027] If yes, determining that the targets corresponding to the target detection frame information are the same target;

[0028] If not, it is determined that the target corresponding to the target detection frame information is a different target.

[0029] Optionally, the correcting the trajectory prediction result of the target to be tracked according to the tracking chain information of the target to be tracked to obtain the tracking trajectory of the target to be tracked includes:

[0030] Obtaining a target detection frame corresponding to the target to be tracked and a predicted target detection frame corresponding to the target to be tracked in the to-be-tracked chain information;

[0031] Correcting the predicted target detection frame based on the target detection frame to obtain an updated target detection frame;

[0032] The target to be tracked is tracked based on the updated target detection frame to obtain a tracking trajectory of the target to be tracked.

[0033] A target tracking device, comprising:

[0034] a detection unit, configured to detect a target in a target video stream and obtain a detection feature of the target to be tracked, wherein the target video stream includes multiple frames of images, and the detection feature includes at least two sub-features, each of the sub-features being different;

[0035] a determining unit, configured to determine a weight of each sub-feature in the detection feature based on change information of each sub-feature in each frame of image;

[0036] a processing unit, configured to process each sub-feature according to a weight of each sub-feature to obtain target feature information, wherein the weight of each sub-feature is used to control the influence of each sub-feature on generating the target feature information;

[0037] The tracking unit is configured to perform tracking processing on the target to be tracked in the target video stream based on the target feature information to obtain a tracking trajectory of the target to be tracked.

[0038] An electronic device, comprising:

[0039] Memory, used to store programs;

[0040] The processor is configured to call and execute the program in the memory, and implement the various steps of the target tracking method as described above by executing the program.

[0041] A readable storage medium stores a computer program thereon, wherein when the computer program is executed by a processor, the steps of the target tracking method described in any one of the above items are implemented.

[0042] Through the above technical solutions, it can be seen that the present application discloses a target tracking method, device, electronic device and storage medium, which detects the target in the target video stream and obtains the detection features of the target to be tracked. The detection features include at least two sub-features. Based on the change information of each sub-feature in each frame image, the weight of each sub-feature is determined; each sub-feature is processed according to the weight of each sub-feature to obtain target feature information; based on the target feature information, the target to be tracked in the target video stream is tracked and processed to obtain the tracking trajectory of the target to be tracked. The detection features are adaptively adjusted based on the weights to make the obtained target feature information more suitable for complex tracking scenarios, such as scenarios where the target is obscured, thereby improving the efficiency and accuracy of target tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0044] Figure 1 A flowchart of a target tracking method provided in an embodiment of the present application;

[0045] Figure 2 A schematic diagram of a target tracking application scenario provided in an embodiment of the present application;

[0046] Figure 3 A schematic diagram of the structure of a target tracking device provided in an embodiment of the present application;

[0047] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0049] A target tracking method is provided in an embodiment of the present application. The method can be applied to the field of computer vision technology to analyze and track a given target in a video to determine the position of the target in the video. The method can be specifically applied in multiple fields such as human-computer interaction, virtual reality, autonomous driving, and video surveillance.

[0050] See also Figure 1, is a flow chart of a target tracking method provided in an embodiment of the present application, the method may include the following steps:

[0051] S101: Detect a target in a target video stream to obtain detection features of the target to be tracked.

[0052] The target video stream is composed of multiple frames of video images. It can be a real-time video stream captured by a video acquisition device such as a camera, or a continuously playable video stream composed of images corresponding to frames with specific time relationships. For example, the target video stream can be video data of the target surveillance area captured in real time by a camera installed in the target surveillance area, or it can be video data uploaded by a user.

[0053] The target video stream contains multiple targets, which can be static or dynamic. Since target tracking is required, the focus is on dynamic targets, meaning that targets can have different positions in the images corresponding to different video frames. For example, targets can be moving vehicles or people. The number of targets to be tracked can be one or multiple, such as tracking a specific person, or vehicles with specific characteristics, such as tracking red vehicles.

[0054] Detecting targets in a target video stream means detecting features that can extract each target, such as detecting facial information that can distinguish each target. The detection features obtained refer to identification features of each target, such as facial features of each target can be used as detection features, or a detection frame for detecting the target can be used as detection features, or the detection frame and the information in the detection frame can be used as detection features, such as a detection frame and feature points in the detection frame. In one possible implementation, the target in the target video stream can be detected based on a target detection model to obtain detection features, such as obtaining a detection image with a detection frame, wherein the target detection model can be a model with a neural network architecture. It should be noted that by including at least one target in the target video stream, each target is detected, and the detection features corresponding to each target will be obtained. In the embodiment of the present application, the main purpose is to obtain the tracking trajectory of the target to be tracked. Therefore, after obtaining the detection features of each target, it is necessary to obtain the detection features of the target to be tracked. The detection features include at least two sub-features, each sub-feature is different. The attributes of the features may be different, such as features belonging to different categories, such as body feature points and target detection frames belonging to different types of detection features. They may also be features representing different parts. For example, when the target is a person, the sub-features corresponding to the detection features may be facial feature points, shoulder feature points, and hand feature points, etc. In another possible implementation, the detection features may be specified first, and then target detection may be performed, that is, target detection may be performed based on the specified detection features, so that each target is determined by the specified detection features. At this time, the detection features of the target to be tracked obtained are the associated information of the detection features of the target to be tracked in each video frame image, such as the coordinate information of the detection features relative to the target point, etc.

[0055] In the embodiment of the present application, the detection feature includes at least two sub-features, which can avoid the problem of selecting a single feature as the detection feature during target tracking, resulting in the single feature being unclear or easily lost when the target is occluded. At the same time, in the embodiment of the present application, the detection feature can also be obtained during the target detection process in the target video, so that target detection and feature acquisition are carried out simultaneously, reducing the use of processing resources.

[0056] S102 : Determine the weight of each sub-feature based on the change information of each sub-feature in the detection feature in each frame of image.

[0057] S103: Process each sub-feature according to its weight to obtain target feature information.

[0058] Since the target video stream includes multiple frames of images, corresponding sub-features will be detected in each frame of image, and the targets to be tracked are mostly moving targets, which will be affected by their own motion and may also be affected by the movement of other targets, so that the detected sub-features will have certain changes in each frame of image, which can be mainly reflected in the change of the position or size information of the sub-features. For example, the sub-features include a right shoulder feature point, which can be detected in the first frame of image, or the relative coordinates of the right shoulder feature point and the target point are detected as the first coordinates. When the right shoulder feature point is not detected in the tenth frame of image, or the relative coordinates of the right shoulder feature point and the target point are detected as the second coordinates, the change information of the right shoulder feature point can be determined by the above information. Similarly, the change information of each sub-feature in each frame of image can be obtained.

[0059] Based on this change information, a weight is determined for each sub-feature. The weight of each sub-feature controls its influence on the generated target feature information. This weight is used to adjust the weight of each sub-feature in the generated target feature information. This allows each sub-feature to be adaptively integrated based on the change information detected in each frame, better suiting the current tracking scenario.

[0060] In the embodiment of the present application, when detecting a target, information about the changes in each sub-feature in different image frames is obtained, and this change information is used as prior information, so that with the help of the prior information, adaptive integration of each sub-feature is obtained. Specifically, the information about the changes in each sub-feature in the image frame before the prediction frame in the target video stream can be used as prior information. For example, when tracking a target person, the sub-features may include head feature points, shoulder feature points, and target detection frames. Usually, head feature points are more important for tracking, and they are assigned a higher weight during the initialization process. However, in actual tracking scenarios, the movement of the person will cause the feature points and target detection frames to change. In scenarios where there may be occlusion, the smaller the occlusion range of the target person, the larger the target detection frame, and the more obvious the feature points of the target itself. Therefore, the weight of each sub-feature can be determined to carry the prior information, and the weight can be applied to the initial weight adjustment of the feature points and target detection frames. This makes the final target feature information more consistent with the current tracking scenario.

[0061] S104 : Tracking the target to be tracked in the target video stream based on the target feature information to obtain a tracking trajectory of the target to be tracked.

[0062] Since the target feature information is a feature that balances the information amount of each sub-feature according to the weight, a more effective tracking feature is obtained. When performing tracking processing based on the target feature information, a general tracking processing method can be used to implement it, except that the tracking feature is replaced by the target feature information obtained by processing. For example, tracking of the target to be tracked can be achieved by means of Hungarian matching and Kalman filtering. However, since a more effective tracking feature is obtained by determining the weight of each sub-feature in the embodiment of the present application, the commonly used cascade matching can be converted to a single-stage Hungarian matching, thereby speeding up the tracking speed. Specifically, the tracking processing process will be described with specific feature information in the subsequent embodiments of the present application and will not be described in detail here.

[0063] The present application discloses a target tracking method that detects a target in a target video stream to obtain a detection feature of the target to be tracked. The detection feature includes at least two sub-features. The weight of each sub-feature is determined based on the change information of each sub-feature in each frame of the image; each sub-feature is processed according to the weight of each sub-feature to obtain target feature information; and the target to be tracked in the target video stream is tracked based on the target feature information to obtain a tracking trajectory of the target to be tracked. The detection feature is adaptively adjusted based on the weight to make the obtained target feature information more suitable for complex tracking scenarios, such as scenarios where the target is occluded, thereby improving the efficiency and accuracy of target tracking.

[0064] In another embodiment of the present application, a method for determining sub-feature weights is provided, which may include the following steps:

[0065] S201 : Obtain the size and position information of each sub-feature in each frame of image, as well as the corresponding relationship information of each sub-feature in each frame of image.

[0066] S202 : Determine change information of each sub-feature based on the size and position information of each sub-feature in each frame of image and the corresponding relationship information of each sub-feature in each frame of image.

[0067] S203: Obtain an initial weight for each sub-feature.

[0068] S204: Adjust the initial weight based on the change information of each sub-feature to determine the weight of each sub-feature.

[0069] When obtaining the size and position information of each sub-feature in each image frame, it is necessary to obtain it based on the characteristics of the sub-feature. When the sub-feature is a feature point, its size information is not obvious, so its position information can be obtained. This position information can be coordinate information in a specific coordinate system or relative to a target point. When the sub-feature is a detection box, the size information is the size of the detection box. If the detection box is a rectangular detection box, the position information can be the coordinates of the detection box vertices.

[0070] The correspondence information of each sub-feature in each frame of image may include at least one of the relative position relationship of each sub-feature, the inclusion relationship between each sub-feature, or the position change information of each sub-feature relative to the target observation point. For example, the sub-features include head feature points and shoulder feature points respectively. In the first few frames of the video stream, the relative distance and angle between the head feature points and the shoulder feature points may be a fixed value. However, in the subsequent frames of images, the shoulder feature points may not be detected because the person is blocked by other objects. At this time, the corresponding relative distance cannot be calculated, and the corresponding correspondence information in different image frames is determined based on this. In an implementation method of an embodiment of the present application, the sub-features include feature points and target detection frames, and obtaining the correspondence information of each sub-feature in each frame of image includes: obtaining the number of feature points and the size of the feature points in the target detection frame in each frame of image. The target detection frame is an identification frame for detecting an object. Typically, in target tracking scenarios, the image or video acquisition device is fixed. For example, in surveillance scenarios, the position of the surveillance camera remains unchanged. When the object is in motion, the farther away it is from the surveillance camera, the smaller the target detection frame becomes. Conversely, when the object is occluded by other objects, the target detection frame will also be affected by the occlusion and become smaller. Accordingly, the feature points in the target detection frame will also change. For example, the feature points in the initial target detection frame include the head feature point, the left shoulder feature point, and the right shoulder feature point. When there is occlusion, the right shoulder feature point may be occluded, and only the head feature point and the left shoulder feature point will remain in the current target detection frame. This serves as the correspondence information between the feature points and the target detection frame.

[0071] After obtaining the size and position information of each sub-feature, as well as the corresponding relationship information of each sub-feature, the changes in this information between different image frames can be determined as change information. Then, based on the learning and verification of this change information, the initial weights of the sub-features can be adjusted to obtain the weight of each sub-feature. The initial weight is a fixed weight value initially set based on the importance of the sub-feature. However, since tracking scenes are usually complex, using a fixed weight value may not reflect the actual motion state of the tracked object, making the tracking results prone to deviation.

[0072] For example, the sub-features include the target detection box and the target feature points, and the weights are obtained by the following formula:

[0073] ω=sigmoid(mean(w,h)).(0.4.J(Point1)+0.3.J(Point2)+0.3.J(Point3))(1)

[0074] Among them, Point represents a feature point. In this formula, three feature points are selected, namely Point1, Point2 and Point3, which can represent the head feature point, left shoulder feature point and right shoulder feature point respectively. J(Pointx) represents whether the feature point Pointx exists, and the value of x can be 1, 2 or 3. If Pointx exists, the value of J(Pointx) is 1, and if it does not exist, it is 0. Because head information is more important for tracking, when initially allocating weights, a higher weight is assigned to Point1 representing the head feature point, and mean(w,h) represents the size of the target detection box. After calculation using the above formula (1), the weight ω obtained can effectively reflect the prior information such as target size and occlusion status in the matching process, so that the amount of information from the target features and target position can be intelligently balanced, thereby obtaining more effective tracking features, and the commonly used two-step cascade can be converted into a single-step cascade to speed up the tracking speed, which will be specifically explained in subsequent embodiments.

[0075] Among them, the sigmoid function in formula (1) is the activation function of the neural network, which is used for the output of hidden layer neurons and has a value range of (0, 1). It can map a real number to the interval of (0, 1) and can be used for binary classification. Therefore, when determining the weight value in the embodiment of the present application, the mode of the neural network model can be used for processing, and the information of the feature points and detection frames in each frame of the video stream and the corresponding weight values ​​​​of the annotations are used as training samples to train the neural network model so that the neural network model can automatically calculate the corresponding weights based on the relevant information of the feature points and detection frames. In addition, the weight obtained by formula (1) can be the weight corresponding to one of the sub-features, and the weight of another sub-feature can be adjusted according to the weight obtained by the calculation, or calculated using a fixed mode, such as the weight of another sub-feature can be (1-ω).

[0076] Correspondingly, in the actual tracking process, calculations and processing are performed based on the feature matrix of the tracking features (i.e., each sub-feature). Correspondingly, the feature matrix corresponding to the sub-feature can be processed based on the weight to obtain the target feature matrix corresponding to the target feature information. In one possible implementation of the embodiment of the present application, processing each sub-feature according to the weight of each sub-feature to obtain the target feature information includes: obtaining a feature matrix corresponding to each sub-feature; and determining the target feature matrix based on the weight of each sub-feature and the feature matrix corresponding to each sub-feature.

[0077] The feature matrix of the sub-feature is constructed based on the feature information corresponding to each sub-feature. The corresponding feature matrix can be constructed based on the binarized data in the sub-feature information. The binarized data corresponds to the pixel points of the sub-feature, that is, the binarized data is obtained by binarizing the pixel points of the sub-feature. The feature matrix can be an N*M matrix, where N represents the number of rows of the binarized data in the sub-feature information, and M represents the number of columns of the binarized data in the sub-feature information. Correspondingly, the feature matrix can also be a matrix obtained by converting the neural network model. For example, the feature information corresponding to the sub-feature is input into a pre-created neural network model, and the corresponding feature matrix is ​​obtained through processing of the convolution layer in the neural network model. It can also be a cost matrix constructed based on the sub-feature information of the current frame and the target detection information of the current frame. The cost matrix can be applied in the Hungarian matching and Kalman filtering processes used in the target tracking process.

[0078] Then, the target feature matrix is ​​determined according to the obtained weight of each sub-feature and the feature matrix of each sub-feature.

[0079] For example, Features represents the feature matrix composed of feature points, Simalicy(boxes) represents the feature matrix of the detection box, and the target feature matrix is:

[0080] A=Features·ω+(1-ω)·Simalirity(boxes) (2)

[0081] It should be noted that the above formula (2) is only a specific application example, and can be combined with actual tracking scenarios. After the weights of each sub-feature are calculated, flexible adjustments can be made when calculating the target feature information.

[0082] The tracking process in the embodiment of the present application is described below. In another embodiment of the present application, tracking the target to be tracked in the target video stream based on the target feature information to obtain the tracking trajectory of the target to be tracked includes:

[0083] S301. Match each target in the target video stream based on target feature information and determine the matching result of each target.

[0084] S302: Based on the matching results of each target, obtain the tracking chain information of the target to be tracked.

[0085] S303: Correct the trajectory prediction result of the target to be tracked according to the tracking chain information of the target to be tracked to obtain the tracking trajectory of the target to be tracked.

[0086] Since the target video stream may include multiple targets, it is necessary to match each target in each frame with the target in the previous frame image based on the target feature information to determine the matching results of each target. That is, if the first frame image includes target A, target B and target C, it is necessary to determine in the second frame image, the third frame image, etc. which target in the current image frame is target A, which is target B, which is target C, etc.

[0087] When the target feature information includes target detection frame information, in one possible implementation, matching each target in the target video stream based on the target feature information and determining a matching result for each target includes:

[0088] Determine whether the target detection frame information corresponding to the (n+1)th frame image and the target detection frame information corresponding to the (n)th frame image meet a matching condition; if so, determine that the targets corresponding to the target detection frame information are the same target; if not, determine that the targets corresponding to the target detection frame information are different targets.

[0089] Where n is the sequence number of the image frame in the target video stream, with a minimum value of 1. The matching condition refers to the conditions under which the target can be matched, such as whether the number of feature points and features in the target detection frame are the same, or whether the feature information in the target detection frame is the same. The corresponding targets in the target detection frame that meet the target condition are determined to be the same target; otherwise, they are determined to be different targets. Matching continues until the targets in the previous and next frames are matched one-to-one. The target detection frame information can include the target detection frame and information about the feature points included in the target detection frame.

[0090] Specifically, the Hungarian matching algorithm can be used to optimally match the target detection frame of the target in the current frame image with the target detection frame of the target in the next frame image. For example, based on the center point coordinates of the target detection frame detected in the current frame image and the center coordinates of each target detection frame in the next frame image, the Euclidean distance or cosine distance similarity is used to find the best match to achieve target matching. Then, the next frame image data is acquired and the target detection frame and the features in the target detection frame are matched in sequence. The above matching process is repeated. Based on the target detection frame information matching results of each target in the consecutive frame images, the tracking chain of each target can be output, thereby obtaining the tracking chain of the target to be tracked.

[0091] In another possible implementation, the trajectory prediction result of the target to be tracked is corrected according to the tracking chain information of the target to be tracked to obtain the tracking trajectory of the target to be tracked, including: obtaining a target detection frame corresponding to the target to be tracked and a predicted target detection frame corresponding to the target to be tracked in the tracking chain information; correcting the predicted target detection frame based on the target detection frame to obtain an updated target detection frame; tracking the target to be tracked based on the updated target detection frame to obtain the tracking trajectory of the target to be tracked.

[0092] Through Hungarian matching, we can determine whether a target in the current frame is the same as a target in the previous frame. Based on this, we obtain the tracking chain information of the target to be tracked, that is, the position of the target to be tracked in each image frame, and then form a tracking chain with these tracking information.

[0093] It should be noted that in the embodiment of the present application, similarity matching is performed on the targets between each video frame. Specifically, the corresponding matching model can be used to perform similarity matching on the target feature information of the object between any two video frames. The matching process is not limited to this. In addition to the Hungarian matching model, the corresponding matching model can also adopt a greedy matching model or other matching models. Examples are not given here one by one.

[0094] Then, based on the position of the target detection frame of the target in the previous frame, the position of the target detection frame in the next frame can be predicted. The predicted target detection frame needs to be corrected based on the above information to obtain a corrected target detection frame. This makes target tracking using the corrected target detection frame, or the tracking trajectory and tracking information obtained after target position prediction, more accurate.

[0095] Specifically, the target detection frame position can be predicted through Kalman filtering. Kalman filtering can predict the current position of the target based on its position at the previous moment, and can estimate the target position more accurately than the sensor. When the Kalman filter filters the detection frame position, for the current video frame, the target position can be obtained through the detection model to determine the position of the target detection frame. The position of the target detection frame marked in the video frame before the current video frame, combined with the target's movement speed, can be predicted in the current video frame. Then, the position of the previously marked target detection frame is corrected according to the predicted detection frame, thereby improving the accuracy of the target detection frame position.

[0096] See also Figure 2 , is a schematic diagram of a target tracking application scenario provided by an embodiment of the present application, in Figure 2 The target tracking process in the illustrated embodiment adopts a combination of Hungarian matching and Kalman filtering.

[0097] Figure 2 The picture in is an image of any frame in the target video stream, where the sub-features include the target frame size and target feature points. The target frame size is represented by (w, h), where w represents the width of the target frame and h represents the height of the target frame. Figure 2 The feature points selected in the example include head feature points, left shoulder feature points, and right shoulder feature points. In the embodiment of the present application, after obtaining the target frame size and target feature points, target tracking is not performed directly. Instead, the target frame size and target feature points are used as prior information, and the adaptive weight ω is obtained by the above formula (1). Based on the adaptive weight, the target frame size and target feature point information are adaptively integrated. Figure 2 Adaptive prior extraction in is used to obtain the target features. Input to the first-order Hungarian matching, such as Figure 2 As shown in the figure, the first-order Hungarian matching mainly matches the targets in each image frame. The target frame is then corrected through Hungarian matching to obtain the final tracking trajectory of the target to be tracked. Specifically, the information input to the first-order Hungarian matching is the target feature information, including the target feature vector after weight adjustment, the target frame size, and the prediction results of the feature points in the target frame. The output of Hungarian matching is the result of a one-to-one matching between targets. The input to the Kalman filter is the tracking chain information, which can correspond to the position of the target center in the tracking chain, the aspect ratio, the height, and the change in these variables. The output of the Kalman filter is the tracking information after the filter iteration update.

[0098] It should be noted that the Hungarian matching in the embodiments of this application is a first-level Hungarian matching, which is different from the commonly used cascade matching processing method. Among them, cascade matching first uses the target features and the previous frame to perform feature Hungarian matching and Kalman wave, and then uses the overlap between the detected target and the previous frame to perform a second Hungarian matching and Kalman filter. Therefore, this application can adaptively integrate each feature based on a determined weight, use the comprehensive features to perform sequential matching to obtain the tracking matching result, speeding up the tracking speed and improving the tracking performance.

[0099] See also Figure 3 , is a schematic diagram of the structure of a target tracking device provided by another embodiment of the present application. The technical solution in this embodiment is mainly used to improve tracking performance and accuracy.

[0100] Specifically, the device in this embodiment may include the following units:

[0101] A detection unit 401 is configured to detect a target in a target video stream and obtain a detection feature of the target to be tracked, wherein the target video stream includes multiple frames of images, and the detection feature includes at least two sub-features, each of which is different;

[0102] a determining unit 402, configured to determine a weight of each sub-feature in the detection feature based on change information of each sub-feature in each frame of the image;

[0103] a processing unit 403 configured to process each sub-feature according to a weight of each sub-feature to obtain target feature information, wherein the weight of each sub-feature is used to control the influence of each sub-feature on generating the target feature information;

[0104] The tracking unit 404 is configured to perform tracking processing on the target to be tracked in the target video stream based on the target feature information to obtain a tracking trajectory of the target to be tracked.

[0105] As can be seen from the above technical solution, in a target tracking device provided by an embodiment of the present application, a target in a target video stream is detected to obtain a detection feature of the target to be tracked. The detection feature includes at least two sub-features. The weight of each sub-feature is determined based on the change information of each sub-feature in each frame image; each sub-feature is processed according to the weight of each sub-feature to obtain target feature information; and the target to be tracked in the target video stream is tracked based on the target feature information to obtain the tracking trajectory of the target to be tracked. Adaptive adjustment of the detection feature based on the weight is achieved to make the obtained target feature information more suitable for complex tracking scenarios, such as scenarios where the target is obscured, thereby improving the efficiency and accuracy of target tracking.

[0106] In one implementation, the determining unit 402 includes:

[0107] The information acquisition subunit is used to obtain the size and position information of each sub-feature in each frame image, and

[0108] The corresponding relationship information of each sub-feature in each frame of the image;

[0109] A first determining subunit is configured to determine change information of each subfeature based on the size and position information of each subfeature in each frame of image and the corresponding relationship information of each subfeature in each frame of image;

[0110] The weight acquisition subunit is used to obtain the initial weight of each sub-feature;

[0111] The adjusting subunit is configured to adjust the initial weight based on the change information of each sub-feature to determine the weight of each sub-feature.

[0112] Optionally, the sub-features include feature points and target detection frames, and obtaining corresponding relationship information of each sub-feature in each frame of the image includes:

[0113] Get the number and size of feature points in the target detection frame in each frame of the image.

[0114] Furthermore, the processing unit 403 is specifically configured to:

[0115] Obtain the feature matrix corresponding to each sub-feature;

[0116] A target feature matrix is ​​determined based on the weight of each sub-feature and the feature matrix corresponding to each sub-feature.

[0117] In one implementation, the tracking unit 404 includes:

[0118] a matching subunit, configured to match each target in the target video stream based on the target feature information and determine a matching result for each target;

[0119] an acquisition subunit, configured to obtain tracking chain information of the target to be tracked based on the matching results of the respective targets;

[0120] The correction subunit is used to correct the trajectory prediction result of the target to be tracked according to the tracking chain information of the target to be tracked, so as to obtain the tracking trajectory of the target to be tracked.

[0121] Optionally, the target feature information includes target detection frame information, and the matching subunit is specifically configured to:

[0122] Determine whether the target detection frame information corresponding to the (n+1)th frame image and the target detection frame information corresponding to the (n)th frame image meet the matching conditions;

[0123] If yes, determining that the targets corresponding to the target detection frame information are the same target;

[0124] If not, it is determined that the target corresponding to the target detection frame information is a different target.

[0125] Optionally, the syndrome unit is specifically configured to:

[0126] Obtaining a target detection frame corresponding to the target to be tracked and a predicted target detection frame corresponding to the target to be tracked in the tracking chain information;

[0127] Correcting the predicted target detection frame based on the target detection frame to obtain an updated target detection frame;

[0128] The target to be tracked is tracked based on the updated target detection frame to obtain a tracking trajectory of the target to be tracked.

[0129] It should be noted that the specific implementation of each unit in this embodiment can refer to the corresponding content in the previous text and will not be described in detail here.

[0130] See also Figure 4 , is a structural diagram of an electronic device provided in another embodiment of the present application. The technical solution in this embodiment is mainly used to improve the effectiveness and accuracy of target tracking.

[0131] Specifically, the electronic device in this embodiment may include the following structure:

[0132] Memory 501, used for storing programs;

[0133] The processor 502 is configured to call and execute the program in the memory, and to implement the following by executing the program:

[0134] Detecting a target in a target video stream to obtain a detection feature of the target to be tracked, wherein the target video stream includes multiple frames of images, and the detection feature includes at least two sub-features, each of the sub-features being different;

[0135] Determining a weight of each sub-feature based on change information of each sub-feature in the detection feature in each image;

[0136] Processing each sub-feature according to a weight of each sub-feature to obtain target feature information, wherein the weight of each sub-feature is used to control the influence of each sub-feature on generating the target feature information;

[0137] The target to be tracked in the target video stream is tracked based on the target feature information to obtain a tracking trajectory of the target to be tracked.

[0138] As can be seen from the above technical solution, in an electronic device provided by an embodiment of the present application, a target in a target video stream is detected to obtain a detection feature of the target to be tracked. The detection feature includes at least two sub-features. The weight of each sub-feature is determined based on the change information of each sub-feature in each frame image; each sub-feature is processed according to the weight of each sub-feature to obtain target feature information; and the target to be tracked in the target video stream is tracked based on the target feature information to obtain a tracking trajectory of the target to be tracked. Adaptive adjustment of the detection feature based on the weight is achieved to make the obtained target feature information more suitable for complex tracking scenarios, such as scenarios where the target is obscured, thereby improving the efficiency and accuracy of target tracking.

[0139] It should be noted that the specific implementation of the processor in this embodiment can refer to the corresponding content in the previous text and will not be described in detail here.

[0140] In another embodiment of the present application, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, each step of the target tracking method as described above is implemented.

[0141] Each embodiment in this specification focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to in detail. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0142] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0143] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0144] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A target tracking method, comprising: Detecting a target in a target video stream to obtain a detection feature of the target to be tracked, wherein the target video stream includes multiple frames of images, and the detection feature includes at least two sub-features, each of the sub-features being different; Determining a weight of each sub-feature based on change information of each sub-feature in the detection feature in each frame of image; wherein the change information of each sub-feature is determined based on size or position information of each sub-feature in each frame of image; Processing each sub-feature according to a weight of each sub-feature to obtain target feature information, wherein the weight of each sub-feature is used to control the influence of each sub-feature on generating the target feature information; The target to be tracked in the target video stream is tracked based on the target feature information to obtain a tracking trajectory of the target to be tracked.

2. The method according to claim 1, wherein determining the weight of each sub-feature based on change information of each sub-feature in the detection feature in each frame of image comprises: Get the size and position information of each sub-feature in each frame image, and The corresponding relationship information of each sub-feature in each frame of the image; Determining change information of each sub-feature based on the size and position information of each sub-feature in each frame of image and the corresponding relationship information of each sub-feature in each frame of image; Get the initial weight of each sub-feature; The initial weight is adjusted based on the change information of each sub-feature to determine the weight of each sub-feature.

3. The method according to claim 2, wherein the sub-features include feature points and target detection frames, and obtaining corresponding relationship information of each sub-feature in each frame of the image comprises: Get the number and size of feature points in the target detection frame in each frame of the image.

4. The method according to claim 1, wherein processing each sub-feature according to the weight of each sub-feature to obtain target feature information comprises: Obtain the feature matrix corresponding to each sub-feature; A target feature matrix is ​​determined based on the weight of each sub-feature and the feature matrix corresponding to each sub-feature.

5. The method according to claim 1, wherein tracking the target to be tracked in the target video stream based on the target feature information to obtain the tracking trajectory of the target to be tracked comprises: Matching each target in the target video stream based on the target feature information to determine a matching result for each target; Based on the matching results of the respective targets, obtaining tracking chain information of the target to be tracked; According to the tracking chain information of the target to be tracked, the trajectory prediction result of the target to be tracked is corrected to obtain the tracking trajectory of the target to be tracked.

6. The method according to claim 5, wherein the target feature information includes target detection frame information, and the matching of each target in the target video stream based on the target feature information to determine the matching result of each target comprises: Determine whether the target detection frame information corresponding to the (n+1)th frame image and the target detection frame information corresponding to the (n)th frame image meet the matching conditions; If yes, determining that the targets corresponding to the target detection frame information are the same target; If not, it is determined that the target corresponding to the target detection frame information is a different target.

7. The method according to claim 6, wherein the step of correcting the trajectory prediction result of the target to be tracked based on the tracking chain information of the target to be tracked to obtain the tracking trajectory of the target to be tracked comprises: Obtaining a target detection frame corresponding to the target to be tracked and a predicted target detection frame corresponding to the target to be tracked in the tracking chain information; Correcting the predicted target detection frame based on the target detection frame to obtain an updated target detection frame; The target to be tracked is tracked based on the updated target detection frame to obtain a tracking trajectory of the target to be tracked.

8. A target tracking device comprising: a detection unit, configured to detect a target in a target video stream and obtain a detection feature of the target to be tracked, wherein the target video stream includes multiple frames of images, and the detection feature includes at least two sub-features, each of the sub-features being different; a determining unit, configured to determine a weight of each sub-feature in the detection feature based on change information of each sub-feature in each frame of image; wherein the change information of each sub-feature is determined based on size or position information of each sub-feature in each frame of image; a processing unit, configured to process each sub-feature according to a weight of each sub-feature to obtain target feature information, wherein the weight of each sub-feature is used to control the influence of each sub-feature on generating the target feature information; The tracking unit is configured to perform tracking processing on the target to be tracked in the target video stream based on the target feature information to obtain a tracking trajectory of the target to be tracked.

9. An electronic device comprising: Memory, used to store programs; The processor is configured to call and execute the program in the memory, and implement the various steps of the target tracking method according to any one of claims 1 to 7 by executing the program.

10. A readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program implements the steps of the target tracking method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target tracking method, device and equipment and storage medium

    CN111709973A

  • Target tracking method and device, electronic equipment and storage medium

    CN113177968A