A target tracking method and device, electronic equipment and storage medium

CN122122628APending Publication Date: 2026-05-29GUANGZHOU SHIYUAN ELECTRONICS CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU SHIYUAN ELECTRONICS CO LTD
Filing Date
2024-09-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In densely populated scenes with many people, existing multi-target tracking technologies suffer from low tracking accuracy due to problems such as occlusion, motion blur, deformation, and interference from similar targets, especially at low frame rates.

Method used

The targets are divided into two categories: standing and sitting. The sitting targets are tracked first, followed by the standing targets. This step-by-step tracking strategy reduces interference and improves accuracy and stability.

Benefits of technology

It improves the accuracy and efficiency of target tracking, maintains the continuity and consistency of the tracking trajectory, and adapts to target tracking in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122122628A_ABST
    Figure CN122122628A_ABST
Patent Text Reader

Abstract

The application provides a target tracking method and device, electronic equipment and storage medium, the method is applied to the computer vision technical field, and the method comprises the following steps: target detection is carried out on the current frame image, and the detection result of each target is obtained;The detection result includes the detection frame and the posture of each target; the first tracking result is determined according to the detection frame of the sitting posture target and each historical trajectory; the sitting posture target is the target in each target, and the posture is sitting posture; each historical trajectory is the trajectory of target tracking in the last frame image; the second tracking result is determined according to the detection frame of the standing posture target and each remaining historical trajectory; the standing posture target is the target in each target, and the posture is standing posture; each remaining historical trajectory includes other historical trajectories in each historical trajectory except the historical trajectory tracked to the sitting posture target; the tracking trajectory of each target in the current frame image is determined according to the first tracking result and the second tracking result.The method can improve the tracking precision of multi-target tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Target tracking method and device, electronic device and storage medium TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, and more particularly, to a target tracking method, device, electronic device and storage medium. BACKGROUND

[0002] Currently, the mainstream technical route of multi-target tracking is detection and then association, that is, target detection is performed based on a single frame image, and then a correlation algorithm based on a similarity matrix is used to associate the target detection boxes between different frames to form a target track. In a multi-person dense scene, due to problems such as occlusion, motion blur, deformation and similar target interference, the accuracy of the multi-person tracking algorithm is not high.

[0003] SUMMARY

[0004] The present application provides a target tracking method, device, electronic device and storage medium, which can improve tracking accuracy.

[0005] In a first aspect, a target tracking method is provided, which includes: performing target detection on a current frame image to obtain detection results of targets; wherein the detection results include detection boxes and poses of the targets, and the poses are standing poses or sitting poses; determining a first tracking result according to the detection boxes of sitting pose targets and historical tracks; wherein the sitting pose targets are targets with sitting poses among the targets; the historical tracks are tracks obtained by tracking targets in a previous frame image; determining a second tracking result according to the detection boxes of standing pose targets and remaining historical tracks; wherein the standing pose targets are targets with standing poses among the targets; the remaining historical tracks include historical tracks other than the historical tracks of the sitting pose targets among the historical tracks; and determining tracking tracks of the targets in the current frame image according to the first tracking result and the second tracking result.

[0006] The technical solution has the advantages that when tracking and matching are performed, the target is divided into a standing posture type target and a sitting posture type target, the first tracking result is determined according to the detection frame of the sitting posture type target and each historical trajectory, which is beneficial to preferentially complete tracking of the sitting posture type target, and when the first tracking result is determined, the standing posture type target is not considered, thereby avoiding interference of the standing posture type target. After the first tracking result is determined, the second tracking result is determined according to the detection frame of the standing posture type target and each remaining historical trajectory. Since each remaining historical trajectory is a historical trajectory other than the historical trajectory tracking the sitting posture type target in each historical trajectory, when the second tracking result is determined, the historical trajectory tracking the sitting posture type target is not considered, but the detection frame of the standing posture type target and each remaining historical trajectory are considered, which is beneficial to avoid interference of the historical trajectory tracking the sitting posture type target as much as possible in the process of determining the second tracking result. Meanwhile, since the sitting posture type target is generally more stable, tracking of this type of target is performed first, which can reduce interference of this type of target on subsequent target tracking and improve accuracy of target tracking. The strategy of determining the tracking result in steps helps to reduce complexity of target tracking, preferentially process the target that is easy to track, that is, the sitting posture type target, thereby improving overall tracking efficiency and accuracy. The tracking trajectory of each target in the current frame image is determined according to the first tracking result and the second tracking result, which is beneficial to maintain continuity and consistency of the tracking trajectory and helps to track the target for a long time and improve tracking precision.

[0007] In combination with the first aspect, in some possible implementation manners, the determining the first tracking result according to the detection frame of the sitting posture type target and each historical trajectory includes: performing first tracking matching on the detection frame of the sitting posture type target and each historical trajectory to obtain the first tracking result; the first tracking result includes a first detection frame matching success set, a first trajectory non-matching success set, and a first detection frame non-matching success set; each remaining historical trajectory includes a historical trajectory in the first trajectory non-matching success set; and the determining the second tracking result according to the detection frame of the standing posture type target, the detection frame in the first detection frame non-matching success set, and the historical trajectory in the first trajectory non-matching success set includes: performing second tracking matching on the detection frame of the standing posture type target, the detection frame in the first detection frame non-matching success set, and the historical trajectory in the first trajectory non-matching success set to obtain the second tracking result; the second tracking result includes a second detection frame matching success set, a second trajectory non-matching success set, and a second detection frame non-matching success set.

[0008] The technical solution has the advantages that the detection frame of the sitting posture target is matched with each historical trajectory in the first tracking matching, which is beneficial to preferentially complete the track matching of the sitting posture target and avoid the interference of the standing posture target in the first tracking matching process. After the first tracking matching is completed, the detection frame of the standing posture target, the detection frame in the first detection frame unmatched successful set and the historical trajectory in the first track unmatched successful set are matched in the second tracking matching. Since the detection frame in the first detection frame matched successful set and the historical trajectory matched with the detection frame are not matched in the second tracking matching process, but the detection frame of the standing posture target, the detection frame in the first detection frame unmatched successful set and the historical trajectory in the first track unmatched successful set are matched, the interference of the sitting posture target that has completed the matching can be avoided in the second tracking matching process as much as possible. Meanwhile, since the sitting posture target is generally more stable, the matching of the target is performed first, which can reduce the interference in the subsequent matching and improve the matching accuracy. The step-by-step matching strategy can help reduce the complexity of the matching, preferentially process the target that is easy to match, that is, the sitting posture target, thereby improving the overall matching efficiency and accuracy, and further improving the tracking accuracy.

[0009] In combination with the first aspect, in some possible implementation manners, the detection result further includes: a detection score corresponding to the detection frame; and the determining the tracking trajectory of each target in the current frame image according to the first tracking result and the second tracking result includes: for each detection frame in the first detection frame matched successful set and each detection frame in the second detection frame matched successful set, updating the historical trajectory matched with the detection frame according to the detection frame to obtain the tracking trajectory of the target located in the detection frame in the current frame image; determining whether the detection score corresponding to each detection frame in the second detection frame unmatched successful set is greater than a first score threshold; and for the detection frame with the detection score greater than the first score threshold, newly establishing the track of the target in the detection frame to obtain the tracking trajectory of the target located in the detection frame in the current frame image.

[0010] The technical solution has the advantages that each detection frame in the first detection frame matched successful set and each detection frame in the second detection frame matched successful set are directly updated according to the detection frame to update the historical trajectory matched with the detection frame, which is beneficial to ensure the continuity and accuracy of the target tracking. The first score threshold is set to determine whether the track of the target in each detection frame in the second detection frame unmatched successful set is newly established, which is helpful to filter out the low-quality detection result. If the detection score corresponding to the detection frame is greater than the first score threshold, it indicates that the accuracy of the detection result of the target surrounded by the detection frame is high, and the tracking trajectory of the target in the detection frame is newly established, which ensures the necessity of newly establishing the tracking trajectory.

[0011] With reference to the first aspect, in some possible implementation manners, after the second tracking matching of the detection box of the standing and sitting posture target, the detection boxes in the first detection box unmatched successful set and the historical trajectories in the first trajectory unmatched successful set is performed to obtain the second tracking result, the method further includes: adding 1 to the continuous missing frame number corresponding to each historical trajectory in the second trajectory unmatched successful set to obtain an updated continuous missing frame number corresponding to each historical trajectory; if the updated continuous missing frame number corresponding to each historical trajectory is greater than a preset discard threshold, the historical trajectory is discarded; and if the updated continuous missing frame number corresponding to each historical trajectory is less than or equal to the preset discard threshold and greater than a preset inactivation threshold, the state information of the historical trajectory is updated to an inactivation state.

[0012] The technical solution has the advantages that for the historical trajectory in the second trajectory unmatched successful set, based on the relationship between the continuous missing frame number and the discard threshold and the inactivation threshold, it is determined whether to discard the historical trajectory or update the state of the historical trajectory to the inactivation state, which helps to reduce the misdiscard situation caused by the temporary disappearance of the historical target tracked by the historical trajectory, can avoid the resource waste caused by the over tracking of the historical target that is not tracked for a long time, can better adapt to the situation of the temporary disappearance of the target tracked by the historical trajectory, and improves the performance of the system in a complex environment.

[0013] With reference to the first aspect, in some possible implementation manners, after the target detection on the current frame image is performed to obtain the detection result of each target, the method further includes: predicting a prediction box corresponding to each historical trajectory in the current frame image by using a state estimator according to the historical trajectory information corresponding to each historical trajectory; the first tracking matching of the detection box of the standing and sitting posture target and each historical trajectory to obtain the first tracking result includes: performing the first tracking matching of the detection box of the standing and sitting posture target and each historical trajectory according to the prediction box corresponding to each historical trajectory in the current frame image to obtain the first tracking result; and the second tracking matching of the detection box of the standing and sitting posture target, the detection boxes in the first detection box unmatched successful set and the historical trajectories in the first trajectory unmatched successful set to obtain the second tracking result includes: performing the second tracking matching of the detection box of the standing and sitting posture target, the detection boxes in the first detection box unmatched successful set and the historical trajectories in the first trajectory unmatched successful set according to the prediction box corresponding to the historical trajectory in the first trajectory unmatched successful set in the current frame image to obtain the second tracking result.

[0014] The above technical solution, when tracking and matching, uses a state estimator to predict the prediction box of each historical trajectory of the previous frame image in the current frame image, and then performs first tracking matching and second tracking matching in combination with the prediction box. The state estimator can effectively process noise and uncertainty, provide more accurate position prediction, and thus improve the success rate of matching.

[0015] In combination with the first aspect, in some possible implementations, the historical trajectory information corresponding to each of the historical trajectories includes: a sitting position box of a target tracked by each of the historical trajectories in a sitting position; and the first tracking matching of the detection box of the sitting class target with each of the historical trajectories according to the prediction box corresponding to each of the historical trajectories in the current frame image to obtain the first tracking result includes: for each of the detection boxes of the sitting class target, calculating a first intersection-over-union between the detection box and the prediction box corresponding to each of the historical trajectories, and calculating a second intersection-over-union between the detection box and the sitting position box corresponding to each of the historical trajectories; calculating a weighted intersection-over-union between the detection box and each of the historical trajectories according to the first intersection-over-union and the second intersection-over-union; if the weighted intersection-over-union between the detection box and any one of the historical trajectories is greater than or equal to a preset intersection-over-union threshold, adding the detection box to the first detection box matching success set; if the weighted intersection-over-union between the detection box and each of the historical trajectories is less than the preset intersection-over-union threshold, adding the detection box to the first detection box non-matching success set; and if the weighted intersection-over-union between the historical trajectory and the detection box of all the sitting class targets is less than the preset intersection-over-union threshold, adding the historical trajectory to the first trajectory non-matching success set.

[0016] The above technical solution considers that the relative position of a target in a sitting position is relatively fixed, and therefore, when matching the detection box of the sitting class target and the prediction box, not only the first intersection-over-union between the detection box and the prediction box corresponding to each of the historical trajectories is calculated, but also the second intersection-over-union between the detection box and the sitting position box corresponding to each of the historical trajectories is further calculated. Finally, the weighted intersection-over-union is calculated according to the first intersection-over-union and the second intersection-over-union, and the historical trajectory successfully matched with the detection box of the sitting class target is determined based on the weighted intersection-over-union. This can obviously avoid the tracking identifier of the sitting class target from being taken away by the standing class target, that is, avoid the case that the detection box of the sitting class target and the prediction box of the historical trajectory of the standing class target are successfully matched. Especially for a classroom scene, the tracking identifier of a student in a sitting position can be obviously avoided from being taken away by a walking student or a teacher, and the tracking accuracy is improved.

[0017] In some possible implementation manners of the first aspect, the detection result further includes: a detection score corresponding to the detection frame; the historical trajectory information corresponding to each historical trajectory includes: state information of each historical trajectory; the first tracking matching of the detection frame of the sitting posture target and each historical trajectory to obtain the first tracking result includes: dividing the detection frame of the sitting posture target into a first high-score detection frame set and a first low-score detection frame set based on a second score threshold; a detection frame in the first high-score detection frame set corresponds to a detection score greater than the second score threshold, and a detection frame in the first low-score detection frame set corresponds to a detection score less than or equal to the second score threshold; dividing each historical trajectory into a first active trajectory set and a first non-active trajectory set based on the state information of each historical trajectory; a historical trajectory in the first active trajectory set is in an active state, and a historical trajectory in the first non-active trajectory set is in a non-active state; performing first matching of the detection frame in the first high-score detection frame set and the historical trajectory in the first active trajectory set to obtain a first detection frame matching success sub-set, a first trajectory non-matching success sub-set and a first detection frame non-matching success sub-set in the first tracking matching; performing second matching of the detection frame in the first low-score detection frame set and the historical trajectory in the first trajectory non-matching success sub-set to obtain a second detection frame matching success sub-set, a second trajectory non-matching success sub-set and a second detection frame non-matching success sub-set in the first tracking matching; performing third matching of the detection frame in the first detection frame non-matching success sub-set and the historical trajectory in the first non-active trajectory set to obtain a third detection frame matching success sub-set, a third trajectory non-matching success sub-set and a third detection frame non-matching success sub-set in the first tracking matching; performing fourth matching of the detection frame in the second detection frame non-matching success sub-set and the historical trajectory in the third trajectory non-matching success sub-set to obtain a fourth detection frame matching success sub-set, a fourth trajectory non-matching success sub-set and a fourth detection frame non-matching success sub-set in the first tracking matching; combining the first detection frame matching success sub-set, the second detection frame matching success sub-set, the third detection frame matching success sub-set and the fourth detection frame matching success sub-set in the first tracking matching to obtain the first detection frame matching success set; combining the second trajectory non-matching success sub-set and the fourth trajectory non-matching success sub-set in the first tracking matching to obtain the first trajectory non-matching success set; and combining the third detection frame non-matching success sub-set and the fourth detection frame non-matching success sub-set in the first tracking matching to obtain the first detection frame non-matching success set.

[0018] The technical solution above divides the first tracking and matching into four times of matching, each time of matching has specific matching conditions and objects. The first time of matching gives priority to the matching between the detection boxes in the first high-score detection box set and the historical trajectories in the first active trajectory set. The second time of matching considers the matching between the detection boxes in the first low-score detection box set and the historical trajectories in the first time of trajectory unmatched sub-set. The third time of matching considers the matching between the detection boxes in the first time of detection box unmatched sub-set and the historical trajectories in the first non-active trajectory set. The fourth time of matching considers the matching between the detection boxes in the second time of detection box unmatched sub-set and the historical trajectories in the third time of trajectory unmatched sub-set. By setting different matching stages, the matching degree between the detection boxes and the trajectories can be more accurately evaluated. The high-score detection boxes are preferentially matched with the historical trajectories in the active state, which can improve the accuracy of matching and reduce the possibility of false matching. Moreover, although the low-score detection boxes have low confidence, they still have the opportunity to be matched, which helps to reduce the missed detection. At the same time, the historical trajectories in the non-active state are used for matching, which can maintain the continuity of target tracking in the case of temporary loss of the target, improve the stability of target tracking, especially in the case of the reappearance of the target after a short disappearance.

[0019] In conjunction with the first aspect, in some possible implementations, the above detection result further includes: the detection score corresponding to the above detection box; the historical trajectory information corresponding to each historical trajectory in the above set of unmatched first trajectories includes: the state information of each of the above historical trajectories; the above-mentioned second tracking matching of the detection boxes of the standing target, the detection boxes in the above set of unmatched first detection boxes, and the historical trajectories in the above set of unmatched first trajectories to obtain the second tracking result includes: classifying the detection boxes of the standing target and the detection boxes in the above set of unmatched first detection boxes into a second high-scoring detection box set and a second low-scoring detection box set based on a second score threshold; wherein, the detection box corresponding to the detection box in the above set of second high-scoring detection boxes... If the detection score is greater than the second score threshold, the detection score of the detection box in the second low-scoring detection box set is less than or equal to the second score threshold. Based on the state information of each historical trajectory in the first trajectory unmatched set, each historical trajectory in the first trajectory unmatched set is divided into a second active trajectory set and a second inactive trajectory set. The historical trajectories in the second active trajectory set are in an active state, and the historical trajectories in the second inactive trajectory set are in an inactive state. The detection boxes in the second high-scoring detection box set and the historical trajectories in the second active trajectory set are matched for the first time to obtain the first successful detection box matching subset and the first unmatched trajectory matching subset in the second tracking matching. The first set of low-scoring detection boxes is matched with the second set of historical trajectories in the set of low-scoring detection boxes. The second set of historical trajectories in the set of low-scoring detection boxes is then matched with the first set of historical trajectories in the set of low-scoring ... The detection boxes and historical trajectories in the third unmatched trajectory subset are matched a fourth time to obtain the fourth successful detection box matching subset, the fourth unmatched trajectory subset, and the fourth unmatched detection box subset in the second tracking match. The first successful detection box matching subset, the second successful detection box matching subset, the third successful detection box matching subset, and the fourth successful detection box matching subset in the second tracking match are combined to obtain the second successful detection box matching set. The second unmatched trajectory subset and the fourth unmatched trajectory subset in the second tracking match are combined to obtain the second unmatched trajectory set.

[0020] Combining the third detection frame unsuccessful matching sub-set and the fourth detection frame unsuccessful matching sub-set in the second tracking matching to obtain a second detection frame unsuccessful matching set.

[0021] In combination with the first aspect, in some possible implementations, the historical trajectory information corresponding to each historical trajectory includes: a sitting position frame of a target tracked by each historical trajectory when the target is in a sitting position; each historical trajectory corresponds to an initialized state estimator; the predicted frame corresponding to each historical trajectory in the current frame image is obtained through the state estimator according to the historical trajectory information corresponding to each historical trajectory, including: for each historical trajectory in the historical trajectories, determining a tracking interval corresponding to the historical trajectory; the tracking interval represents a distance between a latest image frame tracked by the historical trajectory and the current frame image; in a case where the tracking interval is greater than a first preset interval, reinitializing the state estimator corresponding to the historical trajectory through the sitting position frame corresponding to the historical trajectory, and predicting the predicted frame corresponding to the historical trajectory in the current frame image through the reinitialized state estimator; in a case where the tracking interval is greater than a second preset interval and less than or equal to the first preset interval, reinitializing the state estimator corresponding to the historical trajectory through a target detection frame corresponding to the historical trajectory, and predicting the predicted frame corresponding to the historical trajectory in the current frame image through the reinitialized state estimator; the second preset interval is less than the first preset interval, and the target detection frame is a detection frame corresponding to the historical trajectory in the latest image frame tracked by the historical trajectory; in a case where the tracking interval is less than or equal to the second preset interval, predicting the predicted frame corresponding to the historical trajectory in the current frame image through the initialized state estimator corresponding to the historical trajectory.

[0022] The technical solution considers that under a low frame rate of a video, most tracking targets show irregular motion, and when the tracking targets reappear after being occluded, if a linear prediction of a predicted frame is still performed by using a previous Kalman filter, the predicted frame will be obviously distorted, leading to a failure of subsequent association matching. Therefore, the technical solution designs two thresholds (that is, the first preset interval and the second preset interval) for judging an initialization operation, when the tracking interval is too large, causing a target tracked by the historical trajectory to be occluded for a long time, the sitting position frame of the target is used for reinitializing the Kalman filter. When the tracking interval is small, causing the target tracked by the historical trajectory to be occluded for a short time, the latest position frame of the target is used for reinitializing the Kalman filter; when the tracking interval is very small, the prediction value of the Kalman filter is normally used. In this way, the accuracy of the predicted frame obtained by prediction can be improved, and the failure of matching caused by prediction distortion of the Kalman filter can be reduced.

[0023] With reference to the first aspect, in some possible implementation manners, after the tracking trajectories of the targets in the current frame image are determined according to the first tracking result and the second tracking result, the method further includes: determining a number of the tracking trajectories present in the current frame image; and if the number of the tracking trajectories is greater than a preset number, determining to-be-merged tracking trajectories according to the tracking trajectories present in the current frame image, and performing trajectory merging on the to-be-merged tracking trajectories.

[0024] With reference to the first aspect, in some possible implementation manners, the determining of the to-be-merged tracking trajectories according to the tracking trajectories present in the current frame image includes: traversing the tracking trajectories present in the current frame image, taking a currently traversed tracking trajectory as a candidate trajectory, calculating a first coincidence degree in a time dimension and a second coincidence degree in a space dimension between the candidate trajectory and each reference trajectory; wherein each reference trajectory includes a tracking trajectory other than the candidate trajectory among the tracking trajectories present in the current frame image; calculating a spatiotemporal coincidence degree between the candidate trajectory and each reference trajectory according to the first coincidence degree and the second coincidence degree; determining a maximum spatiotemporal coincidence degree among the spatiotemporal coincidence degrees between the candidate trajectory and each reference trajectory, and determining a target reference trajectory having the maximum spatiotemporal coincidence degree with the candidate trajectory among each reference trajectory; and if the maximum spatiotemporal coincidence degree is greater than a preset coincidence degree threshold, determining the candidate trajectory and the target reference trajectory as the to-be-merged tracking trajectories.

[0025] The technical solution described above, when performing tracking trajectory merging, simultaneously considers the first coincidence degree in the time dimension and the second coincidence degree in the space dimension between different tracking trajectories, which is beneficial to accurately determining which tracking trajectories actually track the same target, and thus more accurately determining the to-be-merged tracking trajectories, which is beneficial to ensuring that the tracking trajectory of the seated posture target can maintain continuity after the seated posture target disappears for a period of time and then returns to the seat.

[0026] With reference to the first aspect, in some possible implementation manners, the track information of the track existing in the current frame image includes: a sitting position box of a target tracked by the track; and the calculating the first coincidence degree in the time dimension and the second coincidence degree in the space dimension between the candidate track and each reference track includes: determining a number of frames in which the candidate track and each reference track track to the same frame, and determining a total number of frames tracked by the candidate track; determining a ratio of the number of frames to the total number of frames as the first coincidence degree in the time dimension between the candidate track and each reference track; and determining an intersection-over-union between the sitting position box corresponding to the candidate track and a sitting position box corresponding to each reference track as the second coincidence degree in the space dimension between the candidate track and each reference track.

[0027] With reference to the first aspect, in some possible implementation manners, the method further includes: in a case where the current frame image is a first frame image, screening, from the detection result of the current frame image, a detection box with a detection score greater than a third score threshold; and initializing a track of the current frame image according to the detection box with the detection score greater than the third score threshold; wherein the track information of each initialized track includes: a current pose of a target tracked by the track, a number of tracking frames corresponding to the track, a current detection box of the target tracked by the track, a state estimator configured for the track, and a sitting position box corresponding to the track recorded in a case where the current pose is a sitting pose.

[0028] The second aspect provides a target tracking apparatus, which includes: a detection module configured to perform target detection on a current frame image to obtain detection results of targets; wherein the detection results include detection boxes and poses of the targets, and the poses are standing poses or sitting poses; a first tracking module configured to determine a first tracking result according to detection boxes of sitting pose targets and historical tracks; wherein the sitting pose targets are targets with sitting poses among the targets; and the historical tracks are tracks of target tracking in a previous frame image; a second tracking module configured to determine a second tracking result according to detection boxes of standing pose targets and remaining historical tracks; wherein the standing pose targets are targets with standing poses among the targets; and the remaining historical tracks include historical tracks other than the historical tracks tracking the sitting pose targets among the historical tracks; and a determination module configured to determine tracks of the targets in the current frame image according to the first tracking result and the second tracking result.

[0029] In a third aspect, an electronic device is provided, comprising: a memory configured to store executable program code; and a processor configured to invoke and run the executable program code from the memory, so that the electronic device executes the method in the first aspect or any possible implementation manner of the first aspect.

[0030] In a fourth aspect, a computer program product is provided, comprising: computer program code which, when executed on a computer, causes the computer to execute the method in the first aspect or any possible implementation manner of the first aspect.

[0031] In a fifth aspect, a computer-readable storage medium is provided, which stores computer program code which, when executed on a computer, causes the computer to execute the method in the first aspect or any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0032] FIG. 1 is a schematic flowchart of a target tracking method according to an embodiment of the present application;

[0033] FIG. 2 is a schematic diagram of a double reset prediction method according to an embodiment of the present application;

[0034] FIG. 3 is a schematic diagram of twice internal tracking matching according to an embodiment of the present application;

[0035] FIG. 4 is a schematic diagram of updating, discarding and newly creating a track according to an embodiment of the present application;

[0036] FIG. 5 is a schematic diagram of four times of matching involved in a first tracking matching according to an embodiment of the present application;

[0037] FIG. 6 is a schematic diagram of four times of matching involved in a second tracking matching according to an embodiment of the present application;

[0038] FIG. 7 is a schematic flowchart of track merging according to an embodiment of the present application;

[0039] FIG. 8 is a schematic diagram of a target tracking apparatus according to an embodiment of the present application;

[0040] FIG. 9 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0041] The technical solutions in the present application will be described clearly and exhaustively below with reference to the drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B: "and / or" in the text only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent: A alone, A and B together, and B alone, in addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0042] Hereinafter, the terms "first" and "second" are used only for descriptive purposes and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more features.

[0043] In the multi-target tracking scene, the mainstream technical route of multi-target tracking is to detect and then associate, that is, first target detection based on a single frame image, and then use a similarity matrix-based association algorithm to associate the target detection frame between different frames to form the tracking trajectory of the target.

[0044] For the classroom scene, in order to analyze the behavior of students and teachers, multi-target tracking is usually needed, and the multi-target tracking algorithm needs to detect the motion trajectory of each target (student and teacher) in the classroom scene, which is an upstream basic task for student and teacher behavior analysis. However, the inventors of the present application found the following problems when using existing multi-target tracking algorithms in the classroom scene:

[0045] 1) In the classroom scene, due to the frequent occurrence of teacher patrol and movement, student seat concentration, student standing and going up and down the platform, etc., local occlusion events are prone to occur frequently, and the occlusion duration is not fixed, ranging from several seconds to several minutes. That is, in the classroom scene, occlusion events occur frequently and the duration is not fixed, which may lead to inaccurate association of targets between different frames, thereby affecting the accuracy of target tracking.

[0046] 2) In the classroom scene, most people (such as students) are in a sitting position and have a fixed position, and a few people (such as teachers or individual students) are in a walking state. This may make it difficult for traditional tracking algorithms to distinguish between students in a sitting position and teachers or individual students in a walking state, and in the tracking process, students in a sitting position and teachers or individual students in a walking state will interfere with each other, thereby easily causing tracking errors.

[0047] 3) In the classroom scenario, a large amount of video data can be generated every day. If a high frame rate is used for shooting, the storage requirement will be greatly increased due to the larger storage space occupied by the high frame rate video file. Moreover, the cost of a high frame rate camera is higher, and schools or educational institutions may consider cost-effectiveness when purchasing monitoring cameras and choose low frame rate cameras with higher cost performance. The main purpose of video collection in the classroom scenario is to record classroom activities and behaviors, rather than to capture fast actions, so a lower frame rate is sufficient to meet the monitoring requirements. Based on the above, due to the influence of storage limitations, monitoring requirements, cost-effectiveness, etc., low frame rate cameras are usually used for video collection in the classroom scenario. However, when the frame rate of the video is low, the time interval between frames is large, and students will appear in the video with motion blur, jumping and missing details. These factors together make most students appear as irregular motion, which makes target detection and tracking difficult and makes it difficult to capture the motion trajectory of students, increasing the tracking difficulty.

[0048] From the above analysis, it can be seen that in the classroom scenario, the effect of directly using existing multi-target tracking algorithms is poor, the tracking accuracy is low, and it is difficult to meet the target tracking requirements in the classroom scenario.

[0049] Based on this, in order to improve the tracking accuracy of multi-target tracking, so that the tracking algorithm can be better applied to target tracking in the classroom scenario, the embodiments of the present application provide a target tracking method applied to an electronic device, which can be a computer, a server, etc. Considering that the target in the classroom scenario is usually in a sitting position or walking state, and most of the targets are in a sitting position, and a small part is in a walking state, i.e. a standing position. Therefore, in the embodiments of the present application, the target is divided into two categories: standing position target and sitting position target, and the tracking of the sitting position target is performed first, and then the tracking of the standing position target is performed. Through classification tracking and step-by-step tracking, the mutual interference between the standing position target and the sitting position target in the tracking process can be effectively reduced, and the accuracy and stability of the tracking can be improved.

[0050] The application scenario of the embodiments of the present application can be a tracking scene in which a sitting posture target and a standing posture target can exist simultaneously. The tracking scene can be a classroom scene, a conference scene, a cinema scene, etc. In the classroom scene, the sitting posture target can be a student sitting on a seat, and the standing posture target can be a teacher standing, or the standing posture target can also be a student standing, rising or going up and down the podium. In the conference scene, the sitting posture target can be a conference participant sitting on a seat, and the standing posture target can be a conference host standing or a conference participant standing and speaking. In the cinema scene, the sitting posture target can be a cinema viewer sitting on a seat, and the standing posture target can be a staff standing or a cinema viewer not yet sitting. It should be noted that the embodiments of the present application are only examples of the above-mentioned classroom scene, conference scene, cinema scene, etc., and are not limited thereto in the specific implementation.

[0051] FIG. 1 is a schematic flowchart of a target tracking method provided by the embodiments of the present application. As an example, as shown in FIG. 1, the method comprises the following steps:

[0052] Step 101: performing target detection on a current frame image to obtain detection results of targets; wherein the detection results comprise detection boxes and postures of the targets, and the postures are standing postures or sitting postures.

[0053] Step 102: determining a first tracking result according to the detection boxes of the sitting posture targets and historical trajectories; wherein the sitting posture targets are targets with sitting postures among the targets, and the historical trajectories are trajectories obtained by tracking the targets in a previous frame image.

[0054] Step 103: determining a second tracking result according to the detection boxes of the standing posture targets and remaining historical trajectories; wherein the standing posture targets are targets with standing postures among the targets, and the remaining historical trajectories comprise historical trajectories other than the historical trajectories tracking the sitting posture targets among the historical trajectories.

[0055] Step 104: determining tracking trajectories of the targets in the current frame image according to the first tracking result and the second tracking result.

[0056] In the embodiment shown in FIG. 1, when performing tracking matching, the targets are divided into two categories of standing posture type targets and sitting posture type targets, and first, according to the detection frame of the sitting posture type target and each historical trajectory, a first tracking result is determined, which is conducive to preferentially completing tracking of the sitting posture type target, and when determining the first tracking result, the standing posture type target is not considered, thereby avoiding interference of the standing posture type target. After the first tracking result is determined, according to the detection frame of the standing posture type target and each remaining historical trajectory, a second tracking result is determined. Since each remaining historical trajectory is other historical trajectory in each historical trajectory except the historical trajectory tracking the sitting posture type target, that is, when determining the second tracking result, the historical trajectory tracking the sitting posture type target is not considered, but the detection frame of the standing posture type target and each remaining historical trajectory are considered, thus it is conducive to avoiding as much as possible interference of the historical trajectory tracking the sitting posture type target in the process of determining the second tracking result. At the same time, since the sitting posture type target is generally more stable, tracking of this type of target first can reduce interference of this type of target on subsequent target tracking, and improve the accuracy of target tracking. The strategy of determining tracking results in steps helps to reduce the complexity of target tracking, preferentially processes the target that is easy to track, that is, the sitting posture type target, thereby improving the overall tracking efficiency and accuracy. According to the first tracking result and the above-mentioned second tracking result, the tracking trajectory of each target in the current frame image is determined, which is conducive to maintaining the continuity and consistency of the tracking trajectory, and helps to track the target for a long time and improve the accuracy of tracking.

[0057] In summary, by dividing the targets into two categories of standing posture type targets and sitting posture type targets, and adopting the strategy of tracking in steps, the mutual interference of the standing posture type targets and the sitting posture type targets in the tracking process can be effectively reduced, and the accuracy and stability of tracking can be improved. Therefore, in terms of tracking accuracy and stability in a multi-person dense scene, there is a clear advantage, which can effectively deal with problems such as occlusion and similar target interference, and improve the accuracy of tracking.

[0058] The specific implementation of each step in the embodiment shown in FIG. 1 will be described below.

[0059] In step 101, the current frame image can be understood as: in each video frame of the video collected in the above tracking scene, the current video frame to be processed. The to-be-processed can be understood as: to be subjected to target detection processing.

[0060] In one possible implementation, the tracking scene is a classroom scene, and the video in the classroom scene, that is, the classroom video, can be collected by the camera (such as the camera in the classroom) arranged in the classroom scene. In this case, the current frame image can be understood as: the current video frame to be processed in the classroom video.

[0061] In a possible implementation, the tracking scene is a conference scene, and a video under the conference scene, i.e., a conference video, can be captured by a camera arranged in the conference scene, such as a camera in a conference room. In this case, the current frame image can be understood as a video frame currently to be processed in the conference video.

[0062] In a possible implementation, the tracking scene is a cinema scene, and a video under the cinema scene, i.e., a cinema video, can be captured by a camera arranged in the cinema scene, such as a camera in a cinema. In this case, the current frame image can be understood as a video frame currently to be processed in the cinema video.

[0063] In the embodiments of the present application, the manner of performing target detection on the current frame image includes but is not limited to using a target detection model. The target detection model can be a mature model with good detection effect (for example, a yolov8 model). The input of the target detection model is the current frame image, and the output includes the detection result of each target in the current frame image. The detection result includes the detection box and the posture of each target in the current frame image. The posture is a standing posture or a sitting posture.

[0064] In the case of tracking a human body, the detection box of each target in the current frame image can be a human body detection box, which is used to enclose a human body target in the current frame image. The posture of each target in the current frame image is used to represent whether each target is in a sitting posture or a standing posture in the current frame image.

[0065] In some embodiments, the detection result can further include a detection score corresponding to the detection box. That is, the target detection model can further output the detection score corresponding to the detection box while outputting the detection box and the posture of each target in the current frame image. The detection score corresponding to the detection box is a quantitative index for measuring the credibility of the detection result of the target detection model. The detection score can also be understood as a confidence score or a probability score, which refers to the confidence degree of the target detection model on the target enclosed in the detection box belonging to a specific category (such as a sitting posture, a standing posture, etc.). The detection score is usually a value between 0 and 1, where 1 represents the highest confidence and 0 represents the lowest confidence. That is, the higher the detection score corresponding to the detection box, the more accurate the detection result of the target enclosed in the detection box by the target detection model, and vice versa. The lower the detection score corresponding to the detection box, the less accurate the detection result of the target enclosed in the detection box by the target detection model.

[0066] It can be understood that the current video frame to be processed can be a first frame image or a non-first frame image. The first frame image is a first frame image in each video frame of the video collected in the tracking scene, and the non-first frame image includes other frame images in each video frame of the video collected in the tracking scene except the first frame image. Since the first frame image does not have a previous frame image, in the embodiments of the present application, different processing methods are used for processing in the two cases that the current frame image is a first frame image and the current frame image is a non-first frame image. The different processing methods are introduced as follows:

[0067] In the case that the current frame image is a first frame image, the detection boxes with detection scores greater than a third score threshold are filtered out from the detection results of the current frame image. The tracking trajectories corresponding to the current frame image are initialized according to the detection boxes with detection scores greater than the third score threshold. The trajectory information of each initialized tracking trajectory includes: the current pose of the target tracked by the tracking trajectory, the tracking frame number corresponding to the tracking trajectory, the current detection box of the target tracked by the tracking trajectory, the state estimator configured for the tracking trajectory, and the sitting position box corresponding to the tracking trajectory recorded in the case that the current pose is a sitting pose.

[0068] It can be understood that in the case that the current frame image is a first frame image, each target in the current frame image does not have trajectory information, and therefore the tracking trajectories corresponding to the current frame image need to be initialized. Initializing the tracking trajectories corresponding to the current frame image can be understood as newly creating tracking trajectories for the targets meeting the requirements in the current frame image, and the targets meeting the requirements include the targets enclosed by the detection boxes with detection scores greater than the third score threshold. The third score threshold can be pre-calibrated to measure whether the detection score corresponding to the detection box is high. As described above, the higher the detection score corresponding to the detection box, the more accurate the detection result of the target enclosed by the detection box by the target detection model, and therefore, in the embodiments of the present application, the tracking trajectories are initialized for the detection boxes with detection scores greater than the third score threshold, which is beneficial to improving the accuracy of the initialized tracking trajectories.

[0069] Each initialized tracking trajectory is used to track a target enclosed by a detection box with a detection score greater than the third score threshold. Each tracking trajectory can be provided with a unique trajectory identifier, so that different tracking trajectories can be distinguished by different trajectory identifiers.

[0070] The current pose of the target tracked by the tracking trajectory can be a sitting pose or a standing pose. When the trajectory information of the tracking trajectory is updated based on each image frame after the first frame image, the poses of the target tracked by the tracking trajectory in each image frame appeared can be recorded.

[0071] The tracking frame number corresponding to the tracking trajectory can be understood as the sequence number of the image frame in which the target tracked by the tracking trajectory appears, indicating in which image frame the target is tracked. When updating the track information of the tracking trajectory in each image frame based on the first image frame, the tracking frame number corresponding to the tracking trajectory can be updated, and the sequence numbers of each image frame in which the target tracked by the tracking trajectory appears are recorded. For example, if the tracking frame number corresponding to the tracking trajectory includes id1, id2, id4, and id6, it indicates that the image frames in which the target tracked by the tracking trajectory appears include the first image frame (the sequence number is id1), the second image frame (the sequence number is id2), the fourth image frame (the sequence number is id4), and the sixth image frame (the sequence number is id6).

[0072] The current detection box of the target tracked by the tracking trajectory can be understood as a detection box surrounding the target detected in the current image frame.

[0073] The state estimator configured for the tracking trajectory is used to predict the position of the target in the next image frame. In the embodiments of the present application, a corresponding state estimator is configured for each initialized tracking trajectory to predict the corresponding prediction box of the tracking trajectory in the next image frame. The current detection box of the target tracked by the tracking trajectory and the speed of the current detection box are recorded in the state estimator, so that after the historical track information corresponding to the tracking trajectory in the last image frame is input into the state estimator, the state estimator can output the corresponding prediction box of the tracking trajectory in the next image frame.

[0074] In some embodiments, the above-mentioned state estimator can select a Kalman filter. The input of the Kalman filter is the historical track information corresponding to each historical track in the last image frame, and the output is the prediction box corresponding to each historical track in the current image frame. The Kalman filter adopts a Kalman filtering algorithm, which follows a uniform straight-line motion mathematical model in the scenario of predicting the prediction box corresponding to each historical track in the current image frame. The Kalman filter predicts the prediction box corresponding to each historical track in the current image frame based on the recorded detection box and the speed of the detection box, and assumes that the speed of the detection box is constant.

[0075] When initializing the tracking trajectory, the sitting position box corresponding to the tracking trajectory can also be recorded in the case of the current posture being a sitting posture. At this time, the sitting position box is essentially the same as the current detection box, but in the case of the current posture being a sitting posture, the current detection box is also recorded as a sitting position box in the embodiments of the present application. When updating the track information of the tracking trajectory in each image frame based on the first image frame, the sitting position box of the target tracked by the tracking trajectory in the sitting posture can be recorded.

[0076] In a specific implementation, in order to record the sitting position bounding box, a sliding window can be set in advance for each tracking trajectory, and if it is detected that the target tracked by the tracking trajectory is in a sitting position, the current detection bounding box of the target is stored in the sliding window as a sitting position bounding box. The length of the sliding window can be set in advance, for example, it can be set to be able to store at most 100 sitting position bounding boxes of the target tracked by the tracking trajectory, and the 100 sitting position bounding boxes are actually detected in 100 image frames. In a classroom scene, the sitting class target is usually a student, and since the student has a seat, the sitting position bounding boxes of the same student detected in 100 image frames are less different.

[0077] In step 102, the current frame image is not the first frame image, and therefore, there is a previous frame image of the current frame image. Each historical trajectory is a trajectory for tracking a target in the previous frame image, and each historical trajectory corresponds to a historical target, which includes a target detected when performing target detection on the previous frame image. It should be noted that in the embodiments of the present application, the target detected in the previous frame image is referred to as a historical target, only for the purpose of distinguishing it from the target detected in the current frame image. In fact, the historical target in the previous frame image and the target in the current frame image can be the same target or different targets.

[0078] In this step, the first tracking is performed according to the detection bounding box of the sitting class target and each historical trajectory, and a first tracking result is obtained. The object of the first tracking is the sitting class target, and the corresponding first tracking result can be the tracking result of the sitting class target. The first tracking is essentially to determine whether the sitting class target in the current frame image and a historical target in the previous frame image belong to the same target. For example, if the detection bounding box of a sitting class target in the current frame image is successfully matched with a historical trajectory in the previous frame image, it can be determined that the sitting class target in the detection bounding box and the historical target corresponding to the historical trajectory essentially belong to the same target, that is, the historical trajectory is a historical trajectory that tracks the sitting class target among the historical trajectories. If a historical trajectory is not successfully matched with the detection bounding box of all sitting class targets in the current frame image, it can be determined that the historical trajectory does not track the sitting class target. In the embodiments of the present application, after the first tracking, the historical trajectories that track the sitting class target among the historical trajectories and the historical trajectories that do not track the sitting class target among the historical trajectories can be obtained, and the historical trajectories that do not track the sitting class target among the historical trajectories are taken as the remaining historical trajectories.

[0079] Exemplarily, the implementation of the step 102 can comprise: performing first tracking matching of the detection frame of the sitting posture target and each historical trajectory to obtain a first tracking result; wherein the first tracking result comprises: a first detection frame matching success set, a first trajectory non-matching success set, and a first detection frame non-matching success set; each remaining historical trajectory comprises a historical trajectory in the first trajectory non-matching success set.

[0080] After the first tracking matching of the detection frame of the sitting posture target and each historical trajectory, all the detection frames matching successfully are added to the first detection frame matching success set (denoted as M1), all the detection frames not matching successfully are added to the first detection frame non-matching success set (denoted as U_D1), and all the historical trajectories not matching successfully are added to the first trajectory non-matching success set (denoted as U_T1). The detection frames in M1 are all detection frames of the sitting posture target, and there is a historical trajectory matching successfully with the detection frame. The detection frames in U_D1 are all detection frames of the sitting posture target, and there is no historical trajectory matching successfully with the detection frame in each historical trajectory. The historical trajectories in U_T1 do not match successfully with all the detection frames of the sitting posture target. Each historical trajectory in U_T1 is taken as a remaining historical trajectory, and the sitting posture target in each detection frame in U_D1 is taken as a sitting posture target not tracked after the first tracking.

[0081] Exemplarily, the detection frame of the sitting posture target in the current frame image can be matched with each historical trajectory through a preset matching algorithm, and the matching success means that the detection frame of the sitting posture target is considered to belong to a certain historical trajectory, that is, the sitting posture target in the detection frame belongs to the same target as the historical target corresponding to the historical trajectory. The preset matching algorithm can be: Hungarian algorithm, greedy algorithm, etc.

[0082] In the step 103, second tracking is performed according to the detection frame of the sitting posture target and each remaining historical trajectory (each historical trajectory in U_T1) to obtain a second tracking result. The object of the second tracking includes the sitting posture target and the sitting posture target not tracked after the first tracking (the sitting posture target in each detection frame in U_D1). The second tracking is essentially to determine whether the sitting posture target in the current frame image and the sitting posture target not tracked after the first tracking belong to the same target as a certain historical target in the previous frame image. For example, if the detection frame of a certain sitting posture target in the current frame image matches successfully with a certain remaining historical trajectory in the previous frame image, it can be determined that the sitting posture target in the detection frame and the historical target corresponding to the remaining historical trajectory essentially belong to the same target.

[0083] Exemplarily, the implementation of the step 103 can include: performing second tracking matching on the detection box of the standing posture target, the detection boxes in the first detection box unmatched success set and the historical trajectories in the first trajectory unmatched success set to obtain a second tracking result; the second tracking result includes: a second detection box matched success set, a second trajectory unmatched success set and a second detection box unmatched success set.

[0084] After performing the second tracking matching on the detection box of the standing posture target, the detection boxes in U_D1 and the historical trajectories in U_T1, all the detection boxes matched successfully are added to the second detection box matched success set (denoted as M2), all the detection boxes unmatched successfully are added to the second detection box unmatched success set (denoted as U_D2), and all the historical trajectories unmatched successfully are added to the second trajectory unmatched success set (denoted as U_T2). The detection boxes in M2 can be the detection boxes of the standing posture target or the detection boxes of the sitting posture target, and there is a historical trajectory matched successfully with the detection box in the first trajectory unmatched success set. The detection boxes in U_D2 can be the detection boxes of the standing posture target or the detection boxes of the sitting posture target, and there is no historical trajectory matched successfully with the detection box in the first trajectory unmatched success set. The historical trajectories in U_T2 are all unmatched successfully with all the detection boxes in the current frame image.

[0085] Exemplarily, the second tracking matching on the detection box of the standing posture target, the detection boxes in the first detection box unmatched success set and the historical trajectories in the first trajectory unmatched success set can be performed by a preset matching algorithm. The preset matching algorithm can be a Hungarian algorithm, a greedy algorithm, etc.

[0086] In a possible implementation manner, after the step 101, the method further includes: predicting a prediction box corresponding to each historical trajectory in the current frame image by a state estimator according to historical trajectory information corresponding to each historical trajectory. Correspondingly, the first tracking matching on the detection box of the sitting posture target and each historical trajectory to obtain the first tracking result includes: performing the first tracking matching on the detection box of the sitting posture target and each historical trajectory according to the prediction box corresponding to each historical trajectory in the current frame image to obtain the first tracking result.

[0087] In combination with the above description, each tracking trajectory is configured with a corresponding state estimator. Based on this, for each historical trajectory of the previous frame image, the historical trajectory information corresponding to the historical trajectory can be input into the state estimator corresponding to the historical trajectory, so as to output a prediction box corresponding to the historical trajectory in the current frame image. The prediction box corresponding to the historical trajectory in the current frame image is actually a prediction box corresponding to a historical target tracked by the historical trajectory in the current frame image.

[0088] It should be noted that, in order to distinguish, the track in the current frame image is referred to as a tracking track in the embodiments of the present application, and the track in the previous image frame is referred to as a historical track. With the tracking of the target, the current frame image also dynamically changes. When the next image frame of the current frame image is tracked, the tracking track in the current frame image will also become a historical track, and the target in the current frame image will also become a historical target.

[0089] In some embodiments, the state estimator configured for each historical track is a Kalman filter. The Kalman filter corresponding to each historical track is used to predict a prediction box of a historical target tracked by the historical track in the current frame image based on historical track information corresponding to the historical track. The prediction box represents the predicted position of the historical target tracked by the historical track in the current frame image. Through the Kalman filter, the motion track of the target can be more accurately predicted, the noise influence can be reduced, and a smooth tracking effect can be achieved.

[0090] In some embodiments, when the detection box of the sitting posture class target is matched with each historical track based on the prediction box corresponding to each historical track in the current frame image, if a certain detection box in the current frame image has sufficient similarity with the prediction box corresponding to a certain historical track, it is considered that the detection box and the historical track are matched successfully, and the detection box can be added to the first detection box matched successfully set M1. The remaining detection boxes of all detection boxes of the sitting posture class target except those added to M1 are added to the first detection box unmatched successfully set U_D1. The remaining historical tracks of all historical tracks except those matched successfully with the detection boxes in M1 are added to the first track unmatched successfully set U_T1.

[0091] In some embodiments, in a classroom scene, considering that most students, i.e., tracked targets, exhibit irregular motion under low frame rate of a video, only a few students with specific moving purposes (such as going up and down the platform, standing up, sitting down, etc.) exhibit linear motion. When the student with irregular motion is occluded and then reappears, if the linear prediction of the prediction box is still performed by using the state estimator (such as the Kalman filter), the prediction box will be obviously distorted, resulting in failure of subsequent association matching. Based on this, the application provides a double reset prediction mode of the state estimator to predict the prediction box corresponding to each historical track in the current frame image. The double reset prediction mode is specifically described as follows:

[0092] In some embodiments, the historical track information corresponding to each historical track includes a sitting posture position box of a target tracked by each historical track in a sitting posture; each historical track corresponds to an initialized state estimator; and the prediction box corresponding to each historical track in the current frame image is predicted by the state estimator based on the historical track information corresponding to each historical track, and includes the following steps S01 to S04:

[0093] S01: For each historical trajectory, determine a tracking interval corresponding to the historical trajectory.

[0094] The tracking interval represents the interval between the image frame newly tracked by the historical trajectory and the current image frame. For example, the current frame image is the 5th frame, and suppose the target tracked by historical trajectory 1 is target a. If target a is tracked in the 4th frame, it means that the image frame newly tracked by historical trajectory 1 is the 4th frame, and the tracking interval corresponding to historical trajectory 1 is the distance between the 4th frame and the 5th frame. The distance between the 4th frame and the 5th frame can be represented by the difference in sequence numbers between the frames, i.e., the distance between the 4th frame and the 5th frame can be 1. If neither the 4th frame nor the 3rd frame tracks target a, and target a is tracked in the 2nd frame, it means that the image frame newly tracked by historical trajectory 1 is the 2nd frame, and the tracking interval corresponding to historical trajectory 1 is the distance between the 2nd frame and the 5th frame, which is 5-2=3.

[0095] In some embodiments, in combination with the above description, the historical trajectory information includes the tracking frame number corresponding to the historical trajectory. The tracking frame number includes the sequence numbers of the image frames in which the target tracked by the historical trajectory appears. Based on the tracking frame number corresponding to each historical trajectory, the tracking interval corresponding to the historical trajectory can be determined. For example, if the tracking frame number corresponding to historical trajectory 2 includes id1, id2, id4, and id6, and the current image frame is the 8th frame (sequence number id8), based on the tracking frame number corresponding to historical trajectory 2, it can be determined that the image frame newly tracked by historical trajectory 2 is the 6th frame (sequence number id6). Further, it can be determined that the tracking interval corresponding to historical trajectory 2 is the distance between the 6th frame and the 8th frame, which is 8-6=2.

[0096] S02: In the case where the tracking interval is greater than a first preset interval, reinitializing the state estimator corresponding to the historical trajectory by using the sitting position box corresponding to the historical trajectory, and predicting the prediction box corresponding to the historical trajectory in the current frame image by using the reinitialized state estimator.

[0097] The first preset interval can be preset according to actual needs, and is used to measure whether the tracking interval is too large. If the tracking interval is greater than the first preset interval, it indicates that the tracking interval is too large, and the target tracked by the historical trajectory can be blocked for a long time, so that the target tracked by the historical trajectory is not detected in many image frames before the current image frame. If the state estimator initialized by the first image is still used for the historical trajectory, the accuracy of the output bounding box can be reduced. Therefore, in this case, the state estimator corresponding to the historical trajectory is reinitialized based on the pose position box corresponding to the historical trajectory, so that the reinitialized state estimator can continue to predict the prediction box corresponding to the historical trajectory in the current frame image based on the pose position box corresponding to the historical trajectory.

[0098] Reinitializing the state estimator based on the pose position box corresponding to the historical trajectory can be understood as updating the state quantity of the state estimator based on the pose position box corresponding to the historical trajectory. The state quantity includes position and speed, and these state quantities are used for the state estimator to predict the prediction box corresponding to the historical trajectory in the current frame image. In the embodiment of the present application, the position in the state quantity can be updated to the pose position box corresponding to the historical trajectory, and the speed in the state quantity can be updated to 0, so that the reinitialized state estimator can predict the prediction box based on the updated state quantity and historical trajectory information.

[0099] The pose position box corresponding to the historical trajectory can be understood as a target pose position box obtained based on each recorded pose position box when the target tracked by the historical trajectory is in a sitting position before the current frame image. The position coordinates of the target pose position box can be the average of the position coordinates of each recorded pose position box.

[0100] In combination with the above description, a sliding window is preset for each historical trajectory. If it is detected that the target tracked by the historical trajectory is in a sitting position, the current detection box of the target is stored in the sliding window as a pose position box. Based on this, the pose position box corresponding to the historical trajectory can be understood as calculating the average of the position coordinates of each pose position box recorded in the sliding window corresponding to the historical trajectory to obtain the position coordinates of the target pose position box.

[0101] It can be understood that for each historical trajectory, even if there is a difference between the recorded respective sitting position boxes, the position of the tracked target is relatively fixed when the target is in a sitting position. For example, students in a classroom each have their own seat, and even if there is a difference in the position of the students when they are in a sitting position at different times, the difference is small, so the difference between the sitting position boxes of the same student detected in different frame images is also relatively small. That is, the sitting position box can relatively accurately measure the true position of the same target, so when the tracking interval is large, reinitializing the state estimator corresponding to the historical trajectory using the sitting position box corresponding to the historical trajectory is beneficial to achieving accurate tracking of the target that has been occluded for a long time through the sitting position box corresponding to the historical trajectory.

[0102] S03: In the case where the tracking interval is greater than the second preset interval and less than or equal to the first preset interval, reinitializing the state estimator corresponding to the historical trajectory using the target detection box corresponding to the historical trajectory, and predicting the prediction box corresponding to the historical trajectory in the current frame image through the reinitialized state estimator.

[0103] The second preset interval can be pre-marked according to actual needs, aiming to measure whether the tracking interval is relatively small, and the second preset interval is smaller than the first preset interval.

[0104] If the tracking interval is greater than the second preset interval and less than or equal to the first preset interval, it means that the tracking interval is relatively small, and further means that the target tracked by the historical trajectory may only be short-time occluded, resulting in that the target tracked by the historical trajectory is not detected in few image frames before the current image frame. At this time, if the initialized state estimator configured for the trajectory in the first frame image is still used, the accuracy of the output prediction box may be reduced. Therefore, in this case, the state estimator corresponding to the historical trajectory is reinitialized using the target detection box corresponding to the historical trajectory, and the target detection box is the detection box of the target corresponding to the historical trajectory in the image frame newly tracked by the historical trajectory.

[0105] The target detection box can also be understood as: the position box of the target tracked by the historical trajectory detected in the image frame newly tracked by the historical trajectory before the current image frame. For example, the target tracked by the historical trajectory 3 is the target b, the current image frame is the 5th frame, and the image frame newly tracked by the historical trajectory 3 is the 2nd frame, and the target detection box can be the detection box of the target b detected in the 2nd frame.

[0106] The way of reinitializing the state estimator using the target detection box is similar to the way of reinitializing the state estimator using the sitting position box, and the difference is that the position in the state quantity of the state estimator is updated to the target detection box corresponding to the historical trajectory. Therefore, to avoid repetition, this will not be repeated here.

[0107] When the target tracked by the historical trajectory is only blocked for a short time, the detection box corresponding to the historical trajectory in the latest tracked image frame can represent the real position of the target tracked by the historical trajectory. Therefore, when the tracking interval is relatively small, reinitializing the state estimator corresponding to the historical trajectory by using the target detection box corresponding to the historical trajectory is beneficial to accurately track the target blocked for a short time through the target detection box corresponding to the historical trajectory.

[0108] S04: In a case where the tracking interval is less than or equal to the second preset interval, predicting, by the state estimator corresponding to the historical trajectory, a prediction box corresponding to the trajectory in the current frame image.

[0109] The tracking interval less than or equal to the second preset interval indicates that the tracking interval is very small, and further indicates that the object tracked by the historical trajectory is probably not blocked, or is blocked for a very short time, which can be ignored. In this case, the state estimator corresponding to the historical trajectory can be directly predicted by using the previously initialized state estimator without reinitialization.

[0110] To facilitate understanding of the double-reset prediction method, further description is made in combination with FIG. 2. FIG. 2 is a schematic diagram of the principle of the double-reset prediction method in the embodiment of the present application. For ease of description, the first preset interval is denoted as reset_long and the second preset interval is denoted as reset_short in FIG. 2. In FIG. 2, the state estimator is taken as a Kalman filter for example, and the double-reset prediction can also be referred to as double-reset prediction of the Kalman filter.

[0111] For each historical trajectory strack, the tracking interval time_since_updata corresponding to the trajectory is acquired.

[0112] When the tracking interval time_since_updata>reset_long, the sitting position box corresponding to the historical trajectory is acquired, and the Kalman filter is reinitialized by using the sitting position box, and then the prediction box corresponding to the historical trajectory in the current frame image is output by using the reinitialized Kalman filter.

[0113] When the tracking interval time_since_updata≤reset_long and time_since_updata>reset_short, the latest position box (i.e., the target detection box) corresponding to the historical trajectory is acquired, and the Kalman filter is reinitialized by using the latest position box, and then the prediction box corresponding to the historical trajectory in the current frame image is output by using the reinitialized Kalman filter.

[0114] When the tracking interval time_since_updata≤reset_long and time_since_updata≤reset_short, the Kalman filter does not need to be reinitialized, and the historical trajectory is directly output according to the previously initialized Kalman filter to obtain a prediction box corresponding to the historical trajectory in the current frame image.

[0115] In FIG. 2, mean represents position information of the prediction box, which can include a center point position, a width, and a height of the prediction box. Covariance represents uncertainty of state estimation of the Kalman filter, and can be used to evaluate accuracy of the mean output by the Kalman filter.

[0116] In the embodiments of the present application, double-reset prediction of the state estimator (such as the Kalman filter) is designed. Considering that most tracking targets exhibit irregular motion under low frame rate of a video, when the tracking targets are occluded and then reappear, if linear prediction of the prediction box is still performed by using the previous Kalman filter, the prediction box will be obviously distorted, which leads to failure of subsequent association matching. Therefore, two threshold values (i.e., double-reset threshold values, including the first preset interval and the second preset interval) for judging initialization operation are designed in the embodiments of the present application. When the tracking interval is too large, causing the target tracked by the historical trajectory to be occluded for a long time, the Kalman filter is reinitialized by using the sitting position box of the target. When the tracking interval is small, causing the target tracked by the historical trajectory to be occluded for a short time, the Kalman filter is reinitialized by using the last position box of the target. When the tracking interval is very small, the prediction box of the previously initialized Kalman filter is normally used. In this way, accuracy of the prediction box obtained by prediction can be improved, and matching failure caused by prediction distortion of the Kalman filter can be reduced.

[0117] In a possible implementation manner, the historical trajectory information corresponding to each historical trajectory includes a sitting position box of a target tracked by each historical trajectory in a sitting state; and the detection box of the sitting class target is matched with each historical trajectory for first tracking to obtain a first tracking result, which can include the following S11 to S15:

[0118] S11: For each detection box of the sitting class target, a first intersection-over-union (IOU) between the detection box and a prediction box corresponding to each historical trajectory is calculated, and a second IOU between the detection box and a sitting position box corresponding to each historical trajectory is calculated.

[0119] The prediction box corresponding to each historical trajectory is obtained based on the state estimator, and the detection box of each sitting class target is obtained in step 101. Based on this, the first IOU can be calculated by the following formula: iou_box=IOU(strack_cur_tlbr,det_cur_tlbr)

[0120] wherein, iou box represents the first intersection over union, strack cur tlbr represents the prediction box corresponding to the historical track, and det cur tlbr represents the detection box of the sitting posture target.

[0121] The sitting posture position box corresponding to the historical track can refer to the description above, that is, the target sitting posture position box obtained by the recorded each sitting posture position box when the target tracked based on the historical track is in a sitting posture before the current frame image. The position coordinates of the target sitting posture position box can be the average of the position coordinates of the recorded each sitting posture position box. That is, the sitting posture position box corresponding to the historical track reflects the average level of the sitting posture position box in which the target tracked based on the historical track is located before the current frame image. Based on this, the second intersection over union can be calculated by the following formula: iou sit = IOU(strack sit tlbr, det cur tlbr)

[0122] wherein, iou sit represents the second intersection over union, strack sit tlbr represents the sitting posture position box corresponding to the historical track, and det cur tlbr represents the detection box of the sitting posture target.

[0123] S12: calculating the weighted intersection over union between the detection box and each historical track according to the first intersection over union and the second intersection over union.

[0124] Specifically, the corresponding weights of the first intersection over union and the second intersection over union can be set respectively, so as to calculate the weighted intersection over union between the detection box and each historical track according to the first intersection over union, the second intersection over union and the corresponding weights thereof.

[0125] In some embodiments, considering that the current matched object is a sitting posture target, the weight corresponding to the second intersection over union related to the sitting posture position box can be set relatively large, and the weight corresponding to the first intersection over union can be set relatively small, so as to increase the proportion of the second intersection over union in the finally calculated weighted intersection over union, and avoid that the tracking identifier corresponding to the sitting posture target is taken away by the standing posture target. In the classroom scene, the tracking identifier of the student in a sitting posture can be avoided from being taken away by the walking student or the teacher.

[0126] Exemplarily, the weighted intersection over union can be calculated by the following formula: IOU dist = iou box x (1-sit weight) + iou sit x sit weight

[0127] wherein, IOU dist represents the weighted intersection over union, sit weight represents the weight corresponding to the second intersection over union, and 1-sit weight represents the weight corresponding to the first intersection over union.

[0128] S13: If the weighted intersection-over-union between the detection box and any historical trajectory is greater than or equal to the preset intersection-over-union threshold, the detection box is added to the first detection box matching success set.

[0129] Specifically, if the weighted intersection-over-union between the detection box and a historical trajectory is greater than or equal to the preset intersection-over-union threshold, the detection box and the historical trajectory are taken as a matching success detection box and historical trajectory, that is, the weighted intersection-over-union between the matching success detection box and the historical trajectory is greater than or equal to the preset intersection-over-union threshold. If the weighted intersection-over-union between the detection box and any historical trajectory is greater than or equal to the preset intersection-over-union threshold, it indicates that the detection box and the historical trajectory match successfully, and the detection box is added to the first detection box matching success set M1.

[0130] S14: If the weighted intersection-over-union between the detection box and each historical trajectory is less than the preset intersection-over-union threshold, the detection box is added to the first detection box non-matching success set.

[0131] That is, if the weighted intersection-over-union between the detection box and each historical trajectory is less than the preset intersection-over-union threshold, it indicates that the detection box and any one of the historical trajectories all match unsuccessfully, and the detection box is added to the first detection box non-matching success set U_D1.

[0132] S15: If the weighted intersection-over-union between the historical trajectory and all detection boxes of the sitting posture class target is less than the preset intersection-over-union threshold, the historical trajectory is added to the first trajectory non-matching success set.

[0133] That is, if the weighted intersection-over-union between a historical trajectory and all detection boxes of the sitting posture class target is less than the preset intersection-over-union threshold, it indicates that the historical trajectory and any one of the detection boxes of the sitting posture class target all match unsuccessfully, and the historical trajectory is added to the first trajectory non-matching success set U_T1.

[0134] It can be understood that after the first tracking matching, U_D1 and U_T1 exist, based on which the historical trajectories in U_T1 and the detection boxes in U_D1 can be added to the second tracking matching for matching again. Thus, in step 103, the detection boxes of the sitting posture class target, the detection boxes in the first detection box non-matching success set U_D1, and the historical trajectories in the first trajectory non-matching success set U_T1 are subjected to the second tracking matching to obtain the second tracking result. That is, in step 103, the second tracking matching is performed, and the matching objects include the detection boxes of the sitting posture class target, the detection boxes in U_D1, and the historical trajectories in U_T1.

[0135] In a possible implementation manner, the implementation manner of the second tracking matching of step 103 can include:

[0136] The detection frame of the standing posture class target and the detection frame in U_D1 are both regarded as the detection frame to be matched this time, and the historical track in U_T1 is regarded as the historical track to be matched this time. For each detection frame to be matched, the intersection-over-union of the detection frame and the prediction frame corresponding to the historical track in U_T1 is calculated. If the intersection-over-union between the detection frame and any historical track in U_T1 is greater than or equal to a preset intersection-over-union threshold, the detection frame is added to M2. If the intersection-over-union between the detection frame and all historical tracks in U_T1 is less than the preset intersection-over-union threshold, the detection frame is added to U_D2. If the intersection-over-union between all detection frames to be matched this time and a historical track in U_T1 is less than the preset intersection-over-union threshold, the historical track is added to U_T2.

[0137] To facilitate the understanding of the above-mentioned two internal tracking matches (i.e., the first tracking match and the second tracking match), further description is made in combination with FIG. 3. FIG. 3 is a schematic diagram of the principle of the two internal tracking matches in the embodiment of the present application.

[0138] As shown in FIG. 3, first, the targets Dets detected in the current frame image are divided into the sitting posture class targets Sit-Dets and the standing posture class targets Up-Dets according to the postures. Then, the two internal tracking matches, i.e., the first tracking match (InnerTrack1) and the second tracking match (InnerTrack2) are performed. The two internal tracking matches are described in combination with FIG. 3 as follows:

[0139] The first tracking match is to match the sitting posture class targets Sit-Dets with the prediction frames corresponding to the historical tracks Stracks output by the Kalman filter, to obtain the first tracking result. The first tracking result is shown as M1, U_D1 and U_T1 in FIG. 3, M1 represents the first detection frame matching successful set, U_D1 represents the first detection frame non-matching successful set, and U_T1 represents the first track non-matching successful set.

[0140] It can be understood that, when target tracking is performed on each image frame, M1, U_D1 and U_T1 are all empty, and after the first tracking match is completed, M1, U_D1 and U_T1 are filled with corresponding data.

[0141] The second tracking matching is to take U_T1 as the historical track (U-Stracks) to be matched this time, take U_D1 as the sit-posed class target (U-Sit-Dets) which is not matched successfully after the first tracking matching, and match the prediction box corresponding to U-Stracks with the detection box of U-Sit-Dets and the detection box of Up-Dets to obtain the second tracking result. The second tracking result is represented as M2, U_D2 and U_T2 in FIG. 3, M2 represents a second detection box matching successful set, U_D2 represents a second detection box not matching successful set, and U_T2 represents a second track not matching successful set.

[0142] It can be understood that, when target tracking is performed on each image frame, M2, U_D2 and U_T2 are all empty, and M2, U_D2 and U_T2 are filled with corresponding data after the second tracking matching is completed.

[0143] After the first tracking result and the second tracking result are obtained, step 104 is performed, in which the tracking tracks of the targets in the current frame image are determined according to the first tracking result and the second tracking result.

[0144] The targets in the current frame image include existing targets and newly appeared targets. The existing target can be understood as a target detected in the last frame image and the current frame image, and the newly appeared target can be understood as a target not detected in the last frame image but detected in the current frame image. Determining the tracking tracks of the targets in the current frame image can include updating the historical track information corresponding to the historical track of the existing target to obtain the tracking track of the target in the current frame image, and newly creating a tracking track for tracking the newly appeared target.

[0145] The track information corresponding to each tracking track includes the tracking frame number corresponding to the tracking track, the current detection box of the target tracked by the tracking track, the sit-posed position box corresponding to the tracking track, and information about whether the tracking track is in an active state or an inactive state. The matching priority corresponding to the track in the active state is higher than the matching priority corresponding to the track in the inactive state, that is, when the tracking track and the detection box are matched, the track in the active state is matched preferentially.

[0146] For example, the implementation of step 104 is as follows:

[0147] Step 1041: For each detection box in the first detection box matching successful set and each detection box in the second detection box matching successful set, the historical track matched with the detection box is updated according to the detection box to obtain the tracking track of the target in the current frame image located in the detection box.

[0148] It can be understood that, since each detection box in M1 and each detection box in M2 has a history trajectory matched successfully therewith, it indicates that the target in the detection box and the history target tracked by the history trajectory matched successfully with the detection box belong to the same target, and then the history trajectory matched successfully with the detection box can be updated according to the detection box to obtain the tracking trajectory of the target in the detection box in the current image frame. Specifically, since the detection box is matched to the history trajectory, the detection box can be updated to the history trajectory matched successfully therewith, such as updating the state information of the detection box into the history trajectory information corresponding to the history trajectory matched successfully therewith to obtain the trajectory information corresponding to the tracking trajectory of the target in the detection box in the current image frame. The state information of the detection box can include the position information of the detection box, the sequence number of the current image frame in which the detection box is located (indicating that the target in the detection box is tracked in the current image frame), the posture of the target in the detection box, and the like. When the posture of the target in the detection box is a sitting posture, the detection box can also be recorded as the sitting posture position box of the target tracked by the tracking trajectory in the current image frame, so as to facilitate subsequent determination of the sitting posture position box corresponding to the tracking trajectory.

[0149] The above updating the history trajectory matched successfully with the detection box according to the detection box can be understood as updating the history trajectory of the existing target.

[0150] Step 1042: determining whether the detection scores corresponding to the detection boxes in the second detection box unmatched successful set are greater than the first score threshold.

[0151] Step 1043: for the detection box with the detection score greater than the first score threshold, a trajectory for tracking the target in the detection box is newly built to obtain the tracking trajectory of the target in the detection box in the current image frame.

[0152] The detection boxes in the second detection box unmatched successful set U_D2 are the detection boxes that are not matched successfully after the second tracking matching, which indicates that there is no history trajectory matched successfully with the detection boxes in U_D2 in the history trajectories of the previous image frame, that is, the target in the detection box in U_D2 can not have appeared in the previous image frame. Therefore, in the current image frame, the target in the detection box in U_D2 is likely to be a newly appeared target. Based on this, it is considered to newly build a tracking trajectory for the target in the detection box in U_D2.

[0153] To ensure the necessity of newly establishing a tracking trajectory, the embodiments of the present application do not directly newly establish a tracking trajectory for the target in each detection box in U D2, but first determine whether the detection scores corresponding to each detection box in U D2 are greater than a first score threshold. If the detection score corresponding to the detection box is greater than the first score threshold, it indicates that the detection result of the target surrounded by the detection box by the target detection model has a high accuracy, and at this time, a tracking trajectory for tracking the target in the detection box is newly established. If the detection score corresponding to the detection box is less than or equal to the first score threshold, it indicates that the detection result of the target surrounded by the detection box by the target detection model has a low accuracy, and at this time, the tracking trajectory is not newly established, but the detection box is discarded, so as to avoid newly establishing a tracking trajectory for a target with low credibility, which is beneficial to improving the accuracy of target tracking.

[0154] The newly established tracking trajectory for tracking the target in the detection box can be understood as initializing the tracking trajectory corresponding to the target in the detection box. The track information of the initialized tracking trajectory includes the current pose of the target tracked by the tracking trajectory, the tracking frame number corresponding to the tracking trajectory, the detection box of the target tracked by the tracking trajectory, the state estimator configured for the tracking trajectory, and the sitting pose position box recorded in the case that the current pose of the target in the detection box is a sitting pose. Moreover, the state information of the initialized tracking trajectory is in an active state.

[0155] The first score threshold can be pre-set and can be greater than the third score threshold. However, the embodiments of the present application do not make specific limitations in this regard.

[0156] In some embodiments, after the second tracking matching of the detection box of the standing pose target, the detection boxes in the first detection box unmatched successful set and the historical trajectories in the first trajectory unmatched successful set is performed to obtain the second tracking result, the method can further include: adding 1 to the continuous loss frame number corresponding to each historical trajectory in the second trajectory unmatched successful set to obtain the updated continuous loss frame number corresponding to each historical trajectory; if the updated continuous loss frame number corresponding to each historical trajectory is greater than a preset discard threshold, the historical trajectory is discarded; and if the updated continuous loss frame number corresponding to each historical trajectory is less than or equal to the preset discard threshold and greater than a preset inactivation threshold, the state information of the historical trajectory is updated to an inactivation state.

[0157] The historical trajectories in the second trajectory unmatched successful set U T2 are historical trajectories that are not matched successfully after the second tracking matching, which indicates that there is no detection box matched successfully with each historical trajectory in U T2 in each detection box of the current frame image, that is, the historical target tracked by each historical trajectory in U T2 can not be detected in the current frame image.

[0158] For each historical trajectory in U_T2, the number of consecutive missing frames corresponding to the historical trajectory is used to record how many frames the historical trajectory has not matched with the detection box before the current image frame. The number of consecutive missing frames can be determined based on the tracking frame number of the historical trajectory. As described above, the tracking frame number of the historical trajectory includes the serial numbers of the image frames in which the historical target tracked by the historical trajectory appears, so that based on the tracking frame number of the historical trajectory, the serial number of the latest image frame tracked by the historical trajectory can be determined, and based on the serial number of the latest image frame tracked by the historical trajectory and the serial number of the current image frame, the number of consecutive missing frames can be determined. For example, if the tracking frame number of the historical trajectory includes id1, id2, id4, and id6, and the current image frame is the ninth image (the serial number is id9), it can be determined that the latest image frame tracked by the historical trajectory is the sixth image (the serial number is id6), and then it can be determined that the number of consecutive missing frames is two, i.e., the seventh and eighth images between the ninth image and the sixth image, that is, the number of consecutive missing frames is equal to 2. Since the historical trajectory does not successfully match the detection box this time, it means that the historical target tracked by the historical trajectory is not tracked in the ninth image, and therefore the number of consecutive missing frames corresponding to the updated historical trajectory is 3. That is, the number of consecutive missing frames corresponding to the updated historical trajectory is the number of consecutive missing frames of the historical trajectory before the current image frame plus 1.

[0159] The discard threshold is a threshold for measuring whether to discard the historical trajectory, and the inactivation threshold is a threshold for measuring whether to update the state of the historical trajectory to the inactivation state. The discard threshold is greater than the inactivation threshold, and the specific values of the discard threshold and the inactivation threshold are not limited in the embodiments of the present application.

[0160] After obtaining the updated number of consecutive missing frames, the updated number of consecutive missing frames is compared with the discard threshold first. If the updated number of consecutive missing frames is greater than the discard threshold, it means that the number of consecutive missing frames is too large, i.e., the historical target corresponding to the historical trajectory has not been tracked for a plurality of consecutive frames, and at this time the historical trajectory is discarded.

[0161] If the updated number of consecutive missing frames is less than or equal to the discard threshold and greater than the preset inactivation threshold, it means that there is a certain number of consecutive missing frames, but the number of consecutive missing frames is within the allowable range and does not trigger the discard of the historical trajectory, and at this time the state information of the historical trajectory can be updated to the inactivation state to reduce the matching priority of the historical trajectory when tracking and matching based on the next image frame.

[0162] In order to facilitate the understanding of the above updating, discarding, and newly building trajectory, the following further describes the principle of updating, discarding, and newly building trajectory with reference to FIG. 4.

[0163] In FIG. 4, the corresponding historical tracks Stracks are updated by M1 and M2. The corresponding historical tracks updated by M1 can be understood as: according to the detection boxes in M1, the historical tracks matched successfully with the detection boxes in M1 are updated. The corresponding historical tracks updated by M2 can be understood as: according to the detection boxes in M2, the historical tracks matched successfully with the detection boxes in M2 are updated.

[0164] For the detection boxes in the second detection box unmatched set U_D2 with detection scores greater than the first score threshold, a tracking track for tracking the target in the detection box is newly created, and the detection boxes in U_D2 with detection scores less than or equal to the first score threshold are discarded.

[0165] For the historical tracks in the second track unmatched set U_T2 with continuous loss frame numbers greater than the preset discard threshold lost_buffer_size, the historical tracks are discarded. For the historical tracks in U_T2 with continuous loss frame numbers less than or equal to the preset discard threshold and greater than the preset non-activation threshold non_act_buffer_size, the state information of the historical tracks is updated to the non-activation state None-Act-Stracks.

[0166] In FIG. 4, the historical tracks updated by M1 and M2, the tracking tracks newly created by U_D2, and the historical tracks in the non-activation state None-Act-Stracks updated by U_T2 are all taken as the tracking tracks Stracks existing in the current frame image.

[0167] In the embodiments of the present application, for M1 and M2, the historical tracks matched successfully with M1 and M2 are directly updated according to the detection boxes in M1 and M2, which is beneficial to guarantee the continuity and accuracy of target tracking. By setting the first score threshold to determine whether a tracking track is newly created for the target in the detection box in U_D2, it is helpful to filter out low-quality detection results, and the detection boxes in U_D2 with detection scores less than or equal to the first score threshold are discarded, which can save computing resources and improve the processing efficiency of the system. For the historical tracks in U_T2, based on the relationship between the continuous loss frame number and the discard threshold and the non-activation threshold, it is determined whether to discard the historical track or to update the state of the historical track to the non-activation state, which can better adapt to the situation that the target tracked by the historical track disappears temporarily and improve the performance of the system in a complex environment.

[0168] In some embodiments, the detection result further includes: a detection score corresponding to the detection box; and the historical trajectory information corresponding to each historical trajectory includes: state information of each historical trajectory, which can be an active state or an inactive state. The historical trajectory in the active state includes a tracking trajectory newly established in the previous frame of image and a historical trajectory successfully matched with the detection box in the previous frame of image, and the historical trajectory in the inactive state corresponds to a continuous missing frame number less than or equal to a preset discard threshold and greater than a preset inactivation threshold. On this basis, the embodiment of the present application described that the detection box of the sitting posture target is matched with each historical trajectory for the first time to obtain the first tracking result, including the following S21 to S29:

[0169] S21: The detection box of the sitting posture target is divided into a first high-score detection box set and a first low-score detection box set based on a second score threshold.

[0170] The second score threshold can be preset, and the second score threshold can be between the first score threshold and the third score threshold. The second score threshold is used to measure whether the detection box of each sitting posture target belongs to a high-score detection box. Specifically, the detection box of the sitting posture target with a detection score greater than the second score threshold can be divided into the first high-score detection box set, and the detection box of the sitting posture target with a detection score less than or equal to the second score threshold can be divided into the first low-score detection box set. That is, the detection box in the first high-score detection box set corresponds to a detection score greater than the second score threshold, and the detection box in the first low-score detection box set corresponds to a detection score less than or equal to the second score threshold.

[0171] S22: Each historical trajectory is divided into a first active trajectory set and a first inactive trajectory set based on the state information of each historical trajectory.

[0172] Specifically, if the state information of a historical trajectory is in the active state, the historical trajectory is divided into the first active trajectory set. If the state information of a historical trajectory is in the inactive state, the historical trajectory is divided into the first inactive trajectory set. That is, the historical trajectory in the first active trajectory set is in the active state, and the historical trajectory in the first inactive trajectory set is in the inactive state.

[0173] S23: The detection box in the first high-score detection box set and the historical trajectory in the first active trajectory set are matched for the first time to obtain a first detection box matching successful sub-set, a first trajectory unmatched successful sub-set and a first detection box unmatched successful sub-set in the first tracking matching.

[0174] For convenience of description, the first high-score bounding box set is denoted as high_score1, the first activated track set is denoted as activated_stracks, the first time bounding box matching successful subset in the first track matching is denoted as M1-1, the first time track unmatched successful subset is denoted as U_T1-1, and the first time bounding box unmatched successful subset is denoted as U_D1-1.

[0175] M1-1 includes the bounding boxes in high_score1 that are matched successfully with the historical tracks in activated_stracks. U_D1-1 includes the bounding boxes in high_score1 that are not matched successfully with all the historical tracks in activated_stracks. U_T1-1 includes the historical tracks in activated_stracks that are not matched successfully with all the bounding boxes in high_score1.

[0176] For example, for each bounding box in high_score1, a first intersection over union between the bounding box and the predicted box corresponding to each historical track in activated_stracks is calculated, and a second intersection over union between the bounding box and the sitting position box corresponding to each historical track in activated_stracks is calculated. According to the first intersection over union and the second intersection over union, a weighted intersection over union between the bounding box and each historical track in activated_stracks is calculated, and the bounding box with a weighted intersection over union greater than a preset intersection over union threshold is added to M1-1.

[0177] For the bounding box in high_score1, if the weighted intersection over union between the bounding box and all the historical tracks in activated_stracks is less than or equal to the preset intersection over union threshold, the bounding box is added to U_D1-1.

[0178] For the historical track in activated_stracks, if the weighted intersection over union between the historical track and all the bounding boxes in high_score1 is less than or equal to the preset intersection over union threshold, the historical track is added to U_T1-1.

[0179] Optionally, the distance cost matrix between the bounding boxes in high_score1 and the historical tracks in activated_stracks can be calculated first, and then the M1-1, U_D1-1 and U_T1-1 in the first track matching can be obtained by using the Hungarian algorithm.

[0180] S24: performing second matching on the detection boxes in the first low-score detection box set and the historical trajectories in the second unsuccessful trajectory sub-set in the first tracking matching, to obtain a second successful detection box sub-set in the first tracking matching, a second unsuccessful trajectory sub-set in the first tracking matching, and a second unsuccessful detection box sub-set in the first tracking matching.

[0181] For ease of description, the first low-score detection box set is denoted as low_score1, the second successful detection box sub-set in the first tracking matching is denoted as M1-2, the second unsuccessful trajectory sub-set in the first tracking matching is denoted as U_T1-2, and the second unsuccessful detection box sub-set in the first tracking matching is denoted as U_D1-2.

[0182] M1-2 includes the detection boxes in low_score1 that are successfully matched with the historical trajectories in U_T1-1. U_D1-2 includes the detection boxes in low_score1 that are unsuccessfully matched with all the historical trajectories in U_T1-1. U_T1-2 includes the historical trajectories in U_T1-1 that are unsuccessfully matched with all the detection boxes in low_score1.

[0183] For example, for each detection box in low_score1, a first intersection-over-union between the detection box and the prediction box corresponding to each historical trajectory in U_T1-1 is calculated, and a second intersection-over-union between the detection box and the position box corresponding to each historical trajectory in U_T1-1 is calculated. According to the first intersection-over-union and the second intersection-over-union, a weighted intersection-over-union between the detection box and each historical trajectory in U_T1-1 is calculated, and the detection box with a weighted intersection-over-union greater than a preset intersection-over-union threshold is added to M1-2.

[0184] For the detection box in low_score1, if the weighted intersection-over-union between the detection box and all the historical trajectories in U_T1-1 is less than or equal to the preset intersection-over-union threshold, the detection box is added to U_D1-2.

[0185] For the historical trajectory in U_T1-1, if the weighted intersection-over-union between the historical trajectory and all the detection boxes in low_score1 is less than or equal to the preset intersection-over-union threshold, the historical trajectory is added to U_T1-2.

[0186] Optionally, a distance cost matrix between the detection boxes in low_score1 and the historical trajectories in U_T1-1 can be calculated first, and then the M1-2, U_D1-2, and U_T1-2 in the first tracking matching can be obtained by using the Hungarian algorithm.

[0187] S25: performing third matching on the bounding boxes in the third detection frame unmatched successful sub-set and the historical trajectories in the first non-activated trajectory set, to obtain a third detection frame matched successful sub-set in the first tracking matching, a third trajectory unmatched successful sub-set and a third detection frame unmatched successful sub-set.

[0188] For convenience of description, the first non-activated trajectory set is denoted as non_activated_stracks, the third detection frame matched successful sub-set in the first tracking matching is denoted as M1-3, the third trajectory unmatched successful sub-set is denoted as U_T1-3, and the third detection frame unmatched successful sub-set is denoted as U_D1-3.

[0189] M1-3 includes the bounding boxes in U_D1-1 matched successfully with the historical trajectories in non_activated_stracks. U_D1-3 includes the bounding boxes in U_D1-1 unmatched successfully with all the historical trajectories in non_activated_stracks. U_T1-3 includes the historical trajectories in non_activated_stracks unmatched successfully with all the bounding boxes in U_D1-3.

[0190] For example, for each bounding box in U_D1-1, a first intersection over union between the bounding box and the predicted box corresponding to each historical trajectory in non_activated_stracks is calculated, and a second intersection over union between the bounding box and the sitting position box corresponding to each historical trajectory in non_activated_stracks is calculated. According to the first intersection over union and the second intersection over union, a weighted intersection over union between the bounding box and each historical trajectory in non_activated_stracks is calculated, and the bounding box with a weighted intersection over union greater than a preset intersection over union threshold is added to M1-3.

[0191] For the bounding box in U_D1-1, if the weighted intersection over union between the bounding box and all the historical trajectories in non_activated_stracks is less than or equal to the preset intersection over union threshold, the bounding box is added to U_D1-3.

[0192] For the historical trajectory in non_activated_stracks, if the weighted intersection over union between the historical trajectory and all the bounding boxes in U_D1-1 is less than or equal to the preset intersection over union threshold, the historical trajectory is added to U_T1-3.

[0193] Optionally, the distance cost matrix between the detection boxes in U_D1-1 and the historical trajectories in non_activated_stracks can be calculated first, and then the M1-3, U_D1-3 and U_T1-3 in the first tracking matching can be obtained by using the Hungarian algorithm.

[0194] S26: The detection boxes in the second detection box unmatched sub-set and the historical trajectories in the third trajectory unmatched sub-set are matched for the fourth time to obtain a fourth detection box matched sub-set, a fourth trajectory unmatched sub-set and a fourth detection box unmatched sub-set in the first tracking matching.

[0195] For convenience of description, the fourth detection box matched sub-set in the first tracking matching is denoted as M1-4, the fourth trajectory unmatched sub-set is denoted as U_T1-4, and the fourth detection box unmatched sub-set is denoted as U_D1-4.

[0196] M1-4 includes the detection boxes in U_D1-2 that are matched with the historical trajectories in U_T1-3. U_D1-4 includes the detection boxes in U_D1-2 that are not matched with all the historical trajectories in U_T1-3. U_T1-4 includes the historical trajectories in U_T1-3 that are not matched with all the detection boxes in U_D1-2.

[0197] For example, for each detection box in U_D1-2, a first intersection over union between the detection box and the prediction box corresponding to each historical trajectory in U_T1-3 is calculated, and a second intersection over union between the detection box and the sitting position box corresponding to each historical trajectory in U_T1-3 is calculated. According to the first intersection over union and the second intersection over union, a weighted intersection over union between the detection box and each historical trajectory in U_T1-3 is calculated, and the detection box with a weighted intersection over union greater than a preset intersection over union threshold is added to M1-4.

[0198] For the detection box in U_D1-2, if the weighted intersection over union between the detection box and all the historical trajectories in U_T1-3 is less than or equal to the preset intersection over union threshold, the detection box is added to U_D1-4.

[0199] For the historical trajectory in U_T1-3, if the weighted intersection over union between the historical trajectory and all the detection boxes in U_D1-2 is less than or equal to the preset intersection over union threshold, the historical trajectory is added to U_T1-4.

[0200] Optionally, the distance cost matrix between the detection boxes in U_D1-2 and the historical trajectories in U_T1-3 can be calculated first, and then the M1-4, U_D1-4 and U_T1-4 in the first tracking matching can be obtained by using the Hungarian algorithm.

[0201] S27: Combining the first-time detection box matching success subset M1-1, the second-time detection box matching success subset M1-2, the third-time detection box matching success subset M1-3 and the fourth-time detection box matching success subset M1-4 in the first tracking matching to obtain a first detection box matching success set M1.

[0202] S28: Combining the second-time trajectory unmatched success subset U_T1-2 and the fourth-time trajectory unmatched success subset U_T1-4 in the first tracking matching to obtain a first trajectory unmatched success set U_T1.

[0203] S29: Combining the third-time detection box unmatched success subset U_D1-3 and the fourth-time detection box unmatched success subset U_D1-4 in the first tracking matching to obtain a first detection box unmatched success set U_D1.

[0204] To facilitate the understanding of the four times of matching involved in the first tracking matching, the following further describes in combination with FIG. 5. FIG. 5 is a schematic diagram of the four times of matching involved in the first tracking matching.

[0205] As shown in FIG. 5, based on the second score threshold, the detection boxes of the sit-dets are divided into a first high-score detection box set high_score1 and a first low-score detection box set low_score1. Based on the state information of each historical trajectory, each historical trajectory Stracks is divided into a first activated trajectory set activated_stracks and a first non-activated trajectory set non_activated_stracks.

[0206] The first-time matching in FIG. 5 is to match the detection boxes in high_score1 and the historical trajectories in activated_stracks to obtain M1-1, U_T1-1 and U_D1-1. The first-time matching is the first-time matching in S23 described above.

[0207] The second-time matching in FIG. 5 is to match the detection boxes in low_score1 and the historical trajectories in U_T1-1 to obtain M1-2, U_T1-2 and U_D1-2. The second-time matching is the second-time matching in S24 described above.

[0208] The third-time matching in FIG. 5 is to match the detection boxes in U_D1-1 and the historical trajectories in non_activated_stracks to obtain M1-3, U_T1-3 and U_D1-3. The third-time matching is the third-time matching in S25 described above.

[0209] The fourth matching in FIG. 5 is based on the bounding box in U_D1-2 and the historical trajectory in U_T1-3, and the fourth matching is performed to obtain M1-4, U_T1-4 and U_D1-4. The fourth matching is the fourth matching in S26 described above.

[0210] After the four matchings in FIG. 5 are completed, M1-1, M1-2, M1-3 and M1-4 are combined to obtain M1. U_T1-2 and U_T1-4 are combined to obtain U_T1. U_D1-3 and U_D1-4 are combined to obtain U_D1.

[0211] In combination with FIG. 3 and FIG. 5, M1 in FIG. 3 includes M1-1, M1-2, M1-3 and M1-4 in FIG. 5. U_T1 in FIG. 3 includes U_T1-2 and U_T1-4 in FIG. 5. U_D1 in FIG. 3 includes U_D1-3 and U_D1-4 in FIG. 5.

[0212] For example, the detection result further includes: a detection score corresponding to the bounding box; and historical trajectory information corresponding to each historical trajectory in the set of historical trajectories that are not matched successfully, including: state information of each historical trajectory. On this basis, the bounding box of the standing and posturing target, the bounding boxes in the set of bounding boxes that are not matched successfully, and the historical trajectories in the set of historical trajectories that are not matched successfully are subjected to second tracking matching to obtain a second tracking result, including the following S31 to S39:

[0213] S31: Based on a second score threshold, the bounding box of the standing and posturing target and the bounding boxes in the set of bounding boxes that are not matched successfully are divided into a second high-score bounding box set and a second low-score bounding box set.

[0214] Specifically, the bounding box of the standing and posturing target with a detection score greater than the second score threshold and the bounding boxes in U_D1 are divided into the second high-score bounding box set, and the bounding box of the standing and posturing target with a detection score less than or equal to the second score threshold and the bounding boxes in U_D1 are divided into the second low-score bounding box set. That is, the detection score of the bounding box in the second high-score bounding box set is greater than the second score threshold, and the detection score of the bounding box in the second low-score bounding box set is less than or equal to the second score threshold.

[0215] S32: Based on the state information of each historical trajectory in the set of historical trajectories that are not matched successfully, each historical trajectory in the set of historical trajectories that are not matched successfully is divided into a second active trajectory set and a second non-active trajectory set.

[0216] Specifically, if the state information of the first track does not match a certain historical track in the success set is an active state, the historical track is divided into the second active track set. If the state information of the first track does not match a certain historical track in the success set is an inactive state, the historical track is divided into the second inactive track set, that is, the historical track in the second active track set is in an active state, and the historical track in the second inactive track set is in an inactive state.

[0217] S33: performing first matching on the detection boxes in the second high-score detection box set and the historical tracks in the second active track set to obtain a first detection box matching success subset, a first track non-matching success subset and a first detection box non-matching success subset in the second tracking matching.

[0218] For convenience of description, the second high-score detection box set is denoted as high_score2, the second active track set is denoted as activated_U_stracks, the first detection box matching success subset in the second tracking matching is denoted as M2-1, the first track non-matching success subset is denoted as U_T2-1, and the first detection box non-matching success subset is denoted as U_D2-1.

[0219] M2-1 includes the detection boxes in high_score2 that match the historical tracks in activated_U_stracks. U_D2-1 includes the detection boxes in high_score2 that do not match any historical track in activated_U_stracks. U_T2-1 includes the historical tracks in activated_U_stracks that do not match any detection box in high_score2.

[0220] For example, for each detection box in high_score2, the intersection over union between the detection box and the prediction box corresponding to each historical track in activated_U_stracks is calculated, and the detection box with an intersection over union greater than a preset intersection over union threshold is added to M2-1.

[0221] For the detection box in high_score2, if the intersection over union between the detection box and the prediction box corresponding to each historical track in activated_U_stracks is less than or equal to the preset intersection over union threshold, the detection box is added to U_D2-1.

[0222] For each history track in the activated_U_stracks, if the intersection-over-union between the prediction bounding box corresponding to the history track and all the detection bounding boxes in the high_score2 is less than or equal to a preset intersection-over-union threshold, the history track is added to the U_T2-1.

[0223] S34: The detection bounding boxes in the second low-score detection bounding box set and the history tracks in the first track unmatched successful subset are matched for a second time to obtain a second track matching, a second time detection bounding box matched successful subset, a second time track unmatched successful subset and a second time detection bounding box unmatched successful subset.

[0224] For convenience of description, the second low-score detection bounding box set is denoted as low_score2, the second time detection bounding box matched successful subset in the second track matching is denoted as M2-2, the second time track unmatched successful subset is denoted as U_T2-2, and the second time detection bounding box unmatched successful subset is denoted as U_D2-2.

[0225] The M2-2 includes the detection bounding boxes in the low_score2 matched with the history tracks in the U_T2-1. The U_D2-2 includes the detection bounding boxes in the low_score2 not matched with all the history tracks in the U_T2-1. The U_T2-2 includes the history tracks in the U_T2-1 not matched with all the detection bounding boxes in the low_score2.

[0226] For example, for each detection bounding box in the low_score2, the intersection-over-union between the detection bounding box and the prediction bounding boxes corresponding to the history tracks in the U_T2-1 is calculated, and the detection bounding box with the intersection-over-union greater than a preset intersection-over-union threshold is added to the M1-2.

[0227] For the detection bounding box in the low_score2, if the intersection-over-union between the detection bounding box and the prediction bounding boxes corresponding to all the history tracks in the U_T2-1 is less than or equal to a preset intersection-over-union threshold, the detection bounding box is added to the U_D2-2.

[0228] For each history track in the U_T2-1, if the intersection-over-union between the prediction bounding box corresponding to the history track and all the detection bounding boxes in the low_score2 is less than or equal to a preset intersection-over-union threshold, the history track is added to the U_T2-2.

[0229] S35: The detection bounding boxes in the first time detection bounding box unmatched successful subset and the history tracks in the second non-activated track set are matched for a third time to obtain a third track matching, a third time detection bounding box matched successful subset, a third time track unmatched successful subset and a third time detection bounding box unmatched successful subset.

[0230] For the convenience of description, the second non-activated track set is denoted as non_activated_U_stracks, the third detection frame matched successfully sub-set in the second track matching is denoted as M2-3, the third track unmatched successfully sub-set is denoted as U_T2-3, and the third detection frame unmatched successfully sub-set is denoted as U_D2-3.

[0231] M2-3 includes the detection frame in U_D2-1 matched successfully with the historical track in non_activated_U_stracks. U_D2-3 includes the detection frame in U_D2-1 unmatched successfully with all historical tracks in non_activated_U_stracks. U_T2-3 includes the historical track in non_activated_U_stracks unmatched successfully with all detection frames in U_D2-3.

[0232] For example, for each detection frame in U_D2-1, the intersection-over-union between the detection frame and the prediction frame corresponding to each historical track in non_activated_U_stracks is calculated, and the detection frame with the intersection-over-union greater than the preset intersection-over-union threshold is added to M2-3.

[0233] For the detection frame in U_D2-1, if the intersection-over-union between the detection frame and the prediction frame corresponding to each historical track in non_activated_U_stracks is less than or equal to the preset intersection-over-union threshold, the detection frame is added to U_D2-3.

[0234] For the historical track in non_activated_U_stracks, if the intersection-over-union between the prediction frame corresponding to the historical track and each detection frame in U_D2-1 is less than or equal to the preset intersection-over-union threshold, the historical track is added to U_T2-3.

[0235] S36: The detection frame in the second detection frame unmatched successfully sub-set and the historical track in the third track unmatched successfully sub-set are matched for the fourth time to obtain the fourth detection frame matched successfully sub-set, the fourth track unmatched successfully sub-set and the fourth detection frame unmatched successfully sub-set in the second track matching.

[0236] For the convenience of description, the fourth detection frame matched successfully sub-set in the second track matching is denoted as M2-4, the fourth track unmatched successfully sub-set is denoted as U_T2-4, and the fourth detection frame unmatched successfully sub-set is denoted as U_D2-4.

[0237] M2-4 includes: the detection box in U_D2-2 that matches successfully with the historical trajectory in U_T2-3. U_D2-4 includes: the detection box in U_D2-2 that does not match successfully with all the historical trajectories in U_T2-3. U_T2-4 includes: the historical trajectory in U_T2-3 that does not match successfully with all the detection boxes in U_D2-2.

[0238] For each detection box in U_D2-2, the intersection over union between the detection box and the prediction box corresponding to each historical trajectory in U_T2-3 is calculated, and the detection box with the intersection over union greater than the preset intersection over union threshold is added to M2-4.

[0239] For the detection box in U_D2-2, if the intersection over union between the detection box and the prediction box corresponding to all the historical trajectories in U_T1-3 is less than or equal to the preset intersection over union threshold, the detection box is added to U_D2-4.

[0240] For the historical trajectory in U_T2-3, if the intersection over union between the prediction box corresponding to the historical trajectory and all the detection boxes in U_D2-2 is less than or equal to the preset intersection over union threshold, the historical trajectory is added to U_T2-4.

[0241] S37: Combining the first-time detection box matching successful sub-set M2-1, the second-time detection box matching successful sub-set M2-2, the third-time detection box matching successful sub-set M2-3 and the fourth-time detection box matching successful sub-set M2-4 in the second tracking matching to obtain a second detection box matching successful set M2.

[0242] S38: Combining the second-time trajectory unmatched successful sub-set U_T2-2 and the fourth-time trajectory unmatched successful sub-set U_T2-4 in the second tracking matching to obtain a second trajectory unmatched successful set U_T2.

[0243] S39: Combining the third-time detection box unmatched successful sub-set U_D2-3 and the fourth-time detection box unmatched successful sub-set U_D2-4 in the second tracking matching to obtain a second detection box unmatched successful set U_D2.

[0244] For the convenience of understanding the four times of matching involved in the above-mentioned second tracking matching, the following will be further described in combination with FIG. 6. FIG. 6 is a schematic diagram of the principle of the four times of matching involved in the second tracking matching.

[0245] As shown in FIG. 6, based on the second score threshold, the bounding boxes of the station pose class target Up-Dets and the bounding boxes in U_D1 are divided into a second high-score bounding box set high_score2 and a second low-score bounding box set low_score2. Based on the state information of each historical trajectory in the first trajectory unmatched successful set (shown in the figure as U_Stracks), each historical trajectory in U_Stracks is divided into a second activated trajectory set activated_U_stracks and a second non-activated trajectory set non_activated_U_stracks.

[0246] The first match in FIG. 6 is to match the bounding boxes in high_score2 and the historical trajectories in activated_U_stracks, to obtain the second tracking match M2-1, U_T2-1 and U_D2-1. This first match is the first match in S33 described above.

[0247] The second match in FIG. 6 is to match the bounding boxes in low_score2 and the historical trajectories in U_T2-1, to obtain the second tracking match M2-2, U_T2-2 and U_D2-2. This second match is the second match in S34 described above.

[0248] The third match in FIG. 6 is to match the bounding boxes in U_D2-1 and the historical trajectories in non_activated_U_stracks, to obtain the second tracking match M2-3, U_T2-3 and U_D2-3. This third match is the third match in S35 described above.

[0249] The fourth match in FIG. 6 is to match the bounding boxes in U_D2-2 and the historical trajectories in U_T2-3, to obtain the second tracking match M2-4, U_T2-4 and U_D2-4. This fourth match is the fourth match in S36 described above.

[0250] After completing the four matches in FIG. 6, M2-1, M2-2, M2-3 and M2-4 are combined to obtain M2. U_T2-2 and U_T2-4 are combined to obtain U_T2. U_D2-3 and U_D2-4 are combined to obtain U_D2.

[0251] In combination with FIG. 3 and FIG. 6, M2 in FIG. 3 includes M2-1, M2-2, M2-3 and M2-4 in FIG. 6. U_T2 in FIG. 3 includes U_T2-2 and U_T2-4 in FIG. 5. U_D2 in FIG. 3 includes U_D2-3 and U_D2-4 in FIG. 6.

[0252] In view of the fixed number of people in tracking scenarios such as a classroom and a conference room, the embodiments of the present application can also perform trajectory merging according to actual needs after trajectory updating. Based on this, after determining the tracking trajectories of each target in the current frame image according to the first tracking result and the second tracking result, the embodiments of the present application further include: determining the number of trajectories of the tracking trajectories present in the current frame image; if the number of trajectories is greater than a preset number, determining the tracking trajectories to be merged according to the tracking trajectories present in the current frame image, and performing trajectory merging on the tracking trajectories to be merged.

[0253] In combination with the foregoing description, the tracking trajectories of each target in the current frame image can be newly created tracking trajectories or can be a historical trajectory discarded, and therefore, the number of trajectories present in the current frame image can change. Therefore, the embodiments of the present application determine the number of trajectories of the tracking trajectories present in the current frame image after determining the tracking trajectories of each target in the current frame image according to the first tracking result and the second tracking result.

[0254] In a possible implementation, each tracking trajectory is provided with a unique trajectory identifier, so that the number of current existing trajectory identifiers can be determined as the number of trajectories corresponding to the current frame image. The preset number can be set according to actual needs, and in the embodiments of the present application, the preset number can be set as the maximum value of the number of detected targets in the current image frame and each image frame before the current image frame. If the number of trajectories is greater than the preset number, it is indicated that multiple trajectories can be established for a target, resulting in that the number of trajectories exceeds the number of targets, and therefore, when the number of trajectories is greater than the preset number, the tracking trajectories to be merged in the current existing tracking trajectories can be determined, and the tracking trajectories to be merged are subjected to trajectory merging.

[0255] The tracking trajectories to be merged are tracking trajectories of the same target, and by merging multiple tracking trajectories of the same target, the continuity of target tracking can be improved.

[0256] For example, determining the tracking trajectories to be merged according to the tracking trajectories present in the current frame image includes the following S41 to S44:

[0257] S41: traversing the tracking trajectories present in the current frame image, taking the currently traversed tracking trajectory as a candidate trajectory, and calculating the first coincidence degree in the time dimension and the second coincidence degree in the space dimension between the candidate trajectory and each reference trajectory.

[0258] In the embodiments of the present application, the current traversed tracking trajectory is taken as a candidate trajectory, and the other tracking trajectories in the current existing tracking trajectories except the candidate trajectory are taken as reference trajectories. For the current traversed candidate trajectory, the first coincidence degree in the time dimension and the second coincidence degree in the space dimension between the candidate trajectory and each reference trajectory can be calculated. The first coincidence degree represents the correlation degree between the candidate trajectory and the reference trajectory in the time dimension, and the second coincidence degree represents the correlation degree between the candidate trajectory and the reference trajectory in the space dimension.

[0259] In a possible implementation manner, the manner of calculating the first coincidence degree can include: determining the number of frames in which the candidate trajectory and each reference trajectory track to the same frame, and determining the total number of frames tracked by the candidate trajectory. Then, the ratio of the number of frames to the total number of frames is determined as the first coincidence degree in the time dimension between the candidate trajectory and each reference trajectory.

[0260] In the embodiments of the present application, the current traversed tracking trajectory is taken as a candidate trajectory, and the other tracking trajectories in the current existing tracking trajectories except the candidate trajectory are taken as reference trajectories. For the current traversed candidate trajectory, the first coincidence degree in the time dimension and the second coincidence degree in the space dimension between the candidate trajectory and each reference trajectory can be calculated. The first coincidence degree represents the correlation degree between the candidate trajectory and the reference trajectory in the time dimension, and the second coincidence degree represents the correlation degree between the candidate trajectory and the reference trajectory in the space dimension.

[0261] In a specific implementation, assuming that the current traversed candidate trajectory is the i th trajectory, and the reference trajectory is the j th trajectory, the first coincidence degree can be calculated by the following formula: T i o l = I n t e r s e c t i o n t ( i, j ) / L i, j ≠ i

[0262] wherein, Tiol represents the first coincidence degree, Intersection_t(i,j) represents the time intersection of the i-th trajectory and the j-th trajectory in the tracking frame number, and L_i is the tracking record length of the i-th trajectory. In combination with the foregoing description, Intersection_t(i,j) is the frame number in which the candidate trajectory and the reference trajectory track to the same frame, and L_i is the total frame number tracked by the candidate trajectory. In combination with the calculation manner of the first coincidence degree, it can be seen that the smaller the first coincidence degree is, the more likely the i-th trajectory and the j-th trajectory are merged in the time dimension.

[0263] In a possible implementation manner, the trajectory information of the currently existing tracking trajectory includes a sitting position box of a target tracked by the tracking trajectory in a sitting posture. On this basis, the manner of calculating the second coincidence degree can include: determining the intersection-over-union between the trajectory sitting position box corresponding to the candidate trajectory and the trajectory sitting position box corresponding to each reference trajectory as the second coincidence degree between the candidate trajectory and each reference trajectory in the spatial dimension.

[0264] In combination with the foregoing description, it can be known that if a target tracked by a trajectory exists in a sitting posture, the trajectory corresponds to a sitting position box, and therefore, if the target tracked by the candidate trajectory and the target tracked by the reference trajectory exist in a sitting posture, the candidate trajectory and the reference trajectory both correspond to a sitting position box.

[0265] The sitting position box corresponding to the candidate trajectory can be understood as a first target sitting position box obtained based on each sitting position box recorded when the target tracked by the candidate trajectory is in a sitting posture before the current frame image. The position coordinates of the first target sitting position box can be an average value of the position coordinates of each sitting position box recorded by the candidate trajectory.

[0266] The sitting position box corresponding to the reference trajectory can be understood as a second target sitting position box obtained based on each sitting position box recorded when the target tracked by the reference trajectory is in a sitting posture before the current frame image. The position coordinates of the second target sitting position box can be an average value of the position coordinates of each sitting position box recorded by the reference trajectory.

[0267] In a specific implementation, assuming that the candidate trajectory currently traversed is the i-th trajectory and the reference trajectory is the j-th trajectory, the second coincidence degree can be calculated by the following formula: Siou=IOU(sit_tl br_i,sit_tlbr_j),j≠i

[0268] wherein Siou represents the second coincidence degree, sit tlbr i represents the sitting position box corresponding to the i th trajectory, and sit tlbr j represents the sitting position box corresponding to the j th trajectory. The greater the second coincidence degree is, the more likely the i th trajectory and the j th trajectory are merged in the spatial dimension.

[0269] S42: According to the first coincidence degree and the second coincidence degree, a spatio-temporal coincidence degree between the candidate trajectory and each reference trajectory is calculated.

[0270] The smaller the first coincidence degree is, the more likely the candidate trajectory and the reference trajectory are merged in the time dimension, and the greater the second coincidence degree is, the more likely the candidate trajectory and the reference trajectory are merged in the spatial dimension. Therefore, in order to measure whether the candidate trajectory and the reference trajectory can be merged by a unified numerical value, in the embodiments of the present application, a spatio-temporal coincidence degree is further calculated according to the first coincidence degree and the second coincidence degree, so as to measure whether the candidate trajectory and the reference trajectory can be merged by the spatio-temporal coincidence degree.

[0271] For example, assuming that the candidate trajectory currently traversed is the i th trajectory, and the reference trajectory is the j th trajectory, the spatio-temporal coincidence degree between the i th trajectory and the j th trajectory can be calculated by the following formula: T Siou = (1-Tiol) x Siou

[0272] wherein T Siou represents the spatio-temporal coincidence degree, Tiol represents the first coincidence degree, and Siou represents the second coincidence degree.

[0273] S43: The maximum spatio-temporal coincidence degree in the spatio-temporal coincidence degrees between the candidate trajectory and each reference trajectory is determined, and a target reference trajectory having the maximum spatio-temporal coincidence degree with the candidate trajectory is determined among the reference trajectories.

[0274] It can be understood that the spatio-temporal coincidence degree is calculated between the candidate trajectory and each reference trajectory, and based on this, if there are 10 reference trajectories, the spatio-temporal coincidence degrees between the candidate trajectory and the 10 reference trajectories are calculated, i.e. 10 spatio-temporal coincidence degrees are obtained. Further, the maximum spatio-temporal coincidence degree can be selected from the 10 spatio-temporal coincidence degrees, and the reference trajectory having the maximum spatio-temporal coincidence degree with the candidate trajectory is determined as the target reference trajectory. The target reference trajectory is the tracking trajectory most likely to be merged with the candidate trajectory among the reference trajectories.

[0275] S44: If the maximum spatio-temporal coincidence degree is greater than a preset coincidence degree threshold, the candidate trajectory and the target reference trajectory are determined as the tracking trajectories to be merged.

[0276] The preset coincidence degree threshold can be preset as a criterion for whether two tracking trajectories can be merged. In the embodiment of the present application, the spatio-temporal coincidence degree between the candidate trajectory and the target reference trajectory, i.e., the maximum spatio-temporal coincidence degree, is greater than the preset coincidence degree threshold. If the spatio-temporal coincidence degree between the candidate trajectory and the target reference trajectory is less than or equal to the preset coincidence degree threshold, it is determined that the candidate trajectory and the target reference trajectory cannot be merged.

[0277] In a specific implementation, after the candidate trajectory and the target reference trajectory are merged, the state of the target reference trajectory can be updated to a deletion state, and the spatio-temporal coincidence degree between the reference trajectory in the deletion state and each of the tracking trajectories is set to 0, so as to avoid the target reference trajectory that has been merged from continuing to participate in the calculation of the spatio-temporal coincidence degree between the tracking trajectories that are subsequently traversed, thereby facilitating the saving of calculation amount.

[0278] In the embodiment of the present application, when the tracking trajectories are merged, the first coincidence degree in the time dimension and the second coincidence degree in the space dimension between different tracking trajectories are considered simultaneously, which is beneficial to accurately determine which tracking trajectories actually track the same target, thereby more accurately determining the tracking trajectories to be merged, which is beneficial to guarantee the continuity of the tracking trajectories after the sitting posture target disappears for a period of time and then returns to the seat.

[0279] To facilitate the understanding of the above trajectory merging process, further description is made in combination with FIG. 7. FIG. 7 is a schematic flowchart of trajectory merging in the embodiment of the present application.

[0280] For example, as shown in FIG. 7, the flowchart of trajectory merging includes:

[0281] Step 701: Determine the currently existing tracking trajectories Stracks.

[0282] Step 702: len(Stracks)≥max_num.

[0283] Specifically, len(Stracks) is the number of the currently existing tracking trajectories, and max_num is a preset number. In this step, if the number of the currently existing tracking trajectories is greater than or equal to the preset number, step 703 is performed, otherwise the flowchart ends.

[0284] Step 703: Take the tracking trajectories strack=Stracks[i] in sequence.

[0285] That is, the currently existing tracking trajectories are traversed in sequence to obtain the candidate trajectory that is currently traversed, and the candidate trajectory is denoted as trajectory i.

[0286] Step 704: Calculate the first coincidence degree Tiol in the time dimension between trajectory i and the remaining tracking trajectories.

[0287] Step 705: calculating a second coincidence degree Siol between the trajectory i and the rest of the tracking trajectories in the spatial dimension.

[0288] Step 706: calculating a spatio-temporal coincidence degree TSiol between the trajectory i and the rest of the tracking trajectories according to the first coincidence degree and the second coincidence degree.

[0289] Step 707: setting the spatio-temporal coincidence degree between the tracking trajectory in the deletion state and the rest of the tracking trajectories to 0.

[0290] That is, the spatio-temporal coincidence degree between the tracking trajectory in the deletion state and the rest of the tracking trajectories is all set to 0.

[0291] Step 708: selecting a maximum value, and recording the value and the index as tsiou and j respectively.

[0292] Specifically, the maximum value tsiou is the maximum spatio-temporal coincidence degree between the candidate trajectory (trajectory i) and the reference trajectories (the rest of the tracking trajectories). j is the identification of the target reference trajectory having the maximum spatio-temporal coincidence degree with the candidate trajectory.

[0293] Step 709: tsiou>th_ts.

[0294] Specifically, in this step, it is judged whether the maximum spatio-temporal coincidence degree tsiou is greater than a preset coincidence degree threshold th_ts. If yes, step 710 is executed, otherwise step 711 is executed.

[0295] Step 710: merging the trajectory i and the trajectory j, and updating the state of the trajectory j to the deletion state.

[0296] Step 711: i≥len(Stracks)-1.

[0297] Specifically, this step is equivalent to judging whether the trajectory i is the last trajectory in the current existing trajectories that has been traversed. If yes, the flow ends, otherwise i+=1 is set and step 703 is executed to continue taking the next tracking trajectory.

[0298] In the embodiment of the application, based on the characteristics of the fixed number of students in the classroom, a tracking trajectory merging strategy is designed. The merging of the tracking trajectories based on the spatio-temporal coincidence degree of the tracking trajectories can solve the problem of discontinuity of the tracking trajectory of a student who disappears from the picture for a period of time after going to the podium and then returns to the seat, which is conducive to realizing the continuous tracking of the student and avoiding the interruption of the trajectory.

[0299] To verify the advantages of the target tracking method (hereinafter referred to as ClassTracker) in the embodiments of the present application compared with the existing target tracking methods, the inventors of the present application conducted tests on the tracking test dataset CMOT24-10 of the self-built classroom scene. The test results can be seen in Table 1:

[0300] Table 1

[0301] As shown in Table 1, compared with DeepSort (Deep Learning-based Sort, a target tracking algorithm based on deep learning), the ClassTracker in the embodiments of the present application has a large increase in HOTA (Higher Order Tracking Accuracy, an overall performance index for evaluating a multi-target tracking algorithm) and IDF1 (Identity F1 Score, an index for evaluating the identity accuracy of a multi-target tracking algorithm), which are increased by 26.4 and 25.5 percentage points, respectively, and IDSW (Identity Switches, an index for measuring the number of identity exchanges in the tracking process of a multi-target tracking algorithm) is reduced by 52%. MOTA (Multiple Object Tracking Accuracy, an index for evaluating the accuracy of a multi-target tracking algorithm) is also improved to a certain extent. Compared with Bytetrack (a real-time multi-target tracking framework), the ClassTracker in the embodiments of the present application also has obvious improvements in HOTA, MOTA, IDF1 and other indexes, and IDSW is also significantly reduced. Therefore, the multi-target tracking method provided in the embodiments of the present application greatly improves the accuracy and precision of multi-target tracking.

[0302] FIG. 8 is a structural schematic diagram of a target tracking device provided in an embodiment of the present application.

[0303] For example, as shown in FIG. 8, the target tracking device 800 includes:

[0304] The detection module 801 is configured to perform target detection on a current frame image to obtain detection results of targets, wherein the detection results include detection boxes and postures of the targets, and the postures are standing postures or sitting postures.

[0305] The first tracking module 802 is configured to determine a first tracking result according to the detection box of a sitting posture target and historical trajectories, wherein the sitting posture target is a target with a sitting posture among the targets, and the historical trajectories are trajectories obtained by tracking targets in a previous frame image.

[0306] The second tracking module 803 is configured to determine a second tracking result according to the bounding box of the standing posture target and each remaining historical trajectory; the standing posture target is a target with a standing posture among the targets; and the remaining historical trajectory includes a historical trajectory other than the historical trajectory tracking the sitting posture target among the historical trajectories.

[0307] The determining module 804 is configured to determine a tracking trajectory of each target in the current frame image according to the first tracking result and the second tracking result.

[0308] In a possible implementation, the first tracking module 802 is specifically configured to perform first tracking matching on the bounding box of the sitting posture target and each historical trajectory to obtain a first tracking result; the first tracking result includes a first bounding box matching success set, a first trajectory non-matching success set, and a first bounding box non-matching success set; and the remaining historical trajectory includes a historical trajectory in the first trajectory non-matching success set. The second tracking module 803 is specifically configured to perform second tracking matching on the bounding box of the standing posture target, the bounding box in the first bounding box non-matching success set, and the historical trajectory in the first trajectory non-matching success set to obtain a second tracking result; the second tracking result includes a second bounding box matching success set, a second trajectory non-matching success set, and a second bounding box non-matching success set.

[0309] In a possible implementation, the detection result further includes a detection score corresponding to the bounding box; and the determining module 804 is specifically configured to: for each bounding box in the first bounding box matching success set and each bounding box in the second bounding box matching success set, update a historical trajectory matched with the bounding box according to the bounding box to obtain a tracking trajectory of a target located in the bounding box in the current frame image; determine whether a detection score corresponding to each bounding box in the second bounding box non-matching success set is greater than a first score threshold; and for the bounding box with the detection score greater than the first score threshold, newly create a trajectory for tracking the target in the bounding box to obtain the tracking trajectory of the target located in the bounding box in the current frame image.

[0310] In a possible implementation, the target tracking apparatus 800 further includes an updating module configured to add 1 to a continuous loss frame number corresponding to each historical trajectory in the second trajectory non-matching success set to obtain an updated continuous loss frame number corresponding to each historical trajectory; if the updated continuous loss frame number corresponding to each historical trajectory is greater than a preset discard threshold, discard the historical trajectory; and if the updated continuous loss frame number corresponding to each historical trajectory is less than or equal to the preset discard threshold and greater than a preset inactivation threshold, update state information of the historical trajectory to an inactivation state.

[0311] In a possible implementation, the target tracking apparatus 800 further includes a prediction module, configured to predict, according to historical trajectory information corresponding to each of the historical trajectories, a prediction box corresponding to each of the historical trajectories in the current frame image by using a state estimator; the first tracking module 802 is specifically configured to perform first tracking matching between the detection box of the standing posture target and each of the historical trajectories according to the prediction box corresponding to each of the historical trajectories in the current frame image, to obtain a first tracking result; and the second tracking module 803 is specifically configured to perform second tracking matching between the detection box of the standing posture target, the detection box in the first detection box unmatched success set and the historical trajectory in the first trajectory unmatched success set according to the prediction box corresponding to the historical trajectory in the first trajectory unmatched success set in the current frame image, to obtain a second tracking result.

[0312] In a possible implementation, the historical trajectory information corresponding to each of the historical trajectories includes a standing posture position box of a target tracked by each of the historical trajectories in a standing posture; the first tracking module 802 is specifically configured to: for each of the detection boxes of the standing posture target, calculate a first intersection over union between the detection box and the prediction box corresponding to each of the historical trajectories, and calculate a second intersection over union between the detection box and the standing posture position box corresponding to each of the historical trajectories; calculate a weighted intersection over union between the detection box and each of the historical trajectories according to the first intersection over union and the second intersection over union; if the weighted intersection over union between the detection box and any one of the historical trajectories is greater than or equal to a preset intersection over union threshold, add the detection box to the first detection box matched success set; if the weighted intersection over union between the detection box and each of the historical trajectories is less than the preset intersection over union threshold, add the detection box to the first detection box unmatched success set; and if the weighted intersection over union between the historical trajectory and all the detection boxes of the standing posture target is less than the preset intersection over union threshold, add the historical trajectory to the first trajectory unmatched success set.

[0313] In a possible implementation, the detection result further includes a detection score corresponding to the detection frame; the historical trajectory information corresponding to each historical trajectory includes state information of each historical trajectory; the first tracking module 802 is specifically configured to: divide the detection frames of the sitting and standing target into a first high-score detection frame set and a first low-score detection frame set based on a second score threshold; the detection frames in the first high-score detection frame set correspond to a detection score greater than the second score threshold, and the detection frames in the first low-score detection frame set correspond to a detection score less than or equal to the second score threshold; divide each historical trajectory into a first active trajectory set and a first non-active trajectory set based on the state information of each historical trajectory; the historical trajectories in the first active trajectory set are in an active state, and the historical trajectories in the first non-active trajectory set are in a non-active state; perform first matching on the detection frames in the first high-score detection frame set and the historical trajectories in the first active trajectory set, to obtain a first detection frame matching success subset, a first trajectory non-matching success subset, and a first detection frame non-matching success subset in the first tracking matching; perform second matching on the detection frames in the first low-score detection frame set and the historical trajectories in the first trajectory non-matching success subset, to obtain a second detection frame matching success subset, a second trajectory non-matching success subset, and a second detection frame non-matching success subset in the first tracking matching; perform third matching on the detection frames in the first detection frame non-matching success subset and the historical trajectories in the first non-active trajectory set, to obtain a third detection frame matching success subset, a third trajectory non-matching success subset, and a third detection frame non-matching success subset in the first tracking matching; perform fourth matching on the detection frames in the second detection frame non-matching success subset and the historical trajectories in the third trajectory non-matching success subset, to obtain a fourth detection frame matching success subset, a fourth trajectory non-matching success subset, and a fourth detection frame non-matching success subset in the first tracking matching; combine the first detection frame matching success subset, the second detection frame matching success subset, the third detection frame matching success subset, and the fourth detection frame matching success subset in the first tracking matching, to obtain the first detection frame matching success set; combine the second trajectory non-matching success subset and the fourth trajectory non-matching success subset in the first tracking matching, to obtain the first trajectory non-matching success set; and combine the third detection frame non-matching success subset and the fourth detection frame non-matching success subset in the first tracking matching, to obtain the first detection frame non-matching success set.

[0314] In a possible implementation, the detection result further includes: a detection score corresponding to the detection frame; and the historical trajectory information corresponding to each historical trajectory in the first trajectory unsuccessful matching set includes: state information of each historical trajectory; the second tracking module 803 is specifically configured to: based on a second score threshold, divide the detection frame of the standing posture target and the detection frame in the first detection frame unsuccessful matching set into a second high-score detection frame set and a second low-score detection frame set; the detection frame in the second high-score detection frame set corresponds to a detection score greater than the second score threshold, and the detection frame in the second low-score detection frame set corresponds to a detection score less than or equal to the second score threshold; based on the state information of each historical trajectory in the first trajectory unsuccessful matching set, divide each historical trajectory in the first trajectory unsuccessful matching set into a second active trajectory set and a second non-active trajectory set; the historical trajectory in the second active trajectory set is in an active state, and the historical trajectory in the second non-active trajectory set is in a non-active state; perform first matching on the detection frame in the second high-score detection frame set and the historical trajectory in the second active trajectory set, to obtain a first detection frame successful matching sub-set, a first trajectory unsuccessful matching sub-set, and a first detection frame unsuccessful matching sub-set in the second tracking matching; perform second matching on the detection frame in the second low-score detection frame set and the historical trajectory in the first trajectory unsuccessful matching sub-set, to obtain a second detection frame successful matching sub-set, a second trajectory unsuccessful matching sub-set, and a second detection frame unsuccessful matching sub-set in the second tracking matching; perform third matching on the detection frame in the first detection frame unsuccessful matching sub-set and the historical trajectory in the second non-active trajectory set, to obtain a third detection frame successful matching sub-set, a third trajectory unsuccessful matching sub-set, and a third detection frame unsuccessful matching sub-set in the second tracking matching; perform fourth matching on the detection frame in the second detection frame unsuccessful matching sub-set and the historical trajectory in the third trajectory unsuccessful matching sub-set, to obtain a fourth detection frame successful matching sub-set, a fourth trajectory unsuccessful matching sub-set, and a fourth detection frame unsuccessful matching sub-set in the second tracking matching; combine the first detection frame successful matching sub-set, the second detection frame successful matching sub-set, the third detection frame successful matching sub-set, and the fourth detection frame successful matching sub-set in the second tracking matching, to obtain the second detection frame successful matching set; combine the second trajectory unsuccessful matching sub-set and the fourth trajectory unsuccessful matching sub-set in the second tracking matching, to obtain the second trajectory unsuccessful matching set; and combine the third detection frame unsuccessful matching sub-set and the fourth detection frame unsuccessful matching sub-set in the second tracking matching, to obtain the second detection frame unsuccessful matching set.

[0315] In a possible implementation, the historical trajectory information corresponding to each historical trajectory includes: a sitting position frame of a target tracked by each historical trajectory when the target is in a sitting position; each historical trajectory corresponds to an initialized state estimator; the prediction module is specifically configured to: for each historical trajectory in the historical trajectories, determine a tracking interval corresponding to the historical trajectory; the tracking interval represents a distance between a latest tracked image frame of the historical trajectory and the current frame image; in a case where the tracking interval is greater than a first preset interval, reinitialize the state estimator corresponding to the historical trajectory by using the sitting position frame corresponding to the historical trajectory, and predict a prediction frame corresponding to the historical trajectory in the current frame image by using the reinitialized state estimator; in a case where the tracking interval is greater than a second preset interval and less than or equal to the first preset interval, reinitialize the state estimator corresponding to the historical trajectory by using a target detection frame corresponding to the historical trajectory, and predict a prediction frame corresponding to the historical trajectory in the current frame image by using the reinitialized state estimator; the second preset interval is less than the first preset interval, and the target detection frame is a detection frame corresponding to the historical trajectory in a latest tracked image frame of the historical trajectory; in a case where the tracking interval is less than or equal to the second preset interval, predict a prediction frame corresponding to the historical trajectory in the current frame image by using the initialized state estimator corresponding to the historical trajectory.

[0316] In a possible implementation, the target tracking apparatus 800 further includes a trajectory merging module configured to: after determining the tracking trajectories of each target in the current frame image according to the first tracking result and the second tracking result, determine a number of trajectories of the tracking trajectories present in the current frame image; if the number of trajectories is greater than a preset number, determine tracking trajectories to be merged according to the tracking trajectories present in the current frame image, and perform trajectory merging on the tracking trajectories to be merged.

[0317] In a possible implementation, the trajectory merging module is specifically configured to: traverse the tracking trajectories present in the current frame image, take a currently traversed tracking trajectory as a candidate trajectory, and calculate a first coincidence degree in a time dimension and a second coincidence degree in a space dimension between the candidate trajectory and each reference trajectory; each reference trajectory includes a tracking trajectory present in the current frame image and different from the candidate trajectory; calculate a spatio-temporal coincidence degree between the candidate trajectory and each reference trajectory according to the first coincidence degree and the second coincidence degree; determine a maximum spatio-temporal coincidence degree in the spatio-temporal coincidence degrees between the candidate trajectory and each reference trajectory, and determine a target reference trajectory having the maximum spatio-temporal coincidence degree with the candidate trajectory in each reference trajectory; and if the maximum spatio-temporal coincidence degree is greater than a preset coincidence degree threshold, determine the candidate trajectory and the target reference trajectory as tracking trajectories to be merged.

[0318] In a possible implementation, the trajectory information of the tracking trajectory present in the current frame image includes a sitting position box of a target tracked by the tracking trajectory in a sitting posture; and the trajectory merging module is specifically configured to: determine a number of frames in which the candidate trajectory and each reference trajectory track the same frame, and determine a total number of frames tracked by the candidate trajectory; determine a ratio of the number of frames to the total number of frames as the first coincidence degree in the time dimension between the candidate trajectory and each reference trajectory; and determine an intersection-over-union between the sitting position box corresponding to the candidate trajectory and the sitting position box corresponding to each reference trajectory as the second coincidence degree in the space dimension between the candidate trajectory and each reference trajectory.

[0319] In a possible implementation, the target tracking apparatus 800 further includes an initialization module configured to, in a case where the current frame image is a first frame image, filter, from detection results of the current frame image, a detection box having a detection score greater than a third score threshold; and initialize a tracking trajectory corresponding to the current frame image according to the detection box having the detection score greater than the third score threshold; wherein the trajectory information of each initialized tracking trajectory includes a current posture of a target tracked by the tracking trajectory, a tracking frame number corresponding to the tracking trajectory, a current detection box of the target tracked by the tracking trajectory, a state estimator configured for the tracking trajectory, and a sitting position box corresponding to the tracking trajectory recorded in a case where the current posture is a sitting posture.

[0320] FIG. 9 is a structural schematic diagram of an electronic device provided by an embodiment of the present application.

[0321] As shown in FIG. 9, the electronic device 900 includes a memory 901 and a processor 902, wherein the memory 901 stores executable program code 9011, and the processor 902 is configured to invoke and execute the executable program code 9011 to perform a target tracking method.

[0322] In addition, the embodiments of the present application also protect a device, which can include a memory and a processor, wherein the memory stores executable program code, and the processor is configured to invoke and execute the executable program code to perform a target tracking method provided by the embodiments of the present application.

[0323] The embodiments can divide the device into functional modules according to the above method examples, for example, corresponding to each functional module, or two or more functions can be integrated into one processing module, and the integrated module can be implemented in the form of hardware. It should be noted that the division of modules in the embodiments is illustrative, and is only a logical function division. In actual implementation, another division mode can be used.

[0324] In the case of dividing each functional module corresponding to each function, the device can further include a detection module, a first tracking module, a second tracking module, and a determination module, etc. It should be noted that all related contents of each step involved in the above method embodiments can be referred to the function description of the corresponding functional module, which will not be repeated here.

[0325] It should be understood that the device provided by the embodiments is used to perform the above target tracking method, and thus the same effect as the above implementation method can be achieved.

[0326] In the case of using an integrated unit, the device can include a processing module and a storage module. When the device is applied to an electronic device, the processing module can be used to control and manage the actions of the electronic device. The storage module can be used to support the electronic device to execute related program codes, etc.

[0327] The processing module can be a processor or a controller, which can realize or execute various exemplary logical blocks, modules and circuits shown in combination with the disclosure of the present application. The processor can also be a combination of computing functions, such as one or more microprocessor combinations, digital signal processing (DSP) and microprocessor combinations, etc. The storage module can be a memory.

[0328] In addition, the device provided by the embodiments of the present application can be a chip, a component or a module. The chip can include a connected processor and a memory. The memory is used to store instructions, and when the processor invokes and executes the instructions, the chip can perform the target tracking method provided by the above embodiments.

[0329] The embodiment further provides a computer readable storage medium, which stores computer program codes, and when the computer program codes are run on a computer, the computer is caused to execute the above-mentioned related method steps to realize the target tracking method provided by the above-mentioned embodiment.

[0330] The embodiment further provides a computer program product, which, when run on a computer, causes the computer to execute the above-mentioned related steps to realize the target tracking method provided by the above-mentioned embodiment. The apparatus, the computer readable storage medium, the computer program product or the chip provided by the embodiment are used to execute the corresponding method provided above, and thus the beneficial effects achieved by the apparatus, the computer readable storage medium, the computer program product or the chip can refer to the beneficial effects in the corresponding method provided above, which will not be described here.

[0331] Through the description of the above embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above.

[0332] In the embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are only schematic, and the division of the modules or units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection between the apparatuses or units can be indirect coupling or communication connection through some interfaces, and can be electrical, mechanical or other forms.

[0333] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited to this, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A target tracking method characterized by, The method comprises: target detection is performed on a current frame image to obtain detection results of each target; wherein the detection results comprise detection boxes and postures of each target, and the postures are standing postures or sitting postures; a first tracking result is determined according to the detection boxes of the sitting posture targets and each historical trajectory; wherein the sitting posture targets are targets with a sitting posture among the targets; and each historical trajectory is a trajectory obtained by tracking a target in a previous frame image; a second tracking result is determined according to the detection boxes of the standing posture targets and each remaining historical trajectory; wherein the standing posture targets are targets with a standing posture among the targets; and each remaining historical trajectory comprises historical trajectories other than the historical trajectories tracking the sitting posture targets among the historical trajectories; tracking trajectories of each target in the current frame image are determined according to the first tracking result and the second tracking result.

2. The method of claim 1, wherein, The first tracking result is determined according to the detection boxes of the sitting posture targets and each historical trajectory, and comprises: first tracking matching is performed on the detection boxes of the sitting posture targets and each historical trajectory to obtain the first tracking result; wherein the first tracking result comprises a first detection box matching success set, a first trajectory non-matching success set and a first detection box non-matching success set; and the each remaining historical trajectory comprises historical trajectories in the first trajectory non-matching success set. The second tracking result is determined according to the detection boxes of the standing posture targets, the detection boxes in the first detection box non-matching success set and the historical trajectories in the first trajectory non-matching success set, and comprises: second tracking matching is performed on the detection boxes of the standing posture targets, the detection boxes in the first detection box non-matching success set and the historical trajectories in the first trajectory non-matching success set to obtain the second tracking result; wherein the second tracking result comprises a second detection box matching success set, a second trajectory non-matching success set and a second detection box non-matching success set.

3. The method of claim 2, wherein, The detection result further comprises a detection score corresponding to the detection box; The tracking trajectories of each target in the current frame image are determined according to the first tracking result and the second tracking result, and comprise: for each detection box in the first detection box matching success set and each detection box in the second detection box matching success set, a historical trajectory matched with the detection box is updated according to the detection box to obtain a tracking trajectory of a target located in the detection box in the current frame image; it is judged whether the detection score corresponding to each detection box in the second detection box non-matching success set is greater than a first score threshold; for a detection box with a detection score greater than the first score threshold, a new trajectory is created for tracking a target in the detection box to obtain a tracking trajectory of a target located in the detection box in the current frame image.

4. The method of claim 2, wherein, After the second tracking matching is performed on the detection boxes of the standing posture targets, the detection boxes in the first detection box non-matching success set and the historical trajectories in the first trajectory non-matching success set to obtain the second tracking result, the method further comprises: a continuous loss frame number corresponding to each historical trajectory in the second trajectory non-matching success set is added by 1 to obtain an updated continuous loss frame number corresponding to each historical trajectory. if the updated number of consecutive lost frames corresponding to each of the historical trajectories is greater than a preset discard threshold, discarding the historical trajectories; if the updated number of consecutive lost frames corresponding to each of the historical trajectories is less than or equal to the preset discard threshold and greater than a preset inactivation threshold, updating state information of the historical trajectories to an inactivated state.

5. The method of claim 2, wherein, After the target detection on the current frame image is performed to obtain the detection result of each target, the method further comprises: predicting, by a state estimator, a prediction box corresponding to each of the historical trajectories in the current frame image according to historical trajectory information corresponding to each of the historical trajectories; the first tracking matching of the detection box of the sitting posture target and each of the historical trajectories to obtain the first tracking result comprises: the first tracking matching of the detection box of the sitting posture target and each of the historical trajectories according to the prediction box corresponding to each of the historical trajectories in the current frame image to obtain the first tracking result; the second tracking matching of the detection box of the standing posture target, the detection box in the first detection box unmatched successful set and the historical trajectory in the first trajectory unmatched successful set to obtain the second tracking result comprises: the second tracking matching of the detection box of the standing posture target, the detection box in the first detection box unmatched successful set and the historical trajectory in the first trajectory unmatched successful set according to the prediction box corresponding to the historical trajectory in the first trajectory unmatched successful set in the current frame image to obtain the second tracking result.

6. The method of claim 5, wherein, the historical trajectory information corresponding to each of the historical trajectories comprises a sitting posture position box of a target tracked by each of the historical trajectories in a sitting posture; the first tracking matching of the detection box of the sitting posture target and each of the historical trajectories according to the prediction box corresponding to each of the historical trajectories in the current frame image to obtain the first tracking result comprises: for each of the detection boxes of the sitting posture target, calculating a first intersection over union between the detection box and the prediction box corresponding to each of the historical trajectories, and calculating a second intersection over union between the detection box and the sitting posture position box corresponding to each of the historical trajectories; calculating a weighted intersection over union between the detection box and each of the historical trajectories according to the first intersection over union and the second intersection over union; if the weighted intersection over union between the detection box and any one of the historical trajectories is greater than or equal to a preset intersection over union threshold, adding the detection box to the first detection box matched successful set; if the weighted intersection over union between the detection box and each of the historical trajectories is less than the preset intersection over union threshold, adding the detection box to the first detection box unmatched successful set; if the weighted intersection over union between the historical trajectory and all the detection boxes of the sitting posture target is less than the preset intersection over union threshold, adding the historical trajectory to the first trajectory unmatched successful set.

7. The method according to any one of claims 2 to 4, characterized in that, the detection result further comprises a detection score corresponding to the detection box; and the historical trajectory information corresponding to each of the historical trajectories comprises state information of each of the historical trajectories. the first tracking matching of the detection box of the sitting posture target and each of the historical trajectories to obtain the first tracking result comprises: dividing, based on a second score threshold, the detection boxes of the sitting posture type target into a first high-score detection box set and a first low-score detection box set, wherein a detection box in the first high-score detection box set corresponds to a detection score greater than the second score threshold, and a detection box in the first low-score detection box set corresponds to a detection score less than or equal to the second score threshold; dividing, based on state information of each historical trajectory, each historical trajectory into a first active trajectory set and a first non-active trajectory set, wherein a historical trajectory in the first active trajectory set is in an active state, and a historical trajectory in the first non-active trajectory set is in a non-active state; performing first-time matching on the detection boxes in the first high-score detection box set and the historical trajectories in the first active trajectory set to obtain a first-time detection box matching successful sub-set, a first-time trajectory non-matching successful sub-set, and a first-time detection box non-matching successful sub-set in the first tracking matching; performing second-time matching on the detection boxes in the first low-score detection box set and the historical trajectories in the first-time trajectory non-matching successful sub-set to obtain a second-time detection box matching successful sub-set, a second-time trajectory non-matching successful sub-set, and a second-time detection box non-matching successful sub-set in the first tracking matching; performing third-time matching on the detection boxes in the first-time detection box non-matching successful sub-set and the historical trajectories in the first non-active trajectory set to obtain a third-time detection box matching successful sub-set, a third-time trajectory non-matching successful sub-set, and a third-time detection box non-matching successful sub-set in the first tracking matching; performing fourth-time matching on the detection boxes in the second-time detection box non-matching successful sub-set and the historical trajectories in the third-time trajectory non-matching successful sub-set to obtain a fourth-time detection box matching successful sub-set, a fourth-time trajectory non-matching successful sub-set, and a fourth-time detection box non-matching successful sub-set in the first tracking matching; combining the first-time detection box matching successful sub-set, the second-time detection box matching successful sub-set, the third-time detection box matching successful sub-set, and the fourth-time detection box matching successful sub-set in the first tracking matching to obtain a first detection box matching successful set; combining the second-time trajectory non-matching successful sub-set and the fourth-time trajectory non-matching successful sub-set in the first tracking matching to obtain a first trajectory non-matching successful set; combining the third-time detection box non-matching successful sub-set and the fourth-time detection box non-matching successful sub-set in the first tracking matching to obtain a first detection box non-matching successful set.

8. The method according to any one of claims 2 to 4, characterized in that, The detection result further includes: a detection score corresponding to the detection box; and historical trajectory information corresponding to each historical trajectory in the first trajectory non-matching successful set includes state information of each historical trajectory; The second tracking matching on the detection boxes of the sitting posture type target, the detection boxes in the first detection box non-matching successful set, and the historical trajectories in the first trajectory non-matching successful set to obtain a second tracking result includes: divide, based on the second score threshold, the detection boxes of the station posture class target and the detection boxes in the first detection box unmatched successful set into a second high-score detection box set and a second low-score detection box set; wherein the detection boxes in the second high-score detection box set correspond to detection scores greater than the second score threshold, and the detection boxes in the second low-score detection box set correspond to detection scores less than or equal to the second score threshold; divide, based on the state information of each historical trajectory in the first trajectory unmatched successful set, each historical trajectory in the first trajectory unmatched successful set into a second active trajectory set and a second non-active trajectory set; wherein the historical trajectories in the second active trajectory set are in an active state, and the historical trajectories in the second non-active trajectory set are in a non-active state; perform first-time matching on the detection boxes in the second high-score detection box set and the historical trajectories in the second active trajectory set, to obtain a first-time detection box matched successful sub-set, a first-time trajectory unmatched successful sub-set and a first-time detection box unmatched successful sub-set in the second tracking matching; perform second-time matching on the detection boxes in the second low-score detection box set and the historical trajectories in the first-time trajectory unmatched successful sub-set, to obtain a second-time detection box matched successful sub-set, a second-time trajectory unmatched successful sub-set and a second-time detection box unmatched successful sub-set in the second tracking matching; perform third-time matching on the detection boxes in the first-time detection box unmatched successful sub-set and the historical trajectories in the second non-active trajectory set, to obtain a third-time detection box matched successful sub-set, a third-time trajectory unmatched successful sub-set and a third-time detection box unmatched successful sub-set in the second tracking matching; perform fourth-time matching on the detection boxes in the second-time detection box unmatched successful sub-set and the historical trajectories in the third-time trajectory unmatched successful sub-set, to obtain a fourth-time detection box matched successful sub-set, a fourth-time trajectory unmatched successful sub-set and a fourth-time detection box unmatched successful sub-set in the second tracking matching; combine the first-time detection box matched successful sub-set, the second-time detection box matched successful sub-set, the third-time detection box matched successful sub-set and the fourth-time detection box matched successful sub-set in the second tracking matching, to obtain the second detection box matched successful set; combine the second-time trajectory unmatched successful sub-set and the fourth-time trajectory unmatched successful sub-set in the second tracking matching, to obtain the second trajectory unmatched successful set; combine the third-time detection box unmatched successful sub-set and the fourth-time detection box unmatched successful sub-set in the second tracking matching, to obtain the second detection box unmatched successful set.

9. The method of claim 5, wherein, the historical trajectory information corresponding to each historical trajectory includes a sitting posture position box of a target tracked by each historical trajectory in a sitting posture; each historical trajectory corresponds to an initialized state estimator; The predicted box corresponding to each historical trajectory in the current frame image is predicted by a state estimator according to historical trajectory information corresponding to each historical trajectory, and the historical trajectory information comprises: For each historical trajectory in the historical trajectories, a tracking interval corresponding to the historical trajectory is determined, wherein the tracking interval represents a distance between a latest tracked image frame of the historical trajectory and the current frame image; In a case where the tracking interval is greater than a first preset interval, a state estimator corresponding to the historical trajectory is reinitialized by using a sitting position box corresponding to the historical trajectory, and a predicted box corresponding to the historical trajectory in the current frame image is predicted by using the reinitialized state estimator; In a case where the tracking interval is greater than a second preset interval and less than or equal to the first preset interval, a state estimator corresponding to the historical trajectory is reinitialized by using a target detection box corresponding to the historical trajectory, and a predicted box corresponding to the historical trajectory in the current frame image is predicted by using the reinitialized state estimator; the second preset interval is less than the first preset interval, and the target detection box is a detection box corresponding to the historical trajectory in a latest tracked image frame of the historical trajectory; In a case where the tracking interval is less than or equal to the second preset interval, a predicted box corresponding to the historical trajectory in the current frame image is predicted by using the initialized state estimator corresponding to the historical trajectory.

10. The method according to any one of claims 1 to 4, characterized in that, After the tracking trajectories of each target in the current frame image are determined according to the first tracking result and the second tracking result, the method further comprises: A number of tracking trajectories existing in the current frame image is determined; If the number of tracking trajectories is greater than a preset number, tracking trajectories to be merged are determined according to the tracking trajectories existing in the current frame image, and the tracking trajectories to be merged are merged.

11. The method of claim 10, wherein, The tracking trajectories to be merged are determined according to the tracking trajectories existing in the current frame image, and the method comprises: Each tracking trajectory existing in the current frame image is traversed, a currently traversed tracking trajectory is taken as a candidate trajectory, a first coincidence degree in a time dimension and a second coincidence degree in a space dimension between the candidate trajectory and each reference trajectory are calculated; each reference trajectory comprises a tracking trajectory other than the candidate trajectory among the tracking trajectories existing in the current frame image; A spatio-temporal coincidence degree between the candidate trajectory and each reference trajectory is calculated according to the first coincidence degree and the second coincidence degree; A maximum spatio-temporal coincidence degree among the spatio-temporal coincidence degrees between the candidate trajectory and each reference trajectory is determined, and a target reference trajectory having the maximum spatio-temporal coincidence degree with the candidate trajectory is determined among each reference trajectory; If the maximum spatio-temporal coincidence degree is greater than a preset coincidence degree threshold, the candidate trajectory and the target reference trajectory are determined as tracking trajectories to be merged.

12. The method of claim 11, wherein, The trajectory information of the tracking trajectory existing in the current frame image comprises a sitting position box of a target tracked by the tracking trajectory in a sitting state. The calculating the first coincidence degree between the candidate trajectory and each reference trajectory in the time dimension and the second coincidence degree in the space dimension comprises: determining the number of frames in which the candidate trajectory and each reference trajectory track to the same frame, and determining the total number of frames tracked by the candidate trajectory; determining the ratio of the number of frames to the total number of frames as the first coincidence degree between the candidate trajectory and each reference trajectory in the time dimension; determining the intersection-over-union between the position box corresponding to the candidate trajectory and the position box corresponding to each reference trajectory as the second coincidence degree between the candidate trajectory and each reference trajectory in the space dimension.

13. The method of claim 1, wherein, The method further comprises: in the case that the current frame image is a first frame image, screening, from the detection results of the current frame image, a detection box with a detection score greater than a third score threshold; initializing a tracking trajectory corresponding to the current frame image according to the detection box with the detection score greater than the third score threshold; wherein the track information of each initialized tracking trajectory comprises: the current pose of the target tracked by the tracking trajectory, the tracking frame number corresponding to the tracking trajectory, the current detection box of the target tracked by the tracking trajectory, a state estimator configured for the tracking trajectory, and a position box corresponding to the tracking trajectory recorded in the case that the current pose is a sitting pose.

14. A target tracking device, characterized by comprise: a detection module configured to perform target detection on a current frame image to obtain detection results of targets, wherein the detection results comprise detection boxes and poses of the targets, and the poses are standing poses or sitting poses; a first tracking module configured to determine first tracking results according to detection boxes of sitting-class targets and historical trajectories, wherein the sitting-class targets are targets with sitting poses among the targets, and the historical trajectories are trajectories obtained by performing target tracking on a previous frame image; a second tracking module configured to determine second tracking results according to detection boxes of standing-class targets and remaining historical trajectories, wherein the standing-class targets are targets with standing poses among the targets, and the remaining historical trajectories comprise historical trajectories other than the historical trajectories tracking the sitting-class targets among the historical trajectories; a determination module configured to determine tracking trajectories of the targets in the current frame image according to the first tracking results and the second tracking results.

15. An electronic device, comprising: The electronic device comprises: a memory configured to store executable program code; a processor configured to call and run the executable program code from the memory, so that the electronic device performs the method according to any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed, implements the method according to any one of claims 1 to 13.