A pedestrian target tracking method and device in a complex scenario

By establishing the movement trajectory of pedestrian targets in complex scenarios, and combining Kalman linear prediction and feature similarity matching methods, the problem of pedestrian target tracking in complex scenarios is solved, improving the tracking effect and accuracy.

CN113822163BActive Publication Date: 2025-06-10BEIJING ZIYAN LIANHE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110981011.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-25
Publication Date
2025-06-10
Estimated Expiration
2041-08-25

AI Technical Summary

Technical Problem

In complex scenarios, it is difficult for the prior art to effectively track pedestrian targets, especially when the characteristics of the target person do not change significantly and the target crosses, resulting in target loss and poor tracking effect.

Method used

By selecting the target person in the video, establishing his or her motion trajectory, and using the preset trajectory disappearance time threshold to determine whether the target is lost. If not lost, Kalman linear prediction is used to calculate the predicted position of the target and form a candidate set through distance screening. Extract face features and embedded features from the candidate set, perform similarity matching, and update the target motion trajectory.

Benefits of technology

Effective tracking of pedestrian targets in complex scenarios is achieved, the tracking accuracy and stability of target characters is improved, and the situation of target loss is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113822163B_ABST
    Figure CN113822163B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian target tracking method in a complex scenario. The method includes selecting a target person, obtaining the starting moment of the target image, and establishing a motion trajectory; calculating the time interval between the moment when the target image is detected and the last update moment of the motion trajectory, and if the time interval exceeds a preset trajectory disappearance time threshold, the target has been lost; otherwise, calculating the distance between the predicted target image position information and the actual target image position information, screening the targets that meet the distance condition to form a candidate set, extracting the target face features and embedded features, and performing person similarity matching. The present invention realizes the tracking and positioning of the target person in a short time interval, combines the facial information and the upper body physical characteristics of the target person for auxiliary discrimination and association, and uses effective features to verify and correct the matching result, which is beneficial to improving the tracking effect of the target person in a complex scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a pedestrian target tracking method and device in a complex scenario. Background Art

[0002] The target tracking of key figures has very important application value. In the real world, the motion scenarios are relatively complex, and simple linear motion models or traditional target tracking technologies are all inadequate. In some application scenarios, pedestrian targets overlap, and the characteristics of pedestrians change insignificantly, resulting in the crossing and mis-matching of the trajectories of two target figures, and ultimately the loss of the target figures. This will bring a bad experience to the intelligent monitoring and tracking of target figures.

[0003] Although single-target tracking based on deep learning and other correlation filtering has also achieved relatively good tracking effects, the efficiency is still not good, and the tracking effect for pedestrian targets in complex scenarios is not very good either.

[0004] Facial features and embedded features of re-identification of identities have achieved very good results in the fields of biometric recognition and identity discrimination. Using these stable features for auxiliary target figure tracking is an effective technical approach. Summary of the Invention

[0005] The present invention provides a pedestrian target tracking method in a complex scenario, including:

[0006] Select a target figure, obtain the starting moment of the target image, and establish a motion trajectory for the target figure;

[0007] Calculate the time interval between the moment when the target image is detected next time and the last update moment of the motion trajectory, and determine whether this time interval is within a preset trajectory disappearance time threshold:

[0008] If this time interval exceeds the preset trajectory disappearance time threshold, it means that the target figure has been lost for too long, and it is determined that the target has been lost;

[0009] If this time interval is within the preset trajectory disappearance time threshold, then calculate the distance between the predicted target image position information and the actual target image position information, screen the targets that meet the distance condition to form a candidate set, extract the target face features and embedded features from the candidate set, perform person similarity matching. If the target figure can be matched, update the motion trajectory of the target figure, otherwise complete the current tracking and wait for the target figure in the next frame.

[0010] The pedestrian target tracking method in the complex scenario as described above. When tracking a certain target person in the complex scenario, first find the starting time t of the target image of this target person in the video. This moment represents obtaining a new detected target, and a new motion trajectory Trackobj is established for this target, and record the target image at this time as image t , and its position is position obj_t (x obj_t , y obj_t , width obj_t , height obj_t ). The image at time t + k is denoted as image t+k , and the position is position obj_t+k (x obj_t+k , y obj_t+k , width obj_t+k , height obj_t+k ). k is a natural number related to video frame skipping processing, and k ∈ {1, 2,..}.

[0011] The pedestrian target tracking method in the complex scenario as described above. If this time interval is within the preset trajectory disappearance time threshold, then specifically execute the following sub-steps:

[0012] Predict the position of the target image position corresponding to the last update time t in the motion trajectory obj_t to obtain the predicted position position predict at time t + k;

[0013] Calculate the distance between the predicted position position predict at time t + k and the actual position position obj_t+k ;

[0014] Form a set of candidate objects objSet{obj0, obj1,..., objm} at time t + k that meet the distance condition.

[0015] The pedestrian target tracking method in the complex scenario as described above. Predict the position of the target image position corresponding to the last update time t in the motion trajectory obj_t , specifically: Adopt Kalman linear prediction, and calculate the Kalman state parameters according to the current frame position position obj_t+k and the position information of the newly updated target position obj_t . Predict the position and size position predict (x pre , ypre , width pre , height pre )。

[0016] The pedestrian target tracking method in the complex scenario as described above, wherein the Kalman state estimation uses an 8-dimensional space to characterize the state of the trajectory at a certain moment respectively representing the center position, aspect ratio, height of the target, and the corresponding speed information in the image coordinates; the Kalman filter adopts a constant velocity model and a linear observation model, and the corresponding observation variables are the center position (x, y), aspect ratio r, and height h, and finally the predicted position position is obtained predict (x pre , y pre , width pre , height pre ), where x pre = x, y pre = y, width pre = h * r, height pre = h.

[0017] The pedestrian target tracking method in the complex scenario as described above, wherein the predicted position position at time t + k is calculated predict and the distance between the predicted position position and the actual position position obj_t+k is calculated by the formula:

[0018]

[0019] where xcen predict = x pre + width pre * 0.5, ycen predict = y pre + height pre * 0.5, if the position distance dist(position predict , position obj_t+k ) < 1.5 * (width pret + width obj_t+k ), then the target that meets the distance condition is considered, saved to the candidate set objSet, and the target image image at time t + k is determined by screening through the distance condition t+k and whether the target image image at time t and the target image image t are in the same target set objSet.

[0020] The pedestrian target tracking method in the complex scenario as described above, wherein the embedded features refer to the comprehensive features of the target person's clothing, face, and hairstyle.

[0021] The pedestrian target tracking method in the complex scenario as described above, wherein, for person similarity matching, it specifically includes the following sub-steps:

[0022] Perform face detection on all members of the target image candidate set objSet at time t + k;

[0023] If the face region exists, extract the face feature FaceFea corresponding to time t + k t+k , calculate the face feature FaceFea corresponding to the target person object at time t t and the face feature FaceFea corresponding to time t + k t+k The similarity simi(FaceFea t , FaceFea t+k ), if the similarity is greater than the given threshold faceThr, then it is considered the same target person, and the target person's movement trajectory is updated; otherwise, this tracking is completed and wait for the target person in the next frame;

[0024] If the face region does not exist, extract the embedded feature ReidFea t+k , calculate the embedded feature ReidFea corresponding to the target person object at time t t and the embedded feature ReidFea corresponding to time t + k t+k The similarity simi(ReidFea t , ReidFea t+k ), if the similarity threshold is greater than the given threshold AttriThr, then it is considered the same target person, and the target person's movement trajectory is updated; otherwise, this tracking is completed and wait for the target person in the next frame.

[0025] The pedestrian target tracking method in the complex scenario as described above, wherein, the calculation formula for similarity is:

[0026]

[0027] Wherein, the target feature of the buffered trajectory is fea0(fea0 1 , fea0 2 ,..., fea0 N ), the feature of the target obj new is fea1(fea1 1 , fea1 2 ,..., fea1 N ), N is the dimension of the corresponding feature; when using this formula to calculate similarity, the target feature fea0 represents the face feature FaceFea corresponding to time t t, fea1 represents the face feature FaceFea corresponding to the time t + k t+k ; When using this formula to calculate the similarity, the target feature fea0 represents the embedded feature ReidFea corresponding to the time t t , fea1 represents the embedded feature ReidFea corresponding to the time t + k t+k .

[0028] The pedestrian target tracking method in the complex scenario as described above, wherein, if there are multiple targets with similarity greater than the given threshold faceThr among all members of the target image candidate set objSet, then select the target with the highest similarity as the movement trajectory of the target person, and add the corresponding target object record to the trajectory Trackobj to complete the matching.

[0029] The beneficial effects achieved by the present invention are as follows: The present invention proposes an improved method for real-time tracking by means of multiple feature attribute constraints. In the scenario of target person tracking and positioning within a short time interval, the face information and the upper body embedded feature of the target person are combined for auxiliary discrimination and association, and the results of effective feature pair matching are verified and corrected, which is beneficial to improving the tracking effect of the target person in the complex scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.

[0031] Figure 1 is a flowchart of a pedestrian target tracking method in a complex scenario provided by Embodiment 1 of the present invention;

[0032] Figure 2 is a flowchart of the method executed when it is determined that the time interval is within the preset trajectory disappearance time threshold;

[0033] Figure 3 is a flowchart of the method for matching the similarity of people;

[0034] Figure 4 is a schematic diagram of a pedestrian target tracking device in a complex scenario provided by Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] Combined with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present invention.

[0036] Embodiment 1

[0037] As Figure 1 shown, Embodiment 1 of the present invention provides a pedestrian target tracking method in a complex scenario, including:

[0038] Step 110: Select a target person, obtain the starting moment of the target image, and establish a motion trajectory for the target person;

[0039] When tracking a certain target person in a complex scenario, first find the starting moment t of the target image of the target person in the video. This moment represents obtaining a new detection target, and a new motion trajectory Track obj is established for this target. Record the target image at this time as image t , and its position is position obj_t . The image at time t + k is denoted as image t+k , and the position is position obj_t+k .

[0040] It is set that for any target object, its position is expressed as position(x, y, width, height), where x and y are the coordinates of the upper left corner of the target bounding rectangle, which are used to represent the position of the target, and width and height are the width and height of the corresponding rectangle, which are used to represent the size of the target. Therefore, the position of the target person object corresponding to time t is position obj_t (x obj_t , y obj_t , width obj_t , height obj_t ). Similarly, the image position at time t + k is position obj_t+k (x obj_t+k , y obj_t+k , width obj_t+k , height obj_t+k ). k is related to video frame skipping processing and is a natural number, k ∈ {1, 2,..}. The larger k is, the smaller the corresponding algorithm operation amount is, but it will affect the tracking experience. The present invention preferably selects k = 6.

[0041] Step 120: Calculate the time interval between the moment when the target image is next detected and the last update moment of the motion trajectory, and determine whether this time interval is within the preset trajectory disappearance time threshold. If this time interval exceeds the preset trajectory disappearance time threshold, it means that the target person has been lost for too long and the target is determined to be lost; if this time interval is within the preset trajectory disappearance time threshold, then execute Step 130;

[0042] In the embodiment of the present invention, for the target image image detected at the moment t + k t+k , the position is position obj_t+k , calculate the time interval between the moment t + k and the time of the last update of the motion trajectory, and determine whether the time interval is within the time threshold endTrackThr of the trajectory disappearance (preferably, endTrackThr can be set to 60). If the time interval is greater than this threshold, the person target has been lost for too long, the target is lost, and the tracking ends;

[0043] As Figure 2 shown, if the time interval is within the time threshold, then execute the following sub-steps:

[0044] Step 210: Predict the position position of the target image corresponding to the last update moment t in the motion trajectory obj_t to obtain the predicted position position at the moment t + k predict ;

[0045] Preferably, the present invention adopts Kalman linear prediction. According to the current frame position position obj_t+k and the position information position of the newly updated target obj_t , calculate the Kalman state parameters, and predict the position and size position predict (x pre , y pre , width pre , height pre ) where the target should appear at the next moment through the Kalman state parameters;

[0046] Among them, the specific method of Kalman prediction is:

[0047] The Kalman state estimation uses an 8-dimensional space to describe the state of the trajectory at a certain moment which respectively represent the center position, aspect ratio, height, and speed information corresponding in the image coordinates of the target; the Kalman filter adopts a uniform motion model and a linear observation model, and the corresponding observation variables are the center position (x, y), aspect ratio r, height h, and finally obtain the predicted position position predict (xpre , y pre , width pre , height pre ), where x pre = x, y pre = y, width pre = h * r, height pre = h. The Kalman prediction performance is relatively stable and can reliably predict the position where the next-frame target appears, which is conducive to improving the trajectory connection effect.

[0048] Step 220: Calculate the predicted position position at time t + k predict and the actual position position obj_t+k at time t + k, and form a set of candidate objects objSet{obj0, obj1,..., objm} for the targets that meet the distance condition at time t + k;

[0049] Specifically, the formula for calculating the distance between the predicted position position predict at time t + k and the actual position position obj_t+k is as follows:

[0050]

[0051] where xcen predict = x pre + width pre * 0.5, ycen predict = y pre + height pre * 0.5. If the position distance dist(position predict , position obj_t+k ) < 1.5 * (width pre + width obj_t+k ), then it is considered a target that meets the distance condition and is saved in the candidate set objSet. By screening through the distance condition, it can be determined whether the target image image t+k at time t + k and the target image image t at time t belong to the same target set objSet.

[0052] If the candidate set objSet obtained after distance condition screening is empty, continue with target tracking for the next frame. Additionally, calculate the time interval timeInterval since the last update of the distance trajectory Trackobj, and determine whether the interval timeInterval is greater than the time threshold for trajectory disappearance. Until the interval timeInterval is greater than the trajectory disappearance time threshold endTrackThr, end this trajectory and establish a new trajectory.

[0053] Return to Figure 1 , step 130: Calculate the distance between the predicted target image position information and the actual target image position information, screen the targets that meet the distance condition to form a candidate set, extract the target face features and embedded features from the candidate set, perform person similarity matching. If the target person can be matched, update the target person's movement trajectory; otherwise, complete this tracking and wait for the target person in the next frame.

[0054] The embedded feature refers to the comprehensive features of the target person's clothing, face, hairstyle, etc. This application comprehensively considers the face features and embedded features to achieve accurate tracking and positioning of the target task.

[0055] In the embodiments of the present invention, as Figure 3 shown, extract the target face features and embedded features from the candidate set and perform person similarity matching, which specifically includes the following sub-steps:

[0056] Step 310: Perform face detection on all members of the target image candidate set objSet at time t + k. If a face region exists, execute step 320; otherwise, execute step 330.

[0057] Step 320: Extract the face feature FaceFea corresponding to time t + k t+k , calculate the similarity simi(FeaceFea t corresponding to the target person object at time t and the face feature FeaceFea t+k corresponding to time t + k t ,FeaceFea t+k ). If the similarity is greater than the given threshold faceThr (preferably, the threshold faceThr = 0.7), then it is considered the same target person, and update the target person's movement trajectory; otherwise, complete this tracking and wait for the target person in the next frame.

[0058] If there are multiple targets in all members of the target image candidate set objSet with similarity greater than the given threshold faceThr, then select the target with the highest similarity as the movement trajectory of the target person, add the corresponding target object record to the trajectory Trackobj, and complete the matching.

[0059] Step 330: Extract the embedded feature ReidFea t+k (Preferably extract the upper body attribute features), calculate the similarity simi(ReidFea t+k of the embedded feature ReidFeat corresponding to the target person object at time t and the embedded feature ReidFea corresponding to time t + k t , ReidFea t+k ). If the similarity threshold is greater than the given threshold AttriThr (preferably AttriThr = 0.75), then it is considered the same target person, and the movement trajectory of the target person is updated; otherwise, this tracking is completed, and wait for the target person in the next frame.

[0060] Specifically, the calculation formula for similarity is:

[0061]

[0062] where the target feature of the buffered trajectory is fea0(fea0 1 , fea0 2 ,..., fea0 N ), the feature of the target obj new is fea1(fea1 1 , fea1 2 ,..., fea1 N ), and N is the dimension of the corresponding feature; when using this formula to calculate similarity in step 320, the target feature fea0 represents the face feature FaceFea t corresponding to time t, and fea1 represents the face feature FaceFea t+k corresponding to time t + k; when using this formula to calculate similarity in step 330, the target feature fea0 represents the embedded feature ReidFea t corresponding to time t, and fea1 represents the embedded feature ReidFea t+k corresponding to time t + k;

[0063] If there are multiple targets in all members of the target image candidate set objSet with similarity greater than the given threshold AttriThr, then select the target with the highest similarity as the movement trajectory of the target person, add the corresponding target object record to the trajectory Trackobj, complete the matching, and thus complete the tracking, waiting for the input of a new target in the next frame.

[0064] Example Two

[0065] As Figure 4 shown, Example Two of the present invention provides a pedestrian target tracking device 40 in a complex scenario, including a motion trajectory creation module 41, a target motion trajectory detection module 42, and a target motion trajectory determination module 43;

[0066] The motion trajectory creation module 41 is used to select a target person, obtain the starting moment of the target image, and establish a motion trajectory for the target person;

[0067] Specifically, when tracking a certain target person in a complex scenario, first find the starting moment t of the target image of the target person in the video. This moment represents obtaining a new detection target. Re-establish a motion trajectory Trackobj for the target, and record the target image at this time as image t , and its position is position obj_t (x obj_t , y obj_t , width obj_t , height obj_t ). The image at the moment t + k is recorded as image t+k , and the position is position obj_t+k (x obj_t+k , y obj_t+k , width obj_t+k , height obj_t+k ). k is related to the video frame skipping process and is a natural number, k ∈ {1, 2,..}.

[0068] The target motion trajectory detection module 42 is used to calculate the time interval between the moment when the target image is detected next time and the last update moment of the motion trajectory, and determine whether this time interval is within the preset trajectory disappearance time threshold: If this time interval exceeds the preset trajectory disappearance time threshold, it means that the target person has been lost for too long and the target is determined to be lost; If this time interval is within the preset trajectory disappearance time threshold, the target motion trajectory determination module 43 is triggered;

[0069] Specifically, if this time interval is within the preset trajectory disappearance time threshold, it specifically includes: predicting the position position obj_t corresponding to the last update moment t in the motion trajectory to obtain the predicted position position predict at the moment t + k; calculating the predicted position position predict at the moment t + k and the actual position position obj_t+kThe distance between; form a candidate set objSet{obj0, obj1,..., objm} with several targets at time t + k that meet the distance condition.

[0070] Among them, for the target image position position corresponding to the last update time t in the motion trajectory obj_t Make a prediction, specifically: adopt Kalman linear prediction, according to the current frame position position obj_t+k And the position information position of the newly updated target obj_t Calculate the Kalman state parameters, and predict the position and size postion where the target should appear at the next moment through the Kalman state parameters predict (x pre , y pre , width pre , height pre ).

[0071] The target motion trajectory determination module 43 is used to calculate the distance between the predicted target image position information and the actual target image position information, screen the targets that meet the distance condition to form a candidate set, extract the target face features and embedded features from the candidate set, perform person similarity matching, if the target person can be matched, then update the target person's motion trajectory, otherwise complete this tracking and wait for the target person in the next frame.

[0072] Specifically, for person similarity matching, it specifically includes: perform face detection on all members of the target image candidate set objSet at time t + k; if the face area exists, then extract the face feature FaceFea corresponding to time t + k t+k , calculate the similarity simi(FaceFea t corresponding to the target person object at time t and the face feature FaceFea t+k corresponding to time t + k t , FaceFea t+k ), if the similarity is greater than the given threshold faceThr (preferably, the threshold faceThr = 0.7), then it is considered the same target person, and the target person's motion trajectory is updated, otherwise this tracking is completed and wait for the target person in the next frame; if the face area does not exist, then extract the embedded feature ReidFea t+k (preferably extract the upper body embedded feature), calculate the similarity simi(ReidFea t corresponding to the target person object at time t and the embedded feature ReidFea t+k corresponding to time t + k t,ReidFea t+k ), if the similarity threshold is greater than the given threshold AttriThr (preferably AttriThr = 0.75), then it is considered the same target person, and the movement trajectory of the target person is updated; otherwise, this tracking is completed, and wait for the target person in the next frame.

[0073] Corresponding to the above embodiments, an embodiment of the present invention provides a computer storage medium, including: at least one memory and at least one processor;

[0074] The memory is used to store one or more program instructions;

[0075] The processor is used to run one or more program instructions to execute the pedestrian target tracking method in complex scenarios.

[0076] Corresponding to the above embodiments, an embodiment of the present invention provides a computer-readable storage medium, and the computer storage medium contains one or more program instructions, and the one or more program instructions are used to be executed by the processor for the pedestrian target tracking method in complex scenarios.

[0077] An embodiment disclosed by the present invention provides a computer-readable storage medium, and computer program instructions are stored in the computer-readable storage medium. When the computer program instructions run on a computer, the computer is caused to execute the above-mentioned pedestrian target tracking method in complex scenarios.

[0078] In an embodiment of the present invention, the processor may be an integrated circuit chip with signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0079] It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. The processor reads the information in the storage medium and combines its hardware to complete the steps of the above method.

[0080] The storage medium can be a memory, for example, it can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.

[0081] Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory.

[0082] The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).

[0083] The storage medium described in the embodiments of the present invention is intended to include but not be limited to these and any other suitable types of memories.

[0084] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by a combination of hardware and software. When applying software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0085] The specific embodiments described above further elaborate on the object, technical solution and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention shall be included in the protection scope of the present invention.

Claims

1. A pedestrian target tracking method in a complex scenario, characterized in that, it includes: Select a target person, obtain the starting moment of the target image, and establish a motion trajectory for the target person; Calculate the time interval between the moment when the target image is detected next time and the last update moment of the motion trajectory, and determine whether this time interval is within the preset trajectory disappearance time threshold: If this time interval exceeds the preset trajectory disappearance time threshold, it means that the target person has been lost for too long and the target is determined to be lost; If this time interval is within the preset trajectory disappearance time threshold, calculate the distance between the predicted position and the actual position of the target image, screen the targets that meet the distance condition to form a candidate set, extract the target face features and embedded features from the candidate set, perform person similarity matching. If the target person can be matched, update the motion trajectory of the target person, otherwise complete this tracking and wait for the target person in the next frame; The embedded feature refers to the comprehensive feature of the target person's clothing, face and hairstyle; Specifically, according to the actual position position of the target image corresponding to the last update time t in the motion trajectory obj_t perform prediction to obtain the predicted position position of the target image at time t + k predict , where k is a natural number related to video frame skipping processing, k ∈ {1, 2,..}; calculate the predicted position position of the target image at time t + k predict and the actual position position of the target image at time t + k obj_t+k calculate the distance between them; form a set of candidates objSet{obj0, obj1,..., objm} consisting of several targets at time t + k that meet the distance condition; Among them, Kalman linear prediction is adopted, and according to position obj_t+k and position obj_t calculate the Kalman state parameters, and predict the predicted position position of the target image through the Kalman state parameters predict (x pre , y pre , width pre , height pre ); Calculate the predicted position position of the target image at time t + k predict and the actual position position of the target image at time t + k obj_t+k The calculation formula for the distance between them is as follows: where xcen predict = x pre + width pre * 0.5, ycen predict = y pre + height pre * 0.5, if the distance dist(position predict , position obj_t+k ) < 1.5 * (width pre + width obj_t+k ), width obj_t+k is the width of the bounding matrix of the target image at the actual position position obj_t+k in the target image, then it is considered as a target that meets the distance condition and is saved to the candidate set.

2. The pedestrian target tracking method in a complex scenario according to claim 1, characterized in that, Performing person similarity matching specifically includes the following sub-steps: Perform face detection on all members of the target image candidate set objSet at time t + k; If the face region exists, extract the face feature FaceFea corresponding to the moment t + k t+k , calculate the face feature FaceFea corresponding to the target person object at the moment t t and the face feature FaceFea corresponding to the moment t + k t+k of the similarity simi(FaceFea t , FaceFea t+k ). If the similarity is greater than the given threshold faceThr, then it is considered the same target person, and the movement trajectory of the target person is updated. Otherwise, this tracking is completed, and the target person in the next frame is waited for; If the face region does not exist, extract the embedded feature ReidFea t+k , calculate the embedded feature ReidFea corresponding to the target person object at time t t and the embedded feature ReidFea corresponding to time t + k t+k The similarity simi(ReidFea t , ReidFea t+k ). If the similarity threshold is greater than the given threshold AttriThr, then it is considered the same target person, and the target person's movement trajectory is updated. Otherwise, this tracking is completed, and wait for the target person in the next frame.

3. A pedestrian target tracking device in a complex scenario, characterized in that, it includes: A motion trajectory creation module (21), a target motion trajectory detection module (22), and a target motion trajectory determination module (23); The motion trajectory creation module (21) is used to select a target person, obtain the starting moment of the target image, and establish a motion trajectory for the target person; The target motion trajectory detection module (22) is used to calculate the time interval between the moment when the target image is detected next time and the last update moment of the motion trajectory, and determine whether this time interval is within the preset trajectory disappearance time threshold: If this time interval exceeds the preset trajectory disappearance time threshold, it means that the target person has been lost for too long and the target is determined to be lost; If this time interval is within the preset trajectory disappearance time threshold, trigger the target motion trajectory determination module 23; The target motion trajectory determination module (23) is used to calculate the distance between the predicted position and the actual position of the target image, screen the targets that meet the distance condition to form a candidate set, extract the target face features and embedded features from the candidate set, perform person similarity matching. If the target person can be matched, update the motion trajectory of the target person, otherwise complete this tracking and wait for the target person in the next frame; The embedded feature refers to the comprehensive feature of the target person's clothing, face and hairstyle; Specifically, according to the actual position position of the target image corresponding to the last update time t in the motion trajectory obj_t perform prediction to obtain the predicted position position of the target image at time t + k predict , where k is a natural number related to video frame skipping processing, k ∈ {1, 2,..}; calculate the predicted position position of the target image at time t + k predict and the actual position position of the target image at time t + k obj_t+k calculate the distance between them; form a set of candidates objSet{obj0, obj1,..., objm} for several targets at time t + k that meet the distance condition; Among them, Kalman linear prediction is adopted, and according to position obj_t+k and position obj_t calculate the Kalman state parameters, and predict the predicted position position of the target image through the Kalman state parameters predict (x pre , y pre , width pre , height pre ); Calculate the predicted position position of the target image at time t + k predict and the actual position position of the target image at time t + k obj_t+k The formula for calculating the distance between them is as follows: where xcen predict = x pre + width pre * 0.5, ycen predict = y pre + height pre * 0.5, if the distance dist(position predict , position obj_t+k ) < 1.5 * (width pre + width obj_t+k ), width obj_t+k is the width of the bounding matrix of the target image at the actual position position obj_t+k in the target image, then it is considered as a target that meets the distance condition and is saved to the candidate set.

4. The pedestrian target tracking device in a complex scenario according to claim 3, characterized in that, The target motion trajectory determination module (23) is specifically used for: performing face detection on all members of the target image candidate set objSet at time t + k; if a face region exists, extracting the face feature FaceFea corresponding to time t + k t+k , calculating the face feature FaceFea corresponding to the target person object at time t t and the face feature FaceFea corresponding to time t + k t+k for similarity simi(FaceFea t , FaceFea t+k ). If the similarity is greater than the given threshold faceThr, then it is considered the same target person, and the target person's motion trajectory is updated; otherwise, this tracking is completed, and the target person in the next frame is awaited. If the face region does not exist, the embedded feature ReidFea is extracted t+k , calculating the embedded feature ReidFea corresponding to the target person object at time t t and the embedded feature ReidFea corresponding to time t + k t+k for similarity simi(ReidFea t , ReidFea t+k ). If the similarity threshold is greater than the given threshold AttriThr, then it is considered the same target person, and the target person's motion trajectory is updated; otherwise, this tracking is completed, and the target person in the next frame is awaited.

5. A computer storage medium, characterized in that, it includes: At least one memory and at least one processor; The memory is used to store one or more program instructions; The processor is used to run one or more program instructions to execute the pedestrian target tracking method in a complex scenario according to any one of claims 1-2.

Citation Information

Patent Citations

  • Multi-camera linkage multi-target tracking method and system for smart community

    CN110619657A

  • Monitoring video pedestrian recognition and tracking method and device and storage medium

    CN112257502A