Deepsort-based multi-object face tracking method and apparatus, computer device, and medium
By determining the trajectory of the upper body detection frame in the multi-objective tracking algorithm DeepSort and associating the ID of the face detection frame, the problem of poor target tracking continuity when the face is blocked or the pose changes are large, and the target face tracking in these scenarios is achieved.
Patent Information
- Application Number
- PCT/CN2023/139399
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-26
AI Technical Summary
In the prior art, the multi-objective tracking algorithm DeepSort has poor continuity of target tracking and difficult to detect a face in scenarios such as the face being blocked and the face pose changes greatly.
By determining the trajectory corresponding to the upper body detection box, and determining the face detection box corresponding to the upper body detection box based on the positional relationship between the face detection box and the designated area in the upper body detection box, the ID of the face detection box is associated with the target ID bound to the trajectory matching the upper body detection box.
In scenes where the face is blocked and the face position changes greatly, the target face can also be tracked to improve the continuity of target tracking.
Smart Images

Figure CN2023139399_26062025_PF_FP_ABST
Abstract
Description
Multi-target face tracking method, device, computer equipment and medium based on DeepSort Technical Field
[0001] The present disclosure relates to the field of computer vision, and in particular to a DeepSort-based multi-target face tracking method, apparatus, computer equipment, and medium. Background Art
[0002] Object tracking technology is an important branch of the field of computer vision and has broad application prospects in video surveillance, visual navigation, human-computer interaction, etc.
[0003] In related technologies, the multi-target tracking algorithm DeepSort has poor target tracking continuity in scenarios where faces are difficult to detect, such as when the face is occluded or when the face posture changes significantly.
[0004] Summary of the Invention
[0005] In view of this, the embodiments of the present disclosure propose a multi-target face tracking method, apparatus, computer equipment and medium based on DeepSort to solve the technical problems in the related art.
[0006] According to a first aspect of an embodiment of the present disclosure, a multi-target face tracking method based on DeepSort is proposed, comprising:
[0007] Perform target detection on the current video frame to obtain a set of face detection frames and a set of upper body detection frames;
[0008] Determine the trajectory corresponding to each upper body detection frame and the predicted upper body detection frame matching the upper body detection frame using the DeepSort algorithm; the trajectory is bound to the target ID;
[0009] Based on the positional relationship between the face detection frame and the specified area in the specified upper body detection frame, the face detection frame corresponding to each upper body detection frame is determined; the specified upper body detection frame is an upper body detection frame or a predicted upper body detection frame; the ID of the face detection frame is the target ID bound to the corresponding trajectory.
[0010] Optionally, the step of determining the face detection frame corresponding to each upper body detection frame based on a positional relationship between the face detection frame and a specified area in a specified upper body detection frame includes:
[0011] When all or part of the face detection frame is within the range of the specified upper body detection frame, calculating the matching degree between the face detection frame and the specified area of the specified upper body detection frame, and determining the face detection frame with the greatest matching degree as the face detection frame corresponding to the upper body detection frame;
[0012] When the area of the face detection frame is outside the range of the designated upper body detection frame, a face detection frame is generated in the designated area and is determined as the face detection frame corresponding to the upper body detection frame.
[0013] Optionally, the designated area is determined based on the size of the face detection frame in the historical data set and the position of the face detection frame in the upper body detection frame.
[0014] Optionally, after determining the upper body detection box that matches the trajectory in the DeepSort algorithm, a customized confidence-based moving average formula is used to update the feature vector corresponding to the trajectory; the confidence is the confidence of the upper body detection box that matches the trajectory.
[0015] According to a second aspect of an embodiment of the present disclosure, a multi-target face tracking device based on DeepSort is proposed, comprising:
[0016] The target detection module is used to perform target detection on the current video frame and obtain a set of face detection frames and a set of upper body detection frames;
[0017] A trajectory determination module is used to determine the trajectory corresponding to each upper body detection frame and the predicted upper body detection frame matching the upper body detection frame using the DeepSort algorithm; the trajectory is bound to the target ID;
[0018] A face determination module is configured to determine the face detection frame corresponding to each upper body detection frame based on the positional relationship between the face detection frame and a specified area in a specified upper body detection frame; the specified upper body detection frame is an upper body detection frame or a predicted upper body detection frame; and the ID of the face detection frame is the target ID bound to the corresponding trajectory.
[0019] Optionally, determining the face detection frame corresponding to each upper body detection frame based on a positional relationship between the face detection frame and a specified area in a specified upper body detection frame includes:
[0020] When all or part of the face detection frame is within the range of the specified upper body detection frame, calculating the matching degree between the face detection frame and the specified area of the specified upper body detection frame, and determining the face detection frame with the greatest matching degree as the face detection frame corresponding to the upper body detection frame;
[0021] When the area of the face detection frame is outside the range of the designated upper body detection frame, a face detection frame is generated in the designated area and is determined as the face detection frame corresponding to the upper body detection frame.
[0022] Optionally, the designated area is determined based on the size of the face detection frame in the historical data set and the position of the face detection frame in the upper body detection frame.
[0023] Optionally, after determining the upper body detection box that matches the trajectory in the DeepSort algorithm, a customized confidence-based moving average formula is used to update the feature vector corresponding to the trajectory; the confidence is the confidence of the upper body detection box that matches the trajectory.
[0024] According to a third aspect of an embodiment of the present disclosure, a computer device is proposed, comprising a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor is prompted by the machine-executable instructions to execute the method described in the first aspect above.
[0025] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect is implemented.
[0026] It can be seen from the above technical solutions that the technical solution disclosed in the present invention determines the trajectory corresponding to the upper body detection frame, and based on the positional relationship between the face detection frame and the specified area in the upper body detection frame, determines the face detection frame corresponding to the upper body detection frame, and associates the ID of the face detection frame with the target ID bound to the trajectory matched by the upper body detection frame. In this way, the target face can be tracked even in scenarios where it is difficult to detect the face, such as when the face is occluded or the face posture changes greatly, thereby improving the continuity of target tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0028] FIG1 is a flowchart illustrating a multi-target face tracking method based on DeepSort according to an embodiment of the present disclosure;
[0029] FIG2 is a block diagram of a multi-target face tracking device based on DeepSort according to an embodiment of the present disclosure;
[0030] FIG3 is a schematic diagram showing the hardware structure of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present disclosure.
[0032] The terms used in the embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the embodiments of the present disclosure. The singular forms "a," "an," and "the" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0033] It should be understood that although the terms first, second, third, etc. may be used to describe various information in the embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the embodiments of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0034] For the purpose of brevity and ease of understanding, the terms "greater than," "less than," "higher than," and "lower than" are used herein to describe size relationships. However, those skilled in the art will understand that the term "greater than" also encompasses the meaning of "greater than or equal to," and "less than" also encompasses the meaning of "less than or equal to," and the term "higher than" also encompasses the meaning of "higher than or equal to," and "lower than" also encompasses the meaning of "lower than or equal to."
[0035] In related technologies, the multi-target tracking algorithm DeepSort identifies the trajectory corresponding to the face detection frame of the target person in the video frame and binds the face detection frame ID to the trajectory ID to track the target face. However, in scenarios where face detection is difficult, such as when the face is occluded or the face's posture changes significantly, it is difficult to detect the face detection frame in the video frame. Even when the target person is still in the video frame, it is impossible to match their face with the trajectory, resulting in poor target tracking continuity.
[0036] In response to the above technical problems, the present disclosure provides a multi-target face tracking method based on DeepSort, which determines the trajectory corresponding to the upper body detection frame and the predicted upper body detection frame matching the upper body detection frame, and determines the face detection frame corresponding to the upper body detection frame based on the positional relationship between the face detection frame and the specified area in the specified upper body detection frame, and associates the ID of the face detection frame with the target ID bound to the trajectory matching the upper body detection frame. In this way, the target face can be tracked even in scenarios where the face is blocked or the face posture changes greatly, thereby improving the continuity of target tracking.
[0037] The invention will be described in detail through one or more of the following embodiments.
[0038] FIG1 is a flowchart illustrating a multi-target face tracking method based on DeepSort according to an embodiment of the present disclosure.
[0039] As shown in FIG1 , the multi-target face tracking method based on DeepSort may include the following steps:
[0040] Step S101: perform target detection on the current video frame to obtain a face detection frame set and an upper body detection frame set.
[0041] Step S102: Determine the trajectory corresponding to each upper body detection frame and the predicted upper body detection frame matching the upper body detection frame using the DeepSort algorithm; the trajectory is bound to the target ID.
[0042] Step S103: Determine the face detection frame corresponding to each upper body detection frame based on the positional relationship between the face detection frame and the specified area in the specified upper body detection frame; the specified upper body detection frame is an upper body detection frame or a predicted upper body detection frame; and the ID of the face detection frame is the target ID bound to the corresponding trajectory.
[0043] In this embodiment, the face area of the target person can be determined through the face detection frame. The face area can refer to the area enclosed by the left ear edge, the right ear edge, the chin and the forehead, and may not include hair to reduce noise when identifying facial key points in the image area of the face detection frame; the upper body area of the target person can be determined through the upper body detection frame. The upper body area can refer to the area formed above the chest, below the top of the head and between the shoulders of the human body; of course, the face area and the upper body area can also refer to other areas, and the present disclosure does not limit this.
[0044] In this embodiment, the designated upper body detection frame can be the upper body detection frame, and its corresponding face detection frame can be determined through the upper body detection frame. Since the face detection frame is the detection frame actually detected in the current video frame, the position of the face detection frame determined by it is more accurate.
[0045] In this embodiment, the designated upper body detection frame can be a predicted upper body detection frame that matches the upper body detection frame, and the face detection frame corresponding to the upper body detection frame can be determined by predicting the face detection frame. The trajectories formed by the face detection frames of the same target in several determined video frames are relatively smooth, and the display effect is better.
[0046] Since the face detection frame is a detection frame actually detected in the current video frame, the position of the face detection frame determined thereby is more accurate.
[0047] In one embodiment, step S101, performing target detection on the current video frame to obtain a face detection frame set and an upper body detection frame set may specifically include:
[0048] Step S1011: pre-process the current video frame:
[0049] In the present disclosure, preprocessing may include resolution adjustment, format conversion, and normalization of video frames, and the present disclosure does not impose any restrictions on this.
[0050] Among them, resolution adjustment can be implemented based on image scaling algorithms such as bilinear interpolation and bicubic interpolation. The adjusted video frame resolution can be set according to actual conditions, for example, according to the input image size requirements of the target detection model used, or according to actual application scenarios. For example, when the target detection model is deployed on a processor with weaker performance, the input image resolution can be reduced to reduce computational complexity and improve processing efficiency. Appropriately reducing the image resolution can take into account both image clarity and processing efficiency. The converted image format can be set according to the input image format requirements of the target detection model used, for example, it can be converted to RGB format, and then further normalized, such as scaling the value of each pixel in the image to a specified range (e.g., [0,1]) to reduce computational complexity and improve processing efficiency.
[0051] Step S1012: Detecting a plurality of candidate face detection frames and a plurality of candidate upper body detection frames of the target person in the preprocessed video frame based on the target detection model;
[0052] In this embodiment, the target detection model can be a YOLO model, an R-CNN model, or an SSD model, etc., which is not limited in this disclosure.
[0053] In this embodiment, a target detection model can be used to detect video frames containing several target persons. For each target person, several candidate face detection frames can be detected to determine the target person's facial region, as well as several candidate upper body detection frames to determine the target person's upper body region. The detection frame information can include the center coordinates, length, and width of the detection frame, which represent the size of the detection frame and its position in the target image. The detection frame information can also include a confidence level, which represents the probability that the target is present in the detection frame.
[0054] Step S1023 : Determine a face detection frame and an upper body detection frame for each target person from the plurality of candidate face detection frames and the plurality of candidate upper body detection frames, and obtain a face detection frame set and an upper body detection frame set.
[0055] In this embodiment, based on the NMS (Non-Maximum Suppression) algorithm, for each target person, a face detection frame and an upper body detection frame can be determined from several candidate face detection frames and several candidate upper body detection frames corresponding to the target person. In this way, a face detection frame set and an upper body detection frame set can be obtained for several target tasks.
[0056] In one embodiment, step S102, determining the trajectory corresponding to each upper body detection frame using the DeepSort algorithm, and the predicted upper body detection frame matching the upper body detection frame may specifically include:
[0057] Step S1021: If the current video frame is the first video frame, a corresponding track is created for each upper body detection frame in the current video frame.
[0058] Step S1022: If the current video frame is not the first video frame, the position of the upper body detection frame corresponding to the trajectory of the previous video frame is predicted in the current video frame based on the Kalman filter algorithm; the position can be represented by the predicted upper body detection frame.
[0059] Step S1023 , matching the upper body detection frame of the current video frame with the predicted upper body detection frame based on the Hungarian algorithm, determining the predicted upper body detection frame that matches the upper body detection frame, and the trajectory corresponding to the upper body detection frame is the trajectory corresponding to the predicted upper body detection frame.
[0060] In this embodiment, trajectories can be divided into confirmed trajectories and unconfirmed trajectories. When a trajectory is created, it is in the unconfirmed state by default. For an unconfirmed trajectory, it can be converted to a confirmed trajectory when the corresponding predicted upper body detection frame successfully matches the upper body detection frame of the video frame three times.
[0061] During the matching process, if there are no confirmed tracks, the predicted upper body detection frame corresponding to the unconfirmed track can be matched with the upper body detection frame of the video frame through Intersection over Union (IoU). After IoU matching, there may be unsuccessful unconfirmed tracks and unsuccessful upper body detection frames. In this case, the unsuccessful unconfirmed tracks can be deleted, and new tracks can be created for the unsuccessful upper body detection frames.
[0062] When a confirmed track exists, the predicted upper body detection frame corresponding to the confirmed track can be cascade matched with the upper body detection frame of the video frame. After cascade matching, there may be unsuccessful confirmed tracks and unsuccessful upper body detection frames. In this case, the unsuccessful confirmed tracks and the predicted upper body detection frames corresponding to the non-confirmed tracks can be matched with the unsuccessful upper body detection frames. At this time, after IoU matching, there may still be unsuccessful confirmed tracks. If the number of consecutive unsuccessful matches for a confirmed track exceeds a specified threshold, the confirmed track can be deleted.
[0063] In one embodiment, when the number of consecutive unsuccessful matches of the confirmation trajectory does not exceed a specified threshold (for example, 3 times), the designated upper body detection frame may be a predicted upper body detection frame corresponding to the unsuccessful matching confirmation trajectory, that is, a predicted upper body detection frame that is not matched with an upper body detection frame in the current video frame. In the case where a target object exists in some video frames but its upper body detection frame is not detected normally, the face detection frame can be determined through its corresponding predicted upper body detection frame. At the same time, setting a specified threshold can avoid the situation where the target object does not exist in too many video frames but the face detection frame is still determined through its corresponding predicted upper body detection frame, thereby improving the continuity of target tracking while ensuring the accuracy of target tracking.
[0064] During the cascade matching process, it is necessary to obtain the feature vector of each upper body detection frame in the video frame, calculate the similarity between each feature vector and the feature vector corresponding to the trajectory, determine the upper body detection frame corresponding to the feature vector with the largest similarity and exceeding the specified threshold and successfully match the trajectory, and update the feature vector corresponding to the trajectory.
[0065] In one embodiment, a custom confidence-based moving average formula can be used to update the feature vector corresponding to the trajectory. The formula can be: n+1 =(1–c n+1 )*f n +c n+1 *d n+1
[0066] Among them, f n+1 represents the feature vector corresponding to the trajectory in the n+1th video frame, c n+1 represents the confidence of the upper body detection box matching the trajectory in the n+1th video frame, f n represents the feature vector corresponding to the trajectory in the nth video frame, d n+1 Represents the feature vector of the upper body detection box that matches the trajectory in the n+1th video frame, where n is a natural number.
[0067] In this embodiment, during target detection on video frames, the confidence of the upper body detection frame can be determined based on the target detection model; the appearance features of the upper body detection frame can be extracted through the feature extraction network to determine the feature vector of the upper body detection frame.
[0068] The DeepSort algorithm in related art directly uses the feature vector of the currently matched face detection frame as the feature vector corresponding to the trajectory. This embodiment updates the feature vector of the trajectory based on the feature vector of the upper body detection frame. Compared to the face detection frame, the upper body detection frame contains richer information, which can improve matching accuracy. Furthermore, after determining the upper body detection frame that matches the trajectory, this embodiment uses a customized confidence-based moving average formula to update the feature vector corresponding to the trajectory. This allows the feature vector corresponding to the trajectory to incorporate the feature vectors of all previously successfully matched upper body detection frames. Consequently, when matching the trajectory to the upper body detection frame based on the feature vector, compared to the trajectory feature vector in related art that only includes the feature vector of the most recently matched face detection frame, the trajectory feature vector updated using the confidence-based moving average formula disclosed in this disclosure can significantly reduce the number of ID switches in scenarios with complex background interference and obscured figures, further improving matching accuracy.
[0069] In one embodiment, step S103, based on the positional relationship between the face detection frame and the designated area in the designated upper body detection frame, determining the face detection frame corresponding to each upper body detection frame may specifically include:
[0070] Step S1031, when all or part of the face detection frame is within the range of the specified upper body detection frame, calculate the matching degree between the face detection frame and the specified area in the specified upper body detection frame, and determine that the face detection frame with the greatest matching degree is the face detection frame corresponding to the upper body detection frame.
[0071] In this embodiment, IoU matching may be performed between the face detection frame and the designated area in the designated upper body detection frame, and the overlap ratio between the face detection frame and the designated area may be used as the matching degree.
[0072] Step S1032: When the area of the face detection frame is outside the range of the designated upper body detection frame, a face detection frame is generated in the designated area and determined as the face detection frame corresponding to the upper body detection frame.
[0073] In this embodiment, the size and position of a designated area within a designated upper body detection frame can be determined based on the size of the face detection frames and the positions of the face detection frames within the upper body detection frame in the historical dataset. For example, the historical dataset may include several images containing face detection frames and upper body detection frames containing people. The sizes of the face detection frames and the distribution positions of the face detection frames within the upper body detection frame are statistically analyzed, and the size of the most frequently appearing face detection frame, or the average of the sizes of the face detection frames, is used as the size of the designated area. The distribution positions of the most frequently appearing face detection frames within the upper body detection frame, or the average of the distribution positions, are used as the position of the designated area.
[0074] In this embodiment, it is possible to first determine whether there is a complete face detection frame in the specified upper body detection frame area. If so, the complete face detection frame with the greatest matching degree can be determined as the face detection frame corresponding to the upper body detection frame; if not, it is possible to further determine whether there is a partial face detection frame in the specified upper body detection frame area. If so, the partial face detection frame with the greatest matching degree can be determined as the face detection frame corresponding to the upper body detection frame; if not, a face detection frame can be further generated in the specified area of the specified upper body detection frame and determined as the face detection frame corresponding to the upper body detection frame.
[0075] This embodiment determines that the face detection frame with the greatest matching degree is the face detection frame corresponding to the upper body detection frame when all or part of the area of the face detection frame is within the range of the specified upper body detection frame; when the area of the face detection frame is outside the range of the specified upper body detection frame, a face detection frame is generated in the specified area and determined as the face detection frame corresponding to the upper body detection frame. Therefore, when there is a face detection frame within the range of the specified upper body detection frame, the face detection frame matching the upper body detection frame can be accurately determined; when there is no face detection frame within the range of the specified upper body detection frame, the corresponding face detection frame can also be generated. Therefore, in scenarios where it is difficult to detect the face, such as when the face is occluded or the face posture is large, the face detection frame can be determined based on the upper body detection frame, and the ID of the face detection frame is associated with the target ID bound to the trajectory matched by the upper body detection frame, thereby improving the continuity of target tracking.
[0076] FIG2 is a block diagram of a multi-target face tracking apparatus based on DeepSort according to an embodiment of the present disclosure, comprising:
[0077] The target detection module 11 is used to perform target detection on the current video frame to obtain a face detection frame set and an upper body detection frame set;
[0078] A trajectory determination module 12 is configured to determine a trajectory corresponding to each upper body detection frame and a predicted upper body detection frame that matches the upper body detection frame using a DeepSort algorithm; the trajectory is bound to a target ID;
[0079] The face determination module 13 is used to determine the face detection frame corresponding to each upper body detection frame based on the positional relationship between the face detection frame and the specified area in the upper body detection frame; the specified upper body detection frame is an upper body detection frame or a predicted upper body detection frame; the ID of the face detection frame is the target ID bound to the corresponding trajectory.
[0080] In one embodiment, the face determination module 13 may determine the face detection frame corresponding to each upper body detection frame based on the positional relationship between the face detection frame and the specified area in the specified upper body detection frame, including:
[0081] When all or part of the face detection frame is within the range of the specified upper body detection frame, calculating the matching degree between the face detection frame and the specified area of the specified upper body detection frame, and determining the face detection frame with the greatest matching degree as the face detection frame corresponding to the upper body detection frame;
[0082] When the area of the face detection frame is outside the range of the designated upper body detection frame, a face detection frame is generated in the designated area and is determined as the face detection frame corresponding to the upper body detection frame.
[0083] In one embodiment, the designated area may be determined based on the size of the face detection frame in the historical dataset and the position of the face detection frame in the upper body detection frame.
[0084] In one embodiment, after determining the upper body detection box that matches the trajectory in the DeepSort algorithm, a customized confidence-based moving average formula can be used to update the feature vector corresponding to the trajectory; the confidence can be the confidence of the upper body detection box that matches the trajectory.
[0085] Figure 3 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. The computer device may include a processor 301 and a machine-readable storage medium 302 storing machine-executable instructions. The processor 301 and the machine-readable storage medium 302 may communicate via a system bus 303. Furthermore, by reading and executing the machine-executable instructions corresponding to the DeepSort-based multi-target face tracking logic in the machine-readable storage medium 302, the processor 301 may execute the DeepSort-based multi-target face tracking method described above.
[0086] The machine-readable storage medium 302 referred to herein can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium 302 can include at least one of the following types of storage media: volatile memory, non-volatile memory, or other types of storage media. Volatile memory can be RAM (Random Access Memory), and non-volatile memory can be flash memory, a storage drive (such as a hard disk drive), a solid-state drive, or a storage disk (such as a CD, DVD, etc.).
[0087] Based on the method described in any of the above embodiments, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can be used to execute the beautification method described in any of the above embodiments.
[0088] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.
[0089] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0090] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0091] The phrases "specific examples" or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the present disclosure, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0092] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not claimed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0093] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
[0094] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A multi-object face tracking method based on DeepSort, characterized in that Including: Performing object detection on the current video frame to obtain a set of face detection boxes and a set of upper body detection boxes; Determining the trajectory corresponding to each upper body detection box and the predicted upper body detection box matching the upper body detection box through the DeepSort algorithm; the trajectory is bound with a target ID; Determining the face detection box corresponding to each upper body detection box based on the positional relationship between the face detection box and the specified area in the specified upper body detection box; the specified upper body detection box is an upper body detection box or a predicted upper body detection box; the ID of the face detection box is the target ID bound to the corresponding trajectory.
2. The method according to claim 1, characterized in that, The step of determining the face detection box corresponding to each upper body detection box based on the positional relationship between the face detection box and the specified area in the specified upper body detection box includes: When all or part of the area of the face detection box is within the range of the specified upper body detection box, calculating the matching degree between the face detection box and the specified area in the specified upper body detection box, and determining the face detection box with the maximum matching degree as the face detection box corresponding to the upper body detection box; When the area of the face detection box is outside the range of the specified upper body detection box, generating a face detection box in the specified area and determining it as the face detection box corresponding to the upper body detection box.
3. The method according to claim 1, wherein The specified area is determined based on the size of the face detection box in the historical dataset and the position of the face detection box in the upper body detection box.
4. The method according to claim 1, wherein After determining the upper body detection box matching the trajectory in the DeepSort algorithm, using a custom confidence-based moving average formula to update the feature vector corresponding to the trajectory; the confidence is the confidence of the upper body detection box matching the trajectory.
5. A multi-object face tracking device based on DeepSort, characterized in that, Including: An object detection module for performing object detection on the current video frame to obtain a set of face detection boxes and a set of upper body detection boxes; A trajectory determination module for determining the trajectory corresponding to each upper body detection box and the predicted upper body detection box matching the upper body detection box through the DeepSort algorithm; the trajectory is bound with a target ID; A face determination module for determining the face detection box corresponding to each upper body detection box based on the positional relationship between the face detection box and the specified area in the specified upper body detection box; the specified upper body detection box is an upper body detection box or a predicted upper body detection box; the ID of the face detection box is the target ID bound to the corresponding trajectory.
6. The device according to claim 5, wherein The determining the face detection box corresponding to each upper body detection box based on the positional relationship between the face detection box and the specified area in the specified upper body detection box includes: When all or part of the area of the face detection box is within the range of the specified upper body detection box, calculating the matching degree between the face detection box and the specified area in the specified upper body detection box, and determining the face detection box with the maximum matching degree as the face detection box corresponding to the upper body detection box; When the area of the face detection box is outside the range of the specified upper body detection box, generating a face detection box in the specified area and determining it as the face detection box corresponding to the upper body detection box.
7. The device according to claim 5, characterized in that, The specified area is determined based on the size of the face detection box in the historical dataset and the position of the face detection box in the upper body detection box.
8. The device according to claim 5, characterized in that, After determining the upper body detection box that matches the trajectory in the DeepSort algorithm, a custom confidence-based moving average formula is used to update the feature vector corresponding to the trajectory; the confidence is the confidence of the upper body detection box that matches the trajectory.
9. A computer device, characterized in that, The computer device includes a processor and a machine-readable storage medium, and the machine-readable storage medium stores machine-executable instructions that can be executed by the processor. The processor is prompted by the machine-executable instructions to execute the method according to any one of claims 1-4.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the medium, and when the computer program is executed by a processor, the method according to any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Object tracking method, object tracking device and computer-readable storage medium
CN108875488A
Pedestrian tracking method and device, and terminal
CN110427905A
Object tracking method and apparatus and storage medium
CN111797652A
Target tracking method and device and computer storage medium
CN112037247A
Multi-person identity and action association recognition method and device and readable medium
CN114419480A
Cited By
Target tracking method, electronic equipment, readable medium and program product
CN120279056A