Light Tracking: System and Method for Online Top-Down Human Pose Tracking

A lightweight online framework for human pose tracking using CNNs and twin graph convolutional networks addresses the lack of real-time multi-person tracking by efficiently updating object boundaries and poses, achieving high accuracy and frame rates.

CN114787865BActive Publication Date: 2025-07-15BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080062563.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-20
Filing Date
2020-09-18
Publication Date
2025-07-15
Estimated Expiration
2040-09-18

AI Technical Summary

Technical Problem

Most existing human posture tracking methods are offline, lack real-time, making it difficult to accurately estimate multi-person human postures in videos and assign unique instance IDs to key points across frames.

Method used

The top-down multi-person online pose tracking system is adopted, and key points are estimated using convolutional neural networks and twin graph convolutional networks (SGCNs), real-time tracking is achieved through object state judgment and bounding box inference, and identity identification is recognized in combination with the pose matching module.

Benefits of technology

It realizes real-time and accurate tracking of multiple human postures in video, improves the real-time and accuracy of posture estimation, reduces the computing burden of data association, and is suitable for online environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114787865B_ABST
    Figure CN114787865B_ABST
Patent Text Reader

Abstract

A system and method for pose tracking, particularly for top-down online multi-person pose tracking. The system includes: a computing device including a processor and a storage device storing computer-executable code, wherein the computer-executable code, when executed at the processor, is configured to: provide a plurality of consecutive frames of a video, the plurality of consecutive frames including at least one key frame and a plurality of non-key frames; for each non-key frame of the plurality of non-key frames: receive a previous inferred bounding box of an object inferred from a previous frame; estimate key points from the non-key frame in a region defined by the previous inferred bounding box to obtain estimated key points; determine an object state based on the estimated key points, wherein the object state includes a "tracked" state and a "lost" state; and when the object state is "tracked", infer an inferred bounding box based on the estimated key points to process the next frame of the non-key frame.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] In the description of the present disclosure, some references are cited and discussed, which may include patents, patent applications, and various publications. The citation and / or discussion of such references are provided only to clarify the description of the present disclosure and do not admit that any such reference is "prior art" of the disclosure described herein. All references cited and discussed in this specification are incorporated herein by reference in their entirety and to the same extent as each reference is incorporated by reference individually. Technical Field

[0003] The present disclosure relates to human pose tracking, and more particularly, to a general lightweight framework for online top-down human pose tracking. Background Art

[0004] The background description provided herein is to generally present the context of the present disclosure. To the extent within the scope of this background section, the work of the inventors, as well as aspects of the description that may not conform to the prior art at the time of application, are neither expressly nor implicitly admitted as prior art against the present disclosure.

[0005] Pose tracking is a task for estimating the human poses of multiple persons in a video and assigning a unique instance ID to each key point across frames. The accurate estimation of human key point trajectories is useful for human action recognition, human interaction understanding, motion capture, and animation, etc. Recently, the publicly available Pose-Track dataset and the MPII Video Pose dataset have pushed the research of human motion analysis into real-world scenarios, and two pose tracking challenges have been held. However, most of the existing methods are offline and thus lack the potential for real-time.

[0006] Therefore, there is a need in the art to address the above-mentioned deficiencies and drawbacks. Summary of the Invention

[0007] In some aspects, the present disclosure relates to a system for pose tracking, particularly a system for online top-down multi-person pose tracking. The system includes a computing device, which includes a processor and a storage device storing computer-executable code. The computer-executable code, when executed at the processor, is configured to:

[0008] Provide a plurality of consecutive frames of a video, the consecutive frames including at least one key frame and a plurality of non-key frames;

[0009] For each of the plurality of non-key frames: receive a previous inferred bounding box of an object inferred from a previous frame; estimate key points from the non-key frame in a region defined by the previous inferred bounding box to obtain estimated key points; determine an object state based on the estimated key points, where the object state includes a "tracked" state and a "lost" state; and when the object state is "tracked", infer an inferred bounding box based on the estimated key points to process the next frame of the non-key frame.

[0010] In some embodiments, the computer-executable code is configured to estimate key points from the non-key frame using a convolutional neural network.

[0011] In some embodiments, when the estimated key points have an average confidence greater than a threshold score, the object state is "tracked", and when the estimated key points have an average confidence less than or equal to the threshold score, the object state is "lost".

[0012] In some embodiments, the computer-executable code is configured to infer the inferred bounding box by: defining a bounding box enclosing the estimated key points; and magnifying the bounding box by 20% in the horizontal and vertical directions of the bounding box, respectively.

[0013] In some embodiments, the computer-executable code is further configured to, when the object state is "lost": detect an object from the non-key frame, where each detected object is defined by a detected bounding box; estimate key points from each detected bounding box to obtain detected key points; identify each detected object by comparing the detected key points of the detected object with stored key points of a stored object, each of the stored objects having an object identification ID; and assign the object ID of the corresponding object in the stored object to the detected object when the detected key points match the stored key points of the corresponding object from the stored objects.

[0014] In some embodiments, the computer-executable code is configured to detect an object using a convolutional neural network. In some embodiments, the computer-executable code is configured to use a convolutional neural network to estimate the key points.

[0015] In some embodiments, the step of comparing the detected key points of the detected object with the stored key points is performed using a Siamese Graph Convolutional Network (SGCN), where the SGCN includes two Graph Convolutional Networks (GCNs) with shared network weights, and each GCN includes: a first GCN layer; a first ReLU unit connected to the first GCN layer; a second GCN layer connected to the first ReLU unit; a second ReLU unit connected to the second GCN layer; an average pooling layer connected to the second GCN layer; a fully connected network (FCN); and a feature vector transformation layer. The first GCN layer is configured to receive the detected key points of one of the detected objects, and the feature vector transformation layer is configured to generate a feature vector representing the pose of one of the detected objects.

[0016] In some embodiments, the SGCN is configured to perform the comparing step by: running the estimated key points through one of the two GCNs to obtain an estimated feature vector for the estimated key points; running the stored key points of one of the stored objects through the other of the two GCNs to obtain a feature vector for the stored key points; and determining that the estimated key points match the stored key points when the distance between the estimated feature vector and the stored feature vector is less than a predetermined threshold.

[0017] In some embodiments, for each non-key frame among the non-key frames, the computer-executable code is configured to, when the estimated key points of the object do not match the key points of any of the stored objects: assign a new object ID to the object.

[0018] In some embodiments, for each key frame among the key frames, the computer-executable code is configured to:

[0019] detect the objects in the key frame, where each detected object is defined by a bounding box;

[0020] estimate multiple detected key points of each detected object from the bounding box corresponding to the detected object;

[0021] identify each detected object by comparing the detected key points of the detected object with the stored key points of the stored objects, where each of the stored objects has an object identification ID; and

[0022] when the detected key points match the stored key points of the corresponding object from the stored objects, assign the object ID of the corresponding object in the stored objects to the detected object.

[0023] In some embodiments, for each of the key frames, the computer-executable code is configured to, when the estimated key points of the object do not match the key points of any of the stored objects: assign a new object ID to the object.

[0024] In some aspects, the present disclosure relates to a method for pose tracking, particularly for top-down multi-person online pose tracking. In some embodiments, the method includes:

[0025] Providing a plurality of consecutive frames of a video, the consecutive frames including at least one key frame and a plurality of non-key frames;

[0026] For each of the plurality of non-key frames:

[0027] Receiving a previously inferred bounding box of an object inferred from a previous frame;

[0028] In the region defined by the previously inferred bounding box, estimating key points from the non-key frame to obtain estimated key points;

[0029] Based on the estimated key points, determining an object state, wherein the object state includes a "tracked" state and a "lost" state; and

[0030] When the object state is "tracked", inferring an inferred bounding box based on the estimated key points to process the next frame of the non-key frame.

[0031] In some embodiments, when the estimated key points have an average confidence greater than a threshold score, the object state is "tracked", and when the estimated key points have a confidence less than or equal to the threshold score, the object state is "lost".

[0032] In some embodiments, the step of inferring the inferred bounding box includes: defining a bounding box surrounding the estimated key points; and magnifying the bounding box by 20% in the horizontal and vertical directions of the bounding box respectively.

[0033] In some embodiments, the method further includes, when the object state is "lost":

[0034] Detecting an object from the non-key frame, wherein each detected object is defined by a detected bounding box;

[0035] Estimating key points of each detected object from the corresponding bounding box of the detected bounding box to obtain detected key points;

[0036] Each of the detected objects is identified by comparing the detected key points of the detected object with the stored key points of each stored object, each of the stored objects having an object identification ID; and

[0037] When the detected key points match the stored key points from one of the stored objects, the object ID of one of the stored objects is assigned to the detected object.

[0038] In some embodiments, the steps of detecting the object and estimating the key points are both performed using a convolutional neural network CNN.

[0039] In some embodiments, the step of comparing the detected key points of the detected object with the stored key points is performed using a Siamese graph convolutional network SGCN, the SGCN including two graph convolutional networks GCNs with shared network weights, each GCN including: a first graph convolutional network GCN layer; a first Relu unit connected to the first GCN layer; a second GCN layer connected to the first Relu unit; a second Relu unit connected to the second GCN layer; an average pooling layer connected to the second GCN layer; a fully connected network FCN; and a feature vector transformation layer, wherein the first GCN layer is configured to receive the detected key points of one of the detected objects, and the feature vector transformation layer is configured to generate a feature vector representing the pose of one of the detected objects.

[0040] In some embodiments, the method further includes, for each key frame of the key frames:

[0041] Detecting objects from the key frame, each detected object being defined by a detected bounding box;

[0042] Estimating detected key points from each of the detected bounding boxes to obtain detected key points;

[0043] Each of the detected objects is identified by comparing the detected key points of the detected object with the stored key points of the stored objects, each of the stored objects having an object identification ID; and

[0044] When the detected key points match the stored key points of the corresponding object from the stored objects, the object ID of the corresponding object in the stored objects is assigned to the detected object.

[0045] In some aspects, the present disclosure relates to a non - transitory computer - readable medium storing computer - executable code. The computer - executable code, when executed at a processor of a computing device, is configured to perform the above - described method.

[0046] These and other aspects of the present disclosure will become apparent from the following description of the preferred embodiments in conjunction with the accompanying drawings and their captions, although variations and modifications therein may be effected without departing from the spirit and scope of the novel concepts of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings illustrate one or more embodiments of the present disclosure and, together with the written description, are used to explain the principles of the present disclosure. Wherever possible, the same reference numerals are used throughout the drawings to refer to the same or like elements of the embodiments.

[0048] Figure 1 Schematically depicts a system for online top - down human pose tracking according to certain embodiments of the present disclosure.

[0049] Figure 2 Schematically depicts a pose tracking module according to certain embodiments of the present disclosure.

[0050] Figure 3 Schematically depicts a sudden camera movement in a video and a sudden camera zoom - in in a video according to certain embodiments of the present disclosure.

[0051] Figure 4 Schematically depicts a Siamese Graph Convolutional Network (SGCN) according to certain embodiments of the present disclosure.

[0052] Figure 5 Schematically depicts a spatial configuration strategy for graph sampling and weighting to construct graph convolution operations according to certain embodiments of the present disclosure.

[0053] Figure 6 Schematically depicts a process of online human pose tracking according to certain embodiments of the present disclosure.

[0054] Figure 7A Schematically depicts a process of online human pose tracking on key frames or specific non - key frames according to certain embodiments of the present disclosure.

[0055] Figure 7B Schematically depicts a process of online human pose tracking on non - key frames according to certain embodiments of the present disclosure.

[0056] Figure 8 Pose pairs collected from the PoseTrack′18 dataset are shown in Table 1.

[0057] Figure 9 Table 2 shows a comparison of the detectors in Table 2 according to certain embodiments of the present disclosure.

[0058] Figure 10 Table 3 shows a comparison of the offline pose tracking results using various detectors on the PoseTrack′17 validation set according to certain embodiments of the present disclosure.

[0059] Figure 11 Table 4 shows a comparison of the offline and online pose tracking results with various key frame intervals on the PoseTrack′18 validation set according to certain embodiments of the present disclosure.

[0060] Figure 12 Table 5 shows a performance comparison of the tracking using lightweight GCN and SC on the PoseTrack′18 validation set according to certain embodiments of the present disclosure.

[0061] Figure 13 Table 6 shows a performance comparison on the PoseTrack dataset according to certain embodiments of the present disclosure. Detailed Description

[0062] The present disclosure is described more specifically in the following examples, which are intended to be illustrative only, since many modifications and variations will be apparent to those skilled in the art. Various embodiments of the present disclosure are now described in detail. Referring to the accompanying drawings, like numerals indicate like components throughout the views. Unless the context clearly dictates otherwise, the meanings of "a", "an", and "the" as used in the description herein and throughout the claims include pluralities. Additionally, as used in the description and claims of the present disclosure, unless the context clearly dictates otherwise, the meaning of "in" includes "in" and "on". Also, the specification may use headings or subheadings for the convenience of the reader, which does not affect the scope of the present disclosure. Additionally, some terms used in this specification are defined more specifically below.

[0063] The terms used in this specification generally have their ordinary meanings in the art, in the context of the present disclosure, and in the particular context in which each term is used. Certain terms used to describe the present disclosure are discussed below or elsewhere in the specification to provide additional guidance to the practitioner regarding the description of the present disclosure. It is understood that the same thing can be expressed in more than one way. Accordingly, alternative language and synonyms may be used for any one or more of the terms discussed herein, and no special significance is to be attached to whether or not a term is elaborated or discussed herein. The present disclosure provides synonyms for certain terms. The recitation of one or more synonyms does not exclude the use of other synonyms. The use of examples anywhere in this specification, including examples of any of the terms discussed herein, is illustrative only and in no way limits the scope and meaning of the present disclosure or of any exemplary term. Likewise, the present disclosure is not limited to the various embodiments given in this specification.

[0064] It should be understood that when an element is referred to as being “on” another element, it can be directly on the other element or intervening elements may be present therebetween. In contrast, when an element is referred to as being “directly on” another element, there are no intervening elements. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.

[0065] It should be understood that although the terms first, second, third, etc. may be used herein to describe various elements, components, regions, layers, and / or sections, these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer, or section from another element, component, region, layer, or section. Thus, a first element, component, region, layer, or section discussed below could be termed a second element, component, region, layer, or section without departing from the teachings of the present disclosure.

[0066] In addition, relative terms such as “below” or “bottom” and “above” or “top” may be used herein to describe the relationship of one element to another element, as illustrated. It should be understood that relative terms are intended to encompass different orientations of the device in addition to the orientations depicted in the figures. For example, if the device in one of the figures is flipped, an element described as on the “below” side of another element will be oriented on the “above” side of the other element. Thus, the exemplary term “below” can include both the directions of “below” and “above,” depending on the particular orientation of the figure. Similarly, if the device in one of the figures is flipped, an element described as “under” or “beneath” another element will be oriented “over” the other element. Thus, the exemplary terms “under” or “beneath” can include the directions of both up and down.

[0067] Unless otherwise defined, all terms (including technical and scientific terms) used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It should also be understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the relevant art and the context of this disclosure, and should not be interpreted in an idealized or overly formal sense.

[0068] As used herein, "substantially", "about", "essentially" or "approximately" generally means within 20% of a given value or range, preferably within 10%, and more preferably within 5%. The numerical values given herein are approximate values, meaning that the terms "substantially", "about", "essentially" or "approximately" can be inferred if not explicitly stated.

[0069] As used herein, "plurality" means two or more.

[0070] As used herein, the terms "comprising", "including", "carrying", "having", "containing", "involving", etc. should be understood to be open-ended, i.e., meaning including but not limited to.

[0071] As used herein, the phrase "at least one of A, B, and C" should be interpreted to mean logic (A or B or C), using non-exclusive logical OR. It should be understood that, without changing the principles of this disclosure, one or more steps within a method can be performed in a different order (or simultaneously).

[0072] As used herein, the term "module" may refer to belonging to or including an application specific integrated circuit (ASIC); electronic circuitry; combinatorial logic circuitry; a field programmable gate array (FPGA); a processor (shared, dedicated, or group) that executes code; other suitable hardware components that provide the described functionality; or a combination of some or all of the above, such as in a system-on-chip. A module may include a memory (shared, dedicated, or group) that stores code executed by a processor.

[0073] As used herein, the term "code" may include software, firmware, and / or microcode, and may refer to programs, routines, functions, classes, and / or objects. The term shared as used above means that a single (shared) processor can be used to execute some or all of the code from multiple modules. Additionally, some or all of the code from multiple modules can be stored in a single (shared) memory. The term group as used above means that a group of processors can be used to execute some or all of the code from a single module. Additionally, a group of memories can be used to store some or all of the code from a single module.

[0074] As described herein, the term "interface" generally refers to a communication tool or device used to perform data communication between components at the interaction point between components. Generally speaking, interfaces can be applicable at both the hardware and software levels and can be one-way or two-way interfaces. Examples of physical hardware interfaces can include electrical connectors, buses, ports, cables, terminals, and other I / O devices or components. The components communicating with the interface can be, for example, multiple components of a computer system or peripheral devices.

[0075] The present disclosure relates to computer systems. As shown in the accompanying drawings, computer components can include physical hardware components, which are shown as solid blocks, and virtual software components, which are shown as dashed blocks. Those of ordinary skill in the art will understand that unless otherwise stated, these computer components can be implemented in the form of software, firmware, or hardware components or combinations thereof, but are not limited to these forms.

[0076] The apparatuses, systems, and methods described herein can be implemented by one or more computer programs executed by one or more processors. The computer programs include processor-executable instructions stored on a non-transitory tangible computer-readable medium. The computer programs may also include stored data. Non-limiting examples of non-transitory tangible computer-readable media are non-volatile memories, magnetic storage, and optical storage.

[0077] The present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, in which embodiments of the present disclosure are shown. However, the present disclosure can be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0078] Figure 1 A computing system for online top-down human pose tracking according to certain embodiments of the present disclosure is schematically depicted. Here, top-down means that pose estimation is performed after a candidate is detected, and the tracking is truly online. The system combines single pose tracking (SPT) with multi-person identity association and links keypoint tracking with object tracking.

[0079] The visual features of the Visual Object Tracking (VOT) method are implicitly represented by the kernel or Convolutional Neural Network (CNN) feature maps. In contrast, the present disclosure tracks each human pose by recursively updating the object / bounding box and its corresponding pose in an explicit manner. The bounding box region of the target object is inferred from explicit features, i.e., human key points. Human key points can be considered as a series of special visual features. The advantages of using pose as an explicit feature include the following. (1) The explicit features are human-related and interpretable, and have a very strong and stable relationship with the bounding box position. The human pose directly constrains the bounding box region. (2) The tasks of pose estimation and tracking first require predicting human key points. The predicted key points can be used to effectively track the region of interest (ROI) that is almost idle. This mechanism enables online tracking. (3) The identity of the candidates is naturally preserved, greatly reducing the burden of data association in the system. Even when data association is required, the present disclosure can reuse the pose features for skeleton-based pose matching. Therefore, SPT and single VOT are combined into a unified functional entity and can be easily implemented through a replaceable single-person human pose estimation module.

[0080] The contributions of the present disclosure particularly include the following. (1) The present disclosure proposes a general online pose tracking framework applicable to the top-down method of human pose estimation. Both the human pose estimator and the Re-ID (re-identification) module are replaceable. Compared with the Multi-Object Tracking (MOT) framework, the framework of the present disclosure is specifically designed for the pose tracking task. To our knowledge, this is the first public disclosure of an online human pose tracking system in a top-down manner. (2) The present disclosure provides a Siamese Graph Convolution Network (SGCN) for human pose matching as the Re-ID module. Different from the existing Re-ID modules, the present disclosure uses the graphical representation of human joints for matching. The skeleton-based representation effectively captures human pose similarity and has a low computational cost. It is robust to sudden camera movements that cause human drift. (3) Extensive experiments are conducted on various settings and ablation studies. The online pose tracking method of the present disclosure outperforms the existing online methods and is competitive with the offline state-of-the-art techniques, but with a much higher frame rate.

[0081] The present disclosure provides a novel top - down pose tracking network. In this framework, we consider both accurate human position and human pose estimation, where: (1) the rough human position can be refined into body key points by a single - person pose estimator; (2) the positions of human joints can be directly used to indicate the rough positions of candidates; and (3) thus, estimating one by one repeatedly is a feasible strategy for single - person pose tracking (SPT).

[0082] In addition, in the present disclosure, the multi - target pose tracking (MPT) problem is not merely considered as a repeated SPT problem for multiple individuals. The reason is that certain constraints need to be satisfied. For example, in a certain frame, two different IDs should not belong to the same person; two candidates should not have the same identity. Therefore, the present disclosure provides methods for tracking multiple individuals simultaneously and preserving / updating their identities using an additional Re - ID module.

[0083] The Re - ID module is essential because it is usually difficult to always maintain the correct identity. It is impossible to effectively track individual poses across the frames of an entire video. For example, in the following situations, the identity must be updated: (1) some people disappear from the camera's field of view or are occluded; (2) new candidates join or previous candidates reappear; (3) people cross each other (two identities may merge into one if not carefully handled); (4) rapid movement or zoom causes tracking failure.

[0084] The present disclosure first processes each candidate separately so that their corresponding identities are maintained across frames. In this way, the present disclosure avoids time - consuming offline optimization procedures. If the tracked candidates are lost due to occlusion or camera movement, the present disclosure then calls the detection module to recover the candidates and associate them with the tracking targets in the previous frame through pose matching. In this way, the present disclosure completes multi - target pose tracking through the SPT module and the pose matching module.

[0085] As Figure 1 shown, system 100 includes a computing device 110. In certain embodiments, Figure 1 the computing device 110 shown in may be a server computer, a cluster, a cloud computer, a general - purpose computer, a headless computer, or a dedicated computer that provides pose tracking services. The computing device 110 may include, but is not limited to, a processor 112, a memory 114, and a storage device 116. In certain embodiments, the computing device 110 may include other hardware components and software components (not shown) to perform its corresponding tasks. Examples of these hardware and software components may include, but are not limited to, other required memories, interfaces, buses, input / output (I / O) modules or devices, network interfaces, and peripherals.

[0086] The processor 112 can be a central processing unit (CPU) configured to control the operation of the computing device 110. The processor 112 can execute an operating system (OS) or other applications of the computing device 110. In some embodiments, the computing device 110 can have more than one CPU as the processor, such as two CPUs, four CPUs, eight CPUs, or any suitable number of CPUs.

[0087] The memory 114 can be a volatile memory, such as random access memory (RAM), for storing data and information during the operation of the computing device 110. In some embodiments, the memory 114 can be a volatile memory array. In some embodiments, the computing device 110 can operate on more than one memory 114.

[0088] In some embodiments, the computing device 110 can also include a graphics card to assist the processor 112 and the memory 114 in image processing and display.

[0089] The storage device 116 is a non-volatile data storage medium for storing an operating system (not shown) and other applications of the computing device 110. Examples of the storage device 116 can include non-volatile memories such as flash memory, memory cards, USB drives, hard disk drives, floppy disks, optical drives, or any other type of data storage device. In some embodiments, the computing device 110 can have multiple storage devices 116, which can be the same storage device or different types of storage devices, and the applications of the computing device 110 can be stored in one or more storage devices 116 of the computing device 110.

[0090] In this embodiment, the processor 112, the memory 114, and the storage device 116 are components of the computing device 110, such as a server computing device. In other embodiments, the computing device 110 can be a distributed computing device, and the processor 112, the memory 114, and the storage device 116 are shared resources from multiple computers in a predetermined area.

[0091] The storage device 116 particularly includes an attitude tracking application 118. In some embodiments, the attitude tracking application 118 is an online attitude tracking application program. The attitude tracking application 118 includes a scheduler 120, an object detection module 140, an attitude tracking module 160, a re-identification module 180, and an optional user interface 190. In some embodiments, the storage device 116 may include other applications or modules necessary for the operation of the attitude tracking application 118. It should be noted that the modules 120, 140, 160, 180, and 190 are all implemented by computer-executable code or instructions, or data tables or databases, and together form an application program. In some embodiments, each module may further include sub-modules. Alternatively, some modules may be combined into a stack. In other embodiments, certain modules may be implemented as circuits rather than executable code.

[0092] The scheduler 120 is configured to receive or retrieve a series of frames or images for processing. In some embodiments, the series of frames is a time series of frames, such as an online video or a live video. The video can be a Red Green Blue (RGB) video, a black and white video, or a video in any other format. The scheduler 120 is configured to define key frames from every N frames, where N is a positive integer. The quantity N can be predefined based on the complexity of the series of frames, the computing power of the computing device 110, the speed requirements for tracking, the accuracy requirements for tracking, and the like. In some embodiments, N ranges from 5 to 100. In some embodiments, N ranges from 5 to 30. In some embodiments, N ranges from 10 to 20. In one embodiment, N equals 10, i.e., the 1st, 11th, 21st, 31st,... frames of the video are key frames. Unless otherwise specified, the following embodiments are described with N = 10, but the present disclosure is not limited to this definition of N. In some embodiments, the scheduler 120 includes a counter for counting the frames to know what the current frame to be processed is, what the previous frame of the current frame is, and what the upcoming frame or the next frame after the current frame is. In some embodiments, the scheduler 120 is configured to load the video into the memory 114 in real time. In some embodiments, the scheduler 120 can load a predetermined number M of video frames into the memory 114, for example, in a sliding window fashion, and remove the oldest frame from the memory 114 in a first-in first-out (FIFO) manner. In some embodiments, the quantity M is the total length of the video, and the oldest frame is removed only when the usage of the memory 114 exceeds a predetermined percentage, such as 30%, 50%, 70%, etc., as long as the memory 114 is sufficient to hold the video. In some embodiments, the quantity M is a positive integer from 1 to 1000. In some embodiments, the quantity M is a positive integer from 1 to 100. In some embodiments, the quantity M is a positive integer from 1 to 10. In some embodiments, the quantity M is the positive integer 10, and the frames from one key frame to the next key frame are loaded into the memory 114. In some embodiments, the quantity M is 1, and the memory 114 only retains the current frame to be processed. The scheduler 120 is also configured to hold certain inputs and outputs of the frame processing steps in the memory 114. The input and output information can include object or target ID, bounding box ID (optional), key points of each bounding box, and the status of the target in the corresponding bounding box. In some embodiments, the scheduler 120 can also store those information in the storage device 116 or retain a copy of the information in the storage device 116. The scheduler 120 is configured to call modules 140, 160, and 180 to perform certain functions at different times.

[0093] After the video is loaded into the memory 114, the scheduler 120 is configured to process the frames of the video sequentially. When the first frame of the video is to be processed, the scheduler 120 is configured to call the object detection module 140 and the pose tracking module 160 to operate. After being called, the object detection module 140 is configured to detect objects (or targets, or candidates) from the first frame. Each object is defined with a bounding box, each object is labeled with an object ID and is associated with the frame (or optionally labeled with a frame ID), and includes the coordinates of the bounding box in the first frame. In some embodiments, each object is linked to the corresponding frame, and the frame ID is not required for object labeling. After the bounding box of the object is detected, the pose tracking module 160 is configured to estimate the pose of each object in parallel (or optionally estimate the pose of the object sequentially). For each bounding box of the object, the pose tracking module 160 is configured to determine the key points in the bounding box, such as using a convolutional neural network (CNN) and the heatmap generated by the CNN. Each key point corresponds to a confidence level based on the heatmap, indicating the likelihood that the key point is located in the frame. The key points define the pose of the object. In some embodiments, the number of key points for each object is 15, and these 15 key points are, for example, "right knee", "left knee", "right pelvis", "left pelvis", "right wrist", "left wrist", "right ankle", "left ankle", "right shoulder", "left shoulder", "left elbow", "right elbow", "neck", "nose", "head". In some embodiments, 15 key points are defined and stored in a specific order. It should be noted that the number of categories of key points is not limited to this, and the categories of key points can be more or less than those described in the embodiments. The key points define the pose of the object, and the determination of the key points as described above is also called pose estimation. After pose estimation, each object in the first frame is defined as: object ID and optional frame ID, coordinates of the bounding box, coordinates of the key points, and the confidence level of each key point or the average confidence level of all key points. In some embodiments, those explicit and simple information is stored in the memory 114. The pose tracking module 160 is then configured to infer the inferred bounding box of each object from the determined key points. In some embodiments, the inferred bounding box is inferred by: defining a bounding box (or namely the minimum bounding box or the smallest bounding box) with the minimum metric where all points are located, magnifying the rectangular box by 20% in the horizontal and vertical directions, and obtaining the inferred bounding box. The bounding box is also called the minimum bounding box or the smallest bounding box. In some embodiments, the bounding box is provided by selecting the topmost key point, the lowest key point, the leftmost key point, and the rightmost key point from the 15 key points and drawing a rectangular box. The top of the rectangle is determined by the y coordinate of the highest point, the bottom of the rectangle is determined by the y coordinate of the lowest point, the left side of the rectangle is determined by the x coordinate of the leftmost point, and the right side of the rectangle is determined by the x coordinate of the rightmost point.

[0094] At this time, the scheduler 120 determines that the first frame has been processed, optionally records the first frame as being processed by the counter, and is configured to call the pose tracking module 160 to process the second frame. In some embodiments, the pose tracking module 160 may continue to process frames until a frame count is met, at which point the scheduler 120 interrupts the operation of the pose tracking module 160. Based on the assumption of the present disclosure that the position and pose of an object are generally similar in two consecutive frames, the pose tracking module 160 is configured to perform a CNN using the inferred bounding box and the second frame, generate a heatmap, and determine key points by using the heatmap, wherein all key points in the second frame are located in the area enclosed by the inferred bounding box. Each key point has a confidence level, and the pose tracking module 160 is further configured to determine the object state in the second frame based on the average confidence level of the key points. If the average confidence level is greater than a threshold, the object state in the second frame is "tracked", and if the average confidence level is equal to or lower than the threshold, the object state in the second frame is "lost". When the object state is "tracked", the pose tracking module 160 is configured to generate a bounding box using four key points from the second frame and infer the inferred bounding box from the bounding box. When the object state in the second frame is determined to be "tracked", the inferred bounding box from the first frame is considered to be the bounding box of the corresponding object in the second frame. At this time, for an object determined to be "tracked" in the second frame, the pose tracking module 160 stores the frame ID, object ID, coordinates of the bounding box (inferred bounding box), coordinates of the key points, and the confidence levels of the key points. Similar to the above process, the pose tracking module 160 uses the inferred bounding box from the second frame to estimate the key points in the third frame, determine the object state, and when the object state passes the check, generate a bounding box using four key points and infer the inferred bounding box. The pose tracking module 160 is configured to continue processing consecutive frames until the next key frame (e.g., frame 11) or a specific non-key frame (when the object state in the non-key frame is "lost"). For the frame preceding the next key frame, e.g., frame 10, the pose tracking module 160 determines the object state, and if the object state is "tracked", generates a bounding box based on the key points as the bounding box of frame 10. However, it is not necessary to infer the inferred bounding box because the scheduler 120 will schedule a new object detection for the next key frame.

[0095] As described above, the first frame (or Frame 1) is defined as the first key frame, and Frame 11 is defined as the second key frame. When the scheduler 120 counts that Frame 11 is to be processed, the scheduler 120 is configured to call the object detection module 140, the pose tracking module 160, and the re-identification module 180 to process the key frame 11. Specifically, the object detection module 140 detects objects from the key frame 11 and outputs a bounding box for each object. For each bounding box corresponding to one of the detected objects, the pose tracking module 160 determines key points and estimates the pose of the object. The re-identification module 180 compares the key points of each object with the key points of the stored objects from the previous frame to see if there is a match. If there is a match, the same object ID as the matching object in Frame 10 is assigned to the object in key frame 11. If no match for the object is found, a new object ID is assigned to the object. In this embodiment, matching is performed between the objects from the current frame and those from the previous frame. In other embodiments, matching can also be performed between the objects from the current frame and the objects from multiple previous frames, such as 2, 5, or 10 previous frames. After re-identification, the pose tracking module 160 continues to infer the bounding box, estimate the key points, and calculate the object state in Frame 12. The subsequent processing of Frames 12 to 19 is substantially the same as the processing of Frames 2 to 9. Similarly, the processing of Frames 20 and 21 is substantially the same as the processing of Frames 10 and 11.

[0096] In some embodiments, each object is tracked independently of other objects by the pose tracking module 160. When an object is "lost", the scheduler 120 is configured to direct object detection, pose estimation, and re-identification for all objects immediately from the next frame. In some embodiments, if the number of "lost" objects is greater than a predetermined threshold or the percentage of "lost" objects is greater than a predetermined percentage, the scheduler 120 starts new object detection from the next frame. In some embodiments, if only one or several objects among multiple objects are "lost", the scheduler 120 may abandon the tracking of the "lost" objects and continue the tracking of the "tracked" objects until new object detection starts at the next key frame.

[0097] When processing a specific non-key frame (the object state in the frame is "lost"), the processing of this frame is similar to the processing of a key frame. In other words, the "lost" state of the object triggers new object detection in a specific non-key frame. In some embodiments, when the quality of a specific non-key frame is low, the process can also skip this frame and start processing the next frame. During re-identification, the objects (key points) of the next frame are compared with one of the previous frames (not necessarily the adjacent previous frame).

[0098] After being called, each of the object detection module 140, the pose tracking module 160, and the re-identification module 180 can remain active in the memory 114 to process upcoming frames according to the instructions of the scheduler 120. In some embodiments, the pose tracking application 118 may not include the scheduler 120, and the functions of the scheduler 120 may be assigned to the object detection module 140, the pose tracking module 160, and the re-identification module 180.

[0099] The object detection module 140 is configured to detect objects from the first frame (key frame), subsequent key frames, and any specific non-key frames with the object state being "lost". The object detection module 140 in the pose tracking application 118 can be any object detector and is replaceable. In some embodiments, the object detection module 140 uses a convolutional neural network (CNN) for object detection. In some embodiments, the object detection module 140 includes a deformable ConvNet with ResNet101 as the backbone, a feature pyramid network (FRN) for feature extraction, and a fast R-CNN scheme as the detection head. In some embodiments, the object detection module 140 can use other types of detection processes, such as using R-FCN as the backbone, or using ground truth (GT) detection.

[0100] In some embodiments, the input of the object detection module 140 is an RGB frame, and the output of the object detection module 140 is multiple objects, each object being defined with a bounding box. In some embodiments, the object can be only a human body. As described above, each object can be characterized by a frame ID and bounding box coordinates. In some embodiments, the bounding box of an object can be represented by a feature vector or simple text. In some embodiments, the features of the bounding box are preferably stored in the memory 114, or alternatively stored in at least one of a graphics card or video memory, a storage device 116, or a remote storage device accessible by the computing device 110.

[0101] The pose tracking module 160 is used to determine the key points of the object in the corresponding bounding box, estimate the pose of the object in the current frame, assign an object ID to the object, and store the object with the frame ID, object ID, coordinates of the bounding box, coordinates of the key points, and confidence of the key points into the memory 114 after receiving the object from the first frame (first key frame) of the object detection module 140. The pose tracking module 160 is also configured to generate a bounding box surrounding the key points and infer the inferred bounding box to process the second frame.

[0102] The pose tracking module 160 is configured to, when receiving an object in a key frame (except the first frame) or an object in a specific non-key frame (with a "lost" object state), determine the key points of the object in the corresponding bounding box and estimate the pose of the object in the frame, re-identify the object through the re-identification module 180, and store the object with the frame ID, object ID, coordinates of the bounding box, coordinates of the key points, and confidence of the key points in the memory 114. The pose tracking module 160 is further configured to generate a bounding box surrounding the bounding box based on the key points, and infer the inferred bounding box for processing the next frame.

[0103] The pose tracking module 160 is configured to, when receiving an inferred bounding box (inferred based on the key points of the object in the previous frame and marked with the object ID) of a non-key frame (the current frame), determine the key points of the object in the current frame and estimate the pose of the object in the current frame, and store the object with the frame ID, object ID, bounding box coordinates, key point coordinates, and key point confidence in the memory 114. The pose tracking module 160 is further configured to determine the object state in the current frame. If the object state is "tracked", a bounding box surrounding the bounding box is generated based on the key points, and the inferred bounding box is inferred for processing the next frame. Then, the pose tracking module 160 is further used to process the next frame according to the inferred bounding box by: determining the key points in the next frame surrounded by the inferred bounding box, estimating the pose of the object, calculating the confidence of the key points, and storing the key points and their bounding boxes. By repeating the above process, the recursive update of the bounding box and pose of the object in consecutive frames is achieved. Therefore, the tracking of the object in a series of frames is achieved. Since the tracking uses bounding boxes and key points instead of kernels or feature maps, the tracking speed is very fast and can be applied to online tracking. Otherwise, if the object state in the current frame is "lost", the scheduler 120 will instruct to restart determining the object in the current frame.

[0104] As Figure 2 shown, the pose tracking module 160 includes a pose estimation module 162, an object state module 164, and a bounding box inference module 166. The pose estimation module 162 is configured to determine the key points and estimate the pose of each object when receiving the detection result from the object detection module 140. In some embodiments, the pose estimation module 162 is a single-person pose estimator or a single-person pose tracking module. If the object key points do not match the stored object key points, the object with the bounding box and the corresponding key points will be stored with a new object ID; if there is a match between the object key points and the stored object key points, the object will be stored with the old ID.

[0105] The pose estimation module 162 is configured to, when receiving an inferred bounding box from the bounding box inference module 166 (based on a previous frame), determine key points and estimate the pose of each object (in the current frame), obtain the object state calculated by the object state module 164, and when the object state is "tracked", store the object together with the bounding box and key point information. If the object state is "lost", the scheduler 120 will instruct to restart object detection for the object in the current frame.

[0106] As described above, the pose estimation module 162 can receive the detected bounding box from the object detection module 140 or the inferred bounding box from the bounding box inference module 166 to repeatedly perform pose estimation and key point determination. In some embodiments, the inferred bounding box is resized before being input to the pose estimation module 162 such that the input detected bounding box or inferred bounding box has a fixed size. In some embodiments, the output of the pose estimation module 162 is a series of heatmaps for subsequently predicting key points. Each of the consecutive key points corresponds to a specific part of the human body and represents the pose of the human object. The pose estimation module 162 is also used to store the result after obtaining the pose estimation and key point determination results. The result may include the frame ID, object ID, coordinates of the bounding box, coordinates of the key points, and the confidence of the key points (or average confidence). The pose estimation module 162 uses the rough position of the object defined by the bounding box to help determine the key points of the object, and the key points of the object indicate the rough position of the object in the frame. By repeatedly estimating the object position and the key points of the object, the pose estimation module 162 can estimate the object pose and accurately and effectively determine the object key points.

[0107] The object state module 164 is configured to determine the state of the object in the current frame when receiving the key points determined based on the current frame and the inferred bounding box (inferred from the previous frame). The state of the object is "tracked" or "lost". In some embodiments, the object state module 164 is configured to use the confidence to determine the state. Specifically, using the inferred bounding box from the previous frame and the determined key points from the current frame, the object state module 164 determines the likelihood that the determined key points are located in the region of the current frame covered by the inferred bounding box. The likelihood is represented by the confidence s, and each estimated key point has a confidence. Calculate the average value of the confidences of the estimated key points of the object and compare it with the standard error τ s If the average confidence is greater than the standard error τ s , it indicates that the state of the object in the current frame is "tracked", that is, the object is considered to be located in the inferred bounding box in the current frame. Otherwise, the state of the object is "lost", that is, the object is considered not to be in the current frame surrounded by the inferred bounding box. The state of the object is defined as:

[0108]

[0109] In the above embodiments, the average of the confidence levels of all key points (e.g., 15 key points) is used. In other embodiments, several key points with the highest confidence levels may also be used to calculate the average of the confidence levels. For example, the average confidence level may be the average of the confidence levels of the top 3, top 4, or top 5 key points.

[0110] In certain embodiments, the confidence level is from a heat map generated by operating a convolutional neural network, and the obtained confidence level has a value ranging from 0 to approximately 1.5. In certain embodiments, the obtained confidence level is the sigmoid value of the confidence level from the CNN, where the confidence level is in the range of 0 to 1. In certain embodiments, when the confidence level is in the range of 0 to 1, the standard error τ s or the threshold is predefined as a value between 0.3 and 1.0. In certain embodiments, the standard error τ s is in the range of 0.6 to 0.9. In certain embodiments, the standard error τ s is in the range of 0.7 to 0.8. In one embodiment, the standard error τ s is set to 0.75.

[0111] The bounding box inference module 166 is configured to infer an inference bounding box for processing the next frame when receiving the key points of the object in the current non-key frame. The inference bounding box is regarded as the local area of the object in the next frame. In certain embodiments, the inference bounding box is defined by determining a bounding box based on the topmost, bottommost, leftmost, and rightmost key points and enlarging the bounding box. In certain embodiments, the inference bounding box is obtained by enlarging the bounding box by 5 to 50% in the x and y directions. In certain embodiments, it is enlarged by about 10 to 30%. In certain embodiments, it is enlarged by 20%; for example, if the coordinates of the four corners of the current bounding box are (x1, y1), (x2, y1), (x1, y2), (x2, y2) respectively, where x2 > x1 and y2 > y1, then the coordinates of the four corners of the inference bounding box will be ((1.1x1 - 0.1x2), (1.1y1 - 0.1y2)), ((1.1x2 - 0.1x1), (1.1y1 - 0.1y2)), ((1.1x1 - 0.1x2), (1.1y2 - 0.1y1)), ((1.1x2 - 0.1x1), (1.1y2 - 0.1y1)).

[0112] Note that when an object or target is lost, the present disclosure provides two correction methods: (1) Fixed Keyframe Interval (FKI) mode: Ignore this target until the next predetermined keyframe. The detection module regenerates candidates and then associates their IDs with the tracking history. (2) Adaptive Keyframe Interval (AKI) mode: Immediately restore the lost target through candidate detection and identity association. In some embodiments, the advantage of the FKI mode is that since the interval of keyframes is fixed, the frame rate of pose tracking is stable. The advantage of the AKI mode is that the average frame rate for non-complex videos can be higher. In some embodiments shown in the experimental section, the present disclosure combines them by adopting keyframes with a fixed interval, and at the same time, once the target is lost, the detection module can be called before the next arranged keyframe arrives. This makes the tracking accuracy higher because when the target is lost, it will be processed immediately.

[0113] The re-identification module 180 is configured to compare the key points of each detected object with the previously stored key points of each stored object after the object detection module 140 detects an object (bounding box) from a keyframe (but not the first frame) or a specific non-keyframe (in which the object state is "lost"), and the pose tracking module 160 determines the key points of each object. When one of the detected objects matches one of the previously stored objects, the re-identification module 180 assigns the same object ID as the matching object to the detected object, and assigns a new object ID to the detected object when there is no match. In some embodiments, the comparison is made with the bounding box / key points of the object from the previous frame, or with the bounding box / key points of the object from several previous frames. In some embodiments, the pose sequences from multiple previous frames are used together with the spatio-temporal SGCN to provide more robust results. In some embodiments, when the object detection module 140 detects an object from the first frame (which is also the first keyframe), the re-identification module 180 does not need to be called because there are no stored objects yet, and each object is assigned a new object ID.

[0114] In some embodiments, the re-identification module 180 considers two complementary pieces of information: spatial consistency and pose consistency for re-identifying the bounding box. In some embodiments, the match between the bounding box from the current frame and the bounding boxes from one or more previous frames is determined by their adjacency. There is an Intersection over Union (IOU) between two bounding boxes. When the IOU between two bounding boxes is greater than a threshold, the two bounding boxes are considered to belong to the same object (or target). In some embodiments, for keyframe k, if the tracked object t k ∈T k corresponds to the detection d k ∈Dk The maximum IOU overlap rate o(t k , D i,k ) is higher than the threshold τ o , then the matching flag m(t k , d k ) is set to 1. Otherwise, m(t k , d k ) is set to 0:

[0115]

[0116] In some embodiments, the above criterion is based on the assumption that the tracked target (bounding box) from the previous frame has a significant overlap with the actual position of the target (bounding box) in the current frame.

[0117] However, such an assumption is not always reliable, especially when the camera moves rapidly. In some embodiments, we need to match new observations to the candidates being tracked. In the Re-ID problem, this is typically done by a visual feature classifier. However, visually similar candidates with different identities may confuse these classifiers. In an online tracking system, the computational cost of extracting visual features may also be high. Therefore, to overcome these drawbacks, the present disclosure designs a graph convolutional network (GCN) to utilize the graphical representation of human joints or human key points. The present disclosure observes that in two adjacent frames, the position of a person may drift due to a sudden movement of the camera, but the human pose will remain almost the same because people's movements are usually not that fast, as Figure 3 shown. Figure 3 The left frame shows consecutive adjacent frames with a sudden movement of the camera, where the movement in the third frame from the top is obvious and the tracking of the human target is unlikely to be successful. Figure 3 The right frame shows a sudden zoom, where the 3rd - 5th frames from the top suddenly shrink and it is unlikely to successfully track the human target using the same / similar sized bounding boxes. In some embodiments, the present disclosure uses the graphical representation of the human skeleton for candidate matching, which is named pose matching, as described below.

[0118] In some embodiments, the re-identification module 180 uses a Siamese graph convolutional network (SGCN) for pose matching, i.e., the re-identification module 180 uses the pose of the object in the bounding box to match two objects. Figure 4 Schematically shows an SGCN according to some embodiments of the present disclosure. As Figure 4As shown, the SGCN includes two repeated sub-models 410 and 430 for receiving input key points 411 from the detected object and input key points 431 from the object (bounding box) dataset. The sub-model 410 includes a first graph convolutional layer 412, a first Relu (Rectified Linear) unit 413, a second graph convolutional layer 414, a second Relu unit 415, a first average pooling layer 416, and a first fully connected layer 417. Similarly, the sub-model 430 includes a third graph convolutional layer 432, a third Relu unit 433, a fourth graph convolutional layer 434, a fourth Relu unit 435, a second average pooling layer 436, and a second fully connected layer 437. In some embodiments, the layers 412 to 417 are arranged in sequence, the layers 432 to 437 are arranged in sequence, and the model structures and parameters of the sub-models 410 and 430 are the same.

[0119] When pose matching is required, the re-identification module 180 is configured to input the key points 411 from the target object (bounding box) into the first GCN layer 412 and input the key points 431 from the stored object (bounding box) into the third GCN layer 432. In some embodiments, the key points have two-dimensional coordinates in the detected figure (current frame), and the key points 411 are in the form of coordinate vectors. The key points 411 are the input to the first GCN layer 412. In some embodiments, the first GCN layer 412 constructs a spatial graph with the key points as graph nodes and the connectivity in the human body structure as graph edges, and inputs the key points 411 on the graph nodes. As Figure 4 shown, in some embodiments, the first GCN layer 412 has a dual-channel input and outputs an output of 64 channels. The output of 64 channels passes through the first Relu unit 413 and is input to the second GCN layer 414. Both the input and output of the second GCN layer 414 have 64 channels. The output of the second GCN layer 414 passes through the second Relu unit 415 and the average pooling layer 417 to form a 128-dimensional feature representation vector 418 as the conceptual summary of the human pose. Through the same process, the key points 431 in the stored object (bounding box) are input to the sub-model 430, and after being processed by the sub-model 430, a 128-dimensional feature representation vector 438 is output. Then, the re-identification module 180 matches the feature representation vectors 418 and 438 in the vector space. If these two vectors are close enough in the vector space, it is determined that these two vectors represent the same pose. In some embodiments, a threshold is predefined for judgment. When the distance between the two feature vectors is less than the threshold, the two feature vectors are considered to match. It should be noted that the pose of the human body is represented by the latent feature vector generated by the GCN and can correspond to standing, sitting, or any other possible human pose.

[0120] In some embodiments, the SGCN network is optimized with a contrastive loss L because the present disclosure wants the network to generate a feature representation that is close enough for positive pairs and at least far from the minimum for negative pairs. The present disclosure employs a Margin Contrastive Loss:

[0121]

[0122] where D = ||f(p j ) - f(pk)||2 is the Euclidean distance between two l2-norm normalized latent representations, and y jk ∈ {0, 1} indicates whether p j and p k are the same pose, and ∈ is the minimum distance margin that pairs describing different poses should satisfy.

[0123] The graph convolutional layers 412, 414, 432, 434 perform convolutions on the skeleton, i.e., human joints or key points. In some embodiments, for standard 2D convolutions on natural images, the output feature map can have the same size as the input feature map, with a stride of 1 and appropriate padding. In some embodiments, the graph convolution operation is similarly designed to output a graph with the same number of nodes. The attribute dimension of these nodes, similar to the number of feature map channels in standard convolutions, can change after the graph convolution operation.

[0124] In some embodiments, the standard convolution operation is defined as follows: Given a convolution operator with a kernel size of K×K and an input feature map f in with c channels, the output value of a single channel at spatial position x can be written as:

[0125]

[0126] where the sampling function s: Z 2 ×Z 2 →Z 2 enumerates the neighbors of position x. The weight function W: has a vector at each node of the graph. The next step in the extension is to redefine the sampling function p and the weight function w. The present disclosure follows the method proposed by Yan (Yan, S et al., Spatial temporal graph convolutional networks for skeleton-based action recognition, AAAI, 2018, which is incorporated herein by reference in its entirety). For each node, only its adjacent nodes are sampled. The neighbor set of node v i is: B(v i ) = {vj d(v j , v i ) ≤ 1}. The sampling function p: B(v i ) → V can be written as: p(v i , v j ) = v j . In this way, the number of adjacent nodes is not fixed, nor is the weight order. To obtain a fixed number of samples and a fixed weighted order, the present disclosure labels the neighbor nodes around the root node into a fixed number of partitions, and then weights these nodes based on their partition categories. Figure 5 Schematically shows a spatial configuration partitioning strategy for graph sampling and weighting to construct a graph convolution operation. As Figure 5 shown, for two skeletons 510 and 530, the centers of gravity 512 and 532 of the skeletons are determined. Nodes are labeled according to the comparison of the distance from the node to the center of gravity 512 / 532 of the skeleton and the distance from the root node 514 / 534 to the center of gravity 512 / 532 of the skeleton. Compared with the root node, the centripetal nodes 516 / 536 have shorter distances, while the centrifugal nodes 518a / 518b / 538 have longer distances.

[0127] Therefore, equation (4) of the graph convolution is rewritten as:

[0128]

[0129] where the normalization term Z i (v j ) = |{v k |l i (v k ) = l i (v j )}| is to balance the contributions of different subsets to the output. According to the above partitioning method, the present disclosure has:

[0130] where r i is the average distance from the center of gravity to joint or key point i for all frames in the training set.

[0131] In some embodiments, the re-identification module 180 uses both spatial consistency and pose consistency for bounding box (or object) matching. When spatial consistency is determined, a match between two bounding boxes is determined and further pose matching may not be required. When spatial consistency is determined to be mismatched, further pose matching is performed because a person's pose is unlikely to change between two adjacent frames. In some embodiments, other combinations of spatial consistency and pose consistency are possible. For example, matching bounding boxes with mismatched poses can also be considered mismatched. In some embodiments, when the re-identification module 180 determines that a detected bounding box matches a stored bounding box, the re-identification module 180 is used to assign the same object ID (or target ID) to the detected bounding box as the stored bounding box.

[0132] Return reference Figure 1 , the pose tracking application 118 may further include a user interface 190. The user interface 190 is configured to provide a usage interface or a graphical user interface in the computing device 110. In some embodiments, the user is able to configure parameters for training or using the pose tracking application 118.

[0133] In some embodiments, the pose tracking application 118 may further include a database, which may be configured to store at least one of the training data, online videos, and parameters of the pose tracking application 118 and the input / output from the pose tracking application 118. However, the pose tracking application 118 preferably loads the videos, bounding boxes, etc. to be processed into the memory 114 for fast processing. In some embodiments, the sub-modules of the pose tracking application 118 are designed as different layers of an integrated network, with each layer corresponding to a specific function.

[0134] Figure 6 is a flowchart showing the process of executing the pose tracking application 118. The solid lines indicate the operations of the scheduler instructing programs 601-609, and the dashed lines indicate the possible data flows coordinated by the scheduler 120. The scheduler 120 also counts and tracks the frames of the video to be processed.

[0135] As Figure 6As shown, when processing a video, the scheduler 120 loads the video into the memory 114, retrieves the first frame (which is also the first key frame), and invokes the operations of object detection 601, pose estimation 603, and bounding box inference 609. Specifically, the scheduler 120 invokes the object detection module 140 to perform object detection 601. Each object is labeled with a frame ID (frame 1), an object ID (a newly assigned ID, e.g., from object 1), and coordinates (the two-dimensional coordinates of the four corners of the bounding box or the two diagonal corners of the bounding box). The scheduler 120 then invokes the pose estimation module 162 of the pose tracking module 160 to perform pose estimation 603, specifically to determine key points from each bounding box. The sequential key points determine the pose of the object. In some embodiments, the number of key points is 15. Heatmaps generated by the pose estimation module 162 can be used to determine the key points, and each key point has a confidence level indicating the likelihood that the key point exists within the bounding box. For the first frame, the scheduler 120 may not indicate the operations of state determination 605 and re-identification 607, but indicates the operation of bounding box inference 609. The bounding box inference module 166 infers a bounding box from each bounding box, e.g., by determining an enclosing bounding box using four key points among the topmost, bottommost, leftmost, and rightmost, and enlarging the enclosing bounding box. The inferred bounding box has the same object label as the corresponding current bounding box.

[0136] At this time, the scheduler 120 retrieves the second frame and counts the number of frames. In this embodiment, frame 2 is a non-key frame, and the key frames are every 10 frames, i.e., frames 1, 11, 21, 31,.... The scheduler 120 invokes the operations of pose estimation 603, state determination 605, and bounding box inference 609. Specifically, the scheduler 120 invokes the pose estimation module 162 to determine the key points in the second frame based on the inferred bounding box and the second frame. The object state module 164 then uses the confidence levels of the determined key points to determine the object state in frame 2. When the object state is "tracked", the key points and the bounding box (enclosing bounding box) are stored, and the process proceeds to bounding box inference 609 to generate an inferred bounding box. Similarly, as described above, the inferred bounding box based on the second frame is used for the processing of the third frame. As long as the result of state determination 605 is "tracked", the process of bounding box inference 609, pose estimation 603, and state determination 605 is repeated for non-key frames. In each repetition, the frame ID, object ID, bounding box coordinates, and key point coordinates with confidence levels are stored in the memory 114. For the non-key frame (e.g., non-key frame 10) preceding a key frame (e.g., key frame 11), the process may only include bounding box inference 609 based on frame 9, and pose estimation (key point determination) based on the inferred bounding box and frame 10, but not include state determination 605, because the scheduler 120 will in any case indicate new object detection for frame 11.

[0137] For key frames other than the first frame, such as frame 11, the scheduler 120 invokes the operations of object detection 601, pose estimation 603, re-identification 607, and bounding box inference 609. Specifically, when frame 10 is being processed and the pose estimation 603 for frame 10 has been completed, the scheduler fetches frame 11, invokes the object detection module 140 to resume detecting the bounding box, invokes the pose estimation module 162 to determine the key points, and invokes the re-identification module 180 to perform pose matching of the detected bounding box / key points with the stored bounding box / key points. In some embodiments, the detected bounding box / key points in frame 11 and frame 10 are compared. In some embodiments, the detected bounding box / key points in frame 11 and those in several previous frames are compared, such as frames 8 - 10, frames 6 - 10, frames 2 - 10, the frames from the previous key frame to frame 10, or more previous frames. When there is a match, the detected bounding box / key points are labeled with the same object ID as the matching bounding box / key points. When there is no match, the object corresponding to the detected bounding box is considered a new object, and the detected bounding box / key points are labeled with a new object ID. Thereafter, the scheduler 120 instructs the bounding box inference module 166 to perform bounding box inference.

[0138] In the above embodiments, when an object is lost in the current frame, object detection 601 is performed to identify all objects in the current frame. In some embodiments, when an object is determined to be lost in the current frame due to its low confidence, this object may correspond to a different object. Therefore, in some embodiments, the key points of this object can also be directly compared with the objects stored in the system to determine whether this object is actually another object. If this object is determined to be another object, the label of that other object is assigned to this object.

[0139] Furthermore, for each frame, each object has a corresponding bounding box, and the key points of the object are located within the bounding box. In some embodiments, due to the one-to-one correspondence between the object and its bounding box (where the object is located), the object and the bounding box can be used interchangeably.

[0140] In summary, the scheduler 120 coordinates the execution of the different functions of the object detection module 140, the pose tracking module 160, and the re-identification module 180. The object detection module 140 includes a CNN or any other neural network, with the input being a key frame or a specific non-key frame, and the output being the detected objects with bounding boxes. The pose tracking module 160 has a CNN for cyclically tracking objects. The pose tracking module 160 is a single-person tracking module that can operate on different objects separately. The input is the detected bounding box or the inferred bounding box, and the output is the key points of the bounding box derived from the heat map. The input bounding box can be preprocessed to have a fixed size. The re-identification module 180 has an SGCN structure, with the input being the bounding boxes and key points for comparison, and the output being two pose feature vectors or specifically the distance between the two pose feature vectors, and determining whether the distance is small enough for the two objects defined by the bounding boxes and key points to be the same object. During operation, the objects in the frame are stored in the memory 114, and each stored entry can include: frame ID, object ID, coordinates of the bounding box, coordinates of the key points in the bounding box, and confidence of the key points.

[0141] Note that most of the above processes are for one object. When processing multiple objects in each frame, the tracking process for each object is similar. In some embodiments, the tracking of each object is performed independently. In other embodiments, a certain synchronization is performed. For example, if all objects are tracked, the object state is defined as "tracked". That is, if the pose tracking application 118 loses tracking on any object, new object detection is performed to restart the detection of multiple objects.

[0142] Figure 7A and FIG. Figure 7B schematically depicts an online top-down human pose tracking process according to certain embodiments of the present disclosure. In some embodiments, the tracking process is performed by a computing device, such as Figure 1 the computing device 110 shown in, and specifically by the pose tracking application 118. The network in the pose tracking application 118 is pre-trained before performing this process. In some embodiments, Figure 7A and Figure 7B the processes shown in Figure 6 are similar or the same as the process shown in Figure 7A and Figure 7B It should be noted that, unless otherwise specified in the present disclosure, the steps of the pose tracking process or method can be arranged in different orders, and thus are not limited to the order shown in

[0143] After the pose tracking application 118 is trained, the pose tracking application 118 is ready to perform pose tracking on multiple objects, especially in an online environment. In the following process, the i-th frame is a non-critical frame, and the j-th frame is a critical frame or a specific non-critical frame where the object state is "lost", where i and j are positive integers. Figure 7A Schematically shows the process of processing the j-th frame, Figure 7B Schematically shows the process for processing the i-th frame.

[0144] As Figure 7A shown, in step 702, an online video is provided, the online video includes a plurality of consecutive frames arranged in sequence, and the scheduler 120 loads the video into the memory 114 for processing. In the following steps, the scheduler 120 invokes the operations of other modules, such as the operations of the object detection module 140, the pose tracking module 160 or its sub-modules, and the re-identification module 180. In some embodiments, these operations may also not be invoked by the scheduler 120, but are encoded in the modules 140, 160, and 180.

[0145] In step 704, for the j-th frame of the video, that is, the critical frame or a specific non-critical frame where the object state is "lost", the object detection module 140 detects objects from the j-th frame and defines a bounding box surrounding each detected object. The object has a frame ID (e.g., j) and the coordinates of the corresponding bounding box (e.g., the 2D coordinates of two diagonal corner points).

[0146] In step 706, the pose estimation module 162 determines the key points of each bounding box. Now each object has a frame ID and its bounding box coordinates, and the key points have a class ID and their coordinates. The class ID of the key points corresponds to various parts of the human body, such as 15 key points corresponding to the head, shoulders, elbows, wrists, waist, knees, ankles, etc. Each key point has a confidence level.

[0147] In step 708, the re-identification module 180 matches the pose of the detected object (with a bounding box / key points) with the pose of the stored object (with a bounding box / key points). When there is a match, the re-identification module 180 assigns the same object ID as the object ID of the matching object to the detected object. If there is no match, the re-identification module 180 assigns a new object ID to the detected object. The detected object now has the characteristics of a frame ID, an object ID, key points, the coordinates of the bounding box / key points, and the key point confidence level, which are stored in the memory 114 or other predefined storage locations.

[0148] In some embodiments, the poses of the detected object and the stored object are determined using a siamese graph convolutional network, such as Figure 4As shown. In some embodiments, when the first frame is processed and there is no stored object information, step 708 does not need to be performed. Instead, the object (with a bounding box and key points) refined by the pose estimation module 162 is assigned a new object ID and stored in the memory 114 for later use.

[0149] In step 710, the bounding box inference module 166 infers a bounding box based on the key points in the current frame. In some embodiments, the bounding box inference module 166 first uses the four key points at the topmost, lowest, leftmost, and rightmost positions to define an enclosing bounding box, and then enlarges the enclosing bounding box by 20% in the horizontal and vertical directions, expanding equally from both the left and right sides and from both the top and bottom sides.

[0150] In step 712, the pose estimation module 162 uses the inferred bounding box and the (j + 1)-th frame to estimate the key points in the (j + 1)-th frame in the area covered by the inferred bounding box. The pose estimation module 162 outputs a heat map for estimating the key points, and each key point has a confidence level based on the heat map.

[0151] In step 714, the object state module 164 calculates the object state based on the confidence levels of the key points estimated in the (j + 1)-th frame. When the average confidence level is large, the object state is "tracked". This can indicate that the inferred bounding box and the estimated key points fit the image of the (j + 1)-th frame, and steps 760 - 754 - 756 - 758 described below are performed on the (j + 1)-th frame. When the average confidence level is small, the object state is "lost". This can indicate that the inferred bounding box and the estimated key points do not fit the image of the (j + 1)-th frame, and steps 704 to 714 are repeated on the (j + 1)-th frame.

[0152] When the state is "tracked", the pose tracking application 118 can also store the detected object (with a bounding box and key points) in the memory 114 for the j-th frame after step 708.

[0153] As Figure 7B shown, in step 752 which is the same as step 702 above, an online video is provided.

[0154] In step 754, for the i-th frame of the video, the pose estimation module 162 receives the inferred bounding box based on the (i - 1)-th frame, and the i-th frame is a non-key frame (where the object state is subsequently determined to be "tracked").

[0155] In step 756, the pose estimation module 162 estimates the key points based on the inferred bounding box and the i-th frame, and obtains the estimated key points with confidence levels.

[0156] In step 758, the object state module 164 calculates the object state in the i-th frame using the estimated key points. When the average confidence is high, the object state is "tracked". This can indicate that the inferred bounding box and the estimated key points fit the image of the i-th frame, and then steps 760-754-756-758 are performed on the (i+1)-th frame. When the average confidence is low, the object state is "lost". This can indicate that the inferred bounding box and the estimated key points do not fit the image of the i-th frame, and then steps 704-714 are performed on the i-th frame.

[0157] In step 760, the bounding box inference module 166 infers the inferred bounding box based on the estimated key points.

[0158] When the state is "tracked", the pose tracking application 118 can also store the inferred bounding box and key points in the memory 114 in step 758 for the i-th frame, or, for the i-th frame, store the bounding box and key points in the memory 114 in step 760.

[0159] By Figure 7A and Figure 7B the above operations described in

[0160] the pose tracking application 118 is able to provide online tracking of an object in a video. Advantages of the present disclosure include the overall model design, object re-identification based on pose matching, object state determination based on the current frame and upcoming frames, and decision-making based on pose matching and object state. Since object detection is only performed on key frames and specific non-key frames, the performance of the application is more efficient. In addition, using key points based on the enlarged region for tracking is both fast and accurate. In addition, the present disclosure uses GCN to encode the spatial relationship between human joints as a latent representation of human pose. The latent representation robustly encodes the pose, which is not affected by the position or perspective of the person, and then measures the similarity of such encodings to match the human pose.

[0161] In some aspects, the present disclosure relates to a non-transitory computer-readable medium storing computer-executable code. In some embodiments, the computer-executable code can be the software stored in the storage device 116 as described above. The computer-executable code, when executed, can perform one of the above methods.

[0162] Experiment: Experiments have been conducted using a model according to some embodiments of the present disclosure.

[0163] Experiment: 1. Dataset. PoseTrack is a large-scale benchmark for human pose estimation and joint tracking in videos. It provides publicly available training and validation sets, as well as an evaluation server for benchmarking on a held-out test set. This benchmark was the basis for the challenges at the ICCV′17 and ECCV′18 workshops. The dataset consists of over 68,000 frames for the ICCV′17 challenge and was extended to twice the number of frames for the ECCV′18 challenge. It now includes 593 training videos, 74 validation videos, and 375 test videos. For the held-out test set, for the same method, up to four submissions per task are allowed. There is no submission limit for the evaluation of the validation set. Thus, the ablation study described below was conducted on the validation set. Since the PoseTrack′18 test set has not been opened, we compare our results with other methods in the experimental performance comparison section of the PoseTrack′17 test set. Figure 8 The pose pairs collected from the PoseTrack′18 dataset are shown in Table 1.

[0164] Experiment: 2. Evaluation Metrics. The evaluation includes pose estimation accuracy and pose tracking accuracy. Pose estimation accuracy is evaluated using the standard mAP metric, while the evaluation of pose tracking is based on the explicit MOT metric, which is the evaluation criterion for multi-object tracking.

[0165] Experiment: 3. Implementation Details. We adopted the state-of-the-art key-frame object detector trained on the ImageNet and COCO datasets. Specifically, we used the pre-trained model from the deformable convolutional networks (ConvNets). We conducted experiments on the validation set to select the object detector with better recall. For the object detector, we compared the deformable convolutional versions of the RFCN network and the FPN network with the ResNet101 backbone. The FPN feature extractor was connected to the Fast R-CNN head for detection. We compared the detection results with the ground truth according to the precision and recall of the PoseTrack′17 validation set. To eliminate redundant candidates, we discarded the candidates with lower probabilities. As Figure 9 (Table 2) shows, the precision and recall of the detector for various discard thresholds are given. Since the FPN network performed better, we selected the FPN network as our candidate detector. During training, we inferred the ground truth bounding boxes of the candidates from the annotated key points because the bounding box positions were not provided in the annotations in the PoseTrack′17 dataset. Specifically, we located the bounding box from the minimum and maximum coordinates of 15 key points and then enlarged the box by 20% horizontally and vertically.

[0166] For the single-person human pose estimator, we adopt the slightly modified CPN101 and MSRA152. We first train the network for 260 epochs using the fused dataset of Pose-Track′17 and COCO. Then we fine-tune the network on PoseTrack′17 only for 40 epochs to mitigate the inaccurate regression of the head and neck. For COCO, the bottom head and top head positions are not given. We infer these key points by interpolating the annotated key points. We find that the prediction of the head key points will be improved by fine-tuning on the PoseTrack dataset. During fine-tuning, we use the online hard key point mining technique and only focus on the losses of the 7 most difficult key points among the 15 key points. Pose inference is performed online using a single thread.

[0167] For the pose matching module, we train a Siamese graph convolutional network with 2 GCN layers and 1 convolutional layer using contrastive loss. We take the normalized key point coordinates as the input; the output is a 128-dimensional feature vector. Referring to the method of Yan (Yan, S et al., Spatial temporal graph convolutional networks for skeleton-based action recognition, AAAI, 2018), we use spatial configuration partitioning as the sampling method for graph convolution and use learnable edge importance weighting. To train the Siamese network, we generate training data from the Pose-Track dataset. Specifically, we extract the people with the same ID within adjacent frames as the positive pairs and the people with different IDs within the same frame and across frames as the negative pairs. The hard negative pairs only include poses with spatial overlap. The number of pairs collected is shown in Figure 8 (Table 1). We use the SGD optimizer to train the model with a batch size of 32 for a total of 200 epochs. The initial learning rate is set to 0∶001 and decays by 0∶1 at epochs 40, 60, 80, 100. The weight decay is 10 -4 .

[0168] Experiment 4: Ablation study. We conduct a series of ablation studies to analyze the contribution of each component to the overall performance.

[0169] Figure 9 , Table 2 shows the comparison of the precision-recall of the detector on the PoseTrack 2017 validation set. A bounding box is correct if the IoU of the bounding box with GT is higher than a certain threshold (set to 0.4 for all experiments). Figure 10, Table 3 shows a comparison of the offline pose tracking results using various detectors on the PoseTrack′17 validation set. Figure 11 , Table 4 shows a comparison of the offline and online pose tracking results with various keyframe intervals on the PoseTrack′18 validation set. Figure 12 , Table 5 shows a performance comparison of lightweight tracking using GCN and SC on the PoseTrack′18 validation set.

[0170] Detectors: We experimented with several detectors and decided to use a deformable convolutional network with ResNet101 as the backbone, a Feature Pyramid Network (FPN) for feature extraction, and a Fast R-CNN scheme as the detection head. As shown in Table 2, the performance of this detector is better than that of deformable R-FCN with the same backbone. Undoubtedly, a better detector will produce better performance in pose estimation and pose tracking, as Figure 10 , shown in Table 3.

[0171] Offline vs. Online: We investigated the effect of the keyframe interval of the online method and compared it with the offline method. For a fair comparison, we used the same candidate detector and pose estimator for both methods. For the offline method, we pre-computed candidate detections and estimated the pose of each candidate, and then we employed a flow-based pose tracker, where the pose flow is constructed by associating poses indicating the same person across frames. For the online method, we performed true online pose tracking. Since candidate detections are only made at keyframes, the online performance varies with the time interval. In Table 4, we illustrate the performance of the offline method, compared with the online method given various keyframe intervals. The offline method performs better than the online method. However, when the detections (DET) at keyframes are more accurate, we can see the great potential of the online method, whose upper limit is achieved by ground truth (GT) detections. As expected, frequent keyframes help improve performance. Note that the online method only uses spatial consistency for data association at keyframes. We report ablation experiments of the pose matching module below. GCN vs. Spatial Consistency (SC): Next, we report the results when performing pose matching during the data association phase, compared with using only spatial consistency. As can be seen from Table 5, the tracking performance improves with GCN-based pose matching. However, in some cases, different people may have almost identical poses. To mitigate this ambiguity, spatial consistency is considered before pose similarity.

[0172] GCN vs. Euclidean Distance (ED): We investigated whether the GCN network outperforms a simple pose matching scheme. By performing the same normalization on the key points, ED, as a dissimilarity metric for pose matching, achieved 85% accuracy on the validation pairs generated from the PoseTrack dataset, while GCN achieved 92% accuracy. We verified the positive pairs and the difficult negative pairs.

[0173] Experiment: 5. Performance Comparison. Figure 13 , Table 6 shows the performance comparison on the PoseTrack dataset. The last column shows the speed in frames per second (* indicates excluding the pose inference time). For our online method, mAP is provided after key point dropping. For our offline method, mAP is provided both before (left) and after (right) key point dropping.

[0174] Since the PoseTrack′18 test set has not been released, we compared our method with other online and offline methods on the PoseTrack′17 test set. For a fair comparison, we only used the PoseTrack′17 training set and the training and validation sets of COCO to train the pose estimator. Instead of using auxiliary data, we performed an ablation study on the validation set using CPN-101 as the pose estimator. During the testing process, in addition to CPN-101, we also conducted experiments using MSRA-152.

[0175] Accuracy: As shown in Table 6, our LightTrack method outperforms other online methods while maintaining a higher frame rate and is very competitive with the offline state-of-the-art methods. For our offline method, we used the same detector and pose estimator as LightTrack, except that we replaced LightTrack with the official version of PoseFlow for performance comparison. Although the PoseFlow algorithm is conceptually online, the processing is performed in multiple stages and requires pre-computing the key point matches between frames, which is computationally expensive. In contrast, our LightTrack is truly online processing.

[0176] Speed: Tested on a single Tesla P40 GPU, the pose matching for each pair takes an average of 2.9 milliseconds. Since pose matching only occurs at key frames, its occurrence frequency depends on the number of candidates and the length of the key frame interval. Therefore, we tested the average processing time on the PoseTrack'18 validation set, which consists of 74 videos and a total of 8,857 frames. The online algorithm CPN101-LightTrack took 11,638 seconds to process, of which 11,450 seconds were used for pose estimation. The frame rate of the entire system was 0.76 fps. The framework runs at approximately 47.11 fps, excluding the pose inference time. A total of 57,928 people were encountered. On average, 6.54 people were tracked per frame. CPN101 takes 140 milliseconds to process each candidate, including 109 milliseconds for pose inference and 31 milliseconds for preprocessing and postprocessing. With other options optimized by the pose estimator and parallel inference, there is potential room for improvement in the actual frame rate and tracking performance. We saw an improvement in the performance of MSRA152-LightTrack, but the frame rate was slightly slower due to its 133-millisecond inference time.

[0177] Experiments: 6. Discussion.

[0178] Accuracy: Since the components in our framework are easy to replace and expand, the methods using this framework can become faster, more accurate, or both. It should be noted that the pose estimators mentioned in Section 3 can be replaced with more accurate or faster counterparts. The improved performance in general object detectors or methods focusing on detecting people (e.g., using auxiliary datasets) should also improve the pose tracking performance. The ablation study in Section 4 shows that better detection improves the MOTA score regardless of which detector is used.

[0179] Speed: The pose estimation network can prioritize speed while sacrificing some accuracy. For example, we used YOLOv3 and MobileNetv1-deconv (YoloMD) as the detector and pose estimator respectively. This achieved an average of 2 FPS on the PoseTrack'18 validation set with a 70.4 mAP and a MOTA score of 55.7%. In addition to network structure design, faster networks can also refine the heatmap from previous frames. Recently, refinement-based networks have received great attention.

[0180] Flexibility: The advantage of our top-down approach in pose tracking is that we can conveniently track specific targets without having to track all candidates. This can be simply achieved by selecting the target in the first frame and providing the target location at key frames. As a side effect, this further reduces the computational complexity. If the target has a specific visual appearance, the framework can be conveniently extended to ensure that only the target is matched at key frames and tracked at the remaining frames.

[0181] In summary, the present disclosure provides an effective and general lightweight framework for online human pose tracking. The present disclosure also provides a baseline adopting this framework and provides a siamese graph convolutional network for human pose matching as the Re-ID module in the pose tracking system. The skeleton-based representation effectively captures human pose similarity and has a low computational cost. The method of the present disclosure significantly outperforms other online methods and is very competitive compared with the state-of-the-art in the offline state, with a higher frame rate.

[0182] The foregoing description of the exemplary embodiments of the present disclosure is presented only for purposes of illustration and description and is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings.

[0183] The embodiments are chosen and described in order to explain the principles of the present disclosure and its practical application, thereby enabling others skilled in the art to utilize the present disclosure and various embodiments and various modifications suitable for the particular purposes contemplated. Alternative embodiments will become apparent to those skilled in the art of the present disclosure without departing from the spirit and scope of the present disclosure. Accordingly, the scope of the present disclosure is defined by the appended claims rather than the foregoing description and the exemplary embodiments described therein.

[0184] References (incorporated herein by reference in their entirety):

[0185] 1. LAW, H and DENG, J, CornerNet: Detecting objects as paired keypoints, 2019, arXiv:1808.01244.

[0186] 2. XIAO, B; WU, H and WEI, Y, Simple baselines for human pose estimation and tracking, 2018, arXiv:1804.06208.

[0187] 3. YAN, S; XIONG, Y; and LIN D, Spatial temporal graph convolutional networks for skeleton-based action recognition, 2018, arXiv:1801.07455.

[0188] 4. ZHU, J; YANG, H, et al., Online multi-object tracking with dual matching attention networks, 2019, arXiv: 1902.00749.

Claims

1. A system for attitude tracking, comprising: A computing device, including a processor and a storage device storing computer-executable code, wherein the computer-executable code, when executed at the processor, is configured to: Provide multiple consecutive frames of a video, the consecutive frames including at least one key frame and multiple non-key frames; For each non-key frame of the multiple non-key frames: Receive a previously inferred bounding box of an object inferred from a previous frame; In the region defined by the previously inferred bounding box, estimate key points from the non-key frame to obtain estimated key points; Based on the estimated key points, determine an object state, wherein the object state includes a "tracked" state and a "lost" state; When the object state is "tracked", based on the estimated key points, infer an inferred bounding box to process the next frame of the non-key frame; wherein the bounding box is determined based on the topmost, bottommost, leftmost, and rightmost key points and the bounding box is enlarged to determine the inferred bounding box; When the object state is "lost": Detect an object from the non-key frame, wherein each detected object is defined by a detected bounding box; Estimate key points of each detected object from the corresponding bounding box in the detected bounding box to obtain detected key points; Identify each detected object by comparing the detected key points of the detected object with the stored key points of a stored object, each stored object having an object identification ID; When the detected key points match the stored key points of the corresponding object from the stored objects, assign the object ID of the corresponding object in the stored objects to the detected object.

2. The system according to claim 1, wherein The computer-executable code is configured to estimate key points from the non-key frame using a convolutional neural network.

3. The system according to claim 1 or 2, wherein, When the estimated key points have an average confidence greater than a threshold score, the object state is "tracked", and when the estimated key points have an average confidence less than or equal to the threshold score, the object state is "lost".

4. The system according to claim 1, wherein The computer-executable code is configured to infer the inferred bounding box by: Defining a bounding box enclosing the estimated key points; and Enlarging the bounding box by 20% in the horizontal and vertical directions of the bounding box respectively.

5. The system according to claim 1, wherein The computer-executable code is configured to detect an object using a convolutional neural network.

6. The system according to claim 1, wherein, The computer-executable code is configured to use a convolutional neural network to estimate the key points.

7. The system according to claim 1, wherein The step of comparing the detected key points of the detected object with the stored key points is performed using a Siamese Graph Convolutional Network SGCN, the SGCN including two Graph Convolutional Networks GCN with shared network weights, each GCN including: A first Graph Convolutional Network GCN layer; A first Relu unit, connected to the first GCN layer; A second GCN layer, connected to the first Relu unit; A second Relu unit, connected to the second GCN layer; An average pooling layer, connected to the second GCN layer; Fully Connected Network FCN; and a feature vector conversion layer, wherein the first GCN layer is configured to receive the detected key points of one of the detected objects, and the feature vector conversion layer is configured to generate a feature vector representing the pose of one of the detected objects.

8. The system according to claim 7, wherein The SGCN is configured to perform the comparison step by: running the estimated key points through one of the two GCNs to obtain an estimated feature vector for the estimated key points; running the stored key points of one of the stored objects through the other of the two GCNs to obtain a feature vector for the stored key points; and when the distance between the estimated feature vector and the stored feature vector is less than a predetermined threshold, determining that the estimated key points match the stored key points.

9. The system according to claim 1, wherein, For each non-key frame among the non-key frames, the computer-executable code is configured to, when the estimated key points of the object do not match the key points of any of the stored objects: assign a new object ID to the object.

10. The system according to claim 1, wherein, For each key frame among the key frames, the computer-executable code is configured to: detect objects in the key frame, where each detected object is defined by a bounding box; estimate a plurality of detected key points of each detected object from the bounding box corresponding to the detected object; identify each detected object by comparing the detected key points of the detected object with the stored key points of the stored objects, each of the stored objects having an object identification ID; and when the detected key points match the stored key points of the corresponding object from the stored objects, assign the object ID of the corresponding object in the stored objects to the detected object.

11. The system according to claim 10, wherein, For each of the key frames, the computer-executable code is configured to, when the estimated key points of the object do not match the key points of any of the stored objects: assign a new object ID to the object.

12. A method for pose tracking, comprising: providing a plurality of consecutive frames of a video, the consecutive frames including at least one key frame and a plurality of non-key frames; for each non-key frame among the plurality of non-key frames: receiving a previous inferred bounding box of an object inferred from a previous frame; estimating key points from the non-key frame in a region defined by the previous inferred bounding box to obtain estimated key points; determining an object state based on the estimated key points, where the object state includes a "tracked" state and a "lost" state; when the object state is "tracked", inferring an inferred bounding box based on the estimated key points to process the next frame of the non-key frame; wherein the bounding box is determined based on the topmost, lowest, leftmost, and rightmost key points and the bounding box is enlarged to determine the inferred bounding box; when the object state is "lost": detecting objects from the non-key frame, where each detected object is defined by a detected bounding box; Estimate the key points of each detected object from the corresponding bounding boxes in the detected bounding boxes to obtain the detected key points; Identify each of the detected objects by comparing the detected key points of the detected objects with the stored key points of each stored object, each of the stored objects having an object identification ID; When the detected key points match the stored key points from one of the stored objects, assign the object ID of one of the stored objects to the detected object.

13. The method according to claim 12, wherein, When the estimated key points have an average confidence greater than a threshold score, the object status is "tracked", and when the estimated key points have a confidence less than or equal to the threshold score, the object status is "lost".

14. The method according to claim 12 or 13, wherein The step of the inference inferring the bounding box includes: Defining a bounding box surrounding the estimated key points; and Enlarging the bounding box by 20% in the horizontal and vertical directions of the bounding box respectively.

15. The method according to claim 12, wherein The steps of detecting the object and estimating the key points are both performed using a convolutional neural network CNN.

16. The method according to claim 12, wherein, The step of comparing the detected key points of the detected object with the stored key points is performed using a Siamese graph convolutional network SGCN, the SGCN including two graph convolutional networks GCN with shared network weights, each GCN including: A first graph convolutional network GCN layer; A first Relu unit connected to the first GCN layer; A second GCN layer connected to the first Relu unit; A second Relu unit connected to the second GCN layer; An average pooling layer connected to the second GCN layer; A fully connected network FCN; and A feature vector conversion layer, wherein the first GCN layer is configured to receive the detected key points of one of the detected objects, and the feature vector conversion layer is configured to generate a feature vector representing the pose of one of the detected objects.

17. The method according to claim 12, further comprising, for each key frame in the key frames: Detect an object from the key frames, wherein, Each detected object is defined by a detected bounding box; Estimate the detected key points from each of the detected bounding boxes to obtain the detected key points; Identify each of the detected objects by comparing the detected key points of the detected objects with the stored key points of the stored objects, each of the stored objects having an object identification ID; And When the detected key points match the stored key points of the corresponding object among the stored objects, assign the object ID of the corresponding object among the stored objects to the detected object.

18. A non-transitory computer-readable medium storing computer-executable code, wherein, When the computer-executable code is executed at a processor of a computing device, it is configured to implement the method for pose tracking according to any one of claims 12 to 17.

Citation Information

Patent Citations

  • Systems and methods for object tracking

    US20160267325A1

  • Generating labeled data for deep object tracking

    US20190050693A1

  • Object tracking for neural network systems

    US20190114804A1