Position estimation based on sensor data
The system effectively addresses the challenges of tracking deforming objects by using depth camera and IMU data to dynamically adjust the tracking area's pose and size, ensuring accurate and stable object tracking.
Patent Information
- Application Number
- PCT/CN2023/134964
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-05
AI Technical Summary
Existing object tracking techniques face challenges in accurately tracking objects that undergo deformation, such as changes in size and pose, and are unstable when the object moves rapidly or has complex depth profiles.
The system employs a method for position estimation of target objects based on sensor data, including depth camera images and IMU data. It determines a tracking area in the image, adjusts its pose and size based on the device's pose and distance to the target, and uses depth maps to refine the tracking area's dimensions.
This approach enables accurate and stable tracking of objects across frames, even when the object undergoes deformation or moves rapidly, by dynamically adjusting the tracking area's pose and size based on real-time sensor data.
Smart Images

Figure CN2023134964_05062025_PF_FP_ABST
Abstract
Description
POSITION ESTIMATION BASED ON SENSOR DATAFIELD
[0001] The present disclosure generally relates to image processing. For example, aspects of the present disclosure relate to systems and techniques for performing position estimation of an object based on sensor data, such as depth camera images and inertial measurement unit (IMU) data.BACKGROUND
[0002] Many devices and systems allow a scene to be captured by generating frames and / or video data (including multiple frames) of the scene. For example, a camera or a computing device including a camera can capture a sequence of frames of a scene. The image and / or video data can be captured and processed by such devices and systems and can be output for consumption (e.g., displayed on the device and / or other device) . In some cases, the image and / or video data can be captured by such devices and systems and output for processing and / or consumption by other devices.
[0003] A frame or image can be processed to determine objects that are present in the frame, which can be useful for many applications. For instance, a model can be determined for representing an object in a frame, and can be used to facilitate effective operation of various systems. Examples of such applications and systems include robotics, automotive and aviation applications, object tracking, in addition to many other applications and systems.SUMMARY
[0004] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary has the sole purpose to present certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.
[0005] Disclosed are systems and techniques for performing position estimation of a target object based on sensor data. According to some aspects, a method is provided for object tracking at a computing device. The method includes: obtaining an image of a scene including an object; determining a tracking area in the image to represent the object; determining a point of the tracking area in the image is a target point; determining a distance between the computing device and the target point; and adjusting a pose and a size of the tracking area in the image based on a pose of the computing device and the distance between the computing device and the target point.
[0006] In some aspects, an apparatus for object tracking is provided. The apparatus includes at least one memory and at least one processor coupled to the at least one memory and configured to: obtain an image of a scene including an object; determine a tracking area in the image to represent the object; determine a point of the tracking area in the image is a target point; determine a distance between the apparatus and the target point; and adjust a pose and a size of the tracking area in the image based on a pose of the apparatus and the distance between the apparatus and the target point.
[0007] In some aspects, computer-readable memory is provided storing instructions which, when executed by a computing device, cause the computing device to: obtain an image of a scene including an object; determine a tracking area in the image to represent the object; determine a point of the tracking area in the image is a target point; determine a distance between the apparatus and the target point; and adjust a pose and a size of the tracking area in the image based on a pose of the apparatus and the distance between the apparatus and the target point.
[0008] In some aspects, an apparatus for object tracking is provided. The apparatus includes: means for obtaining an image of a scene including an object; means for determining a tracking area in the image to represent the object; means for determining a point of the tracking area in the image is a target point; means for determining a distance between the apparatus and the target point; and means for adjusting a pose and a size of the tracking area in the image based on a pose of the apparatus and the distance between the apparatus and the target point.
[0009] Aspects generally include a method, apparatus, system, computer program product, non-transitory computer-readable medium, user device, user equipment, wireless communication device, and / or processing system as substantially described with reference to and as illustrated by the drawings and specification.
[0010] In some aspects, one or more of the apparatuses described herein is, can be part of, or can include a mobile device (e.g., a mobile telephone or so-called “smart phone” , a tablet computer, or other type of mobile device) , an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device) , a vehicle (or a computing device, component, or system of a vehicle) , a smart or connected device (e.g., an Internet-of-Things (IoT) device) , a wearable device, a personal computer, a laptop computer, a video server, a television (e.g., a network-connected television) , a robotics device or system, or other device. In some aspects, each apparatus can include an image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each apparatus can include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each apparatus can include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each apparatus can include one or more sensors. In some cases, the one or more sensors can be used for determining a location of the apparatuses, a state of the apparatuses (e.g., a tracking state, an operating state, a temperature, a humidity level, and / or other state) , and / or for other purposes.
[0011] Some aspects include a device having a processor configured to perform one or more operations of any of the methods summarized above. Further aspects include processing devices for use in a device configured with processor-executable instructions to perform operations of any of the methods summarized above. Further aspects include a non-transitory processor-readable storage medium having stored thereon processor-executable instructions configured to cause a processor of a device to perform operations of any of the methods summarized above. Further aspects include a device having means for performing functions of any of the methods summarized above.
[0012] The foregoing has outlined rather broadly the features and technical advantages of examples according to the disclosure in order that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The conception and specific examples disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. Characteristics of the concepts disclosed herein, both their organization and method of operation, together with associated advantages will be better understood from the following description when considered in connection with the accompanying figures. Each of the figures is provided for the purposes of illustration and description, and not as a definition of the limits of the claims. The foregoing, together with other features and aspects, will become more apparent upon referring to the following specification, claims, and accompanying drawings.
[0013] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Illustrative aspects of the present application are described in detail below with reference to the following figures:
[0015] FIG. 1 is a graph illustrating an example of a camera coordinate system, in accordance with some aspects of the present disclosure.
[0016] FIG. 2 is a diagram illustrating an example of the principle of binocular parallax with two monocular cameras, in accordance with some aspects of the present disclosure.
[0017] FIG. 3 is a diagram illustrating an example of a camera coordinate system for a binocular camera, in accordance with some aspects of the present disclosure.
[0018] FIG. 4 is a diagram illustrating an example of a process for tracking an object using correlation filtering, in accordance with some aspects of the present disclosure.
[0019] FIG. 5 is a diagram illustrating an example of tracking an object using distance and angle estimation, in accordance with some aspects of the present disclosure.
[0020] FIG. 6 is a diagram illustrating an example for determining a height of a tracked object in an image, in accordance with some aspects of the present disclosure.
[0021] FIG. 7 is a flow chart illustrating an example of a process for human following based on binocular depth camera and IMU data, in accordance with some aspects of the present disclosure.
[0022] FIG. 8 is a diagram illustrating an example of forming a tracking area in the form of a bounding box, in accordance with some aspects of the present disclosure.
[0023] FIG. 9 is a diagram illustrating an example of a pinhole model of a camera, in accordance with some aspects of the present disclosure.
[0024] FIG. 10 is a graph illustrating an example of a pose of a camera, in accordance with some aspects of the present disclosure.
[0025] FIG. 11 are graphs illustrating an example of eliminating errors in a pose of a camera, in accordance with some aspects of the present disclosure.
[0026] FIG. 12 is a diagram illustrating an example of adjusting (e.g., to correct) a size of a tracking area for a human, in accordance with some aspects of the present disclosure.
[0027] FIG. 13 is a diagram illustrating an example of further adjusting (e.g., to correct) a size of a tracking area for a human by using a depth map, in accordance with some aspects of the present disclosure.
[0028] FIG. 14 is a diagram illustrating an example of further adjusting (e.g., to correct) a size of a tracking area for a human, in accordance with some aspects of the present disclosure.
[0029] FIG. 15 is a flow diagram illustrating an example of a process for human following based on binocular depth camera and IMU data, in accordance with some aspects of the present disclosure.
[0030] FIG. 16 is a diagram illustrating an example of a system for implementing certain aspects described herein.DETAILED DESCRIPTION
[0031] Certain aspects of this disclosure are provided below for illustration purposes. Alternate aspects may be devised without departing from the scope of the disclosure. Additionally, well-known elements of the disclosure will not be described in detail or will be omitted so as not to obscure the relevant details of the disclosure. Some of the aspects described herein can be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive.
[0032] The ensuing description provides example aspects only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the example aspects will provide those skilled in the art with an enabling description for implementing an example aspect. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims.
[0033] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration. ” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term “aspects of the disclosure” does not require that all aspects of the disclosure include the discussed feature, advantage or mode of operation.
[0034] As noted previously, a system can process frames of data (e.g., by performing object detection, recognition, segmentation, or other operation) to determine one or more objects (e.g., a human, a vehicle, or other object) that are present in the images. As used herein, the term frame can refer to a frame of sensor data, such as an image frame (also referred as an image) , a frame of radar or light-detection and ranging (LIDAR) data, depth frames including depth data, any combination thereof, and / or other type of sensor data.
[0035] Processing frames to determine objects in the frames can be useful for many applications. For instance, a model can be determined for representing an object in a frame, and can be used to facilitate effective operation of various systems. Examples of such applications and systems include robotics (e.g., human-following robots or other types of robots) , automotive and aviation applications (e.g., for autonomous or semi-autonomous driving or navigation systems) , object tracking, and other applications and systems. In some cases, multiple frames including detected objects can be processed to track the objects across the frames.
[0036] In one illustrative example, robots programmed to follow humans (referred to as human-following robots) have many applications in everyday life, such as applications related to home use, manufacturing, commerce, etc. For instance, a human-following robot may be employed within an airport to follow a passenger around the airport, such as to carry heavy luggage for the passenger. In another example, a human-following robot can be employed in a factory to follow an employee around the factory, such as to carry large, heavy equipment for the employee. When a human-following robot is following a human, it can be necessary to perform human tracking (e.g., across multiple frames captured by the robot, such as images, depth information, etc. ) , positioning of the robot, and motion control of the robot, so that the human-following robot can effectively follow the human.
[0037] In some cases, stereo vision may be used for determining depth of points in space, which can be useful in performing tracking of objects (e.g., humans, vehicles, or other objects) . For instance, to perform stereo vision, a stereo camera in the form of two monocular cameras or in the form of a single binocular camera may be employed. The stereo camera may be mounted onto a device or system, such as a robot, a vehicle, or other object. In one illustrative example, for tracking of a human, a robot with a stereo camera can be moved such that an angle between a projection line (e.g., extending from the stereo camera to a projection point located on the human) and a horizontal axis goes to zero. When this angle reaches zero, the distance from the stereo camera to the target becomes a fixed value (e.g., 0.9 meters or other value) . When using stereo vision (e.g., by either employing two monocular cameras or one binocular camera for the stereo camera) , the depth (distance) of a point located on the human (from the stereo camera) can be calculated.
[0038] Various techniques exist for performing object tracking. However, there are disadvantages associated with existing object tracking techniques. For example, one existing technique for object tracking includes a correlation filter-based tracking method, which may employ a Kernelized Correlation Filter (KCF) . Correlation filter-based tracking can achieve a good balance between accuracy and real-time performance, but performs poorly in scenarios where the tracked object (e.g., a human) undergoes a deformation (e.g., the size and / or pose of the human changes (is deformed) as the human walks) .
[0039] Another example of an object tracking technique includes calculating depths of all points imaged in a rectangle used to represent the target (e.g., a human, a vehicle, a building, etc. ) . In some cases, point with a minimum value from the points is chosen as the target point. However, such a technique can be unstable when the tracked object moves, such as when the human is swinging his hands while walking. In other cases, points whose depths exceed a set threshold are first eliminated. The median value of the depths of the remaining points is then calculated, and the point with the median value is chosen as the target point. The background area of the scene is then filtered by setting a fixed depth threshold. However, using such a fixed depth threshold restricts the algorithm to being able to only perform satisfactorily for some scenes and, as such, cannot operate universally for all types of scenes.
[0040] Another example of an object tracking technique is a common distance estimation method that directly uses the absolute distance from a target point to the camera given by a depth value. However, some devices or systems with sensors (e.g., wheeled robots) can only move on a two-dimensional (2D) horizontal ground, while tracked objects (e.g., the human body) may be three-dimensional (3D) . The common distance estimation method must thus project the tracked object (e.g., the 3D human body) onto a horizontal plane for the positioning of the object.
[0041] Another example technique for object tracking includes using a pool of bounding boxes. For example, a system utilizing such a technique can preselect bounding boxes of different sizes from the pool of bounding boxes to match to a deformed target object. However, this technique can result in a significant decrease in computational speed, and the bounding box pool may not cover all sizes for the target effectively. As such, an improved technique for tracking objects can be beneficial.
[0042] Systems, apparatuses, methods (also referred to as processes) , and computer-readable media (collectively referred to herein as “systems and techniques” ) are described herein for performing position estimation of target objects based on sensor data, which can be used for tracking the target objects. In some aspects, the sensor data includes depth camera images (e.g., images from a binocular depth camera) and inertial measurement unit (IMU) data. The systems and techniques can provide solutions for various types of systems, such as human-following robots, vehicles, extended reality (XR) systems, among others.
[0043] In some aspects, the systems and techniques employ the use of a tracking area (e.g., in the form of a bounding region, such as a bounding box) to represent an object (e.g., a human, a vehicle, etc. ) being tracked (e.g., across frames, such as image frames) within a scene. Positioning is an important part of object tracking, where the accuracy of the positioning of the human will significantly affect the tracking results. In one illustrative example, human positioning is an indispensable aspect for the tracking of a human. For instance, as a human walks and the pose of the human changes, the size and pose of a bounding box used for tracking the human across frames may no longer accurately represent the human. The systems and techniques can use sensor data (e.g., images from a binocular depth camera and IMU data) to correct for the errors in the bounding box size and pose to be able to accurately estimate the position of the object (e.g., a target object, such as an object being tracked by a tracking device, such as a human-following robot, a vehicle, etc. ) .
[0044] In one illustrative example, the systems and techniques be used by a device to track an object, such as a human, a vehicle, or other object. During operation, the device can obtain one or more images (e.g., a 2D image from a binocular depth camera) of a scene including the object. After the device obtains the an image of the scene, one or more processors (e.g., implemented within the device and / or located remote from the device) can determine a tracking area (e.g., a bounding region, such as a bounding box) in the image to represent the human. In some cases, the tracking region can be a bounding box that includes four corners, where each respective corner of the bounding box represents certain portions of the object. In one illustrative example, a top-left corner of the bounding box represents a left shoulder of a human, a top-right corner of the bounding box represents a right shoulder of the human, a bottom-left corner of the bounding box represents a left hip of the human, and a bottom-right corner of the bounding box represents a right hip of the human.
[0045] After the one or more processors determine the tracking area, the one or more processors can determine a center of the tracking area (e.g., the bounding box) in the image is a target point. The one or more processors can obtain the coordinates of the target point in the coordinate system of the device (e.g., camera coordinate system) . After the one or more processors determine the target point and the coordinates of the target point in the coordinate system of the device, the one or more processors can then determine a distance between the device (e.g., a binocular depth camera) and the target point.
[0046] The one or more processors can then adjust a pose and a size of the tracking area in the image based on a pose of the device (e.g., obtained from IMU data of the device) and the determined distance between the device and the target point. In one or more examples, the one or more processors may further adjust the size of the tracking area based on a depth map of the scene (e.g., based on an amount of discontinuity between the depth of the human and the depth of background of the scene in the depth map) .
[0047] While some examples provided herein describe certain devices or systems tracking certain objects (e.g., a human-following robot tracking a human, a vehicle tracking another vehicle, a human, or other object, etc. ) for illustrative purposes, the systems and techniques are not limited thereto. For example, the systems and techniques described herein can be used by any type of device or system for tracking any type of object.
[0048] Additional aspects of the present disclosure are described in more detail below.
[0049] target FIG. 1 shows an example, represented in a camera coordinate system, of an object (or target) being tracked by a device or system (e.g., a human being followed by a robot, a vehicle or human being tracked by another vehicle, etc. ) with a camera mounted onto the device or system. In particular, FIG. 1 is a graph 100 illustrating an example of a camera coordinate system including X, Y, and Z axes. In FIG. 1, a camera (e.g., a stereo camera) is located at point C, and may be mounted onto a robot. The object (e.g., the human, vehicle being tracked, etc. ) is represented by target point P1 and projection point P2. The height of the object (e.g., human, vehicle being tracked, etc. ) is the distance between points P1 and P2. It can be assumed that the optical axis of the camera is horizontal (e.g., the X-axis) . A target line L1 is shown to extend from the camera at point C to the target point P1 on the object, and a projection line L2 is shown to extend from the camera at point C to the projection point P2 on the object.
[0050] For the tracking of the object (e.g., human, vehicle being tracked, etc. ) , the device or system (e.g., human-following robot, vehicle performing the tracking, etc. ) is moved such that an angle between the projection line L2 (e.g., extending from the camera at point C to the projection point P2 located on the object) and the horizontal axis (e.g., the X-axis) goes to zero (0) . When the angle reaches zero, the distance from the camera to the object (e.g., human) becomes a fixed value (e.g., 0.9 meters) .
[0051] When using stereo vision (e.g., by either employing two monocular cameras or one binocular camera for the stereo camera) , the depth (distance) of a point located on the object (e.g., from the stereo camera) can be calculated. FIG. 2 shows an example of calculating the depth r of a point P located on an object when using two monocular cameras. In particular, FIG. 2 is a diagram illustrating an example 200 of the principle of binocular parallax with two monocular cameras (e.g., camera 1 210a and camera 2 210b) . According to the principle of binocular parallax, the depth of a world point P can be calculated by:
[0052] where f is the focal length, b is the baseline distance between the two monocular cameras 210a and 210b, xl is the x coordinate of the image point Pl on the left, and xr is the x coordinate of the image point Pr on the right.
[0053] FIG. 3 shows an example of calculating the depth r of a point P located on an object (e.g., a human or other object or target) when using a single binocular camera. In particular, FIG. 3 is a diagram illustrating an example of a camera coordinate system 300 (including XC, YC, and ZC axes) for a binocular camera. In the binocular camera, the left camera may be chosen as the main camera, and the center of the lens of the left camera may be chosen as the origin of the coordinate system. The optical axis is the XC axis. The line where the centers of each of the two lenses are located is the YC axis. The ZC axis is perpendicular to the XC-YC plane. The red, green, blue (RGB) image captured by the left camera may be used as the source for processing. Since b << r, then r ≈ d (e.g., depth from point P to the center of the left lens) .
[0054] As noted previously, various techniques can be utilized for object tracking. One example of a technique for object tracking is a correlation filter-based tracking method, which may employ a Kernelized Correlation Filter (KCF) . FIG. 4 shows an example of a correlation filter-based tracking method. In particular, FIG. 4 is a diagram illustrating an example of a process 400 for tracking an object (e.g., shown as a human in FIG. 4) using correlation filtering. In FIG. 4, an initial image 410a (e.g., initial frame) of a scene including the human is obtained by a stereo camera. A tracking area 412 (e.g., in the form of a bounding box) for the human is then determined in the initial image 410a. The initial image 410a is cropped 415 according to the tracking area 412 to produce a subframe 420. Features are then extracted from the subframe 420 to obtain features 430a. The extracted features 430a are then used to train 435a a correlation filter (e.g., a KCF) .
[0055] Preselection 445 can be performed on a next (e.g., subsequent) image frame 410b (e.g., next frame) to determine a plurality of preselection boxes 450 (e.g., multiple bounding boxes) for the next image frame 410b. Features are then extracted to obtain features 430b. The trained correlation filter can perform a corresponding graph calculation 455 to generate a response map 460. The region of the response map with the highest response value (e.g., max value) is then selected 465 to produce an output, which is then used to update the filter 470.
[0056] First, features are extracted from the initial frame to train the correlation filter. Then, in the next frame, some candidate regions are selected, and the trained correlation filter is used to generate a response map. The region with the highest response value is identified as the output.
[0057] However, as described above, the correlation filter-based tracking method cannot handle scenarios where the tracked object (e.g., a human) undergoes a deformation (e.g., the size and / or pose of the human changes (is deformed) as the human walks) .
[0058] As previously noted, another example of an object tracking technique is a common distance estimation tracking method. FIG. 5 shows an example of a common distance estimation tracking method. In particular, FIG. 5 is a diagram illustrating an example 500 of tracking an object using distance and angle estimation. In FIG. 5, a robot 510 is shown to be following a human 520. A camera 570 is shown to be mounted onto a robot 510. The camera 570 is located at point C. A target point P1 530 is selected on the body of the human 520. The target point P1 530 is projected onto the horizontal plane, where the camera 570 on the robot 510 lies. A target line L1 550 extends from the camera 570 at point C to a target point P1 530 on the human 520. A projection line L2 560 extends from the camera 570 at point C to a projection point P2 540 on the human 520.
[0059] The length of the projection line L2 560 is dH. The distance from the target point P1 530 to the camera 570 at point C is d, which can also be referred to as the depth. The angle between the direction of the robot 510 and the projection line L2 560 is the angle δ. The common distance estimation tracking method uses dH and the angle δ to determine the position of the human 520.
[0060] Other existing techniques for object tracking also have deficiencies, as described above.
[0061] As noted previously, systems and techniques are described herein for performing position estimation of target objects based on sensor data, such as depth camera images (e.g., images from a binocular depth camera) and inertial measurement unit (IMU) data. The systems and techniques can be used for tracking the target objects and can provide solutions for various types of systems, such as human-following robots, vehicles, extended reality (XR) systems, among others.
[0062] FIG. 6 shows an example of a method for human following based on binocular depth camera and IMU data. In particular, FIG. 6 is a flow chart illustrating an example of a process 600 for human following based on binocular depth camera and IMU data. During operation of the process 600 of FIG. 6, at block 610, the process 600 starts. At block 620, one or more image sensors of a device (e.g., a binocular depth camera) can obtain an image (e.g., an RGB image) of a scene including one or more humans. Also at block 620, one or more IMU sensors of the device can obtain IMU data (e.g., including a pose of the device) associated with the device.
[0063] At decision block 630, one or more processors (e.g., which may be implemented within the device or may be located remote from the device) can determine whether a tracking area (e.g., a tracker in the form of a bounding box) for a human to be tracked within the scene has been initialized. When the one or more processors determine that a tracking area (e.g., bounding box) has not been initialized, the process 600 can proceed to block 640. At block 640, a tracking area (e.g., bounding box) for a human to be tracked can be initialized.
[0064] For the initialization of a tracking area of block 640, at block 642, the one or more processors can select a human (e.g., person) from the one or more humans in the scene to be tracked (e.g., by a robot) . At block 644, the one or more processors can detect (e.g., determine) a pose of the selected human to be tracked. At block 646, the one or more processors can locate the upper portion of the body of the selected human to be tracked. At block 648, the one or more processors can determine (e.g., initialize) a tracking area (e.g., a tracker in the form of a bounding box) of the upper portion of the body of the selected human to be tracked, where the tracking area can be used to represent the selected human. At block 648, the one or more processors can also determine a center of the tracking area to be a target point.
[0065] However, when the one or more processors determine that a tracking area (e.g., bounding box) has been initialized, the process 600 can proceed to block 650. At block 650, the one or more processors can use the following obtained image (e.g., a next image) for tracking. At block 660, the one or more processors can detect a human (e.g., person) in the following image, and select a target point for the detected human. At block 660, the one or more processors can calculate (e.g., determine) the pose (e.g., roll, pitch, and yaw) of the device (e.g., camera) based on the IMU data obtained in block 620. At block 680, the one or more processors can estimate a position of the detected human in the following image.
[0066] At block 690, the tracking area (e.g., tracker in the form of a bounding box) can be updated (e.g., adjusted) . For the updating (e.g., adjusting) of the tracking area of block 690, at block 692, one or more processors can detect a target (e.g., a human) in an image. At decision block 694, the one or more processors can determine whether an intersection over union (IoU) for the detected target is greater than a predetermined IoU threshold. Based on determining that the IoU is greater than the IoU threshold (e.g., a Yes decision at decision block 694) , the one or more processors can adjust the size (and / or pose) of the tracking area (e.g., bounding box) at block 696. However, based on a determination that the IoU is not greater than the IoU threshold (e.g., a No decision at decision block 694) , the one or more processors can determine that a tracking failure has not be initialized at block 698. After the tracking area (e.g., tracker) has been updated at block 690, the process 600 can proceed back to block 620.
[0067] FIG. 7 shows an example of a tracking area and a target point selected for a human to be tracked. In particular, FIG. 7 is a diagram illustrating an example 700 of forming a tracking area 730 in the form of a bounding box and determining a target point 740 for a human to be tracked. In FIG. 7, two images 710a and 710b are shown.
[0068] For image 710a, one or more processors (e.g., which may be implemented within a device, such as a binocular depth camera, or may be located remote from the device) can detect a human in the image 710a. In one or more examples, the one or more processors can run an algorithm (e.g., AI post net) to detect a pose of a human. The algorithm can generate a rough outline 720 for the detected pose of the human. Some of the points (e.g., nodes) within the rough outline 720 are shown to denote the left shoulder, right shoulder, left hip, and right hip of the human.
[0069] After the rough outline 720 of the pose of the human is obtained, as shown in image 710b, the one or more processors can generate a tracking area 730 based on the locations of the nodes (e.g., denoting the left shoulder, right shoulder, left hip, and right hip) of the rough outline 720 for the pose of the human. The tracking area 730 is in the form of a bounding box, where the four corners of the bounding box are located at the left shoulder, right shoulder, left hip, and right hip of the human. The one or more processors can then determine a center point (e.g., located at the cross section of two diagonal lines drawn within the bounding box) of the tracking area 730 to be the target point 740 for the human.
[0070] The torso area of the upper body of the human body is a more stable tracking target, than the whole body because of less deformation occurs for the torso area. The target point 740 is selected to be located at the center of the torso of the human body and, as such, this selection of the target point 740 avoids mistakenly selecting points in the background of the image to estimate the position of the human.
[0071] FIG. 8 shows an example of estimating, for any camera pose, the distance between a tracked human and a camera, and an angle between the projection line to the human and the horizontal axis of the camera coordinate system. In particular, FIG. 8 is a diagram illustrating an example of a pinhole model 800 of a camera (represented as “C” in FIG. 8) . FIG. 8 shows the calculation of 3D coordinates of a point (e.g., a target point 820 of a human 810) in the camera coordinate system.
[0072] In the camera coordinate system, the field of view (FOV) of the camera and the pixel coordinate of the image point (e.g., corresponding to the target point 820) can be used to convert the depth information into 3D coordinates in the camera coordinate system.
[0073] In the pinhole model 800 of FIG. 8, a section perpendicular to the X-Z plane of the camera can be used to give: XC=d·cosθ·cosφ (equation 4) YC=d·cosθ·sinφ (equation 5) ZC=d·sinθ (equation 6)
[0074] where f is the focal length, HI is the pixel height of the image, HP is the pixel height from the image point to the horizonal centerline of the image, v is the vertical FOV of the camera, φ is the angle shown in FIG. 1 and FIG. 8, θ is the angle shown in FIG. 1, and d is the distance of the target point 820 from the camera (e.g., which can be obtained from the binocular camera system) . The 3D coordinates of the target point 820 in the camera coordinate system are XC, YC, ZC.
[0075] FIG. 9 shows an example 900 calculating the height di of a tracked object (e.g., a human) in an image. As shown in FIG. 9, an image plane 910, a pinhole plane 920, and an object surface 930 (e.g., surface of the human) , and a camera optical axis 940 are shown. When the camera optical axis 940 lies within the horizontal plane (e.g., optical axis 940) , a vertical section can be analyzed to determine the actual height do of the object (e.g., the actual height of the human) and the height di of the object as depicted in the image (e.g., the height of the human as they appear in the image) . To obtain these values, the following formula can be used: (di / do) = (f / D) (equation 2)
[0076] where do is the actual height of the tracked object body area, di is the height of the object in the image, f is the focal length (e.g., the distance between the image plane 910 and the pinhole plane 920) , and D is the distance between the pinhole plane 920 and the object surface 930.
[0077] FIG. 10 is a graph 1000 illustrating an example of a pose of a camera 1010 (e.g., which may be mounted on a robot) . In the graph 1000 of FIG. 10, a horizontal coordinate system with XW, YW, ZW axes is shown. In FIG. 10, the camera 1010 is shown to be rotated (e.g., by angles α and β) to the XC, YC, ZC axes such that the camera 1010 is not aligned with the horizontal coordinate system XW, YW, ZW axes.
[0078] FIG. 11 shows an example of eliminating the pose errors (e.g., rotation errors of angles α and β) of the camera 1010 as shown in FIG. 10. In particular, FIG. 11 are graphs illustrating an example 1100 of eliminating errors in a pose of a camera. For the elimination of the pose errors, the angle α between XC and XW and the angle β between YC and YW can be directly obtained from the IMU data. As such, it possible to transform the XC-YC-ZC coordinate system to XW-YW-ZW coordinate system by using the following equations:
[0079] where the horizontal distance dH and angle φ (e.g., calculated by using the IMU) and the pinhole model can be used to represent the position of the human body. Each image (e.g., image frame) will get associated IMU data in real time, and the IMU data can be used to synchronously correct the error caused by the camera pose (e.g., which can adapt for a scene of a bumpy road) .
[0080] FIG. 12 shows an example for adjusting (e.g., correcting) the size of a tracking area (e.g., bounding box) during the following of a human. In particular, FIG. 12 is a diagram illustrating an example 1200 of adjusting (e.g., to correct) a size of a tracking area for a human. When the camera optical axis forms an angle α with the horizontal plane, the relationship between the image height and the actual height is shown in the following equations, where ε is the angle between the optical center of the lens and the center of the tracking frame before correction and the horizontal plane.
[0081] According to the pinhole model:
[0082] In the right triangle OEF, OF=OE·cos∠EOF=OE·cos (ε-α) (equation 13)
[0083] In the triangle ABC, according to the Law of Sines,
[0084] In the human-following cases, it can be considered that: ∠A≈∠ODB=90°-ε (equation 15) OE≈OD (equation 16)
[0085] Using equations 12, 13, 14, 15, and 16, α ε D (xw, yw, zw) Xw-Yw-Zw (equation 21)
[0086] The pixel height of bounding box can be recorded in the initial frame (e.g., initial image) as A0B0. For the following frames (e.g., following images or next images) , the pixel height of bounding box can be denoted as AiBi. AB is the target (e.g., human) height in real world of the bounding box.
[0087] OD0 is the distance from the camera to the target point of the initial frame (e.g., initial image) . ODi is the distance from the camera to the target point for the following camera frames (e.g., following images) , and ODi is to be solved for.
[0088] Similarly, ε is the angle for the initial frame, and ε’ is the angle for the following camera frames. Also, α is the angle for the initial frame, and α’ is the angle for the following camera frames. ODi can be solved for using the following equations:
[0089] FIG. 13 shows an example of adjusting (e.g., correcting) the size of the tracking area (e.g., bounding box) during human following. In particular, FIG. 13 is a diagram illustrating an example 1300 of further adjusting (e.g., to correct) a size of a tracking area 1330 (e.g., bounding box) for a human by using a depth map. Using the binocular depth camera, a depth map of the scene (e.g., including the human body and the background) can be obtained.
[0090] Within the tracking area 1330, the depth values along the horizontal centerline of the depth map can be analyzed. The depth values are plotted in graph 1310, which has an x-axis denoting the number of pixels and a y-axis denoting the depth. When viewing the depth map (or the depth values plotted in graph 1310) , there should be a noticeable depth discontinuity between the human body and the background. As shown in the graph 1310, the distance measured between two such discontinuity points can be denoted as the width. The graph 1320, which has an x-axis denoting the number of pixels and a y-axis denoting the |d(depth) / d (pixel) |, can be obtained by differentiating and calculating the absolute value for the depth values of graph 1310.
[0091] FIG. 14 shows an example of adjusting (e.g., correcting) the size of the tracking area (e.g., bounding box) as a human pose changes. In particular, FIG. 14 is a diagram illustrating an example 1400 of further adjusting (e.g., to correct) of a size of a tracking area for a human. In FIG. 14, an initial image 1410 is shown to include a tracked human. In the initial image 1410, a tracking area (e.g., bounding box) for the tracked human is shown. A following image 1420 is shown where the human has walked and changed his pose. The size of the tracking area (e.g., bounding box) shown in the following image 1420 is not correct for the changed pose of the human (e.g., the bounding box in the following image 1420 is too large for the updated pose of the human) . As such, the tracking area (e.g. bounding box) in the following image needs to be adjusted (e.g., corrected) .
[0092] The initial image 1430 is the same as the initial image 1410. A following image 1440 is shown where the human has walked and changed his pose. The size of the tracking area (e.g., bounding box) shown in the following image 1430 has been adjusted (e.g., the bounding box size has been reduced) and is correct for the changed pose of the human (e.g., the bounding box in the following image 1430 is the correct for the updated pose of the human) .
[0093] FIG. 15 is a flow chart illustrating an example of a process 1500 for object tracking using the techniques described herein. The process 1500 can be performed by a device (e.g., a binocular depth camera, such as camera 1010 of FIG. 10, a device including the binocular depth camera, and / or other device and / or camera) or by a component or system (e.g., a chipset) of the device. The operations of the process 1500 may be implemented as software components that are executed and run on one or more processors (e.g., processor 1610 of FIG. 16 or other processor (s) ) . Further, the transmission and / or reception of signals by the device in the process 1500 may be enabled, for example, by one or more antennas and / or one or more transceivers (e.g., wireless transceiver (s) ) .
[0094] At block 1510, the device (or component thereof) can obtain an image (e.g., a two-dimensional (2D) image) of a scene including an object (e.g., image 710a and / or image 710b of FIG. 7) . The object can include any type of object, such as a human, a vehicle, a robotic device or system, and / or other object.
[0095] At block 1520, the device (or component thereof) can determine a tracking area in the image to represent the object. In one illustrative example, the tracking area is a bounding box (e.g., the tracking area 730 of FIG. 7 in the form of a bounding box) . In some aspects, the bounding box includes four corners (e.g., as shown in FIG. 7) . In some cases, each respective corner of the four corners of the bounding box represents a first portion of the object, a second portion of the object, a third portion of the object, and a further portion of the object. In some examples, the four corners of the bounding box may include a top-left corner, a top-right corner, a bottom-left corner, and a bottom-right corner. In one illustrative example using a human as an example of the object, the top-left corner represents a left shoulder of the human, the top-right corner representing a right shoulder of the human, the bottom-left corner representing a left hip of the human, and the bottom-right corner representing a right hip of the human.
[0096] At block 1530, the device (or component thereof) can determine a point of the tracking area in the image is a target point (e.g., target point 740 of FIG. 7, target point 820 of FIG. 8, etc. ) . In some cases, the point of the tracking area is a center of the tracking area.
[0097] At block 1540, the device (or component thereof) can determine a distance between the apparatus and the target point. In some cases, the distance between the apparatus and the target point is a distance between a camera of the apparatus used to capture the image and the target point (e.g., the distance d of the target point 820 from the camera illustrated in FIG. 8) .
[0098] At block 1550, the device (or component thereof) can adjust a pose and a size of the tracking area in the image based on a pose of the apparatus and the distance between the apparatus and the target point (e.g., as described with respect to FIG. 6) . In some aspects, the device (or component thereof) can obtain the pose of the apparatus from inertial measurement unit (IMU) data of the apparatus.
[0099] In some cases, the device (or component thereof) can obtain a depth map of the scene comprising a depth of the object and a depth of background of the scene. In some aspects, the device (or component thereof) can further adjust the size of the tracking area based on the depth map. In some examples, the device (or component thereof) can adjust the size of the tracking area further based on an amount of discontinuity between the depth of the object and the depth of the background in the depth map.
[0100] In some examples, the process 1500 may be performed by one or more computing devices or apparatuses. In some illustrative examples, the process 1500 can be performed by the camera 1010 of FIG. 10 and / or one or more computing devices or systems (e.g., the computing system 1600 of FIG. 16) . In some cases, such a computing device or apparatus may include a processor, microprocessor, microcomputer, or other component of a device that is configured to carry out the steps of the process 1500. Such computing device may further include a network interface configured to communicate data.
[0101] The components of the computing device can be implemented in circuitry. For example, the components can include and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs) , digital signal processors (DSPs) , central processing units (CPUs) , and / or other suitable electronic circuits) , and / or can include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein. The computing device may further include a display (as an example of the output device or in addition to the output device) , a network interface configured to communicate and / or receive the data, any combination thereof, and / or other component (s) . The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other type of data.
[0102] The process 1500 is illustrated as a logical flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.
[0103] Additionally, the process 1500 may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0104] FIG. 16 is a block diagram illustrating an example of a computing system 1600, which may be employed by the disclosed system for a human following method based on binocular depth camera and IMU data. In particular, FIG. 16 illustrates an example of computing system 1600, which can be for example any computing device making up internal computing system, a remote computing system, a camera, or any component thereof in which the components of the system are in communication with each other using connection 1605. Connection 1605 can be a physical connection using a bus, or a direct connection into processor 1610, such as in a chipset architecture. Connection 1605 can also be a virtual connection, networked connection, or logical connection.
[0105] In some aspects, computing system 1600 is a distributed system in which the functions described in this disclosure can be distributed within a datacenter, multiple data centers, a peer network, etc. In some aspects, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some aspects, the components can be physical or virtual devices.
[0106] Example system 1600 includes at least one processing unit (CPU or processor) 1610 and connection 1605 that communicatively couples various system components including system memory 1615, such as read-only memory (ROM) 1620 and random access memory (RAM) 1625 to processor 1610. Computing system 1600 can include a cache 1612 of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1610.
[0107] Processor 1610 can include any general purpose processor and a hardware service or software service, such as services 1632, 1634, and 1636 stored in storage device 1630, configured to control processor 1610 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 1610 may essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.
[0108] To enable user interaction, computing system 1600 includes an input device 1645, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 1600 can also include output device 1635, which can be one or more of a number of output mechanisms. In some instances, multimodal systems can enable a user to provide multiple types of input / output to communicate with computing system 1600.
[0109] Computing system 1600 can include communications interface 1640, which can generally govern and manage the user input and system output. The communication interface may perform or facilitate receipt and / or transmission wired or wireless communications using wired and / or wireless transceivers, including those making use of an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, an AppleTM LightningTM port / plug, an Ethernet port / plug, a fiber optic port / plug, a proprietary wired port / plug, 3G, 4G, 5G and / or other cellular data network wireless signal transfer, a BluetoothTM wireless signal transfer, a BluetoothTM low energy (BLE) wireless signal transfer, an IBEACONTM wireless signal transfer, a radio-frequency identification (RFID) wireless signal transfer, near-field communications (NFC) wireless signal transfer, dedicated short range communication (DSRC) wireless signal transfer, 802.11 Wi-Fi wireless signal transfer, wireless local area network (WLAN) signal transfer, Visible Light Communication (VLC) , Worldwide Interoperability for Microwave Access (WiMAX) , Infrared (IR) communication wireless signal transfer, Public Switched Telephone Network (PSTN) signal transfer, Integrated Services Digital Network (ISDN) signal transfer, ad-hoc network signal transfer, radio wave signal transfer, microwave signal transfer, infrared signal transfer, visible light signal transfer, ultraviolet light signal transfer, wireless signal transfer along the electromagnetic spectrum, or some combination thereof.
[0110] The communications interface 1640 may also include one or more range sensors (e.g., LIDAR sensors, laser range finders, RF radars, ultrasonic sensors, and infrared (IR) sensors) configured to collect data and provide measurements to processor 1610, whereby processor 1610 can be configured to perform determinations and calculations needed to obtain various measurements for the one or more range sensors. In some examples, the measurements can include time of flight, wavelengths, azimuth angle, elevation angle, range, linear velocity and / or angular velocity, or any combination thereof. The communications interface 1640 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers that are used to determine a location of the computing system 1600 based on receipt of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US-based GPS, the Russia-based Global Navigation Satellite System (GLONASS) , the China-based BeiDou Navigation Satellite System (BDS) , and the Europe-based Galileo GNSS. There is no restriction on operating on any particular hardware arrangement, and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0111] Storage device 1630 can be a non-volatile and / or non-transitory and / or computer-readable memory device and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, a floppy disk, a flexible disk, a hard disk, magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, a compact disc read only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, digital video disk (DVD) optical disc, a blu-ray disc (BDD) optical disc, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a Memory card, a smartcard chip, a EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, random access memory (RAM) , static RAM (SRAM) , dynamic RAM (DRAM) , read-only memory (ROM) , programmable read-only memory (PROM) , erasable programmable read-only memory (EPROM) , electrically erasable programmable read-only memory (EEPROM) , flash EPROM (FLASHEPROM) , cache memory (e.g., Level 1 (L1) cache, Level 2 (L2) cache, Level 3 (L3) cache, Level 4 (L4) cache, Level 5 (L5) cache, or other (L#) cache) , resistive random-access memory (RRAM / ReRAM) , phase change memory (PCM) , spin transfer torque RAM (STT-RAM) , another memory chip or cartridge, and / or a combination thereof.
[0112] The storage device 1630 can include software services, servers, services, etc., that when the code that defines such software is executed by the processor 1610, it causes the system to perform a function. In some aspects, a hardware service that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1610, connection 1605, output device 1635, etc., to carry out the function. The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction (s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD) , flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.
[0113] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects, the methods may be performed in a different order than that described.
[0114] For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects.
[0115] Further, those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0116] Individual aspects may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed, but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.
[0117] Processes and methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.
[0118] In some aspects the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bitstream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0119] Those of skill in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof, in some cases depending in part on the particular application, in part on the desired design, in part on the corresponding technology, etc.
[0120] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor (s) may perform the necessary tasks. Examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
[0121] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.
[0122] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM) , read-only memory (ROM) , non-volatile random access memory (NVRAM) , electrically erasable programmable read-only memory (EEPROM) , FLASH memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.
[0123] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs) , general purpose microprocessors, an application specific integrated circuits (ASICs) , field programmable logic arrays (FPGAs) , or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor, ” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.
[0124] One of ordinary skill will appreciate that the less than ( “<” ) and greater than ( “>” ) symbols or terminology used herein can be replaced with less than or equal to ( “≤” ) and greater than or equal to ( “≥” ) symbols, respectively, without departing from the scope of this description.
[0125] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof.
[0126] The phrase “coupled to” or “communicatively coupled to” refers to any component that is physically connected to another component either directly or indirectly, and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.
[0127] Claim language or other language reciting “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B, C and C, A and A and B, and so on) , or any other ordering, duplication, or combination of A, B, and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B”or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein.
[0128] Claim language or other language reciting “at least one processor configured to, ” “at least one processor being configured to, ” “one or more processors configured to, ” “one or more processors being configured to, ” or the like indicates that one processor or multiple processors (in any combination) can perform the associated operation (s) . For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a certain subset of operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z.
[0129] Where reference is made to one or more elements performing functions (e.g., steps of a method) , one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function) . Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.
[0130] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method) , the entity may be configured to cause one or more elements (individually or collectively) to perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and / or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and / or each function need not be performed in whole by only one component (e.g., different components may perform different sub-functions of a function) .
[0131] The various illustrative logical blocks, modules, engines, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, engines, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0132] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as engines, modules, or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM) , read-only memory (ROM) , non-volatile random access memory (NVRAM) , electrically erasable programmable read-only memory (EEPROM) , FLASH memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.
[0133] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs) , general purpose microprocessors, an application specific integrated circuits (ASICs) , field programmable logic arrays (FPGAs) , or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor, ” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated software modules or hardware modules configured for encoding and decoding, or incorporated in a combined video encoder-decoder (CODEC) .
[0134] Illustrative aspects of the disclosure include:
[0135] Aspect 1. An apparatus for object tracking, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: obtain an image of a scene including an object; determine a tracking area in the image to represent the object; determine a point of the tracking area in the image is a target point; determine a distance between the apparatus and the target point; and adjust a pose and a size of the tracking area in the image based on a pose of the apparatus and the distance between the apparatus and the target point.
[0136] Aspect 2. The apparatus of Aspect 1, wherein the tracking area is a bounding box.
[0137] Aspect 3. The apparatus of Aspect 2, wherein the bounding box comprises four corners, and wherein each respective corner of the four corners of the bounding box represents a first portion of the object, a second portion of the object, a third portion of the object, and a further portion of the object.
[0138] Aspect 4. The apparatus of Aspect 3, wherein the object is a human.
[0139] Aspect 5. The apparatus of Aspect 4, wherein the four corners of the bounding box comprise a top-left corner, a top-right corner, a bottom-left corner, and a bottom-right corner, the top-left corner representing a left shoulder of the human, the top-right corner representing a right shoulder of the human, the bottom-left corner representing a left hip of the human, and the bottom-right corner representing a right hip of the human.
[0140] Aspect 6. The apparatus of any one of Aspects 1 to 5, wherein the object is a vehicle.
[0141] Aspect 7. The apparatus of any one of Aspects 1 to 6, wherein the at least one processor is configured to obtain the pose of the apparatus from inertial measurement unit (IMU) data of the apparatus.
[0142] Aspect 8. The apparatus of any one of Aspects 1 to 7, wherein the at least one processor is configured to obtain a depth map of the scene comprising a depth of the object and a depth of background of the scene.
[0143] Aspect 9. The apparatus of Aspect 8, wherein the at least one processor is configured to further adjust the size of the tracking area based on the depth map.
[0144] Aspect 10. The apparatus of Aspect 9, wherein the at least one processor is configured to adjust the size of the tracking area further based on an amount of discontinuity between the depth of the object and the depth of the background in the depth map.
[0145] Aspect 11. The apparatus of any one of Aspects 1 to 10, wherein the apparatus is a binocular depth camera.
[0146] Aspect 12. The apparatus of any one of Aspects 1 to 11, wherein the image is a two-dimensional (2D) image.
[0147] Aspect 13. The apparatus of any one of Aspects 1 to 12, wherein the point of the tracking area is a center of the tracking area.
[0148] Aspect 14. The apparatus of any one of Aspects 1 to 13, wherein the distance between the apparatus and the target point is a distance between a camera of the apparatus used to capture the image and the target point.
[0149] Aspect 15. A method for object tracking at a computing device, the method comprising: obtaining an image of a scene including an object; determining a tracking area in the image to represent the object; determining a point of the tracking area in the image is a target point; determining a distance between the computing device and the target point; and adjusting a pose and a size of the tracking area in the image based on a pose of the computing device and the distance between the computing device and the target point.
[0150] Aspect 16. The method of Aspect 15, wherein the tracking area is a bounding box.
[0151] Aspect 17. The method of Aspect 16, wherein the bounding box comprises four corners, and wherein each respective corner of the four corners of the bounding box represents a first portion of the object, a second portion of the object, a third portion of the object, and a further portion of the object.
[0152] Aspect 18. The method of Aspect 17, wherein the object is a human.
[0153] Aspect 19. The method of Aspect 18, wherein the four corners of the bounding box comprise a top-left corner, a top-right corner, a bottom-left corner, and a bottom-right corner, the top-left corner representing a left shoulder of the human, the top-right corner representing a right shoulder of the human, the bottom-left corner representing a left hip of the human, and the bottom-right corner representing a right hip of the human.
[0154] Aspect 20. The method of any one of Aspects 15 to 19, wherein the object is a vehicle.
[0155] Aspect 21. The method of any one of Aspects 15 to 20, further comprising obtaining the pose of the computing device from inertial measurement unit (IMU) data of the computing device.
[0156] Aspect 22. The method of any one of Aspects 15 to 21, further comprising obtaining a depth map of the scene comprising a depth of the object and a depth of background of the scene.
[0157] Aspect 23. The method of Aspect 22, further comprising further adjusting the size of the tracking area based on the depth map.
[0158] Aspect 24. The method of Aspect 23, wherein adjusting the size of the tracking area is further based on an amount of discontinuity between the depth of the object and the depth of the background in the depth map.
[0159] Aspect 25. The method of any one of Aspects 15 to 24, wherein the computing device is a binocular depth camera.
[0160] Aspect 26. The method of any one of Aspects 15 to 25, wherein the image is a two-dimensional (2D) image.
[0161] Aspect 27. The method of any one of Aspects 15 to 26, wherein the point of the tracking area is a center of the tracking area.
[0162] Aspect 28. The method of any one of Aspects 15 to 27, wherein the distance between the computing device and the target point is a distance between a camera of the computing device used to capture the image and the target point.
[0163] Aspect 29. An apparatus comprising one or more means for performing operations according to any of Aspects 15 to 28.
[0164] Aspect 30. A computer-readable memory storing instructions which, when executed by a computing device, cause the computing device to perform operations according to any of Aspects 15 to 28.
[0165] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but is to be accorded the full scope consistent with the language claims, wherein reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more. ”
Claims
1.An apparatus for object tracking, the apparatus comprising:at least one memory; andat least one processor coupled to the at least one memory and configured to:obtain an image of a scene including an object;determine a tracking area in the image to represent the object;determine a point of the tracking area in the image is a target point;determine a distance between the apparatus and the target point; andadjust a pose and a size of the tracking area in the image based on a pose of the apparatus and the distance between the apparatus and the target point.2.The apparatus of claim 1, wherein the tracking area is a bounding box.3.The apparatus of claim 2, wherein the bounding box comprises four corners, and wherein each respective corner of the four corners of the bounding box represents a first portion of the object, a second portion of the object, a third portion of the object, and a further portion of the object.4.The apparatus of claim 3, wherein the object is a human.5.The apparatus of claim 4, wherein the four corners of the bounding box comprise a top-left corner, a top-right corner, a bottom-left corner, and a bottom-right corner, the top-left corner representing a left shoulder of the human, the top-right corner representing a right shoulder of the human, the bottom-left corner representing a left hip of the human, and the bottom-right corner representing a right hip of the human.6.The apparatus of claim 1, wherein the object is a vehicle.7.The apparatus of claim 1, wherein the at least one processor is configured to obtain the pose of the apparatus from inertial measurement unit (IMU) data of the apparatus.8.The apparatus of claim 1, wherein the at least one processor is configured to obtain a depth map of the scene comprising a depth of the object and a depth of background of the scene.9.The apparatus of claim 8, wherein the at least one processor is configured to further adjust the size of the tracking area based on the depth map.10.The apparatus of claim 9, wherein the at least one processor is configured to adjust the size of the tracking area further based on an amount of discontinuity between the depth of the object and the depth of the background in the depth map.11.The apparatus of claim 1, wherein the apparatus is a binocular depth camera.12.The apparatus of claim 1, wherein the image is a two-dimensional (2D) image.13.The apparatus of claim 1, wherein the point of the tracking area is a center of the tracking area.14.The apparatus of claim 1, wherein the distance between the apparatus and the target point is a distance between a camera of the apparatus used to capture the image and the target point.15.A method for object tracking at a computing device, the method comprising:obtaining an image of a scene including an object;determining a tracking area in the image to represent the object;determining a point of the tracking area in the image is a target point;determining a distance between the computing device and the target point; andadjusting a pose and a size of the tracking area in the image based on a pose of the computing device and the distance between the computing device and the target point.16.The method of claim 15, wherein the tracking area is a bounding box.17.The method of claim 16, wherein the bounding box comprises four corners, and wherein each respective corner of the four corners of the bounding box represents a first portion of the object, a second portion of the object, a third portion of the object, and a further portion of the object.18.The method of claim 17, wherein the object is a human.19.The method of claim 18, wherein the four corners of the bounding box comprise a top-left corner, a top-right corner, a bottom-left corner, and a bottom-right corner, the top-left corner representing a left shoulder of the human, the top-right corner representing a right shoulder of the human, the bottom-left corner representing a left hip of the human, and the bottom-right corner representing a right hip of the human.20.The method of claim 15, wherein the object is a vehicle.21.The method of claim 15, further comprising obtaining the pose of the computing device from inertial measurement unit (IMU) data of the computing device.22.The method of claim 15, further comprising obtaining a depth map of the scene comprising a depth of the object and a depth of background of the scene.23.The method of claim 22, further comprising further adjusting the size of the tracking area based on the depth map.24.The method of claim 23, wherein adjusting the size of the tracking area is further based on an amount of discontinuity between the depth of the object and the depth of the background in the depth map.25.The method of claim 15, wherein the computing device is a binocular depth camera.26.The method of claim 15, wherein the image is a two-dimensional (2D) image.27.The method of claim 15, wherein the point of the tracking area is a center of the tracking area.28.The method of claim 15, wherein the distance between the computing device and the target point is a distance between a camera of the computing device used to capture the image and the target point.
Citation Information
Patent Citations
Adaptive resizing of manipulatable and readable objects
WO2023219612A1