Motion tracking using neural networks and depth images
Patent Information
- Application Number
- CN202480088829.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2026-09-25
Smart Images

Figure CN122826596A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to neural networks (also known as “deep neural networks” or “DNNs”), and more specifically, to motion tracking utilizing neural networks and depth images. Background Technology
[0002] The past decade has witnessed the rapid rise of artificial intelligence (AI)-based data processing technologies, particularly those based on deep neural networks (DNNs). DNNs have been widely adopted in computer vision, speech recognition, image and video processing primarily due to their ability to achieve accuracy surpassing human levels. Many DNNs (e.g., convolutional networks) can extract features from images to predict the classification or state of objects captured within them. Attached Figure Description
[0003] The various embodiments will be readily understood from the following detailed description taken in conjunction with the accompanying drawings. For ease of description, similar reference numerals denote similar structural elements. In the accompanying figures, embodiments are shown by way of example rather than limitation.
[0004] Figure 1 This is a block diagram of a motion tracking system according to various embodiments.
[0005] Figure 2 This is a block diagram of an imaging module according to various embodiments.
[0006] Figure 3 Example local areas are shown for displaying content to a person based on estimated human movement according to various embodiments.
[0007] Figure 4 Example motion tracking processes according to various embodiments are shown.
[0008] Figure 5 Example template feature diagrams according to various embodiments are shown.
[0009] Figure 6 The 3D joint positions representing an estimated three-dimensional (3D) pose of a person are shown according to various embodiments.
[0010] Figure 7 Example DNNs according to various embodiments are shown.
[0011] Figure 8 An AI-based motion tracking environment according to various embodiments is shown.
[0012] Figure 9 This is a flowchart illustrating motion tracking methods according to various embodiments.
[0013] Figure 10This is a block diagram of an example computing device according to various embodiments. Detailed Implementation
[0014] Overview DNNs typically consist of a sequence of layers. A DNN layer can include one or more deep learning operations (also known as "neural network operations"), such as matrix multiplication, convolution, pooling, element-wise operations, linear operations, non-linear operations, etc. The input or output data of a deep learning operation can be arranged in a data structure called a tensor. A tensor is a data structure with multiple elements in one or more dimensions. Examples of tensors include vectors (which are one-dimensional (1D) tensors), matrices (which are two-dimensional (2D) tensors), three-dimensional (3D) tensors, four-dimensional (4D) tensors, and even higher-dimensional tensors. The dimensions of a tensor can correspond to axes, such as the axes in a coordinate system. Dimensions can be measured by the number of data points along an axis. The dimensions of a tensor can define the shape of the tensor. A DNN layer can receive one or more input tensors and compute an output tensor based on those one or more input tensors. The input to a convolutional layer can include an input tensor (also known as an "activation tensor" or "input feature map (IFM)") (which includes one or more activation values (also known as "input elements")) and a weight tensor. Weight tensors can be kernels (2D weight tensors), filters (3D weight tensors), or filter banks (4D weight tensors). The output of a convolutional layer can be an output tensor, also known as an output feature map (OFM). The output tensor can be used as the input tensor for the next layer in a DNN.
[0015] Immersive projection systems are widely used in many applications. In one example, a room can be equipped with one or more projectors that can project computer-generated content onto one or more screens in the room, allowing people in the room to have an immersive stereoscopic experience. The projectors or screens can be mounted on the ceiling, walls, or floor of the room. In some scenarios, the ceiling, walls, or floor of the room can be used as screens. However, due to the challenges of human motion capture in this environment, many currently available immersive projection systems have little or no interactivity. Projection systems with strong interactivity allow people to interactively control virtual content in real time through their body posture or movement, which would greatly enhance the user experience and increase product value. For body motion tracking, many currently available motion capture methods are rarely considered because they restrict user movement and are often expensive. Furthermore, some currently available projection methods suffer from low tracking accuracy because camera images can be severely affected by the projection effect, making accurate human detection and posture tracking difficult.
[0016] Several currently available techniques can estimate 3D human pose from point clouds. For example, three unsupervised losses can be used to learn 3D human keypoints from field point clouds without any human labels. Cascaded architectures can be used to enhance point feature extraction from challenging point clouds. 3D pose regression networks can be trained end-to-end to extract body features and regress 3D keypoint locations. These techniques can produce appropriate 3D keypoint results when the input human point cloud is complete and noise-free or minimally noisy. However, in many environments, people are captured by multiple cameras overhead, and the extracted human point cloud is often incomplete due to severe self-occlusion. Therefore, these currently available techniques often produce incorrect 3D results.
[0017] Embodiments of this disclosure can improve at least some of the challenges and problems described above by using a DNN to track the motion of an object in a local region based on a depth image captured in the local region. The motion tracking process may include a process of estimating the 3D pose of the object, such as a process of regressing the 3D position of the deformable skeleton of the object.
[0018] In various embodiments of this disclosure, a motion tracking system is used to track the motion of objects in a local region. Examples of local regions include rooms, buildings, open areas, or other relatively small graphical regions of other types. Examples of objects include people, robots, vehicles, animals, etc. The motion tracking system may include hardware (e.g., a depth camera) and software (e.g., a module for estimating 3D pose using a DNN based on depth images captured by the depth camera). In one example, a depth camera may be positioned within a local region (e.g., a room) to capture depth images of objects within the local region. The depth camera may be mounted on the ceiling of the room and pointed downwards. One or more depth cameras may be tilted.
[0019] Point clouds can be extracted from depth images. Point clouds extracted from different depth images can be fused into a single point cloud. A DNN (e.g., a point-based DNN) can be used to encode the point cloud into a low-dimensional feature map. This feature map can be a feature vector. This feature map can be further combined with a template feature map encoding a template structure of the object. The template structure can represent a reference skeleton of the object, including the object's joints and the connections between them. As the object moves and changes pose, the object's skeleton may change. A skeleton different from the reference skeleton can be called a deformable skeleton. The combined feature map can be processed by another DNN (e.g., a graph convolutional network) to regress the 3D positions of the joints. The 3D positions of the joints can be the coordinates of the joints in 3D space. Using template feature maps can improve the accuracy and robustness of learning the 3D positions of joints from the point cloud. Motion tracking systems can predict complete and plausible 3D skeleton joints from incomplete and noisy point clouds from a depth camera. The 3D positions of the object's joints indicate the estimated 3D pose of the object at a given time. The estimated 3D poses at different times can constitute the estimated motion of the object.
[0020] The estimated motion of an object can be used to dynamically present content in a localized area. In the example where the object is a person, the estimated motion can be used to facilitate interaction between the person and the presented content. One or more projectors may also be present in the localized area to project computer-generated content. The depth camera and (one or more) projectors can be fixed in a box-like device, which may be mounted in the center of the ceiling. The depth camera captures depth images of the object.
[0021] This disclosure provides a more efficient method for tracking motion and estimating 3D pose. This method simplifies hardware setup and enables accurate 3D human motion tracking to enhance real-time interactivity. It can improve the robustness and usability of various 3D motion tracking applications, including immersive projection, augmented reality, virtual reality, motion analytics, telepresence, film and game production, motion recognition, etc.
[0022] For illustrative purposes, specific figures, materials, and configurations have been set forth to provide a thorough understanding of the illustrative implementation. However, it will be apparent to those skilled in the art that this disclosure may be practiced without specific details, and / or may be practiced using only some of the aspects described. In other instances, well-known features have been omitted or simplified so as not to obscure the illustrative embodiments.
[0023] Furthermore, reference has been made to the accompanying drawings, which form part of this disclosure, and practical embodiments are illustrated in the drawings by way of illustration. It should be understood that other embodiments may be utilized and structural or logical changes may be made without departing from the scope of this disclosure. Therefore, the following detailed description should not be construed as limiting.
[0024] Various operations can be described sequentially as a plurality of discrete actions or operations in a manner most conducive to understanding the claimed subject matter. However, the order of description should not be construed as implying that these operations must depend on the order. In particular, these operations may not be performed in the order presented. The described operations may be performed in a different order than in the described embodiments. Various additional operations may be performed, or the described operations may be omitted in additional embodiments.
[0025] For the purposes of this disclosure, the phrase "A or B" or the phrase "A and / or B" refers to (A), (B), or (A and B). For the purposes of this disclosure, the phrase "A, B, or C" or the phrase "A, B, and / or C" refers to (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). When used to refer to a measurement range, the term "between" includes the endpoints of the measurement range.
[0026] This description uses the phrases "in one embodiment" or "in an embodiment," both of which can refer to one or more of the same or different embodiments. Terms such as "comprising," "including," "having," etc., used with respect to embodiments of this disclosure are synonyms. This disclosure may use perspective-based descriptions such as "above," "below," "top," "bottom," and "side" to interpret various features of the drawings; however, these terms are merely for ease of discussion and do not imply any desired or required direction. The drawings are not necessarily drawn to scale. Unless otherwise stated, the use of ordinal adjectives such as "first," "second," and "third" to describe common objects indicates only different instances of the similar objects referred to and is not intended to imply that the objects described must be arranged in a given order, whether temporally, spatially, in rank, or otherwise.
[0027] In the following detailed description, terms commonly used by those skilled in the art will be used to describe various aspects of the illustrative implementations in order to convey the substance of their work to others skilled in the art.
[0028] The terms “substantially,” “close to,” “approximately,” “near,” and “about” generally refer to within + / -20% of the target value as described herein or known in the art. Similarly, terms indicating the orientation of various elements, such as “coplanar,” “perpendicular,” “orthogonal,” “parallel,” or any other angle between elements, generally refer to within + / -5-20% of the target value based on the description herein or known in the art.
[0029] Furthermore, the terms “comprising,” “including,” “having,” or any other variations thereof are intended to cover non-exclusive inclusion. For example, a method, process, apparatus, or DNN accelerator that includes a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such a method, process, apparatus, or DNN accelerator. Additionally, the term “or” refers to inclusive “or” rather than exclusive “or.”
[0030] The systems, methods, and apparatuses disclosed herein are innovative in several ways, but none of them alone is responsible for all the desired properties disclosed herein. Details of one or more implementations of the subject matter described herein are set forth in the following description and figures.
[0031] Example motion tracking system Figure 1 This is a block diagram of a motion tracking system 100 according to various embodiments. The motion tracking system 100 uses deep learning methods to estimate the 3D pose of an object from a depth image. The motion tracking system 100 includes an interface module 110, a depth camera 120 (solely referred to as "depth camera 120"), an imaging module 130, a point cloud generator 140, a neural network 150, a stitcher 160, another neural network 170, an output module 180, and a data storage device 190. In other embodiments, the motion tracking system 100 may include alternative configurations, different, or additional components. Furthermore, the functionality of components attributed to the motion tracking system 100 may be performed by different components included in the motion tracking system 100 or by different modules.
[0032] Interface module 110 facilitates communication between motion tracking system 100 and other systems, devices, or modules. For example, interface module 110 can receive images from online systems (e.g., social media systems, online image libraries, online search tools, etc.), devices, etc. As another example, interface module 110 can receive one or more data points for training or testing neural networks 150 or 170. Yet another example, interface module 110 can send data generated by motion tracking system 100 to other systems, devices, or modules. For example, interface module 110 can send estimated 3D poses of objects, object animations, motion analysis results, or other types of motion tracking information to other systems, devices, or modules.
[0033] Depth camera 120 captures depth images. Depth camera 120 can be positioned within a local region and infer the depth of points within that region. Depth camera 120 can output a depth image. The depth image can include depth pixels, each of which can encode the distance of a point from depth camera 120. In some embodiments, depth camera 120 is placed at a specific location within the local region. Depth camera 120 can be configured to capture depth images at different angles. For example, one or more depth cameras 120 can be tilted.
[0034] Imaging module 130 can control and manage depth cameras 120. In some embodiments, imaging module 130 can determine how many depth cameras 120 to place in a local area. Imaging module 130 can also determine the placement location of depth cameras 120. Imaging module 130 can also synchronize depth cameras 120 so that depth cameras 120 can capture depth images simultaneously. Imaging module 130 can deploy depth cameras, for example, by commanding depth cameras to capture depth images. Imaging module 130 can receive depth images from depth cameras. Different depth cameras can capture depth images simultaneously. These depth images can capture different portions of a local area. In some embodiments, there can be overlap between depth images captured by different depth cameras. For example, a first depth image and a second depth image from two adjacently placed depth cameras can both capture a first portion of a local area, while the first depth image can also capture a second portion of the local area (which is not captured by the second depth image), and another depth image can capture a third portion of the local area (which is not captured by the first depth image). Imaging module 130 can provide depth images for further processing. Certain aspects of imaging module 130 are described below in conjunction with... Figure 2 Describe it.
[0035] Point cloud generator 140 generates a point cloud based on a depth image received from imaging module 130. In some embodiments, point cloud generator 140 may segment each depth image into depth pixels representing a target object (“object segmentation pixels”) and background depth pixels. Background depth pixels may represent other objects or features in a local region. Point cloud generator 140 may remove background depth pixels from the depth image. In one embodiment, point cloud generator 140 may define a depth threshold and select depth pixels whose values do not meet the depth threshold (e.g., above or below the depth threshold). In another embodiment, point cloud generator 140 may compare the depth image with a 3D reference model of the local region. The 3D reference model of the local region may represent the local region when no object is present. Point cloud generator 140 may label depth pixels in the depth image that match one or more features of the 3D reference model as background depth pixels. Additionally or alternatively, point cloud generator 140 may use Euclidean clustering extraction to further refine the segmentation process.
[0036] After the point cloud generator 140 extracts object segmentation pixels, it can convert the extracted pixels into a 3D point cloud in camera space. The point cloud generator 140 can define the camera space based on camera intrinsic parameters, which can be determined by the imaging module 130. In some embodiments, the point cloud generator 140 can generate multiple point clouds, each generated from depth pixels extracted from images at different depths. Each point cloud can reside in a camera space determined based on the intrinsic parameters of the depth camera that captured the corresponding depth image. The point cloud generator 140 can also convert these point clouds to a unique world space based on the extrinsic parameters of the depth camera. The point cloud generator 140 can also fuse the point clouds into a single point cloud. The fused point cloud can be a complete point cloud of the object.
[0037] In some embodiments, the point cloud generator 140 can reduce the size of the merged point cloud. For example, the point cloud generator 140 can determine whether the total number of points in the merged point cloud exceeds a predetermined number. In response to determining that the total number of points in the merged point cloud exceeds the predetermined number, the point cloud generator 140 can downsample the merged point cloud to generate a point cloud with a predetermined number of points. In one example, the predetermined number could be 5000. Each point in the point cloud can have 3D coordinates indicating the point's position in world space. The point cloud generator 140 can also perform a normalization operation on the 3D coordinates of the points, such that the normalized 3D coordinates of the points indicate the point's position in a unit cube space. The point cloud in the unit cube space can be represented as follows: Point cloud generator 140 can provide point clouds to neural network 150 for further processing.
[0038] Neural network 150 generates a feature map based on the point cloud generated by point cloud generator 140. In some embodiments, neural network 150 may be a point-based DNN. Neural network 150 may receive the point cloud as input and extract features from the point cloud to generate the feature map. In some embodiments, the feature map has a dimension smaller than that of the point cloud. For example, the feature map may have 1024 dimensions. The feature map may include low-dimensional global features extracted from the point cloud by neural network 150. Neural network 150 may output the feature map and provide it to stitcher 160.
[0039] The stitcher 160 receives a feature map from the neural network 150 and a template feature map of an object. The template feature map can encode a reference structure of the object, such as a reference skeleton structure. The reference structure can include keypoints of the object and connections between keypoints. Keypoints can be joints of the object. In the example where the object is a person, keypoints can include skeletal joints of a person. Connections between keypoints can be determined based on the connections between the corresponding skeletal joints. The reference structure can represent a reference pose of the object. In some embodiments, the template feature map can have a spatial dimension of J × 3, where J is the number of keypoints, and each keypoint has 3D coordinates indicating its location. The stitcher 160 can perform a stitching operation on the feature map from the neural network 150 and the template feature map of the object to generate a combined feature map.
[0040] In some embodiments, the combined feature map can be a tensor of spatial size J × (3 + M), where M is the dimension of the feature map from neural network 150. In the example where M = 1024, the combined feature map can have a spatial size of J × 1027. The feature map from neural network 150 can be embedded into the template feature map through a splicing operation. The combined feature map can include a graph structure containing nodes and edges. Nodes can correspond to keypoints in the template feature map. Edges can represent connections between two or more keypoints in the template feature map. Splicer 160 can provide the combined feature map to neural network 170.
[0041] Neural network 170 receives a combined feature map as input and uses the combined feature map to estimate the 3D pose of an object. In some embodiments, neural network 170 may be a convolutional network, such as a graph convolutional network. In one example, neural network 170 may be a semantic graph convolutional network including one or more semantic graph convolutional layers. Neural network 170 can extract features of the object from the combined feature map and regress 3D keypoint locations. Neural network 170 can output a tensor of spatial size J × 3, which encodes the 3D locations of J keypoints of the object. The 3D locations of the J keypoints can define the estimated 3D pose of the object. In some embodiments, the 3D locations of the keypoints can be represented as... .
[0042] Despite Figure 1Neural networks 150 and 170 are shown as separate neural networks, but they can be two parts of a single DNN. In some embodiments, the DNN can be trained end-to-end. For example, training samples can be input into neural network 150 and processed by neural networks 150 and 170, and the internal parameters of neural networks 150 and 170 can be tuned based on a loss determined from the output of neural network 170 and the ground truth labels of the training samples. The training samples can be point clouds representing objects. The ground truth labels of the training samples can be known or verified 3D poses of the objects. The ground truth labels can be 3D keypoint labels. In some embodiments, training samples are generated from depth images captured by a depth camera mounted on the ceiling of a local area, while ground truth labels can be generated using depth images captured by depth cameras mounted on the ceiling and sidewalls of the local area. Ground truth labels can be generated by fitting a parameterized 3D model to the point cloud extracted from these depth images and estimating the 3D shape and pose of the object using an iterative optimization algorithm. 3D keypoint coordinates can be calculated based on the 3D shape and pose of the object, and then used as ground truth labels.
[0043] Output module 180 processes the 3D keypoint coordinates regressed by neural network 170. In some embodiments, output module 180 may transform the 3D keypoint coordinates from canonical space to world space. Additionally or alternatively, output module 180 may perform a filter (e.g., an inline filter) to smooth the 3D coordinates. Output module 180 may also calculate 3D keypoint angles or 3D rotations based on the keypoint positions. Output module 180 may output an estimated 3D pose of the object.
[0044] The output module 180 can also control the projector in a local area based on the motion of the tracked object. In some embodiments, the output module 180 can control the content items presented in the local area based on the estimated 3D pose of the object. For example, the output module 180 can generate or modify content items based on the estimated 3D pose of the object. The output module 180 can command the projector in the local area to present the generated or modified content items. Content items may include images, videos, audio, text, other types of content, or some combination thereof. The output module 180 can facilitate user interaction with the content items presented by the projector. In this way, motion tracking results can be used to facilitate real-time interactive control of the projection system.
[0045] In some embodiments, users can use their gestures to switch content displays, use body gestures to control digital characters to generate animations, etc. In one example, output module 180 can zoom in on an image or video presented to a person after it detects a movement indicating zoom. As another example, output module 180 can present a different image or video to a person after it detects a gesture indicating movement to the next content item. Yet another example, a game player can use their body movements to control the actions of one or more virtual characters in a game. Output module 180 can determine various parameters of the game player's estimated movement (e.g., speed, direction, timing, jump height, trajectory, etc.) and replicate the game player's movement in the game.
[0046] Data storage device 190 stores data associated with motion tracking system 100, such as data received, generated, or used by components of motion tracking system 100. For example, data storage device 190 may store camera parameters (e.g., intrinsic parameters, extrinsic parameters, etc.) of annotated networks. Data storage device 190 may also store training or validation data used to train or validate neural networks 150 and 170. Data storage device 190 may also store images received by interface module 110, outputs of neural networks 150 and 170, graphics representing estimated 3D poses generated by output module 180, content items generated by output module 180, etc. In some embodiments, motion tracking system 100 may include or be associated with more than one data storage device. Data storage device 190 may be implemented as random access memory (RAM), such as static RAM (SRAM), disk storage, nearline storage, online storage, offline storage, etc.
[0047] Figure 2 This is a block diagram of an imaging module 200 according to various embodiments. The imaging module 200 can control and manage a depth camera in a local area. The imaging module 200 is... Figure 1 An example of the imaging module 130. (e.g.) Figure 2 As shown, the imaging module 200 includes a configuration module 210, a placement module 220, a calibration module 230, and a synchronization module 240. In other embodiments, the imaging module 200 may include alternative configurations, different, or additional components. Furthermore, the functionality of the components attributed to the imaging module 200 may be performed by different components included in the imaging module 200 or by different modules.
[0048] Configuration module 210 can configure an immersive projection device for a local area. For example, configuration module 210 can define an immersive space for placing the immersive projection device in the local area. The immersive space can occupy the entire local area or a portion of the local area. The immersive space can have a shape, such as a cuboid, cube, cylinder, hexagonal prism, or other shapes. The immersive projection device may include a depth camera (e.g., Figure 1 The immersive space includes a depth camera 120 and one or more projectors to be placed at various locations within the immersive space. In some embodiments, the configuration module 210 can determine the dimensions of the immersive space. In an example where the immersive space is a cuboid or cubic space, the configuration module 210 can determine the length, width, or height of the space. For a cylindrical or hexagonal space, the configuration module 210 can determine the radius or height of the space. In some embodiments, the configuration module 210 can determine that the length and width of the immersive space should be the same or similar. In some embodiments, the configuration module 210 can determine that the length of the immersive space is between approximately three meters and approximately eight meters.
[0049] Configuration module 210 can determine the spatial constraints of the immersive projection device. Configuration module 210 can determine an area ratio of at least approximately 5%. The area ratio can be the ratio of the area of the immersive space to the area of a local area (e.g., the area of the local area ceiling). Configuration module 210 can also determine a height ratio between approximately 5% and approximately 20%. The height ratio can be the ratio of the height of the immersive space to the height of the local area. Configuration module 210 can also determine a minimum distance from the immersive space to one or more edges of the local area. In some embodiments, the minimum distance is expressed as... Where L represents the length of the immersive space and W represents the width of the immersive space. Configuration module 210 can also determine that the ratio of a person's height to the height of the local area is at least approximately 2. This person can be an object whose movement is tracked using depth images captured by a depth camera, or an object to which a projector presents computer-generated content.
[0050] In some embodiments, configuration module 210 can create a 3D immersive space by employing synchronous multi-channel video or projection technology and stereoscopic optoelectronic technology. Users can immerse themselves in this immersive space. The immersive space provides users with an immersive projection environment, which can be defined by a stereoscopic interface. Immersive projection devices can be integrated with speaker systems and posture capture systems to produce realistic, high-resolution 3D visual experiences and 360-degree interactive experiences. This setup allows users to immerse themselves in a natural environment, enhancing the overall sensory experience.
[0051] Placement module 220 controls the placement of depth cameras within a space defined by configuration module 210. In some embodiments, placement module 220 can determine the number of depth cameras to be placed in a local area based on the shape or size of the immersive space. For example, placement module 220 can determine that a cubic room requires at least 5 depth cameras, a cylindrical room may require at least 2 depth cameras, or a hexagonal room may require at least 6 depth cameras. Placement module 220 can also determine the position and orientation (e.g., direction or location) of each depth camera. In some embodiments, placement module 220 can determine that N-1 depth cameras should have a downward orientation (e.g., from the ceiling to each wall) and a top-down view. N can be the total number of depth cameras in the local area. In an example where the immersive space is a cuboid or cube, placement module 220 can distribute the depth cameras across the five faces of the immersive space to ensure complete coverage of the space.
[0052] The placement module 220 can also set the angle of the depth cameras to ensure that the depth cameras can capture the entire immersive space. In an example where the immersive space is cubic in shape, the placement module 220 can determine that the camera angle on the side of the cube is tilted downwards between approximately 30 degrees and approximately 55 degrees. The placement module 220 can adjust the angle of the depth cameras as the immersive space changes. The angle between depth cameras (e.g., depth cameras mounted on a wall) can be 360° / (N-1). In some embodiments, the placement module 220 can ensure that there is sufficient overlap between the fields of view of adjacent depth cameras. The overlap between adjacent depth cameras can facilitate the fusion of point clouds extracted from the depth images generated by these depth cameras.
[0053] In some embodiments, the placement module 220 can control the placement of the projectors within a space defined by the configuration module 210. For example, the placement module 220 can determine the number of projectors or projector characteristics (e.g., projection range) based on the shape or size of the immersive space. In one example, the placement module 220 can configure the projectors to ensure that the projection of content can cover the entire space of a local area or a specific portion of a local area.
[0054] Calibration module 230 calibrates depth cameras in a local area. In some embodiments, calibration module 230 can perform internal or external calibration. Calibration module 230 can calculate internal and external parameters for each depth camera. Calibration module 230 can use calibration tools (e.g., calibration plates) to determine precise camera parameters. The camera parameters determined by calibration module 230 can be used, for example, by point cloud generator 140, to extract point clouds from depth images captured by depth cameras. Calibration module 230 can also calibrate projectors in a local area. In some embodiments, calibration module 230 can calibrate projectors to correct projection range and geometric distortion.
[0055] Synchronization module 240 synchronizes the depth camera and projector in an immersive projection device. In some embodiments, synchronization module 240 can synchronize the depth cameras so that they can simultaneously capture depth images. This ensures that the depth cameras capture the same pose of the same object in a local area, which can help track the object's motion. For multi-camera synchronization, synchronization module 240 can use a depth camera with external synchronization capabilities. In some embodiments, synchronization module 240 can use a hardware clock, network time protocol, or other tools to ensure time synchronization between the depth camera and the projector.
[0056] Example of immersive projection based on motion tracking Figure 3 An example local area 300 for displaying content to a person 310 based on estimated motion of the person 310, according to various embodiments, is shown. For illustrative and simplified purposes, the local area 300 is a room with a cuboid shape. In other embodiments, the local area 300 may be a space within a larger area. For example, the local area 300 may be an immersive space identified by an imaging module from a larger area.
[0057] like Figure 3 As shown, the partial area 300 has a ceiling 301 and a floor 302. The partial area 300 also has a sidewall located between the ceiling 301 and the floor 302. A depth camera 320 (referred to separately as "depth camera 320") and one or more projectors ( Figure 3 (Not shown in the image) is placed in local region 300. Depth camera 320 can be... Figure 1 An example of a medium depth camera 120. Figure 3 In this design, depth cameras 320 are mounted on the ceiling 301 to capture depth images of a person. The depth cameras 320 are positioned at different locations on the ceiling 301 so that they can detect the entire local area 300. The depth cameras 320 can capture depth images in a downward direction. In some embodiments, the orientation of the different depth cameras 320 can be different. For example, the depth cameras 320 can be tilted to capture depth images using a tilt angle, rather than perpendicular to the ceiling 301 or the floor 302. Although... Figure 3 Three depth cameras 320 are shown, but a different number of depth cameras 320 can be used. The fields of view of the depth cameras 320 can overlap, which can facilitate the fusion of point clouds extracted from depth images captured by the depth cameras 320. The number of depth cameras 320 mounted in the local region 300, the position of the depth cameras 320, and the configuration of the depth cameras 320 (e.g., field of view, angle, etc.) can be determined by [the specific configuration]. Figure 1 The imaging module 130 is defined. The depth cameras 320 can be synchronized so that they can capture depth images simultaneously.
[0058] The motion of a person 310 can be detected based on depth images captured by a depth camera 320. The detected motion of the person can be used to control the projector. Figure 3 In this embodiment, a projector projects image 330 onto the side wall of a local area 300, allowing person 310 to view image 330. Image 330 displays a scene in a bar, where a person is sitting at the counter and a bartender is making a drink. Image 330 can be a frame from a video (e.g., a movie, animation, game, etc.). Person can interact with the image by making gestures or other types of movement. For example, person 310 waves at image 330. After detecting the wave, image 330 can change to display a new image showing the bartender responding to person 310's wave. In other embodiments, person 310 can make other movements to control the projection differently.
[0059] Example motion tracking Figure 4 An example motion tracking process according to various embodiments is illustrated. The process includes estimating the 3D pose of an object in a local region. In some embodiments, the process may include estimating multiple 3D poses of the object within a time window. This process can be performed by… Figure 1 The motion tracking system 100 in the middle is executed. Figure 4 In some embodiments, the process begins with a depth image 401. The depth image 401 may be captured by multiple depth cameras placed within an immersive space in a localized area. In some embodiments, the immersive space may include at least a portion of the ceiling of the localized area.
[0060] The depth image 401 is converted into a point cloud 402 by the point extractor 410. The point cloud 402 can capture at least a portion of an object. The point cloud 402 can include multiple points in a 3D space with a cubic shape. An example of the point extractor 410 could be... Figure 1 The point cloud generator 140 in the diagram. Point cloud 402 is converted into a global feature map 403 by point encoder 420. Compared to point cloud 402, global feature map 403 can be a lower-dimensional feature map. For example, the total dimension in global feature map 403 can be less than the total number of points in point cloud 402. An example of point encoder 420 could be... Figure 1 The neural network 150 in the example.
[0061] The splicer 430 combines the global feature map 403 with the template feature map 404 to generate a spliced feature map 405. An example of the splicer 430 could be... Figure 1The splicer 160 is used. Template feature map 404 can encode the 3D positions of J keypoints of an object. The 3D positions of the keypoints can represent the reference structure of the object. Splicer 430 can embed global feature map 403 into template feature map 404. In some embodiments, splicer 430 can combine global feature map 403 and template feature map 404 in a dimension. This dimension of the spliced feature map 405 can be equal to the sum of the corresponding dimensions of the global feature map 403 and the corresponding dimensions of the template feature map 404. Another dimension of the spliced feature map 405 can be J.
[0062] The 3D keypoint location 406 is regressed from the stitched feature map 405 by DNN 440. An example of DNN 440 could be... Figure 1 The neural network 170 in the image. In some embodiments, 3D keypoint locations 406 can represent the deformable structure of an object. The deformable structure can be a structure different from the reference structure of the object represented by the template feature map 404. When the object moves or its pose changes, the 3D position of at least one keypoint may change. The 3D keypoint locations 406 can constitute an estimate of the pose of the object captured in the depth image 401. Although in Figure 4 As not shown in the diagram, the 3D key point position 406 can be further used in content presentation applications, such as immersive projection, virtual reality, augmented reality, mixed reality, etc.
[0063] Figure 5 Example template feature diagrams 500 according to various embodiments are shown. Template feature diagram 500 represents a human skeletal structure. Template feature diagram 500 can be used as a reference structure or reference posture for tracking human movement. In some embodiments, template feature diagram 500 can be created to track the movement of a specific person or a specific group of people. In other embodiments, template feature diagram 500 can be created to track the movement of any person. Template feature diagram 500 includes a plurality of keypoints 510, individually referred to as "keypoints 510". In some embodiments, each keypoint can correspond to a human bony joint. Keypoints 510 are connected using lines. The connections between keypoints 510 can be determined based on the connections of the corresponding bony joints.
[0064] The template feature map 500 can be used as a representation of a person's 3D pose and reference skeletal structure, which can be defined by the 3D positions of keypoints 510. As the person moves, the position of at least one keypoint 510 may change, resulting in different skeletal structures. Skeletal structures different from the template feature map 500 can be referred to as deformable structures. For illustrative purposes, the template feature map 500 is... Figure 5 The template feature map 500 includes 21 key points 510. In other embodiments, the template feature map 500 may include different, fewer, or more key points 510. Furthermore, the connections between the key points may differ. Figure 5The connection shown is shown.
[0065] Figure 6 The diagram illustrates 3D joint positions representing an estimated 3D pose of a person according to various embodiments. The 3D joint positions are determined by... Figure 6 The graphic representation 610 shows an estimate of a human posture 620. The graphic representation 610 includes joints 630 (referred to separately as "joints 630"). Joints 630 are formed by... Figure 6 The position of joint 630 can be the (X, Y, Z) coordinates of joint 630 in 3D space defined by the X, Y, and Z axes. Joint 630 can be a skeletal joint of a human. Pose 620 can be captured in one or more depth images, and the 3D position of joint 630 can be determined from one or more depth images using one or more DNNs. In some embodiments, the graphical representation 610 can be derived from a DNN (e.g., Figure 1 Neural network 170 or Figure 4 The output of DNN 440 in the middle.
[0066] Example DNN Figure 7 An example DNN 700 according to various embodiments is shown. At least a portion (or a part thereof) of the DNN 700 may be Figure 1 The neural network in the 150 or 170, Figure 4 The dot encoder 420 or Figure 4 An example of DNN 440. In Figure 7 In one embodiment, the DNN 700 includes a sequence of layers including multiple convolutional layers 710 (referred to individually as "convolutional layer 710"), multiple pooling layers 720 (referred to individually as "pooling layer 720"), and multiple fully connected layers 730 (referred to individually as "fully connected layer 730"). In other embodiments, the DNN 700 may include fewer, more, or different layers. During inference in the DNN 700, the layers of the DNN 700 perform tensor computations including a number of tensor operations, such as convolution (e.g., multiply-accumulate (MAC) operations, pooling operations, element-wise operations (e.g., element-wise addition, element-wise multiplication, etc.), other types of tensor operations, or certain combinations of these operations.
[0067] Convolutional layer 710 summarizes the presence of features in the input of DNN 700. Convolutional layer 710 acts as a feature extractor. The first layer of DNN 700 is convolutional layer 710. In the example, convolutional layer 710 performs a convolution operation on the input tensor 740 (also known as IFM 740) and filter 750. Figure 7As shown, the IFM 740 is represented by a 7×7×3 three-dimensional (3D) matrix. The IFM 740 includes three input channels, each represented by a 7×7 two-dimensional (2D) matrix. Each row of the 7×7 2D matrix contains 7 input elements (also called input points), and each column contains 7 input elements. The filter 750 is represented by a 3×3×3 3D matrix. The filter 750 includes three kernels, each corresponding to a different input channel of the IFM 740. The kernel is a 2D matrix of weights, where the weights are arranged by column and row. The kernel can be smaller than the IFM. Figure 7 In this embodiment, each kernel is represented by a 3×3 2D matrix. Each row of the 3×3 kernel contains 3 weights, and each column also contains 3 weights. The weights can be initialized and updated using gradient descent via backpropagation. The magnitude of the weights can represent the importance of filter 750 in extracting features from IFM 740.
[0068] The convolution involves a MAC operation on the input elements in IFM 740 and the weights in filter 750. The convolution can be either a standard convolution 763 or a depthwise convolution 783. In a standard convolution 763, the entire filter 750 slides over the IFM 740. All input channels are combined to produce an output tensor 760 (also known as OFM 760). OFM 760 is represented by a 5×5 2D matrix. Each row of the 5×5 2D matrix contains 5 output elements (also called output points), and each column also contains 5 output elements. For illustration, in... Figure 7 In one embodiment, the standard convolution includes a filter. In an embodiment with multiple filters, the standard convolution can produce multiple output channels in the OFM 760.
[0069] The multiplication applied between a kernel-sized local patch of IFM 740 and the kernel can be a dot product. A dot product is an element-wise multiplication between a kernel-sized local patch of IFM 740 and the corresponding kernel, then summed, always producing a single value. Because it produces a single value, this operation is often called a "scalar product." Using a kernel smaller than IFM 740 is intentional because it allows the same kernel (a set of weights) to be multiplied multiple times at different points on IFM 740. Specifically, the kernel is systematically applied from left to right and top to bottom to each overlapping portion or kernel-sized local patch of IFM 740. Multiplying the kernel by IFM 740 once results in a single value. Since the kernel is applied multiple times to IFM 740, the result of the multiplication is the output element of a 2D matrix. Thus, the 2D output matrix from the standard convolution 763 (i.e., OFM 760) is called OFM.
[0070] In depthwise convolution 783, the input channels are not combined. Instead, a MAC operation is performed on individual input channels and individual kernels to produce the output channel. For example... Figure 7 As shown, depthwise convolution 783 produces a depth output tensor 780. The depth output tensor 780 is represented by a 5×5×3 3D matrix. The depth output tensor 780 includes three output channels, each represented by a 5×5 2D matrix. Each row of the 5×5 2D matrix contains 5 output elements, and each column also contains 5 output elements. Each output channel is the result of a MAC operation performed on the input channels of the IFM 740 and the kernel of the filter 750. For example, the first output channel (dot pattern) is the result of a MAC operation on the first input channel (dot pattern) and the first kernel (dot pattern); the second output channel (horizontal stripe pattern) is the result of a MAC operation on the second input channel (horizontal stripe pattern) and the second kernel (horizontal stripe pattern); and the third output channel (diagonal stripe pattern) is the result of a MAC operation on the third input channel (diagonal stripe pattern) and the third kernel (diagonal stripe pattern). In such depthwise convolution, the number of input channels equals the number of output channels, and each output channel corresponds to a different input channel. The input channels and output channels are collectively referred to as depth channels. After depthwise convolution, pointwise convolution 793 is performed on the depth output tensor 780 and the 7×1×3 tensor 790 to produce OFM 760.
[0071] OFM 760 is then passed to the next layer in the sequence. In some embodiments, OFM 760 is passed through an activation function. An example activation function is the Rectified Linear Unit (ReLU). ReLU is a computation that directly returns the value provided as input, or returns 0 if the input is 0 or less. Convolutional layer 710 can receive several images as input and compute the convolution of each of them with each kernel. This process can be repeated several times. For example, OFM 760 is passed to a subsequent convolutional layer 710 (i.e., the convolutional layer 710 in the sequence that produces OFM 760). The subsequent convolutional layer 710 performs convolution on OFM 760 with a new kernel and generates a new feature map. The new feature map can also be normalized and resized. The new feature map can be kernelized again by further subsequent convolutional layers 710, and so on.
[0072] In some embodiments, the convolutional layer 710 has four hyperparameters: the number of kernels, the kernel size (e.g., the kernel size is F×F×D pixels), the stride S of dragging the window corresponding to the kernel on the image (e.g., a stride of 1 means moving the window one pixel at a time), and zero padding P (e.g., adding a black outline of P pixels thickness to the input image of the convolutional layer 710). The convolutional layer 710 can perform various types of convolutions, such as 2D convolution, dilated or dilated convolution, spatially separable convolution, depthwise separable convolution, transposed convolution, etc. The DNN 700 includes 76 convolutional layers 710. In other embodiments, the DNN 700 may include a different number of convolutional layers.
[0073] Pooling layer 720 downsamples the feature map generated by the convolutional layer, for example, by downsampling the presence of features in local blocks that summarize the feature map. Pooling layer 720 is positioned between two convolutional layers 710: a pre-convolutional layer 710 (the convolutional layer 710 preceding pooling layer 720 in the layer sequence) and a post-convolutional layer 710 (the convolutional layer 710 following pooling layer 720 in the layer sequence). In some embodiments, pooling layer 720 is added after convolutional layer 710, for example, after an activation function (e.g., ReLU, etc.) has been applied to OFM 760.
[0074] Pooling layer 720 receives feature maps generated by the preceding convolutional layer 710 and applies pooling operations to these feature maps. Pooling operations reduce the size of the feature maps while preserving their important characteristics. Therefore, pooling operations improve the efficiency of the DNN and avoid overlearning. Pooling layer 720 can perform pooling operations using average pooling (calculating the average value of each local block on the feature map), max pooling (calculating the maximum value of each local block on the feature map), or a combination of both. The size of the pooling operation is smaller than the size of the feature map. In various embodiments, the pooling operation is applied with a stride of 2×2 pixels, thus reducing the size of the feature map by a factor of 2, for example, reducing the number of pixels or values in the feature map to one-quarter of its original size. In one example, pooling layer 720 applied to a 6×6 feature map produces a 3×3 output pooled feature map. The output of pooling layer 720 is fed into the subsequent convolutional layer 710 for further feature extraction. In some embodiments, pooling layer 720 operates on each feature map separately to create a new set of the same number of pooled feature maps.
[0075] Fully connected layer 730 is the last layer of the DNN. Fully connected layer 730 may or may not be convolutional. Fully connected layer 730 may also be referred to as a linear layer. In some embodiments, fully connected layer 730 (e.g., a fully connected layer in DNN 700) may receive input operands. The input operands may define the outputs of convolutional layer 710 and pooling layer 720 and include the values of the final feature map generated by the last pooling layer 720 in the sequence. Fully connected layer 730 applies a linear combination and activation function to the input operands and generates a vector. Fully connected layer 730 may apply a linear transformation to the input operands via a weight matrix. This weight matrix may be the kernel of fully connected layer 730. The linear transformation may include tensor multiplication between the input operands and the weight matrix. The result of the linear transformation may be an output operand. In some embodiments, the fully connected layer may further apply a nonlinear transformation (e.g., by using a nonlinear activation function) to the result of the linear transformation to generate an output operand. This output operand may contain as many elements as there are categories: element i represents the probability that an image belongs to category i. Therefore, each element is between 0 and 7, and the sum of all elements is 7. These probabilities are calculated by the final fully connected layer 730 using either a logistic function (for binary classification) or a SoftMax function (for multi-class classification) as the activation function.
[0076] Example of an AI-based motion tracking environment Figure 8 An AI-based motion tracking environment 800 according to various embodiments is illustrated. The AI-based motion tracking environment 800 includes a motion tracking system 810, a client device 820 (solely referred to as client device 820), and a third-party system 830. In other embodiments, the AI-based motion tracking environment 800 may include fewer, more, or different components. For example, the AI-based motion tracking environment 800 may include different numbers of client devices 820 or more than one third-party system 830.
[0077] Motion tracking system 810 tracks the motion of an object in a local region. For example, motion tracking system 810 can track the motion of an object by estimating its 3D pose based on a depth image captured by a depth camera placed in the local region where the object is located. Motion tracking system 810 can receive depth images from a depth camera, one or more client devices 820, or a third-party system 830. Furthermore, motion tracking system 810 can send the estimated 3D pose information of the object to one or more client devices 820 or third-party systems 830. Additionally or alternatively, motion tracking system 810 can send content items generated using the estimated 3D pose of the object to one or more client devices 820 or third-party systems 830. An example of motion tracking system 810 is... Figure 1 The motion tracking system 100 in the middle.
[0078] Client device 820 communicates with motion tracking system 810. For example, client device 820 can receive a 3D pose graph representation from motion tracking system 810 and display the 3D pose graph representation to one or more users associated with client device 820. As another example, client device 820 can facilitate an interface connection with one or more depth cameras in a local area and can send commands to the depth cameras to capture depth images to be used by motion tracking system 810. Additionally or alternatively, client device 820 can facilitate an interface connection with one or more projectors in the local area and can provide content items to the projectors for the projectors to render the content items in the local area. Client device 820 can generate content items using motion tracking results from motion tracking system 810. The client device can have one or more users whose motion can be tracked by motion tracking system 810.
[0079] In some embodiments, client device 820 may execute one or more applications that allow one or more users of client device 820 to interact with motion tracking system 810. For example, client device 820 executes a browser application to enable interaction between client device 820 and motion tracking system 810. In another embodiment, client device 820 interacts with motion tracking system 810 through an application programming interface (API) running on client device 820's native operating system (such as iOS® or Android™).
[0080] Client device 820 may be one or more computing devices capable of receiving user input and sending and / or receiving data via network 840. In one embodiment, client device 820 is a conventional computer system, such as a desktop or laptop computer. Alternatively, client device 820 may be a device with computer functionality, such as a personal digital assistant (PDA), mobile phone, smartphone, autonomous vehicle, or other suitable device. Client device 820 is configured to communicate via network 840. In one embodiment, client device 820 is an integrated computing device operating as a standalone network support device. For example, client device 820 includes a display, speakers, microphone, camera, and input devices. In another embodiment, client device 820 is a computing device for coupling to an external media device, such as a television or other external display and / or audio output system. In this embodiment, client device 820 may be coupled to the external media device via a wireless or wired interface and may utilize various functions of the external media device, such as its display, speakers, microphone, camera, and input devices. Here, the client device 820 can be configured to be compatible with the following general external media devices: these general external media devices do not have dedicated software, firmware or hardware specifically for interacting with the client device 820.
[0081] The third-party system 830 is an online system capable of communicating with the motion tracking system 810 or at least one client device 820. In some embodiments, the third-party system 830 may provide data to the motion tracking system 810 for 3D pose estimation. This data may include depth images, data for training a DNN, data for validating a DNN, etc. The third-party system 830 may be a social media system, an online image library, an online search system, etc. Additionally or alternatively, the third-party system 830 may use the results of 3D pose estimation in various applications. For example, the third-party system 830 may use motion tracking results from the motion tracking system 810 for action recognition, motion analysis, virtual reality, augmented reality, film and game production, telepresence, etc.
[0082] The motion tracking system 810, client device 820, and third-party system 830 are connected via network 840. Network 840 may include any combination of local area networks (LANs) and / or wide area networks (WANs), using wired and / or wireless communication systems. In one embodiment, network 840 may use standard communication technologies and / or protocols. For example, network 840 may include communication links using technologies such as Ethernet, 8010.11, WiMAX, 3G, 4G, CDMA, and Digital Subscriber Line (DSL). Examples of network protocols used for communication via network 840 may include Multiprotocol Label Switching (MPLS), Transmission Control Protocol / Internet Protocol (TCP / IP), Hypertext Transfer Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), and File Transfer Protocol (FTP). Data exchanged via network 840 may be represented in any suitable format, such as Hypertext Markup Language (HTML) or Extensible Markup Language (XML). In some embodiments, all or part of the communication links of network 840 may be encrypted using any suitable technology or a combination of technologies.
[0083] Example motion tracking method Figure 9 This is a flowchart illustrating a motion tracking method 900 according to various embodiments. Method 900 can be used for 3D motion tracking. Method 900 can be... Figure 1 The motion tracking system 100 in the middle is executed. Although method 900 is referenced Figure 9 The flowchart shown illustrates this method, but many other motion tracking methods can be used alternatively. For example, the method can be changed... Figure 9 The execution order of the steps in the diagram. As another example, some steps can be changed, eliminated, or combined.
[0084] Motion tracking system 100 generates a point cloud of an object using one or more depth images capturing an object in a local region. In some embodiments, the one or more depth images are multiple depth images. Motion tracking system 100 converts the multiple depth images into multiple point clouds using depth pixels extracted from the multiple depth images. Motion tracking system 100 generates a fused point cloud using the multiple point clouds. Motion tracking system 100 generates the point cloud by reducing the total number of points in the fused point cloud to a predetermined number.
[0085] In some embodiments, one or more depth images are multiple depth images captured by multiple depth cameras in a local region. The positions of the multiple depth cameras in the local region are determined based on the shape of the local region. In some embodiments, the multiple depth cameras are synchronized and capture multiple depth images simultaneously.
[0086] The motion tracking system 100 extracts 920 feature maps from a point cloud using a first neural network. In some embodiments, the first neural network is a point-based neural network. In some embodiments, the first neural network is... Figure 1 The neural network 150 in the example.
[0087] The motion tracking system 100 generates a combined feature map 930 using the extracted feature map and the template feature map. The template feature map represents a template structure including joints and one or more connections between joints. In some embodiments, the motion tracking system 100 stitches together the extracted feature map and the template feature map. The dimension of the combined feature map is the sum of the dimensions of the extracted feature map and the template feature map.
[0088] In some embodiments, the combined feature map is a graph including nodes and one or more edges connecting the nodes. Nodes correspond to keypoints in a template structure encoded by the template feature map, and the second neural network is a graph convolutional network.
[0089] The motion tracking system 100 estimates the pose of an object in a local region 940 using a second neural network combined with feature maps. In some embodiments, the first neural network is a graph convolutional network. In some embodiments, the second neural network is... Figure 1 The motion tracking system 100 uses a neural network 170. In some embodiments, the motion tracking system 100 estimates the pose of an object by determining the positions of joints in a local region using a second neural network. In some embodiments, the motion tracking system 100 generates content items based on the estimated object pose and provides the content items for display in the local region.
[0090] Example computing device Figure 10 This is a block diagram of an example computing device 1000 according to various embodiments. In some embodiments, the computing device 1000 may be used as... Figure 1 At least a portion of the motion tracking system 100. The computing device 1000 may be... Figure 8 Client device 820 or Figure 8 An example of a third-party system 830. Multiple components in... Figure 10 The components are shown as being included in computing device 1000, but any one or more of these components may be omitted or copied to suit the application. In some embodiments, some or all of the components included in computing device 1000 may be attached to one or more motherboards. In some embodiments, some or all of these components are manufactured on a single system-on-a-chip (SoC) die. Furthermore, in various embodiments, computing device 1000 may not include... Figure 10The computing device 1000 may include one or more of the components shown, but may include interface circuitry for coupling to said one or more components. For example, the computing device 1000 may not include display device 1006, but may include display device interface circuitry (e.g., connector and driver circuitry) to which display device 1006 may be coupled. In another set of examples, the computing device 1000 may not include audio input device 1018 or audio output device 1008, but may include audio input or output device interface circuitry (e.g., connector and support circuitry) to which audio input device 1018 or audio output device 1008 may be coupled.
[0091] Computing device 1000 may include processing device 1002 (e.g., one or more processing devices). Processing device 1002 processes electronic data from registers and / or memory to convert the electronic data into other electronic data that can be stored in registers and / or memory. Computing device 1000 may include memory 1004, which itself may include one or more memory devices, such as volatile memory (e.g., DRAM), non-volatile memory (e.g., read-only memory (ROM)), high-bandwidth memory (HBM), flash memory, solid-state memory, and / or hard disk drive. In some embodiments, memory 1004 may include memory sharing a die with processing device 1002. In some embodiments, memory 1004 includes one or more non-transitory computer-readable media storing instructions executable to perform operations for performing convolution (e.g., in conjunction with the above). Figure 9 The described method 900) or some operations performed by the motion tracking system 100. Instructions stored in one or more non-transitory computer-readable media can be executed by the processing device 1002.
[0092] In some embodiments, computing device 1000 may include communication chip 1012 (e.g., one or more communication chips). For example, communication chip 1012 may be configured to manage wireless communication for transmitting data to and from computing device 1000. The term "wireless" and its derivatives can be used to describe circuits, devices, systems, methods, technologies, communication channels, etc., which can transmit data through a non-solid medium using modulated electromagnetic radiation. This term does not imply that the associated device does not contain any wires, however, in some embodiments they may not.
[0093] The 1012 communication chip can implement any of many wireless standards or protocols, including but not limited to Institute of Electrical and Electronics Engineers (IEEE) standards, such as Wi-Fi (IEEE 802.10 series), IEEE 802.16 standards (e.g., IEEE 802.16-2005 amendments), the Long Term Evolution (LTE) project, and any amendments, updates, and / or revisions (e.g., improved LTE projects, Ultra Mobile Broadband (UMB) projects (also known as "3GPP2"), etc.). Broadband Wireless Access (BWA) networks compatible with IEEE 802.16 are often referred to as WiMAX networks, an abbreviation for Global Microwave Access Interoperability, which is a certification mark for products that have passed conformance and interoperability testing of the IEEE 802.16 standard. The communication chip 1012 can operate according to Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High SMAC-enhanced Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE networks. The communication chip 1012 can also operate according to Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). The communication chip 1012 can operate according to Code-Division Multiple Access (CDMA), Time-Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunication (DECT), Evolution-Data Optimized (EV-DO) and its derivatives, as well as any other wireless protocol specified as 3G, 4G, 5G, etc. In other embodiments, the communication chip 1012 can operate according to other wireless protocols.The computing device 1000 may include an antenna 1022 to facilitate wireless communication and / or receive other wireless communications (e.g., AM or FM radio transmissions).
[0094] In some embodiments, the communication chip 1012 can manage wired communications such as electrical, optical, or any other suitable communication protocol (e.g., Ethernet). As described above, the communication chip 1012 may include multiple communication chips. For example, a first communication chip 1012 may be dedicated to short-range wireless communications such as Wi-Fi or Bluetooth, and a second communication chip 1012 may be dedicated to long-range wireless communications such as GPS, EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some embodiments, the first communication chip 1012 may be dedicated to wireless communications, and the second communication chip 1012 may be dedicated to wired communications.
[0095] The computing device 1000 may include a battery / power circuit 1014. The battery / power circuit 1014 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 1000 to a power source (e.g., AC line power) that is separate from the computing device 1000.
[0096] The computing device 1000 may include a display device 1006 (or the corresponding interface circuitry described above). For example, the display device 1006 may include any visual indicator, such as a head-up display, computer monitor, projector, touch screen display, liquid crystal display (LCD), light-emitting diode display, or flat panel display.
[0097] The computing device 1000 may include an audio output device 1008 (or the corresponding interface circuitry described above). For example, the audio output device 1008 may include any device that generates audible indicators, such as a speaker, headphones, or earphones.
[0098] The computing device 1000 may include an audio input device 1018 (or the corresponding interface circuitry described above). The audio input device 1018 may include any device that generates a signal representing sound, such as a microphone, microphone array, or digital musical instrument (e.g., a musical instrument with a Musical Instrument Digital Interface (MIDI) output).
[0099] The computing device 1000 may include a GPS device 1016 (or a corresponding interface circuit as described above). As is known in the art, the GPS device 1016 can communicate with a satellite-based system and can receive the location of the computing device 1000.
[0100] The computing device 1000 may include other output devices 1010 (or corresponding interface circuits as described above). Examples of other output devices 1010 may include audio codecs, video codecs, printers, wired or wireless transmitters for providing information to other devices, or additional storage devices.
[0101] The computing device 1000 may include other input devices 1020 (or corresponding interface circuits as described above). Examples of other input devices 1020 may include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a barcode reader, a Quick Response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.
[0102] The computing device 1000 can have any desired form factor, such as a handheld or mobile computer system (e.g., a mobile phone, smartphone, mobile internet device, music player, tablet computer, laptop computer, netbook computer, ultrabook computer, PDA, ultraportable personal computer, etc.), a desktop computer system, a server or other networked computing component, a printer, scanner, monitor, set-top box, entertainment control unit, vehicle control unit, digital camera, digital video recorder, or wearable computer system. In some embodiments, the computing device 1000 can be any other electronic device that processes data.
[0103] Select Example The following paragraphs provide various examples of the embodiments disclosed herein.
[0104] Example 1 provides a method comprising: generating a point cloud of an object using one or more depth images of an object captured in a local region; extracting a feature map from the point cloud using a first neural network; generating a combined feature map using the extracted feature map and a template feature map, the template feature map representing a template structure including joints and one or more connections between the joints; and estimating the pose of the object in the local region using the combined feature map using a second neural network.
[0105] Example 2 provides the method of Example 1, wherein the one or more depth images are multiple depth images, and generating a point cloud of the object includes: converting the multiple depth images into multiple point clouds using depth pixels extracted from the multiple depth images; generating a fused point cloud using the multiple point clouds; and generating the point cloud by reducing the total number of points in the fused point cloud to a predetermined number.
[0106] Example 3 provides the method of Example 1 or 2, wherein generating the combined feature map includes: concatenating the extracted features and the template feature map, wherein the dimension of the combined feature map is the sum of the dimensions of the extracted feature map and the template feature map.
[0107] Example 4 provides a method of any one of Examples 1-3, wherein the combined feature map is a graph including nodes and one or more edges connecting the nodes, the nodes corresponding to key points in the template structure encoded by the template feature map, and the second neural network is a graph convolutional network.
[0108] Example 5 provides a method of any one of Examples 1-4, wherein estimating the pose of the object in the local region by means of the second neural network includes: determining the position of the joint in the local region by means of the second neural network.
[0109] Example 6 provides a method of any one of Examples 1-5, wherein the one or more depth images are multiple depth images captured by multiple depth cameras in the local region, and the positions of the multiple depth cameras in the local region are determined based on the shape of the local region.
[0110] Example 7 provides a method of any one of Examples 1-6, further comprising: generating a content item based on an estimated pose of the object; and providing the content item for display in the local area.
[0111] Example 8 provides one or more non-transitory computer-readable media storing instructions executable to perform operations including: generating a point cloud of an object using one or more depth images of an object captured in a local region; extracting a feature map from the point cloud via a first neural network; generating a combined feature map using the extracted feature map and a template feature map, the template feature map representing a template structure including joints and one or more connections between the joints; and estimating the pose of the object in the local region using the combined feature map via a second neural network.
[0112] Example 9 provides one or more non-transitory computer-readable media of Example 8, wherein the one or more depth images are multiple depth images, and generating a point cloud of the object includes: converting the multiple depth images into multiple point clouds using depth pixels extracted from the multiple depth images; generating a fused point cloud using the multiple point clouds; and generating the point cloud by reducing the total number of points in the fused point cloud to a predetermined number.
[0113] Example 10 provides one or more non-transitory computer-readable media of Example 8 or 9, wherein generating the combined feature map includes concatenating the extracted features and the template feature map, wherein the dimension of the combined feature map is the sum of the dimensions of the extracted feature map and the template feature map.
[0114] Example 11 provides one or more non-transitory computer-readable media of any of Examples 8-10, wherein the combined feature map is a graph including nodes and one or more edges connecting the nodes, the nodes corresponding to key points in the template structure encoded by the template feature map, and the second neural network is a graph convolutional network.
[0115] Example 12 provides one or more non-transitory computer-readable media of any of Examples 8-11, wherein estimating the pose of the object in the local region by means of the second neural network includes: determining the position of the joint in the local region by means of the second neural network.
[0116] Example 13 provides one or more non-transitory computer-readable media of any of Examples 8-12, wherein the one or more depth images are multiple depth images captured by multiple depth cameras in the local region, and the positions of the multiple depth cameras in the local region are determined based on the shape of the local region.
[0117] Example 14 provides one or more non-transitory computer-readable media of any of Examples 8-13, wherein the operation further includes: generating a content item based on an estimated pose of the object; and providing the content item for display in the local area.
[0118] Example 15 provides an apparatus comprising: a computer processor for executing computer program instructions; and one or more non-transitory computer-readable media storing executable instructions for performing operations, the operations including: generating a point cloud of an object using one or more depth images of an object captured in a local region; extracting feature maps from the point cloud via a first neural network; generating a combined feature map using the extracted feature maps and a template feature map, the template feature map representing a template structure including joints and one or more connections between the joints; and estimating the pose of the object in the local region using the combined feature map via a second neural network.
[0119] Example 16 provides the apparatus of Example 15, wherein the one or more depth images are multiple depth images, and generating a point cloud of the object includes: converting the multiple depth images into multiple point clouds using depth pixels extracted from the multiple depth images; generating a fused point cloud using the multiple point clouds; and generating the point cloud by reducing the total number of points in the fused point cloud to a predetermined number.
[0120] Example 17 provides an apparatus similar to that of Example 15 or 16, wherein generating the combined feature map includes: concatenating the extracted features and the template feature map, wherein the dimension of the combined feature map is the sum of the dimensions of the extracted feature map and the template feature map.
[0121] Example 18 provides an apparatus of any of Examples 15-17, wherein the combined feature map is a graph including nodes and one or more edges connecting the nodes, the nodes corresponding to joints in the template structure encoded by the template feature map, and the second neural network is a graph convolutional network.
[0122] Example 19 provides an apparatus of any one of Examples 15-18, wherein estimating the pose of the object in the local region by means of the second neural network includes: determining the position of the joint in the local region by means of the second neural network.
[0123] Example 20 provides an apparatus of any one of Examples 15-19, wherein the one or more depth images are multiple depth images captured by multiple depth cameras in the local region, and the positions of the multiple depth cameras in the local region are determined based on the shape of the local region.
[0124] The foregoing description of the implementations of this disclosure (including those described in the abstract) is not intended to be exhaustive or to limit this disclosure to its exact form. While specific implementations and examples of this disclosure have been described herein for illustrative purposes, various equivalent modifications can be made within the scope of this disclosure, as will be recognized by those skilled in the art. These modifications can be made to this disclosure based on the foregoing detailed description.
Claims
1. A method comprising: A point cloud of the object is generated using one or more depth images of the object captured in a local region; Feature maps are extracted from the point cloud using a first neural network; A combined feature map is generated using the extracted feature map and the template feature map, wherein the template feature map represents a template structure including joints and one or more connections between the joints; as well as The pose of the object in the local region is estimated using the combined feature map via a second neural network.
2. The method according to claim 1, wherein, The one or more depth images are multiple depth images, and generating the point cloud of the object includes: The multiple depth images are converted into multiple point clouds using depth pixels extracted from the multiple depth images; Generate a fused point cloud using the multiple point clouds; and The point cloud is generated by reducing the total number of points in the fused point cloud to a predetermined number.
3. The method according to claim 1, wherein, Generating the combined feature map includes: By concatenating the extracted features and the template feature map, The dimension of the combined feature map is the sum of the dimension of the extracted feature map and the dimension of the template feature map.
4. The method according to claim 1, wherein, The combined feature map is a graph including nodes and one or more edges connecting the nodes, the nodes corresponding to joints in the template structure encoded by the template feature map, and the second neural network is a graph convolutional network.
5. The method according to claim 1, wherein, Estimating the pose of the object in the local region using the second neural network includes: The location of the joint in the local region is determined by the second neural network.
6. The method of claim 1, wherein the one or more depth images are multiple depth images captured by multiple depth cameras in the local region, and the positions of the multiple depth cameras in the local region are determined based on the shape of the local region.
7. The method according to claim 1, further comprising: Content items are generated based on the estimated pose of the object; as well as Provide the content items to be displayed in the local area.
8. One or more non-transitory computer-readable media storing instructions executable to perform operations, said operations including: A point cloud of the object is generated using one or more depth images of the object captured in a local region; Feature maps are extracted from the point cloud using a first neural network; A combined feature map is generated using the extracted feature map and the template feature map, wherein the template feature map represents a template structure including joints and one or more connections between the joints; as well as The pose of the object in the local region is estimated using the combined feature map via a second neural network.
9. One or more non-transitory computer-readable media according to claim 8, wherein, The one or more depth images are multiple depth images, and generating the point cloud of the object includes: The multiple depth images are converted into multiple point clouds using depth pixels extracted from the multiple depth images; Generate a fused point cloud using the multiple point clouds; and The point cloud is generated by reducing the total number of points in the fused point cloud to a predetermined number.
10. One or more non-transitory computer-readable media according to claim 8, wherein, Generating the combined feature map includes: By concatenating the extracted features and the template feature map, The dimension of the combined feature map is the sum of the dimension of the extracted feature map and the dimension of the template feature map.
11. One or more non-transitory computer-readable media according to claim 8, wherein, The combined feature map is a graph including nodes and one or more edges connecting the nodes, the nodes corresponding to joints in the template structure encoded by the template feature map, and the second neural network is a graph convolutional network.
12. One or more non-transitory computer-readable media according to claim 8, wherein, Estimating the pose of the object in the local region using the second neural network includes: The location of the joint in the local region is determined by the second neural network.
13. One or more non-transitory computer-readable media according to claim 8, wherein, The one or more depth images are multiple depth images captured by multiple depth cameras in the local region, and the positions of the multiple depth cameras in the local region are determined based on the shape of the local region.
14. The one or more non-transitory computer-readable media of claim 8, wherein the operation further comprises: Content items are generated based on the estimated pose of the object; as well as Provide the content items to be displayed in the local area.
15. An apparatus comprising: A computer processor is used to execute computer program instructions; as well as One or more non-transitory computer-readable media storing executable instructions to perform operations, said operations including: A point cloud of the object is generated using one or more depth images of the object captured in a local region. Feature maps are extracted from the point cloud using a first neural network. A combined feature map is generated using the extracted feature map and a template feature map, wherein the template feature map represents a template structure including joints and one or more connections between the joints, and The pose of the object in the local region is estimated using the combined feature map via a second neural network.
16. The apparatus according to claim 15, wherein, The one or more depth images are multiple depth images, and generating the point cloud of the object includes: The multiple depth images are converted into multiple point clouds using depth pixels extracted from the multiple depth images; Generate a fused point cloud using the multiple point clouds; and The point cloud is generated by reducing the total number of points in the fused point cloud to a predetermined number.
17. The apparatus according to claim 15, wherein, Generating the combined feature map includes: By concatenating the extracted features and the template feature map, The dimension of the combined feature map is the sum of the dimension of the extracted feature map and the dimension of the template feature map.
18. The apparatus according to claim 15, wherein, The combined feature map is a graph including nodes and one or more edges connecting the nodes, the nodes corresponding to joints in the template structure encoded by the template feature map, and the second neural network is a graph convolutional network.
19. The apparatus according to claim 15, wherein, Estimating the pose of the object in the local region using the second neural network includes: The location of the joint in the local region is determined by the second neural network.
20. The apparatus according to claim 15, wherein, The one or more depth images are multiple depth images captured by multiple depth cameras in the local region, and the positions of the multiple depth cameras in the local region are determined based on the shape of the local region.