Method for capturing and optimizing human motion in a scene and motion capture system
Patent Information
- Application Number
- CN202280006556.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-02-25
- Filing Date
- 2022-03-03
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-03-03
AI Technical Summary
然而,HPS需要预先构建的地图和大型图像数据库进行自我定位,在捕捉大型场景时,这些条件并不理想
[0026]在一些实施例中,基于所述人体的优化动作和所述优化的场景图,可以创建元宇宙。
Smart Images

Figure CN117178298B_ABST
Abstract
Description
[0001] Related patent applications
[0002] This application is a continuation application of International Patent Application No. PCT / CN2022 / 079151, filed on March 3, 2022, which claims the benefit and priority of International Patent Application No. PCT / CN2022 / 078083, filed on February 25, 2022, entitled "System and Method for Capturing Three-Dimensional Human Motion Using LiDAR". The foregoing references are incorporated herein by reference in their entirety. Technical Field
[0003] This application relates to the field of motion capture technology, and more specifically, to a method and motion capture system for capturing and optimizing human movements in a scene. Background Technology
[0004] In the digital world, technologies such as augmented reality, virtual reality, smart cities, robotics, and autonomous driving have enriched people's lives, making them of significant importance. Humans and the environment are two major components in creating the digital world. Current research tends to separate dynamic human movements from static environments to improve the accuracy of both. Inertial measurement unit (IMU) sensors are widely used to capture human movements, mounted on different parts of the body, such as arms, legs, feet, and head, to capture precise short-term actions. However, sensor drift occurs as the acquisition time increases. Conventional methods often use external cameras as a remedy to improve accuracy, but these methods may limit the capture space, human activity, and interaction. For example, Human Positioning and Attitude System (HPS) uses a head-mounted camera that looks outward like the human eye, supplementing the IMU sensors in global positioning. Without the constraints of an external camera, HPS can recover full-body posture and register the HPS wearer with large 3D scans of real-world scenes. However, HPS requires pre-built maps and large image databases for self-localization, conditions that are not ideal when capturing large scenes. Therefore, conventional methods are not well-suited for capturing scenes from large-scale spaces. Summary of the Invention
[0005] The embodiments of this application provide a method and motion capture system for capturing and optimizing human movements in a scene, which is at least to some extent suitable for capturing scenes in large-scale spaces.
[0006] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part by practice of this application.
[0007] This invention describes a method for capturing human motion in a scene. Multiple IMU devices and one LiDAR sensor can be mounted on the human body. The IMU devices capture IMU data, and the LiDAR sensor captures LiDAR data. Human motion can be estimated based on the IMU data and the LiDAR data. A 3D scene map can be constructed based on the LiDAR data. Optimization can be performed to obtain optimized human motion and an optimized scene map.
[0008] In some embodiments, the self-motion of the LiDAR sensor can be estimated. The LiDAR sensor may be mounted on the hip of a human body.
[0009] In some embodiments, the LiDAR data and IMU data can be calibrated based on the self-motion of the LiDAR sensor.
[0010] In some embodiments, human jumping can be performed during the step of obtaining IMU data captured by the IMU device and LiDAR data captured by the LiDAR sensor. The LiDAR data and IMU data can be synchronized based on the peak values derived from the LiDAR data and the peak values derived from the IMU data.
[0011] In some embodiments, graph-based optimization may be performed to fuse LiDAR and IMU trajectories. The LiDAR trajectory may include the movement of the human body center derived from the LiDAR data. The IMU trajectory may include the movement of the human body center derived from the IMU data.
[0012] In some embodiments, the optimization may be based on contact constraints and sliding constraints.
[0013] In some embodiments, the contact constraint can define the contact loss as the distance from a body part (derived from human motion) to the nearest surface (derived from the 3D scene graph). The sliding constraint can define the sliding loss as the distance between two consecutive body parts derived from human motion. Using a gradient descent algorithm, the sum of the contact loss and sliding loss can be minimized to iteratively optimize human motion.
[0014] In some embodiments, the body part may be the feet of a human body. The surface may be the ground.
[0015] In some embodiments, multiple second IMU devices and one second LiDAR sensor may be mounted on a second human body. Second IMU data captured by the second IMU devices and second LiDAR data captured by the second LiDAR sensor can be obtained. The movements of the second human body can be estimated based on the second IMU data. A three-dimensional second scene map can be constructed based on the second LiDAR data. The three-dimensional scene map and the three-dimensional second scene map can be fused to obtain a combined scene map.
[0016] In some embodiments, a metaverse can be created based on the optimized movements of the human body and the optimized scene graph.
[0017] This invention introduces a motion capture system. The system may include multiple wearable IMU devices, a LiDAR sensor that can be mounted on the human body, a processor, and a memory storing instructions. When the processor executes the instructions, the instructions cause the system to perform a human motion capture method in a scene. The method includes: acquiring IMU data captured by the IMU devices and LiDAR data captured by the LiDAR sensor.
[0018] In some embodiments, an L-shaped bracket can mount the LiDAR sensor on the hip of a human body. The LiDAR sensor and the IMU device may have a fundamentally rigid transformation.
[0019] In some embodiments, a wireless receiver may be coupled to the system. The wireless receiver may receive IMU data captured by the IMU device.
[0020] In some embodiments, human motion can be estimated based on the IMU data and LiDAR data. A 3D scene map can be constructed based on the LiDAR data. Optimization can be performed to obtain optimized human motion and an optimized scene map.
[0021] In some embodiments, graph-based optimization may be performed to fuse LiDAR and IMU trajectories. The LiDAR trajectory may include the movement of the human body center derived from the LiDAR data. The IMU trajectory may include the movement of the human body center derived from the IMU data.
[0022] In some embodiments, the optimization may be based on contact constraints and sliding constraints.
[0023] In some embodiments, the contact constraint can define the contact loss as the distance from a body part (derived from human motion) to the nearest surface (derived from the 3D scene graph). The sliding constraint can define the sliding loss as the distance between two consecutive body parts derived from human motion. Using a gradient descent algorithm, the sum of the contact loss and sliding loss can be minimized to iteratively optimize human motion.
[0024] In some embodiments, the body part may be the feet of a human body. The surface may be the ground.
[0025] In some embodiments, multiple second IMU devices may be worn by a second human body. A second LiDAR sensor may be mounted on the second human body. Second IMU data captured by the second IMU devices and second LiDAR data captured by the second LiDAR sensor can be obtained. The movements of the second human body can be estimated based on the second IMU data. A three-dimensional second scene map can be constructed based on the second LiDAR data. The three-dimensional scene map and the three-dimensional second scene map can be fused to obtain a combined scene map. Optimization can be performed to obtain optimized movements of the human body and the second human body in the optimized combined scene map.
[0026] In some embodiments, a metaverse can be created based on the optimized movements of the human body and the optimized scene graph.
[0027] This invention introduces a method for optimizing human motion in a scene. It can obtain a 3D scene graph and human motion. Graph-based optimization can be performed to fuse LiDAR and IMU trajectories. Joint optimization based on multiple physical constraints can be performed to obtain optimized human motion and an optimized scene graph.
[0028] In some embodiments, the three-dimensional scene map can be obtained from LiDAR data captured by a LiDAR sensor mounted on the human body. Human motion can be obtained from IMU data captured by multiple IMU devices mounted on the human body.
[0029] In some embodiments, the 3D scene graph and human motion can be calibrated.
[0030] In some embodiments, the 3D scene graph and human motion can be synchronized.
[0031] In some embodiments, the LiDAR trajectory may include the movement of the human body center derived from LiDAR data. The IMU trajectory may include the movement of the human body center derived from IMU data.
[0032] In some embodiments, the joint optimization may be based on contact constraints and sliding constraints.
[0033] In some embodiments, the contact constraint can define the contact loss as the distance from a body part (derived from human motion) to the nearest surface (derived from the 3D scene graph). The sliding constraint can define the sliding loss as the distance between two consecutive body parts derived from human motion. Using a gradient descent algorithm, the sum of the contact loss and sliding loss can be minimized to iteratively optimize human motion.
[0034] In some embodiments, the body part may be the feet of a human body, and the surface may be the ground.
[0035] In some embodiments, multiple second IMU devices and a second LiDAR sensor can be mounted on a second human body. Second IMU data captured by the second IMU devices and second LiDAR data captured by the second LiDAR sensor can be obtained. The movements of the second human body can be estimated based on the second IMU data. A three-dimensional second scene map can be constructed based on the second LiDAR data. The three-dimensional scene map and the three-dimensional second scene map can be fused to obtain a combined scene map. Optimization can be performed to obtain optimized movements of the human body and the second human body in the optimized combined scene map.
[0036] In some embodiments, a metaverse can be created based on the optimized movements of the human body and the optimized scene graph.
[0037] These and other features of the devices, systems, methods, and non-transitory computer-readable media disclosed in this invention, as well as the operational methods and functions of structurally related components and parts combinations, and the economics of manufacture, will become more apparent after considering the following description and appended claims, and referring to the accompanying drawings, all of which form part of this specification, wherein similar reference numerals denote corresponding parts in the various drawings. However, it is to be clear that the drawings are for illustrative and descriptive purposes only and are not intended to define any limitations on the invention.
[0038] According to the technical solution of the embodiments of this application, by installing multiple IMU devices and a LiDAR sensor on the human body, corresponding IMU data and LiDAR data are obtained, so as to estimate human movements based on the IMU data, construct a three-dimensional scene map based on the LiDAR data, and obtain optimized human movements and scene maps through optimization, thereby making it suitable for capturing scenes in large-scale spaces.
[0039] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0040] Certain features of various embodiments of the present invention are specifically set forth in the appended claims. The features and advantages of the present invention will be better understood by referring to the following detailed description, which illustrates illustrative embodiments based on the principles of the invention, with reference to the accompanying drawings:
[0041] Figure 1 Various human movements and large-scale indoor and outdoor scenes captured using a motion and scene capture system according to various embodiments of the present disclosure are shown.
[0042] Figure 2 An action and scene capture system according to various embodiments of the present disclosure is shown.
[0043] Figure 3A The data processing pipeline for determining pose estimation and LiDAR mapping based on data captured by motion and scene capture systems according to various embodiments of the present disclosure is illustrated.
[0044] Figure 3B A comparison of IMU trajectories and LiDAR trajectories according to various embodiments of this disclosure is shown.
[0045] Figure 3C An example HSC4D dataset is shown according to various embodiments of this disclosure.
[0046] Figure 3D-3E Various comparison tables are shown that compare the HSC4D dataset with other datasets according to various embodiments of this disclosure.
[0047] Figure 3F A graphical comparison of the HSC4D dataset described according to various embodiments of the present disclosure with baselines 1 and 2 is shown.
[0048] Figure 4A A method for capturing human motion in a scene is illustrated according to various embodiments of the present disclosure.
[0049] Figure 4B A computing component according to various embodiments of the present disclosure is illustrated, the computing component including one or more hardware processors and a machine-readable storage medium (stored a set of machine-readable / machine-executable instructions), which, when executed, cause the hardware processor to perform a method for optimizing human motion in a scene.
[0050] Figure 5 A block diagram is shown of a computer system that can implement the various embodiments described in this invention.
[0051] The accompanying drawings illustrate various embodiments of the disclosed technology, but are for illustrative purposes only, wherein the same reference numerals are used to identify the same components. Based on the following discussion, those skilled in the art will readily recognize that alternative embodiments of the structures and methods shown in the drawings may be employed without departing from the principles of the disclosed technology. Detailed Implementation
[0052] Light detection and ranging (LiDAR) sensors are best suited for precise localization and mapping. LiDAR sensors are a commonly used sensor type and are widely applied in mobile robotics and autonomous vehicles. They are also widely used for large-scale scene capture. While many LiDAR capture datasets exist, including indoor and large outdoor scenes, these datasets generally focus on scene understanding and 3D perception, neglecting accurate human pose. For example, the PedX dataset provides 3D poses of pedestrians using skinned multi-person linear model (SMPL) parameterization of joint positions of instances (e.g., objects, humans) in third-person view images. The 3D poses from the PedX dataset are not as accurate as those measured by IMU sensors. Furthermore, the PedX dataset focuses on traffic scenes and is therefore not suitable for generating diverse 3D human movements.
[0053] The invention addresses the aforementioned problems. In various embodiments, the invention may include a motion and scene capture system. This system may include multiple IMU sensors and a LiDAR sensor. The multiple IMU sensors and the LiDAR sensor can be worn by a human body to capture human motion while providing localization and scene capture. Utilizing IMU-based motion capture and LiDAR-based localization and scene capture, a dataset, namely a human-centered 4D scene capture (HSC4D) dataset, can be generated to accurately and efficiently create dynamic digital worlds with continuous human motion in indoor and outdoor scenes. The HSC4D dataset is not limited or dependent on specific spaces, poses, or interactions between people if only human-mounted sensors are used. Furthermore, the HSC4D dataset can capture most real-world scenes involving the human body. Compared to camera-based localization, LiDAR-based localization is more accurate, achieving global localization and significantly reducing IMU sensor drift. Moreover, unlike camera-based localization, LiDAR-based localization does not require pre-built maps. Additionally, the IMU sensor can improve the accuracy of local trajectories captured by LiDAR, where errors are primarily caused by body tremors. Therefore, considering several physical constraints, joint optimization can be performed using the IMU sensor and the LiDAR sensor to improve motion estimation and human-scene mapping. The invention described herein uses only human-mounted IMU and LiDAR sensors. Therefore, unlike conventional methods, capturing human motion and large-scale scenes is not limited by device constraints and any pre-built maps, thus enabling long-term human motion capture. In some cases, motion capture time can last for more than an hour, depending on the battery power of the human-mounted IMU and LiDAR sensors. In some embodiments, the LiDAR sensor can be designed as a backpack or handheld device for human object localization. Since LiDAR-based positioning systems are typically not portable, thus affecting human motion capture, a lightweight LiDAR sensor rigidly connected to the human body and mounted on the hip is designed. This allows for human self-localization in both large indoor and outdoor scenes. In some embodiments, to make the motion and scene capture system lighter and enable wireless communication, this invention discloses a method for localization and mapping in a scene using only LiDAR. In such embodiments, the joint optimization of scene and IMU pose can further improve LiDAR mapping results. Based on the invention disclosed herein, a metaverse can be created based on optimized human motion and optimized scene. These and other features of the invention are further described in detail herein.
[0054] Figure 1Various human movements and large-scale indoor and outdoor scenes captured using a motion and scene capture system according to various embodiments of the present disclosure are illustrated. According to the various embodiments, a dataset (i.e., HSC4D)100 containing large scenes, ranging from 1 km to 5 km in size, can be generated using the IMU sensor and LiDAR sensor of the motion and scene capture system. 2 Between them, there are precise dynamic human body movements 102 and positions 104. For example, such as Figure 1 As shown, the dataset includes diverse scenarios, such as a climbing gym 104a, a multi-story building 104b, an indoor staircase 104c, etc., as well as challenging human activities, such as exercise 102a, climbing stairs 102b, rock climbing 102c, etc. Figure 1 It can be seen that it can effectively capture precise human posture and natural interactions between human activities (e.g., 102a-102c) and the environment (e.g., 104a-104c). Furthermore, the effectiveness and generalization ability of the dataset are demonstrated.
[0055] In large, unknown scenes, the motion and scene capture system can utilize LiDAR sensors (such as 3D rotating LiDAR) and IMU sensors to estimate 3D human motion and construct a map of the large, unknown scene. The estimated 3D human motion can include data related to local 3D pose and global localization. The N frames of the estimated 3D human motion can generally be represented as M = (T, θ, β), where T is an N×3 translation parameter, θ is an N×24×3 pose parameter, and β is an N×10 shape parameter. During motion capture, β can be assumed to be constant. The 3D point cloud scene captured by the LiDAR sensor can be represented as S. The subscript k, k∈Z + This represents the exponent of the point cloud frame. The skinned multi-person linear (SMPL) body model Φ(·) can be used to represent M... k Mapping to human body mesh model V k V k ∈R 6890×3Generally, data captured using LiDAR and IMU sensors can be stored in three coordinate systems: the IMU coordinate system, the LiDAR coordinate system, and the global coordinate system. In the IMU coordinate system {I}, the origin is set at the hip joint of the SMPL model representing the human body, with the X / Y / Z axes pointing to the right, top, and front directions of the human body, respectively. In the LiDAR coordinate system {L}, the origin is set at the center of the LiDAR sensor, with the X / Y / Z axes pointing to the right, front, and top directions of the LiDAR sensor, respectively. In the global coordinate system {W}, the coordinates are set to be the same as the coordinates of the first point cloud frame captured by the LiDAR sensor. Generally, the task to be performed by the motion and scene capture system can be defined as follows: Let {L} (i.e., the LiDAR coordinate system) contain the LiDAR scan sequence (i.e., point cloud frames). k∈Z + 3D human motion sequences in {I} (i.e., IMU coordinate system) Calculate human motion in {W} (i.e., the global coordinate system). And in The 3D scene S is constructed below.
[0056] Figure 2 A motion and scene capture system 200 according to various embodiments of the present disclosure is shown. For example... Figure 2 As shown, the motion and scene capture system 200 may include multiple IMU sensors 202 and one LiDAR sensor 204. The multiple IMU sensors 202 and the LiDAR sensor 202 can be mounted or worn on the human body. In some embodiments, the multiple IMU sensors 202 may be embedded in clothing worn by the human body. In other embodiments, the multiple IMU sensors 202 may be embedded in a wearable strip, such as... Figure 2 The IMU sensor shown measures human knee movement. In some embodiments, the LiDAR sensor 204 can be mounted near the hip of the human body, and the human body can wear a dedicated plate 206. In some embodiments, the LiDAR sensor 204 can be a 64-line Ouster LiDAR sensor, and multiple IMU sensors 202 can be Noitom's PN Studio inertial motion capture product. The LiDAR sensor 204 can acquire 3D point clouds. Multiple IMU sensors 202 can capture human movements. like Figure 2As shown, in some embodiments, the plurality of IMU sensors may include 17 IMU sensors connected to or associated with various limbs of the human body. In some embodiments, the plurality of IMU sensors may be associated with a receiver. The receiver can acquire human motion data captured by the plurality of IMU sensors 202. In some embodiments, the receiver may be physically coupled to the plurality of IMU sensors 202. For example, the receiver may be connected to each of the plurality of IMU sensors 202 via a data cable. In some embodiments, the receiver may be wirelessly coupled to the plurality of IMU sensors 202. For example, the receiver may wirelessly communicate with the plurality of IMU sensors 202 via Wi-Fi, Bluetooth, near-field communication, or other suitable types of wireless communication technologies. In some embodiments, to enable the LiDAR sensor 204 to function properly in a wireless state, it may be connected to a computing system 208, such as a DJI Manifold 2-C mini-computer. In some embodiments, the computing system 208 may be portable and powered by a 24V mobile energy storage device such as a battery, thereby providing power to the LiDAR sensor 204 and the computing system 208. In some embodiments, to ensure the motion and scene capture system is lightweight and highly accurate, the wiring of the LiDAR sensor 204 can be modified and routed to a dedicated board 206. In some embodiments, the dedicated board can be a custom-designed 10cm×10cm L-shaped bracket for mounting the LiDAR sensor 204. In some embodiments, the battery and computing system 208 can be placed in a small encapsulated item 210 (e.g., a bag, pouch, etc.) carried on the back of the user. The LiDAR sensor 204 can be carried close to the hip of the user, so that the origins of {I} and {L} are as close as possible. In this way, the point cloud data captured by the LiDAR sensor 204 and the human motion data captured by multiple IMU sensors 202 can be transformed into the same coordinate system with near-rigidity. In some embodiments, such as Figure 2As shown, the LiDAR sensor 204 can have a 360° horizontal field of view and a 45° vertical field of view. However, in this type of embodiment, the horizon field of view is reduced from 360° to between 150° and 200° due to occlusion caused by the human body's back and arm movements. Furthermore, in this type of embodiment, to avoid the laser beam emitted by the LiDAR sensor 204 hitting the nearby ground, the LiDAR sensor 204 is tilted upwards by 30° on a dedicated plate 206 to obtain a good vertical scanning view. In some embodiments, the motion and scene capture system 200 can be worn by multiple people to obtain 3D human motion data and point cloud data. For example, a first human body can wear the motion and scene capture system 200 to capture the first human body's 3D human motion while simultaneously capturing the environment measured by the first human body. In this example, a second human body can also wear the motion and scene capture system 200 and measure the same environment. In this example, the data obtained by the first human body and the data acquired by the second human body can be jointly optimized through a data processing pipeline. The data processing pipeline will be described herein. Figure 3A Further detailed discussion is needed.
[0057] Figure 3A The diagram illustrates a data processing pipeline 300 for determining pose estimation and LiDAR mapping based on data captured by a motion and scene capture system 200 according to various embodiments of the present disclosure. As shown in FIG3, this can be achieved by a computing system 208 (e.g., Figure 3A (302) Obtains 3D human motion data output from multiple IMU sensors 202 and point cloud data output from LiDAR sensor 204. Based on the 3D human motion, the self-motion of LiDAR sensor 204 can be estimated (e.g., Figure 3A (304). Based on point cloud data, a 3D scene graph S can be constructed (e.g., Figure 3A (306). Data initialization can be performed to prepare point cloud data for further optimization (e.g., Figure 3A (308). Graph-based optimization can be performed to fuse LiDAR trajectories determined based on 3D scene graphs and IMU trajectories determined based on 3D human motion (e.g., Figure 3A (310). Finally, by combining point cloud data, 3D human motion data, and 3D scene, joint optimization can be performed to generate human motion M and optimized scene (e.g., Figure 3A (312).
[0058] In some embodiments, the data processing pipeline 300 can perform IMU pose estimation. As shown below, the data processing pipeline 300 can estimate human motion M in IMU coordinates {I}: M I =(T I ,θ I ,,β), where T Iand θ I Provided by multiple IMU sensors 202. Attitude parameter θ I Hip joint orientation R relative to the starting cloud frame I It consists of rotations relative to its parent joint and other joints. This represents the translation of the k-th frame relative to the starting cloud frame. Since 3D human motion data captured by multiple IMU sensors is relatively accurate in the short term, T can be used. I and R I Optimize relative values.
[0059] In some embodiments, the data processing pipeline 300 can perform LiDAR localization and mapping. Generally, LiDAR jitter caused by human movement (such as walking) and occlusion caused by human bodies obstructing the field of view of the LiDAR sensor 204 make constructing a 3D map using point cloud data challenging. By employing a LiDAR-based SLAM method, the self-motion of the LiDAR sensor 204 can be estimated, and point cloud data can be utilized... (in {L}, k∈Z) + Construct a 3D scene graph S. For each LiDAR point cloud frame... Planner and edge feature points can be extracted from point cloud data and used to update the feature map. Since the mapping process can run offline, frame-to-frame odometry can be skipped, requiring only frame-to-map registration. Finally, the self-motion T of the LiDAR sensor 204 is calculated. W and R W And the 3D scene graph S. The graph construction function can be expressed as:
[0060]
[0061] In some embodiments, the data processing pipeline 300 can perform coordinate calibration. To obtain a rigid offset between the point cloud data captured by the LiDAR sensor 204 and the 3D human motion captured by multiple IMU sensors 202, aligning the coordinate systems of the point cloud data and the 3D motion data, the data processing pipeline 300 performs the following steps: First, before capture, the human body stands in an A-shape pose at the starting position in the scene, with the human face direction registered to the Y-axis of the scene. After capture, the point cloud data of the scene is rotated to the Z-axis of the scene, perpendicular to the ground at the starting position. Then, the scene is translated so that the origin of the scene matches the origin of the SMPL model corresponding to the human body standing on the ground. Then, the self-motion T of the LiDAR sensor 204... W and R W Translation and rotation are performed based on the scene. The point cloud data is calibrated as {W}. The pitch, roll, and translation of the 3D human motion are calibrated as a global coordinate system.
[0062] In some embodiments, the data processing pipeline 300 can perform time synchronization. In some embodiments, it can be based on T W and T I The peak timestamps are used to synchronize data captured from LiDAR sensor 204 and multiple IMU sensors 202 based on their z-values. These peaks are generated when a human body jumps from a standing position during motion capture. In some embodiments, the 3D human motion captured by the multiple IMU sensors 202 (100Hz) can be resampled at the same frame rate as the point cloud data captured by LiDAR sensor 204 (20Hz).
[0063] In some embodiments, the data processing pipeline 300 can perform graphics-optimized data fusion. For example... Figure 3B A comparison 320 of IMU and LiDAR trajectories according to various embodiments of this disclosure is shown. Generally, the IMU trajectory is a movement of the human body center derived from 3D human motion data, while the LiDAR trajectory is a movement of the human body center derived from point cloud data (rotation and translation starting from the center of the LiDAR sensor). Figure 3B It is known that 3D human motion data drifts significantly over time, becoming invalid when the scene height changes, while point cloud data, although correctly positioned, exhibits jitter during localized movements. Therefore, to estimate a smoother and more stable trajectory, the advantages of both types of data are utilized. As shown below, the data processing pipeline 300 can estimate the trajectory: 1) First, T... W Point cloud frames exceeding the IMU velocity are marked as outliers, where the IMU velocity is obtained by multiplying the coefficient by the local fit value, 2) the remaining (R W T W ) is regarded as a boundary marker, and then T is determined according to the boundary marker. I 3) Based on the constraints of the landmarks, make T... I and T W Alignment, 4) Optimize the graphics to make T I and T W coupling.
[0064] In some embodiments, the data processing pipeline 300 can perform joint optimization. To obtain accurate human motion M = (T, θ) in a natural scene, and higher quality point cloud data of the scene S, the data processing pipeline 300 can perform joint optimization using the point cloud data and physical constraints of the scene S. The optimized human motion and point cloud data T are sent as initial values to the back-construction graph function F to create a new scene S. optIn one particular implementation, translation from one scene to a new scene can be achieved using four constraints. These four constraints may include: a foot contact constraint L that encourages the human to stand on the ground. cont Eliminating the sliding constraint L for human walking sld From R I Orientation constraint L for smooth rotation ort And the smoothing constraint L that makes the translation smooth. smt The optimization can be represented as follows:
[0065]
[0066]
[0067]
[0068] Where, λ cont , λ sld , λ ort , λ smt Let L be the coefficient of the loss term. Gradient descent is used to minimize L, iteratively optimizing M. (i) =(T (i) θ (i) ), where (i) represents the iteration. Let M... (0) Set to (T) M θ I ).
[0069] In some embodiments, the data processing pipeline 300 can perform plane detection. In plane detection, to improve the effectiveness of foot contact, it is necessary to detect the plane near the human body. In some embodiments, cloth simulation filtering (CSF) can be used to extract ground points S in S. g Then, you can in S g Search T in China W The adjacent points. Unlike methods based on dense mesh models, discrete point clouds have blank areas, which may cause foot-touch constraints to be invalid. To address this issue, RANSAC can be used to fit the plane of adjacent points. The plane function can be represented as p k .
[0070] In some embodiments, the data processing pipeline 300 can determine foot contact constraints. The foot contact loss can be defined as the distance from the stable foot to the nearest point on the ground. Unlike HPS, which requires information about which foot is on the ground, the current method detects foot status based on movement. First, 3D human body data captured by multiple IMU sensors 202 can be used to determine foot contact constraints. For each consecutive foot vertex in the frame, the movement of the left and right feet is compared. In the current method, if the movement distance of one foot is less than 2cm and less than the movement distance of the other foot, the former is marked as a stable foot. The list of stable foot vertex indices in the k-th frame is written as V. j , denoted as S k Foot contact loss L cont It can be represented as follows:
[0071]
[0072] in, For v c Homogeneous coordinates. Represented as S in j The foot apex is derived from the 3D human motion that needs optimization.
[0073] In some embodiments, the data processing pipeline 300 can determine foot slippage constraints. These constraints reduce slippage of the 3D human body on the ground, resulting in more natural and smooth 3D human body movements. Slippage loss can be defined as the distance between every two consecutive stable feet. Slippage loss can be expressed as follows:
[0074]
[0075] Where IE(·) is the average function.
[0076] In some embodiments, the data processing pipeline 300 may determine orientation constraints. These orientation constraints allow the 3D human motion M to rotate smoothly like an IMU and have the same orientation as the previously described landmark A. The orientation loss may be expressed as follows:
[0077]
[0078] In some embodiments, the data processing pipeline 300 can determine smoothing constraints. These constraints enable 3D human motion to move as smoothly as IMU motion, thereby minimizing the translational distance difference between the LiDAR sensor 204 and the plurality of IMU sensors 202. The smoothing loss term can be expressed as follows:
[0079]
[0080] In general, the datasets published in this paper (i.e., the HSC4D dataset) can be evaluated in large indoor and outdoor 3D scenes. The results of these evaluations demonstrate the effectiveness and generalization ability of the HSC4D dataset. For example, Figure 3C An example HSC4D dataset 340 according to various embodiments of this disclosure is shown. Figure 3CAs shown, the example HSC4D dataset 340 can include three large scenes: a climbing gym 342, a multi-story building 344, and an outdoor closed-loop road 346. The climbing gym 342 has walls 20 meters high, and the ground and climbing area exceeds 1200m². 2 The multi-story building 344 includes a scene with indoor and outdoor areas up to 5000m². 2 This scene features various heights and environments, including multi-level structures, slopes, and stairs. The outdoor closed-loop road 346 has an area of 70m × 65m and includes a slope. Within these scenes, captured 3D human activities can include walking, exercising, going up / down stairs, rock climbing, etc. Since the 3D maps generated by LiDAR are coarse and lack color information, a terrestrial laser scanner (TLS) can be used to scan the scene in color for a more intuitive effect. In various embodiments, the HSC4D dataset can provide 250K IMU human motion frames (100Hz), 50k time-synchronized LiDAR point cloud frames (20Hz), various SLAM results, and scene-specific colored ground point clouds.
[0081] Figure 3D-3E Various comparison tables comparing the HSC4D dataset with other datasets according to various embodiments of this disclosure are shown. Table 362 shows the comparison between the HSC4D dataset and the HPS dataset. As can be seen from Table 366, the HSC4D dataset has many advantages compared to the HPS dataset. Table 364 shows a quantitative evaluation of the HSC4D dataset. In the table, column 2 shows the L values of the four sequences. cont and L sld Columns 3-5 show the ablation comparison of the loss term, columns 6-8 show the analysis of the neighborhood radius r in the pruning scenario, and columns 9-10 show the analysis of the optimized sequence length l. Table 366 shows the comparison of global localization errors between the baseline and the HSC4D dataset. Baseline 1 depicts the IMU results. Baseline 2 depicts the IMU pose + LiDAR localization. In the table, error is measured according to the distance (cm) from the selected checkpoint (CP) to the foot of the SMPL model. Table 368 shows the comparison of local pose errors between the baseline and the HSC4D dataset. L cont It is the average foot contact distance error (cm) to the nearest point on the ground. sld This is the average foot slip distance error (cm) between two consecutive feet on the ground. Building 01 / 03: 1.5min / 1min sequence on the second / first floor corridor. Building 02 / 04: 3min / 1min outdoor sequence, including stairs and ramps. Training Hall 01 / 02 / 03: Three one-minute walking and warm-up sequences. Road: 3.5-minute walking sequence.
[0082] These tables show that, compared to the HPS dataset, which uses IMU sensors to estimate 3D human pose and a head-mounted camera for localization in large scenes, the HSC4D dataset achieves more accurate localization. In this respect, the HSC4D dataset was generated using IMU sensors and a hip-mounted LiDAR instead of a camera. Furthermore, and more importantly, the HSC4D dataset can acquire pre-built maps without the need for a LiDAR sensor, allowing the capture system to directly capture scenes without any constraints. As shown in Table 362, the HSC4D dataset can extend scene capture to multi-story buildings and vertical routes, incorporating more human movements and interactions with the environment. As shown in Table 366, the errors in all methods increase linearly with increasing distance within the scene. IMUs drift over time, resulting in baseline 1 having an error more than ten times greater than other methods. Baseline 2 has a smaller global localization error, but the cumulative error is still between 8cm and 90cm. The last column shows that the HSC4D dataset has the smallest global localization error in multi-story buildings and sloping roads. More specifically, the accuracy of the HSC4D dataset is improved by 78.3% compared to baseline 1 and by 25.4% compared to baseline 2. The improvements are as follows: Figure 3F As shown. Figure 3F A graphical comparison of the HSC4D dataset described according to various embodiments of this disclosure with baselines 1 and 2 is shown. Figure 3F In the image, the colored spheres represent the IMU and LiDAR trajectories. (By...) Figure 3F As can be seen, the HSC4D dataset results are natural and accurate in both cases. For example, in the left image, baseline 1 and baseline 2 float in the air, while in the right image, baseline 1 floats in the air and the lower leg of baseline 2 penetrates the ground. As shown in Table 368, the foot contact loss of baseline 1 is much larger than other methods, especially in scenarios with changing height. Among all methods, the L... sld Maximum. In the first three sequences, baseline 1 did not drift with altitude, and baseline 2's L... cont Much larger than baseline 1. These cases indicate that LiDAR increases local errors. As shown in the previous column, in all cases, the HSC4D dataset makes L... cont Significantly reduced, achieving L comparable to baseline 1. sld Smoothness. These comparisons demonstrate that the HSC4D dataset can achieve smooth results in local poses and is robust across a wide range of highly diverse scenes.
[0083] Figure 4AA method for capturing human motion in a scene is illustrated according to various embodiments of this disclosure. It should be understood that, unless otherwise stated, additional, fewer, or alternative steps may be performed in a similar or alternative order or in parallel within the scope of the various embodiments discussed herein.
[0084] At frame 402, multiple IMU devices and a LiDAR sensor can be mounted on the human body. In some embodiments, the self-motion of the LiDAR sensor can be estimated. The LiDAR sensor can be mounted on the hip of the human body.
[0085] At box 404, IMU data captured by the IMU device and LiDAR data captured by the LiDAR sensor can be obtained. In some embodiments, the LiDAR data and the IMU data can be calibrated based on the self-motion of the LiDAR sensor. In some embodiments, the LiDAR data and IMU data can be synchronized based on peak values derived from the LiDAR data and peak values derived from the IMU data. During data acquisition, peak values can be determined when a person jumps.
[0086] At frame 406, human movements can be estimated based on the IMU data.
[0087] At frame 408, a three-dimensional scene map can be constructed based on the LiDAR data.
[0088] At box 410, optimization can be performed to obtain optimized human motion and optimized scene graph. In some embodiments, graph-based optimization can be performed to fuse LiDAR and IMU trajectories. The LiDAR trajectory may include human center movement derived from the LiDAR data, and the IMU trajectory may include human center movement derived from the IMU data. In some embodiments, the optimization can be based on contact and sliding constraints. The contact constraint can define the contact loss as the distance from a body part (derived from the human motion) to the nearest surface (derived from the 3D scene graph). The sliding constraint can define the sliding loss as the distance between two consecutive body parts derived from the human motion. In some embodiments, a gradient descent algorithm can be used to minimize the sum of the contact and sliding losses to iteratively optimize the human motion.
[0089] Figure 4B A computing component 450 according to various embodiments of the present disclosure is illustrated. The computing component 450 includes one or more hardware processors 452 and a machine-readable storage medium 454 (stored with a set of machine-readable / machine-executable instructions). When executed, the instructions cause the hardware processors 452 to perform a method for optimizing human motion in a scene. The computing component 450 may be... Figure 5 The computing system 500 in the middle. The hardware processor 452 may include Figure 5 The processor 504 or any other processing unit described in this invention. Machine-readable storage medium 454 may include main memory 506, read-only memory (ROM) 508, ... Figure 5 The memory 510 and / or any other suitable machine-readable storage medium described in this invention.
[0090] At frame 456, a 3D scene image and human motion can be obtained. In some embodiments, the 3D scene image can be obtained from LiDAR data captured by a LiDAR sensor mounted on the human body, and the human motion can be obtained from IMU data captured by multiple IMU devices mounted on the human body.
[0091] At box 458, graph-based optimization can be performed to fuse the LiDAR trajectory and the IMU trajectory. In some embodiments, the LiDAR trajectory may include the human center of gravity movement derived from LiDAR data, and the IMU trajectory may include the human center of gravity movement derived from IMU data.
[0092] At frame 460, joint optimization based on multiple physical constraints can be performed to obtain optimized human motion and optimized scene graph. In some embodiments, the joint optimization can be based on contact constraints and sliding constraints. The contact constraint can define the contact loss as the distance from the body part (derived from the human motion) to the nearest surface (derived from the 3D scene graph), and the sliding constraint can define the sliding loss as the distance between two consecutive body parts derived from the human motion.
[0093] For example, the techniques described in this invention can be implemented using one or more dedicated computing devices. These dedicated computing devices may execute the techniques in a hard-wired manner, or include circuitry or digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) (which can be continuously programmed to perform the techniques), or include one or more hardware processors (which, after programming, can execute the techniques according to program instructions in firmware, memory, other storage devices, or combinations thereof).
[0094] Figure 5 This diagram illustrates a computer system 500 that can implement the various embodiments described in this invention. The computer system 500 includes an information transport bus 502 or other communication mechanism, and one or more hardware processors 504 coupled to the bus 502 for information processing. The description of tasks performed by the device is intended to represent tasks performed by the one or more hardware processors 504.
[0095] Computer system 500 also includes a main memory 506, such as random access memory (RAM), cache memory, and / or other dynamic storage devices. Main memory 506 is coupled to bus 502 and is used to store information and instructions to be executed by processor 504. During the execution of instructions to be executed by processor 504, main memory 506 can also be used to store temporary variables or other intermediate information. When these instructions are stored in storage media accessible to processor 504, they transform computer system 500 into a customized special-purpose machine for performing the operations specified in the instructions.
[0096] The computer system 500 further includes a read-only memory (ROM) 508 or other static storage device coupled to a bus 502 for storing static information and instructions required by the processor 504. A storage device 510, such as a disk, optical disk, or USB thumb drive (flash drive), coupled to the bus 502, is provided for storing information and instructions.
[0097] Computer system 500 can be coupled to output device 512 (such as a cathode ray tube (CRT) or LCD display (or touch screen)) via bus 502 for displaying information to the computer user. Input device 514, including alphanumeric and other keys, is coupled to bus 502 for transmitting information and command selection to processor 504. Another user input device is cursor controller 516. Computer system 500 also includes a communication interface 518 coupled to bus 502.
[0098] Unless the context otherwise requires, the term “comprise” and its variations, such as “comprises” and “comprising”, shall be interpreted in an open, inclusive sense, meaning “including but not limited to”, throughout this specification and claims. Statements of numerical ranges are intended as shorthand notations to refer individually to each individual value within a range (including the range-defined value), and each individual value is part of this specification as it is individually enumerated in this invention. Furthermore, unless the context explicitly specifies otherwise, the singular forms “a,” “an,” and “the” include plural references. Phrases such as “at least one of,” “at least one selected from a set,” or “at least one selected from a set comprising” require interpretation within the context of disjunctive concepts (e.g., they should not be interpreted as at least one of A and at least one of B).
[0099] The reference to "an embodiment" or "an embodiment" in this specification means that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Therefore, the phrases "in one embodiment" or "in an embodiment" appearing in various parts of this specification do not necessarily all refer to the same embodiment, but may refer to the same embodiment in some cases. Furthermore, in one or more embodiments, particular features, structures, or characteristics may be combined in any suitable manner.
[0100] A component implemented as another component can be interpreted as a component that operates in the same or similar manner as another component, and / or includes the same or similar features, characteristics, and parameters as another component.
Claims
1. A method for capturing human motion in a scene, the method comprising the following steps: Multiple IMU devices and a LiDAR sensor are installed on the human body; Obtain IMU data captured by the IMU device and LiDAR data captured by the LiDAR sensor; Human movements are estimated based on the IMU data; A 3D scene map is constructed based on the LiDAR data; Optimize to obtain optimized human body movements and optimized scene images; The optimization is based on contact constraints and sliding constraints; The contact constraint defines contact loss as the distance from a body part derived from human motion to the nearest surface derived from the 3D scene graph, and the sliding constraint defines sliding loss as the distance between two consecutive body parts derived from human motion. The method further includes the following steps: The gradient descent algorithm is used to minimize the sum of the contact loss and the sliding loss in order to iteratively optimize human motion.
2. The method according to claim 1, further comprising the following steps: Estimate the self-motion of the LiDAR sensor, wherein the LiDAR sensor is mounted on the hip of a human body.
3. The method according to claim 2, further comprising the following steps: The LiDAR data and IMU data are calibrated based on the self-motion of the LiDAR sensor.
4. The method according to claim 1, further comprising the following steps: Human jumping is performed in the step of obtaining IMU data captured by the IMU device and LiDAR data captured by the LiDAR sensor; The LiDAR data and IMU data are synchronized based on the peak values derived from the LiDAR data and the peak values derived from the IMU data.
5. The method according to claim 1, further comprising the following steps: Graph optimization is performed to fuse LiDAR and IMU trajectories, wherein the LiDAR trajectory includes the human body center movement derived from the LiDAR data, and the IMU trajectory includes the human body center movement derived from the IMU data.
6. The method according to claim 5, wherein, The body part is the human foot, and the surface is the ground.
7. The method of claim 1, further comprising the following steps: Multiple second IMU devices and one second LiDAR sensor were installed on the second human body; Obtain the second IMU data captured by the second IMU device and the second LiDAR data captured by the second LiDAR sensor; The movements of the second human body are estimated based on the second IMU data; A three-dimensional second scene map is constructed based on the second LiDAR data; The three-dimensional scene map and the three-dimensional second scene map are merged to obtain a combined scene map; Optimize to obtain the optimized actions of the human body and the second human body in the optimized combined scene diagram.
8. The method of claim 1, further comprising the following steps: A metaverse is created based on the optimized human body movements and the optimized scene graph.
9. A motion capture system, comprising: Multiple wearable IMU devices; A human-mounted LiDAR sensor; A processor; A memory storing instructions that, when executed by a processor, cause the system to perform a human motion capture method in a scene, the method comprising: Obtain IMU data captured by the IMU device and LiDAR data captured by the LiDAR sensor; When an instruction is executed, it further causes the system to perform the following: Human motion is estimated based on the IMU and LiDAR data; A 3D scene map is constructed based on the LiDAR data; Optimize to obtain optimized human body movements and optimized scene images; The optimization is based on contact constraints and sliding constraints; Wherein, the contact constraint defines contact loss as the distance from a body part derived from human motion to the nearest surface derived from the 3D scene graph, and the sliding constraint defines sliding loss as the distance between two consecutive body parts derived from human motion. The method further includes: The gradient descent algorithm is used to minimize the sum of the contact loss and the sliding loss in order to iteratively optimize human motion.
10. The system according to claim 9, further comprising: An L-shaped bracket allows the LiDAR sensor to be mounted on the hip of a human body, wherein the LiDAR sensor and the IMU device may have a basic rigidity transformation.
11. The system of claim 9, further comprising: A wireless receiver is coupled to the system, wherein the wireless receiver can receive IMU data captured by the IMU device.
12. The system according to claim 9, wherein, When an instruction is executed, it further causes the system to perform the following: Graph-based optimization is performed to fuse LiDAR and IMU trajectories, wherein the LiDAR trajectory includes the human body center movement derived from the LiDAR data, and the IMU trajectory includes the human body center movement derived from the IMU data.
13. The system according to claim 12, wherein, The body part is the human foot, and the surface is the ground.
14. The system of claim 9, further comprising: Multiple secondary IMU devices can be worn by a secondary human body; A second LiDAR sensor can be installed on a second human body; The method further includes: Obtain the second IMU data captured by the second IMU device and the second LiDAR data captured by the second LiDAR sensor; The movements of the second human body are estimated based on the second IMU data; A three-dimensional second scene map is constructed based on the second LiDAR data; The three-dimensional scene map and the three-dimensional second scene map are merged to obtain a combined scene map; Optimize to obtain the optimized actions of the human body and the second human body in the optimized combined scene diagram.
15. The system according to claim 12, wherein, The system further includes: A metaverse is created based on the optimized human body movements and the optimized scene graph.
16. A method for optimizing human motion in a scene, the method comprising: Obtain 3D scene images and human animations; Graph-based optimizations are performed to fuse LiDAR and IMU trajectories; Perform joint optimization based on multiple physical constraints to obtain optimized human motion and optimized scene graph; The joint optimization is based on contact constraints and sliding constraints; Wherein, the contact constraint defines contact loss as the distance from a body part derived from human motion to the nearest surface derived from the 3D scene graph, and the sliding constraint defines sliding loss as the distance between two consecutive body parts derived from human motion. The method further includes: The gradient descent algorithm is used to minimize the sum of the contact loss and the sliding loss in order to iteratively optimize human motion.
17. The method according to claim 16, wherein, The three-dimensional scene map is obtained from LiDAR data captured by LiDAR sensors installed on the human body, and the human motion is obtained from IMU data captured by multiple IMU devices installed on the human body.
18. The method of claim 17, further comprising: The 3D scene diagram and human motion are calibrated.
19. The method of claim 16, further comprising: Synchronize the 3D scene diagram and human movements.
20. The method of claim 16, wherein, The LiDAR trajectory includes the movement of the human body center derived from LiDAR data, and the IMU trajectory includes the movement of the human body center derived from IMU data.
21. The method according to claim 20, wherein, The body part is the human foot, and the surface is the ground.
22. The method of claim 16, further comprising: Multiple second IMU devices and one second LiDAR sensor were installed on the second human body; Obtain the second IMU data captured by the second IMU device and the second LiDAR data captured by the second LiDAR sensor; The movements of the second human body are estimated based on the second IMU data; A three-dimensional second scene map is constructed based on the second LiDAR data; The three-dimensional scene map and the three-dimensional second scene map are merged to obtain a combined scene map; Optimize to obtain the optimized actions of the human body and the second human body in the optimized combined scene diagram.
23. The method of claim 16, further comprising: A metaverse is created based on the optimized human body movements and the optimized scene graph.