Method, device, medium and equipment for capturing human motion in a scene

By combining IMU, point cloud and image data, and using camera external parameters and lidar IMU data to optimize the three-dimensional scene mesh, the problem of low accuracy in outdoor human motion capture is solved, and motion capture with higher precision and accuracy is achieved.

CN116543457BActive Publication Date: 2025-10-03XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310483296.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2025-10-03
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in capturing human motion in outdoor scenes. Sensor drift causes unstable data, and the lack of depth information makes it difficult to provide global translation, affecting capture precision and data accuracy.

Method used

Combining IMU data, point cloud data and image data, through camera external parameter alignment, using lidar IMU data to estimate the motion trajectory, constructing a three-dimensional scene mesh, and optimizing it through human-scene contact constraints, smoothness constraints, posture prior constraints and mesh-to-point constraints to improve the accuracy of human motion and scene capture.

Benefits of technology

It improves the accuracy of capturing scenes and human movements, ensures the accuracy of data, and is suitable for motion capture in outdoor environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543457B_ABST
    Figure CN116543457B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a method, device, medium and equipment for capturing human motion in a scene. The method includes: obtaining the first IMU data of the target object, and the point cloud data and image captured by the observer for the target object; estimating the motion of the target object based on the first IMU data to obtain the human motion corresponding to the target object; aligning the point cloud data with the image; estimating the motion trajectory information of the laser radar based on the second IMU data of the laser radar itself, and constructing a three-dimensional scene mesh based on the motion trajectory information and the aligned point cloud data; optimizing the human motion and the three-dimensional scene mesh according to the human scene contact constraint item, smoothness constraint item, posture prior constraint item and grid-to-point constraint item to obtain the optimized human motion and three-dimensional scene mesh. The technical solution of the embodiment of the present application improves the capture accuracy of the scene and human motion, and ensures the accuracy of the captured data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, medium, and device for capturing human motion in a scene. Background Art

[0002] With the rapid development of computer technology, the digital world has enriched people's lives. Humans and the environment are the two major components of building the digital world. Current research tends to separate the dynamic human body from the static environment, thereby helping to improve the accuracy of capturing human motion and the environment. In current technical solutions, in order to capture human motion, an inertial measurement unit (IMU) is often used to collect human motion data. However, as the acquisition time increases, the sensor will drift. Even if an external camera is used as a remedial measure to improve accuracy, it will be unstable in outdoor scenes due to the lack of depth information, and it will be difficult to provide global translation. Therefore, how to improve the accuracy of capturing scenes and human motions and ensure the accuracy of the captured data has become a technical problem that needs to be solved urgently. Summary of the Invention

[0003] The embodiments of the present application provide a method, device, medium and electronic device for capturing human motion in a scene, which can improve the capture accuracy of the scene and human motion to at least a certain extent, and ensure the accuracy of the captured data.

[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.

[0005] According to one aspect of an embodiment of the present application, a method for capturing human motion in a scene is provided, comprising:

[0006] Obtaining first IMU data of a target object, and point cloud data and an image captured by an observer for the target object, wherein the first IMU data is obtained by a first IMU device worn by the target object, and the point cloud data and the image are respectively obtained by a lidar and a camera worn by the observer;

[0007] Estimating the motion of the target object according to the first IMU data to obtain a human body motion corresponding to the target object;

[0008] Converting the point cloud data into an image coordinate system corresponding to the camera according to the external parameters of the camera to align the point cloud data with the image;

[0009] Based on the second IMU data of the laser radar itself, estimate the motion trajectory information of the laser radar, and construct a three-dimensional scene mesh according to the motion trajectory information and the aligned point cloud data;

[0010] The human body motion and the three-dimensional scene mesh are optimized according to the human body scene contact constraint item, the smoothness constraint item, the posture prior constraint item and the grid-to-point constraint item to obtain the optimized human body motion and three-dimensional scene mesh, wherein the human body scene contact constraint item is used to constrain the contact between the human and the scene without penetration, the smoothness constraint item is used to ensure the temporal continuity of translation, direction and joints, the posture prior constraint item is used to ensure the consistency of human body motion and IMU data at the beginning of optimization, and the grid-to-point constraint item is used to minimize the distance from the human body vertex to the resampled human body point.

[0011] According to one aspect of an embodiment of the present application, a human motion capture device in a scene is provided, comprising:

[0012] an acquisition module, configured to acquire first IMU data of a target object, and point cloud data and an image captured by an observer of the target object, wherein the first IMU data is acquired by a first IMU device worn by the target object, and the point cloud data and the image are acquired by a lidar and a camera worn by the observer, respectively;

[0013] A first estimation module is used to estimate the motion of the target object according to the first IMU data to obtain a human body motion corresponding to the target object;

[0014] a conversion module, configured to convert the point cloud data into an image coordinate system corresponding to the camera according to external parameters of the camera, so as to align the point cloud data with the image;

[0015] A second estimation module is configured to estimate motion trajectory information of the laser radar based on the second IMU data of the laser radar itself, and construct a three-dimensional scene mesh according to the motion trajectory information and the aligned point cloud data;

[0016] A processing module is used to optimize the human body motion and the three-dimensional scene mesh according to the human body scene contact constraint item, the smoothness constraint item, the posture prior constraint item and the grid-to-point constraint item to obtain the optimized human body motion and three-dimensional scene mesh, wherein the human body scene contact constraint item is used to constrain the contact between the human and the scene without penetration, the smoothness constraint item is used to ensure the temporal continuity of translation, direction and joints, the posture prior constraint item is used to ensure the consistency of human body motion and IMU data at the beginning of optimization, and the grid-to-point constraint item is used to minimize the distance from the human body vertex to the resampled human body point.

[0017] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for capturing human motion in a scene as described in the above embodiment is implemented.

[0018] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the method for capturing human motion in a scene as described in the above embodiments.

[0019] According to one aspect of an embodiment of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method for capturing human motion in a scene provided in the above-described embodiment.

[0020] In the technical solutions provided in some embodiments of the present application, first IMU data of a target object and point cloud data and images captured by an observer for the target object are obtained, where the first IMU data is obtained by a first IMU device worn by the target object itself, and the point cloud data and images are respectively obtained by a lidar and a camera worn by the observer. The target object's motion is estimated based on the first IMU data to obtain a human motion corresponding to the target object. The point cloud data is converted into an image coordinate system corresponding to the camera based on external parameters of the camera to align the point cloud data and the image. The motion trajectory information of the lidar is then estimated based on the second IMU data of the lidar itself, and a three-dimensional scene mesh is constructed based on the motion trajectory information and the aligned point cloud data. The human motion and the three-dimensional scene mesh are optimized based on human-scene contact constraints, smoothness constraints, posture prior constraints, and grid-to-point constraints to obtain optimized human motion and three-dimensional scene mesh. Therefore, by combining the IMU data, images, and point cloud data, and optimizing the three-dimensional scene mesh and human motion based on the above constraints, the accuracy of scene and human motion capture can be improved, and the accuracy of the captured data can be guaranteed.

[0021] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, explaining the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:

[0023] Figure 1 A schematic flow chart of a method for capturing human motion in a scene according to an embodiment of the present application is shown;

[0024] Figure 2 A schematic diagram of a collection system that can be applied to the technical solution of an embodiment of the present application is shown;

[0025] Figure 3 A block diagram of a human motion capture device in a scene according to an embodiment of the present application is shown;

[0026] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0027] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.

[0028] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0029] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0030] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0031] Figure 1 A schematic flow chart of a method for capturing human motion in a scene according to one embodiment of the present application is shown. This method can be applied to a terminal device or a server. The terminal device may include, but is not limited to, one or more of a smartphone, a tablet computer, a laptop computer, and a desktop computer. The server may be a physical server or a cloud server, and this application does not impose any specific restrictions on this.

[0032] It is worth noting that this application does not limit the number of terminal devices and / or servers, that is, the number of terminal devices and / or servers can be one or more than one. For example, the server can be a single server or a server cluster composed of multiple servers.

[0033] Please refer to Figure 1 The method for capturing human motion in a scene at least includes steps S110 to S150, which are described in detail as follows:

[0034] In step S110, first IMU data of the target object and point cloud data and images captured by the observer for the target object are obtained. The first IMU data is obtained by a first IMU device worn by the target object itself, and the point cloud data and images are respectively obtained by a lidar and a camera worn by the observer.

[0035] The target object may be a human object whose motion is to be captured, which may be one or more. Each target object may be equipped with a first IMU device to collect motion information of the target object and generate corresponding first IMU data.

[0036] The observer can wear a lidar and a camera. When performing motion capture, the observer can use the lidar and the camera to obtain point cloud data and images including the environment and the target object.

[0037] like Figure 2As shown, in one embodiment of the present application, an acquisition system is also provided, which consists of a 128-beam Ouster-os1 laser radar, a DJI-Action2 wide-angle camera installed under the laser radar, and Noitom's inertial motion capture (Mocap) product PN Studio to obtain human body movement. PN Studio is worn by the target object, which uses 17 wireless IMUs to connect to the limbs, and the IMU receiver and LiDAR are connected to the Intel NUC1 microcomputer, and are charged using a 24V mobile power supply. The observer can wear a helmet, the LiDAR is rigidly connected to the helmet, and the data is transmitted to the NUC11 through a cable behind the helmet. The camera data can be stored offline in a local storage space. Thus, through the above-mentioned acquisition system, the first IMU data, point cloud data and image corresponding to the target object can be obtained.

[0038] Based on the aforementioned acquisition system, three coordinate systems can be defined:

[0039] 1) IMU coordinate system {I}: The origin is located at the base of the spine of the wearer in the first LiDAR frame, and the X / Y / Z axes point to the left / upper / front of the human body.

[0040] 2) LiDAR coordinate system {L}: The origin is at the center of the LiDAR, and the X / Y / Z axes point to the right / front / top of the LiDAR.

[0041] 3) Global / World Coordinate System {W}: The origin is located on the floor at the starting position of the LiDAR wearer, and the X / Y / Z axes point to the right / front / above the LiDAR wearer.

[0042] In one embodiment of the present application, before motion capture, a terrestrial three-dimensional scanning system (Terrestrial Laser Scanning, TLS) is also used to calibrate the device position. Specifically, the high-precision color point cloud map scanned by TLS is used as a reference, and the 2D image and the 3D point cloud of LiDAR are aligned to the high-precision color point cloud map to obtain the accurate positional relationship between the camera and LiDAR.

[0043] In addition, for each scene, the LiDAR to world coordinate calibration matrix R WL They are all manually set so that the ground z-axis of the first frame of the LiDAR data is upward and its height is zero. The calibration matrix R from {I} to {W} is calculated by the similarity of the LiDAR trajectory and the IMU trajectory in the XY plane using singular value decomposition (SVD) LI .

[0044] In one embodiment of the present application, after obtaining the first IMU data of the target object and the point cloud data and image captured by the observer for the target object, the method further includes:

[0045] The peak values ​​of the target object caused by jumping at the beginning and end of motion capture are determined in the first IMU data and the point cloud data respectively to synchronize the first IMU device and the lidar.

[0046] In this embodiment, synchronization between the IMU and LiDAR is achieved by detecting jump peaks across multiple sensors. Specifically, the subject is asked to jump at the beginning and end of motion capture. The peak heights are then detected in the IMU and manually found in the LiDAR and camera. Finally, the higher-frequency IMU and camera data are resampled to the LiDAR frame rate (e.g., 20Hz).

[0047] In step S120, the motion of the target object is estimated according to the first IMU data to obtain the human body motion corresponding to the target object.

[0048] In this embodiment, based on the first IMU data, the motion of the target object can be estimated to obtain the human body motion corresponding to the target object. It should be noted that the motion estimation can adopt the existing algorithm, which will not be repeated here.

[0049] In one embodiment of the present application, performing motion estimation on the target object according to the first IMU data to obtain a human motion corresponding to the target object includes:

[0050] Motion estimation is performed based on the first IMU data to obtain an SMPL model for describing a human body motion corresponding to the target object.

[0051] In this embodiment, SMPL (Skinned Multi-Person Linear Model) is a skinned, vertex-based three-dimensional human model that can accurately represent different shapes and poses of the human body. By using the SMPL model to represent the human motion M in the IMU coordinates {I} I =(T I ,θ I , β), where T I and θ I Provided by MoCap. Pose parameter θ I The direction R of the pelvic joint relative to the starting frame Iand the rotations of the other 23 joints relative to their parent joints. β is a 10×1 constant representing a person's body shape. By scanning the texture models of all subjects (i.e., target objects), the SMPL model shape information of the subjects is calculated and fitted using IPNet (which combines implicit function learning and parametric models for 3D human body reconstruction). k I represents the k-th frame translation in the global coordinate system, due to the T in the IMU I and R I The relative value of is accurate, but it will drift over time. Therefore, the position data provided by the IMU is only used for the initial calibration of {I} to {W}.

[0052] Please continue to refer to Figure 1 In step S130, the point cloud data is converted into an image coordinate system corresponding to the camera according to the external parameters of the camera, so as to align the point cloud data with the image.

[0053] In this embodiment, for point cloud data, the external parameter K between the camera and the lidar is used. ex The point cloud data is converted to the image coordinate system so that the image and point cloud data can be aligned to enhance the map information.

[0054] In one embodiment of the present application, the method further includes:

[0055] Recognize the target object in the image and determine its skeleton and bounding box;

[0056] Based on the point cloud data and the SMPL model, respectively determining a bounding box corresponding to the target human body in the point cloud data and a skeleton corresponding to the target human body in the SMPL model;

[0057] The camera's extrinsic parameters are iteratively optimized using a gradient descent algorithm, by minimizing the intersection-over-union loss of corresponding points between the skeleton in the image and the skeleton in the SMPL model, and the sum of the mean squared errors of corresponding nodes between the bounding box in the image and the bounding box in the point cloud data.

[0058] In this embodiment, in order to reduce the impact of the dynamic rotation of the head and the error in time synchronization between devices, two pairs of 2D-3D correspondences can be used to optimize the K between the camera and the radar. exSpecifically, image recognition can be performed on the image to identify the 2D skeleton corresponding to the target object in the image and the 2D bounding box of the human body, and based on the point cloud data and the SMPL model, the 3D bounding box of the target object in the point cloud data and the 3D skeleton corresponding to the target object in the SMPL model are determined respectively. Then, the intersection loss L between the corresponding points of the two skeletons is calculated. IoU and the mean square error (MSE) L of the corresponding nodes on the two bounding boxes jts , minimize L jts +L IoU , and the gradient descent method is used to iteratively optimize the camera's external parameters.

[0059] Please continue to refer to Figure 1 In step S140, based on the second IMU data of the laser radar itself, the motion trajectory information of the laser radar is estimated, and a three-dimensional scene grid is constructed according to the motion trajectory information and the aligned point cloud data.

[0060] In this embodiment, due to dynamic head rotation and crowded urban environments, if only LiDAR is used, the commonly used uniform velocity compensation model is difficult to correct the distortion of the point cloud within the frame, so it often fails in scene mapping. Combining IMU can solve this problem by compensating for the motion distortion in the LiDAR scan and providing a good initial posture. Using the IMU built into the LiDAR, the second IMU data is processed by integrating an iterative Kalman filter, and the pose changes between the point cloud data are optimized by loop closure based on a factor graph to estimate the motion trajectory information of the LiDAR, that is, observe the user's own motion trajectory to achieve more accurate scene reconstruction. Therefore, a three-dimensional scene mesh can be constructed based on the motion trajectory information and the aligned point cloud data. In one example, in order to meet the optimization needs, VDB-Fusion is used to generate a clean three-dimensional scene mesh S, which can be used for subsequent joint optimization and visualization.

[0061] In step S150, the human body motion and the three-dimensional scene mesh are optimized according to the human body scene contact constraint item, the smoothness constraint item, the posture prior constraint item and the grid-to-point constraint item to obtain the optimized human body motion and three-dimensional scene mesh, wherein the human body scene contact constraint item is used to constrain the contact between the human and the scene without penetration, the smoothness constraint item is used to maintain the time continuity of translation, direction and joints, the posture prior constraint item is used to ensure the consistency of human body motion and IMU data at the beginning of optimization, and the grid-to-point constraint item is used to minimize the distance from the human body vertex to the resampled human body point.

[0062] In this embodiment, in order to obtain accurate and credible human motion M W, performs a joint optimization by utilizing the scene mesh S and several physics terms. The following optimization terms are used:

[0063] Human scene contact term L sf Make the person contact with the scene without penetrating, the smoothing term L smt Keep translation, direction and joints continuous in time, and the pose prior L prior To ensure that the body pose remains close to the IMU result at the beginning of the optimization, the grid-to-point term L m2p Minimize the distance from the body vertex to the clipped body point.

[0064] The optimization expression is shown in the following formula (1), where λ sf ,λ smt ,λ prior ,λ m2p is the loss term coefficient. L is minimized using the gradient descent algorithm to iteratively optimize M (i) =(T (i) ,θ (i) ), where (i) represents the iterative term.

[0065] L=λ sf L sf +λ smt L smt +λ prior L prior +λ m2p L m2p (1)

[0066]

[0067] Human scene contact loss L sf It is used in experiments to measure the distance between the vertices of the foot SMPL model and the surrounding scene. By detecting the state of the foot according to the action, the V k I =Φ(T k I ,θ k I ,β), the left and right foot movements of each consecutive foot vertex in,in one example, if the movement of a foot vertex is less than 2cm, it is marked as a stable vertex,,i.e., the part should be in direct contact with the ground.

[0068] The smoothing term includes the translation term L trans , direction item L orit and joint term L joints , see formula (2), λ trans ,λ orit ,λ jointsis the corresponding loss term coefficient. The root joint is defined as a translation, and the root relative joint is defined by subtracting it from the remaining joints. The goal of the translation term and the joint term is to minimize acceleration and smooth motion, while the goal of the orientation term is to minimize the relative rotation of the motion.

[0069] L smt =λ trans L trans +λ orit L orit +λ joints L joints (2)

[0070] in,

[0071]

[0072]

[0073]

[0074] The pose estimated by IMU may cause some misalignment of the limb ends due to drift, but it is relatively accurate in the short term. prior Constrain the root relative body pose θ, with the goal of making the optimized pose as close to it as possible, where θ is provided by Mocap, θ imu Recorded by the IMU on the subject. This constraint is regularized by equation (3):

[0075]

[0076] Point cloud data P provides strong prior depth information for human pose, which can be used for optimization. However, the SMPL grid is dense, but the points of the human body are sparse and local. Therefore, even after optimization using methods such as ICP, the local pose, global translation, and even the orientation error from the IMU are not as ideal as expected.

[0077] To this end, in one embodiment of the present application, the method further includes:

[0078] Removing hidden vertices in the SMPL model that are blocked by objects and cannot be scanned by the laser radar, simulating the resolution of the laser radar, and resampling the hidden vertices based on the viewpoint of the laser radar;

[0079] The vertices other than the hidden vertices in the SMPL model are regarded as visible human body vertices, and the distances between human body points and visible human body vertices are minimized to obtain the mesh-to-point constraint item.

[0080] In this embodiment, a viewpoint-based mesh-point loss term is provided. First, the hidden vertices of the SMPL mesh in the LiDAR viewpoint are removed, that is, the vertices in the SMPL mesh that are located in positions blocked by objects and cannot be scanned by the LiDAR. Then, the vertices are resampled like ray casting, simulating the LiDAR resolution. The obtained SMPL vertex is defined as vertex P i ′. Finally, use L m2p Minimize the human point cloud P and P i ′. This constraint is regularized by the following equation:

[0081]

[0082] Thus, by acquiring first IMU data of the target object and point cloud data and images captured by the observer for the target object, the first IMU data is acquired by a first IMU device worn by the target object itself, and the point cloud data and images are acquired by a laser radar and a camera worn by the observer, respectively, the motion of the target object is estimated based on the first IMU data to obtain the human motion corresponding to the target object, the point cloud data is converted into an image coordinate system corresponding to the camera according to the external parameters of the camera to align the point cloud data and the image, and then the motion trajectory information of the laser radar is estimated based on the second IMU data of the laser radar itself, and a three-dimensional scene mesh is constructed based on the motion trajectory information and the aligned point cloud data, and the human motion and the three-dimensional scene mesh are optimized according to the human-scene contact constraint, the smoothness constraint, the posture prior constraint, and the grid-to-point constraint to obtain an optimized human motion and three-dimensional scene mesh. Thus, by combining the IMU data, the image, and the point cloud data, and optimizing the three-dimensional scene mesh and the human motion according to the above constraints, the accuracy of scene and human motion capture can be improved, and the accuracy of the captured data can be guaranteed.

[0083] Based on the aforementioned embodiments, the optimized human body motion and three-dimensional scene mesh can be added to the dataset for storage. For the obtained dataset, the dataset can be evaluated to determine the reliability of the dataset. First, a stereotyped evaluation can be performed on the dataset to show that the dataset is reliable enough to benchmark new tasks. Then, a cross-dataset evaluation is performed to further evaluate the innovation of the dataset on two tasks, including lidar-based 3DHPE and camera-based 3D HPE. Finally, a new benchmark, global human pose estimation (GHPE), is introduced and experiments are conducted on GLMR.

[0084] Specifically, the data is divided into training and test sets for LiDAR / camera-based pose estimation. The training set contains 80,000 LiDAR frames and corresponding RGB frames. The test set contains approximately 20,000 LiDAR frames and corresponding RGB frames. For global human pose estimation, this module selects three challenging scenes for evaluation. The first is a single-player soccer training scene with highly dynamic movements on a playground. The second is a running scene along a coastal track. The third is a park activity involving everyday human activities.

[0085] For 3D HPE, mean per-joint position error (MPJPE) and mean per-joint position error using Procruste analysis (PA-MPJPE) are used for evaluation. MPJPE is the average Euclidean distance between the true and predicted joints.

[0086] PA-MPJPE first performs a rigid transformation based on Procrustes analysis to align the predicted joints with the true joints. It then calculates the Euclidean distance between each joint position between the predicted and true poses and averages these distances to obtain the average joint position error (MPJPE). For global trajectory evaluation, combining the absolute trajectory error (ATE) and relative pose error (RPE) (here, pose refers to human body positioning) in a visual SLAM system provides a more comprehensive performance evaluation. This considers both the overall system error and the pose error between adjacent frames, helping to assess the system's accuracy and stability. For ATE, the error between the pose estimate and the true pose at each moment in the system's trajectory is calculated and averaged to obtain the average error. For RPE, the relative pose estimation error between each adjacent frame is calculated and averaged to obtain the average error. Finally, the average ATE and RPE errors are summed to obtain the overall system error metric. ATE is well-suited for measuring global positioning, while RPE is suitable for measuring system drift, for example, per second.

[0087] For qualitative evaluation of human pose, the method is to project SMPL into the image and visualize the 3D human body in 3D space using the corresponding LiDAR points. The results show that the 3D human body mesh is well aligned with the 3D environment and 2D image. As a large-scale city-level human pose dataset, it not only provides multimodal capture data and rich human scene annotations, but also includes a variety of challenging human activities in large scenes. To evaluate this joint optimization method, the results of the joint optimization method are first compared with those of ICP. The scene-aware constraints and the human mesh-to-point constraints effectively optimize the local pose, global translation, and even the orientation error of the IMU. After optimization, the 2D projection error is significantly reduced.

[0088] When performing cross-dataset evaluation, this application uses different modalities (i.e., lidar and camera) to evaluate the 3D human pose estimation associated with the root node. 3DPW is the outdoor human motion dataset most relevant to this application. Through VIBE, this module cross-evaluates the camera modality of the dataset using 3DPW. LiDARHuman26M is a lidar-based long-range human pose estimation dataset that can be used to cross-evaluate the LiDAR modality of the dataset. When the model is trained from only the other dataset, the error is the largest. But when trained on the other dataset and the dataset of this patent, the error is further reduced. This shows that there is a field gap between different lidar sensors and the two datasets complement each other.

[0089] The following describes an embodiment of the device of the present application, which can be used to perform the method for capturing human motion in a scene in the above-mentioned embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method for capturing human motion in a scene in the above-mentioned embodiment of the present application.

[0090] Figure 3 A block diagram of a human motion capture device in a scene according to an embodiment of the present application is shown.

[0091] Reference Figure 3 As shown, a human motion capture device in a scene according to one embodiment of the present application includes:

[0092] An acquisition module 310 is configured to acquire first IMU data of a target object, and point cloud data and an image captured by an observer of the target object, wherein the first IMU data is acquired by a first IMU device worn by the target object, and the point cloud data and the image are acquired by a lidar and a camera worn by the observer, respectively.

[0093] A first estimation module 320 is configured to estimate the motion of the target object based on the first IMU data to obtain a human motion corresponding to the target object;

[0094] a conversion module 330 for converting the point cloud data into an image coordinate system corresponding to the camera according to the external parameters of the camera, so as to align the point cloud data with the image;

[0095] A second estimation module 340 is configured to estimate the motion trajectory information of the laser radar based on the second IMU data of the laser radar itself, and construct a three-dimensional scene mesh according to the motion trajectory information and the aligned point cloud data;

[0096] The processing module 350 is used to optimize the human body motion and the three-dimensional scene mesh according to the human body scene contact constraint item, the smoothness constraint item, the posture prior constraint item and the grid-to-point constraint item to obtain the optimized human body motion and three-dimensional scene mesh, wherein the human body scene contact constraint item is used to constrain the contact between the human and the scene without penetration, the smoothness constraint item is used to ensure the temporal continuity of translation, direction and joints, the posture prior constraint item is used to ensure the consistency of human body motion and IMU data at the beginning of optimization, and the grid-to-point constraint item is used to minimize the distance from the human body vertex to the resampled human body point.

[0097] In one embodiment of the present application, the first estimation module 320 is configured to perform motion estimation based on the first IMU data to obtain an SMPL model for describing a human motion corresponding to the target object.

[0098] In one embodiment of the present application, the conversion module 330 is further used to: perform recognition based on the image to determine the skeleton and bounding box corresponding to the target object in the image; determine the bounding box corresponding to the target object in the point cloud data and the skeleton corresponding to the target object in the SMPL model based on the point cloud data and the SMPL model; minimize the intersection loss of corresponding points between the skeleton in the image and the skeleton in the SMPL model and the sum of the mean square errors of corresponding nodes between the bounding box in the image and the bounding box in the point cloud data, and use a gradient descent algorithm to iteratively optimize the external parameters of the camera.

[0099] In one embodiment of the present application, after acquiring the first IMU data of the target object, and the point cloud data and image captured by the observer for the target object, the acquisition module 310 is also used to: respectively determine the peak values ​​of the target object in the first IMU data and the point cloud data caused by jumping at the beginning and end of motion capture, so as to synchronize the first IMU device and the lidar.

[0100] In one embodiment of the present application, the processing module 350 is further used to: remove hidden vertices in the SMPL model that cannot be scanned by the lidar due to being blocked by objects, simulate the resolution of the lidar, and resample the hidden vertices based on the viewpoint of the lidar; regard vertices other than the hidden vertices in the SMPL model as visible human body vertices, and minimize the distance between the human body points and the visible human body vertices to obtain the mesh-to-point constraint item.

[0101] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown.

[0102] It should be noted that Figure 4 The computer system of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0103] like Figure 4 As shown, the computer system includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage part 408 into the random access memory (RAM) 403, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 403. The CPU 401, ROM 402 and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0104] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, and the like; an output section 407 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 408 including a hard disk and the like; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. Removable media 411, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 410 as needed, so that computer programs read therefrom can be installed into the storage section 408 as needed.

[0105] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 409, and / or installed from a removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the various functions defined in the system of the present application are executed.

[0106] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0108] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.

[0109] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device implements the method described in the above embodiments.

[0110] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0111] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0112] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.

[0113] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for capturing human motion in a scene, characterized in that: include: Obtaining first IMU data of a target object, and point cloud data and an image captured by an observer for the target object, wherein the first IMU data is obtained by a first IMU device worn by the target object, and the point cloud data and the image are respectively obtained by a lidar and a camera worn by the observer; Estimating the motion of the target object according to the first IMU data to obtain a human body motion corresponding to the target object; Converting the point cloud data into an image coordinate system corresponding to the camera according to the external parameters of the camera to align the point cloud data with the image; Based on the second IMU data of the laser radar itself, estimate the motion trajectory information of the laser radar, and construct a three-dimensional scene mesh according to the motion trajectory information and the aligned point cloud data; The human body motion and the three-dimensional scene mesh are optimized according to a human-scene contact constraint, a smoothness constraint, a posture prior constraint, and a mesh-to-point constraint to obtain an optimized human body motion and three-dimensional scene mesh, wherein the human-scene contact constraint is used to constrain the contact between the human and the scene without penetration, the smoothness constraint is used to ensure that the translation, direction, and joints maintain time continuity, the posture prior constraint is used to ensure the consistency of the human body motion and the IMU data at the beginning of the optimization, and the mesh-to-point constraint is used to minimize the distance from the human body vertex to the resampled human body point; The step of estimating the motion of the target object according to the first IMU data to obtain a human motion corresponding to the target object includes: Performing motion estimation based on the first IMU data to obtain an SMPL model for describing a human body motion corresponding to the target object; The method further comprises: Recognize the target object in the image and determine its skeleton and bounding box; Based on the point cloud data and the SMPL model, determining a bounding box corresponding to the target object in the point cloud data and a skeleton corresponding to the target object in the SMPL model; The camera's extrinsic parameters are iteratively optimized using a gradient descent algorithm, by minimizing the intersection-over-union loss of corresponding points between the skeleton in the image and the skeleton in the SMPL model, and the sum of the mean squared errors of corresponding nodes between the bounding box in the image and the bounding box in the point cloud data.

2. The method according to claim 1, characterized in that After acquiring the first IMU data of the target object and the point cloud data and image captured by the observer for the target object, the method further includes: The peak values ​​of the target object caused by jumping at the beginning and end of motion capture are determined in the first IMU data and the point cloud data respectively, so as to synchronize the first IMU device and the lidar.

3. The method according to claim 1, characterized in that The method further comprises: Removing hidden vertices in the SMPL model that are blocked by objects and cannot be scanned by the laser radar, simulating the resolution of the laser radar, and resampling the hidden vertices based on the viewpoint of the laser radar; The vertices other than the hidden vertices in the SMPL model are regarded as visible human body vertices, and the distances between human body points and visible human body vertices are minimized to obtain the mesh-to-point constraint item.

4. A human motion capture device in a scene, characterized in that: include: an acquisition module, configured to acquire first IMU data of a target object, and point cloud data and an image captured by an observer of the target object, wherein the first IMU data is acquired by a first IMU device worn by the target object, and the point cloud data and the image are acquired by a lidar and a camera worn by the observer, respectively; A first estimation module is used to estimate the motion of the target object according to the first IMU data to obtain a human body motion corresponding to the target object; a conversion module, configured to convert the point cloud data into an image coordinate system corresponding to the camera according to external parameters of the camera, so as to align the point cloud data with the image; A second estimation module is configured to estimate motion trajectory information of the laser radar based on the second IMU data of the laser radar itself, and construct a three-dimensional scene mesh according to the motion trajectory information and the aligned point cloud data; a processing module, configured to optimize the human motion and the three-dimensional scene mesh according to a human-scene contact constraint, a smoothness constraint, a posture prior constraint, and a mesh-to-point constraint to obtain an optimized human motion and three-dimensional scene mesh, wherein the human-scene contact constraint is used to constrain the contact between the human and the scene to prevent penetration, the smoothness constraint is used to ensure temporal continuity of translation, direction, and joints, the posture prior constraint is used to ensure consistency between the human motion and the IMU data at the start of optimization, and the mesh-to-point constraint is used to minimize the distance from the human vertex to the resampled human point; Wherein, the first estimation module is used for: Performing motion estimation based on the first IMU data to obtain an SMPL model for describing a human body motion corresponding to the target object; Wherein, the conversion module is further used for: Recognize the target object in the image and determine its skeleton and bounding box; Based on the point cloud data and the SMPL model, determining a bounding box corresponding to the target object in the point cloud data and a skeleton corresponding to the target object in the SMPL model; The camera's extrinsic parameters are iteratively optimized using a gradient descent algorithm, by minimizing the intersection-over-union loss of corresponding points between the skeleton in the image and the skeleton in the SMPL model, and the sum of the mean squared errors of corresponding nodes between the bounding box in the image and the bounding box in the point cloud data.

5. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for capturing human motion in a scene according to any one of claims 1 to 3 is implemented.

6. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the human motion capture method in a scene as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Motion capture method based on inertia and optical measurement fusion

    CN104658012A

  • Motion capture method and device, and server

    CN113705520A