A 3D human body reconstruction method and device
Through the parameterized rigid transformation of the human body model and the alignment of point cloud data, and automatic binding with the preset distance strategy, the problems of low efficiency and low accuracy of human body three-dimensional reconstruction in the existing technology are solved, and efficient and accurate model binding and driving effects are achieved.
Patent Information
- Application Number
- CN202111273524.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-10-29
AI Technical Summary
In the prior art, the motion matching and binding process of human body three-dimensional reconstruction has problems such as low efficiency and low accuracy. Manual binding takes a long time and the automatic binding effect is not ideal, resulting in poor fidelity and driving effect of the reconstruction model.
By obtaining the bone node motion logic data of the pre-constructed parameterized human model, using multiple preset key points for rigid transformation and fitting, combining point cloud data alignment to the driven model, using the preset distance strategy for automatic binding, and determining the vertex motion logic data and skin weights based on the skin weight and the motion logic data of the bone node, to achieve fast and reasonable model binding.
It improves the efficiency of three-dimensional reconstruction, reduces binding time, avoids "brows" and "slips" phenomena, enhances the accuracy and authenticity of the model, and ensures that the driving model is consistent with the posture and movement of the reconstruction object.
Smart Images

Figure CN116071485B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of computer vision and computer graphics, and particularly relates to a method and device for three-dimensional human body reconstruction. Background Art
[0002] With the gradual expansion of virtual reality (VR) technology and augmented reality (AR) technology from the military and industrial fields to the entertainment and life fields, people's social interaction methods have changed. Realistic virtual avatars conduct social interactions in virtual spaces, reproducing the immersive feeling of face-to-face communication in the first generation of social interaction methods. Virtual / augmented reality social interaction methods may become the fifth generation of social media after the fourth generation of the mobile Internet era. As a key technology for virtual social interaction, three-dimensional human body reconstruction has important research significance.
[0003] Generally, three-dimensional human body reconstruction involves shape, pose, and texture data. During the reconstruction process, acquisition information is first obtained from various sensors, and then a three-dimensional reconstruction method is used to process the acquisition information, thereby reconstructing a three-dimensional human body model.
[0004] Due to network bandwidth limitations, currently, in social scenarios such as 3D holographic communication and virtual live streaming, the ″pre-modeling + human pose″ method is mostly used to achieve real-time driving to meet bandwidth requirements. The human pose is obtained through motion capture technology to real-time drive a pre-generated three-dimensional human body model, thereby completing the dynamic three-dimensional reconstruction of the human body. During the reconstruction process, the motion matching and binding between the human pose and the three-dimensional human body model, as the key technology for real-time driving, directly affect the reconstruction quality of the model.
[0005] Currently, there are mainly two methods for motion matching and binding: one is manual binding by animators, which has high precision but often takes dozens of hours, resulting in low reconstruction efficiency of the model; the other is automatically binding a defined human skeleton to the three-dimensional human body model, which saves binding time but has low precision. Improper binding will cause poor driving effects, thereby reducing the fidelity of the reconstructed model. Summary of the Invention
[0006] The embodiments of this application provide a method and device for three-dimensional human body reconstruction to improve the reconstruction efficiency and accuracy of three-dimensional models.
[0007] In a first aspect, the embodiments of this application provide a method for three-dimensional human body reconstruction, including:
[0008] Obtaining the motion logic data of the bone nodes in a pre-constructed parametric human body model;
[0009] Perform a rigid transformation on the corresponding key points in the parametric human body model according to the three-dimensional coordinates of multiple preset key points in the pre-constructed driven model to obtain an initial parametric human body model;
[0010] Extract the point cloud data of the driven model and the initial parametric human body model from the front view respectively, and determine the rigid transformation relationship between the models according to the extracted point cloud data;
[0011] Align and fit the driven model and the initial parametric model according to the rigid transformation relationship to obtain a target parametric human body model;
[0012] Bind each vertex of the driven model to the corresponding node of the target parametric human body model according to a preset distance strategy, and determine the motion logic data and skinning weights of the corresponding vertices according to the skinning weights of the bound nodes and the motion logic data of the bone nodes; the skinning weights of the nodes are pre-determined based on the bone nodes;
[0013] Drive the driven model to move according to the motion logic data and skinning weights of each vertex in the driven model to obtain a three-dimensional human body model.
[0014] In a second aspect, an embodiment of the present application provides a reconstruction terminal, including a processor, a memory, a display, and at least one external communication interface. The at least one external communication interface, the display, the memory, and the processor are connected by a bus. The memory stores computer program instructions, and the processor is configured to perform the following operations based on the computer program instructions:
[0015] Obtain the motion logic data of the bone nodes in the pre-constructed parametric human body model through the at least one external communication interface;
[0016] Perform a rigid transformation on the corresponding key points in the parametric human body model according to the three-dimensional coordinates of multiple preset key points in the pre-constructed driven model to obtain an initial parametric human body model;
[0017] Extract the point cloud data of the driven model and the initial parametric human body model from the front view respectively, and determine the rigid transformation relationship between the models according to the extracted point cloud data;
[0018] Align and fit the driven model and the initial parametric model according to the rigid transformation relationship to obtain a target parametric human body model;
[0019] According to a preset distance strategy, each vertex of the driven model is bound to the corresponding node of the target parametric human model, and according to the skinning weights of the bound nodes and the motion logic data of the bone nodes, the motion logic data and skinning weights of the corresponding vertices are determined; the skinning weights of the nodes are determined in advance based on the bone nodes;
[0020] According to the motion logic data and skinning weights of each vertex in the driven model, the driven model is driven to move, and a three-dimensional human model is obtained and displayed by the display.
[0021] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute a three-dimensional human body reconstruction method.
[0022] In the above embodiments of the present application, based on a plurality of preset key points, a parametric human model is rigidly transformed to obtain a fitted initial parametric human model, and according to the point cloud data of the driven model and the initial parametric human model in the front view, the driven model and the initial parametric model are aligned and non-rigidly fitted to obtain a target parametric human model. Further, according to a preset distance strategy, each vertex of the driven model is automatically and quickly and reasonably bound to the corresponding node of the target parametric human model, reducing the binding time and improving the reconstruction efficiency. And according to the skinning weights of the bound nodes and the motion logic data of the bone nodes of the obtained parametric human model, the motion logic data and skinning weights of the corresponding vertices are determined, so that the captured motion logic data can better fit the driven model, avoiding the phenomena of "revealing flaws" and "sliding steps", thereby improving the accuracy of the reconstructed model; according to the motion logic data and skinning weights of each vertex determined after binding, the driven model is driven to move, so that the driven model is consistent with the posture and actions of the reconstruction object, ensuring the authenticity of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 Exemplarily shows the architecture diagram of the real-time driving system provided by the embodiment of the present application;
[0025] Figure 2 Exemplarily shows the overall framework diagram of the three-dimensional human body reconstruction method provided by the embodiment of the present application;
[0026] Figure 3 Exemplarily shows the logic diagram of motion matching and binding in the human body three-dimensional reconstruction method provided by the embodiments of the present application;
[0027] Figure 4 Exemplarily shows the method flowchart of the human body three-dimensional reconstruction provided by the embodiments of the present application;
[0028] Figure 5 Exemplarily shows the schematic diagram of the key points of the model provided by the embodiments of the present application;
[0029] Figure 6 Exemplarily shows the schematic diagram of the extraction of the frontal point cloud of the model provided by the embodiments of the present application;
[0030] Figure 7 Exemplarily shows the structural diagram of the reconstruction terminal provided by the examples of the present application. Detailed implementation manners
[0031] To make the objectives, implementation manners, and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part rather than all of the embodiments of the present application.
[0032] Based on the exemplary embodiments described in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope protected by the claims of the present application. In addition, although the disclosure in the present application is introduced according to one or several exemplary examples, it should be understood that each aspect of these disclosures can also constitute a complete implementation manner alone.
[0033] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the subsequent described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise stated, these terms should be understood in their ordinary and general meanings.
[0034] The terms "first", "second", "third", etc. in the description, claims, and the above-mentioned drawings of the present application are used to distinguish similar or the same kind of objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise indicated. It should be understood that such terms can be interchanged under appropriate circumstances, for example, they can be implemented in an order other than those given in the illustration or description of the embodiments of the present application.
[0035] In addition, the terms "comprise" and "have" and any variations thereof are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to those components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0036] The term "module" as used in this application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0037] In order to clearly describe the embodiments of the present application, the terms of the present application are explained below.
[0038] Bone node: The point at the human body joint in the 3D model.
[0039] Vertex: A 3D model is composed of multiple polygonal meshes (such as triangles or quadrilaterals). Each polygon can be considered a vertex. The more vertices there are, the finer the model.
[0040] Node: It is the point obtained by spatial sampling of the vertices of the three-dimensional model.
[0041] At present, the core technologies of remote social interaction systems involve real-time 3D reconstruction technology, 2D or 3D data encoding and decoding and transmission technology, immersive rendering and display technology, etc. Among them, the 3D data involved in 3D reconstruction technology include vertex data, patch data and texture data that describe 3D geometry. Vertex data includes vertex position, vertex normal and vertex color. Generally, the higher the accuracy of the reconstructed model, the higher the voxel resolution required for 3D reconstruction, the more dramatic the growth of data volume, and the longer the time required for data transmission, which will cause rendering delays and freezes. Therefore, in the absence of mature, efficient and high-fidelity 3D data compression technology, cloud transmission technology has an important impact on the quality of the reconstructed model and the display effect in dynamic 3D reconstruction methods.
[0042] For example, taking a network speed of 30FPS as an example, a voxel with a resolution of 192*192*128 requires a transmission bit rate of 256Mbps. The imaging effect of the model corresponding to this resolution size on the rendering display end is very poor, while a voxel with a resolution of 384*384*384 requires a transmission bit rate of 1120Mbps. The amount of data is growing explosively, and it is difficult to transmit in real time even under ideal network bandwidth conditions.
[0043] With the development of motion capture technology and network technology, in order to reduce the amount of data transmitted, lower the requirements for network bandwidth, and reduce the rendering delay while ensuring the accuracy of the reconstructed model, a method of "pre-modeling + human pose" real-time driving can be adopted to achieve dynamic 3D reconstruction of the human body. First, a 3D human body model is pre-constructed. By using motion matching and binding methods, the motion logic data of the driving model is aligned with the motion logic data of the driven 3D human body model, and the model data and binding relationship are saved at the rendering display end. Then, with the help of 2D or 3D human pose detection algorithms with or without markers, the motion logic data for driving the 3D human body model is extracted and transmitted in real time. Compared with transmitting 3D reconstruction data, the amount of data transmitted is reduced, and the link pressure on the transmission network is reduced. In this way, 3D model holographic communication can be achieved based on a 5G wide area network, meeting the requirements of model accuracy and real-time performance.
[0044] Figure 1 The system architecture diagram of the "pre-modeling + human pose" real-time driving provided by the embodiments of this application is shown in Figure 1 As shown, the system includes three modules: a motion capture module, a motion matching and binding module, and a motion driving module.
[0045] (1) Motion capture module: According to the different motion data capture methods, it can be divided into the motion capture method with markers and the motion capture method without markers.
[0046] The motion capture method with markers is an "intrusive" motion capture method. This method requires the acquisition object to wear different types of sensors at the nodes where the motion logic data is to be obtained, and then by real-time obtaining the signal data of the sensors, the signal data is converted into motion logic data for driving by using a capture algorithm. Among them, the types of sensors worn can be divided into electromechanical, electromagnetic, inertial, and optical. Each type of sensor has its own advantages and disadvantages, but overall, it can provide relatively accurate, robust, and real-time motion capture performance. Since the motion capture method with markers does not require the aid of external imaging equipment, the acquisition of motion data is not restricted by the venue.
[0047] Currently, the commonly used motion capture methods with markers for human body include the optical motion capture method and the inertial motion capture method. With the development of Micro-Electro-Mechanical System (MEMS) technology, the volume and cost of Inertial Measurement Unit (IMU) have been reduced, promoting the application of the inertial motion capture method. By wearing IMUs on different parts of the human body, the spatial acceleration and posture of that part can be obtained in real time, so as to quickly capture the human motion logic data.
[0048] The motion capture method without markers generally obtains the 2D images or 3D point clouds of the acquisition object in real time, uses machine learning or deep learning algorithms to obtain the geometric information of the human skeleton in real time, and then uses the spatial transformation method to convert the obtained geometric information into motion logic data for driving. The motion capture method without markers does not require any sensors to be worn on the acquisition object and will not cause excessive interference or restriction to human movement. This method often uses a single or multiple cameras as sensors and adopts the motion capture method based on computer vision and computer graphics to capture human actions. Since the collected images are taken by cameras, compared with the sensors sparsely attached to human joints, denser information can be obtained. Therefore, the motion capture method without markers is often used to reconstruct the dense surface model of the human body with a higher degree of freedom of non-rigid motion and obtain the human rendering material information.
[0049] Currently, the commonly used motion capture methods without markers can be divided into two categories: markerless action capture based on image recognition and markerless action capture based on template tracking.
[0050] (2) Motion matching and binding module: The commonly used motion matching and binding methods include manual embedding binding and automatic embedding binding.
[0051] Manual embedding binding is to bind the driving nodes of the pre-defined driving model to the vertices on the driven model by professional animators to obtain the binding relationship and skinning weights between each vertex and the driving nodes. Specifically, professional animators use professional binding software to manually bind the nodes of the driving model that provides driving data to the corresponding vertices of the driven model; then verify the driving results, adjust the bound nodes according to the verification results to obtain the best model embedding and binding relationship, and then calculate the motion logic data of each vertex in the driven model according to the adjusted nodes.
[0052] Taking the 3D production software Maya as an example, animators can directly use the software interface or plug-ins for animation production. Specifically, animators bind the human skeleton nodes to the driven model through graphical clicking and dragging operations to embed the human skeleton into the driven model. The software will automatically determine the initial skinning weights. Animators simulate the driving process according to the motion logic data and update and correct the initial skinning weights according to the driving effect to complete the optimization of motion matching and binding.
[0053] Automatic embedding binding is to fit the skeleton information of the preset driving model that generates motion logic data obtained in real time (the bone nodes in the skeleton are connected by the hinge structure of the human body) into the driven model according to algorithms such as machine learning and graphics, and then calculate the skinning weights and motion logic data of each vertex in the driven model through metrics such as distance, so as to drive the entire human body 3D model to move.
[0054] (3) Motion driving module: According to the skinning algorithm and the skinning weights of the vertices, convert the motion logic data of the driving model into the motion logic data of the driven model, that is, obtain the real-time posture and actions of the driving model, so that the driven model makes the same posture and actions as the driving model, achieving the effect of remote driving.
[0055] Based on Figure 1 The real-time driving system architecture shown can ensure accurate motion capture, reasonable matching and binding of the motion logic data of the driving skeleton and the driven model, and a high-efficiency and high-fidelity skinning animation algorithm. It is the core technology for realizing holographic communication based on real-time driving. Motion capture and motion driving are relatively mature, and the key to implementing driving lies in motion matching and binding.
[0056] Currently, in a high-precision real-time driving system, it is necessary to use sensors in the motion capture module to obtain the motion logic data of the marker points, manually embed and bind the marker points to the driven model, and finally apply physical simulation and skinning algorithms to convert the motion logic data captured by the sensors into the motion logic data of the driven model, thus completing real-time driving. The advantages of this system are high driving effect accuracy and vividness, and the binding result and driving result can be adjusted and corrected arbitrarily according to needs. However, manual embedding binding requires a large amount of human and material costs, the motion capture sensors are expensive, the motion matching and binding time is relatively long, and it cannot be automated. For a markerless real-time driving system, usually, the defined driving model and the driven model are automatically bound first, then the motion logic data of the driving model is captured in real time using 2D or 3D human bone point detection algorithms, and finally the real-time driving of the driven model is completed through the skinning algorithm. The advantages of this system are that the driving system is simple, easy to build and use, but the driving effect is not realistic enough. Often due to unreasonable motion matching and binding, abnormal actions and "glitches" occur in the driven model.
[0057] To solve the problems in the motion matching and binding module, an embodiment of the present application provides a method and device for three-dimensional human body reconstruction. In the motion matching and binding process of this method, the simplicity and convenience of a markerless system can be achieved, and at the same time, the rationality and accuracy of a marker-based system can be achieved, realizing a real-time driving effect. Specifically, by means of a parametric human body model, the driving model and the driven model are automatically bound quickly and reasonably. Compared with manual embedding binding, it saves manpower and material resources, reduces the binding time, and thus improves the efficiency of the entire three-dimensional reconstruction. Compared with the existing automatic binding, it improves the accuracy and performance of the binding, making the captured motion data better fit the driven model, avoiding the phenomena of "wardrobe malfunction" and "sliding step", and improving the fidelity of the model. Further, when using the captured motion logic data to drive the driven model, the reconstruction accuracy of the model is improved.
[0058] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0059] Figure 2 This is the overall framework diagram of the three-dimensional human body reconstruction method provided by the embodiment of the present application. As Figure 2 shown, first, through motion capture technology, the posture and action information of the skeleton of the driving model corresponding to the object to be reconstructed are extracted in real time and converted into motion logic data for driving. Then, through a three-dimensional reconstruction algorithm, the driven model is reconstructed, and the parametric human body model is fitted with the driven model twice to embed the parametric human body model into the driven model. According to the principle of the human skeleton hinge, the skeleton information of the driven model is obtained. Spatial sampling is performed on the vertices of the fitted parametric human body model to obtain a node deformation map. The vertices of the driven model are bound to the nodes of the parametric human body model to determine the skinning weights of each vertex. Finally, by means of the parametric human body model, the motion logic data of the driving model is converted into the motion logic data of the driven model, and the linear skinning or dual quaternion skinning algorithm is used to convert the motion logic data into the posture and action of the driven model according to the skinning weights of each vertex.
[0060] Figure 3 This is the logic diagram of motion matching and binding in the three-dimensional human body reconstruction method provided by the embodiment of the present application. As Figure 3 shown, motion matching and binding mainly include three parts, namely: 1) Sampling the parametric human body model to obtain a node deformation map; 2) Fitting the parametric human body model and the driven model; 3) Calculating the skinning weights of the vertices and the motion logic data in the driven model. Among them:
[0061] The first part is to define a parametric human model that can be embedded inside the driven model. Uniformly sample all vertices on the surface of the parametric human model to obtain a node deformation map, where each node is associated with multiple bone nodes. Since the parametric human model defines the human skeleton information (including the index information of bone nodes and the parent-child relationship of each bone node), according to the distance between the vertex and the associated bone node, the skinning weight of each vertex on the parametric human model can be obtained. After uniformly sampling the vertices in space, the skinning weights of all nodes can be obtained. When the driving information (motion logic data) of the human skeleton in the parametric human model is obtained, according to the association relationship between the nodes and the bone nodes, the driving information of each node can be directly obtained.
[0062] The second part is the process of fitting and optimization. Embed the parametric human model into the driven model. The entire fitting includes rigid fitting and non-rigid fitting. First, perform rigid fitting on the parametric human model and the driven model. Then, perform non-rigid fitting on the parametric human model and the driven model. After two rounds of fitting and optimization, the shape parameters and pose parameters of the target parametric human model are obtained. The purpose of fitting is to bind each vertex in the driven model to the nodes in the parametric human model, so that after obtaining the driving information of the nodes in the parametric human model, according to the binding relationship, the driving information of the vertices in the driven model can be obtained.
[0063] The third part is to bind each vertex in the driven model to at least one node in the parametric human model that is the closest in distance according to the distance strategy based on the fitted parametric human model and the driven model. Determine the skinning weight of the corresponding vertex according to the distance between the at least one bound node and the corresponding vertex.
[0064] Through the above-mentioned motion matching and binding, after capturing the motion logic data of the skeleton of the acquisition object, drive the parametric human model to obtain the motion logic data of each node in the node deformation map. Further, convert it into the motion logic data of the driven model, and drive the driven model according to the skinning weight of each vertex.
[0065] Based on the above-mentioned logic block diagram of motion matching and binding, Figure 4 Exemplarily shows the method flow chart of the human three-dimensional reconstruction method provided by the embodiment of the present application. This process is executed by the reconstruction terminal and mainly includes the following steps:
[0066] S401: Obtain the motion logic data of the bone nodes in the pre-constructed parametric human model.
[0067] When executing S401, a parameterized human body model is pre-constructed. The skeletal nodes of the constructed parameterized human body model are matched with the driving model of the object to be reconstructed. When the motion capture module obtains the motion logic data of the skeletal nodes of the driving model corresponding to the object to be reconstructed in real time, the motion logic data of the skeletal nodes in the parameterized human body model can be obtained based on the matching relationship. The motion logic data can be captured using a motion capture method with or without markers. This part is not the focus of this application and will not be elaborated on in detail.
[0068] In an optional embodiment, the parameterized human body model can be an SMPL model or a STAR model. Taking the parameterized human body model as an SMPL model as an example, the expression formula of the SMPL model is:
[0069]
[0070]
[0071] Among them, the represents the human body shape parameters, represents the human body posture parameters, W() represents the linear skinning function, J() represents the function of predicting the positions of different human joints, T represents the human body model grid, B s () represents the influence function of human body shape parameters on human body model mesh T, B p () represents the influence function of human posture parameters on the human body model grid T, T p () represents the function that deforms the human body model mesh T under the joint action of human body shape parameters and human body posture parameters, and s, p, and ω represent shape weight, posture weight, and skin weight, respectively.
[0072] S402: performing rigid transformation on corresponding key points in the parameterized human body model according to the three-dimensional coordinates of a plurality of preset key points in the pre-built driven model to obtain an initial parameterized human body model.
[0073] In an optional embodiment, a three-dimensional reconstruction algorithm is used to pre-construct a driven model (which may also be a digital cartoon model) of the object to be reconstructed based on the pose (e.g., Tpose) of a parameterized human body model. When executing S402, the center point of all vertices in the driven model is determined, and the three-dimensional coordinates of the center point in the world coordinate system are determined. The three-dimensional coordinates of the center point can be used as the initial translation value for the rigid transformation of the parameterized human body model. After obtaining the three-dimensional coordinates of the center point of the driven model, the three-dimensional coordinates of multiple preset key points in the driven model are determined based on the three-dimensional coordinates of the center point according to the preset scaling ratio of the multiple preset key points in the driven model to the center point.
[0074] The selection method of key points is shown in Figure 5 , such as Figure 5 shown. The number of preset key points is 7, which are located at the top of the head, left hand, right hand, left calf, right calf, left waist, and right waist of the model respectively.
[0075] Furthermore, according to the three-dimensional coordinates of multiple preset key points in the driven model, a rough fitting of the parametric human model is performed. Specifically, based on multiple preset key points in the driven model, the corresponding key points in the parametric human model are extracted. According to the three-dimensional coordinates of multiple preset key points in the driven model, a rigid transformation is performed on the corresponding key points in the parametric human model to obtain the initial shape parameters and the pose parameters of the root node in the parametric human model. And according to the initial shape parameters and the pose parameters of the root node, the parametric human model is updated to obtain the initial parametric human model.
[0076] S403: Respectively extract the point cloud data of the driven model and the initial parametric human model from the front view, and determine the rigid transformation relationship between the models according to the extracted point cloud data.
[0077] In order to improve the accuracy and solution speed of the rigid transformation relationship between the driven model and the initial parametric human model, in an optional implementation manner, the point cloud data of the driven model and the initial parametric human model from the front view can be used for solution.
[0078] Specifically, when executing S403, according to the three-dimensional space coordinate system of the model, a two-dimensional projection plane of the driven model and the initial parametric human model is selected. Taking one of the models as an example, the selection of the two-dimensional projection plane is as Figure 6 shown. The positive direction of the model is used as the two-dimensional projection plane, and the point cloud data of the driven model and the initial parametric human model in the two-dimensional projection plane are respectively extracted. Furthermore, the Iterative Closest Point (ICP) algorithm is used to determine the rigid transformation relationship (including the rotation matrix and the translation vector) between the driven model and the initial parametric human model according to the extracted front point cloud data, so as to optimize the initial shape parameters and the pose parameters of the root node of the parametric human model.
[0079] S404: According to the rigid transformation relationship, align and fit the driven model and the initial parametric model to obtain the target parametric human model.
[0080] When performing S404, according to the rigid transformation relationship between the driven model and the initial parameterized human model, the vertices in the driven model and the initial parameterized model are rigidly aligned. According to the preset energy function of the established non-rigid transformation, and based on the vertex data (such as coordinates, normals, etc.) in the driven model, the initially parameterized model after rigid alignment is fitted, and the Gauss-Newton method is used to obtain the target shape parameters of the initially parameterized model and the pose parameters of each node. According to the target shape parameters and the pose parameters of each node, the target parameterized human model is obtained.
[0081] In an alternative embodiment, the preset energy function for iterative solution between the point clouds of the driven model and the initial parameterized model is defined as:
[0082] E loss =E sdata +E pri Equation 3
[0083] where E sdata is a data item used to measure the alignment degree between the corresponding point pairs in the point cloud data of the parameterized human model and the point cloud data of the driven model, and E pri is the prior item of the human pose in the parameterized human model.
[0084] S405: According to the preset distance strategy, each vertex of the driven model is bound to the corresponding node of the target parameterized human model, and according to the skinning weights of the bound nodes and the motion logic data of the bone nodes, the motion logic data and skinning weights of the corresponding vertices are determined.
[0085] When performing S405, each node is obtained by uniformly sampling the vertices on the surface of the target parameterized human model in space. According to the preset distance strategy, each vertex of the driven model is bound to the corresponding node of the target parameterized human model.
[0086] Taking any one vertex among the vertices of the driven model as an example, at least one node in the target parameterized human model that is closest to the vertex is determined, and at least one node is bound to one vertex.
[0087] In the target parametric human model, each node is driven by multiple associated bone nodes. After obtaining the motion logic data of the bone nodes, the motion logic data of the corresponding nodes can be obtained. Moreover, the skinning weight of each vertex is determined based on the distance between the vertex and the associated bone nodes, and is a fixed value. The nodes are obtained by uniformly sampling the vertices in space. Therefore, according to the skinning weights of the sampled vertices, the skinning weights of each node can be obtained. The nodes in the target parametric human model are bound to the vertices of the driven model. Furthermore, according to the nodes bound to each vertex, the motion logic data and skinning weights of each vertex can be determined.
[0088] Taking any one of the vertices in the driven model as an example, the calculation process of the motion logic data and skinning weight of the vertex is described. Specifically, the distances from at least one node bound to the vertex to the vertex are respectively determined, the at least one distance is normalized to obtain the skinning weight of the vertex, and according to the motion logic data of each bone node in the target parametric model associated with at least one node bound to the vertex respectively, and the skinning weights of at least one node respectively, the motion logic data of at least one node respectively is determined. According to the skinning weight of the vertex and the motion logic data of at least one node bound to the vertex, the motion logic data of the vertex is determined.
[0089] Taking the example that one vertex binds n nodes, the calculation formula for the skinning weight of the vertex is:
[0090]
[0091] where i represents the i-th node, n represents the number of nodes bound to one vertex, and d i represents the distance from the i-th node to one vertex. Optionally, n = 4.
[0092] S406: Drive the driven model to move according to the motion logic data and skinning weights of each vertex in the driven model, and obtain a three-dimensional human model.
[0093] In an alternative embodiment, a linear skinning or dual quaternion skinning algorithm is adopted. According to the principle of human bone hinges, the driven model is driven to move according to the motion logic data and skinning weights of each vertex, so that the posture and actions of the driven model are consistent with the object to be reconstructed, and a three-dimensional human model is obtained to complete the three-dimensional reconstruction of the human body.
[0094] It should be noted that the three-dimensional reconstruction method of the embodiments of the present application is also applicable to the reconstruction of local human models. For example, when reconstructing a three-dimensional head model, the parametric models that can be selected are the FLAM model and the 3DMM model; when reconstructing a three-dimensional hand model, the parametric model that can be selected is the MANO model.
[0095] In the above embodiments of the present application, a parameterized human model and a driven model are pre-constructed. Through the matching of the parameterized human model and the driven model, the motion logic data of the skeleton of the driven model captured in real time is converted into the motion logic data of the skeleton of the parameterized human model. Further, based on multiple preset key points and frontal point cloud data, rigid fitting is performed on the parameterized human model and the driven model to align the parameterized human model and the driven model. Then, non-rigid fitting is performed on the parameterized human model and the driven model according to a preset energy function to obtain a target parameterized human model. According to a preset distance strategy, the nodes sampled from the parameterized human model are automatically and quickly and reasonably bound to the vertices of the driven model. Compared with manual binding, it saves manpower and material resources and binding time, thereby improving the model reconstruction efficiency; compared with existing automatic binding, it improves the binding accuracy, reduces the phenomena of "revealing the truth" and "sliding step", and improves the model fidelity; further, according to the binding relationship, the skinning weights and motion logic data of each vertex are determined, so as to drive the driven model, making the captured motion logic data better fit the driven model and improving the model reconstruction accuracy.
[0096] It should be noted that the above terminal in the embodiments of the present application can be a smart phone, a tablet computer, a desktop computer, a notebook computer, a smart TV, and an interactive terminal such as a VR head-mounted display device and an AR glasses.
[0097] Based on the same technical concept, the embodiments of the present application provide a reconstruction terminal. The reconstruction terminal can execute the method flow of human body three-dimensional reconstruction provided by the embodiments of the present application and can achieve the same technical effects, which will not be repeated here.
[0098] See Figure 7 , the reconstruction terminal includes a processor 701, a memory 702, a display 703, and at least one external communication interface 704. The at least one external communication interface 704, the display 703, and the memory 702 are connected to the processor 701 through a bus 705; a computer program is stored in the memory 702, and the processor 701 realizes the following operations by executing the computer program:
[0099] Obtain the motion logic data of the bone nodes in the pre-constructed parameterized human model through at least one external communication interface 704;
[0100] According to the three-dimensional coordinates of multiple preset key points in the pre-constructed driven model, perform a rigid transformation on the corresponding key points in the parameterized human model to obtain an initial parameterized human model;
[0101] Extract the point cloud data of the driven model and the initial parameterized human model from the front view respectively, and determine the rigid transformation relationship between the models according to the extracted point cloud data;
[0102] Align and fit the driven model and the initial parameterized model according to the rigid transformation relationship to obtain the target parameterized human model;
[0103] Bind each vertex of the driven model to the corresponding node of the target parameterized human model according to the preset distance strategy, and determine the motion logic data and skinning weights of the corresponding vertices according to the skinning weights of the bound nodes and the motion logic data of the bone nodes; the skinning weights of the nodes are determined in advance based on the bone nodes;
[0104] Drive the driven model to move according to the motion logic data and skinning weights of each vertex in the driven model, obtain the three-dimensional human model and display it on the display 703.
[0105] Optionally, the processor 701 binds each vertex of the driven model to the corresponding node of the target parameterized human model according to the preset distance strategy, and is specifically configured as:
[0106] Sample the points on the surface of the target parameterized human model to obtain each node of the target parameterized human model;
[0107] For any one vertex among each vertex of the driven model, determine at least one node on the target parameterized human model that is closest to the vertex, and bind the at least one node to the vertex.
[0108] Optionally, the processor 701 determines the motion logic data and skinning weights of the corresponding vertices according to the skinning weights of the bound nodes and the motion logic data of the bone nodes, and is specifically configured as:
[0109] For any one vertex among each vertex, perform the following operations:
[0110] Determine the distances from at least one node bound to a vertex to the vertex respectively, normalize the at least one distance to obtain the skinning weight of a vertex;
[0111] Determine the motion logic data of at least one node respectively according to the motion logic data of each bone node associated with at least one node and the skinning weights of at least one node;
[0112] Determine the motion logic data of a vertex according to the motion logic data of at least one node and the skinning weight of a vertex.
[0113] Optionally, the processor 701 determines the three-dimensional coordinates of multiple preset key points in the following manner:
[0114] Determine the three-dimensional coordinates of the center point of the driven model in the world coordinate system;
[0115] According to the preset scaling ratios between multiple preset key points and the center point in the driven model, and based on the three-dimensional coordinates of the center point, determine the three-dimensional coordinates of the multiple preset key points.
[0116] Optionally, the processor 701 performs a rigid transformation on the corresponding key points in the parametric human model to obtain an initial parametric human model, and is specifically configured as:
[0117] Perform a rigid transformation on the corresponding key points in the parametric human model to determine the initial shape parameters of the parametric human model and the pose parameters of the root node;
[0118] Update the parametric human model according to the initial shape parameters and the pose parameters of the root node to obtain an initial parametric human model.
[0119] Optionally, the processor 701 aligns and fits the driven model and the initial parametric model according to the rigid transformation relationship to obtain a target parametric human model, and is specifically configured as:
[0120] According to the preset energy function of the non-rigid transformation, and based on the vertex data of each in the driven model, fit the aligned initial parametric model to obtain the target shape parameters of the initial parametric model and the pose parameters of each node;
[0121] Obtain a target parametric human model according to the target shape parameters and the pose parameters of each node.
[0122] The embodiment of the present application also provides a computer-readable storage medium for storing some instructions, which can complete the method of the foregoing embodiment when executed.
[0123] The embodiment of the present application also provides a computer program product for storing a computer program, and the computer program is used to execute the method of the foregoing embodiment.
[0124] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0125] For ease of explanation, the above description has been presented in connection with specific embodiments. However, the foregoing exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Numerous modifications and variations are possible in light of the above teachings. The selection and description of the above embodiments were made to better explain the principles and the practical application, so that those skilled in the art can better utilize the embodiments and various different variations suitable for specific use considerations.
Claims
1. A method for reconstructing a three-dimensional human body model, characterized in that, Including: Obtaining the motion logic data of the bone nodes in a pre - constructed parametric human body model; Performing a rigid transformation on the corresponding key points in the parametric human body model according to the three - dimensional coordinates of multiple preset key points in a pre - constructed driven model, to obtain an initial parametric human body model; Respectively extracting the point cloud data of the driven model and the initial parametric human body model from the front - view perspective, and determining the rigid transformation relationship between the models according to the extracted point cloud data; Aligning and fitting the driven model and the initial parametric model according to the rigid transformation relationship to obtain a target parametric human body model; Binding each vertex of the driven model to the corresponding node of the target parametric human body model according to a preset distance strategy, and determining the motion logic data and skinning weights of the corresponding vertices according to the skinning weights of the bound nodes and the motion logic data of the bone nodes; the skinning weights of the nodes are pre - determined based on the bone nodes; Driving the motion of the driven model according to the motion logic data and skinning weights of each vertex in the driven model to obtain a three - dimensional human body model.
2. The method according to claim 1, wherein The step of binding each vertex of the driven model to the corresponding node of the target parametric human body model according to a preset distance strategy includes: Sampling the points on the surface of the target parametric human body model to obtain each node of the target parametric human body model; For any one vertex among each vertex of the driven model, determining at least one node on the target parametric human body model that is closest to the one vertex in distance, and binding the at least one node to the one vertex.
3. The method according to claim 1, characterized in that The step of determining the motion logic data and skinning weights of the corresponding vertices according to the skinning weights of the bound nodes and the motion logic data of the bone nodes includes: Performing the following operations for any one vertex among each vertex: Respectively determining the distances from at least one node bound to the one vertex to the one vertex, normalizing the at least one distance to obtain the skinning weight of the one vertex; Determining the motion logic data of each of the at least one node according to the motion logic data of each bone node associated with each of the at least one node and the skinning weights of each of the at least one node; Determining the motion logic data of the one vertex according to the motion logic data of the at least one node and the skinning weight of the one vertex.
4. The method according to claim 1, wherein Determining the three - dimensional coordinates of the multiple preset key points by the following method: Determining the three - dimensional coordinates of the center point of the driven model in the world coordinate system; According to the preset scaling ratios between multiple preset key points and the center point in the driven model, determining the three - dimensional coordinates of the multiple preset key points based on the three - dimensional coordinates of the center point.
5. The method according to any one of claims 1 to 4, characterized in that, Performing a rigid transformation on the corresponding key points in the parametric human body model to obtain an initial parametric human body model, including: Performing a rigid transformation on the corresponding key points in the parametric human body model, and determining the initial shape parameters of the parametric human body model and the pose parameters of the root node; Update the parameterized human body model according to the initial shape parameters and the pose parameters of the root node to obtain an initial parameterized human body model.
6. The method according to any one of claims 1-4, characterized in that, The aligning and fitting of the driven model and the initial parameterized model according to the rigid transformation relationship to obtain a target parameterized human body model includes: Fitting the aligned initial parameterized model according to the preset energy function of non-rigid transformation and the vertex data in the driven model to obtain the target shape parameters of the initial parameterized model and the pose parameters of each node; Obtain a target parameterized human body model according to the target shape parameters and the pose parameters of each node.
7. A reconstruction terminal, characterized in that, It includes a processor, a memory, a display, and at least one external communication interface. The at least one external communication interface, the display, the memory, and the processor are connected by a bus. The memory stores computer program instructions, and the processor is configured to perform the following operations based on the computer program instructions: Obtain the motion logic data of the bone nodes in the pre-constructed parameterized human body model through the at least one external communication interface; Perform a rigid transformation on the corresponding key points in the parameterized human body model according to the three-dimensional coordinates of multiple preset key points in the pre-constructed driven model to obtain an initial parameterized human body model; Extract the point cloud data of the driven model and the initial parameterized human body model from the front view respectively, and determine the rigid transformation relationship between the models according to the extracted point cloud data; Align and fit the driven model and the initial parameterized model according to the rigid transformation relationship to obtain a target parameterized human body model; Bind each vertex of the driven model to the corresponding node of the target parameterized human body model according to a preset distance strategy, and determine the motion logic data and skinning weights of the corresponding vertices according to the skinning weights of the bound nodes and the motion logic data of the bone nodes; the skinning weights of the nodes are determined in advance based on the bone nodes; Drive the driven model to move according to the motion logic data and skinning weights of each vertex in the driven model to obtain a three-dimensional human body model and display it on the display.
8. The reconstruction terminal according to claim 7, wherein The processor binds each vertex of the driven model to the corresponding node of the target parameterized human body model according to a preset distance strategy, and is specifically configured to: Sample the points on the surface of the target parameterized human body model to obtain each node of the target parameterized human body model; For any one vertex among the vertices of the driven model, determine at least one node on the target parameterized human body model that is closest to the one vertex, and bind the at least one node to the one vertex.
9. The reconstruction terminal according to claim 7, characterized in that, The processor determines the motion logic data and skinning weights of the corresponding vertices according to the skinning weights of the bound nodes and the motion logic data of the bone nodes, and is specifically configured to: Perform the following operations for any one vertex among the vertices: Determine the distances from at least one node bound to the one vertex to the one vertex respectively, normalize the at least one distance, and obtain the skinning weight of the one vertex; Determine the motion logic data of each of the at least one node according to the motion logic data of each bone node associated with each of the at least one node and the skinning weight of each of the at least one node; Determine the motion logic data of the one vertex according to the motion logic data of the at least one node and the skinning weight of the one vertex.
10. The reconstruction terminal according to claim 7, wherein Determine the three-dimensional coordinates of the multiple preset key points in the following manner: Determine the three-dimensional coordinates of the center point of the driven model in the world coordinate system; According to the preset scaling ratios of the multiple preset key points and the center point in the driven model, determine the three-dimensional coordinates of the multiple preset key points based on the three-dimensional coordinates of the center point.
Citation Information
Patent Citations
Human body modeling method and device based on point cloud data stream
CN108961393A
Human body geometric reconstruction method based on Euler field deformation constraint
CN110619681A