Three-dimensional digital human skeleton generation and binding method and device, storage medium and equipment

Through deep learning and graph neural networks, the three-dimensional digital human skeletons and binding methods are automatically generated, and the problems of low bone generation efficiency and low accuracy in the existing technology are solved, efficient and accurate bone binding is achieved, and the rapid iteration and large-scale production of three-dimensional digital humans are promoted.

CN120339552AActive Publication Date: 2025-07-18DONGHUA UNIV

Patent Information

Application Number
CN202510829243.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

In the prior art, the generation and binding of three-dimensional digital human bones is inefficient, manual creation is time-consuming and accurate, and automatic methods still need to be improved.

Method used

The mesh segmentation and graph neural network based on deep learning are used to generate a skeleton, combining mesh shrinkage, curve editing and smooth skin weight algorithms to automatically segment the main body and auxiliary parts, and automatically generate a skeleton and bound skin weight that conforms to the human anatomy structure.

Benefits of technology

It improves the efficiency and accuracy of bone creation, reduces manual participation time, and supports the rapid iteration and large-scale production of three-dimensional digital humans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339552A_ABST
    Figure CN120339552A_ABST
Patent Text Reader

Abstract

The invention relates to a three-dimensional digital human skeleton generation and binding method, device and equipment, and a storage medium, and relates to the technical field of computer graphics and artificial intelligence, and the method comprises the steps: obtaining a three-dimensional digital human grid model; identifying a main body part and an auxiliary part; generating a skeleton of the digital human based on a skeleton generation neural network; binding the skeleton main body part to a corresponding grid vertex to obtain a smooth skin weight; binding the skeleton auxiliary part to a corresponding grid vertex and calculating a smooth skin weight; wherein grids of the main body part are shrunk inwards, a skin area is initialized according to the Euclidean distance, a skeleton influence area is finely adjusted by using a curve tool, and a smooth skin weight of the main body part is obtained by a rigid skeleton smooth skin decomposition algorithm; and the smooth skin weight of the auxiliary part is automatically calculated by adopting a nearest distance method and an Euclidean distance. The problems that in the prior art, the efficiency of manually creating and binding bones is low, and the binding precision of an existing automatic method is not high are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer graphics and artificial intelligence technologies. Specifically, it relates to a method, storage medium, device, and electronic device for intelligent generation and binding of 3D digital human skeletons, especially an intelligent bone binding technology based on deep learning and sketch interaction. Background Art

[0002] With the rapid development of technologies such as virtual reality, gaming, and film and television production, the application demand for 3D digital humans has increased significantly. The technology of 3D digital human skeleton generation and binding is a key link in determining the realism of their motion performance. In traditional methods, the creation and skinning binding of 3D digital human skeletons usually need to be manually completed by professional artists. This method has a high learning cost, is cumbersome to operate, and takes a long time, making it difficult to meet the requirements of rapid iteration and large-scale production. Although existing automated methods have shortened the production time, there is still much room for improvement in the quality of automatically generated skeletons and skinning weights. Summary of the Invention

[0003] The present invention provides a method, storage medium, device, and equipment for 3D digital human skeleton generation and binding, so as to at least solve the problems of low efficiency in manually creating and binding skeletons in the prior art and low binding accuracy of existing automatic methods.

[0004] To achieve the above object, in a first aspect, the present invention provides a method for 3D digital human skeleton generation and binding, which includes the following steps: Step S101, obtaining a 3D digital human mesh model to be bound; Step S102, based on a mesh segmentation neural network, identifying the main part and auxiliary part of the 3D digital human mesh model, where the main part includes the body-attached part, and the auxiliary part includes the non-body-attached part; Step S103, generating a digital human skeleton based on a bone generation neural network, where the skeleton includes the main part and the auxiliary part; Step S104, when the main part of the skeleton is obtained, binding it to the mesh vertices corresponding to the main part of the 3D digital human mesh model to obtain smooth skinning weights that can produce natural deformation; Step S105, when the auxiliary part of the skeleton is obtained, binding it to the mesh vertices corresponding to the auxiliary part of the corresponding 3D digital human mesh model and calculating smooth skinning weights; Among them, the smooth skinning weights of the main part that can produce natural deformation are obtained through four steps: shrinking the main part mesh inward, initializing the skinning area according to the Euclidean distance from the shrunk vertices to the bones, fine-tuning the influence area of each bone using a curve tool, and a rigid bone smooth skinning decomposition algorithm; the smooth skinning weights of the auxiliary part are automatically calculated according to the Euclidean distance from the original mesh vertices to the bones using the nearest distance method.

[0005] Preferably, in step S102, a three-dimensional mesh segmentation network based on the MeshCNN architecture is used. For a quadrilateral mesh, first, without changing the number of mesh vertices, it is triangulated into a triangular mesh; then it is input into MeshCNN; the core convolution operation of the network can be expressed as:

[0006] Among them, represents the current edge, , , , represent the four edges adjacent to e, represents the edge 's feature vector, is a learnable parameter matrix, is an activation function.

[0007] Preferably, in step S103, the skeleton generation neural network is based on the graph neural network architecture. Based on the template skeleton, with the help of the joint prediction module and the bone prediction module, the joint positions and bone connection relationships are respectively inferred; the graph neural network is the backbone neural network shared by the two modules, which is used to learn the depth features of the mesh vertices from the three-dimensional mesh. The input of this backbone neural network includes the mesh vertex positions, vertex normals, and edges.

[0008] Preferably, in step S103, the EdgeConv operator is used as the basic operation unit of the backbone neural network, and its definition is as follows:

[0009] Among them, represents the depth feature of the i-th vertex, represents the 's neighbor of the vertex, MLP represents a multi-layer perceptron, l represents the l-th layer of the graph neural network, represents the parameters learned by the neural network.

[0010] Preferably, the joint prediction module is used to regress and predict the joint positions from the depth features of the three-dimensional mesh vertices, including the main joints and the auxiliary joints; for the prediction of the joint positions, first, the main joint heat map and the auxiliary joint heat map are respectively predicted based on the depth features of the mesh vertices, where n represents the number of three-dimensional mesh vertices, m represents the number of main joints, given the depth features from the backbone neural network, a three-layer MLP neural network is used to predict the joint heat map, which can be defined by the following formula:

[0011] Among them, is a matrix composed of the main joint heat map and the auxiliary joint heat map merged; for the joint position prediction of the main joints, the main joint heat map of the first m channels and the positions of the obtained three-dimensional grid vertices are fitted by the following formula:

[0012] Among them, is the heat value of the j-th joint to the i-th vertex after normalization.

[0013] Preferably, the bone prediction module is used to connect each joint to form a skeleton structure. This module first constructs the main skeleton through a predefined skeleton template of the main joints; subsequently, it adaptively generates the auxiliary skeleton; During the generation process of the auxiliary skeleton, it is generated adaptively by using the bone flow field guidance method; the bone flow field is defined as the bone connection direction vector on the three-dimensional grid vertices, and the vector direction on each vertex points from the sub-joint to the parent joint; the bone flow field is predicted by a three-layer MLP neural network, defined as follows:

[0014] Among them, represents the bone connection direction defined on the three-dimensional grid vertices, is the depth feature of the three-dimensional grid vertices, are the learnable parameters of the above three-layer MLP neural network; After obtaining the bone flow field, a skeleton is constructed by combining the Euclidean distance between joints and the bone flow field, defined by the following formula:

[0015] Among them, represents the total cost of connecting the i-th joint and the j-th joint, and represent the positions of the i-th joint and the j-th joint respectively, represents the set of grid vertices affected by the i-th joint, represents the direction of the bone flow field on the i-th grid vertex; the first term in the above formula represents the Euclidean distance between the i-th joint and the j-th joint, then measures the angle deviation between the bone flow field direction vector and the two joint connection directions ;

[0016] Preferably, in step S104, the following steps are included: Based on the grid contraction algorithm, push the main grid vertices inward to make the grid closely fit the skeleton, thereby obtaining a preliminary definition of the skinning area; Initialize the skinning area according to the Euclidean distance between the vertices after inward contraction and the bones; Use the sketch interactive curve tool to finely adjust the bone influence area to improve the naturalness of the skinning effect in complex areas; Use the smooth skinning weight inverse calculation algorithm based on the target deformation sequence to convert the initial rigid weight into a smooth skinning weight suitable for animation driving, ensuring a natural and realistic animation effect when the bones drive the grid.

[0017] Preferably, in step S105, the nearest distance method is used to bind the bones in the auxiliary area, and the distance is the Euclidean distance from the grid vertex to the bone. Among them, the grid vertices close to the bone are generally more affected by the bone, and the grid vertices far from the bone are less affected by the bone.

[0018] In a second aspect, the present invention also provides a storage medium for generating, binding, and storing the bones of a three-dimensional digital human. The storage medium stores computer-readable execution instructions, wherein the execution instructions can be used to execute the above-mentioned method for generating, binding, and storing the bones of a three-dimensional digital human.

[0019] In a third aspect, the present invention also provides a device for generating, binding, and storing the bones of a three-dimensional digital human, which can execute the above-mentioned method for generating, binding, and storing the bones of a three-dimensional digital human.

[0020] In a fourth aspect, the present invention also provides a device for generating, binding, and storing the bones of a three-dimensional digital human, including at least one processor and a memory communicatively connected to the processor. The memory stores instructions executable by the processor, and the instructions are used to cause the processor to execute the above-mentioned method for generating, binding, and storing the bones of a three-dimensional digital human.

[0021] In summary, the present invention includes the following beneficial technical effects: In step S102 of the present invention, by segmenting the three-dimensional digital human mesh model, the body-fitting part and the non-body-fitting part can be efficiently and accurately distinguished, thereby providing a clear boundary division for the subsequent generation and binding of the bone structure, and further improving the preprocessing efficiency in the bone creation process.

[0022] In step S103 of the present invention, the bone generation method based on the graph neural network can not only automatically generate the main skeleton that conforms to the human anatomical structure, but also generate the auxiliary skeleton that adapts to the changes in the attached area, thereby effectively reducing the time required for manually creating the skeleton.

[0023] For the bone binding of the main body area, in step S104 of the present invention, an accurate initial skin weight can be obtained through the inward contraction of the mesh, and combined with curve editing and smooth skin weight solving methods, the time of manual interaction is effectively shortened while ensuring the accuracy.

[0024] For the bone binding of the auxiliary area, in step S105 of the present invention, an automatic binding strategy based on the Euclidean distance is adopted to realize the automatic weight distribution of non-body-fitting areas such as clothing and hair accessories, further reducing the degree of manual participation and improving the overall binding efficiency.

[0025] Finally, the present invention realizes a three-dimensional digital human intelligent bone generation and binding program, which helps to promote the rapid iteration and large-scale production of three-dimensional digital humans at the application level. Description of the Drawings

[0026] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 It is a schematic diagram of the intelligent generation and binding process of the three-dimensional digital human bone; Figure 2 It is a schematic diagram of obtaining the three-dimensional mesh of the digital human; Figure 3 It is a schematic diagram of identifying the main part and auxiliary part of the digital human based on the mesh segmentation neural network; Figure 4 It is a schematic diagram of generating the skeleton of the digital human based on the bone generation neural network, including the main bone and the auxiliary bone; Figure 5 It is a schematic diagram of binding the main bone to the mesh vertices of the main part based on the four-step method; Figure 6 It is a schematic diagram of binding the auxiliary bone to the mesh vertices of the auxiliary part based on the nearest distance method; Figure 7 It is a schematic diagram of the three-dimensional digital human moving after the bone binding is completed; Figure 8 It is a schematic diagram of the three-dimensional digital human bone generation and binding device; Figure 9 It is a schematic diagram of the electronic device. Detailed Embodiments

[0027] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the scope of protection of this application.

[0028] Embodiment 1: For ease of understanding, explanations are made for the three terms of three-dimensional mesh segmentation, skeleton generation, and skeleton binding: Three-dimensional mesh segmentation is the process of automatically dividing a complex three-dimensional model into multiple meaningful sub-regions according to semantic or geometric features. In this embodiment, it is necessary to divide the three-dimensional digital human model into a main body part and an auxiliary non-body-attached part for differential skeleton binding processing. Traditional methods rely on manual annotation, with low efficiency and strong subjectivity. This application adopts an automatic segmentation method based on deep learning, which can intelligently identify the boundaries of the main and auxiliary regions and provide an accurate regional division for subsequent skeleton generation.

[0029] Skeleton generation refers to automatically inferring the internal skeleton structure of a character based on a three-dimensional mesh model. This skeleton is used to control the animation deformation of the surface skin (three-dimensional mesh) of a three-dimensional digital human, equivalent to the motion driver of a three-dimensional model. In this embodiment, based on graph neural network technology, the geometric features and topological relationships of the three-dimensional mesh are automatically learned to predict the joint positions and skeleton connection relationships. Compared with traditional methods based on the geometric central axis, the neural network method can combine the prior knowledge in a large amount of training data to generate a skeleton that more conforms to the human anatomical structure.

[0030] Skeleton binding technology is the process of determining the influence relationship between mesh vertices and skeletons. The influence degree of each vertex affected by each skeleton is described by a skinning weight matrix. Generally speaking, skeleton binding can also be called skinning. This embodiment combines multiple technologies such as mesh contraction, curve editing, and weight inverse solution based on target deformation to achieve high-quality skeleton binding.

[0031] The embodiment of the present invention provides a method for intelligent generation and binding of three-dimensional digital human skeletons, which specifically includes the following steps: Step S101, obtain a three-dimensional digital human mesh model to be bound. The three-dimensional digital human mesh model can be composed of a human body main body, as well as multiple auxiliary components such as clothing, hair, and accessories; Step S102, use a mesh segmentation neural network to automatically segment the input three-dimensional digital human mesh model to identify the human body main body part and the auxiliary part, as Figure 3 shown. Among them, the human body main body includes key regions such as the torso and limbs, and the auxiliary part includes non-body-attached clothing, hair, accessories, etc.

[0032] In this embodiment, a 3D mesh segmentation network based on the MeshCNN architecture is adopted. This network is designed specifically for triangular mesh data and can effectively handle the irregular topological structure of the mesh. For quadrilateral meshes, the following processing method is adopted in this embodiment: without changing the number of mesh vertices, it is triangulated into a triangular mesh; then, it is input into MeshCNN. MeshCNN extends the concept of traditional convolutional neural networks to triangular meshes and extracts local features by defining convolutional operations on the mesh edges. The network input is the edge features of the triangular mesh, specifically including geometric attributes such as edge length, dihedral angle, and the angle between the edge and the normal vector of the adjacent face. The core convolutional operation of the network can be expressed as:

[0033] where, represents the current edge, 、 、 、 represents the four edges adjacent to e , represents the edge 's feature vector, is a learnable parameter matrix, is an activation function. For the binary classification task of the main part and the auxiliary part, the output layer of the last layer of the network adopts the Sigmoid activation function. Since this task belongs to a binary classification task, in training, this embodiment uses binary cross-entropy loss to supervise the learning of MeshCNN.

[0034] Step S103, based on the skeleton generation neural network, intelligently generate the skeleton structure of the digital human. This skeleton structure includes the main skeleton and the auxiliary skeleton, as shown in Figure 4 . The main skeleton usually adopts a predefined template of 65 joint points, and this main skeleton will be generated for different 3D digital humans; the auxiliary skeleton is dynamically and adaptively generated according to the hairstyle, clothing, accessories, etc. of the 3D digital human.

[0035] Specifically, the skeleton generation neural network in this embodiment is based on the graph neural network architecture. Based on the template skeleton, with the help of the joint point prediction module and the bone prediction module, it automatically and efficiently infers the joint point positions and bone connection relationships respectively.

[0036] Among them, the graph neural network is a backbone neural network shared by two modules, which is used to learn the depth features of grid vertices from the 3D grid. Specifically, the input of this backbone neural network includes information such as grid vertex positions, vertex normals, and edges. In this embodiment, the EdgeConv operator is used as the basic operation unit of the backbone neural network, and its definition is as follows:

[0037] Among them, represents the depth feature of the i-th vertex, represents the neighbors of the vertex, MLP represents a multi-layer perceptron, l represents the l layer of the graph neural network, represents the parameters learned by the neural network. In this embodiment, the max operation is used to aggregate the features of the vertices adjacent to the vertex i . Optionally, other technicians in the art can use other custom operators to aggregate the features of adjacent vertices.

[0038] The above-mentioned joint point prediction module is used to regress and predict the joint positions from the depth features of the 3D grid vertices, including the main joints and auxiliary joints as shown in Figure 4 . The main joints include joints that are indispensable to the human body such as the shoulder joint, elbow joint, and knee joint, and the auxiliary joints include joints adaptively generated on non-body-fitting changing parts such as hair, clothing, and accessories.

[0039] For the prediction of joint positions, in this embodiment, the main joint heatmaps and the auxiliary joint heatmaps are first predicted based on the depth features of the grid vertices. Among them, n represents the number of 3D grid vertices, m represents the number of main joints, and in this embodiment, m = 65, representing 65 human main joints. Optionally, other technicians in the art can set different m values according to actual needs. Given the depth features from the backbone neural network, in this embodiment, a three-layer MLP neural network is used to predict the joint heatmaps, which can be defined by the following formula:

[0040] Among them, is a matrix composed of the combination of the main joint heatmap and the auxiliary joint heatmap . For the prediction of the joint positions of the main joints, in this embodiment, the main joint heatmaps of the first m channels and the positions of the obtained 3D grid vertices are used, and are fitted by the following formula:

[0041] Among them, is the heat value of the j -th normalized joint pair to the i -th vertex. Regarding the joint positions of the auxiliary joints, since the number of auxiliary joints is uncertain, in this embodiment, the positions of the auxiliary joints are obtained through a clustering algorithm. First, the heat map of the auxiliary joints is thresholded to screen out important grid vertices. Secondly, the screened vertices are used as the input of the mean shift clustering algorithm. Finally, the clustering center is used as the position of the auxiliary joints. Optionally, other technicians in the art can use other clustering algorithms to obtain the positions of the auxiliary joints. During the training process, this embodiment uses the Dice Loss of the heat map and the mean squared error loss (MSE Loss) of the joint positions for supervision.

[0042] The above-mentioned bone prediction module is used to connect each joint to form a skeleton structure. This module first constructs the main skeleton through a predefined skeleton template of 65 body joints, as Figure 4 shown in the middle; subsequently, it adaptively generates auxiliary skeletons for parts such as hair, clothing, and accessories, as Figure 4 shown on the right. During the generation process of the auxiliary skeletons, this embodiment uses the method guided by the bone flow field to generate adaptively. The bone flow field is defined as the bone connection direction vector on the three-dimensional grid vertices, and the vector direction at each vertex points from the child joint to the parent joint. This embodiment predicts the bone flow field through a three-layer MLP neural network, which is defined as follows:

[0043] Among them, represents the bone connection direction (i.e., the bone flow field) defined on the three-dimensional grid vertices, is the depth feature of the three-dimensional grid vertices, are the learnable parameters of the above-mentioned three-layer MLP neural network. After obtaining the bone flow field, this embodiment then constructs the skeleton by combining the Euclidean distance between joints and the bone flow field, which is defined by the following formula:

[0044] Among them, represents the total cost of connecting the i -th joint and the j -th joint. The smaller the cost, the greater the possibility that the two joints are directly connected to form a bone. and respectively represent the positions of the i -th joint and the j -th joint, represents the set of grid vertices affected by the i -th joint, represents the direction of the bone flow field at the i -th mesh vertex. The first term in the above formula represents the Euclidean distance between the i -th joint and the j -th joint. measures the angle deviation between the direction vector of the bone flow field and the connection direction of the two joints .

[0045] In the embodiment of the present invention, all auxiliary joints are first attached to the corresponding template joints, and then the auxiliary-level joints are reattached to another auxiliary joint according to the connection cost; that is, if the cost of forming a new skeleton connection is less than the current cost, the auxiliary-level joint is reattached to another auxiliary joint, and so on, until the entire skeleton is constructed. During the training process, the MSE Loss of the bone flow field is used for supervision in this embodiment.

[0046] Step S104, when the main skeleton is obtained, bind it to the mesh vertices of the corresponding main part, and use the four steps of "shrinking the main part mesh inward, initializing the skinning area according to the Euclidean distance from the shrunk vertices to the skeleton, fine-tuning the influence area of each bone using the curve tool, and the rigid bone smooth skinning decomposition algorithm" to obtain smooth skinning weights that can generate natural deformations; As Figure 5 shown, this embodiment adopts the following four-step process: The first step, based on the mesh contraction algorithm, push the main mesh vertices inward to make the mesh closely fit the skeleton, so as to obtain a preliminary definition of the skinning area; The second step, initialize the skinning area according to the Euclidean distance between the shrunk vertices and the skeleton; The third step, use the sketch interactive curve tool to finely adjust the bone influence area to improve the naturalness of the skinning effect in complex areas such as the shoulders and armpits; The fourth step, use the smooth skinning weight inverse solution algorithm based on the target deformation sequence to convert the initial rigid weights into smooth skinning weights suitable for animation driving, ensuring a natural and realistic animation effect when the skeleton drives the mesh.

[0047] This skinning data generation method improves the bone binding efficiency and quality, and reduces the need for manual fine-tuning.

[0048] For the mesh contraction algorithm in the above first step, this embodiment adopts a Voronoi-guided mesh contraction algorithm. Given the initial mesh vertex positions, this algorithm finds the optimal shrunk mesh vertex positions by balancing the contraction force, in-situ gravitational force, and medial axis gravitational force 。The contraction energy forces the vertices to move inward along the inverse normal direction of the vertices, promoting the emergence of the skeletal structure. At the same time, the attraction energy anchors the vertices in their original positions as a balancing force to the contraction energy to prevent excessive displacement of the vertex positions during contraction. The medial axis energy encourages the vertices to move towards the medial axis position of the mesh, and the optimized objective formula is expressed as follows:

[0049] where, is the discrete Laplace operator, which is a key component in the mesh contraction process because it is a measure of surface curvature and promotes the inward movement of vertices. In this embodiment, the Cotangent Laplace is adopted. Optionally, other forms of Laplace operators can be adopted by other technicians in the art. represents the Voronoi pole associated with the i -th vertex.

[0050] For the skinning area initialization algorithm in the above second step, in this embodiment, based on the Euclidean distance, the mesh vertices are assigned to the bone closest to it, and the distance formula is defined as follows:

[0051] where, is the angle between the bone and the vector ; the third term in the above formula represents the distance from the mesh vertex to the line segment when is projected onto the line segment; the first two terms in the above formula represent the Euclidean distance from the mesh vertex to the joint point when is projected outside the line segment.

[0052] For the skinning area optimization method based on the sketch curve tool in the above third step, the purpose of this embodiment is to make the control areas of each bone more accurate. Specifically, the implementation of this curve tool consists of two steps: curve initialization and editing the curve to modify the skinning area. First, this embodiment automatically generates the initial boundary curves between the influence areas of different bones, using cubic Hermite spline curves. The mathematical formula of this curve includes a series of basis functions, and each basis function only affects a small part of the curve. The cubic Hermite spline curve segment between each pair of control points is defined as:

[0053] where, represents the parameter varying along the curve, and are the starting control point and the ending control point of the spline curve segment respectively, and is the tangential vector of these control points. During the process of constructing the boundary curve between different influence regions, all the edges located on the boundary of the two regions are first identified to form a continuous edge band. The midpoints of these edges are regarded as potential control points of the spline curve. Since the boundary curve is mostly circular and conforms to the circumference formula of a circle , generally 6 control points are equally divided in this embodiment. Optionally, other technicians in the art can select different numbers of control points according to requirements.

[0054] The boundary spline curve obtained using the above formula is generally not on the 3D mesh surface. Next, the interpolated curve points need to be mapped onto the mesh surface. For each interpolated point, the algorithm initially attempts to directly project it onto the last successfully mapped 3D mesh patch. The projection direction is the opposite direction of the average normal vector of the triangular faces corresponding to the two control points. If no suitable mapping point is found on the current face, the algorithm will iteratively search for adjacent triangular patches until a suitable mapping point is found. Each time a suitable mapping point is found, the corresponding spline curve segment will be adjusted to ensure that the curve smoothly passes through the mesh surface. Finally, a set of curves that closely adhere to the mesh surface will be generated, representing the boundary lines of the bone influence regions. Secondly, this embodiment provides a simple and intuitive curve editing function that allows users to precisely adjust the regions of the skinning weights. Users can edit the curve by adding, deleting, and dragging control points, or create new curves by drawing sketch curves and extension lines. After these modifications are completed, the new influence region of the bone will become the region enclosed by these adjusted curves.

[0055] Regarding the smooth skinning weight inverse solution algorithm based on the target deformation sequence in the above fourth step, in this embodiment, the target deformation sequence is used as the input. This target deformation sequence is obtained by the advanced deformation algorithm Direct Delta Mush, and then based on the SSDR (Smooth Skinning Decomposition for Rigid Bones) algorithm, the linear blend skinning (LBS) parameters are inversely deduced from the target deformation sequence. These parameters include the smooth skinning weights and the affine transformation of each frame of the bone , . Mathematically, LBS can be expressed as:

[0056] where represents the position of the i -th vertex in the initial pose, and represents the position of the i -th vertex after deformation by LBS. This position is affected by the rotation and translation of the j-th bone The influence of represents the influence weight of the j th bone on the i th vertex. Given a high-quality target deformation sequence, the SSDR algorithm searches for the optimal smoothing weight matrix to minimize the difference between the target deformation and the LBS deformation. The optimization objective is expressed as:

[0057] where is the position of the th vertex after target deformation at time

[0058] Step S105, for generating the skinning data of the auxiliary part, the nearest distance method is adopted. The smoothing skinning weights of the auxiliary part are automatically calculated according to the Euclidean distance from the original mesh vertices to the bones, and the auxiliary mesh vertices are automatically bound to the nearest auxiliary bones to ensure that the auxiliary part presents appropriate dynamic effects during movement. Optionally, other technicians in the art can use other automatic skinning algorithms such as heat diffusion and Bounded Bi-harmonic Weights to bind the bones in the auxiliary area. Combining with Step S104, a fully bound 3D digital human is obtained. The bound 3D digital human can be animated-driven to achieve rich and smooth motion performances, such as Figure 7 shown. In this embodiment, the LBS method adopted by the mainstream engine is used to realize the driving of the bones on the mesh. Optionally, other technicians in the art can use other deformation driving methods such as Dual Quaternion Skinning (DQS).

[0059] In Step S102 of the present invention, by segmenting the 3D digital human mesh model, the body-fitting part and the non-body-fitting part can be efficiently and accurately distinguished, thus providing a clear boundary division for the subsequent bone structure generation and binding, and further improving the pre-processing efficiency in the bone creation process.

[0060] In Step S103 of the present invention, the bone generation method based on graph neural network is adopted, which can not only automatically generate the main skeleton that conforms to the human anatomical structure, but also generate the auxiliary skeleton that adapts to the changes in the attached area, thus effectively reducing the time required for manually creating the skeleton.

[0061] For the bone binding of the main body area, in step S104 of the present invention, accurate initial skinning weights can be obtained through the method of grid inward contraction, and combined with curve editing and smooth skinning weight solving methods, which effectively shortens the time of manual interaction while ensuring accuracy.

[0062] For the bone binding of the auxiliary area, in step S105 of the present invention, an automatic binding strategy based on Euclidean distance is adopted to realize automatic weight assignment for non-body-fitting areas such as clothing and hair accessories, further reducing the degree of manual participation and improving the overall binding efficiency.

[0063] Finally, the present invention realizes a three-dimensional digital human intelligent bone generation and binding program, which helps to promote the rapid iteration and large-scale production of three-dimensional digital humans at the application level.

[0064] Embodiment 2: This embodiment provides a computer program product for three-dimensional digital human bone intelligent generation and binding, including a computer program, which can implement the method provided in Embodiment 1 when executed by at least one processor.

[0065] Embodiment 3: This embodiment provides a storage medium for three-dimensional digital human bone intelligent generation and binding. The storage medium is a non-transitory computer-readable storage medium storing computer-readable execution instructions, wherein the execution instructions can be used to make a computer execute the method provided in Embodiment 1.

[0066] Embodiment 4: This embodiment provides a three-dimensional digital human bone intelligent generation and binding device, which can execute the method provided in Embodiment 1, specifically including: A three-dimensional mesh segmentation unit for performing semantic segmentation on the input three-dimensional digital human mesh model to identify the main part and the auxiliary part of the input three-dimensional mesh. The three-dimensional meshes of the main part and the auxiliary part will be used for the subsequent bone generation unit; A bone generation unit for automatically inferring and generating the main bones and auxiliary bones of the digital human based on the segmentation result and the graph neural network. Among them, the main bones are generated based on the three-dimensional mesh of the above-mentioned main part, and the auxiliary bones are generated based on the three-dimensional mesh of the above-mentioned auxiliary part. The main bones are the common skeletons of different digital humans, while the auxiliary bones vary with the hairstyles, clothing, and accessories of the digital human; A bone binding unit for binding the generated bones to the corresponding three-dimensional mesh vertices, including a four-step skinning binding process for the main bones and a nearest distance binding method for the auxiliary bones, so that the bones can drive the surface mesh to deform during movement. Among them, the above-mentioned main bones are bound by the four-step method of "grid inward contraction → initialize the skinning area → fine-tune the bone influence area with the curve tool → rigid bone smooth skinning decomposition algorithm".

[0067] In some alternative embodiments, the mesh segmentation unit may include: a MeshCNN segmentation subunit for extracting edge features and completing binary or multi-class classification; the skeleton generation unit may include a GNN backbone neural network subunit, a joint prediction subunit, and a bone flow field prediction subunit, which are respectively used for predicting the positions of main joints, auxiliary joints, and the bone connection directions; the bone binding unit includes a main body binding subunit and an auxiliary part binding subunit, wherein the main body part binding subunit can be further subdivided into: a mesh contraction subunit for inwardly contracting the main body mesh vertices along the reverse normal direction; a skinning area initialization subunit for dividing the initial binding area according to the Euclidean distance between the vertices after inward contraction and the bones; a curve editing subunit for finely adjusting the boundaries of the influence areas of each bone through sketch interactive curves; and a smooth weight optimization subunit for inversely solving the smooth skinning weights based on the target deformation sequence.

[0068] Embodiment 5: A three-dimensional digital human skeleton intelligent generation and binding electronic device according to an embodiment of the present invention includes: At least one processor, at least one memory communicatively connected to the processor via a bus, the memory storing instructions executable by the at least one processor, so that the at least one processor can execute the method provided in Embodiment 1.

[0069] The electronic device may further include a display, an input / output interface, a communication unit, etc., for realizing human-computer interaction and data network transmission.

[0070] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for generating and binding a three-dimensional digital human skeleton, characterized in that, It includes the following steps: Step S101: Obtain the 3D digital human mesh model to be bound. Step S102: Based on the mesh segmentation neural network, identify the main part and the auxiliary part of the 3D digital human mesh model. The main part includes the body-fitting part, and the auxiliary part includes the non-body-fitting part. Step S103: Based on the bone generation neural network, generate the skeleton of the digital human. The skeleton includes the main part and the auxiliary part. Step S104: When the main part of the skeleton is obtained, bind it to the mesh vertices corresponding to the main part of the 3D digital human mesh model to obtain smooth skinning weights that can generate natural deformation. Step S105: When the auxiliary part of the skeleton is obtained, bind it to the mesh vertices corresponding to the auxiliary part of the corresponding 3D digital human mesh model and calculate the smooth skinning weights. Among them, the smooth skinning weights of the main part that can generate natural deformation are obtained through four steps: inwardly contracting the main part mesh, initializing the skinning area according to the Euclidean distance from the inwardly contracted vertices to the bones, fine-tuning the influence area of each bone using the curve tool, and the rigid bone smooth skinning decomposition algorithm; the smooth skinning weights of the auxiliary part are automatically calculated according to the Euclidean distance from the original mesh vertices to the bones using the nearest distance method.

2. The three-dimensional digital human skeleton generation and binding method according to claim 1, wherein In step S102, the skeleton generation neural network is based on the graph neural network architecture. Based on the template skeleton, with the help of the joint point prediction module and the bone prediction module, the joint point positions and bone connection relationships are respectively inferred; the graph neural network is the backbone neural network shared by the two modules and is used to learn the depth features of the mesh vertices from the 3D mesh. The input of this backbone neural network includes the mesh vertex positions, vertex normals, and edges.

3. The three-dimensional digital human skeleton generation and binding method according to claim 2, wherein, In step S102, a 3D mesh segmentation network based on the MeshCNN architecture is used. For the quadrilateral mesh, first, without changing the number of mesh vertices, it is triangulated into a triangular mesh; then it is input into the MeshCNN; the core convolution operation of the network can be expressed as: Among them, represents the current edge, , , , represent the four edges adjacent to e, represents the edge 's eigenvector, is a learnable parameter matrix, is an activation function.

4. The three-dimensional digital human skeleton generation and binding method according to claim 3, wherein, In step S103, the EdgeConv operator is used as the basic operation unit of the backbone neural network, and its definition is as follows: Among them, represents the depth feature of the i-th vertex, represents the neighbors of the vertex, MLP represents a multi-layer perceptron, and l represents the l-th layer of the graph neural network, represents the parameters learned by the neural network.

5. The three-dimensional digital human skeleton generation and binding method according to claim 4, characterized in that The joint point prediction module is used to regress and predict joint positions from the depth features of 3D mesh vertices, including body joints and auxiliary joints. For the prediction of joint positions, first, the body joint heatmaps are predicted based on the depth features of the mesh vertices and the auxiliary joint heatmaps , where n represents the number of 3D mesh vertices and m represents the number of main joints. Given the depth features from the backbone neural network , a three-layer MLP neural network is used to predict the joint heatmaps, which can be defined by the following formula: Among them, is a matrix composed of the main joint heat map and the auxiliary joint heat map For the joint position prediction of the main joint, the main joint heat map of the first m channels and the position of the obtained three-dimensional grid vertices are fitted by the following formula: Among them, is the heat value of the j-th joint after normalization for the i-th vertex.

6. The three-dimensional digital human skeleton generation and binding method according to claim 5, wherein The bone prediction module is used to connect each joint to form a skeleton structure. This module first constructs the main bones through the skeleton template of the predefined main joints; subsequently, the auxiliary bones are adaptively generated. During the generation process of the auxiliary bones, the bone flow field guidance method is used for adaptive generation. The bone flow field is defined as the bone connection direction vector on the 3D mesh vertices. The vector direction at each vertex points from the sub-joint to the parent joint; the bone flow field is predicted by a three-layer MLP neural network, and the definition is as follows: Among them, represents the bone connection direction defined on the vertices of the three-dimensional mesh, is the depth feature of the vertices of the three-dimensional mesh, are the learnable parameters of the above three-layer MLP neural network; After obtaining the bone flow field, the skeleton is constructed by combining the Euclidean distance between joints and the bone flow field, and the definition is as the following formula: Among them, represents the total cost of connecting the i-th joint and the j-th joint, and represent the positions of the i-th joint and the j-th joint respectively, represents the set of mesh vertices affected by the i-th joint, represents the direction of the bone flow field at the i-th mesh vertex; the first term of the above formula represents the Euclidean distance between the i-th joint and the j-th joint, then measures the bone flow field direction vector and the angular deviation between the two joint connection directions 。 7. The method for generating and binding a three-dimensional digital human skeleton according to claim 6, characterized in that In step S104, it includes the following steps: Based on the mesh contraction algorithm, push the main mesh vertices inward to make the mesh closely fit the skeleton, thereby obtaining the initial definition of the skinning area. Initialize the skinning area according to the Euclidean distance between the inwardly contracted vertices and the bones. Use the sketch interactive curve tool to finely adjust the bone influence area to improve the naturalness of the skinning effect in complex areas. Using the smooth skinning weight inverse calculation algorithm based on the target deformation sequence, the initial rigid weights are converted into smooth skinning weights suitable for animation driving, ensuring a natural and realistic animation effect when the skeleton drives the mesh.

8. A three-dimensional digital human skeleton generation and binding storage medium, characterized in that, The storage medium stores computer-readable execution instructions, wherein the execution instructions can be used to execute the three-dimensional digital human skeleton generation and binding method according to any one of claims 1-7.

9. A three-dimensional digital human skeleton generation and binding device, characterized in that, It can execute the three-dimensional digital human skeleton generation and binding method according to any one of claims 1-7.

10. A three-dimensional digital human skeleton generation and binding device, characterized in that, It includes at least one processor and a memory communicatively connected to the processor. The memory stores instructions executable by the processor, and the instructions are used to cause the processor to execute the three-dimensional digital human skeleton generation and binding method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Three-dimensional virtual character intelligent skinning method

    CN112802161A

  • Automatic skin covering method and device for character grid model based on neural network

    CN113240815A

  • Three-dimensional model skeleton binding method and device, equipment and storage medium

    CN116912433A

Cited By

  • Self-adaptive perception digital human skeleton binding method and system based on dynamic graph convolution

    CN120726193A

  • Adaptive perceptual digital human skeleton binding method and system based on dynamic graph convolution

    CN120726193B

  • Virtual digital human generation method and system based on modular parameters

    CN121437698A