Method and device for generating skeleton binding three-dimensional model
Through the autoregressive bone tree generation model and the bone point cross attention mechanism, the problem of insufficient skin weight prediction accuracy is solved, and high-quality bone binding and animation effects are achieved.
Patent Information
- Application Number
- CN202510251262.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-24
AI Technical Summary
In the existing automatic bone binding method, the prediction accuracy of skin weights is insufficient, and the impact relationship between each vertex and the bone cannot be accurately captured, resulting in unsatisfactory animation effect.
By converting the three-dimensional model to be bound into point cloud data, global and local features are extracted, and input into the autoregressive skeleton tree generation model, a skeleton tree structure that conforms to hierarchical relationships is generated. Then, bone point cross attention calculation is performed based on bone features and vertex features to generate high-quality skin weights.
The accuracy of bone binding is significantly improved, and the generated skin weights can accurately capture the influence relationship between each vertex and the bone, thereby improving the naturalness and coordination of the animation effect.
Smart Images

Figure CN120198554A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method and device for generating a three-dimensional model with bone binding. Background Art
[0002] With the rapid development of three-dimensional content generation technology, three-dimensional models are increasingly widely used in fields such as games, movies, and virtual reality. In these applications, the animation production of three-dimensional models is a key link, and bone binding, as the basic step of animation production, directly affects the final animation effect in terms of quality and efficiency. Traditional bone binding methods mainly rely on manual operations, which require professional animators to spend a large amount of time and effort to complete. Not only is the workload huge, but also the dependence on expert experience is relatively high, making it difficult to meet the needs of large-scale animation production.
[0003] Although existing automatic bone binding methods have improved the binding efficiency to a certain extent, there are still some deficiencies. For example, some methods adopt a template matching strategy, which matches the bone structure of the model through a pre-defined template. However, this method has high requirements for the category and topological structure of the model and is difficult to adapt to three-dimensional models with multiple categories and complex topological structures. Other methods perform binding based on local geometric features, determining the position and direction of the bones by analyzing the local geometric information of the model. However, these methods perform poorly in dealing with problems such as topological errors and inaccurate weight prediction, resulting in unsatisfactory animation effects.
[0004] In existing automatic bone binding methods, the prediction of skinning weights is a key link, which determines the association degree between the model vertices and the bones, thereby affecting the deformation effect of the model in the animation. However, existing skinning weight prediction methods often have problems with insufficient accuracy and cannot accurately capture the influence relationship between each vertex and the bones, resulting in unnatural vertex deformation and uncoordinated bone movement during animation driving. Therefore, how to improve the accuracy of skinning weight prediction and generate high-quality skinning weights is an important challenge faced by current automatic bone binding technology. Summary of the Invention
[0005] The present invention provides a method and device for generating a three-dimensional model with bone binding, aiming to solve the defect of insufficient accuracy in skinning weight prediction in the prior art and realize the generation of a three-dimensional model with relatively high bone binding accuracy.
[0006] The present invention provides a method for generating a three-dimensional model with bone binding, including: Obtain a three-dimensional model to be bound, convert the three-dimensional model to be bound into point cloud data, and extract the global features and local features of the point cloud data; Input the global features and local features of the point cloud data into the autoregressive bone tree generation model to obtain the continuous coordinates of the bones. Based on the continuous coordinates of the bones, perform discretization and decoding to obtain a bone tree structure that conforms to the hierarchical relationship; Extract vertex features from the point cloud data through a point cloud feature extraction network, extract bone features of the bone tree through a bone encoder, and perform bone-point cross-attention calculation based on the bone features and the vertex features to generate skinning weights corresponding to each vertex; Bind the skinning weights to the bone tree structure respectively to generate a 3D model with bone binding.
[0007] According to the method for generating a 3D model with bone binding provided by the present invention, extracting the global features and local features of the point cloud data specifically includes: performing normalization processing on the point cloud data; extracting the global features and local features of the normalized point cloud data through a geometric encoder.
[0008] According to the method for generating a 3D model with bone binding provided by the present invention, based on the continuous coordinates of the bones, perform discretization and decoding to obtain a bone tree structure that conforms to the hierarchical relationship, specifically including: Discretize the continuous coordinates of the bones into discrete coordinate values in topological order and add a special identifier for indicating the bone type to obtain a bone tree token sequence; Decode the bone tree token sequence to obtain a bone tree structure that conforms to the hierarchical relationship.
[0009] According to the method for generating a 3D model with bone binding provided by the present invention, based on the bone features and the vertex features, perform bone-point cross-attention calculation to generate skinning weights corresponding to each vertex, specifically including: Input the bone features and the vertex features into a bone-point cross-attention layer to calculate the correlation weights of each vertex feature to each bone feature; Calculate the geometric distance between the calculated vertex features and the bone features; Generate skinning weights corresponding to each vertex according to the correlation weights of each vertex feature to each bone feature and the geometric distance.
[0010] According to the method for generating a 3D model with bone binding provided by the present invention, bind the skinning weights to the bone tree structure respectively to generate a 3D model with bone binding, specifically including: Decode the correlation weights of each vertex to each bone output by the bone-point cross-attention layer through a parameter decoder to obtain bone parameters; Bind the skinning weights corresponding to each vertex to the bone tree structure respectively, and combine the bone parameters to generate a 3D model with bone binding.
[0011] According to the method for generating a skeleton-bound three-dimensional model provided by the present invention, after generating the skeleton-bound three-dimensional model, the method further includes: driving by action data, and optimizing the skeleton-bound three-dimensional model by using a skinning loss function and an action loss function to obtain a final skeleton-bound three-dimensional model.
[0012] The present invention also provides a device for generating a skeleton-bound three-dimensional model, including the following modules: A point cloud feature extraction module, configured to obtain the three-dimensional model to be bound in the dataset, convert the three-dimensional model to be bound into point cloud data, and extract the global features and local features of the point cloud data; A skeleton tree generation module, configured to input the global features and local features of the point cloud data into an autoregressive skeleton tree generation model to obtain the continuous coordinates of the skeleton, and perform discretization and decoding based on the continuous coordinates of the skeleton to obtain a hierarchical skeleton tree structure; A skinning weight determination module, configured to extract vertex features from the point cloud data through a point cloud feature extraction network, extract the skeleton features of the skeleton tree through a skeleton encoder, and perform bone point cross-attention calculation based on the skeleton features and the vertex features to generate the skinning weights corresponding to each vertex; A three-dimensional model generation module, configured to perform skeleton binding on the skinning weights and the skeleton tree structure respectively to generate a skeleton-bound three-dimensional model.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the method for generating a skeleton-bound three-dimensional model as described in any one of the above.
[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for generating a skeleton-bound three-dimensional model as described in any one of the above.
[0015] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method for generating a skeleton-bound three-dimensional model as described in any one of the above.
[0016] The method and device for generating a bone-bound three-dimensional model provided by the present invention convert the three-dimensional model to be bound into point cloud data, and extract the global features and local features of the point cloud data. The global features describe the overall shape and structure of the model, and the local features describe the details and local geometric information of the model. Then, the global features and local features of the point cloud data are input into an autoregressive bone tree generation model to obtain the continuous coordinates of the bones. Based on the continuous coordinates of the bones, discretization and decoding are performed to obtain a bone tree structure that conforms to the hierarchical relationship, effectively reducing redundant information. Cross-attention calculation of bone points is performed based on bone features and vertex features to generate skinning weights corresponding to each vertex, thereby accurately capturing the influence relationship between each vertex and the bones, and thus generating high-quality skinning weights. Finally, the skinning weights are respectively bone-bound with the bone tree structure to generate a bone-bound three-dimensional model. The method of this embodiment can generate a bone tree that conforms to the topological structure and high-quality skinning weights through the autoregressive bone tree generation model and the cross-attention mechanism of bone points, significantly improving the accuracy of bone binding. Brief Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of the method for generating a bone-bound three-dimensional model provided by the present invention.
[0019] Figure 2 It is a schematic diagram of the generation process of a bone-bound three-dimensional model provided by the present invention.
[0020] Figure 3 It is a schematic structural diagram of the device for generating a bone-bound three-dimensional model provided by the present invention.
[0021] Figure 4 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments
[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0023] First, a schematic explanation is given for the noun terms involved in the embodiments of the present invention.
[0024] Skinning Weight: Skinning Weight refers to the degree of association between each vertex and the bones during the skinning process of a 3D model, representing the weight of the vertex affected by the bones. Skinning Weight is used to determine the deformation method of the vertex during the movement of the bones, enabling the model to deform naturally according to the movement of the bones. Through precise skinning weights, high-quality animation effects can be achieved.
[0025] Point Cloud Data: Point Cloud Data refers to a 3D data set composed of a large number of points. Each point contains its coordinate information in 3D space (usually x, y, z coordinates) and possible other attributes (such as normal vectors, colors, etc.). Point Cloud Data is used to represent the geometric shape and structure of a 3D model and is the basis for extracting model features and subsequent processing. Through point cloud data, operations such as feature extraction, model reconstruction, and bone binding can be performed.
[0026] Global Feature: Global Feature refers to the features extracted from the entire point cloud data, which can reflect the overall shape, structure, and layout information of the model. Global features provide context information about the overall shape and structure of the model for the bone tree generation model, helping the model understand the overall features of the model and thus generate a bone tree that conforms to the topological structure.
[0027] Local Feature: Local Feature refers to the features extracted from the local area of the point cloud data, which can reflect the details and local geometric information of the model. Local features are used to describe the details and local geometric information of the model, helping the model determine the specific position and direction of the bones, and thus generate a more accurate bone tree structure.
[0028] Continuous Coordinate: Continuous Coordinate means that in 3D space, the position and direction of the bones are represented by continuous numerical values. These coordinate values are continuous real numbers and can accurately represent the position and direction of the bones in space. Continuous coordinates are used to describe the precise position and direction of the bones in 3D space and are the basis for generating the bone tree structure and performing bone binding.
[0029] Discrete Coordinate: Discrete Coordinate means discretizing the continuous coordinates into discrete values, usually by quantifying or mapping the continuous coordinates into predefined discrete values. Discrete coordinates are used to reduce redundant information and improve generation efficiency. During the bone tree generation process, discretizing the continuous coordinates into discrete coordinates can more efficiently generate and process the bone tree.
[0030] Skeleton Tree: A Skeleton Tree is a tree structure used to represent the skeletal structure of a 3D model, including the hierarchical relationships, connection methods, and topological structures of the bones. The Skeleton Tree is used to describe the skeletal structure of the model and guide the skinning and animation of the model. By generating a Skeleton Tree that conforms to the topological structure, precise control and natural deformation of the model can be achieved.
[0031] Skeleton-Bound 3D Model: A Skeleton-Bound 3D Model refers to a 3D model that is bound to a Skeleton Tree structure, enabling the model to deform according to the movement of the bones. The Skeleton-Bound 3D Model is used to achieve high-quality animation effects. Through skeleton binding, the model can deform naturally according to the movement of the bones, thereby generating realistic animations.
[0032] In existing automatic skeleton binding methods, the prediction of skinning weights is indeed a crucial step, which directly determines the degree of association between the model vertices and the bones, thus affecting the deformation effect of the model in the animation. However, existing skinning weight prediction methods often suffer from insufficient accuracy. The specific manifestations and reasons are as follows: 1) Model complexity and diversity: Many 3D models have complex geometric shapes and details, such as wrinkles, bumps, etc. These complex geometric features make the relationship between vertices and bones more complex and difficult to accurately describe with simple rules or models. Additionally, the topological structures of different models vary greatly, such as those of humans, animals, objects, etc., with their skeletal structures and connection methods being different. Existing skinning weight prediction methods often struggle to adapt to this diversity, resulting in a decline in prediction accuracy when dealing with different types of models.
[0033] 2) Algorithm limitations: Existing methods adopt a template matching strategy, matching the skeletal structure of the model through pre-defined templates. However, this method has high requirements for the model category and topological structure and is difficult to adapt to 3D models with multiple categories and complex topological structures. When the model does not match the template, the prediction accuracy of the skinning weights will decrease significantly.
[0034] In addition, some methods perform binding based on local geometric features, determining the position and orientation of the bones by analyzing the local geometric information of the model. However, these methods do not perform well in dealing with problems such as topological errors and inaccurate weight prediction. Local geometric features often can only describe the local information of the model and are difficult to capture the global topological structure and skeletal relationships, resulting in insufficient prediction accuracy of the skinning weights.
[0035] 3) Insufficient training data: Skin weight prediction usually requires a large amount of training data to learn the relationship between model vertices and bones. However, in practical applications, it is often difficult to obtain high-quality training data. The number of models in the dataset is limited, and the categories and topological structures of the models are relatively simple, making it difficult to cover all possible situations. This results in the model being unable to learn enough information during training, which affects the accuracy of skin weight prediction.
[0036] In addition, the quality of training data annotation will also affect the accuracy of skin weight prediction. If the annotated skin weights are inaccurate or incomplete, the model will learn wrong information during training, resulting in biased prediction results.
[0037] 4) Existing optimization algorithms often have problems such as slow convergence and easy to fall into local optimality when dealing with skin weight prediction problems. This makes it difficult for the model to find the optimal parameters during training, thus affecting the accuracy of skin weight prediction.
[0038] 5) Due to the insufficient accuracy of skin weight prediction, vertices may deform unnaturally during animation. For example, vertices may be stretched, twisted, overlapped, etc., resulting in unsatisfactory deformation of the model.
[0039] In summary, there is the problem of insufficient precision in the existing skin weight prediction method, and the main reason is the joint effect of many factors such as model complexity and diversity, algorithm limitation, data quality and quantity, computing resources and optimization and animation-driven complexity. These problems cause the prediction result of skin weight to be unable to accurately capture the influence relationship between each vertex and skeleton, so that the problems such as unnatural vertex deformation and uncoordinated skeleton movement occur when animation is driven. Therefore, the method of the present embodiment can generate a skeleton tree and high-quality skin weights that meet the topological structure by an autoregressive skeleton tree generation model and a bone point cross attention mechanism, and significantly improves the precision of skeleton binding.
[0040] Combine the following Figure 1 - Figure 2 The method for generating a skeleton-bound three-dimensional model according to an embodiment of the present invention is described. Figure 1 As shown, the method includes the following: Step 101: Acquire a three-dimensional model to be bound, convert the three-dimensional model to be bound into point cloud data, and extract global features and local features of the point cloud data.
[0041] Get the 3D model to be bound Dataset selection: Get the 3D models to be bound from a pre-built 3D model dataset. This dataset contains 3D models of various types and topologies, such as anime characters, animals, mechanical devices, etc., to ensure the versatility and robustness of the method.
[0042] Model preprocessing: Preprocess the 3D model to be bound, including operations such as removing noise, repairing topological errors, and unifying the model format, to ensure the quality and consistency of the model.
[0043] Convert the 3D model to be bound into point cloud data Sampling algorithm selection: Use the Poisson sampling algorithm to uniformly sample point cloud points from the surface of the 3D model to be bound. The Poisson sampling algorithm can generate uniformly distributed point cloud points while maintaining the geometric features of the model, avoiding the problem of overly concentrated or sparse sampling points.
[0044] Sampling density determination: Determine the appropriate sampling density according to the complexity and detail requirements of the model. For models with rich details, increase the sampling density to capture more geometric information; for simple models, appropriately reduce the sampling density to reduce the computational amount.
[0045] Normal vector calculation: For each sampled point, calculate its normal vector. The calculation of the normal vector can be achieved by solving the geometric information of the points within the neighborhood of the sampled point. For example, use the principal component analysis (PCA) method to determine the direction of the normal vector.
[0046] Point cloud data generation: Combine the coordinates and normal vectors of the sampled points into point cloud data (XYZ&Normal), where XYZ represents the three-dimensional coordinates of the point, and Normal represents the normal vector of the point. The point cloud data contains the geometric shape and local surface information of the model, providing a basis for subsequent feature extraction and skeleton generation.
[0047] In this embodiment, the point cloud data is normalized, and then the global features and local features of the point cloud data are extracted.
[0048] Among them, the coordinate values of the point cloud data are scaled to a preset range, such as [0, 1] or [-1, 1]. The normalization process can eliminate the scale differences between different models, make the point cloud data have a consistent scale, and facilitate subsequent processing and analysis. The normal vectors of the point cloud data are normalized to ensure that the length of the normal vector is 1.
[0049] In this embodiment, for the geometric encoder, a multi-layer perceptron (MLP) can be used as the geometric encoder to extract features from the point cloud data. The MLP has strong non-linear modeling capabilities and can effectively extract the high-dimensional features of the point cloud data.
[0050] Then, the normalized point cloud data is input into the geometric encoder to extract the global features of the point cloud data. The global features describe the overall shape and structure of the model, such as the symmetry of the model, the distribution of the main limbs, etc. The global feature extraction process includes: concatenating the coordinates and normal vectors of the point cloud data into an input vector and inputting it into the geometric encoder. The geometric encoder extracts features from the input vector through a multi-layer MLP network to generate a global feature vector. The global feature vector contains the overall geometric information of the model and provides the overall context information for the generation of the bone tree.
[0051] While extracting the global features, the geometric encoder is used to extract the local features of the point cloud data. The local features describe the details and local geometric information of the model, such as the joint positions of the model, the local shapes of the limbs, etc. The local feature extraction process is as follows: perform a neighborhood search on the point cloud data to determine the neighborhood point set of each point. The neighborhood search can use the K-nearest neighbor algorithm or the fixed-radius search algorithm. Then, input the coordinates and normal vectors of each point and its neighborhood points into the geometric encoder to extract the local feature vector. The local feature vector contains the local geometric information of the model and provides detailed geometric information for the generation of the bone tree.
[0052] Through the above steps, the conversion and feature extraction of the point cloud data of the 3D model to be bound are completed, providing a basis for the subsequent generation of the bone tree and the prediction of skinning weights.
[0053] Step 102: Input the global features and local features of the point cloud data into the autoregressive bone tree generation model to obtain the continuous coordinates of the bones, and perform discretization and decoding based on the continuous coordinates of the bones to obtain a bone tree structure that conforms to the hierarchical relationship.
[0054] In this step, the global features and local features of the point cloud data are integrated to form a complete feature vector. The global features describe the overall shape and structure of the model, and the local features describe the details and local geometric information of the model. By integrating these features, comprehensive input information is provided for the autoregressive bone tree generation model.
[0055] The integrated feature vector is normalized to have a unified scale and distribution. The normalization process can eliminate the dimensional differences between different features and improve the training and prediction efficiency of the model. The normalization formula is as follows in Equation (1): (1) where F′ is the normalized feature vector, F is the original feature vector, F min and F max are the minimum and maximum values of the feature vector, respectively.
[0056] In this embodiment, the autoregressive skeleton tree generation model adopts an autoregressive skeleton tree generation model based on the Transformer technical architecture, and the model is trained with a large amount of training data. The training data includes the feature vectors of point cloud data and the corresponding skeleton tree structures. The model gradually optimizes its parameters by learning the relationship between the input features and the skeleton tree structures, and improves the accuracy of generating the skeleton tree.
[0057] The integrated feature vectors are input into the trained autoregressive skeleton tree generation model, and the model gradually generates the continuous coordinates of the bones according to the input features. The generation process is carried out in topological order, and the generation of the bone coordinates of each layer depends on the results of the previous layer, ensuring the rationality of the hierarchical relationship and topological structure of the skeleton tree.
[0058] Specifically, the continuous coordinates of the bones generated by the model include the starting and ending coordinates of the bones, and these coordinates describe the position and direction of the bones in three-dimensional space. The continuous coordinates provide a basis for subsequent discretization and decoding.
[0059] Then, the generated continuous coordinates of the bones are optimized, for example, by methods such as smoothing filtering and denoising, to improve the accuracy and stability of the coordinates. The optimized coordinates can better reflect the geometric structure of the model and the natural distribution of the bones.
[0060] After obtaining the continuous coordinates of the bones, the continuous coordinates of the bones are discretized into discrete coordinate values in topological order. The discretization process reduces redundant information and improves the generation efficiency by mapping the continuous coordinates into a predefined set of discrete values. The discretization formula is as follows in Equation (2): (2) where, C d is the discrete coordinate value, C c is the continuous coordinate value, Q is the discretization quantization step, and C min is the minimum value in the axis direction.
[0061] At the same time, special identifiers representing bone types (such as spring bones, template bones, etc.) are introduced into the discrete coordinate values to form a skeleton tree token sequence. These identifiers help the model distinguish different types of bones and further optimize the generation of the skeleton tree. For example, a certain skeleton tree tokenization process can be expressed as: <bos> <cls><template_bone>M(x1)M(y1)M(z1)…<spring_bone>M(xt)M(yt)M(zt)… <eos> Among them <bos>Indicates the start symbol, <eos>Indicates the end symbol, <cls>It is a global image category feature marker. <spring_bone> and <template_bone> are different types of bone identifiers, and the subsequent M(xi)M(yi)M(zi) are the discretized coordinates of the bones. In this way, the generated bone tree not only has a reasonable topological structure but also can accurately distinguish bones of different categories according to the type identifier, making the subsequent skinning weight prediction more accurate.
[0062] Then, decode the bone tree token sequence to obtain a bone tree structure that conforms to the hierarchical relationship. The decoding process constructs a complete bone tree structure by restoring the discrete coordinate values and bone type identifiers to specific bone positions and orientations. The decoding formula is as follows in Equation (3): S = D(T) where S is the decoded bone tree structure, T is the bone tree token sequence, and D is the decoding function.
[0063] Through the above steps, the generation from the global and local features of the point cloud data to the bone tree structure that conforms to the hierarchical relationship is completed, providing a basis for subsequent skinning weight prediction and bone binding.
[0064] Step 103: Extract vertex features from the point cloud data through a point cloud feature extraction network, extract bone features of the bone tree through a bone encoder, and perform bone point cross-attention calculation based on the bone features and the vertex features to generate skinning weights corresponding to each vertex.
[0065] In this embodiment, a multi-layer perceptron (MLP) is used as the point cloud feature extraction network to extract vertex features from the point cloud data. The MLP network can effectively capture the local and global features of the point cloud data and provide rich information for subsequent attention calculation.
[0066] Input the coordinates and normal vectors of the point cloud data into the point cloud feature extraction network: extract features from the input data through a multi-layer MLP network to generate vertex feature vectors. The vertex feature vectors contain the local geometric information of each vertex and the relationship information with surrounding points, providing a basis for subsequent attention calculation.
[0067] In this embodiment, a graph neural network (GNN) can be used as the bone encoder to extract features from the bone tree structure. The GNN network can effectively process the graph structure data of the bone tree and capture the topological relationship and hierarchical structure between bones.
[0068] When extracting bone features, first input the bone tree structure into the bone encoder; the encoder extracts features from the bone tree through the GNN network to generate bone feature vectors. The bone feature vectors contain information such as the position, orientation, and hierarchical relationship of the bones, providing key information for subsequent attention calculations.
[0069] Then, input the vertex feature vectors and bone feature vectors into the bone-point cross-attention layer to calculate the association weights between the vertex features and the bone features. This mechanism can accurately capture the influence relationship between each vertex and the bone, generating high-quality skinning weights.
[0070] Specifically, the attention layer calculates the association weights of vertex features to bone features through the Query, Key, and Value mechanisms. The specific formula is as follows in Equation (4): (4) where, A ij represents the association weight of vertex i to bone j, V i represents the vertex feature vector, K j represents the bone feature vector, W q and W k are the weight matrices of the query and the key respectively.
[0071] Calculate the geometric distance (geodesic distance) between the vertex features and the bone features and incorporate it as additional feature information into the attention calculation. The calculation formula of the geodesic distance is as follows in Equation (5): (5) where, d ij represents the geodesic distance between vertex i and bone j, V i represents the vertex position, B j represents the bone position.
[0072] Generate the skinning weights corresponding to each vertex according to the association weights and the geodesic distance. The calculation formula of the skinning weights is as follows in Equation (6): (6) where, W ij represents the skinning weight of vertex i to bone j, A ij represents the association weight, d ij represents the geodesic distance, and f is a fusion function used to fuse the association weight and the geodesic distance to generate the skinning weight.
[0073] Through Step 103, the generation of skinning weights from the point cloud data and the bone tree structure is completed, providing key weight information for subsequent bone binding.
[0074] Step 104: Bind the skin weights to the bone tree structure respectively to generate a 3D model with bone binding.
[0075] Input the association weights of each vertex output by the bone point cross-attention layer to each bone into the parameter decoder. The association weights describe the relationship between each vertex and the bone, providing key information for the decoder.
[0076] Specifically, during the decoding process, the parameter decoder decodes the association weights through a neural network (such as a multi-layer perceptron) to generate specific bone parameters. The decoding process can be expressed as the following formula (7): P = D(A) (7) Where P is the decoded bone parameter, A is the association weight, and D is the decoding function.
[0077] The bone parameters include the length, direction, joint angle, etc. of the bone, which are used to drive the animation of the 3D model.
[0078] During the bone binding process, bind the skin weights to the bone tree structure respectively. The skin weights describe the degree of association between each vertex and the bone, and the bone tree structure provides the hierarchical relationship and topological structure of the bones. The binding process can be expressed as the following formula (8): M = f(W, S, P) (8) Where M is the 3D model with bone binding, W is the skin weight, S is the bone tree structure, and P is the bone parameter.
[0079] Through bone binding, each vertex of the model is connected to the corresponding bone, enabling the model to deform according to the movement of the bone.
[0080] During the model optimization process, drive the 3D model with bone binding through action data. The action data includes the movement information of the bones, such as the angle changes of the joints, the rotation and translation of the bones, etc. The driving process can be expressed as the following formula (9): M′ = g(M, A d ) (9) Where M′ is the driven 3D model, A d is the action data, and g is the driving function.
[0081] The action data driving enables the model to deform according to the predetermined animation to verify the effect of bone binding.
[0082] Then, optimize the 3D model with bone binding using the skin loss function and the action loss function. The skin loss function measures the difference between the model deformation and the expected deformation, and the action loss function measures the difference between the model movement and the expected movement. The optimization process can be expressed as the following formula (10): (10) Among them, L s is the skin loss function, L a is the action loss function, θ is the model parameter, M gt and M gt ' are the ground truths of the skin and the drive respectively, and θ* is the optimized model parameter.
[0083] Through optimization, the parameters of the model are adjusted, making the deformation and movement of the model more in line with expectations, and finally obtaining a 3D model with good bone binding.
[0084] Through step 104, the generation of a 3D model with bone binding from skin weights and bone tree structures is completed, and through action data driving and loss function optimization, the final 3D model with bone binding is obtained.
[0085] The method for generating a 3D model with bone binding provided by the embodiments of the present invention converts the 3D model to be bound into point cloud data, and extracts the global features and local features of the point cloud data. The global features describe the overall shape and structure of the model, and the local features describe the details and local geometric information of the model; then the global features and local features of the point cloud data are input into the autoregressive bone tree generation model to obtain the continuous coordinates of the bones, and based on the continuous coordinates of the bones, discretization and decoding are performed to obtain a bone tree structure that conforms to the hierarchical relationship, effectively reducing redundant information; based on the bone features and vertex features, bone point cross-attention calculation is performed to generate the skin weights corresponding to each vertex, so as to accurately capture the influence relationship between each vertex and the bones, thereby generating high-quality skin weights; finally, the skin weights are respectively bone-bound with the bone tree structure to generate a 3D model with bone binding. The method of this embodiment can generate a bone tree that conforms to the topological structure and high-quality skin weights through the autoregressive bone tree generation model and the bone point cross-attention mechanism, significantly improving the accuracy of bone binding.
[0086] To facilitate the understanding of the solution of the embodiments of the present invention, see Figure 2 . For Figure 2 each unit mentioned in Shape encoder: Used to extract the global features and local features of the point cloud data and provide input for bone tree generation.
[0087] Skeleton Tree GPT: An autoregressive bone tree generation model used to generate a bone tree structure.
[0088] Bone tree token sequence: Discretize the continuous coordinates of the bones into a token sequence for decoding to obtain a bone tree structure.
[0089] Skeletal Forest: Multiple generated skeletal tree structures that ultimately merge into a complete skeletal tree.
[0090] Point-wise Shape Encoder: Used to extract vertex features of point cloud data.
[0091] Skeletal Encoder: Used to extract skeletal features of the skeletal tree.
[0092] Weight Decoder: Used to calculate skinning weights.
[0093] Parameter Decoder: Used to decode skeletal parameters.
[0094] Skinning Weight: Describes the degree of association between vertices and bones and is used for bone binding.
[0095] Skeletal Parameter: Describes the geometric and kinematic properties of bones and is used to drive model animation.
[0096] Mesh Animation Driver: The finally generated 3D model with bone binding that can be animated driven according to motion data.
[0097] Based on Figure 2 , the method of the embodiment of the present invention is as follows: 1) Obtain the 3D model to be bound: Obtain the 3D human model to be bound (Input Mesh) from the dataset. This model has a complex geometric shape and topology, such as the limbs, torso, head, etc. of the human body.
[0098] 2) Convert to point cloud data: Convert the 3D human model to point cloud data (XYZ&Normal). Specifically, extract point cloud points from the model surface through a sampling algorithm and calculate the normal vector of each point. The point cloud data contains the geometric information of the model and provides a basis for subsequent feature extraction and bone generation.
[0099] 3) Normalization processing: Perform normalization processing on the point cloud data and scale the point cloud coordinates to a preset range (such as [0, 1] or [-1, 1]). Normalization processing can eliminate the scale differences between different models, make the point cloud data have a consistent scale, and facilitate subsequent processing and analysis.
[0100] 4) Feature extraction: Use the geometric encoder to extract the global features and local features of the point cloud data. The global features describe the overall shape and structure of the model, and the local features describe the details and local geometric information of the model.
[0101] 5) Input features into the autoregressive skeleton tree generation model: Input the global features and local features of the point cloud data into the autoregressive skeleton tree generation model (Skeleton Tree GPT). Based on the Transformer architecture, this model can gradually generate each layer of the skeleton tree according to the input features, ensuring that the skeleton tree has a reasonable hierarchical relationship and topological structure.
[0102] 6) The model generates continuous coordinates of the bones according to the input features. These coordinates describe the position and orientation of the bones in three-dimensional space and are the key information for generating the skeleton tree structure.
[0103] 7) Discretize and decode based on the continuous coordinates of the bones to obtain a skeleton tree structure that conforms to the hierarchical relationship.
[0104] Specifically, discretize the continuous coordinates of the bones into discrete coordinate values in topological order and add a dedicated identifier for indicating the bone type to obtain the skeleton tree token sequence. Then, decode the skeleton tree token sequence to obtain a skeleton tree structure that conforms to the hierarchical relationship. This process effectively reduces redundant information and improves the generation efficiency.
[0105] 8) Extract vertex features and bone features: After generating the skeleton tree structure, extract vertex features from the point cloud data through the point cloud feature extraction network, and at the same time use the bone encoder to extract the bone features of the skeleton tree. Vertex features describe the local geometric information of each vertex, and bone features describe the geometric and topological information of the bones.
[0106] 9) Bone-point cross-attention calculation: Input the bone features and vertex features into the bone-point cross-attention layer to calculate the association weights of each vertex feature to each bone feature. The cross-attention mechanism can accurately capture the influence relationship between each vertex and the bones and generate preliminary skinning weights.
[0107] 10) Combine geodesic distance information: Calculate the geometric distance (geodesic distance) between the vertex features and the bone features, and combine the association weights to generate the final skinning weights corresponding to each vertex. Geodesic distance information can further optimize the skinning weights to make them more conform to the geometric structure of the model.
[0108] 11) Parameter decoding: Decode the association weights of each vertex to each bone output by the bone-point cross-attention layer through the parameter decoder to obtain bone parameters. Bone parameters include the length, direction, joint angle, etc. of the bones and are used to drive the animation of the 3D model.
[0109] 12) Bone binding: Bind the skinning weights to the skeleton tree structure respectively and combine the bone parameters to generate a bone-bound 3D model. Bone binding enables the model to deform according to the movement of the bones and achieve high-quality animation effects.
[0110] 13) Model optimization: Driven by action data, the three-dimensional model with bone binding is optimized using a skinning loss function and an action loss function. The optimization process aims to make the animation effects generated by the model more realistic and natural, and finally obtain a three-dimensional model with good bone binding.
[0111] The generating device of the three-dimensional model with bone binding provided by the embodiments of the present invention will be described below. The generating device of the three-dimensional model with bone binding described below can be correspondingly referred to the generating method of the three-dimensional model with bone binding described above.
[0112] The generating device of the three-dimensional model with bone binding provided by the embodiments of the present invention is shown in Figure 3 , and includes: A point cloud feature extraction module 301, configured to obtain a three-dimensional model to be bound in a dataset, convert the three-dimensional model to be bound into point cloud data, and extract the global feature and the local feature of the point cloud data; A bone tree generation module 302, configured to input the global feature and the local feature of the point cloud data into an autoregressive bone tree generation model to obtain continuous coordinates of bones, and perform discretization and decoding based on the continuous coordinates of the bones to obtain a bone tree structure that conforms to a hierarchical relationship; A skinning weight determination module 303, configured to extract vertex features from the point cloud data through a point cloud feature extraction network, extract bone features of the bone tree through a bone encoder, and perform bone point cross-attention calculation based on the bone features and the vertex features to generate skinning weights corresponding to each vertex; A three-dimensional model generation module 304, configured to perform bone binding on the skinning weights and the bone tree structure respectively to generate a three-dimensional model with bone binding.
[0113] The device of this embodiment can generate a bone tree that conforms to a topological structure and high-quality skinning weights through an autoregressive bone tree generation model and a bone point cross-attention mechanism, significantly improving the accuracy of bone binding.
[0114] Figure 4 The schematic diagram of the physical structure of an electronic device is shown in Figure 4 As shown in the figure, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute the method for generating a bone-bound three-dimensional model, including: obtaining the three-dimensional model to be bound, converting the three-dimensional model to be bound into point cloud data, and extracting the global features and local features of the point cloud data; inputting the global features and local features of the point cloud data into an autoregressive bone tree generation model to obtain the continuous coordinates of the bones, and performing discretization and decoding based on the continuous coordinates of the bones to obtain a bone tree structure that conforms to the hierarchical relationship; extracting vertex features from the point cloud data through a point cloud feature extraction network, extracting the bone features of the bone tree through a bone encoder, and performing bone-point cross-attention calculation based on the bone features and the vertex features to generate the skinning weights corresponding to each vertex; respectively binding the skinning weights to the bone tree structure, and together with the bone parameters, generating a three-dimensional model with bone binding.
[0115] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0116] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for generating a skeleton-bound three-dimensional model provided by each of the above methods, including: obtaining a three-dimensional model to be bound, converting the three-dimensional model to be bound into point cloud data, and extracting the global features and local features of the point cloud data; inputting the global features and local features of the point cloud data into an autoregressive skeleton tree generation model to obtain the continuous coordinates of the skeleton, and performing discretization and decoding based on the continuous coordinates of the skeleton to obtain a skeleton tree structure that conforms to the hierarchical relationship; extracting vertex features from the point cloud data through a point cloud feature extraction network, extracting the skeleton features of the skeleton tree through a skeleton encoder, and performing bone-point cross-attention calculation based on the skeleton features and the vertex features to generate skinning weights corresponding to each vertex; respectively performing skeleton binding on the skinning weights and the skeleton tree structure, and together with the skeleton parameters, generating a skeleton-bound three-dimensional model.
[0117] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the method for generating a skeleton-bound three-dimensional model provided by each of the above methods, including: obtaining a three-dimensional model to be bound, converting the three-dimensional model to be bound into point cloud data, and extracting the global features and local features of the point cloud data; inputting the global features and local features of the point cloud data into an autoregressive skeleton tree generation model to obtain the continuous coordinates of the skeleton, and performing discretization and decoding based on the continuous coordinates of the skeleton to obtain a skeleton tree structure that conforms to the hierarchical relationship; extracting vertex features from the point cloud data through a point cloud feature extraction network, extracting the skeleton features of the skeleton tree through a skeleton encoder, and performing bone-point cross-attention calculation based on the skeleton features and the vertex features to generate skinning weights corresponding to each vertex; respectively performing skeleton binding on the skinning weights and the skeleton tree structure, and together with the skeleton parameters, generating a skeleton-bound three-dimensional model.
[0118] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / cls> < / eos> < / bos> < / eos> < / cls> < / bos>
Claims
1. A method for generating a skeleton-bound three-dimensional model, characterized in that: include: Acquire a three-dimensional model to be bound, convert the three-dimensional model to be bound into point cloud data, and extract global features and local features of the point cloud data; Inputting the global features and local features of the point cloud data into an autoregressive skeleton tree generation model to obtain continuous coordinates of the skeleton, discretizing and decoding the continuous coordinates of the skeleton to obtain a skeleton tree structure that conforms to a hierarchical relationship; Extracting vertex features from the point cloud data through a point cloud feature extraction network, extracting bone features of the bone tree through a bone encoder, performing bone point cross-attention calculation based on the bone features and the vertex features, and generating skinning weights corresponding to each vertex; The skin weights are bone-bound to the skeleton tree structure respectively to generate a skeleton-bound three-dimensional model.
2. The method for generating a skeleton-bound three-dimensional model according to claim 1, characterized in that: Extracting the global features and local features of the point cloud data includes: Performing normalization processing on the point cloud data; The global and local features of the normalized point cloud data are extracted through the geometric encoder.
3. The method for generating a skeleton-bound three-dimensional model according to claim 1, characterized in that: Discretization and decoding are performed based on the continuous coordinates of the skeleton to obtain a skeleton tree structure that conforms to the hierarchical relationship, specifically including: Discretizing the continuous coordinates of the skeleton into discrete coordinate values according to the topological order, and adding a special identifier for indicating the skeleton type to obtain a skeleton tree token sequence; The skeleton tree token sequence is decoded to obtain a skeleton tree structure that conforms to the hierarchical relationship.
4. The method for generating a skeleton-bound three-dimensional model according to claim 1, characterized in that: Bone point cross attention calculation is performed based on the bone features and the vertex features to generate skin weights corresponding to each vertex, specifically including: Input the bone features and the vertex features into the bone point cross attention layer, and calculate the association weight of each vertex feature to each bone feature; The calculated geometric distance between the vertex feature and the bone feature; The skin weight corresponding to each vertex is generated according to the association weight of each vertex feature to each bone feature and the geometric distance.
5. The method for generating a skeleton-bound three-dimensional model according to claim 4, characterized in that: The skin weights are respectively bound to the skeleton tree structure, and together with the skeleton parameters, a skeleton-bound three-dimensional model is generated, specifically including: Decoding the associated weights of each vertex output by the bone point cross attention layer to each bone through a parameter decoder to obtain bone parameters; The skin weights corresponding to each vertex are respectively bound to the skeleton tree structure, and combined with the skeleton parameters to generate a skeleton-bound three-dimensional model.
6. The method for generating a skeleton-bound three-dimensional model according to claim 1, characterized in that: After generating the skeleton-bound three-dimensional model, the method further includes: Driven by the motion data, the skeleton-bound three-dimensional model is optimized using the skinning loss function and the motion loss function to obtain the final skeleton-bound three-dimensional model.
7. A device for generating a skeleton-bound three-dimensional model, characterized in that: include: A point cloud feature extraction module is used to obtain the three-dimensional model to be bound in the data set, convert the three-dimensional model to be bound into point cloud data, and extract the global features and local features of the point cloud data; A skeleton tree generation module, used for inputting the global features and local features of the point cloud data into an autoregressive skeleton tree generation model to obtain continuous coordinates of the skeleton, and discretizing and decoding the continuous coordinates of the skeleton to obtain a skeleton tree structure that conforms to a hierarchical relationship; A skin weight determination module is used to extract vertex features from the point cloud data through a point cloud feature extraction network, extract bone features of the bone tree through a bone encoder, perform bone point cross attention calculation based on the bone features and the vertex features, and generate a skin weight corresponding to each vertex; The three-dimensional model generation module is used to perform skeleton binding on the skin weights and the skeleton tree structure respectively, so as to generate a skeleton-bound three-dimensional model.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for generating a skeleton-bound three-dimensional model as described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a skeleton-bound three-dimensional model as claimed in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a skeleton-bound three-dimensional model as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Skin weight generation method and related device
CN120451353A