A cloth simulation method and device based on physics-aware deep learning

Through the method based on physical perception deep learning, the existing fabric simulation methods have solved the problem of high computational complexity and low realism, and high accuracy fabric simulation is achieved, enhancing the authenticity and dynamic continuity of clothing.

CN118940596BActive Publication Date: 2025-07-01BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410957177.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-17
Publication Date
2025-07-01
Estimated Expiration
2044-07-17

AI Technical Summary

Technical Problem

When generating clothing deformation effects, existing fabric simulation methods have high computational complexity and have limitations on real-time animation. Learning-based simulation requires a large amount of offline data and ignore the physical properties of fabrics, resulting in low realism and poor dynamic continuity.

Method used

The cloth simulation method based on physical perception deep learning is adopted. By obtaining the motion parameters of the 3D character model, it is converted into static and dynamic descriptors. After encoding processing, the cloth deformation state is generated through the decoder, and the network is trained using the loss function based on the physical simulation to enable it to learn to satisfy the physical constraints of the fabric and human body.

Benefits of technology

It improves the accuracy of fabric simulation, enhances the authenticity and dynamic continuity of clothing, reduces dependence on real ground data, and realizes realistic and natural clothing simulation and human-computer interaction in virtual scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118940596B_ABST
    Figure CN118940596B_ABST
Patent Text Reader

Abstract

This paper provides a fabric simulation method and device based on physics-aware deep learning. The method includes: obtaining the motion action parameters of a skinned 3D human model; converting the motion action parameters into a static descriptor and a dynamic descriptor through a decoupling descriptor; respectively encoding the static descriptor and the dynamic descriptor through an encoder to obtain a static latent variable and a dynamic latent variable, and adding them to obtain an encoded feature vector; decoding the feature vector through a decoder to obtain the deformation state of the local fabric of the 3D human model; training the network model through a physics simulation-based loss function, so that the network learns to satisfy the physical constraints between the fabric and the human body, and outputs the prediction result of the fabric state. The purpose of this paper is to use a physics simulation-based loss function to enable the network to learn to satisfy the physical constraints between the fabric and the human body, and to achieve accurate prediction of fabric dynamics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This article belongs to the field of computer technology, and specifically relates to a cloth simulation method and device based on physically-aware deep learning. Background Art

[0002] Cloth simulation has always been a research focus in the field of computer graphics and plays a key role in fields such as animation and games. Currently, the mainstream cloth simulation methods are physics-based simulation and learning-based deformation. These methods learn from each other, do not conflict with each other, and each has its own advantages and disadvantages.

[0003] Physics-based simulation obtains the corresponding cloth state by constructing a model of cloth movement and solving complex dynamic equations. This method can generate realistic and delicate clothing deformation effects, but has a high computational complexity and requires high computational costs, and still has certain limitations for real-time cloth animation.

[0004] Learning-based simulation usually uses a neural network model to learn the cloth deformation law from a large amount of data, so as to infer new data. Compared with the physical method, this type of method has high efficiency and high automation. However, currently most learning-based simulations rely on supervised learning, and a large number of offline physics-based simulations need to be run to collect the data required for training, and the data of various types of parameters need to be collected repeatedly, which damages the scalability of the supervised solution. In addition, the learning process usually only considers the difference between the vertex position and the target, ignoring the physical properties of the cloth itself, resulting in low clothing realism, poor dynamic continuity, weak detail fidelity, and possibly very different output results for clothing simulation on very similar body movements, and the reliability is average. Summary of the Invention

[0005] Aiming at the above problems of the prior art, the purpose of this article is to provide a cloth simulation method and device based on physically-aware deep learning to improve the accuracy of cloth simulation.

[0006] To solve the above technical problems, the specific technical solutions of this article are as follows:

[0007] On the one hand, this article provides a cloth simulation method based on physically-aware deep learning, and the method includes:

[0008] Obtain the motion action parameters of a skinned 3D human model;

[0009] Convert the motion action parameters into static descriptors and dynamic descriptors through a decoupling descriptor;

[0010] Encode the static descriptor and the dynamic descriptor respectively through an encoder to obtain a static latent variable and a dynamic latent variable, and add the static latent variable and the dynamic latent variable to obtain an encoded feature vector;

[0011] Decode the feature vector through a decoder to obtain the deformation state of the local cloth of the 3D human model;

[0012] Train the network model through a loss function based on physical simulation, so that the network learns to satisfy the physical constraints of the cloth and the human body and outputs a prediction result of the cloth state.

[0013] Further, the obtaining of the motion action parameters of the skinned 3D human model includes:

[0014] Bind the bones of the skinned 3D human model to obtain a bound human model;

[0015] Establish a particle system corresponding to the cloth state, where the particle positions and velocities of each particle in the particle system can change based on time;

[0016] Obtain the motion action parameters of the human model according to the human model and the particle system, and the motion action parameters at least include the position information of each joint.

[0017] Further, the converting the motion action parameters into a static descriptor and a dynamic descriptor through a decoupled descriptor includes:

[0018] Map the angles between each joint through an orthogonal matrix in three-dimensional space to obtain a local static descriptor, denoted as: g GS ([r i,1 ,r i,2 ,r i,3 ) = [r i,1 ,r i,2 , where r i,1 ,r i,2 ,r i,3 is a column vector, and the mapping g GS is the process of converting a three-dimensional rotation matrix into a low-dimensional matrix representation space;

[0019] Establish a unit vector pointing to the non-offset direction of gravity to obtain a global static descriptor, denoted as: where, where is the unit vector of the jth joint, R j is the rotation matrix corresponding to the global joint direction, g is the gravity vector, and |g| is the modulus of the gravity vector;

[0020] Connect the local static descriptor and the global static descriptor to obtain the static descriptor for each joint, expressed as: where α static is the static descriptor of the i-th joint;

[0021] Determine the dynamic descriptor for each joint based on the static descriptor of each joint and the position information of each joint, expressed as: where is the dynamic descriptor of the i-th joint; are the first-order derivatives of the static descriptor respectively, is the acceleration of the joint in the local space.

[0022] Furthermore, the encoder includes a static encoder and a dynamic encoder;

[0023] The static encoder includes four connected fully-connected layers, which have 128, 256, 512, and 1024 neurons respectively, and use the SeLU activation function. The output of the previous fully-connected layer is used as the input of the next fully-connected layer to obtain the static latent variable in the static descriptor

[0024] The dynamic encoder includes a set of fully-connected layers and a long short-term memory network. It is applied to the dynamic descriptor of each joint through two fully-connected layers to obtain the high-level feature array for each joint; the high-level feature data is flattened and fed into another two fully-connected layers, and finally the output is processed by the long short-term memory network to obtain the dynamic latent variable

[0025] Furthermore, the decoder includes three fully-connected layers and a pose space deformation layer; the working process of the decoder is as follows:

[0026] Input the feature vector into the first fully-connected layer to obtain the first intermediate feature, expressed as: h1 = SeLU(W1e + b1), where W1 is the weight matrix, b1 is the bias vector, and SeLU is the activation function;

[0027] Input the first intermediate feature into the second fully-connected layer to obtain the second intermediate feature, expressed as: h2 = SeLU(W2e + b2);

[0028] Input the second intermediate feature into the third fully-connected layer to obtain the final feature, expressed as: h3 = SeLU(W3e + b3);

[0029] Generate the cloth vertex positions through the pose space deformation layer for the final feature to obtain the cloth deformation state, expressed as: vertices = PSD(h3).

[0030] Furthermore, the loss function based on physical simulation is expressed by the following formula:

[0031]

[0032] Wherein, is the final loss function; is the bending loss function; w bending is the bending loss weight; is the collision loss function; w collision is the collision loss weight; is the inertia loss function; w interia is the inertia loss weight; is the gravity loss function; w gravity is the gravity loss weight; is the friction loss function; W friction is the friction weight loss.

[0033] The bending loss function is expressed as: Wherein, N is the number of fixed points of the cloth model, p i is the radius of curvature at the i-th vertex;

[0034] The collision loss function is expressed as: Wherein, N is the number of vertices in the cloth model, x i is the position of the i-th vertex, is the closest position on the surface of the obstacle that collides with the i-th vertex; is the distance between the i-th vertex and its corresponding closest point, and ε is a threshold representing the collision distance. When the distance is less than the threshold, the loss is zero; when the distance is greater than the threshold, the loss increases as the distance increases;

[0035] The inertia loss function is expressed as: Wherein, N is the number of fixed points of the cloth model, m i is the mass of the i-th vertex, v i is the velocity vector of the i-th vertex, v i-1 is the velocity vector of the i-th vertex in the previous time step;

[0036] The gravity loss function is expressed as; Wherein, m is the particle mass and g is the gravity vector;

[0037] The friction loss function is expressed as: Where B is the number of samples, N is the number of particles in each sample, m j m j is the mass of the j-th point in the cloth, g is the gravitational acceleration constant, d jis the distance between the j-th point and the human body, and μ is the friction coefficient.

[0038] On the other hand, this article also provides a cloth simulation device based on physics-aware deep learning. The device includes:

[0039] An acquisition module for acquiring the motion action parameters of a skinned 3D human model;

[0040] A decoupling module for converting the motion action parameters into a static descriptor and a dynamic descriptor through a decoupling descriptor;

[0041] An encoding module for encoding the static descriptor and the dynamic descriptor respectively through an encoder to obtain a static latent variable and a dynamic latent variable, and adding the static latent variable and the dynamic latent variable to obtain an encoded feature vector;

[0042] A decoding module for decoding the feature vector through a decoder to obtain the deformation state of the local cloth of the 3D human model;

[0043] A training module for training the network model through a loss function based on physical simulation, so that the network learns to satisfy the physical constraints of the cloth and the human body and outputs a prediction result of the cloth state.

[0044] On the other hand, this article also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described above is implemented.

[0045] Finally, this article provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.

[0046] Adopting the above technical solution, this paper provides a cloth simulation method based on physically-aware deep learning, and the method includes: obtaining the motion action parameters of a skinned 3D human model; converting the motion action parameters into a static descriptor and a dynamic descriptor through a decoupled descriptor; respectively encoding the static descriptor and the dynamic descriptor through an encoder to obtain a static latent variable and a dynamic latent variable, and adding the static latent variable and the dynamic latent variable to obtain an encoded feature vector; decoding the feature vector through a decoder to obtain the deformation state of the local cloth of the 3D human model; training the network model through a physically-based simulation loss function, so that the network learns to satisfy the physical constraints of the cloth and the human body, and outputs a prediction result of the cloth state. The purpose of this paper is to use a physically-based simulation loss function to enable the network to learn to satisfy the physical constraints of the cloth and the human body, be able to learn the dynamic behavior of the cloth without any ground truth data, achieve accurate prediction of cloth dynamics, and provide a more realistic and natural performance for clothing simulation and human-computer interaction in virtual scenes.

[0047] To make the above and other purposes, features and advantages of this paper more obvious and understandable, the following specifically gives preferred embodiments and cooperates with the attached drawings to make a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of this paper or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of this paper. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 Shows a schematic diagram of the steps of a cloth simulation method based on physically-aware deep learning provided by an embodiment of this paper;

[0050] Figure 2 Shows a schematic diagram of the overall framework of the method provided in an embodiment of this paper;

[0051] Figure 3 Shows a cyclic encoder-decoder architecture diagram in an embodiment of this paper;

[0052] Figure 4 Shows a jumping pose dynamic clothing effect diagram in an embodiment of this paper, where a is a schematic diagram of a human model - dynamic clothing simulation, b is a dynamic clothing effect diagram, and c is a dynamic clothing detail diagram;

[0053] Figure 5Shows the detailed drawing of the running posture dynamic clothing in the embodiments of this article, where a is the schematic diagram of the human model - dynamic clothing simulation, b is the effect drawing of the dynamic clothing, and c is the detailed drawing of the dynamic clothing;

[0054] Figure 6 Shows the structural schematic diagram of the computer device in the embodiments of this article. Detailed implementation manners

[0055] Next, the technical solutions in the embodiments of this article will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this article. Obviously, the described embodiments are only a part of the embodiments of this article, rather than all the embodiments. Based on the embodiments in this article, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this article.

[0056] It should be noted that the terms "first", "second", etc. in the specification and claims of this article and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this article described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0057] In the prior art, in the field of cloth simulation, problems such as low clothing realism, poor dynamic continuity, and weak detail fidelity often occur during the simulation process.

[0058] To solve the above problems, the embodiments of this article provide a cloth simulation method based on physically aware deep learning to improve the accuracy of cloth simulation. Figure 1 Is the step schematic diagram of a cloth simulation method based on physically aware deep learning provided by the embodiments of this article. This specification provides the method operation steps as described in the embodiments or flowcharts, but based on routine or non-creative labor, more or fewer operation steps may be included. The step order listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or device product executes, it can be executed in the order of the embodiments or as shown in the accompanying drawings or executed in parallel. Specifically, as Figure 1 shown, the method may include:

[0059] S101: Obtain the motion action parameters of the skinned 3D human model;

[0060] S102: Convert the motion action parameters into static descriptors and dynamic descriptors through a decoupling descriptor;

[0061] S103: Encode the static descriptor and the dynamic descriptor respectively through an encoder to obtain a static latent variable and a dynamic latent variable, and add the static latent variable and the dynamic latent variable to obtain an encoded feature vector;

[0062] S104: Decode the feature vector through a decoder to obtain the deformation state of the local cloth of the 3D human model;

[0063] S105: Train the network model through a loss function based on physical simulation, so that the network learns to satisfy the physical constraints of the cloth and the human body, and outputs a prediction result of the cloth state.

[0064] It can be understood that the purpose of this application is to propose a cloth neural simulator that integrates physical laws. By parameterizing the skinned 3D human model and using decoupled static and dynamic descriptors, the static and dynamic features of the pose are encoded into the encoding of the cloth subspace. Then, through the neural cloth subspace solver, the encoding of the cloth subspace is predicted, and it is converted into the cloth state through a decoder to complete the prediction and generation of the cloth state. This technology aims to use a loss function based on physical simulation to enable the network to learn to satisfy the physical constraints of the cloth and the human body, achieve accurate prediction of cloth dynamics, and provide a more realistic and natural performance for clothing simulation and human-computer interaction in virtual scenes.

[0065] As Figure 2 shown, it is a framework schematic diagram for implementing this method in an embodiment of this application.

[0066] Further, the obtaining of the motion action parameters of the skinned 3D human model includes:

[0067] Perform bone binding on the skinned 3D human model to obtain a bound human body model;

[0068] Establish a particle system corresponding to the cloth state, where the particle position and velocity of each particle in the particle system can change based on time;

[0069] According to the human body model and the particle system, obtain the motion action parameters of the human body model, and the motion action parameters at least include the position information of each joint.

[0070] It can be understood that inspired by classical computer graphics physical simulation, the cloth is represented as a particle system The cloth solver calculates the cloth configuration at time t from the previous cloth state, which is defined by the particle positions and velocities at t-1. In the present invention, the solver also takes into account the gravity acting on the cloth and the collision of the cloth with the skin 3D model. The solver can be written as x t = f(α t , x t-1 , v t-1 ), where v is the particle velocity and the velocity will depend on the current and previous example positions. Therefore, the expression can be rewritten as Therefore, assuming there is a 3D human body model described by the parameter θ, the state of the cloth can be parameterized by the pose history Θ of the body, Θ t = {α t , α t-1 , α t-2 , …, α0}.

[0071] In the embodiments of the present specification, the conversion of the motion action parameters into static descriptors and dynamic descriptors through decoupled descriptors includes:

[0072] Mapping the angles between each joint through an orthogonal matrix in three-dimensional space to obtain a local static descriptor, expressed as: g GS ([r i,1 , r i,2 , r i,3 ) = [r i,1 , r i,2 , where r i,1 , r i,2 , r i,3 are column vectors, and the mapping g GS is the process of converting a three-dimensional rotation matrix into a representation space of a low-dimensional matrix;

[0073] Establish a unit vector pointing in the non-offset direction of gravity to obtain a global static descriptor, expressed as: where, where is the unit vector of the jth joint, R j is the rotation matrix corresponding to the global joint direction, g is the gravity vector, and |g| is the modulus of the gravity vector;

[0074] Connect the local static descriptor and the global static descriptor to obtain the static descriptor of each joint, expressed as: where, α static is the static descriptor of the ith joint;

[0075] Determine the dynamic descriptor of each joint according to the static descriptor of each joint and the position information of each joint, expressed as: where, is the dynamic descriptor of the i-th joint; are the first-order derivatives of the static descriptor, which is the acceleration of the joint in the local space.

[0076] It can be understood that in the embodiments of this specification, the automated skeleton binding technology is adopted to bind the skeleton of the skinned 3D model, and its pose parameters are parameterized as α. The joint state of this model is represented by a rotation matrix. For each joint i, its rotation matrix is denoted as R i . In order to reduce the required training data, time, and model capacity, and enhance the generalization of the model, the embodiments of this specification adopt decoupled descriptors to extract static information and dynamic information through static descriptors and dynamic descriptors.

[0077] To describe the body pose, the direction of the joint relative to the parent joint, axis angle, or quaternion is usually used, and there are discontinuous problems in the rotation space. Therefore, the present invention selects 6D descriptors to describe the angles between each joint of the human body, and through the 3×3 orthogonal matrix R in the three-dimensional space i mapping:

[0078] g GS ([r i,1 ,r i,2 ,r i,3 ) = [r i,1 ,r i,2

[0079] where r i,1 ,r i,2 ,r i,3 are column vectors, and the mapping g GS represents the process of converting the three-dimensional rotation matrix into a low-dimensional matrix representation space. Since the rotation matrix is an orthogonal matrix, it means that each column vector is a unit vector, and each column vector is orthogonal to each other. Then the third column vector r i,3 can be determined by the first column vector r i,1 and the second column vector r i,2 . To reduce the calculation amount, the third column vector is discarded to maintain the required continuity and orthogonality.

[0080] This descriptor is based on the relative direction and has better local properties, which can better capture the influence of small changes in body pose on clothing. In addition, this application adds a unit vector pointing to the non-offset direction of gravity to the descriptor of each joint to maintain the stability of the descriptor. Where is the unit vector of the j-th joint, and R j ​is the rotation matrix corresponding to the global joint direction, g is the gravity vector, and |g| is the magnitude of the gravity vector. This description contains the global direction information of each joint and is invariant to rotations around the gravity axis. In addition, it will be associated with the local cloth deformation direction due to gravity. In the static pose, the local and global descriptors are connected to obtain a 9-dimensional feature array for each joint. Therefore, the static descriptor can be expressed as:

[0081]

[0082] The dynamic descriptor is used to describe body movements. The time derivatives of the joint directions and positions are calculated. The derivative of the direction is calculated from the static descriptor. These derivatives face a large input space. The present invention solves this problem by removing the derivatives without normalization, greatly reducing the input space. At the same time, a descriptor that is more strongly related to the local cloth dynamic deformation generated by the movement is defined. Assuming no air resistance, the dynamic cloth deformation only appears when the body accelerates. Therefore, the present invention uses the first-order derivative of the joint direction and the second-order derivative of the joint position as the movement descriptors. The first-order derivative reflects the change speed of the joint direction, while the second-order derivative reflects the change rate of the change speed of the joint position. These two descriptors are concatenated into a 12-dimensional descriptor for each joint, giving where K joints, 9 dimensions are the first-order derivatives from the static descriptor, and 3 additional dimensions are for the acceleration of the joints in the local space. In the present invention, joints such as hands, feet, and faces in the 3D model that are not related to cloth dynamics are removed. Therefore, the dynamic descriptor can be expressed as:

[0083]

[0084] where are the first-order derivatives of the static descriptor respectively, is the acceleration of the joint in the local space. After the static features and dynamic features are segmented, during the training process, the gradients of the static part are frozen, and then the dynamic part is randomly rearranged, and the enhanced input data of the dynamic and static parts are returned. This motion enhancement technology can effectively increase the diversity of training data and improve the generalization ability of the model.

[0085] Furthermore, the encoder includes a static encoder and a dynamic encoder;

[0086] The static encoder includes four connected fully-connected layers, and these four fully-connected layers have 128, 256, 512, and 1024 neurons respectively, and use the SeLU activation function. The output of the previous fully-connected layer is used as the input of the next fully-connected layer to obtain the static latent variable in the static descriptor

[0087] The dynamic encoder includes a set of fully connected layers and a long short-term memory network. Two fully connected layers are applied to the dynamic descriptors of each joint respectively to obtain an array of high-level features for each joint; the high-level feature data is flattened and fed into two other fully connected layers, and finally the output is processed by the long short-term memory network to obtain the dynamic latent variable

[0088] It can be understood that the present invention adopts a recurrent encoder-decoder network architecture. The encoder consists of two different modules, a static encoder and a dynamic encoder. Each module receives the corresponding descriptors, such as Figure 3 shown in the schematic diagram of the recurrent encoder-decoder network architecture

[0089] Static encoder, the present invention implements the static encoder as a set of 4 fully connected layers. The encoder only provides the current pose α static , which is first flattened into a 9K-dimensional array. Then it is processed sequentially through four fully connected layers. These fully connected layers have 128, 256, 512, and 1024 neurons respectively, and use the SeLU activation function. The output of each layer is used as the input of the next layer. After a series of linear transformations and non-linear activations, the feature information in the input data is gradually extracted and transformed, and finally a higher-level representation is obtained. The output of the encoder is a static latent code

[0090] Dynamic encoder, this module consists of two parts: a set of fully connected layers and a long short-term memory network (LSTM). First, two fully connected layers are applied to each joint descriptor, as if the joint is a sample, to obtain an array of high-level features for each joint. Then, the array is flattened and fed into two other fully connected layers. Finally, the output is passed through the LSTM, which combines it with the hidden state that encodes the dynamic history, thereby obtaining the dynamic latent variable

[0091]

[0092] where α dynamic is the dynamic descriptor

[0093] It should be noted that is calculated according to the entire action Θ t . All layers of the dynamic encoder have no bias, and zero input will be converted into zero output. This ensures that the static sample will have zero because the time derivative, i.e., the dynamic descriptor, will be zero. Therefore, the addition with will have no effect. In addition, samples with high movement amplitude usually produce higher a value, resulting in high dynamic deformation due to high perturbations. Additionally, as long as the model inputs a constant pose, the hidden state of the LSTM decays to zero.

[0094] After encoding is completed, the static features and dynamic features are added together to synthesize a new feature vector e for subsequent decoding.

[0095] e = e static + e dynamic

[0096] Furthermore, the decoder includes three fully connected layers and a pose space deformation layer; the working process of the decoder is as follows:

[0097] The feature vector is input into the first fully connected layer to obtain a first intermediate feature, expressed as: h1 = SeLU(W1e + b1), where W1 is the weight matrix, b1 is the bias vector, and SeLU is the activation function;

[0098] The first intermediate feature is input into the second fully connected layer to obtain a second intermediate feature, expressed as: h2 = SeLU(W2e + b2);

[0099] The second intermediate feature is input into the third fully connected layer to obtain the final feature, expressed as: h3 = SeLU(W3e + b3);

[0100] The final feature is passed through the pose space deformation layer to generate the cloth vertex positions to obtain the cloth deformation state, expressed as: vertices = PSD(h3).

[0101] It can be understood that the decoder consists of three fully connected layers and activation functions. The encoded features are decoded layer by layer through three fully connected layers with 512 neurons and the SeLU activation function to generate the deformation of the cloth. When the encoding is completed, it is input into the first fully connected layer. Then, the features are processed layer by layer through the second and third fully connected layers, and finally, a pose space deformation layer (PSD) generates the final cloth vertex positions. Each fully connected layer uses the ReLU activation function to ensure non-linear transformation and capture complex cloth deformation characteristics. The input feature vector e passes through the first fully connected layer

[0102] h1 = SeLU(W1e + b1)

[0103] where W1 is the weight matrix, b1 is the bias vector, and SeLU is the activation function

[0104] When the intermediate feature h1 passes through the second fully connected layer:

[0105] h2 = SeLU(W2e + b2)

[0106] When the intermediate feature h2 passes through the third fully connected layer:

[0107] h3 = SeLU(W3e + b3)

[0108] The final feature h3 passes through the PSD layer to generate the cloth vertex positions:

[0109] vertices = PSD(h3)

[0110] Through the step-by-step processing of these layers, the encoded features are decoded into the vertex positions of the cloth, thus realizing the deformation simulation of the cloth.

[0111] The loss function based on physical simulation is represented by the following formula:

[0112]

[0113] Where, is the final loss function; is the bending loss function; w bending is the bending loss weight; is the collision loss function; w collision is the collision loss weight; is the inertia loss function; w interia is the inertia loss weight; is the gravity loss function; w gravity is the gravity loss weight; is the friction loss function; w friction is the friction weight loss.

[0114] It can be understood that the process of model training is to use the energy function of the physical system (including the cloth and the body) as the loss function, and the model learns to predict the clothing state that satisfies the energy constraint during the training process.

[0115] Bending loss. The present invention adopts a bending loss based on surface curvature, and this loss function can be used to constrain the surface of the cloth model:

[0116]

[0117] Where N is the number of fixed points of the cloth model, p iis the radius of curvature at the i-th vertex. The radius of curvature balances the curvature at each point on the fabric surface. For a given point, the smaller the radius of curvature, the greater the surface curvature and the more curved the surface. Therefore, this loss function penalizes a large radius of curvature, making the surface smoother. Collision loss. In computer graphics simulation, the interaction between the fabric and external objects is achieved by detecting and solving collisions. Similarly, in the present invention, the square of the difference between the distance between each fabric vertex and the closest point on the surface of the corresponding obstacle and a threshold is calculated, and the sum of these squared differences is used as a measure of the loss function:

[0118]

[0119] where N is the number of vertices in the fabric model, x i is the position of the i-th vertex, is the closest position on the surface of the obstacle that collides with the i-th vertex; is the distance between the i-th vertex and its corresponding closest point, and ε is a threshold representing the collision distance. When the distance is less than the threshold, the loss is zero; when the distance is greater than the threshold, the loss increases as the distance increases. This can ensure that the fabric model is appropriately penalized when it collides with an obstacle, thus avoiding excessive penetration or unnatural deformation.

[0120] Inertia loss. According to Newton's second law, in the fabric model, each vertex can be regarded as a particle mass point under the action of external forces. We can apply Newton's second law F = ma to each vertex and approximate the acceleration of the object as the velocity change between adjacent times according to the time step Δt, and rewrite it as:

[0121]

[0122] where Δv = v i - v i-1 is the velocity change between adjacent time steps, F is the force acting on the body model, and m is the particle mass.

[0123] To represent the velocity change in the loss function, we consider the square of the velocity difference between adjacent time steps and weight it by the mass of each particle:

[0124]

[0125] where N is the number of fixed points in the fabric model, m i is the mass of the i-th vertex, v i is the velocity vector of the i-th vertex, v i-1is the velocity vector of the i-th vertex in the previous time step. This loss function calculates half of the square of the velocity change of each vertex between adjacent time steps multiplied by its mass, and then sums up the losses of all vertices weighted. Penalizing sharp changes in velocity encourages smoother and more natural motion of objects in the simulation.

[0126] Gravity loss. The influence of gravity is realized through potential energy as a loss:

[0127]

[0128] This formula will push the vertex in the direction of gravity, weighted by the particle mass and gravity.

[0129] Friction loss. The energy loss generated by the relative sliding between the fabric surface and the obstacle under the friction between the fabric and the obstacle. It is obtained by calculating the frictional force between the points on the fabric surface and the contact points of the obstacle, and then multiplying it by factors such as distance and mass, which characterizes the dissipative effect of friction on the kinetic energy of the system:

[0130]

[0131] where B is the number of samples, N is the number of particles in each sample, m j m j is the mass of the j-th point in the fabric, g is the gravitational acceleration constant, d j is the distance between the j-th point and the human body, and μ is the friction coefficient.

[0132] Finally, the total loss function is achieved by summing up each weighted loss function

[0133]

[0134] According to the set weight assignment: w bending is the bending loss weight, w bending = 2×10 -5 ; w collision is the collision loss weight, w collision = 10.0; w int eria is the inertia loss weight, w interia = 0.1; w gravity is the gravity loss weight, w gravity = -10.0; w friction is the frictional force weight loss, w friction = 0.5.

[0135] The purpose of the embodiments of this specification is to propose a cloth neural simulator that integrates physical laws. By parameterizing the 3D human model with skin, using decoupled static and dynamic descriptors, the static and dynamic features of the pose are encoded into the encoding of the cloth subspace. Then, through the neural cloth subspace solver, the encoding of the cloth subspace is predicted, and it is converted into the cloth state through the decoder, completing the prediction and generation of the cloth state. This technology aims to enable the network to learn to meet the physical constraints of the cloth and the human body through unsupervised training using a physical simulation-based loss function, achieve accurate prediction of cloth dynamics, and provide a more realistic and natural performance for clothing simulation and human-computer interaction in virtual scenarios.

[0136] Through the cloth neural simulator that integrates physical laws, the embodiments of this specification can learn the dynamic behavior of cloth without ground truth data, adopt a decoupled architecture, have strong model generalization ability, and improve the robustness of the model through motion enhancement technology, thus achieving a realistic cloth simulation effect.

[0137] As Figure 4 and Figure 5 shown, in the embodiments of this specification, the dynamic clothing effect diagram of the jumping pose and the dynamic clothing detail diagram of the running pose obtained by simulating through the above method are presented.

[0138] Based on the above-provided method, the embodiments of this specification also provide a cloth simulation device based on physics-aware deep learning. The device includes:

[0139] An acquisition module, configured to acquire the motion action parameters of the 3D human model with skin;

[0140] A decoupling module, configured to convert the motion action parameters into a static descriptor and a dynamic descriptor through a decoupled descriptor;

[0141] An encoding module, configured to respectively perform encoding processing on the static descriptor and the dynamic descriptor through an encoder to obtain a static latent variable and a dynamic latent variable, and add the static latent variable and the dynamic latent variable to obtain an encoded feature vector;

[0142] A decoding module, configured to perform decoding processing on the feature vector through a decoder to obtain the deformation state of the local cloth of the 3D human model;

[0143] A training module, configured to train the network model through a physical simulation-based loss function, so that the network learns to meet the physical constraints of the cloth and the human body, and outputs the prediction result of the cloth state.

[0144] The beneficial effects obtained by the above device are the same as those obtained by the above method, and the embodiments of this specification will not elaborate.

[0145] This embodiment provides a computer device, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection.

[0146] Those skilled in the art can understand that Figure 6 the structure shown in

[0147] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0148] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0149] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0150] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.

[0151] It should also be understood that in the embodiments herein, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0152] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this article.

[0153] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0154] In the several embodiments provided in this article, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connection.

[0155] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments in this article.

[0156] Specific embodiments are used in this article to elaborate on the principles and implementation manners of this article. The description of the above embodiments is only used to help understand the method and its core idea of this article; at the same time, for those of ordinary skill in the art, according to the idea of this article, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this article.

Claims

1. A cloth simulation method based on physical perception deep learning, characterized in that: The method comprises: Get the motion parameters of the skinned 3D character model; The motion action parameters are converted into static descriptors and dynamic descriptors through decoupling descriptors, specifically: The angles between the joints are mapped through an orthogonal matrix in three-dimensional space to obtain a local static descriptor, expressed as: GS ([r i,1 ,r i,2 ,r i,3 ])=[r i,1 ,r i,2 ], where r i,1 ,r i,2 ,r i,3 is a column vector, mapping g GS The process of converting a three-dimensional rotation matrix into a low-dimensional matrix representation space; Establishing a unit vector pointing to the direction where gravity is not offset gives a global static descriptor, expressed as: Among them, is the unit vector of the jth joint, R j is the rotation matrix corresponding to the global joint orientation, g is the gravity vector, |g| is the modulus of the gravity vector; Connect the local static descriptor and the global static descriptor to get the static descriptor of each joint, which is expressed as: Among them, α static is the static descriptor of the i-th joint; According to the static descriptor of each joint and the position information of each joint, the dynamic descriptor of each joint is determined, which is expressed as: in, is the dynamic descriptor of the i-th joint; are the first-order derivatives of the static descriptor, is the acceleration of the joint in local space; The static descriptor and the dynamic descriptor are respectively encoded by an encoder to obtain a static latent variable and a dynamic latent variable, and the static latent variable and the dynamic latent variable are added to obtain an encoded feature vector; Decoding the feature vector by a decoder to obtain a deformation state of a local cloth of the 3D character model; The network model is trained through a loss function based on physical simulation, so that the network learns to meet the physical constraints of cloth and human body and outputs the prediction results of cloth state.

2. The method according to claim 1, characterized in that: The step of obtaining the motion parameters of the skinned 3D character model includes: Perform skeleton binding on the skinned 3D character model to obtain a bound human body model; Establish a particle system corresponding to the cloth state, in which the particle position and velocity of each particle in the particle system can change based on time; According to the human body model and the particle system, motion parameters of the human body model are acquired, and the motion parameters at least include position information of each joint.

3. The method according to claim 1, characterized in that The encoder includes a static encoder and a dynamic encoder; The static encoder includes four connected fully connected layers, each of which has 128, 256, 512 and 1024 neurons, respectively, and uses a SeLU activation function. The output of the previous fully connected layer is used as the input of the next fully connected layer to obtain the static latent variables in the static descriptor. The dynamic encoder includes a set of fully connected layers and a long short-term memory network. The two fully connected layers are applied to the dynamic descriptors of each joint to obtain a high-level feature array of each joint. The high-level feature data is flattened and fed to another two fully connected layers. The final output is processed by the long short-term memory network to obtain the dynamic latent variable 4. The method according to claim 1, characterized in that The decoder includes three fully connected layers and one posture space deformation layer; the working process of the decoder is as follows: Input the feature vector into the first fully connected layer to obtain the first intermediate feature, which is expressed as: h1=SeLU(W1e+b1), where W1 is the weight matrix, b1 is the bias vector, and SeLU is the activation function; The first intermediate feature is input into the second fully connected layer to obtain the second intermediate feature, which is expressed as: h2 = SeLU (W2e + b2); The second intermediate feature is input into the third fully connected layer to obtain the final feature, which is expressed as: h3 = SeLU (W3e + b3); The final feature is passed through the posture space deformation layer to generate the cloth vertex position to obtain the cloth deformation state, which is expressed as: vertices = PSD (h3).

5. The method according to claim 1, characterized in that The loss function based on physical simulation is expressed as follows: in, is the final loss function; is the bending loss function; w bending is the bending loss weight; is the collision loss function; w collision is the collision loss weight; is the inertia loss function; w int eria is the inertia loss weight; is the gravity loss function; w gravity is the gravity loss weight; is the friction loss function; w friction is the friction weight loss.

6. The method according to claim 5, characterized in that The bending loss function is expressed as: Where N is the number of fixed points in the cloth model, p i is the radius of curvature at the i-th vertex; The collision loss function is expressed as: Where N is the number of vertices in the cloth model, x i is the position of the i-th vertex, is the closest position on the obstacle surface that collides with the i-th vertex; is the distance between the i-th vertex and its corresponding nearest point, and ε is a threshold representing the collision distance; when the distance is less than the threshold, the loss is zero; when the distance is greater than the threshold, the loss increases with the increase of the distance; The inertia loss function is expressed as: Where N is the number of fixed points in the cloth model, m i is the mass of the ith vertex, v i is the velocity vector of the ith vertex, v i-1 is the velocity vector of the i-th vertex in the previous time step; The gravity loss function is expressed as; Among them, m is the mass of the particle, g is the gravity vector; The friction loss function is expressed as: Where B is the number of samples, N is the number of particles in each sample, and m j m j is the mass of the jth point in the cloth, g is the gravitational acceleration constant, d j is the distance between the jth point and the human body, and μ is the friction coefficient.

7. A cloth simulation device based on physical perception deep learning, characterized in that: The device comprises: An acquisition module is used to obtain motion parameters of a skinned 3D character model; The decoupling module is used to convert the motion action parameters into static descriptors and dynamic descriptors through decoupling descriptors, specifically: The angles between the joints are mapped through an orthogonal matrix in three-dimensional space to obtain a local static descriptor, expressed as: GS ([r i,1 ,r i,2 ,r i,3 ])=[r i,1 ,r i,2 ], where r i,1 ,r i,2 ,r i,3 is a column vector, mapping g GS The process of converting a three-dimensional rotation matrix into a low-dimensional matrix representation space; Establishing a unit vector pointing to the direction where gravity is not offset gives a global static descriptor, expressed as: Among them, is the unit vector of the jth joint, R j is the rotation matrix corresponding to the global joint orientation, g is the gravity vector, |g| is the modulus of the gravity vector; Connect the local static descriptor and the global static descriptor to get the static descriptor of each joint, which is expressed as: Among them, α static is the static descriptor of the i-th joint; According to the static descriptor of each joint and the position information of each joint, the dynamic descriptor of each joint is determined, which is expressed as: in, is the dynamic descriptor of the i-th joint; are the first-order derivatives of the static descriptor, is the acceleration of the joint in local space; An encoding module, configured to encode the static descriptor and the dynamic descriptor respectively through an encoder to obtain a static latent variable and a dynamic latent variable, and add the static latent variable and the dynamic latent variable to obtain an encoded feature vector; A decoding module, used for decoding the feature vector through a decoder to obtain a deformation state of a local cloth of the 3D character model; The training module is used to train the network model through a loss function based on physical simulation, so that the network learning satisfies the physical constraints of the cloth and the human body, and outputs the prediction results of the cloth state.

8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Human body and clothing model collision detection and processing method based on HRBFs

    CN112862956A

  • Cloth static deformation prediction method based on triangular mesh

    CN116401723A