High-generalization clothes-driven animation generation method and device, equipment and medium

By using the combination method of graph network, timing network and regression network in animation generation, the problem of robustness and low efficiency of clothing driving methods in the prior art is solved, high generalization and real-timeness are achieved, processing efficiency is improved and collision problems are reduced.

CN120107424APending Publication Date: 2025-06-06ZHEJIANG TONGHUASHUN INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332231.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, the clothing driving method based on physical simulation is sensitive to initial input, robust and efficient; based on deep learning, the parameterized model method has poor practicality, the graph network method is low in efficiency and depends on the output of the previous frame.

Method used

A highly generalized clothing-driven animation generation method is proposed. By obtaining the initial clothing sequence and human body-driven sequence in the posture space, it is input to the target network model, which includes graph network, timing network and regression network. The graph network is used to extract local features, and the regression network replaces the global feature generation. The timing network learns timing features without relying on the output of the previous frame.

Benefits of technology

It realizes highly generalized clothing drive, which is both practical and real-time, and achieves a balance of performance and efficiency. It can handle the characteristics of a variety of topological structures, improves processing efficiency and reduces collision problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107424A_ABST
    Figure CN120107424A_ABST
Patent Text Reader

Abstract

The invention discloses a high-generalization clothes-driven animation generation method and device, equipment and a storage medium, and relates to the technical field of image processing. The method comprises the following steps: acquiring an initialized clothes sequence and a human body driving sequence in a posture space; inputting the initialized clothes sequence and the human body driving sequence into a target network model; the target network model sequentially comprises a graph network used for extracting graph structure features, a time sequence network used for extracting time sequence features and a regression network in sequence; obtaining the offset of each clothes vertex according to the output of the target network model, and obtaining a learned clothes sequence according to the initialized clothes sequence and the offset; performing geometric post-processing on the learned clothes sequence to obtain a processed clothes sequence, so as to generate an animation according to a processed clothes driving result; the geometric post-treatment comprises wrinkle treatment, and the wrinkle treatment is used for simulating the deformation of the clothes under the constraint of the cloth. The method has practicability and real-time performance, and the performance and the efficiency are balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a method, device, equipment and storage medium for generating animation driven by highly generalized clothing. Background Art

[0002] Clothes driving is an important step in animation generation. Clothes effects are synthesized by inputting clothing information and human motion data. In the prior art, clothes driving based on physical simulation optimizes and solves the clothes under the human body in each frame by constructing physical simulation constraints. Such methods are very sensitive to initialization inputs, are often not robust enough, and are extremely inefficient, and do not meet the requirements of practicality and real-time performance. In the prior art, clothes driving based on deep learning is also used. Such solutions are mainly divided into two types of methods: clothes driving based on parametric models, and clothes driving based on graph networks. However, the methods based on parametric models are less practical and effective, and one model is required for one template garment. The methods based on graph networks are less efficient, and the input of the network of the current frame depends on the output of the network of the previous frame, which limits the practicality of the network to a certain extent. Summary of the invention

[0003] In view of this, the purpose of the present invention is to provide a highly generalized clothing-driven animation generation method, device, equipment and storage medium, which can be both practical and real-time, and achieve a balance between performance and efficiency. The specific scheme is as follows:

[0004] In a first aspect, the present application discloses a highly generalizable clothing-driven animation generation method, comprising:

[0005] Get the initialized clothing sequence and human body drive sequence in the posture space;

[0006] Inputting the initialized clothing sequence and the human body drive sequence into a target network model; the target network model sequentially includes a graph network for extracting graph structure features, a temporal network for extracting temporal features, and a regression network;

[0007] Obtaining a first offset of a clothing vertex according to an output of the target network model, and obtaining a learned clothing sequence according to the initialized clothing sequence and the first offset;

[0008] The learned clothing sequence is subjected to geometric post-processing to obtain a processed clothing sequence, so as to generate an animation according to the processed clothing sequence result; the geometric post-processing includes wrinkle processing, and the wrinkle processing is used to simulate the deformation of clothing under cloth constraints.

[0009] Optionally, inputting the initialized clothing sequence and the human body drive sequence into a target network model includes:

[0010] Based on the initialized clothing sequence and the human body drive sequence, a graph network is used to extract node features and edge features; the node features include vertex velocity vectors, normal vectors, gravity, material parameters, node types, and node levels; the edge features include the length of the edge of the current frame graph structure, the length of the edge of the current frame graph structure after normalization, the length of the edge of the template graph, and the length of the edge of the template graph after normalization;

[0011] According to the features of the edge between the node and the connected clothing vertex, and the features of the edge between the clothing vertex and the human body vertex, the node features are updated to obtain updated node features;

[0012] According to the updated node feature of the first endpoint of the edge and the updated node feature of the second endpoint of the edge, the edge feature is updated to obtain an updated edge feature;

[0013] The graph structure feature is obtained based on the updated node feature and the updated edge feature.

[0014] Optionally, inputting the initialized clothing sequence and the human body drive sequence into a target network model includes:

[0015] According to the output of the graph network, obtaining graph network output data containing graph structure features;

[0016] Performing dimension conversion on the graph network output data to obtain input data representing a time series dimension;

[0017] The temporal network is used to extract temporal features based on the input data.

[0018] Optionally, inputting the initialized clothing sequence and the human body drive sequence into a target network model includes:

[0019] According to the output of the temporal network, temporal network output data including graph structure features and temporal features is obtained;

[0020] Based on the output data of the temporal network, the regression network is used to predict the first offset of each clothing vertex.

[0021] Optionally, inputting the initialized clothing sequence and the human body drive sequence into a target network model includes:

[0022] The network learning of the regression network is performed in combination with a preset loss function; the preset loss function is constructed based on one or more of the clothing-human body collision loss, clothing self-collision loss, gravity loss, clothing material-related loss, clothing stretching loss, clothing-human body friction loss, external force loss, and time series smoothing constraints.

[0023] Optionally, performing geometric post-processing on the learned clothing sequence to obtain a processed clothing sequence includes:

[0024] The second offset of each vertex in the learned clothing sequence is calculated according to the wrinkle formula; the wrinkle formula is constructed based on a constraint set, a cloth yield factor, a Lagrangian operator, a time step and a vertex weight; the constraint set is constructed according to the unidirectional edges and curved edges contained in the triangular mesh corresponding to the template clothing; the vertex weight is determined according to the weight of the curved edge related to the vertex and the weight of the related unidirectional edge;

[0025] The positions of the corresponding vertices are updated according to the second offset to obtain the positions of the vertices after wrinkle processing, and the processed clothing sequence is obtained based on the positions of the vertices after wrinkle processing.

[0026] Optionally, obtaining the processed clothing sequence based on the vertex positions after the wrinkle processing includes:

[0027] Performing collision processing on the wrinkle-processed position in combination with the human body driving sequence to obtain the wrinkle and the vertex position after collision processing;

[0028] The processed clothing sequence is obtained according to the vertex positions after all wrinkles and collision processing.

[0029] In a second aspect, the present application discloses a highly generalizable clothing-driven animation generation device, comprising:

[0030] The sequence acquisition module is used to obtain the initialization clothing sequence and human body drive sequence in the posture space;

[0031] A network learning module, used for inputting the initialized clothing sequence and the human body drive sequence into a target network model; the target network model sequentially includes a graph network for extracting graph structure features, a temporal network for extracting temporal features, and a regression network;

[0032] A first offset acquisition module, used to obtain a first offset of a clothing vertex according to an output of the target network model, and to obtain a learned clothing sequence according to the initialization clothing sequence and the first offset;

[0033] A geometric post-processing module is used to perform geometric post-processing on the learned clothing sequence to obtain a processed clothing sequence, so as to generate an animation according to the processed clothing sequence result; the geometric post-processing includes wrinkle processing, and the wrinkle processing is used to simulate the deformation of clothing under cloth constraints.

[0034] In a third aspect, the present application discloses an electronic device, comprising:

[0035] Memory, used to store computer programs;

[0036] A processor is used to execute the computer program to implement the above-mentioned highly generalized clothing-driven animation generation method.

[0037] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the aforementioned highly generalized clothing-driven animation generation method.

[0038] In the present application, an initialization clothing sequence and a human driving sequence in a posture space are obtained; the initialization clothing sequence and the human driving sequence are input into a target network model; the target network model includes a graph network for extracting graph structure features, a temporal network for extracting temporal features, and a regression network in order; the first offset of the clothing vertex is obtained according to the output of the target network model, and the learned clothing sequence is obtained according to the initialization clothing sequence and the first offset; the learned clothing sequence is subjected to geometric post-processing to obtain a processed clothing sequence, so as to generate an animation according to the processed clothing sequence result; the geometric post-processing includes wrinkle processing, which is used to simulate the deformation of clothing under cloth constraints. It can be seen that by learning local features through a graph network, using a regression network to replace global feature generation, and obtaining overall features through local feature splicing, it is possible to express the features of multiple types of topological structures and improve the generalization of clothing driving. By learning temporal features through a temporal network, there is no need to rely on the output of the previous frame as input, which improves processing efficiency; it can have both practicality and real-time performance, and achieve a balance between performance and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0040] Figure 1 A flow chart of a highly generalized clothing-driven animation generation method provided in this application;

[0041] Figure 2 A schematic diagram of a specific unwrinkled garment and a schematic diagram of a wrinkled garment provided in the present application;

[0042] Figure 3 A specific clothing drive network architecture diagram provided for this application;

[0043] Figure 4 A specific skirt driving effect display diagram provided for this application;

[0044] Figure 5 A specific trousers driving effect display diagram provided for this application;

[0045] Figure 6 A schematic diagram of a specific clothing driving result provided for this application;

[0046] Figure 7 A schematic diagram of the structure of a highly generalized clothing-driven animation generation device provided in this application;

[0047] Figure 8 A structural diagram of an electronic device provided for this application. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] In the prior art, methods based on parameterized models, such as the GAPS algorithm, usually distribute the deformation of clothes in two spaces, the standard space and the posture space. In the standard space, the nonlinear deformation in the standard posture is learned through the network. The network input is the driving parameter and the output is the coordinates of the vertex. It is essentially a generative network based on the global feature vector. In order to transition to the posture space, a linear skinning weight is usually learned to convert the clothes in the standard posture into clothes in the specified shape and posture. This type of method has the following disadvantages: 1. The network relies too much on global features to generate vertex positions. The number of vertices output by the network is usually fixed, so it cannot solve the input of different topological structures, and the practicality is poor. Usually, one model is required for one template clothing. 2. The effect is poor. The network part only considers the vertex information; at the same time, the conversion from the standard space to the posture space is based on the assumption of linear change, which is quite different from the real scene. 3. There is a serious collision problem. It is difficult to ensure that there is no collision problem in the clothes in the two spaces after learning in two spaces. Methods based on graph networks, such as the HOOD algorithm, only operate in the pose space and therefore rely on a preprocessing process to transition the input in the standard space to the initialization input in the pose space. For this type of method, the prediction of the current frame needs to rely on the result of the previous frame, so it is a serial process as a whole. The serial architecture limits the efficiency of the graph network method; in addition, the input requirements are too high, such as requiring the input of three consecutive frames and relying on the output results of the previous frame, which limits the practicality of the network.

[0050] In order to overcome the above technical problems, the present application proposes a highly generalized clothing-driven animation generation method that can combine practicality and real-time performance and achieve a good balance between performance and efficiency.

[0051] The present application embodiment discloses a highly generalized clothing-driven animation generation method, see Figure 1 As shown, the method may include the following steps:

[0052] Step S11: Obtain the initialized clothing sequence and human body drive sequence in the posture space.

[0053] In this embodiment, based on the template clothing, human body driving parameters and the template human body, the initialization clothing sequence and human body driving sequence in the posture space are obtained through linear transformation. The human body driving parameters include shape parameters , attitude parameters Etc. The above-mentioned obtaining of the initialized clothing sequence and human body drive sequence in the posture space may specifically include: obtaining template clothing and template human body, generating a relationship matrix between the template clothing and the template human body; obtaining the clothing weight factor of the template clothing according to the product of the relationship matrix and the human body weight factor of the template human body; obtaining the clothing sequence and human body drive sequence in the posture space through linear transformation based on the template clothing, the human body drive parameters and the clothing weight factor; performing post-collision processing on the clothing sequence in the posture space according to the human body drive sequence to obtain the initialized clothing sequence.

[0054] That is, this stage is a preprocessing stage without network learning, based on the human weight factor , by obtaining the association matrix between the template clothing and the template human body , get the clothing weight factor , the calculation formula of clothing weight factor is as follows:

[0055] (1);

[0056] Among them, the human body parameterization weight Usually includes skinning factor, distortion factor (such as pose-blend-shape, etc.), association matrix It is usually related to the distance between the template human body and the clothes, such as determined by a nearest neighbor weight algorithm, and this application does not limit the specific algorithm. Then, according to the template clothes, human driving parameters and clothes weight factors, a clothes sequence in the posture space is obtained through linear transformation.

[0057] Furthermore, in order to ensure that the initialized clothing sequence does not collide with the human body sequence, a collision post-processing is introduced, which specifically includes two parts: collision detection and collision response. This embodiment provides a specific fast collision post-processing method, including the following steps: first, calculate the nearest neighbor distance from the clothing vertex to the human body , and the normal vector of the projection point of the clothing vertex on the human body ; Then, calculate the vertices of the clothes The initial offset If the value is less than 0, it means there is a collision; for the collision area point Perform offset and obtain offset coordinates . Based on the offset coordinates of all vertices Get the initial clothing sequence.

[0058] Step S12: inputting the initialized clothing sequence and the human body drive sequence into a target network model; the target network model sequentially includes a graph network for extracting graph structure features, a temporal network for extracting temporal features, and a regression network.

[0059] The obtained initialized clothing sequence and human body driving sequence are input into the target network model for network learning. The target network model includes a graph network for extracting graph structure features, a timing network for extracting timing features, and a regression network in order. Therefore, in view of the poor practicality of existing clothing driving based on parameterized models and the low generalization problem of requiring one model for one set of clothes, the present application uses a graph network to extract local features and adopts a regression method to replace the original generation method based on global features. There is no need for one model for one set of clothes, and multiple types of clothes can share one model, which has good generalization. The graph network is used to fully explore the hierarchical graph structure information between clothes and human bodies. In view of the low efficiency of the method of using graph networks for clothing driving in the prior art, the present application introduces a timing network to learn timing features without relying on the output result of the previous frame as input. At the same time, the network architecture provided supports parallel computing, which greatly improves the efficiency compared with the original serial method.

[0060] In some embodiments, after the initialized clothing sequence and the human body drive sequence are input into the target network model, the network learning process of the graph network specifically includes: based on the initialized clothing sequence and the human body drive sequence, using the graph network to extract node features and edge features; the node features include vertex velocity vectors, normal vectors, gravity, material parameters, node types, and node levels; the edge features include the length of the edge of the current frame graph structure, the length of the edge of the current frame graph structure after normalization, the length of the edge of the template graph, and the length of the edge of the normalized template graph; according to the features of the edges between the nodes and the connected clothing vertices, and the features of the edges between the clothing vertices and the human body vertices, the node features are updated to obtain updated node features; according to the updated node features of the first endpoint of the edge and the updated node features of the second endpoint of the edge, the edge features are updated to obtain updated edge features; the graph structure features are obtained based on the updated node features and the updated edge features.

[0061] Extracting node features based on graph network And edge features .

[0062] (2);

[0063] in, is the velocity vector of the node, represents the normal vector, represents gravity, , , is the material parameter, , is the Lame parameter, is the bending parameter; Indicates the node type (e.g. the node belongs to clothing or body, etc.), Represents the node level (such as points from the original model or model points that are downsampled by one time, etc.).

[0064] (3);

[0065] in, is the length of the edge of the current frame graph structure, is the normalized length of the edge of the current frame graph structure, is the length of the edge of the template graph, is the length of the edge of the normalized template graph. The template graph is the template graph corresponding to the above template clothing. The template graph is a predefined geometric model or topological structure used to describe the basic shape and structure of clothing; it is used to provide a basic framework for clothing generation, modification and animation. The template graph is an abstract representation of the geometric structure or topological structure of the template clothing.

[0066] Update nodes and edges through the network, and the node feature update formula is:

[0067] (4);

[0068] in, represents a feature extraction network (such as a multilayer perceptron), The current vertex of the vertex feature, Represents the vertices of the clothes Vertex with clothes The characteristics of the formed edges, Represents the vertices of the clothes Vertex Features of the edges formed.

[0069] The edge feature update formula is:

[0070] (5);

[0071] The update of the edge depends on the current edge characteristics , and the characteristics of the vertices at both ends of the edge , .

[0072] In some embodiments, the inputting of the initialized clothing sequence and the human body driving sequence into the target network model includes: obtaining graph network output data containing graph structure features according to the output of the graph network; performing dimension conversion on the graph network output data to obtain input data representing the time series dimension; and extracting time series features using the time series network based on the input data. It can be understood that the input dimension of the graph network is [BL, V, 3], where B is the batch number, L is the sequence number, V is the number of graph vertices, and 3 represents the feature dimension; the output dimension of the graph network is [BL, V, c], where c represents the feature dimension. In other words, the output dimension of the graph network is: number of sequences, number of vertices, and feature dimension, which indicates that the data is organized according to the sequence, each sequence contains multiple vertices, and each vertex has multiple features. The time series network extracts features between frames and needs to operate on the dimension representing the time series. Therefore, a dimension transformation is performed to convert the dimension into: [BV, L, c], and the transformed dimension is: number of vertices, number of sequences, and feature dimension, which indicates that the data is organized according to vertices, each vertex contains multiple sequences, and each sequence has multiple features.

[0073] Then, feature extraction is performed based on the temporal network, where the temporal network can use GRU (Gated Recurrent Unit, temporal coding network) and the like, and the formula is:

[0074] (6);

[0075] in, Represents a sequential network.

[0076] In some embodiments, inputting the initialized clothing sequence and the human body drive sequence into the target network model may include: obtaining temporal network output data including graph structure features and temporal features according to the output of the temporal network; and predicting the first offset of each clothing vertex using the regression network based on the temporal network output data. The formula is as follows:

[0077] (7);

[0078] in, Represents a regression network, such as a multilayer perceptron network.

[0079] In some embodiments, the inputting of the initialized clothing sequence and the human body drive sequence into the target network model may include: performing network learning of the regression network in combination with a preset loss function; the preset loss function is constructed based on one or more of clothing and human body collision loss, clothing self-collision loss, gravity loss, clothing material-related loss, clothing stretching loss, clothing and human body friction loss, external force loss, and time series smoothing constraints. That is, the present application adopts a self-supervised learning model, which can solve the current problem of difficult production and lack of data for clothing drive data. In order to generate a realistic and natural effect, the driving effect is improved by adding self-collision constraints, friction constraints, and external force constraints on the basis of commonly used constraints.

[0080] A specific loss function is as follows:

[0081] (8);

[0082] in, Represents the weight coefficient, and the loss function L includes the collision loss between clothes and human body , self-collision loss , gravity loss , clothing material related losses such as bending losses , stretch loss , friction loss between clothes and human body , and external force loss (such as random wind force, etc.) and time series smoothing constraints Among them, the gravity constraint simulates the physical gravity field; the material constraint simulates different materials, and the StVK (Saint-Venant-Kirchhoff) elastic constraint can be used.

[0083] The collision loss mainly restricts the penetration of the human body and clothing. The formula is as follows: (9);

[0084] in, represents the sdf value (sign distance function) from the vertex of the clothing to the human body; c represents the threshold constant, such as 2mm; Represents the number of vertices of the cloth; the formula uses the cubic exponent to strengthen the constraint on collision.

[0085] The self-collision constraint is used to constrain the self-collision problem of the clothes themselves. The common self-collision constraint is usually complicated to calculate. Considering the training efficiency problem, this patent provides a simple and effective constraint based on repulsion, which is mainly used to constrain two points that are not connected and close to each other. The formula is as follows:

[0086] (10);

[0087] in, , V and E represent the vertex and edge information of the clothes respectively, , is the top point of the clothes, express and There is no edge relationship, and means "and", d means the distance function, c is the threshold constant, express and The distance between them is less than the threshold c, In other words, this formula can be used to control the distance between two vertices that are close to each other and have no edge relationship, thus avoiding unreasonable wrinkles in clothes caused by the two vertices being too close.

[0088] In order to simulate the friction between clothes and body, for any pair of collision pairs, according to their relative displacement, this application constructs a continuous static and dynamic friction formula, which is specifically constructed based on the local friction factor, local contact normal force, local relative displacement of the collision pair, sliding projection matrix and continuous function. The friction formula is as follows:

[0089] (11);

[0090] in, is the local friction factor, is the local contact normal force, represents the local relative displacement of the collision pair, , that is, a matrix with 2 rows and 1 column; the transpose of T is the sliding projection matrix, , used to convert the spatial relative displacement ( ) is projected onto the plane, that is . is a continuous function that is used to create a smooth mapping between static and kinetic friction through the relative displacement of the collision pair; " ” represents the norm. This formula ensures that when the relative displacement is small, it manifests as static friction, and when the relative displacement is large, it transitions to dynamic friction, thus achieving continuity and smoothness of friction.

[0091] External force constraints are mainly used to simulate external factors such as wind force. To some extent, gravity can also be regarded as an external force. The formula is simplified as follows:

[0092] (12);

[0093] Among them, F represents the external force vector and n represents the direction vector of the force.

[0094] Timing constraints are mainly used to constrain the continuity of timing. The formula is as follows:

[0095] (13);

[0096] in, represents the weight coefficient, vel represents the velocity value of the vertex, and a represents the acceleration value of the vertex. " represents the norm, Indicates the number of vertices of the cloth.

[0097] Step S13: obtaining a first offset of clothing vertices according to the output of the target network model, and obtaining a learned clothing sequence according to the initialized clothing sequence and the first offset.

[0098] Get the first offset of the clothes vertex according to the output of the target network model , according to the vertex positions in the initialized clothing sequence and the first offset , determine the updated vertex position through network learning , thus obtaining the learned clothing sequence.

[0099] Step S14: performing geometric post-processing on the learned clothing sequence to obtain a processed clothing sequence, so as to generate an animation according to the processed clothing sequence result; the geometric post-processing includes wrinkle processing, and the wrinkle processing is used to simulate the deformation of clothing under cloth constraints.

[0100] It is understandable that neural networks usually have difficulty learning high-frequency details. In order to learn more details, a more complex network is often required, but this will sacrifice efficiency. This application proposes an efficient geometric post-processing to improve simulation details. Geometric post-processing can include collision post-processing and wrinkle post-processing. Of course, collision post-processing can be performed first and then wrinkle post-processing. This application does not limit the processing order.

[0101] For wrinkle post-processing, the learned clothing sequence is subjected to wrinkle post-processing to obtain a processed clothing sequence, specifically comprising: calculating the second offset of each vertex in the learned clothing sequence according to the wrinkle formula; the wrinkle formula is constructed based on the constraint set, the cloth yield factor, the Lagrangian operator, the time step and the vertex weight; the constraint set is constructed based on the unidirectional edges and curved edges contained in the triangular mesh corresponding to the template clothing; the unidirectional edge is the edge in the triangular mesh in a straight line shape, and the curved edge is the edge in the triangular mesh in a straight line shape, and the cloth yield factor represents the deformation ability and anti-deformation ability of the cloth when subjected to force. The vertex weight is determined according to the weight of the curved edge related to the vertex and the weight of the related unidirectional edge; the position of the corresponding vertex is updated according to the second offset to obtain the vertex position after wrinkle processing, and the processed clothing sequence is obtained based on the vertex position after wrinkle processing. That is, all calculated displacement increments are accumulated and the vertex position is updated to ensure that the surface of the object is reasonably deformed according to the physical constraints, and a real wrinkle effect is generated through multiple cycles of iteration.

[0102] The fold formula is as follows:

[0103] (14);

[0104] Among them, dL represents the second offset of the vertex, is the yield factor, Represents a set of constraints The i-th constraint in ; h is the time step, such as h=1 / 30 when the frame rate is 30 frames; L is the Lagrangian operator; represents the gradient; is the vertex weight, ; represents the weight of a one-way edge, Represents the weight of the curved edge. For example, if vertex i has 3 one-way edges, the weight of the one-way edge is 1 / 3. Update the position of the vertices in the learned clothing sequence according to the calculated second offset, traverse each vertex through a loop operation, and finally generate clothing with a wrinkled effect. Figure 2 Shown are schematic diagrams of unwrinkled clothing and schematic diagrams of wrinkled clothing.

[0105] In some embodiments, obtaining the processed clothing sequence based on the vertex positions after the wrinkle processing may include: performing collision processing on the processed positions of the wrinkles in combination with the human body drive sequence to obtain wrinkles and vertex positions after collision processing; obtaining the processed clothing sequence according to all wrinkles and vertex positions after collision processing. For example, the post-collision processing may adopt the method mentioned in step S11. Of course, height field collision processing may also be used, that is, by calculating the distance and gradient, determining whether penetration occurs, and pushing the vertex back to the legal position. In addition, cutting plane collision detection may be used to quickly limit the wrinkle depth to prevent self-penetration.

[0106] For example Figure 3 The figure shows a specific clothing drive network architecture diagram provided by the present application, which is both practical and real-time, and achieves a good balance between performance and efficiency. It solves the problem of poor practicality of existing methods based on parameterized models. There is no need for one model for each set of clothes. It can realize that multiple types of clothes share one model, and improve the generalization performance of the model. It improves the problem of poor effect of existing methods based on parameterized models, and fully explores the hierarchical graph structure information of human body sequences and clothes. It improves the collision problem existing in existing methods based on parameterized models, and effectively alleviates the penetration phenomenon of clothes and human bodies and the problem of self-collision of clothes. It improves the low efficiency problem of existing graph network methods, improves network support and formal calculation, and does not need to rely on the output results of the previous frame as input.

[0107] like Figure 4 The figure shows the effect of the skirt obtained by the above method. Figure 5 The figure shows the effect of the pants obtained by the above method. It can be seen that the present application scheme can achieve good results for both loose skirts and tight pants.

[0108] Table 1 Comparison results of tight clothes

[0109]

[0110] Table 2 Comparison results of skirts

[0111]

[0112] In Tables 1 and 2 above, GAPS (Geometry-Aware, Physics-Based, Self-Supervised Neural Garment Draping) is a clothing driving method based on a parameterized model, HOOD (Hierarchical Graphs for Generalized Modelling of Clothing Dynamics) is a clothing driving method based on a graph network, and ours represents the method proposed in this application. Quantitative evaluation indicators include: material loss (bending loss , stretch loss ), timing loss , collision loss between clothes and human body Among these indicators, the smaller the collision loss or timing loss is, the better it is in theory, and the smaller the material loss is, the better it is in theory, but they have no absolute relationship with the results.

[0113] like Figure 6 As shown, Figure 6 a is the clothes obtained by GAPS, Figure 6 b is that the clothes obtained by adopting the present application can alleviate the collision problem existing in GPAS.

[0114] As can be seen from the above, in this embodiment, the initialization clothing sequence and the human body drive sequence in the posture space are obtained; the initialization clothing sequence and the human body drive sequence are input into the target network model; the target network model includes a graph network for extracting graph structure features, a temporal network for extracting temporal features, and a regression network in sequence; the first offset of the clothing vertex is obtained according to the output of the target network model, and the learned clothing sequence is obtained according to the initialization clothing sequence and the first offset; the learned clothing sequence is subjected to geometric post-processing to obtain the processed clothing sequence, so as to generate animation according to the processed clothing sequence result; the geometric post-processing includes wrinkle processing, which is used to simulate the deformation of clothing under cloth constraints. It can be seen that by learning local features through the graph network, using the regression network to replace the global feature generation, and obtaining the overall features through the splicing of local features, it is possible to express the features of various types of topological structures and improve the generalization of clothing driving. By learning temporal features through the temporal network, there is no need to rely on the output of the previous frame as input, which improves processing efficiency; it can have both practicality and real-time performance, and achieve a balance between performance and efficiency.

[0115] Correspondingly, the present application also discloses a highly generalized clothing-driven animation generation device, see Figure 7 As shown, the device comprises:

[0116] A sequence acquisition module 11 is used to acquire an initialization clothing sequence and a human body drive sequence in a posture space;

[0117] A network learning module 12, used for inputting the initialized clothing sequence and the human body drive sequence into a target network model; the target network model sequentially includes a graph network for extracting graph structure features, a temporal network for extracting temporal features, and a regression network;

[0118] A first offset acquisition module 13, used to obtain a first offset of a clothing vertex according to an output of the target network model, and to obtain a learned clothing sequence according to the initialization clothing sequence and the first offset;

[0119] The geometric post-processing module 14 is used to perform geometric post-processing on the learned clothing sequence to obtain a processed clothing sequence, so as to generate an animation according to the processed clothing sequence result; the geometric post-processing includes wrinkle processing, and the wrinkle processing is used to simulate the deformation of clothing under cloth constraints.

[0120] As can be seen from the above, in this embodiment, the initialization clothing sequence and the human body drive sequence in the posture space are obtained; the initialization clothing sequence and the human body drive sequence are input into the target network model; the target network model includes a graph network for extracting graph structure features, a temporal network for extracting temporal features, and a regression network in sequence; the first offset of the clothing vertex is obtained according to the output of the target network model, and the learned clothing sequence is obtained according to the initialization clothing sequence and the first offset; the learned clothing sequence is subjected to geometric post-processing to obtain the processed clothing sequence, so as to generate animation according to the processed clothing sequence result; the geometric post-processing includes wrinkle processing, which is used to simulate the deformation of clothing under cloth constraints. It can be seen that by learning local features through the graph network, using the regression network to replace the global feature generation, and obtaining the overall features through the splicing of local features, it is possible to express the features of various types of topological structures and improve the generalization of clothing driving. By learning temporal features through the temporal network, there is no need to rely on the output of the previous frame as input, which improves processing efficiency; it can have both practicality and real-time performance, and achieve a balance between performance and efficiency.

[0121] In some specific embodiments, the network learning module 12 may specifically include:

[0122] A feature extraction unit, configured to extract node features and edge features using a graph network based on the initialized clothing sequence and the human body drive sequence; the node features include vertex velocity vectors, normal vectors, gravity, material parameters, node types, and node levels; the edge features include the length of an edge of a current frame graph structure, the length of an edge of the current frame graph structure after normalization, the length of an edge of a template graph, and the length of an edge of a normalized template graph;

[0123] A node feature updating unit, used to update the node feature according to the feature of the edge between the node and the connected clothing vertex, and the feature of the edge between the clothing vertex and the human body vertex, to obtain an updated node feature;

[0124] An edge feature updating unit, configured to update the edge feature according to the updated node feature of the first endpoint of the edge and the updated node feature of the second endpoint of the edge to obtain an updated edge feature;

[0125] A graph structure feature determination unit is used to obtain the graph structure feature based on the updated node feature and the updated edge feature.

[0126] In some specific embodiments, the network learning module 12 may specifically include:

[0127] A graph network output data acquisition unit, used to obtain graph network output data containing graph structure features according to the output of the graph network;

[0128] A dimension conversion unit, used to perform dimension conversion on the graph network output data to obtain input data representing a time series dimension;

[0129] A time series feature extraction unit is used to extract time series features based on the input data using the time series network.

[0130] In some specific embodiments, the network learning module 12 may specifically include:

[0131] A time series network output data acquisition unit, used to obtain time series network output data including graph structure features and time series features according to the output of the time series network;

[0132] The first offset calculation unit is used to predict the first offset of each clothing vertex using the regression network based on the output data of the timing network.

[0133] In some specific embodiments, the network learning module 12 may specifically include:

[0134] A loss constraint unit is used to perform network learning of the regression network in combination with a preset loss function; the preset loss function is constructed based on one or more of clothing-human body collision loss, clothing self-collision loss, gravity loss, clothing material-related loss, clothing stretching loss, clothing-human body friction loss, external force loss, and time series smoothing constraint.

[0135] In some specific embodiments, the geometric post-processing module 14 may specifically include:

[0136] A second offset calculation unit is used to calculate the second offset of each vertex in the learned clothing sequence according to a wrinkle formula; the wrinkle formula is constructed based on a constraint set, a cloth yield factor, a Lagrangian operator, a time step and a vertex weight; the constraint set is constructed according to the unidirectional edges and curved edges contained in the triangular mesh corresponding to the template clothing; the vertex weight is determined according to the weight of the curved edge related to the vertex and the weight of the related unidirectional edge;

[0137] A vertex updating unit is used to update the position of the corresponding vertex according to the second offset to obtain the vertex position after wrinkle processing, and obtain the processed clothing sequence based on the vertex position after wrinkle processing.

[0138] In some specific embodiments, the vertex updating unit may specifically include:

[0139] A collision processing unit, used for performing collision processing on the wrinkle-processed position in combination with the human body driving sequence to obtain the wrinkle and the vertex position after collision processing;

[0140] The processed clothing sequence determining unit is used to obtain the processed clothing sequence according to the positions of all wrinkles and vertices after collision processing.

[0141] Furthermore, the present application also discloses an electronic device, see Figure 8 As shown, the contents in the drawings should not be considered as any limitation on the scope of use of the present application.

[0142] Figure 8 The present invention provides a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the highly generalized clothing-driven animation generation method disclosed in any of the aforementioned embodiments.

[0143] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0144] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system 221, a computer program 222, and data 223 including graph structure features, etc. The storage method can be temporary storage or permanent storage.

[0145] The operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, so as to realize the operation and processing of the massive data 223 in the memory 22 by the processor 21, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the animation generation method of highly generalized clothing driven by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.

[0146] Furthermore, an embodiment of the present application also discloses a computer storage medium, in which computer executable instructions are stored. When the computer executable instructions are loaded and executed by a processor, the steps of the highly generalized clothing-driven animation generation method disclosed in any of the aforementioned embodiments are implemented.

[0147] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0148] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0149] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0150] The above is a detailed introduction to a highly generalized clothing-driven animation generation method, device, equipment and storage medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A highly generalizable clothing-driven animation generation method, characterized in that: include: Get the initialized clothing sequence and human body drive sequence in the posture space; Inputting the initialized clothing sequence and the human body drive sequence into a target network model; the target network model sequentially includes a graph network for extracting graph structure features, a temporal network for extracting temporal features, and a regression network; Obtaining a first offset of a clothing vertex according to an output of the target network model, and obtaining a learned clothing sequence according to the initialized clothing sequence and the first offset; The learned clothing sequence is subjected to geometric post-processing to obtain a processed clothing sequence, so as to generate an animation according to the processed clothing sequence result; the geometric post-processing includes wrinkle processing, and the wrinkle processing is used to simulate the deformation of clothing under cloth constraints.

2. The highly generalizable clothing-driven animation generation method according to claim 1, characterized in that: The step of inputting the initialized clothing sequence and the human body drive sequence into a target network model comprises: Based on the initialized clothing sequence and the human body drive sequence, a graph network is used to extract node features and edge features; the node features include vertex velocity vectors, normal vectors, gravity, material parameters, node types, and node levels; the edge features include the length of the edge of the current frame graph structure, the length of the edge of the current frame graph structure after normalization, the length of the edge of the template graph, and the length of the edge of the template graph after normalization; According to the features of the edge between the node and the connected clothing vertex, and the features of the edge between the clothing vertex and the human body vertex, the node features are updated to obtain updated node features; According to the updated node feature of the first endpoint of the edge and the updated node feature of the second endpoint of the edge, the edge feature is updated to obtain an updated edge feature; The graph structure feature is obtained based on the updated node feature and the updated edge feature.

3. The highly generalizable clothing-driven animation generation method according to claim 1, characterized in that: The step of inputting the initialized clothing sequence and the human body drive sequence into a target network model comprises: According to the output of the graph network, obtaining graph network output data containing graph structure features; Performing dimension conversion on the graph network output data to obtain input data representing a time series dimension; The temporal network is used to extract temporal features based on the input data.

4. The highly generalizable clothing-driven animation generation method according to claim 1, characterized in that: The step of inputting the initialized clothing sequence and the human body drive sequence into a target network model comprises: According to the output of the temporal network, temporal network output data including graph structure features and temporal features is obtained; Based on the output data of the temporal network, the regression network is used to predict the first offset of each clothing vertex.

5. The highly generalizable clothing-driven animation generation method according to claim 1, characterized in that: The step of inputting the initialized clothing sequence and the human body drive sequence into a target network model comprises: The network learning of the regression network is performed in combination with a preset loss function; the preset loss function is constructed based on one or more of the clothing-human body collision loss, clothing self-collision loss, gravity loss, clothing material-related loss, clothing stretching loss, clothing-human body friction loss, external force loss, and time series smoothing constraints.

6. The highly generalizable clothing-driven animation generation method according to any one of claims 1 to 5, characterized in that: The step of performing geometric post-processing on the learned clothing sequence to obtain a processed clothing sequence comprises: The second offset of each vertex in the learned clothing sequence is calculated according to the wrinkle formula; the wrinkle formula is constructed based on a constraint set, a cloth yield factor, a Lagrangian operator, a time step and a vertex weight; the constraint set is constructed according to the unidirectional edges and curved edges contained in the triangular mesh corresponding to the template clothing; the vertex weight is determined according to the weight of the curved edge related to the vertex and the weight of the related unidirectional edge; The positions of the corresponding vertices are updated according to the second offset to obtain the positions of the vertices after wrinkle processing, and the processed clothing sequence is obtained based on the positions of the vertices after wrinkle processing.

7. The highly generalizable clothing-driven animation generation method according to claim 6, characterized in that: The step of obtaining the processed clothing sequence based on the vertex positions after the wrinkle processing comprises: Performing collision processing on the wrinkle-processed position in combination with the human body driving sequence to obtain the wrinkle and the vertex position after collision processing; The processed clothing sequence is obtained according to the vertex positions after all wrinkles and collision processing.

8. A highly generalizable clothing-driven animation generation device, characterized in that: include: The sequence acquisition module is used to obtain the initialization clothing sequence and human body drive sequence in the posture space; A network learning module, used for inputting the initialized clothing sequence and the human body drive sequence into a target network model; the target network model sequentially includes a graph network for extracting graph structure features, a temporal network for extracting temporal features, and a regression network; A first offset acquisition module, used to obtain a first offset of a clothing vertex according to an output of the target network model, and to obtain a learned clothing sequence according to the initialization clothing sequence and the first offset; A geometric post-processing module is used to perform geometric post-processing on the learned clothing sequence to obtain a processed clothing sequence, so as to generate an animation according to the processed clothing sequence result; the geometric post-processing includes wrinkle processing, and the wrinkle processing is used to simulate the deformation of clothing under cloth constraints.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the highly generalized clothing-driven animation generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein when the computer programs are executed by a processor, the highly generalized clothing-driven animation generation method as described in any one of claims 1 to 7 is implemented.