Digital human face model face capacity grid editing method, digital human face model determining method, digital human model determining method and digital human video generating method

By generating target facial meshes by receiving explicit and implicit parameters, the problem of template component dependence and complex operation in existing technologies is solved, and personalized facial meshes can be simplified and generated in a coordinated manner.

CN121837539APending Publication Date: 2026-04-10MOFA (SHANGHAI) INFORMATION TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing digital human model facial generation solutions rely on pre-designed template components, which cannot meet users' personalized needs and are complex to operate, making it difficult for ordinary users to understand and adjust facial details.

Method used

By receiving explicit parameters from the user and combining them with implicit parameters from the initial facial mesh, a target facial mesh is generated using a pre-trained encoder and decoder, simplifying user operations and enabling personalized editing of facial features.

Benefits of technology

It reduces the complexity of user operations, increases the editing freedom and overall coordination of facial meshes, generates diverse facial meshes that meet user aesthetics, and improves user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837539A_ABST
    Figure CN121837539A_ABST
Patent Text Reader

Abstract

The invention provides a face capacity grid editing method of a digital human face model, a digital human face model determination method, a digital human model determination method and a digital human video generation method, and relates to the technical field of digital humans. The operation complexity of a user can be reduced, the user does not need professional art knowledge, the use threshold is low, and the editing freedom degree is large. After the target dominant parameter of the user is received, the target dominant parameter is combined with the initial implicit parameter of the initial face grid to generate the target face grid, so that the editing requirement of the user can be met, the diversified target face grid meeting the user requirement can be generated, the overall coordination of the target face grid can be ensured, and the user experience is improved. The situation that the face grid is not coordinated due to the fact that the user edits a large number of face parameters is avoided, the target face grid conforms to the aesthetic appreciation of the user, and the satisfaction degree and the use experience of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital humans, and in particular to a face grid editing method of a digital human face model, a digital human face model determination method, a digital human model determination method and a digital human video generation method. BACKGROUND

[0002] With the continuous development of digital human generation technology and the continuous increase of digital human application scenarios, users are no longer satisfied with using pre-designed digital human models, but hope to be able to design and use personalized digital human models by themselves.

[0003] Generally, the face of a 3D digital human can be divided into different components such as facial features, face shape, hairstyle, skin color, etc. A face generation scheme of a digital human model is to pre-design different template components for each component, and based on the user's selection, use different template components to replace the corresponding part of the general 3D digital human model, thereby designing the face of the personalized 3D digital human model. However, this scheme is extremely dependent on pre-designed template components, and users cannot optimize and adjust the face of the 3D digital human model according to their own needs. Moreover, the existing template components lack diversity and cannot meet the needs of all users. In addition, since the related parameters of the template components are pre-set, there may be a situation that the overall face of the replaced 3D digital human model is not coordinated.

[0004] Another face generation scheme of a digital human model is to use a large number of parameters to define the face of a 3D digital human model, but the facial details correspond to a large number of trivial parameters, and ordinary users have difficulty understanding the association between these trivial parameters and facial details. To modify the facial details, the values of the related parameters need to be modified, which has a high operation threshold. For ordinary users who lack professional art knowledge, after editing the local features of a certain area using the method of modifying related parameters, it is difficult for other areas of the face to coordinate with the edited local features, and corresponding adjustments need to be made to other areas of the face, which is difficult. SUMMARY

[0005] The present application provides a face grid editing method of a digital human face model, a digital human face model determination method, a digital human model determination method and a digital human video generation method to solve the defects in the related art.

[0006] The present application provides a face grid editing method of a digital human face model, comprising: receiving a first editing instruction of a user on an initial face grid of a digital human face model, the first editing instruction comprising a target explicit parameter, the target explicit parameter being used to represent an explicit facial feature of a target face grid edited by the user; generate the target face mesh based on the target explicit parameters and initial implicit parameters of the initial face mesh; the initial implicit parameters are used to represent implicit facial features of the initial face mesh.

[0007] According to the face mesh editing method of the digital human face model provided by the application, the target face mesh is generated based on the target explicit parameters and the initial implicit parameters of the initial face mesh, and the method comprises the following steps: generate an implicit space vector based on the target explicit parameters and the initial implicit parameters; generate the target face mesh by applying a pre-trained decoder based on the implicit space vector.

[0008] The application further provides a face mesh editing method of a digital human face model, and the initial implicit parameters are extracted from the initial face mesh based on an encoder; The encoder comprises a plurality of convolutional pooling layer groups and a fully connected layer; Each convolutional pooling layer group is connected to the fully connected layer in sequence; Each convolutional pooling layer group comprises a convolutional layer and a pooling layer; The convolutional layer is used to apply a spiral convolution kernel to perform feature extraction on the input and generate a feature mesh graph; The pooling layer is used to perform down-sampling on the feature mesh graph based on a pre-determined mapping relationship of mesh graphs with different fineness to obtain a down-sampling result; The fully connected layer is used to integrate the input to obtain the initial implicit parameters.

[0009] The application further provides a face mesh editing method of a digital human face model, and initial explicit parameters of the initial face mesh are extracted from the initial face mesh based on an explicit parameter extractor, and the initial explicit parameters are used to represent explicit facial features of the initial face mesh; The initial face mesh is composed of a plurality of face region meshes; the explicit parameter extractor and the encoder are respectively a region explicit parameter extractor and a region encoder corresponding to each face region mesh; The region explicit parameters of any face region mesh are extracted from the initial face mesh based on the corresponding region explicit parameter extractor; The region implicit parameters of any face region mesh are extracted from the initial face mesh based on the corresponding region encoder.

[0010] The application further provides a face mesh editing method of a digital human face model, and the implicit space vector is generated based on the target explicit parameters and the initial implicit parameters, and the method comprises the following steps: generate a target latent space vector corresponding to at least one target face region mesh in the target face mesh based on the target region explicit parameters and the initial region implicit parameters, and generate a relevant latent space vector based on the region explicit parameters and the region implicit parameters of the relevant face region mesh in the initial face mesh; The target region explicit parameters are used to represent region explicit facial features of at least one target face region mesh in the target face mesh, and the initial region implicit parameters are region implicit parameters of at least one target face region mesh in the initial face mesh.

[0011] The application further provides a face mesh editing method for a digital human face model, and the target face mesh is generated based on the latent space vector and a pre-trained decoder. The target face mesh is generated based on the target latent space vector, the relevant latent space vector and region fusion parameters of the initial face mesh and a pre-trained fusion decoder. The region fusion parameters have a correlation relationship with global position features of each face region mesh in the initial face mesh.

[0012] The application further provides a face mesh editing method for a digital human face model, and the region fusion parameters are extracted from the initial face mesh based on a first fusion parameter extractor. The fusion decoder, the region encoder and the region decoder corresponding to each face region mesh in the initial face mesh are jointly trained based on first face mesh samples before and after face editing, the first fusion parameter extractor and the region explicit parameter extractor corresponding to each face region mesh in the initial face mesh.

[0013] The application further provides a face mesh editing method for a digital human face model, and the target face mesh is generated based on the latent space vector and a pre-trained decoder. The target face region mesh is generated based on the target latent space vector and a pre-trained region decoder corresponding to the target face region mesh. The relevant face region mesh is generated based on the relevant latent space vector and a pre-trained region decoder corresponding to the relevant face region mesh. Each region transformation mesh is obtained by performing position transformation on the target face region mesh and the relevant face region mesh based on the region fusion parameters of each face region mesh. The target face mesh is obtained by fusing each region transformation mesh based on a fusion network.

[0014] The present invention also provides a method for editing the facial mesh of a digital human face model, wherein the region fusion parameters are pre-extracted from the initial facial mesh based on a second fusion parameter extractor; The fusion network, along with the region encoder and region decoder corresponding to each facial region grid in the initial facial mesh, are obtained by joint training using the second fusion parameter extractor and the region explicit parameter extractor corresponding to each facial region grid in the initial facial mesh, based on the second facial mesh samples before and after face editing.

[0015] The present invention also provides a method for editing the facial mesh of a digital human face model, wherein receiving a user's first editing instruction on the initial facial mesh of the digital human face model includes, prior to: Receive the user's second editing instruction on the general facial mesh; the second editing instruction includes specified facial region mesh information, the specified facial region mesh information being used to identify the specified facial region mesh; Based on the explicit and implicit parameters of the specified facial region mesh, the general facial mesh is edited to obtain the initial facial mesh.

[0016] The present invention also provides a method for determining a digital human face model, comprising: Based on the facial mesh editing method of the digital human face model described above, the target facial mesh is determined; Based on the target facial mesh and the target face texture, the target digital human face model is determined.

[0017] The present invention also provides a method for determining a digital human model, comprising: Based on the above-described method for determining digital human face models, the target digital human face model is determined. Based on the target digital human face model and the target body model, the target digital human model is determined.

[0018] The present invention also provides a method for generating digital human videos, comprising: Based on the above-described method for determining digital human models, the target digital human model is determined. Based on the target digital human model, a digital human video is generated.

[0019] The present invention also provides a facial mesh editing device for a digital human face model, comprising: The instruction receiving module is used to receive the user's first editing instruction on the initial facial mesh of the digital human face model. The first editing instruction includes target explicit parameters, which are used to characterize the explicit facial features of the target facial mesh after the user's editing. A mesh generation module is used to generate the target facial mesh based on the target explicit parameters and the initial implicit parameters of the initial facial mesh; the initial implicit parameters are used to characterize the implicit facial features of the initial facial mesh.

[0020] The present invention also provides a digital face model determination device, comprising: The mesh determination module is used to determine the target facial mesh based on the facial mesh editing method of the digital human face model described above. The first face model determination module is used to determine the target digital human face model based on the target facial mesh and the target face texture.

[0021] The present invention also provides a digital human model determination device, comprising: The second face model determination module is used to determine the target digital face model based on the above-described digital face model determination method. The first digital human model determination module is used to determine the target digital human model based on the target digital human face model and the target body model.

[0022] The present invention also provides a digital human video generation apparatus, comprising: The second digital human model determination module is used to determine the target digital human model based on the above-described digital human model determination method. The digital human video generation module is used to generate digital human videos based on the target digital human model.

[0023] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a facial mesh editing method for a digital human face model, or a digital human face model determination method, or a digital human model determination method, or a digital human video generation method as described above.

[0024] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a facial mesh editing method for a digital human face model, a digital human face model determination method, a digital human model determination method, or a digital human video generation method as described above.

[0025] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements a facial mesh editing method for a digital human face model, a digital human face model determination method, a digital human model determination method, or a digital human video generation method as described above.

[0026] This invention provides a method for editing facial meshes of digital human face models, a method for determining digital human face models, a method for determining digital human models, and a method for generating digital human videos. First, it receives a user's initial editing instruction for the initial facial mesh of the digital human face model, which includes target explicit parameters. Then, it generates the target facial mesh using the target explicit parameters and the initial implicit parameters of the initial facial mesh. This method allows users to edit only the explicit facial features of the initial facial mesh, reducing the complexity of user operations. Users do not need professional art knowledge, making it easy to use and offering greater editing freedom. After receiving the user's target explicit parameters, this method combines them with the initial implicit parameters of the initial facial mesh to generate the target facial mesh. This satisfies the user's editing needs, generating diverse target facial meshes that meet user requirements, while also ensuring the overall coordination of the target facial mesh. It avoids facial mesh inconsistencies caused by users editing a large number of facial parameters, ensuring the target facial mesh conforms to the user's aesthetic preferences and improving user satisfaction and experience. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a flowchart illustrating the facial mesh editing method for digital human face models provided by this invention.

[0029] Figure 2 This is a schematic diagram of the joint training process of the encoder and decoder in the facial mesh editing method of the digital human face model provided by the present invention.

[0030] Figure 3 This is one of the schematic diagrams of the shape of the spiral convolution kernel in the facial mesh editing method of the digital human face model provided by the present invention.

[0031] Figure 4 This is the second schematic diagram of the shape of the spiral convolution kernel in the facial mesh editing method for digital human face models provided by this invention.

[0032] Figure 5 This is a schematic diagram illustrating the joint training process of the fusion decoder, the region encoder, and the region decoder corresponding to each facial region mesh in the initial facial mesh in the facial mesh editing method of the digital human face model provided by the present invention.

[0033] Figure 6This is a schematic diagram illustrating the joint training process of the fusion network and the region encoders and region decoders corresponding to each facial region grid in the initial facial grid in the facial mesh editing method of the digital human face model provided by the present invention.

[0034] Figure 7 This is a flowchart illustrating the digital human face model determination method provided by the present invention.

[0035] Figure 8 This is a flowchart illustrating the digital human model determination method provided by the present invention.

[0036] Figure 9 This is a flowchart illustrating the digital human video generation method provided by the present invention.

[0037] Figure 10 This is a schematic diagram of the structure of the facial mesh editing device for digital human face models provided by the present invention.

[0038] Figure 11 This is a schematic diagram of the structure of the digital human face model determination device provided by the present invention.

[0039] Figure 12 This is a schematic diagram of the structure of the digital human model determination device provided by the present invention.

[0040] Figure 13 This is a schematic diagram of the structure of the digital human video generation device provided by the present invention.

[0041] Figure 14 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0043] Digital humans are virtual avatars created using technologies such as computer graphics and artificial intelligence, possessing human-like appearance and behavioral characteristics. They are widely used in brand marketing, intelligent customer service, and cultural tourism services. A digital human exists in digital form in virtual space, possessing anthropomorphic appearance, behavior, and interactive capabilities. Generally speaking, a digital human needs to have the following three main characteristics: a digital appearance, human performance abilities, and interactive capabilities. The digital appearance of a 3D digital human is primarily based on 3D modeling technology, that is, using 3D modeling software to construct a digital human model with 3D data in a virtual 3D space.

[0044] Similar to real people, different digital human models have different appearances, allowing users to distinguish them based on their looks. Among the many components of appearance, the face is crucial, and 3D digital human face editing technology allows users to design personalized 3D digital humans.

[0045] Typically, a 3D digital human's face can be divided into different components such as facial features, face shape, hairstyle, and skin tone. One method for generating a digital human model's face involves pre-designing different template components for each component. Based on the user's choices, different template components are used to replace corresponding parts of a generic 3D digital human model, thereby designing a personalized 3D digital human model's face. However, while this method can design personalized 3D digital human model faces based on user choices, it is extremely dependent on pre-designed template components. The final designed 3D digital human model's face is actually a combination of template components. The more template components there are, the more possible combinations of 3D digital human model faces can be generated.

[0046] Because each user has their own personalized aesthetic and choices, although various different 3D digital human model faces can be designed through permutations and combinations, the number of possible combinations is still limited, resulting in insufficient diversity. If the face desired by the user differs from the face generated by permutations and combinations of pre-designed template components, the user's needs cannot be met, severely impacting the user experience. Moreover, this solution is heavily reliant on pre-designed template components, preventing users from optimizing and adjusting the 3D digital human model's face according to their own requirements.

[0047] Furthermore, since the parameters of the template components are all preset, there may be instances where the overall appearance of the replaced 3D digital human model is inconsistent. For example, the face of a general 3D digital human model is slim, while the template component has wide eyes, resulting in an overall inconsistency in the appearance of the replaced 3D digital human model. Similarly, if multiple different template components are used to replace corresponding parts of the face of a general 3D digital human model, mismatches may occur between the different template components, leading to an overall inconsistency in the appearance of the replaced 3D digital human model.

[0048] Besides template component replacement, another approach to digital human model facial generation involves defining the 3D digital human model's face using numerous parameters. However, facial details correspond to a large number of intricate parameters, making it difficult for ordinary users to understand the relationship between these parameters and facial details. Modifying facial details requires adjusting the values ​​of these parameters, which presents a high barrier to entry. For ordinary users lacking professional art knowledge, editing local features in one area using parameter modification often results in inconsistencies in other facial areas, necessitating further adjustments to those areas, making the process quite challenging.

[0049] Based on this, this embodiment of the invention provides a method for editing the facial mesh of a digital human face model, which assists in the editing and determination of the digital human face model.

[0050] Figure 1 This is a flowchart illustrating a method for editing the facial mesh of a digital human face model provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes: S11, Receive the user's first editing instruction on the initial facial mesh of the digital human face model. The first editing instruction includes target explicit parameters, which are used to characterize the explicit facial features of the target facial mesh after the user's editing. S12, Generate the target face mesh based on the target explicit parameters and the initial implicit parameters of the initial face mesh; the initial implicit parameters are used to characterize the implicit facial features of the initial face mesh.

[0051] Specifically, the face mesh editing method for digital human face models provided in this embodiment of the invention is executed by a face mesh editing device for digital human face models. This device can be configured in a computer, which can be a local computer or a cloud computer. The local computer can be a computer, tablet, etc., and no specific limitation is made here.

[0052] First, step S11 is executed, which can display the user a digital human face model to be edited, or a digital human model containing a digital human face model. Whether it's a digital human face model or a digital human model, the default appearance features are neutral features. The digital human face model can include a facial mesh and a face map. The facial mesh is the basic geometric framework of the digital human face model, defining the facial contours, facial features, and skin undulations through vertex coordinates and topological structure. It also provides support for animation deformation, lighting calculations, and texture mapping, ultimately achieving accurate expression of static shape and dynamic expressions. Here, the digital human face model is a three-dimensional face model used to represent the face of the digital human model. The facial mesh is a three-dimensional structure, and the face map is a two-dimensional image. A mapping relationship is established between each pixel in the face map and each vertex in the facial mesh. Based on this mapping relationship, the face map is then overlaid on the surface of the facial mesh to form a three-dimensional digital human face model. It is understandable that after establishing the aforementioned mapping relationship, changing the position of several vertices in the facial mesh in three-dimensional space will still maintain the mapping relationship between the corresponding pixels in the face texture.

[0053] A facial mesh can be a complex three-dimensional structure composed of a large number of triangles or quadrilaterals, with a hollow interior, and can be represented by the coordinates of each vertex. The smaller the area of ​​each triangle or quadrilateral used to piece together the facial mesh, the more triangles or quadrilaterals the facial mesh contains, and the smoother and more detailed the surface of the digital human face model will look, resulting in a better display effect.

[0054] A facial mesh can include explicit and implicit parameters, which characterize the facial features of the mesh. Explicit parameters characterize the dominant facial features of the mesh and can be extracted from the mesh using a predefined explicit parameter extractor. These dominant facial features refer to the quantifiable and semantically clear main facial features of the mesh, such as facial features like face shape (thickness, length), and features like eyebrow position, spacing, length, and thickness.

[0055] Latent parameters are used to characterize the latent facial features of a facial mesh. These latent facial features can be encoded features obtained by encoding the facial mesh using a pre-trained encoder. These latent facial features refer to secondary facial features that are difficult to quantify or whose semantics are ambiguous, but which are related to complex features such as facial geometric quality and aesthetics. In other words, they are facial features other than explicit facial features.

[0056] As a whole, the different components of a digital human face model will influence each other. A change in one local feature will cause a corresponding change in other local features. For example, if the face is larger, the distribution and spacing of the facial features will also be slightly larger. If the bridge of the nose is higher, the eye sockets will also be deeper. These implicit relationships and other facial features are called latent facial features.

[0057] Understandably, since latent facial features are difficult to quantify or have ambiguous semantics, they cannot be directly described. Therefore, using encoded features to represent latent facial features can make latent facial features explicit, which helps to make the facial mesh explicit.

[0058] The initial facial mesh of the digital human face model is the current facial mesh of the digital human face model. Users can input first editing commands that represent their intention to edit the initial facial mesh of the digital human face model, based on the displayed digital human face model or the visual effect of the digital human model.

[0059] To facilitate user editing, the display interface can be configured with sliders for users to edit initial explicit parameters. Each initial explicit parameter can have its own slider. Each slider can be named according to its corresponding initial display parameter. The slider can be set on the slider bar, and different positions of the slider on the slider bar represent different values ​​of the corresponding initial explicit parameter.

[0060] Furthermore, users can edit the corresponding initial explicit parameters by moving one or more sliders on the slider bar. The operation is convenient and requires no professional art knowledge.

[0061] The first editing instruction may include target explicit parameters, which are the explicit parameters of the target face mesh after user editing, used to characterize the explicit facial features of the target face mesh after user editing. In other words, the user's first editing instruction is used to instruct the explicit parameters of the target face mesh after user editing to be the target explicit parameters. The user only edits the initial explicit parameters in the initial face mesh, while the initial implicit parameters in the initial face mesh remain unchanged.

[0062] Next, step S12 is executed to generate the target face mesh using the target explicit parameters and the initial implicit parameters of the initial face mesh. The generation process of the target face mesh is essentially the process of editing the initial face mesh using the first editing command. During this generation process, the topological relationships between the vertices in the initial face mesh must remain unchanged; only the vertex coordinates related to the first editing command are adjusted. Compared to the initial face mesh, the target face mesh has the same topological relationships between all vertices; the only difference is the vertex coordinates related to the first editing command.

[0063] In the process of generating the target facial mesh, the target explicit parameters are used to determine the main facial features of the target facial mesh, while the initial implicit parameters are used for the overall coordination of the target facial mesh after the target explicit parameters are transformed.

[0064] Since latent parameters are encoded features, the target facial mesh can be obtained by decoding both the target explicit parameters and the initial latent parameters. For example, the target explicit parameters and the initial latent parameters can be input into a pre-trained decoder, which decodes the input and outputs the vertex coordinates of the target facial mesh. The target facial mesh can then be determined using these vertex coordinates.

[0065] In this embodiment of the invention, the encoder used to extract the initial latent parameters and the decoder used to produce the target facial mesh can be jointly trained using an explicit parameter extractor based on third facial mesh samples before and after face editing. Here, the explicit parameter extractor, encoder, and decoder can constitute a facial feature transfer model. Both the third facial mesh samples before and after face editing can be characterized by the coordinates of each vertex they contain.

[0066] The joint training process of the encoder and decoder is as follows: Figure 2 As shown, it can include: using an explicit parameter extractor to extract explicit parameters from the third facial mesh sample before facial editing to obtain explicit parameter samples, and using an encoder to extract implicit parameters from the third facial mesh sample before facial editing to obtain implicit parameter samples.

[0067] Subsequently, the explicit parameter samples and implicit parameter samples are input into the decoder to obtain the decoder output. This decoder output is the prediction result of the third face mesh sample after face editing, and can be characterized by the predicted values ​​of each vertex coordinate.

[0068] Subsequently, the first training loss is calculated by using the difference between the decoding result output by the decoder and the third facial mesh sample after face editing. The structural parameters of the encoder and decoder are then iteratively optimized using the first training loss until the preset number of iterations or the training loss converges, thus completing the training of the encoder and decoder.

[0069] The facial mesh editing method for digital human face models provided in this embodiment of the invention first receives a first editing instruction from a user on the initial facial mesh of the digital human face model, which includes target explicit parameters. Then, using the target explicit parameters and the initial implicit parameters of the initial facial mesh, a target facial mesh is generated. This method allows users to edit only the explicit facial features of the initial facial mesh, reducing the complexity of user operations. Users do not need professional art knowledge, making it easy to use and offering greater editing freedom. After receiving the user's target explicit parameters, this method combines them with the initial implicit parameters of the initial facial mesh to generate the target facial mesh. This satisfies the user's editing needs, generating diverse target facial meshes that meet user requirements, while also ensuring the overall coordination of the target facial mesh. This avoids facial mesh inconsistencies caused by users editing a large number of facial parameters, ensuring the target facial mesh conforms to the user's aesthetic preferences and improving user satisfaction and user experience.

[0070] Based on the above embodiments, a target facial mesh is generated based on the target explicit parameters and the initial implicit parameters of the initial facial mesh, including: Based on the target explicit parameters and the initial implicit parameters, a latent space vector is generated; Based on the latent space vectors, a pre-trained decoder is applied to generate a target facial mesh.

[0071] Specifically, in the process of generating the target facial mesh using the target explicit parameters and the initial implicit parameters, the target explicit parameters and the initial implicit parameters can be fused first to obtain the latent space vector. The latent space is an abstract mathematical concept, and the latent space vector represents the latent features or attributes of the data.

[0072] Subsequently, the latent space vector is input into a pre-trained decoder, which decodes the input and outputs the vertex coordinates of the target face mesh.

[0073] Accordingly, during the joint training process, the first explicit parameter sample and the first implicit parameter sample can be fused to obtain the first latent space vector sample, and the first latent space vector sample can be input into the decoder to obtain the first decoding result output by the decoder.

[0074] In this embodiment of the invention, generating latent space vectors simplifies the decoder's input and improves its processing efficiency. Furthermore, utilizing a pre-trained decoder to generate the target facial mesh not only improves generation efficiency but also enhances the accuracy of the target facial mesh.

[0075] Based on the above embodiments, the initial latent parameters are extracted from the initial facial mesh based on the encoder; The encoder includes multiple groups of convolutional pooling layers and fully connected layers; Each group of convolutional pooling layers is connected sequentially and then connected to the fully connected layer. Each group of convolutional and pooling layers includes convolutional layers and pooling layers; Convolutional layers are used to apply spiral convolution kernels to extract features from the input and generate feature mesh maps; Pooling layers are used to downsample the feature grid map based on a pre-determined mapping relationship between grid maps of different finenesses, and obtain the downsampling result; Fully connected layers are used to integrate the inputs to obtain the initial implicit parameters.

[0076] Specifically, in this embodiment of the invention, the initial latent parameters can be obtained by inputting the initial facial mesh into a pre-trained encoder and extracting it through the encoder.

[0077] In image processing, convolutional neural networks (CNNs) are commonly used to extract features from two-dimensional images. A CNN is a feedforward neural network that typically includes an input layer, convolutional layers, pooling layers, fully connected layers, and an output layer. The input layer receives the two-dimensional image; the convolutional layers extract local features from the image, generating local feature maps; the pooling layers downsample the local feature maps, reducing the number of parameters and improving feature invariance; the fully connected layers perform linear combinations and nonlinear transformations on the downsampled local feature maps to obtain more representative deep features; and the output layer transforms the deep feature maps into prediction results for tasks such as classification or regression.

[0078] It should be understood that a two-dimensional image can be viewed as a rectangular matrix composed of multiple pixels. The positional relationship between pixels is fixed, and each pixel has four neighboring pixels located above, below, left, and right of that pixel. When extracting local features from a two-dimensional image, the convolutional layer in a convolutional neural network slides a convolutional kernel of a preset size across the image with a preset stride, and performs weighted summation operations within a fixed-size window to obtain and output a local feature map.

[0079] The facial mesh addressed in this embodiment of the invention is a three-dimensional mesh, composed of multiple vertices distributed in three-dimensional space. Compared with a two-dimensional image, the topological relationships between different vertices and their surrounding vertices in a facial mesh are not the same, requiring identification of these topological relationships. For example, a Laplacian matrix can be used to identify the topological relationships between all vertices. Due to the differences in data format between facial meshes and two-dimensional images, convolutional kernels used for processing two-dimensional images cannot be directly applied to processing facial meshes. A novel convolutional kernel tailored to the characteristics of facial meshes is required to perform convolution operations on them.

[0080] While facial meshes don't have the regular arrangement and positional relationships of 2D images, they typically use topologically identical 3D facial meshes when constructing different digital face models. That is, although the vertices of different facial meshes may be located in different 3D spaces, the total number of vertices and the topological relationships between them are the same. Therefore, a convolutional processing method similar to that used for 2D images can be used to extract features from facial meshes.

[0081] Based on this, the embodiment of the present invention employs a spiral convolutional neural network as the encoder. This spiral convolutional neural network refers to a convolutional neural network (CNN) in which the convolutional layers apply spiral convolutional kernels to extract features from the input.

[0082] A spiral convolutional neural network can include multiple convolutional pooling layers connected in sequence, as well as fully connected layers. Each convolutional pooling layer group includes a convolutional layer and a pooling layer. Each convolutional layer applies a spiral convolutional kernel to extract features from the input, generating a feature mesh map. Here, the number of convolutional pooling layers in the spiral convolutional neural network can be set as needed.

[0083] like Figure 3 As shown, the spiral convolution kernel is a convolution operator specifically designed for 3D meshes. Similar to a window used in convolution operations on input, the spiral convolution kernel can be a truncated spiral in shape, sampling a fixed number of vertices on its spiral for each point in the input. Further, as... Figure 4 As shown, the spiral convolution kernel can also be a dilated spiral in shape. Instead of directly sampling the continuous vertices on the spiral for each point in the input, it samples at intervals of several vertices to expand the window size.

[0084] Each pooling layer is a downsampling layer. It uses a pre-defined mapping relationship between mesh maps of different fineness to downsample the feature mesh map obtained from the connected convolutional layers, yielding the downsampled result. Here, downsampling reduces the number of parameters in the feature mesh map, improving feature invariance.

[0085] It is understandable that the process of using a spiral convolution kernel to extract features from a facial mesh is consistent with the process of using a traditional square convolution kernel to extract features from a two-dimensional image. Since a traditional convolution kernel is rectangular, the feature map obtained after convolving a two-dimensional image with a traditional kernel is also rectangular, and the result after downsampling through a pooling layer is also rectangular. In other words, the input and output of a convolutional layer using a traditional kernel are both rectangular. Therefore, for the facial mesh of the digital face model in this embodiment of the invention, the feature mesh map obtained after each convolutional layer uses a spiral convolution kernel to extract features from the input has the same topological relationship as the facial mesh.

[0086] However, when pooling layers downsample the feature mesh map, the resolution of the feature mesh map needs to be reduced to extract higher-level, more abstract downsampling results. For 3D mesh maps, reduced resolution means a reduction in the number of vertices and a change in topological relationships. Therefore, it is necessary to pre-design multiple mesh maps of varying fineness and arrange them in descending order of fineness. A mapping relationship is established between the vertices of two adjacent mesh maps, which serve as the input and output of the corresponding pooling layers, respectively.

[0087] Since the total number of vertices and the topological relationships between vertices are the same in the facial meshes of different digital human face models, multiple mesh maps of different levels of detail can be applied to the facial meshes of different digital human face models.

[0088] Preferably, an Exponential Linear Unit (ELU) activation function can be connected between the convolutional layers and pooling layers in each convolutional-pooling layer group. This activation function performs a non-linear mapping on the output of the convolutional layer to enhance the feature representation capability of the facial mesh.

[0089] The last convolutional pooling layer is connected to a fully connected layer. The fully connected layer can use the output of the last convolutional pooling layer as input to perform linear combination and nonlinear transformation and other integration operations, thereby obtaining more representative deep features as the initial latent parameters of the initial facial mesh.

[0090] By comparing existing convolutional neural networks with the spiral convolutional neural network used as an encoder in this embodiment of the invention, it can be seen that the spiral convolutional neural network used in this embodiment of the invention does not contain an output layer, and directly uses the deep features obtained from the fully connected layer as the initial latent parameters of the initial facial mesh.

[0091] In this embodiment of the invention, the extraction efficiency of initial latent parameters can be improved by using an encoder to extract them. The encoder's convolutional layers employ spiral convolution kernels to extract features from the input, ensuring accurate extraction of the feature mesh map and thus improving the accuracy of the initial latent parameters.

[0092] Based on the above embodiments, since the decoding process of the decoder is the reverse process of the encoding process of the encoder, when a spiral convolutional neural network is used as the encoder, the decoder is also implemented using a spiral convolutional neural network. The only difference between the encoder and the decoder is that the encoder includes multiple convolutional pooling layer groups, while the decoder includes multiple convolutional inverse pooling layer groups, each of which includes a convolutional layer and an inverse pooling layer.

[0093] The inverse pooling layer upsamples the input feature grid map based on a pre-determined mapping relationship between grid maps of different finenesses, increasing the resolution of the feature grid map to recover detailed features. The grid maps of different finenesses are arranged in ascending order of fineness, and a mapping relationship is established between the vertices of two adjacent grid maps, which serve as the input and output grid maps for each inverse pooling layer, respectively.

[0094] Based on the above embodiments, the initial dominant parameters of the initial facial mesh are extracted from the initial facial mesh based on the dominant parameter extractor. The initial dominant parameters are used to characterize the dominant facial features of the initial facial mesh. The initial facial mesh consists of multiple facial region meshes; the explicit parameter extractor and encoder are the region explicit parameter extractor and region encoder corresponding to each facial region mesh, respectively; The region explicit parameters of any facial region mesh are extracted from the initial facial mesh based on the corresponding region explicit parameter extractor. The latent parameters of any facial region mesh are extracted from the initial facial mesh based on the corresponding region encoder.

[0095] Specifically, the initial facial mesh also includes initial explicit parameters, which characterize the explicit facial features of the initial facial mesh. The initial explicit parameters and initial implicit parameters together characterize the facial features of the initial facial mesh. These initial explicit parameters can be extracted from the initial facial mesh using a predefined explicit parameter extractor.

[0096] The initial facial mesh can include multiple facial region meshes, each facial region mesh corresponding to a facial region, and each facial region can include the mouth, face, eyes, nose, ears, etc.

[0097] Based on this, each facial region mesh can correspond to a region explicit parameter extractor and a region encoder. The region explicit parameter extractor extracts the region explicit parameters of the corresponding facial region mesh from the initial facial mesh, and the region encoder extracts the region implicit parameters of the corresponding facial region mesh from the initial facial mesh. Here, the region explicit parameters are used to characterize the local explicit facial features of the corresponding facial region mesh, and the region implicit parameters are used to characterize the local implicit facial features of the corresponding facial region mesh.

[0098] Each region dominant parameter extractor can be used individually to extract the initial dominant parameter, in which case each region dominant parameter can be used as the initial dominant parameter alone. Alternatively, several region dominant parameter extractors can be used together to extract the initial dominant parameter, in which case the region dominant parameters extracted by several region dominant parameter extractors can be used together as the initial dominant parameter.

[0099] Similarly, each region encoder can be used individually as an explicit parameter extractor to extract the initial implicit parameters, in which case each region implicit parameter can be used as an initial implicit parameter on its own. Alternatively, several region encoders can be used together as encoders to extract the initial implicit parameters, in which case the region implicit parameters extracted by the several region encoders can be used together as the initial implicit parameters.

[0100] In this embodiment of the invention, by dividing the initial facial mesh into multiple facial region meshes, users can edit each facial region mesh individually, satisfying users' needs for editing facial details and enhancing the ability of local facial features to transfer from the initial facial mesh to the target facial mesh.

[0101] Based on the above embodiments, a latent space vector is generated based on the target explicit parameters and the initial latent parameters, including: Based on the explicit parameters of the target region and the implicit parameters of the initial region, a target latent space vector corresponding to at least one target facial region mesh in the target facial mesh is generated, and a related latent space vector is generated based on the explicit and implicit parameters of the relevant facial region meshes in the initial facial mesh. Among them, the target region explicit parameters are used to characterize the region explicit facial features of at least one target facial region grid in the target facial grid, and the initial region implicit parameters are the region implicit parameters of at least one target facial region grid in the initial facial grid.

[0102] Specifically, when the initial facial mesh is divided into multiple facial region meshes, the target visibility parameter may include one or more target region visibility parameters. Each target region visibility parameter corresponds to a target facial region mesh in the target facial mesh and is used to characterize the regional visibility facial features of the corresponding target facial region mesh in the target facial mesh. Here, the number of target region visibility parameters may include one or more, depending on the user's selection.

[0103] In other words, the user's first editing command is used to instruct the explicit parameters of each target facial region mesh in the edited target facial mesh to be the corresponding explicit parameters of the target region, while the implicit parameters of each facial region mesh in the target facial mesh remain unchanged and are the same as the initial implicit parameters of the corresponding facial region mesh in the initial facial mesh.

[0104] To facilitate user editing, for each facial region mesh, a slider can be configured on the display interface for the explicit parameters of each region of the facial region mesh. The different positions of each slider on the slider bar represent different values ​​of the explicit parameters of the corresponding region.

[0105] Furthermore, users can edit the explicit parameters of the facial region mesh by moving one or more sliders on the slider bar.

[0106] Other facial region meshes in the target facial mesh, excluding the target facial region mesh, are considered as related facial region meshes. Their region explicit parameters are the same as the region display parameters of the corresponding facial region mesh in the initial facial mesh, and their region implicit parameters are the same as the region implicit parameters of the corresponding facial region mesh in the initial facial mesh.

[0107] At this point, the target latent space vector corresponding to the target facial region mesh in each target facial region mesh can be generated by fusing the explicit parameters of the target region and the initial latent parameters of the target region mesh in each target facial mesh. Furthermore, each related latent space vector can be generated by fusing the explicit parameters of the region and the latent parameters of the region of each related facial region mesh in the target facial mesh.

[0108] Therefore, the latent space vector generated using the target explicit parameters and the initial latent parameters can include the target latent space vector corresponding to each target face region mesh in the target face mesh and the related latent space vector corresponding to each related face region mesh.

[0109] In this embodiment of the invention, when the initial facial mesh is divided into multiple facial region meshes, a specific method for calculating the latent space vector is provided, ensuring the feasibility of the solution. Furthermore, by generating corresponding latent space vectors for each facial region mesh in the target facial mesh, the detail representation of each facial region mesh is improved.

[0110] Based on the above embodiments, a target facial mesh is generated by applying a pre-trained decoder based on the latent space vectors, including: Based on the target latent space vector, the relevant latent space vector, and the region fusion parameters of the initial facial mesh, a pre-trained fusion decoder is applied to generate the target facial mesh. Among them, the region fusion parameters are related to the global positional features of each facial region mesh in the initial facial mesh.

[0111] Specifically, in the generation of the target facial mesh, one implementation scheme is to use a pre-trained fusion decoder. This involves inputting the target latent space vector, relevant latent space vectors, and the region fusion parameters of the initial facial mesh into the pre-trained fusion decoder to generate the target facial mesh. It is understood that the region fusion parameters of the initial facial mesh are semantically explicit parameters, correlated with the global positional features of each facial region mesh within the initial facial mesh. That is, the region fusion parameters of the initial facial mesh can be used to characterize the global positional features of each facial region mesh within the initial facial mesh. These global positional features represent the relative positional features between the facial region meshes. For example, if each facial region mesh includes region meshes for the facial features, then this global positional feature can represent the relative positional features of the facial feature region meshes.

[0112] The region fusion parameters of the initial facial mesh can be extracted from the initial facial mesh by a predefined first fusion parameter extractor, or they can be extracted from the initial facial mesh by a trained neural network model as the first fusion parameter extractor. No specific limitation is made here.

[0113] In this embodiment of the invention, introducing region fusion parameters of the initial facial mesh when generating the target facial mesh can ensure the overall coordination of the target facial mesh.

[0114] Based on the above embodiments, the region fusion parameters are pre-extracted from the initial facial mesh based on the first fusion parameter extractor; The fusion decoder, the region encoder and the region decoder corresponding to each facial region in the initial facial mesh are obtained by joint training based on the first facial mesh samples before and after face editing, using the first fusion parameter extractor and the region explicit parameter extractor corresponding to each facial region in the initial facial mesh.

[0115] Specifically, when the initial facial mesh is divided into multiple facial region meshes, each facial region mesh corresponds to an independent facial local feature transfer model. Each facial local feature transfer model includes a region explicit parameter extractor, a region encoder, and a region decoder corresponding to the facial region mesh. The facial feature transfer model may include the facial local feature transfer model corresponding to each facial region mesh, a first fusion parameter extractor, and a fusion decoder.

[0116] The fusion decoder, the region encoder, and the region decoder corresponding to each facial region mesh in the initial facial mesh can be jointly trained based on the first facial mesh samples before and after face editing, using the first fusion parameter extractor and the region explicit parameter extractor corresponding to each facial region mesh in the initial facial mesh. The first facial mesh samples before and after face editing can both be characterized by the coordinates of their respective vertices.

[0117] When using a pre-trained fusion decoder to generate the target facial mesh, the joint training process of the fusion decoder, the region encoders corresponding to each facial region mesh in the initial facial mesh, and the region decoders is as follows: Figure 5 As shown, it can include: using the region explicit parameter extractor corresponding to each facial region grid to extract the region explicit parameters of the first facial grid sample before facial editing, to obtain region explicit parameter samples; and using the region encoder corresponding to each facial region grid to extract the region implicit parameters of the first facial grid sample before facial editing, to obtain region implicit parameter samples.

[0118] The first fusion parameter sample is obtained by using the first fusion parameter extractor to extract fusion parameters from the first facial mesh sample before facial editing.

[0119] The explicit parameter samples and implicit parameter samples of each facial region mesh are fused to obtain the latent space vector sample of each facial region mesh. The latent space vector sample of each facial region mesh is then input into the region decoder corresponding to each facial region mesh to obtain the decoding result output by each region decoder.

[0120] The latent space vector samples of each facial region mesh and the first fusion parameter samples are input into the fusion decoder, and the fusion decoding result is obtained through the fusion decoder. This fusion decoding result is the prediction result of the first facial mesh sample after face editing, and can be characterized by the predicted values ​​of each vertex coordinate.

[0121] Subsequently, the second training loss is calculated by utilizing the differences between the decoding results output by each region decoder and the corresponding facial region meshes in the first facial mesh sample after face editing, as well as the differences between the fusion decoding results and the first facial mesh sample after face editing. The second training loss is then used to iteratively optimize the structural parameters of the fusion decoder, the region encoders and region decoders corresponding to each facial region mesh in the initial facial mesh, until the preset number of iterations is reached or the training loss converges, thus completing the training of the fusion decoder, the region encoders and region decoders corresponding to each facial region mesh in the initial facial mesh.

[0122] Understandable Figure 5 The example only shows the case where the first facial mesh sample contains three facial region meshes.

[0123] It should be noted that the logical relationships that the fusion decoder needs to learn during the joint training phase are far more complex than those of the region decoder. Correspondingly, the quality and quantity requirements for the first facial mesh samples before and after facial editing are also increased exponentially. If the training is insufficient, the trained fusion decoder may be greatly influenced by the local facial features of several facial region meshes, while neglecting global features or the local facial features of a certain facial region mesh. This can lead to the failure of individual local facial feature transfer or an overall inconsistency in the target facial mesh obtained after the transfer of local facial features.

[0124] In this embodiment of the invention, the first fusion parameter extractor and the region explicit parameter extractor corresponding to each facial region grid in the initial facial grid are applied to jointly train the fusion decoder, the region encoder and the region decoder corresponding to each facial region grid in the initial facial grid. This not only improves the training efficiency, but also enhances the collaborative working ability of the fusion decoder, the region encoder and the region decoder corresponding to each facial region grid in the initial facial grid, thereby improving the accuracy of the target facial grid.

[0125] Based on the above embodiments, the generation of the target facial mesh using a pre-trained decoder based on the latent space vectors further includes: Based on the target latent space vector, the target facial region mesh is generated by applying the region decoder corresponding to the pre-trained target facial region mesh. Based on the relevant latent space vectors, the relevant facial region mesh is generated by applying the region decoder corresponding to the pre-trained relevant facial region mesh. Based on the region fusion parameters of each facial region mesh, the target facial region mesh and related facial region meshes are transformed to obtain the transformed meshes of each region. Based on the fusion network, the transformed meshes of each region are fused to obtain the target facial mesh.

[0126] Specifically, in the process of generating the target facial mesh, besides using a fusion decoder, another approach can be adopted: introducing a fusion network to fuse the local meshes output by each region decoder. In this approach, the target latent space vector can be decoded using the pre-trained region decoder corresponding to the target facial region mesh to generate the target facial region mesh. Simultaneously, the relevant latent space vectors can be decoded using the pre-trained region decoders corresponding to related facial region meshes to generate the relevant facial region meshes.

[0127] Finally, by applying the region fusion parameters of each facial region mesh, the target facial region mesh and related facial region meshes can be transformed to obtain transformed meshes for each region. Here, each target facial region mesh and each related facial region mesh corresponds to a transformed mesh through position transformation.

[0128] It is understandable that the region fusion parameter for each facial region mesh refers to the parameter used to transform the corresponding facial region mesh so that the resulting transformed mesh conforms to the global positional characteristics of each facial region mesh. Each facial region mesh's region fusion parameter can include a displacement vector and a scaling factor. The scaling factor is used to scale the corresponding facial region mesh, and the displacement vector is used to translate the corresponding facial region mesh.

[0129] For example, the process of transforming the position of each facial region mesh using the region fusion parameters is equivalent to performing a linear transformation on the vertex coordinates of each facial region mesh, which can be represented as: s is the scaling factor, and t is the displacement vector.

[0130] The region fusion parameters for each facial region mesh can be extracted from the initial facial mesh by a predefined second fusion parameter extractor, or they can be extracted from the initial facial mesh by a trained neural network model as the second fusion parameter extractor; there are no restrictions on this.

[0131] Subsequently, the transformed meshes of each region are input into a fusion network. The fusion network then merges the transformed meshes of each region to reconstruct the target facial mesh. This fusion network can be a U-Net or a neural network with other structures.

[0132] It should be noted that, since the initial facial mesh is divided into multiple facial region meshes and each facial region mesh is independently processed for the transfer and reposition transformation of local facial features, there will be overlapping parts when the facial region meshes are fused. U-Net can fuse the overlapping parts of different facial region meshes to generate the target facial mesh.

[0133] Besides U-Net networks, the fusion network can also use the inpainting method, which specifically prevents the meshes of each facial region from overlapping. The fusion network then fills in the areas outside the facial region meshes to generate the target facial mesh.

[0134] In addition, the fusion network can also use traditional facial feature stitching algorithms to stitch together the meshes of each facial region, without making specific limitations here.

[0135] In this embodiment of the invention, the target facial region mesh and related facial region meshes are transformed by the region fusion parameters of each facial region mesh to obtain each region transformed mesh. Then, the transformation meshes of each region are fused together with the fusion network to ensure the overall coordination of the target facial mesh.

[0136] Based on the above embodiments, the region fusion parameters are pre-extracted from the initial facial mesh based on the second fusion parameter extractor; The fusion network and the region encoders and region decoders corresponding to each facial region in the initial facial mesh are obtained by jointly training the second fusion parameter extractor and the region explicit parameter extractor corresponding to each facial region in the initial facial mesh based on the second facial mesh samples before and after face editing.

[0137] Specifically, when the initial facial mesh is divided into multiple facial region meshes, each facial region mesh corresponds to an independent facial local feature transfer model. Each facial local feature transfer model includes a region explicit parameter extractor, a region encoder, and a region decoder corresponding to the facial region mesh. The facial feature transfer model may include the facial local feature transfer model corresponding to each facial region mesh, a second fusion parameter extractor, and a fusion network.

[0138] The fusion network, along with the region encoders and decoders corresponding to each facial region in the initial facial mesh, can be jointly trained using second facial mesh samples before and after face editing, applying a second fusion parameter extractor and a region explicit parameter extractor corresponding to each facial region in the initial facial mesh. Both the second facial mesh samples before and after face editing can be characterized by the coordinates of each vertex they contain.

[0139] When using a fusion network to generate the target facial mesh by combining regional decoders, the joint training process of the fusion network and the region encoders and region decoders corresponding to each facial region mesh in the initial facial mesh is as follows: Figure 6As shown, it can include: using the region explicit parameter extractor corresponding to each facial region grid to extract the region explicit parameters from the second facial grid sample before facial editing, to obtain region explicit parameter samples; and using the region encoder corresponding to each facial region grid to extract the region implicit parameters from the second facial grid sample before facial editing, to obtain region implicit parameter samples.

[0140] The second fusion parameter extractor is used to extract fusion parameters from the second facial mesh sample before facial editing, resulting in the second fusion parameter sample for each facial region mesh.

[0141] The explicit parameter samples and implicit parameter samples of each facial region mesh are fused to obtain the latent space vector sample of each facial region mesh. The latent space vector sample of each facial region mesh is then input into the region decoder corresponding to each facial region mesh to obtain the decoding result output by each region decoder.

[0142] Using the second fusion parameter sample of each facial region mesh, the decoding result output by the decoder of each region is transformed to obtain each transformation result.

[0143] The transformation results are input into a fusion network, which then fuses them to obtain the final fusion result. This fusion result is a prediction of the second facial mesh sample after face editing, and can be characterized by the predicted values ​​of each vertex coordinate.

[0144] Subsequently, the third training loss is calculated by utilizing the difference between the fusion result and the second facial mesh sample after face editing. The third training loss is then used to iteratively optimize the structural parameters of the fusion network and the region encoders and region decoders corresponding to each facial region mesh in the initial facial mesh until the preset number of iterations is reached or the training loss converges, thus completing the training of the fusion network and the region encoders and region decoders corresponding to each facial region mesh in the initial facial mesh.

[0145] Similarly, Figure 6 The example only shows the case where the second facial mesh sample contains three facial region meshes.

[0146] In this embodiment of the invention, the second fusion parameter extractor and the region explicit parameter extractor corresponding to each facial region grid in the initial facial grid are applied to jointly train the fusion network, the region encoder and region decoder corresponding to each facial region grid in the initial facial grid. This not only improves the training efficiency, but also enhances the collaborative working ability of the fusion network, the region encoder and region decoder corresponding to each facial region grid in the initial facial grid, thereby improving the accuracy of the target facial grid.

[0147] Based on the above embodiments, receiving the user's first editing instruction on the initial facial mesh of the digital human face model, prior to which includes: Receive a second editing instruction from the user on the general facial mesh; the second editing instruction includes specified facial region mesh information, which is used to identify the specified facial region mesh; Based on the explicit and implicit parameters of the specified facial region mesh, the general facial mesh is edited to obtain the initial facial mesh.

[0148] Specifically, before receiving the user's first editing instruction on the initial facial mesh of the digital human face model, a general digital human face model can be generated and displayed to the user before the user enters the facial mesh editing interface. This general digital human face model may include a general facial mesh and a face texture.

[0149] Users can edit the general facial mesh according to their needs, triggering a second editing command. This second editing command can include specified facial region mesh information, which identifies the specified facial region mesh selected by the user. The specified facial region mesh identified by this specified facial region mesh information is the facial region mesh that the user intends to include in the edited facial mesh. To facilitate user editing, multiple specified facial region networks with pre-determined corresponding explicit and implicit parameters can be configured on the display interface. Each specified facial region mesh is a pre-designed panel region template component.

[0150] Subsequently, using the explicit and implicit parameters of the specified facial region mesh identified by the specified facial region mesh information, the explicit and implicit parameters of the corresponding general facial region mesh in the general facial mesh are replaced, thus obtaining the initial facial mesh. At this point, the initial facial mesh is the facial mesh obtained by the user after the parameter replacement operation of the specified facial region mesh. Subsequent editing operations on the initial facial mesh are actually fine-tuning of the local facial features in the initial facial mesh.

[0151] In this embodiment of the invention, by specifying the explicit and implicit parameters of the facial region mesh and replacing the explicit and implicit parameters of the corresponding general facial region mesh in the general facial mesh, the complexity of the user's editing operation can be greatly reduced, and the user's experience in editing the facial mesh of the digital human face model can be improved.

[0152] Based on the above embodiments, such as Figure 7 As shown in the figure, this embodiment of the invention also provides a method for determining a digital face model, including: S21, Based on the facial mesh editing method of the digital human face model provided in the above embodiments, determine the target facial mesh; S22, Based on the target facial mesh and the target face texture, determine the target digital human face model.

[0153] Specifically, the digital face model determination method provided in this embodiment of the invention is executed by a digital face model determination device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, tablet, etc., and no specific limitation is made here.

[0154] First, step S21 is executed, which uses the facial mesh editing method of the digital human face model provided in the above embodiments to determine the target facial mesh after user editing, in conjunction with the user's first editing instruction on the initial facial mesh of the digital human face model.

[0155] Then, step S22 is executed to render the target face texture on the target face mesh, thus obtaining the target digital human face model. Here, the target face texture may include a base texture, detail maps, etc. The base texture contains skin color, basic texture, etc., and can be directly mapped onto the target face mesh to provide a basic appearance for the target face mesh.

[0156] Detail maps can include normal maps, specular maps, and displacement maps. Normal maps are used to simulate facial contours, such as wrinkles and acne, and can enhance the sense of three-dimensionality by perturbing the surface normals of the target facial mesh.

[0157] Specular mapping is used to control the intensity of reflection in skin areas, such as oily areas and wet areas.

[0158] Displacement mapping is used to directly modify the vertex positions of the target facial mesh when high precision is required.

[0159] The digital face model determination method provided in this embodiment of the invention utilizes the facial mesh editing method of the digital face model provided in the above embodiments to determine the target facial mesh, and then combines it with the target face texture to determine the target digital face model. This method can realize rapid editing of the digital face model, reduce the complexity and difficulty of operation for users, and also ensure the overall coordination of the obtained target digital face model, thereby improving user satisfaction and user experience.

[0160] Based on the above embodiments, such as Figure 8As shown in the figure, this embodiment of the invention also provides a method for determining a digital human model, including: S31, Based on the digital face model determination method provided in the above embodiments, determine the target digital face model; S32, Based on the target digital human face model and the target body model, determine the target digital human model.

[0161] Specifically, the digital human model determination method provided in this embodiment of the invention is executed by a digital human model determination device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, tablet, etc., and no specific limitation is made here.

[0162] First, step S31 is executed. Based on the digital human face model determination method provided in the above embodiments, and combined with the user's first editing instruction on the initial facial mesh of the digital human face model, the target digital human face model after user editing can be determined.

[0163] Then, step S32 is executed to stitch the target digital human face model and the target body model together to determine the target digital human model. The target body model can be pre-defined or selected by the user. During the stitching process, key points of the target digital human face model and the target body model can be aligned first. For example, corresponding key points can be marked on the connecting areas such as the neck and shoulders of the target digital human face model and the target body model. Then, an alignment tool can be used to accurately align the key points of the target digital human face model and the target body model.

[0164] Then, a transition mesh is created at the stitching position of the target digital human face model and the target body model to eliminate gaps.

[0165] In addition, a deformer modifier can be used to slightly fit the target digital human face model to the surface of the target body model, enhancing the blending effect.

[0166] The digital human model determination method provided in this embodiment of the invention first determines a target digital human face model based on the digital human face model determination method provided in the above embodiments; then, using the target digital human face model and the target body model, a target digital human model is determined. This method enables rapid editing of the digital human model, reduces the complexity and difficulty of user operations, and also ensures the overall consistency of the obtained target digital human model, improving user satisfaction and user experience.

[0167] Based on the above embodiments, such as Figure 9 As shown, this embodiment of the invention also provides a method for generating digital human videos, including: S41, Based on the digital human model determination method provided in the above embodiments, determine the target digital human model; S42, Generate a digital human video based on the target digital human model.

[0168] Specifically, the digital human video generation method provided in this embodiment of the invention is executed by a digital human video generation device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, tablet, etc., and no specific limitation is made here.

[0169] First, step S41 is executed. Based on the digital human model determination method provided in the above embodiments, and combined with the user's first editing instruction on the initial facial mesh of the digital human face model, the target digital human model after user editing can be determined.

[0170] Next, step S42 is executed to generate a digital human video using the target digital human model. For example, the target digital human model can be used to replace the human figure in an existing video to obtain a digital human video.

[0171] The digital human video generation method provided in this embodiment of the invention first determines a target digital human model based on the digital human model determination method provided in the above embodiments; then, it generates a digital human video using the target digital human model. This method enables rapid editing of the digital human image in the digital human video, reduces the complexity and difficulty of user operations, and also ensures the overall coordination of the digital human image in the digital human video, improving user satisfaction and user experience.

[0172] Based on the above embodiments, such as Figure 10 As shown, this embodiment of the invention also provides a facial mesh editing device for a digital human face model, comprising: The instruction receiving module 101 is used to receive the user's first editing instruction on the initial facial mesh of the digital human face model. The first editing instruction includes target explicit parameters, which are used to characterize the explicit facial features of the target facial mesh after the user's editing. Mesh generation module 102 is used to generate a target face mesh based on the target explicit parameters and the initial implicit parameters of the initial face mesh; the initial implicit parameters are used to characterize the implicit facial features of the initial face mesh.

[0173] Specifically, the functions of each module in the facial mesh editing device for digital human face models provided in this embodiment of the invention correspond one-to-one with the operation flow of each step in the above-described embodiments of the facial mesh editing method for digital human face models, and the achieved effects are also the same. For details, please refer to the above embodiments, and this will not be repeated in this embodiment of the invention.

[0174] Based on the above embodiments, such as Figure 11 As shown, this embodiment of the invention also provides a digital face model determination device, comprising: Mesh determination module 111 is used to determine the target facial mesh based on the facial mesh editing method of the digital human face model provided in the above embodiments; The first face model determination module 112 is used to determine the target digital human face model based on the target facial mesh and the target face texture.

[0175] Specifically, the functions of each module in the digital face model determination device provided in this embodiment of the invention correspond one-to-one with the operation flow of each step in the above-described digital face model determination method embodiment, and the achieved effects are also the same. For details, please refer to the above embodiments, and this will not be repeated in this embodiment of the invention.

[0176] Based on the above embodiments, such as Figure 12 As shown, this embodiment of the invention also provides a digital human model determination device, comprising: The second face model determination module 121 is used to determine the target digital face model based on the digital face model determination method provided in the above embodiments. The first digital human model determination module 122 is used to determine the target digital human model based on the target digital human face model and the target body model.

[0177] Specifically, the functions of each module in the digital human model determination device provided in this embodiment of the invention correspond one-to-one with the operation flow of each step in the above-described digital human model determination method embodiment, and the achieved effects are also the same. For details, please refer to the above embodiments, and this will not be repeated in this embodiment of the invention.

[0178] Based on the above embodiments, such as Figure 13 As shown, this embodiment of the invention also provides a digital human video generation device, comprising: The second digital human model determination module 131 is used to determine the target digital human model based on the digital human model determination method provided in the above embodiments. The digital human video generation module 132 is used to generate digital human videos based on the target digital human model.

[0179] Figure 14 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 14As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute the facial mesh editing method for the digital human face model, or the digital human face model determination method, or the digital human model determination method, or the digital human video generation method provided in the above embodiments.

[0180] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0181] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program that can be stored on a computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the facial mesh editing method for the digital human face model, or the digital human face model determination method, or the digital human model determination method, or the digital human video generation method provided in the above embodiments.

[0182] In another aspect, the present invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the facial mesh editing method for a digital human face model, or the digital human face model determination method, or the digital human model determination method, or the digital human video generation method provided in the above embodiments. This computer-readable storage medium can be either a non-transitory computer-readable storage medium or a transient computer-readable storage medium, and is not specifically limited herein.

[0183] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0184] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0185] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for editing facial meshes in a digital human face model, characterized in that, include: Receive a user's first editing instruction on the initial facial mesh of the digital human face model. The first editing instruction includes a target explicit parameter, which is used to characterize the explicit facial features of the target facial mesh after the user's editing. The target facial mesh is generated based on the target explicit parameters and the initial implicit parameters of the initial facial mesh; the initial implicit parameters are used to characterize the implicit facial features of the initial facial mesh.

2. The facial mesh editing method for a digital human face model according to claim 1, characterized in that, The step of generating the target facial mesh based on the target explicit parameters and the initial implicit parameters of the initial facial mesh includes: Based on the target explicit parameters and the initial implicit parameters, a latent space vector is generated; Based on the latent space vectors, a pre-trained decoder is applied to generate the target facial mesh.

3. The facial mesh editing method for a digital human face model according to claim 2, characterized in that, The initial latent parameters are extracted from the initial facial mesh based on the encoder; The encoder includes multiple groups of convolutional pooling layers and fully connected layers; Each of the convolutional pooling layers is connected sequentially and then connected to the fully connected layer. Each of the aforementioned convolutional-pooling layer groups includes a convolutional layer and a pooling layer; The convolutional layer is used to apply spiral convolution kernels to extract features from the input and generate a feature grid map; The pooling layer is used to downsample the feature grid map based on a pre-determined mapping relationship between grid maps of different finenesses, to obtain the downsampling result; The fully connected layer is used to integrate the input to obtain the initial implicit parameters.

4. The facial mesh editing method for a digital human face model according to claim 3, characterized in that, The initial dominant parameters of the initial facial mesh are extracted from the initial facial mesh based on a dominant parameter extractor. The initial dominant parameters are used to characterize the dominant facial features of the initial facial mesh. The initial facial mesh is composed of multiple facial region meshes; the explicit parameter extractor and the encoder are respectively a region explicit parameter extractor and a region encoder corresponding to each of the facial region meshes; The regional explicit parameters of any of the facial region meshes are extracted from the initial facial mesh based on the corresponding regional explicit parameter extractor; The regional latent parameters of any of the facial region meshes are extracted from the initial facial mesh based on the corresponding region encoder.

5. The facial mesh editing method for a digital human face model according to claim 4, characterized in that, The step of generating a latent space vector based on the target explicit parameters and the initial latent parameters includes: Based on the explicit parameters of the target region and the implicit parameters of the initial region, a target latent space vector corresponding to at least one target facial region mesh in the target facial mesh is generated, and a related latent space vector is generated based on the explicit and implicit parameters of the relevant facial region meshes in the initial facial mesh. The target region explicit parameter is used to characterize the regional explicit facial features of at least one target facial region mesh in the target facial mesh, and the initial region implicit parameter is the regional implicit parameter of at least one target facial region mesh in the initial facial mesh.

6. The facial mesh editing method for a digital human face model according to claim 5, characterized in that, The step of generating the target facial mesh by applying a pre-trained decoder based on the latent space vectors includes: Based on the target latent space vector, the relevant latent space vector, and the region fusion parameters of the initial facial mesh, a pre-trained fusion decoder is applied to generate the target facial mesh; The region fusion parameters are related to the global positional features of each facial region mesh in the initial facial mesh.

7. The facial mesh editing method for a digital human face model according to claim 6, characterized in that, The region fusion parameters are pre-extracted from the initial facial mesh based on a first fusion parameter extractor; The fusion decoder, the region encoder and the region decoder corresponding to each facial region mesh in the initial facial mesh are obtained by joint training based on the first facial mesh samples before and after face editing, using the first fusion parameter extractor and the region explicit parameter extractor corresponding to each facial region mesh in the initial facial mesh.

8. The facial mesh editing method for a digital human face model according to claim 5, characterized in that, The step of generating the target facial mesh by applying a pre-trained decoder based on the latent space vectors further includes: Based on the target latent space vector, the target facial region mesh is generated by applying the pre-trained region decoder corresponding to the target facial region mesh. Based on the relevant latent space vectors, the relevant facial region mesh is generated by applying the pre-trained region decoder corresponding to the relevant facial region mesh. Based on the region fusion parameters of each facial region mesh, the target facial region mesh and the related facial region mesh are transformed to obtain each region transformed mesh; Based on the fusion network, the transformation meshes of each region are fused to obtain the target facial mesh.

9. The facial mesh editing method for a digital human face model according to claim 8, characterized in that, The region fusion parameters are pre-extracted from the initial facial mesh based on a second fusion parameter extractor; The fusion network and the region encoder and region decoder corresponding to each facial region grid in the initial facial mesh are obtained by joint training based on the second facial mesh samples before and after face editing, using the second fusion parameter extractor and the region explicit parameter extractor corresponding to each facial region grid in the initial facial mesh.

10. The facial mesh editing method for a digital human face model according to any one of claims 4-9, characterized in that, The process of receiving the user's first editing instruction on the initial facial mesh of the digital human face model includes, prior to: Receive the user's second editing instruction on the general facial mesh; the second editing instruction includes specified facial region mesh information, the specified facial region mesh information being used to identify the specified facial region mesh; Based on the explicit and implicit parameters of the specified facial region mesh, the general facial mesh is edited to obtain the initial facial mesh.

11. A method for determining a digital face model, characterized in that, include: The facial mesh editing method based on the digital human face model as described in any one of claims 1-10 determines the target facial mesh; Based on the target facial mesh and the target face texture, the target digital human face model is determined.

12. A method for determining a digital human model, characterized in that, include: The target digital face model is determined based on the digital face model determination method as described in claim 11; Based on the target digital human face model and the target body model, the target digital human model is determined.

13. A method for generating digital human videos, characterized in that, include: The target digital human model is determined based on the digital human model determination method as described in claim 12; Based on the target digital human model, a digital human video is generated.

14. A facial mesh editing device for a digital human face model, characterized in that, include: The instruction receiving module is used to receive the user's first editing instruction on the initial facial mesh of the digital human face model. The first editing instruction includes target explicit parameters, which are used to characterize the explicit facial features of the target facial mesh after the user's editing. A mesh generation module is used to generate the target facial mesh based on the target explicit parameters and the initial implicit parameters of the initial facial mesh; the initial implicit parameters are used to characterize the implicit facial features of the initial facial mesh.

15. A digital face model determination device, characterized in that, include: A mesh determination module is used to determine a target facial mesh based on the facial mesh editing method of the digital human face model as described in any one of claims 1-10; The first face model determination module is used to determine the target digital human face model based on the target facial mesh and the target face texture.

16. A device for determining a digital human model, characterized in that, include: The second face model determination module is used to determine the target digital face model based on the digital face model determination method as described in claim 11. The first digital human model determination module is used to determine the target digital human model based on the target digital human face model and the target body model.

17. A digital human video generation device, characterized in that, include: The second digital human model determination module is used to determine the target digital human model based on the digital human model determination method as described in claim 12. The digital human video generation module is used to generate digital human videos based on the target digital human model.

18. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the facial mesh editing method for a digital human face model as described in any one of claims 1-10, or the digital human face model determination method as described in claim 11, or the digital human model determination method as described in claim 12, or the digital human video generation method as described in claim 13.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the facial mesh editing method for a digital human face model as described in any one of claims 1-10, the digital human face model determination method as described in claim 11, the digital human model determination method as described in claim 12, or the digital human video generation method as described in claim 13.

20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the facial mesh editing method for a digital human face model as described in any one of claims 1-10, the digital human face model determination method as described in claim 11, the digital human model determination method as described in claim 12, or the digital human video generation method as described in claim 13.

Citation Information

Patent Citations

  • Face attribute editing model training method and face attribute editing method

    CN113963409A

  • Face image editing method based on KDD-GAN

    CN116739892A

  • Speaking face video generation method, computer equipment and storage medium

    CN117789751A