Data processing method and data processing device

By introducing a micro-renderable module and a massive 2D face picture into the virtual image generation network, and updating network parameters using backpropagation algorithms, the problem of poor virtual image effects in the existing technology is solved, and a higher sense of image reality is achieved.

CN119963699APending Publication Date: 2025-05-09HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410231599.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-07
Filing Date
2024-02-29
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The virtual image generated by the existing virtual image generation scheme is poor in effect and lacks effective methods to enhance the realism of the image.

Method used

By utilizing the differentiable capabilities of the micro-renderable module, combining virtual image generation network and massive 2D face pictures, backpropagation algorithm updates, and optimizing virtual image generation network to enhance the image reality of virtual image.

Benefits of technology

The virtual image realism of virtual image generated by the network is significantly improved, and the generation effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963699A_ABST
    Figure CN119963699A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and a data processing device. The method comprises the following steps: generating three-dimensional assets of a virtual image by a virtual image generation network based on a feature guidance label, then rendering the three-dimensional assets into a two-dimensional first image of the virtual image by means of the micro-ability of a micro-renderable module, and finally, generating a second image of the virtual image by means of the three-dimensional assets. The virtual image generation network is updated by using a back propagation algorithm by using a reference image having rich two-dimensional information of a virtual image, the first image, and a micro-renderable module, and the virtual image generation network can be updated by means of the micro-ability of the micro-renderable module. The reality sense of the two-dimensional image corresponding to the three-dimensional asset of the virtual image output by the virtual image generation network is improved, so that the image effect of the generated virtual image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 7, 2023, with application number 202311473690.8 and application name “Digital Human Generation Method, Apparatus, Computing Device Cluster and Storage Medium”, all contents of which are incorporated by reference in this application. Technical Field

[0002] The embodiments of the present application relate to the field of image processing technology, and in particular, to a data processing method and a data processing device. Background Art

[0003] At present, virtual images (such as digital humans) can be applied to many fields such as television, animation, live broadcast, collaboration, exhibitions, sign language, education, customer service, advertising, finance, culture and tourism.

[0004] However, the virtual images generated by current virtual image generation schemes have poor effects. Summary of the invention

[0005] The present application provides a data processing method and a data processing device, which can make use of the differentiability of a differentiable rendering module and use a two-dimensional image with rich two-dimensional information of a virtual image as a reference image to enhance the realism of a two-dimensional image corresponding to the three-dimensional assets of a virtual image output by a virtual image generation network, so as to enhance the image effect of the generated virtual image.

[0006] In a possible implementation, an embodiment of the present application provides a data processing method. The method includes: obtaining a feature guidance label, the feature guidance label is used to describe the features of the virtual image to be generated; inputting the feature guidance label into the virtual image generation network, and the virtual image generation network performs forward calculation to output the first three-dimensional asset of the virtual image to the micro-rendering module; the micro-rendering module renders the first three-dimensional asset to output a first image, wherein the first image is a two-dimensional image of the virtual image; obtaining multiple second reference images, the multiple second reference images are reference two-dimensional images of the virtual image; based on the first image, the multiple second reference images and the micro-rendering module, back-propagating the virtual image generation network.

[0007] The second reference image may be, for example, a massive 2D face picture. The virtual image corresponding to the second reference image and the virtual image to be generated described by the feature guidance label are not necessarily the same, and the two are independent of each other.

[0008] In the related art, the virtual image generation network is optimized by only calculating the loss for the reference three-dimensional (3D) asset and the first 3D asset obtained by the virtual image generation network based on the feature guidance label, wherein the loss is a 3D loss, which will result in that the 2D loss cannot be directly transmitted back to the virtual image generation network for training optimization. To this end, the embodiment of the present application can obtain multiple two-dimensional images of the virtual image (not necessarily the same as the virtual image to be generated described by the feature guidance label) as a second reference image for calculating the 2D loss, and the second reference image has rich reference 2D information about the virtual image; and use a differentiable rendering module to render the first 3D asset generated by the virtual image generation network to obtain a 2D first image; finally, use the multiple second reference images and the first image and the differentiable rendering module to update the virtual image generation network using the back propagation algorithm. In this way, the differentiability of the differentiable rendering module can be used to transfer the 2D loss between the second reference image with rich reference 2D information of the virtual image and the first image to the virtual image generation network for network optimization, so that the trained virtual image generation network can output a 3D asset with a stronger sense of image realism of the virtual image.

[0009] In a possible implementation, the virtual image generation network is updated using a back propagation algorithm based on the first image, the multiple second reference images and the differentiable rendering module, including: performing a loss calculation based on the first image and the multiple second reference images to obtain a loss function; and the differentiable rendering module passes the loss function to the virtual image generation network to update the virtual image generation network using a back propagation algorithm.

[0010] In the embodiment of the present application, in the forward direction, the virtual image generation network can output the first 3D asset generated based on the label information to the differentiable rendering module; the differentiable rendering module can render each first 3D asset received from the virtual image generation network in the forward direction to obtain multiple two-dimensional first images in the forward direction. In the reverse direction, the loss (such as GAN loss) can be calculated based on the multiple two-dimensional first images and the multiple second reference images that provide rich 2D information of the virtual image, and the loss function can be obtained; then, the loss function can be passed to the virtual image generation network through the differentiable rendering module, so as to use the loss function to update the virtual image generation network using the back propagation algorithm. In this way, the gradient signal corresponding to the loss function can be transmitted back to the virtual image generation network via the differentiable rendering module to learn the network weights, so as to realize iterative optimization of the virtual image generation network until convergence, so that the 2D picture corresponding to the 3D asset of the virtual image generated by the trained virtual image generation network can be closer to the real 2D image taken of the virtual image. For example, if the virtual image is a human face, the 2D picture corresponding to the 3D asset of the virtual image generated by the virtual image generation network can be closer to the real human face picture, thereby improving the realism of the 2D picture.

[0011] In one possible implementation, the first three-dimensional asset includes a first position map and a first material map, and the method further includes: acquiring at least one first reference image, wherein the at least one first reference image is a two-dimensional image that provides a reference for the virtual image to be generated; converting the at least one first reference image into at least one first reference three-dimensional asset, wherein the first reference three-dimensional asset includes a first reference position map and a first reference material map, and the feature guidance label corresponds to the at least one first reference image, or the feature guidance label corresponds to the first reference position map and the first reference material map; using at least one of the first reference position maps to perform loss calculation on the first position map, and using at least one of the first reference material maps to perform loss calculation on the first material map, so as to train the virtual image generation network.

[0012] In an embodiment of the present application, a two-dimensional image of a virtual image to be generated can be converted into a first reference three-dimensional asset to serve as training data for a virtual image generation network. The source of the training data can be a two-dimensional image of a virtual image to be generated (e.g., the aforementioned massive 2D face images), which can enrich the training set of the virtual image generation network to improve the accuracy of the 3D assets of the virtual image generated by the trained virtual image generation network and optimize the training effect of the virtual image generation network. Moreover, compared with the 3D reference data of the virtual image, the acquisition cost of the 2D reference images of the virtual image is lower and the number is larger, thereby reducing the cost and difficulty of network training.

[0013] In one possible implementation, the first three-dimensional asset includes a first position map and a first material map, and the method further includes: acquiring at least one first reference three-dimensional data, wherein the at least one first reference three-dimensional data is three-dimensional data that provides reference for the virtual image to be generated; converting the at least one first reference three-dimensional data into at least one second reference three-dimensional asset, wherein the second reference three-dimensional asset includes a second reference position map and a second reference material map, and the feature guidance label corresponds to the at least one first reference image, or the feature guidance label corresponds to the second reference position map and the second reference material map; using at least one second reference position map to perform loss calculation on the first position map, and using at least one second reference material map to perform loss calculation on the first material map, so as to train the virtual image generation network.

[0014] In an embodiment of the present application, the three-dimensional data of the virtual image can be converted into a second reference three-dimensional asset to serve as training data for the virtual image generation network. The source of the training data can be the three-dimensional data of the virtual image to be generated (e.g., medium-quality 3D face data, high-quality 3D face data). The 2D image rendered with the reference three-dimensional data is closer to the real 2D image. Then, using the three-dimensional data of the virtual image to convert the second reference three-dimensional asset as training data can improve the quality of the reference three-dimensional asset in the training data, so as to enhance the training effect of the virtual image generation network, so that the image rendered by the 3D asset of the virtual image generated after training is closer to the real 2D image of the virtual image.

[0015] In a possible implementation, the virtual image generation network is connected to a super-resolution network, and the method further includes: the virtual image generation network outputs the first three-dimensional asset to the super-resolution network, wherein the resolution of the first three-dimensional asset is a first resolution; the super-resolution network performs super-resolution processing on the first three-dimensional asset to obtain a second three-dimensional asset with a second resolution, wherein the second resolution is higher than the first resolution; based on at least one of the first reference image and the third reference image, and the second three-dimensional asset, the super-resolution network is updated using a back-propagation algorithm, wherein the third reference image is a two-dimensional image obtained by rendering the first reference three-dimensional data and providing a reference for the virtual image to be generated.

[0016] In order to improve the resolution of the first three-dimensional asset of the avatar generated by the avatar generation network, the super-resolution network connected to the avatar generation network can be trained, and the training data of the super-resolution network can be at least one of the first reference image (e.g., a large number of 2D face images) and the third reference image; wherein the third reference image is a two-dimensional image rendered from the first reference three-dimensional data (e.g., at least one of medium-quality 3D face data and high-quality 3D face data), so that the trained super-resolution network can super-resolve the first three-dimensional asset generated by the avatar generation network into a three-dimensional asset with a higher resolution. The resolution after super-resolution can be close to the resolution of at least one of the first reference image and the first reference three-dimensional data used when training the avatar generation network, so that a high-quality 3D asset of the avatar can be obtained.

[0017] In a possible implementation, the method further includes: extracting a first feature describing the virtual image to be generated from the first image; extracting a second feature describing the virtual image to be generated from the feature guidance label; and calculating the loss between the first feature and the second feature to update the virtual image generation network using a back-propagation algorithm.

[0018] In this embodiment, the above-mentioned feature guidance label can be used as supervision data, and the loss between the supervision data and the two-dimensional first image rendered by the differentiable rendering module is calculated. Specifically, by extracting the features of the two (for example, converting text into vectors, converting images into vectors), and then calculating the loss between the two features, the text consistency loss can be obtained to optimize the network parameters in the virtual image generation network, so that the two-dimensional image of the virtual image corresponding to the 3D asset representation output by the optimized converged virtual image generation network can better meet the characteristics of the virtual image described by the feature guidance label, thereby meeting the needs of the user's desired virtual image.

[0019] In a possible implementation, the first reference three-dimensional data includes a first triangular mesh providing a geometric reference for the virtual image to be generated and a first material map providing a material reference for the virtual image to be generated, and the converting of the at least one first reference three-dimensional data into at least one second reference three-dimensional asset includes: converting at least one of the first triangular meshes in the at least one first reference three-dimensional data into at least one second reference position map; converting at least one of the first material maps in the at least one first reference three-dimensional data into at least one second reference material map.

[0020] In this way, when the three-dimensional data of the virtual image is used as the data source to obtain the reference three-dimensional asset, the geometric data represented by the triangular mesh in the three-dimensional data can be converted into geometric data represented by the position map, and the corresponding material map conversion (for example, at least one operation such as unifying the topology, cleaning and repairing the position map) can be performed. Among them, the geometric data represented by the triangular mesh cannot be directly used by the neural network (for example, the virtual image generation network is a neural network), and the converted position map is a two-dimensional expansion of the triangular mesh under the fixed topology. The position map stores the coordinates of each vertex in the geometric data in space. The information expressed by the triangular mesh and the position map is essentially the same, and both can accurately represent the geometric information. However, the position map is in a picture format and can be directly used by the neural network (for example, the virtual image generation network is a neural network). In this way, the present application can convert the reference three-dimensional data into a reference 3D asset that can be directly used by the virtual image generation network, so that it is convenient to use the position map to train the virtual image generation network.

[0021] In a possible implementation, before using at least one of the first reference position maps to perform loss calculation on the first position map, and using at least one of the first reference material maps to perform loss calculation on the first material map to train the virtual image generation network, before using at least one of the second reference position maps to perform loss calculation on the first position map, and using at least one of the second reference material maps to perform loss calculation on the first material map to train the virtual image generation network, the method also includes: performing a style unification operation on the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset to obtain the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset after style unification, wherein the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset after style unification have the same accessory objects described for the virtual image to be generated, or the same image parameters of the features described for the accessory objects.

[0022] In an embodiment of the present application, it is considered that there may be differences between multi-source asset data (such as the first reference image, the first reference three-dimensional data) regarding the accessory objects (such as the presence or absence of the accessory objects) of the virtual images. For example, a first reference image has eyelashes of a human face, and another first reference image does not have eyelashes of a human face. In addition, there may be differences in the image parameters (such as image parameters such as brightness, hue, and contrast) of the features described by the same accessory object between the multi-source asset data. For example, the two first reference images in the multi-source asset data both have human eyes, but the hue of the human eyes between the two first reference images is very different. Then the present application can perform a unified style operation on the reference three-dimensional assets converted from the multi-source asset data, so that the accessory objects possessed by each reference three-dimensional asset as training data are the same, for example, each reference three-dimensional asset has the outline of a human face, has human eyes, has a human mouth, and has no other accessory objects. In order to achieve the unification of the presence or absence of accessory objects. In addition, the image parameters of the same accessory object can be made the same between the reference 3D assets used as training data after the unified style, so that the image parameters of the same accessory object (such as eyes) are the same between different reference 3D assets, so that the image style of the same accessory object in different training samples is unified. In this way, the 2D images corresponding to the 3D assets generated by the virtual image generation network trained by the training data can be of unified style each time.

[0023] In one possible implementation, before using at least one of the first reference position maps to perform loss calculation on the first position map, and using at least one of the first reference material maps to perform loss calculation on the first material map to train the virtual image generation network, before using at least one of the second reference position maps to perform loss calculation on the first position map, and using at least one of the second reference material maps to perform loss calculation on the first material map to train the virtual image generation network, the method further includes: mixing at least two of the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset using mixing weights to add at least one third reference three-dimensional asset, wherein the third reference three-dimensional asset includes a third reference position map and a third reference material map.

[0024] In an embodiment of the present application, data augmentation can be performed on the training data of the virtual image generation network. Specifically, at least two reference three-dimensional assets (such as a first reference three-dimensional asset and a second reference three-dimensional asset) obtained by converting multi-source asset data are mixed using a mixing weight, wherein the mixing weight indicates that the weights of at least two reference three-dimensional assets among the at least two reference three-dimensional assets are different. Through the mixing process, at least one new third reference three-dimensional asset can be generated, which can also be used as training data for training the virtual image generation network. In this way, the data scale of the reference three-dimensional assets and their label information in the training data can be increased, so that the virtual image generation network is easier to converge after training based on rich training data, and can generate more accurate 3D assets about the virtual image.

[0025] In a possible implementation, the first reference three-dimensional data also includes a second triangular mesh that provides a geometric reference for objects other than the virtual image to be generated, and a second material map that provides a material reference for the objects other than the virtual image to be generated, and the method also includes: deleting the second triangular mesh and the second material map of the object in the first reference three-dimensional data to obtain the cleaned first reference three-dimensional data; adding a third triangular mesh and a third material map of missing attachments to the cleaned first reference three-dimensional data, wherein the missing attachments are attachments that are missing in the cleaned first reference three-dimensional data and can provide references for the virtual image to be generated; wherein the first triangular mesh includes the third triangular mesh, and the first material map includes the third material map.

[0026] In an embodiment of the present application, when converting the reference three-dimensional data of the virtual image into a reference three-dimensional asset, the three-dimensional data of objects in the reference three-dimensional data that are not related to the virtual image can be deleted to achieve the cleaning of irrelevant objects; in addition, in order to enable the reference three-dimensional assets used for training the virtual image generation network to provide the complete position map and material map of the virtual image's accessories, the three-dimensional data of the missing accessories can be added to the cleaned three-dimensional data to achieve the repair of the geometry and texture of the missing accessories, thereby obtaining a reference three-dimensional asset that can be used for the training of the virtual image generation network.

[0027] Taking the virtual image as a human face as an example, considering that the multi-source asset data may have missing parts of the 3D human face, such as the human face data in the multi-source asset data has mosaic or empty areas, etc. Then the application can perform texture repair on the missing parts of the human face area of ​​the multi-source asset data, and perform geometric cleaning on the redundant parts (such as hats and hair), so that the geometry and material mapping of the human face in the training data used to train the virtual image generation network are more complete, so that the virtual image generation network trained according to the training data can output a complete 3D asset of the 3D human face.

[0028] In a possible implementation, the converting of the at least one first reference image into at least one first reference three-dimensional asset includes: converting the at least one first reference image into second reference three-dimensional data, wherein the second reference three-dimensional data is three-dimensional data that provides a reference for the virtual image to be generated; adding third reference three-dimensional data of missing attachments to the second reference three-dimensional data, wherein the missing attachments are attachments that are missing in the second reference three-dimensional data and that can provide a reference for the virtual image to be generated; and converting the second reference three-dimensional data after adding the third reference three-dimensional data into a first reference three-dimensional asset.

[0029] In the embodiment of the present application, when converting a two-dimensional reference image of a virtual image into a reference three-dimensional asset, the reference image can be first converted into reference three-dimensional data, and then the reference three-dimensional data of the missing attachments can be added to the reference three-dimensional data, where the missing attachments are attachments that are missing in the reference three-dimensional data and can provide references for the virtual image to be generated. Thus, the three-dimensional data of the missing attachments can be added to the converted reference three-dimensional data to repair the texture of the missing attachments, thereby obtaining a reference three-dimensional asset that can be used for model training.

[0030] In a possible implementation, after the virtual image generation network is updated using a back propagation algorithm based on the first image, the multiple second reference images and the differentiable rendering module, the method further includes: receiving first user input information, wherein the first user input information is used to describe the characteristics of the target virtual image to be generated; generating, by the virtual image generation network, a third three-dimensional asset of the target virtual image based on the first user input information, the third three-dimensional asset including a third position map and a third material map; rendering the third three-dimensional asset to obtain a second image of the target virtual image, wherein the second image is a two-dimensional image of the target virtual image.

[0031] In an embodiment of the present application, the virtual image generation network trained by the above method can be used for reasoning. During the reasoning process, the first user input information of any modality (text, image, or voice, etc.) can be received, and the first user input information can describe the characteristics of the target virtual image that the user expects to generate. The trained virtual image generation network can quickly (for example, in minutes or milliseconds, etc.) generate a 3D asset of the target virtual image with high quality and semantically consistent with the first user input information based on the first user input information. For example, if the virtual image generation network is a neural network model trained based on GAN, the virtual image generation network can generate a 3D asset of the target virtual image with high quality and semantically consistent with the above-mentioned first user information within milliseconds.

[0032] In a possible implementation, the method further includes: generating geometric data that matches features of the target virtual image to be generated based on the first user input information; rendering the geometric data to obtain a texture map; and generating a third three-dimensional asset of the target virtual image based on the first user input information by the virtual image generation network, including: generating a third three-dimensional asset of the target virtual image that matches features of the target virtual image to be generated by the virtual image generation network based on the first user input information and the texture map.

[0033] The geometric data of the human head is rendered to obtain two texture maps of the geometric data of the human head in the texture space.

[0034] The geometric data may include, but is not limited to, position information of key feature points of a human head, and may also include direction information of each point of the human head.

[0035] For example, the texture map obtained by rendering the geometric data may include a landmark map of a key point position of the face and a geometry normal map of the face.

[0036] Among them, the key point position map of the face can describe the position information of the key points of the face, and the geometric normal map of the face can describe the normal direction information of each point on the entire face.

[0037] Among them, the texture map can be used as a 3D control signal (such as a ControlNet signal) to guide the virtual image generation network to generate a 3D asset of the virtual image that matches the user input information.

[0038] In an embodiment of the present application, the texture map provides 3D geometric prior information that conforms to the first user input information, which describes the geometric data of the target virtual image that matches the first user input information in the texture space. Then, the virtual image generation network can, under the guidance of the 3D control signal, combine the first user input information to generate 3D assets of the target virtual image. For example, the virtual image generation network may include a 2D Wensheng map large model (such as Stable Diffusion). The 2D Wensheng map large model has the characteristic of strong diversity of generated results. Then, combined with the above-mentioned prior 3D geometric information, the generation of robust and rich 3D virtual images can be achieved.

[0039] In a possible implementation, before rendering the third three-dimensional asset to obtain the second image of the target virtual image, the method further includes: determining, based on a preset three-dimensional asset library of the virtual image, a fourth three-dimensional asset of at least one accessory that matches the characteristics of the target virtual image, the preset three-dimensional asset library including three-dimensional data of the accessories, the three-dimensional data including geometric data and material parameters, the fourth three-dimensional asset including a fourth position map and a fourth material map; and adding the fourth three-dimensional asset of the at least one accessory to the third three-dimensional asset.

[0040] For example, if the third three-dimensional asset only includes the three-dimensional asset of the face, but does not include the three-dimensional asset of the accessories of the facial features, then in order to obtain the three-dimensional asset of the complete human head of the target virtual image, the three-dimensional data of at least one accessory that matches the characteristics of the target virtual image described by the above user input information (such as the geometric parameters (not necessarily the position map) and material parameters of the hair accessory) can be searched in the preset three-dimensional asset library, and then, based on the three-dimensional data of the at least one accessory found, the fourth three-dimensional asset of the at least one accessory that matches the characteristics described by the user input information is obtained to be added to the third three-dimensional asset generated by the virtual image generation network. In this way, the 3D asset of the complete human head of the target virtual image is obtained.

[0041] In a possible implementation, the preset three-dimensional asset library based on the virtual image determines the fourth three-dimensional asset of at least one accessory that matches the characteristics of the target virtual image, including: based on the preset three-dimensional asset library of the virtual image, determining the three-dimensional data of at least one accessory that matches the characteristics of the target virtual image; rendering the three-dimensional data of the at least one accessory by a micro-rendering module to obtain a third image of the at least one accessory; extracting a third feature describing the target virtual image from the first user input information; extracting a fourth feature describing the target virtual image from the third image; calculating the loss between the third feature and the fourth feature to optimize the material parameters in the three-dimensional data of the at least one accessory to obtain the optimized three-dimensional data of the at least one accessory; converting the optimized three-dimensional data of the at least one accessory into the third three-dimensional asset of the at least one accessory that matches the characteristics described by the first user input information.

[0042] In an embodiment of the present application, the three-dimensional data of at least one attachment matching the feature of the target virtual image described by the user input information can be first searched in the preset three-dimensional asset library. The at least one attachment is an attachment about the human face that is missing in the third three-dimensional asset, such as eyebrows, hair, and other attachments. Considering that the three-dimensional data of the attachment matching the feature in the preset three-dimensional asset library is not the best match for the feature indicated by the user input information, the three-dimensional data of the at least one attachment found can be micro-rendered (can be any micro-rendering module) to obtain a third image of the at least one attachment. Then, the loss between the feature of the third image and the feature of the user input information can be used to optimize the three-dimensional data of the at least one attachment found, and finally the optimized three-dimensional data of the at least one attachment is converted into a three-dimensional asset. The three-dimensional asset of the attachment matching the user input information can be obtained.

[0043] In a possible implementation, after the virtual image generation network generates the third three-dimensional asset of the target virtual image based on the first user input information, the method further includes: receiving second user input information, wherein the second user input information is used to provide feature editing information of the target virtual image; and optimizing the third three-dimensional asset by the virtual image generation network based on the feature difference between the second user input information and the first user input information regarding the target virtual image, to generate a fourth three-dimensional asset of the target virtual image that matches the feature editing information, wherein the fourth three-dimensional asset includes a fourth position map and a fourth material map, and the difference between the fourth three-dimensional asset and the third three-dimensional asset matches the feature difference.

[0044] The virtual image generation network of the present application can support users to input editing information twice, so that a 3D asset that is unrelated to the previously generated 3D asset will not be regenerated based on the editing information inputted twice, but the 3D asset representation will be adjusted based on the semantic difference between the two most recent user input information, so as to obtain the current 3D asset, and the difference between the two generated 3D assets matches the semantic difference between the two most recent user input information. Thus, it can be avoided that a 3D asset that is too different from the previous 3D asset (for example, a 3D asset of two completely different people) is regenerated based on the second user input information inputted later, and thus the virtual image generation network of the present application can support users to edit the generation of virtual images multiple times.

[0045] In a possible implementation, the differentiable rendering module is a differentiable rendering neural network.

[0046] Compared with the traditional differentiable renderer, the differentiable rendering module can be a neural network with differentiable rendering characteristics. The differentiable renderer composed of the neural network has good differentiability. Therefore, using the differentiable rendering neural network for training the virtual image generation network can make the virtual image generation network easier to converge, thereby improving the quality of the 3D assets of the virtual image output by the converged virtual image generation network.

[0047] In a possible implementation, the method further includes: performing super-resolution processing on the third three-dimensional asset to obtain a fifth three-dimensional asset, wherein the resolution of the fifth three-dimensional asset is higher than the resolution of the third three-dimensional asset.

[0048] In this embodiment, during the inference process, the trained super-resolution network may be used to perform super-resolution processing on the third three-dimensional asset generated by the virtual image generation network to improve the resolution of the three-dimensional asset of the obtained virtual image.

[0049] In a possible implementation, before rendering the third three-dimensional asset to obtain the second image of the target virtual image, the method further includes: calculating the loss between the first user input information and the third three-dimensional asset; and optimizing the third three-dimensional asset based on the loss.

[0050] In this embodiment, during the inference process, the loss can be calculated based on the third three-dimensional asset output by the virtual image generation network and the first user input information input by the user, and the third three-dimensional asset can be optimized by combining the loss between the network output and the text. Then, the optimized third three-dimensional asset is rendered to obtain a second image, so that the second image can better match the characteristics of the target virtual image described by the first user input information.

[0051] In a possible implementation, before rendering the third three-dimensional asset to obtain the second image of the target virtual image, the method further includes: rendering the third three-dimensional asset by the differentiable rendering module to obtain a rendering result, wherein the rendering result includes at least one of a fourth image and a normal vector of the target virtual image; determining a loss between the first user input information and the rendering result; and optimizing the third three-dimensional asset based on the loss.

[0052] In this embodiment, during the inference process, a differentiable rendering module can be used to render the third three-dimensional asset output by the virtual image generation network to obtain at least one of a fourth image and a normal vector. Then, the loss between the first user input information and the rendering result of the differentiable rendering module can be calculated, and based on the loss, the third three-dimensional asset can be optimized, so that the image quality of the second image corresponding to the optimized third three-dimensional asset can be higher.

[0053] In a possible implementation, an embodiment of the present application provides a data processing device. The device may include: a first acquisition module, used to acquire a feature guidance label, wherein the feature guidance label is used to describe the features of a virtual image to be generated; a control module, used to input the feature guidance label into a virtual image generation network; the virtual image generation network, used to perform forward calculation based on the feature guidance label to output a first three-dimensional asset of the virtual image to a differentiable rendering module; the differentiable rendering module, used to render the first three-dimensional asset to output a first image, wherein the first image is a two-dimensional image of the virtual image; a second acquisition module, used to acquire a plurality of second reference images, wherein the plurality of second reference images are reference two-dimensional images of the virtual image; a first training module, used to update the virtual image generation network using a back propagation algorithm based on the first image, the plurality of second reference images and the differentiable rendering module.

[0054] In a possible implementation, the first training module is specifically used to: perform loss calculation based on the first image and the multiple second reference images to obtain a loss function; and the differentiable rendering module passes the loss function to the virtual image generation network to update the virtual image generation network using a back-propagation algorithm.

[0055] In a possible implementation, the first three-dimensional asset includes a first position map and a first material map, and the device further includes:

[0056] A third acquisition module is used to acquire at least one first reference image, wherein the at least one first reference image is a two-dimensional image that provides a reference for the virtual image to be generated; a first conversion module is used to convert the at least one first reference image into at least one first reference three-dimensional asset, wherein the first reference three-dimensional asset includes a first reference position map and a first reference material map, and the feature guidance label corresponds to the at least one first reference image, or the feature guidance label corresponds to the first reference position map and the first reference material map; a first loss calculation module is used to use at least one first reference position map to perform loss calculation on the first position map, and use at least one first reference material map to perform loss calculation on the first material map, so as to train the virtual image generation network.

[0057] In one possible implementation, the first three-dimensional asset includes a first position map and a first material map, and the device further includes: a fourth acquisition module, used to acquire at least one first reference three-dimensional data, wherein the at least one first reference three-dimensional data is three-dimensional data that provides reference for the virtual image to be generated; a second conversion module, used to convert the at least one first reference three-dimensional data into at least one second reference three-dimensional asset, wherein the second reference three-dimensional asset includes a second reference position map and a second reference material map, and the feature guidance label corresponds to the at least one first reference image, or the feature guidance label corresponds to the second reference position map and the second reference material map; a second loss calculation module, used to perform loss calculation on the first position map using at least one second reference position map, and perform loss calculation on the first material map using at least one second reference material map, so as to train the virtual image generation network.

[0058] In a possible implementation, the device further includes a super-resolution network connected to the virtual image generation network, the virtual image generation network is further used to output the first three-dimensional asset to the super-resolution network, wherein the resolution of the first three-dimensional asset is a first resolution; the super-resolution network is used to perform super-resolution processing on the first three-dimensional asset to obtain a second three-dimensional asset with a second resolution, wherein the second resolution is higher than the first resolution; the device further includes: a second training module, used to update the super-resolution network using a back-propagation algorithm based on at least one of the first reference image and a third reference image, and the second three-dimensional asset, wherein the third reference image is a two-dimensional image obtained by rendering the first reference three-dimensional data and providing a reference for the virtual image to be generated.

[0059] In a possible embodiment, the device also includes: a first extraction module, used to extract a first feature describing the virtual image to be generated from the first image; a second extraction module, used to extract a second feature describing the virtual image to be generated from the feature guidance label; and a third loss calculation module, used to calculate the loss between the first feature and the second feature, so as to update the virtual image generation network using a back-propagation algorithm.

[0060] In a possible implementation, the first reference three-dimensional data includes a first triangular mesh providing a geometric reference for the virtual image to be generated and a first material map providing a material reference for the virtual image to be generated, and the second conversion module is specifically used to: convert at least one of the first triangular meshes in the at least one first reference three-dimensional data into at least one second reference position map; convert at least one of the first material maps in the at least one first reference three-dimensional data into at least one second reference material map.

[0061] In a possible implementation, the device further includes: a style unification processing module, configured to perform a style unification operation on the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset to obtain the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset after style unification, wherein the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset after style unification have the same accessory objects described in each of the virtual images to be generated, or the same image parameters of the features described in each of the accessory objects.

[0062] In a possible implementation, the device further includes: a blending processing module, configured to perform blending processing on at least two reference three-dimensional assets among the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset using blending weights to add at least one third reference three-dimensional asset, wherein the third reference three-dimensional asset includes a third reference position map and a third reference material map.

[0063] In a possible implementation, the first reference three-dimensional data also includes a second triangular mesh providing a geometric reference for objects other than the virtual image to be generated, and a second material map providing a material reference for the objects other than the virtual image to be generated, and the device also includes: a first cleaning module, used to delete the second triangular mesh and the second material map of the object in the first reference three-dimensional data to obtain the cleaned first reference three-dimensional data; a first repair module, used to add a third triangular mesh and a third material map of missing attachments to the cleaned first reference three-dimensional data, wherein the missing attachments are attachments that are missing in the cleaned first reference three-dimensional data and can provide references for the virtual image to be generated; wherein the first triangular mesh includes the third triangular mesh, and the first material map includes the third material map.

[0064] In a possible implementation, the first conversion module is specifically used to: convert the at least one first reference image into second reference three-dimensional data, wherein the second reference three-dimensional data is three-dimensional data that provides a reference for the virtual image to be generated; add third reference three-dimensional data of missing attachments to the second reference three-dimensional data, wherein the missing attachments are attachments that are missing in the second reference three-dimensional data and can provide a reference for the virtual image to be generated; and convert the second reference three-dimensional data after adding the third reference three-dimensional data into a first reference three-dimensional asset.

[0065] In a possible implementation, the device further includes: a first receiving module, used to receive first user input information, wherein the first user input information is used to describe the characteristics of a target virtual image to be generated; the virtual image generation network, further used to generate a third three-dimensional asset of the target virtual image based on the first user input information, wherein the third three-dimensional asset includes a third position map and a third material map; a first rendering module, used to render the third three-dimensional asset to obtain a second image of the target virtual image, wherein the second image is a two-dimensional image of the target virtual image.

[0066] In a possible implementation, the device further includes: a generation module, used to generate geometric data that matches the features of the target virtual image to be generated based on the first user input information; a second rendering module, used to render the geometric data to obtain a texture map; and the virtual image generation network, specifically used to generate a third three-dimensional asset of the target virtual image that matches the features of the target virtual image to be generated based on the first user input information and the texture map.

[0067] In a possible embodiment, the device also includes: a determination module, used to determine a fourth three-dimensional asset of at least one accessory that matches the characteristics of the target virtual image based on a preset three-dimensional asset library of the virtual image, the preset three-dimensional asset library including three-dimensional data of the accessory, the three-dimensional data including geometric data and material parameters, and the fourth three-dimensional asset including a fourth position map and a fourth material map; an accessory adding module, used to add the fourth three-dimensional asset of the at least one accessory to the third three-dimensional asset.

[0068] In a possible implementation, the determination module is specifically used to: determine the three-dimensional data of at least one accessory that matches the characteristics of the target virtual image based on a preset three-dimensional asset library of the virtual image; render the three-dimensional data of the at least one accessory through a micro-rendering module to obtain a third image of the at least one accessory; extract a third feature that describes the target virtual image from the first user input information; extract a fourth feature that describes the target virtual image from the third image; calculate the loss between the third feature and the fourth feature to optimize the material parameters in the three-dimensional data of the at least one accessory to obtain the optimized three-dimensional data of the at least one accessory; and convert the optimized three-dimensional data of the at least one accessory into a third three-dimensional asset of the at least one accessory that matches the characteristics described by the first user input information.

[0069] In a possible implementation, the device further includes: a second receiving module, used to receive second user input information, wherein the second user input information is used to provide feature editing information of the target virtual image; the virtual image generation network is also used to optimize the third three-dimensional asset based on the feature difference between the second user input information and the first user input information regarding the target virtual image, and generate a fourth three-dimensional asset of the target virtual image that matches the feature editing information, wherein the fourth three-dimensional asset includes a fourth position map and a fourth material map, and the difference between the fourth three-dimensional asset and the third three-dimensional asset matches the feature difference.

[0070] The effects of the data processing devices of the above-mentioned embodiments are similar to the effects of the data processing methods of the above-mentioned embodiments, and will not be described in detail here.

[0071] In a possible implementation, an embodiment of the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory. The processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the data processing method in the first aspect or any possible implementation of the first aspect.

[0072] The effect of the computing device cluster in this embodiment is similar to the effect of the data processing method in the above embodiments, and will not be described in detail here.

[0073] In one possible implementation, an embodiment of the present application provides a computer program product including instructions, which, when executed by a computing device cluster, enables the computing device cluster to execute the data processing method in the first aspect or any possible implementation of the first aspect.

[0074] The effect of the computer program product of this embodiment is similar to the effect of the data processing method in the above-mentioned embodiments, and will not be described in detail here.

[0075] In a possible implementation, an embodiment of the present application provides a computer-readable storage medium including computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the data processing method in any one of the above implementations.

[0076] The effect of the computer-readable storage medium of this embodiment is similar to the effect of the data processing method in the above-mentioned embodiments, and will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1a is a schematic diagram of an exemplary network training process;

[0078] Figure 1b is a schematic diagram of an exemplary network training process;

[0079] Figure 1c A schematic diagram of a data processing process is shown as an example;

[0080] Figure 1d is a schematic diagram of an exemplary network training process;

[0081] Figure 2a A schematic diagram of an exemplary application scenario;

[0082] Figure 2b A schematic diagram of a data processing process is shown as an example;

[0083] Figure 2c A schematic diagram of a data processing process is shown as an example;

[0084] Figure 2d A schematic diagram of a data processing process is shown as an example;

[0085] Figure 2e A schematic diagram of a data processing process is shown as an example;

[0086] Figure 2fA schematic diagram of a data processing process is shown as an example;

[0087] Figure 3 is a schematic structural diagram of a data processing device shown as an example;

[0088] Figure 4 A schematic diagram of the structure of a computing device is shown as an example;

[0089] Figure 5 A schematic diagram of the structure of a computing device is shown as an example;

[0090] Figure 6 The figure is a schematic diagram of the structure of a computing device cluster shown as an example. DETAILED DESCRIPTION

[0091] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0092] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.

[0093] The terms "first" and "second" in the description and claims of the embodiments of the present application are used to distinguish different objects rather than to describe a specific order of objects. For example, a first target object and a second target object are used to distinguish different target objects rather than to describe a specific order of target objects.

[0094] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0095] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" refers to two or more than two. For example, multiple processing units refer to two or more processing units; multiple systems refer to two or more systems.

[0096] Digital humans can be applied to television, animation, live broadcast, collaboration, exhibitions, sign language, education, customer service, advertising, finance, culture and tourism, and many other fields.

[0097] Regarding the generation scheme of digital human, there are two technical schemes in the prior art:

[0098] Prior art 1 mainly uses professional art technicians to manually create geometry and materials of computer graphics (CG) digital humans using professional software to generate traditional three-dimensional (3D) asset representations (specifically geometry and material mapping) to generate 3D digital humans. In this solution, the shape structure of the 3D digital human can be expressed by geometry, and the skin texture of the 3D digital human can be expressed by material mapping.

[0099] Existing rendering scenes such as games and videos generally use CG pipelines to render the geometry, lighting and other characteristics of objects represented by 3D assets. Therefore, the traditional 3D assets (a type of display expression) generated by prior art 1 and represented by geometry and material maps are compatible with the existing CG pipeline, so that the geometry and material maps can be used through the CG pipeline to render various characteristics of 3D digital humans, thereby obtaining a 3D digital human that meets the characteristic requirements.

[0100] However, the digital human generation solution of the prior art 1 needs to rely on professional equipment to collect 2D image data each time a 3D digital human is generated, and professional technicians need to manually create the geometry and material of the CG digital human based on the 2D image data, which makes the generation cost of the 3D digital human relatively high; in addition, the solution also needs to rely on professional artists to use professional software for manual operation, which makes the digital human generation process have a low degree of automation, time-consuming and labor-intensive.

[0101] Prior art 2 uses computer vision technology to encode the 3D model of the digital human into a specific neural representation through a neural network. The generation process of the neural representation is a black box operation. This makes the specific neural representation different from the 3D asset representation (geometry and material) of the display expression in prior art 1, but an implicit expression related to the neural network.

[0102] Therefore, the 3D digital human with specific neural representation generated by the prior art 2 does not have the characteristics of geometry and material, and thus cannot be directly connected to the existing CG pipeline for rendering of various characteristics. However, current rendering scenes such as games and videos all use CG pipelines for characteristic rendering, which limits the application scenarios of the 3D digital human with specific neural representation generated by the prior art 2. Similarly, since the 3D digital human with specific neural representation generated by the prior art 2 is not compatible with existing rendering pipelines such as CG pipelines, the specific neural representation generated by it cannot be edited for secondary lighting, mapping, etc.

[0103] To this end, the present application provides a 3D digital human face generation method (an example of a data processing method) and a 3D digital human face generation system, as well as a training method (an example of a data processing method) of a neural network (e.g., a 3D asset generation network of a digital human) in the face generation system to obtain the 3D digital human face generation system. The face generation system may include a 3D asset generation network for a digital human. The face generation method may generate a 3D asset representation (geometry map and material map) of the digital human that conforms to the user input information (at least one piece of information such as text or image) for instructing the generation of the digital human face through the 3D asset generation network for the digital human, without the need for manual creation of geometry and material, thereby reducing the generation cost of the CG digital human and improving the generation efficiency of the CG digital human; and the 3D asset representation generated by the face generation method is a geometry map and a material map, so that it can be compatible with the traditional CG rendering pipeline, so as to ensure that the face generation method of the present application has a wider range of application scenarios; and the 3D assets of the digital human represented by the geometry data and the material map can be decoupled from each other in the rendering process of various characteristics such as color, geometry, and lighting, so as to facilitate the subsequent editing of the digital human represented by the 3D asset (for example, editing of characteristics such as color, lighting, and geometry).

[0104] In some embodiments, the above-mentioned various methods provided in the present application are not limited to generating 3D digital humans, but can also generate other non-human virtual images, such as animals, plants, cartoon characters and other arbitrary virtual images. The network training process and network reasoning process of this process are similar to the principles of the process of generating 3D digital humans, so they will not be described one by one.

[0105] In addition, the virtual image generated by the method of the present application is not limited to generating a face of the virtual image, but can also generate a body and other parts of the virtual image to generate a complete virtual image.

[0106] The method of the present application is described below using the generation of a 3D digital human face as an example. When generating other virtual images, the principles of the implementation method are similar and will not be repeated here.

[0107] Figure 1a to Figure 1d A schematic diagram of the training process of the neural network in the face generation system 100 is shown to obtain the face generation system 100 to ensure the implementation of the face generation method.

[0108] Figure 2a to Figure 2e A schematic diagram of the use process of the face generation system 100 is shown to implement the face generation method, thereby generating a 3D digital human.

[0109] The following first introduces the training process of the neural network in the face generation system 100 of the present application with examples.

[0110] like Figure 2a As shown, the face generation system 100 may include a 3D asset generation network 101 for a digital human (which may be extended to a virtual image generation network). Figure 1a The training process of the 3D asset generation network 101 is described below.

[0111] Example 1

[0112] like Figure 1a As shown, the training process may include the following steps:

[0113] S101, obtaining multi-source asset data of digital humans.

[0114] The multi-source asset data may include, but is not limited to, at least one of at least one first reference image and at least one first reference three-dimensional data.

[0115] The first reference image is a two-dimensional image that provides a reference for the virtual image to be generated; the first reference three-dimensional data is three-dimensional data that provides a reference for the virtual image to be generated.

[0116] In some embodiments, Figure 1a As shown, the multi-source asset data may include but is not limited to at least one of the following three types of asset data: a massive amount of 2D face images (or texts), a large amount of medium-quality 3D face data, and a small amount of high-quality 3D face data.

[0117] Among them, the massive 2D face pictures are examples of multiple first reference images; the above-mentioned large amount of medium-quality 3D face data and a small amount of high-quality 3D face data are examples of first reference three-dimensional data respectively.

[0118] The 3D face data may include geometric data and material maps of the face. In addition, the 3D face data may be in the form of text or at least one of a picture, which is not limited here.

[0119] Since the acquisition costs of high-quality 3D face data, medium-quality 3D face data, and 2D face images are from high to low, in order to reduce the training cost of the 3D asset generation network 101, a small amount of high-quality 3D face data, a large amount of medium-quality 3D face data, and a large amount of 2D face images can be obtained.

[0120] Regarding the quality of 3D face data, the evaluation criteria can be, but are not limited to, the following: 2D face images obtained by rendering 3D face data under multiple variable illuminations and viewing angles are compared with 2D face images collected (e.g., taken by a camera) under corresponding illuminations and viewing angles to obtain multiple comparison groups. The two 2D face images in each comparison group are images under the same illumination and the same viewing angle. If the multiple comparison groups corresponding to a 3D face data are very similar, it means that the quality of the 3D face data is high, otherwise it means that the quality of the 3D face data is low.

[0121] The quality of the 2D face image is similar to the aforementioned evaluation criteria. The closer the 2D face image is to a real 2D face image taken under the same lighting and viewing angle, the higher the image quality of the 2D face image.

[0122] In this embodiment, the object in the multi-source asset data is a human face. Taking the human face including the geometric shape of the human face and the skin of the human face as an example, the high-quality 3D face data may include the geometric data and material map of the complete 3D face (there are no other objects other than the human face); and the medium-quality 3D face data may further include the geometric data and material map of objects other than the 3D face (such as a hat). For example, the 3D face data includes the geometric data of a portion of the 3D face and the hat, and the material map of the portion of the 3D face and the hat.

[0123] In this way, the purpose of training the 3D asset generation network 101 is to generate a position map and a material map of a complete 3D face through the 3D asset generation network 101 .

[0124] Of course, in other embodiments, the object of each asset data in the above-mentioned multi-source asset data is not limited to a human face, but may be a human head. The human head then includes not only a human face object, but also objects such as facial features (such as hair, mouth, eyes, eyebrows, eyelashes, nose, etc.).

[0125] Alternatively, in some embodiments, the object of each asset data in the multi-source asset data may also be a human body, that is, the object includes not only a human head, but also a torso, limbs and other objects.

[0126] Alternatively, in some embodiments, the object of each asset data in the multi-source asset data may also be a virtual image such as an object, a cartoon character, or a plant.

[0127] Alternatively, in some embodiments, the purpose of training the 3D asset generation network 101 is to generate a 3D asset representation of a certain type of object other than the human body, such as a table, a chair, etc., through the 3D asset generation network 101. Then the multi-source asset data may also include three types of asset data of the corresponding object (e.g., a table), and the implementation principle is the same, which will not be repeated here. It should be noted that since the multi-source asset data needs to be geometrically unified in topology during subsequent preprocessing, so that the trained 3D asset generation network 101 can generate a 3D asset representation of the corresponding object (e.g., a table) under the same topology each time, the multi-source asset data can be asset data of the same type of object (e.g., a class of tables with similar shapes and structures).

[0128] This application does not impose any restrictions on the method of obtaining massive 2D face images, a large amount of medium-quality 3D face data, and a small amount of high-quality 3D face data, which can be achieved through any existing or future developed method.

[0129] For example, a large number of 2D face images can come from the 2D face dataset FFHQ, the Celeb dataset, 2D face images from the Internet, etc. When obtaining a large amount of medium-quality 3D face data, 3D face data can be obtained by performing 3D face reconstruction on 2D face images. When obtaining a small amount of high-quality 3D face data, professional equipment can be used to collect it.

[0130] In this embodiment, at least one of the three types of asset data is used to generate a reference three-dimensional asset for training the 3D asset generation network 101 because high-quality 3D face data itself has less data and a high acquisition cost. If a small amount of the high-quality 3D face data is used to generate a reference three-dimensional asset to train the 3D asset generation network 101, the learning effect of the 3D asset generation network 101 will be poor, resulting in a low accuracy of the 3D asset representation of the face generated by the trained 3D asset generation network 101. If medium-quality 3D face data is used to generate a reference three-dimensional asset to train the 3D asset generation network 101, more 3D face data can be used as a reference three-dimensional asset while ensuring a low cost, so as to improve the learning effect of the 3D asset generation network 101. In addition, the 2D data of the face in the 2D face picture (or text) is relatively accurate. If the 2D face picture is used for the subsequent training of the 3D asset generation network 101, the quality of the 3D face corresponding to the generated face 3D asset can be higher. In this way, by collecting multi-source asset data to obtain reference three-dimensional assets and enriching training data, the 3D asset generation network 101 can be trained to improve the learning effect of the 3D asset generation network 101, so that the 3D asset representation of the face generated by the trained 3D asset generation network 101 is more accurate. With the richness of multi-source asset data, it can be used to generate training data, so that the trained 3D asset generation network 101 can generate a creative 3D asset representation of a digital person that does not exist in real life.

[0131] S102: pre-process the multi-source asset data to obtain training data.

[0132] Since there are differences in topology and style between multi-source asset data, in some embodiments, the multi-source asset data can be pre-processed with at least one of unified topology and unified style; in addition, in some embodiments, in order to increase the number of reference three-dimensional assets in the training data as supervisory data to improve the learning effect of the 3D asset generation network 101, the multi-source asset data can also be pre-processed with data augmentation (i.e., on this basis, the number of reference three-dimensional assets is expanded) to obtain reference three-dimensional assets and their label information.

[0133] The reference three-dimensional assets are 3D asset representations of digital humans, specifically position maps and material maps.

[0134] In some embodiments, after obtaining the reference 3D asset, label information (an example of a feature guidance label) may be added to the reference 3D asset (also referred to as a reference position map and a reference material map) to obtain complete training data. The feature guidance label is used to describe the features of the virtual image to be generated. The label information is used to describe the facial features of the 3D digital human to be generated.

[0135] In some embodiments, for the label information of the 3D asset generation network 101 used for training digital humans in the present application, the multi-source asset data can be labeled separately before converting the multi-source asset data into reference three-dimensional assets, so that the converted reference three-dimensional assets also have corresponding labels; or, after converting the multi-source asset data into reference three-dimensional assets, each reference three-dimensional asset can be labeled for the training of the 3D asset generation network 101 of digital humans.

[0136] Among them, the geometric data in the multi-source asset data is represented by a triangular mesh, but the triangular mesh cannot be directly used by the neural network (such as the 3D asset generation network 101). Therefore, when unifying the topology, the geometric data corresponding to each of the multi-source asset data can be converted from the triangular mesh to a position map so that it can be directly used by the 3D asset generation network 101. In this way, the 3D asset representation of the digital human output by the trained 3D asset generation network 101 also uses a position map to represent the geometric data of the digital human. In this way, the trained 3D asset generation network 101 can output the 3D asset representation of the digital human in the form of a picture.

[0137] The specific implementation process of the preprocessing will be described in detail in Example 3 below and will not be elaborated here.

[0138] S103 , input the tag information to the 3D asset generation network 101 of the digital human, and the 3D asset generation network 101 performs forward calculation to generate and output a 3D asset representation to the micro-rendering network 102 .

[0139] The network structure of the 3D asset generation network 101 of the digital human can be any neural network currently available or developed in the future, such as a generative adversarial network (GAN) or Stable Diffusion, etc. Stable Diffusion is a text-to-image generation model based on latent diffusion models, which can generate high-quality, high-resolution, and highly realistic images based on any text input, and there is no limitation here.

[0140] Optionally, S105 , the 3D asset generation network 101 is optimized according to the reference three-dimensional asset with the label information in the training data and the 3D asset representation obtained in S103 .

[0141] When training and optimizing the network parameters in the 3D asset generation network 101, the network parameters in the 3D asset generation network 101 can be optimized by calculating the loss between the 3D asset representation obtained in S103 and the reference three-dimensional asset, and the optimization goal is to minimize the loss to train the 3D asset generation network 101.

[0142] Specifically, the reference 3D asset with label information may include a reference position map and a reference material map; and the 3D asset representation obtained through S103 is the predicted position map and predicted material map that match the label information and are predicted by the 3D asset generation network 101. When calculating the loss, the loss between the reference position map and the predicted position map corresponding to the same label may be calculated to train the 3D asset generation network 101 (e.g., back propagation); and the loss between the reference material map and the predicted material map corresponding to the same label may be calculated to train the 3D asset generation network 101 (e.g., back propagation).

[0143] However, since the geometric data corresponding to the multi-source asset data is represented by a triangular mesh, which is non-structural information, its information expression in 3D space is a sequence and irregular data, and it contains non-differentiable information. Even if the triangular mesh is converted into a position map, it still contains the non-differentiable information. This will result in poor differential properties of the training data when only the position map and material map represented by the 3D asset are used as supervisory data to train the 3D asset generation network 101, making it difficult for the 3D asset generation network 101 to converge.

[0144] To this end, when training the 3D asset generation network 101, the present application adds a differentiable rendering network 102 with good differentiability. For example, the differentiable rendering network 102 is used to train and optimize the 3D asset generation network 101 through the following S104, so that the 3D asset generation network 101 is easy to converge, thereby improving the image quality of the 3D asset representation output by the 3D asset generation network 101 after training.

[0145] In some embodiments, Figure 1a As shown, the reference three-dimensional assets (reference position map and reference material map) and their label information can be input into the 3D asset generation network 101 for training, so that the trained 3D asset generation network 101 can generate a 3D asset representation (position map and material map) of the face of the digital human to be generated as described by the above label information.

[0146] S104 , the differentiable rendering network 102 combines a large amount of 2D face images and the 3D asset representation from the 3D asset generation network 101 to optimize the 3D asset generation network 101 .

[0147] The micro-renderable network 102 can render the 3D asset representation from the 3D asset generation network 101 to obtain a 2D face image of the 3D digital human to be generated, thereby rendering the 3D data into 2D data. Then, the present application can use a large number of 2D face images (such as the 2D face images mentioned in the above-mentioned multi-source asset data, but not limited to those from the multi-source asset data) as the second reference image to determine the loss between the 2D face image rendered by the micro-renderable network 102 and the second reference image, and optimize the network parameters in the 3D asset generation network 101 in combination with the loss until the 3D asset generation network 101 converges. The face corresponding to the 2D face image as the second reference image does not necessarily match the face described by the above-mentioned label information. The second reference image can be a two-dimensional image of any face (which can be expanded to a virtual image) that is unrelated to the label information.

[0148] In the embodiment of the present application, the network 102 for differentiable rendering provided by the present application is a differentiable renderer for human faces. The network 102 for differentiable rendering is a pre-trained neural network (the training process can refer to the introduction of Example 2 below), so that it has a high differentiable property. Then, the 3D asset generation network 101 of the present application is trained and optimized by the differentiable rendering network 102 with a high differentiable property, so that the 3D asset generation network 101 can be easily converged to improve the image realism of the 3D asset representation outputted by it.

[0149] In addition, by rendering the 3D asset representation output by the 3D asset generation network 101 through the micro-rendering network 102, a high-quality 2D face image can be obtained, so that the 2D face image is closer to the 2D face image corresponding to the aforementioned 3D asset representation; then, the 2D face image obtained by the rendering and the second reference image (such as a large number of real 2D face images) as supervision data and the micro-rendering network 102 are used to optimize the 3D asset generation network 101. Among them, the real massive 2D face images as supervision data can have richer supervision signals of 2D faces, so that the 3D asset generation network 101 can be trained and optimized in the 2D dimension of the face. Thereby, the 2D image corresponding to the 3D asset of the face output by the optimized 3D asset generation network can be made more realistic.

[0150] In an embodiment of the present application, the training data used to train the 3D asset generation network 101 may include reference images (3D assets, such as reference position maps and reference material maps) obtained by converting multi-source asset data, so that the training data is relatively rich and the image quality is relatively high, thereby improving the image quality of the 3D asset representation of the 3D face output by the trained 3D asset generation network 101; in addition, when training the 3D asset generation network 101, since the reference position map representing geometric information has non-differentiable information, then using the differentiable rendering network 102 with good differentiability to train and optimize the 3D asset generation network 101 can also accelerate network convergence, thereby improving the image quality of the 3D asset representation output by it; finally, when using the differentiable rendering network 102 to train the 3D asset generation network 101, the supervision data can be a large amount of high-quality 2D face images, thereby also improving the realism of the face images output by the 3D asset generation network 101. Based on the above, the 3D asset generation network 101 trained in the present application can output a 3D asset representation of a photo-realistic and high-quality 3D face.

[0151] It should be understood that in this example, the differentiable rendering network 102 is taken as a differentiable renderer for a face. In other embodiments, the differentiable rendering network 102 can also be a differentiable rendering network for other required objects, such as a head, a human body, an object, etc., to render the 3D asset representation of the corresponding object into a high-quality 2D image. The principle is the same and will not be repeated here.

[0152] In addition, the present application does not limit the execution order between the above S104 and S105.

[0153] In some scenarios, when the rendering quality of the 3D digital human image is not required to be high due to factors such as low device performance, the differentiable rendering network 102 can also be replaced by a traditional differentiable renderer, such as a rasterization-based differentiable renderer or a physically-based differentiable renderer, for training the 3D asset generation network 101.

[0154] Example 2

[0155] Combine the following Figure 1b , to illustrate the training process of the differentiable rendering network 102 in Example 1.

[0156] Before explaining the training process, let's first introduce the data preparation process:

[0157] The training data for the differentiable rendering network 102 may include: Figure 1b The 3D face position map and the 3D face material map shown. Optionally, the 3D face data may also include Neural.

[0158] Among them, Figure 1b The 3D face position map and the 3D face material map shown can be obtained by converting 3D data from multi-source asset data, or can be 3D asset data converted from high-quality 2D images from multi-source asset data, or can be the highest quality 3D asset data from the aforementioned two sources, and there is no limitation here.

[0159] In addition, the present application can also obtain supervision data, which can be a 2D face image, as a reference image for the above training data.

[0160] For example, the present application can render the above high-quality 3D face data through an offline renderer to obtain a 2D face image as supervision data. In addition, the present application can also use professional camera equipment such as a light cage to shoot the face to obtain a 2D face image as supervision data.

[0161] The supervised data used for training the differentiable rendering network 102 of the present application is high-quality 2D face images.

[0162] In addition, the present application may also perform random sampling of the camera angle of view within a given range of the camera angle of view to obtain camera angle of view parameters, and perform random sampling of the light source within a given range of the light source to obtain light source parameters.

[0163] Below is the data prepared above, refer to Figure 1b , to introduce the training process of the differentiable rendering network 102.

[0164] like Figure 1b As shown, the process may include the following steps:

[0165] S1051, under the camera viewing angle parameters, rendering is performed based on the position map of the 3D face in the training data and the material map of the 3D face to obtain a screen space map.

[0166] The camera viewing angle parameter is the randomly sampled camera viewing angle parameter.

[0167] This application does not limit the specific implementation process of rendering the training data represented by the 3D asset to obtain the screen space map.

[0168] S1052, respectively encode the camera viewing angle parameters and the light source parameters to obtain encoded camera viewing angle parameters and encoded light source parameters.

[0169] The camera viewing angle parameter is the same as the camera viewing angle parameter in S1051.

[0170] For example, the camera viewing angle parameters may be spherical harmonic encoded to obtain encoded camera viewing angle parameters.

[0171] For example, the light source parameters may be light source encoded to obtain encoded light source parameters.

[0172] Among them, the present application does not limit the execution order between S1051 and S1052, which can be executed in parallel or serially and can be flexibly set according to needs.

[0173] S1053, the differentiable rendering network 102 performs differentiable rendering on the screen space map based on the encoded camera view parameters and light source parameters to obtain a 2D face image.

[0174] S1054, performing loss calculation based on the supervision data and the 2D face image output by the differentiable rendering network 102, and optimizing the network parameters of the differentiable rendering network 102 based on the loss until convergence.

[0175] The present application does not limit the network structure of the micro-renderable network 102, which may be any neural network structure.

[0176] The traditional rasterization-based differentiable renderer has high rendering efficiency but low rendering quality, while the traditional physically-based differentiable renderer has high rendering quality but large gradient variance and low rendering efficiency. This makes the differentiability of the renderer in the prior art poor, and when used to train the 3D asset generation network 101, the convergence effect is not optimal.

[0177] In this embodiment, when training the differentiable rendering network 102, the supervisory data used is a high-quality 2D face picture, such as a 2D face picture obtained by rendering high-quality 3D face data, or a 2D face picture obtained by taking a picture with a professional camera such as a light cage. Then, the differentiable rendering network 102 trained with this supervisory data can convert the 3D asset representation (position map and material map) of the face into a high-quality 2D face picture, so that the image rendering quality of the trained differentiable rendering network 102 of the present application is high and has a high differentiable property. Then, when the trained differentiable rendering network 102 is applied to the 3D asset generation network 101 (for example Figure 1a When the 3D asset generation network 101 shown in the figure is optimized, the convergence speed of the 3D asset generation network 101 can be improved; compared with using a traditional differentiable renderer to optimize the 3D asset generation network 101, the present application uses a differentiable rendering network 102 to optimize the 3D asset generation network 101, which can make the convergence effect of the 3D asset generation network 101 better.

[0178] Example 3

[0179] Convert multi-source asset data into reference 3D assets through pre-processing.

[0180] In some embodiments, at least one first reference image may be converted into at least one first reference three-dimensional asset, wherein the first reference three-dimensional asset includes a first reference position map and a first reference material map.

[0181] Wherein, the at least one first reference image is a two-dimensional image providing a reference for the virtual image to be generated;

[0182] Taking the virtual image as a digital human's face as an example, the first reference image may be a massive 2D face picture in multi-source asset data.

[0183] The first reference three-dimensional asset includes a first reference position map and a first reference material map.

[0184] In the embodiment of the present application, the two-dimensional image of the virtual image to be generated can be converted into a first reference three-dimensional asset to serve as training data for the 3D asset generation network 101. The source of the training data can be the two-dimensional image of the virtual image to be generated (such as the aforementioned massive 2D face images).

[0185] In some embodiments, the at least one first reference three-dimensional data may be converted into at least one second reference three-dimensional asset, wherein the second reference three-dimensional asset includes a second reference position map and a second reference material map.

[0186] The at least one first reference three-dimensional data is three-dimensional data providing reference for the virtual image to be generated.

[0187] Taking the face of a digital human as an example, the first reference three-dimensional data may be at least one of a large amount of medium-quality 3D face data and a small amount of high-quality 3D face data in multi-source asset data.

[0188] In the embodiment of the present application, the three-dimensional data of the virtual image can be converted into a second reference three-dimensional asset to serve as training data for the 3D asset generation network 101. The source of the training data can be the three-dimensional data of the virtual image to be generated (such as the above-mentioned medium-quality 3D face data and high-quality 3D face data).

[0189] In some embodiments, the first reference three-dimensional data may include a first triangular mesh providing a geometric reference for the virtual image to be generated and a first material map providing a material reference for the virtual image to be generated. Then, when converting the at least one first reference three-dimensional data into at least one second reference three-dimensional asset, at least one of the first triangular meshes in the at least one first reference three-dimensional data can be converted into at least one second reference position map; and at least one of the first material maps in the at least one first reference three-dimensional data can be converted into at least one second reference material map.

[0190] In this way, when the 3D data of the virtual image is used as the data source to obtain the reference 3D asset, the geometric data represented by the triangle mesh in the 3D data can be converted into geometric data represented by the position map, and the material map can be converted accordingly (for example, at least one operation such as unifying the topology, cleaning and repairing the position map), so that the reference 3D asset obtained after processing can be directly used for the training of the 3D asset generation network 101.

[0191] The first reference three-dimensional data may also include a second triangular mesh that provides a geometric reference for objects other than the virtual image to be generated, and a second material map that provides a material reference for the objects other than the virtual image to be generated, then cleaning and repair operations can be performed. Wherein, in method 1, the cleaning and repair operations can be performed before the first reference three-dimensional data is converted into a second reference three-dimensional asset (then the first reference three-dimensional data is cleaned and repaired); or, in method 2, the cleaning and repair operations can also be performed after the first reference three-dimensional data is converted into a reference three-dimensional asset. For example, the reference three-dimensional asset obtained by conversion regarding the object other than the virtual image to be generated is cleaned and repaired.

[0192] In method 1, when cleaning and repairing the first reference three-dimensional data, the second triangular mesh and the second material map of the object (e.g., the hat worn on a person's head) in the first reference three-dimensional data may be deleted to obtain the cleaned first reference three-dimensional data; a third triangular mesh and a third material map of missing attachments (after the hat is cleaned, the geometry and material data of the part of the face covered by the hat are missing) are added to the cleaned first reference three-dimensional data, wherein the missing attachments are attachments that are missing in the cleaned first reference three-dimensional data and can provide references for the virtual image to be generated;

[0193] The first triangular mesh includes the third triangular mesh, and the first texture map includes the third texture map.

[0194] Method 2 will be Figure 1c is described in the embodiments (e.g., S1022 and S1023 described below).

[0195] In an embodiment of the present application, when converting the reference three-dimensional data of a virtual image into a reference three-dimensional asset, the three-dimensional data of objects in the reference three-dimensional data that are not related to the virtual image can be deleted to achieve the cleaning of irrelevant objects; in addition, in order to enable the reference three-dimensional asset used for model training to provide the complete position map and material map of the virtual image's accessories, the three-dimensional data of the missing accessories can be added to the cleaned three-dimensional data to achieve the repair of the geometry and texture of the missing accessories, thereby obtaining a reference three-dimensional asset (here is the second reference three-dimensional asset) that can be used for model training.

[0196] The following is combined with Figure 1a , refer to Figure 1c , Figure 1c A schematic diagram showing an implementation process of preprocessing multi-source asset data.

[0197] like Figure 1c As shown, the process may include the following steps:

[0198] S1011, performing perspective augmentation on a large number of 2D face images to obtain multi-perspective 2D face images.

[0199] For example, a neural network model (not limited) can be used to augment the perspective of each 2D face image in a large number of 2D face images to obtain 2D face images of various perspectives. For example, 2D face images of various perspectives of a face may include but are not limited to: a front view of a face, a side view of a face, etc.

[0200] S1012, mapping the multi-view 2D face image and the 3D model to obtain the geometry (represented by a triangular mesh) and color map of the 3D face.

[0201] For example, the 3D model can be a unified topology model pre-established for the face object in the present application. In this case, for the position map and material map of the 3D face obtained by converting massive 2D face images obtained in S1015, there is no need to perform the unified topology operation of the following S1041.

[0202] On the contrary, when the 3D model used for modeling in S1012 is not the unified topology model of the present application, it is necessary to perform the unified topology operation of S1041 described below on the 3D asset representation obtained in S1015 (specifically, the position map and material map of the 3D face).

[0203] In S1012, coordinate mapping can be established between the 2D face images of various viewing angles and the preset 3D face model, thereby obtaining 3D assets, which are the triangle mesh and color map of the 3D face.

[0204] S1013, UV unfolding the color map of the 3D face to obtain a 2D color map.

[0205] The obtained 3D face color map may be UV unfolded to unfold the 3D color map into a plane to obtain a 2D color map.

[0206] Through the above S1011, S1012, and S1013, each 2D face image in a large number of 2D face images can be converted into three-dimensional data (such as the second reference three-dimensional data) as a face reference, which may specifically include geometric data represented by a triangular mesh of the 3D face and a color map of the 3D face.

[0207] S1014, performing texture repair on the 2D color map to obtain a repaired 2D color map.

[0208] The 2D color map obtained in S1013 may not be a color map of a complete face, but may contain a color map of a missing part (eg, a chin).

[0209] Then, the present application can segment the missing area in the 2D color map relative to the complete 3D face, and then use an image inpainting algorithm to inpaint the missing area (such as the chin), thereby obtaining a color map of the complete 3D face, which is named as the inpainted 2D color map.

[0210] In some scenarios, there may be mosaic areas or hollow areas in the face areas of massive 2D face images. Therefore, it is necessary to perform geometric and texture repair on the missing object areas to obtain the 3D asset data of a complete 3D face.

[0211] Through the above S1014, the third reference three-dimensional data with missing attachments may be added to the second reference three-dimensional data, wherein the missing attachments are attachments that are missing from the second reference three-dimensional data and can provide a reference for the virtual image to be generated;

[0212] S1015, post-processing the geometry of the 3D face (represented by a triangular mesh) and the repaired 2D color map to obtain a position map of the 3D face and a material map of the 3D face.

[0213] Through the above S1015, the second reference three-dimensional data after adding the third reference three-dimensional data can be converted into a first reference three-dimensional asset.

[0214] The material map may include but is not limited to at least one of the following: a diffuse map, a normal map, a specular map, a roughness map, a subsurface scattering map, etc.

[0215] Specifically, on one hand, the geometric data of a 3D face represented by a triangular mesh can be converted into a position map of the 3D face.

[0216] On the other hand, the repaired 2D color map can be subjected to operations such as highlight removal and shadow removal to obtain a diffuse map; the diffuse map is then processed to obtain a normal map and a specular map. For example, a codec network can be used to generate corresponding normal maps and specular maps end-to-end using a diffuse map as input. The present application does not limit the specific implementation method of processing the diffuse map to obtain the normal map and the specular map, and can be flexibly implemented according to requirements.

[0217] Through the above S1011 to S1015, the 2D face image can be preprocessed into a 3D asset representation, specifically a position map and a material map of the 3D face.

[0218] In the embodiment of the present application, when converting a two-dimensional reference image of a virtual image into a reference three-dimensional asset, the reference image can be first converted into reference three-dimensional data, and then the reference three-dimensional data of the missing attachments can be added to the reference three-dimensional data, where the missing attachments are attachments that are missing in the reference three-dimensional data and can provide references for the virtual image to be generated. Thus, the three-dimensional data of the missing attachments can be added to the converted reference three-dimensional data to repair the texture of the missing attachments, thereby obtaining a reference three-dimensional asset (here, the first reference three-dimensional asset) that can be used for model training.

[0219] S1021, retopology is performed on a large amount of medium-quality 3D face data to obtain position maps and material maps after retopology.

[0220] The medium-quality 3D face data may include geometric data (represented by a triangular mesh) and a material map of the 3D face. For example, the geometric data in the medium-quality 3D face data may include: a first triangular mesh providing a geometric reference for the face of the 3D digital human to be generated. The material map in the medium-quality 3D face data may include: a first material map providing a material reference for the face of the 3D digital human to be generated.

[0221] In one possible implementation, when retopologically processing the medium-quality 3D face data, the unified topological structure of the 3D face of the present application may be used for retopological processing. For example, the unified topological structure of the 3D face may be the same as the topological structure of the 3D model in S1012 above. In this way, there is no need to perform the unified topological operation of S1041 on the position map and material map of the 3D face obtained in S1023.

[0222] S1022, performing geometric cleaning on the position map and the material map after the retopology, to obtain a cleaned position map and a cleaned material map.

[0223] As described in Example 1 above, the medium-quality 3D face data here may include not only the 3D data about the face to be generated, but also the 3D data of objects unrelated to the face (such as a hat worn on a person's head). Therefore, the geometry of the objects unrelated to the face can also be cleaned up for the position map and material map after retopology.

[0224] Take, for example, an object unrelated to a human face (eg, a cleaned object) being a hat worn on a human head.

[0225] In this step, the position map of the hat area in the retopology position map can be deleted, and the material map corresponding to the deleted position map in the retopology material map (here, the material map of the hat area) can be deleted, so as to obtain a cleaned position map and material map that only includes the face object.

[0226] S1023, performing geometry and texture repair on the cleaned position map and material map to obtain a position map of the 3D face and a material map of the 3D face.

[0227] Let's continue to use the hat as an example.

[0228] Since the 3D data of the hat is cleaned, the 3D asset data of the cleaned 3D data about the face is incomplete, for example, the position map and material map of the forehead area are missing in the cleaned position map and material map. Then the missing forehead area can be repaired with geometric data (such as position map) and texture data (such as texture map) to obtain the position map and material map of the complete 3D face.

[0229] Through the above S1021 to S1023, the medium-quality 3D face data can be preprocessed into a 3D asset representation, specifically a position map and a material map of the 3D face.

[0230] S1031, retopology is performed on a small amount of high-quality 3D face data to obtain a position map and a material map of the retopologically 3D face.

[0231] Through the above S1031, the high-quality 3D face data represented by the 3D asset can be retopologically processed to obtain the position map and material map of the retopologically processed 3D face.

[0232] The implementation principle of S1031 is the same as that of the above-mentioned S1021, and will not be described in detail here.

[0233] The above three processing steps S1011 to S1015, S1021 to S1023, and S1031 can be executed in parallel or serially, and the present application does not limit the execution order of the three processing steps.

[0234] After the above three processes, Figure 1c As shown, the method may also include the following steps:

[0235] Optionally, S1041, retopology is performed on three groups of position maps and material maps about 3D faces according to a 3D face model with a unified topology, so as to obtain three groups of position maps and material maps of 3D faces with a unified topology.

[0236] As shown above, in the above S1012, S1021, and S1031, if retopology has been performed according to the unified topological structure of the 3D face of the present application, there is no need to execute S1041. Instead, the above three sets of data need to be unified into 3D asset data of the same topological structure.

[0237] S1042, performing a unified style operation on the 3D asset data (such as position map and material map) of the 3D face after the unified topology.

[0238] Since the training data of the 3D asset generation network 101 of the present application is generated based on multi-source asset data, there may be differences in image styles between the asset data from multiple sources, such as differences in image parameters such as hue and brightness of the image, differences in the presence or absence of certain accessories (such as eyebrows, eyelashes, and hair), or differences in image parameters of certain accessories.

[0239] Then the present application may perform a style unification operation on the first reference three-dimensional asset (e.g., the position map and material map of the 3D face obtained in S1015) and the second reference three-dimensional asset (e.g., the position map and material map of the 3D face obtained in S102 and S1031) to obtain the first reference three-dimensional asset and the second reference three-dimensional asset after style unification, wherein the first reference three-dimensional asset and the second reference three-dimensional asset after style unification have the same accessory objects described for the virtual image to be generated, and / or have the same image parameters for the features described for the accessory objects.

[0240] The image parameters may include but are not limited to: brightness, contrast, etc. The style of the same accessory object between each reference three-dimensional asset in the training data after the style is unified is made unified.

[0241] Among them, the style unification operation may include but is not limited to at least one of the following: unified setting of the presence or absence of accessories, the same image parameters of the same accessory between different reference 3D assets, and the same overall image parameters between different reference 3D assets (for example, the overall brightness, contrast and other image parameters of two training samples are the same).

[0242] Specifically, the present application can uniformly set the presence or absence of attachments, image parameters of the attachments, and image parameters of the overall image of the 3D asset for the three groups of 3D asset data.

[0243] Among them, regarding the presence or absence of attachments, for example, some 2D images corresponding to 3D asset data have eyebrows, while some 2D face images do not have eyebrows. In this case, this application can perform unified settings for attachments for the 3D asset data after unified topology. For example, 3D asset data about eyebrows can be added to 3D assets without eyebrows; and the image style of eyebrow attachments can be uniformly set, such as uniform settings of image parameters such as brightness and contrast; and the image parameters of the overall image of the 3D asset (such as brightness, contrast, etc.) can be uniformly set.

[0244] For example, in the 3D asset data after the unified style of the present application, the human face corresponding to each 3D asset data has a human face outline, a basic hair area, and an eyebrow area, but no eye area. Then, when the subsequent head attachments are generated and aligned, the 3D asset representation output by the trained 3D asset generation network 101 can be added with attachments such as eye attachments and hair attachments, and existing attachments can be modified (such as eyebrow attachments, etc.).

[0245] Among them, this application does not limit the execution order between S1041 and S1042, and can be flexibly set according to needs.

[0246] S1043, performing data augmentation on the 3D asset data (such as position maps and material maps) after the unified style, and obtaining the following Figure 1a The training data is shown.

[0247] In some embodiments, at least two of the first reference 3D asset and the second reference 3D asset may be mixed using a mixing weight to add at least one third reference 3D asset, wherein the third reference 3D asset includes a third reference position map and a third reference material map. In this way, there will be more reference 3D assets in the training sample.

[0248] The mixing weight indicates that at least two of the weights between the at least two reference three-dimensional assets being mixed are different.

[0249] For example, two training samples (ie, reference 3D assets) are mixed, one with a weight of 0.8 and the other with a weight of 0.2. The two training samples are mixed with their respective weights to obtain a new training sample.

[0250] For another example, three training samples (i.e., reference 3D assets) are mixed, with a training sample 1 having a weight of 0.4, a training sample 2 having a weight of 0.4, and a third training sample having a weight of 0.2. The three training samples are mixed with their respective weights to obtain a new training sample.

[0251] exist Figure 1c In the embodiment, the 3D asset data after the unified style includes multiple reference 3D assets, each of which includes a position map of a 3D face and a material map of a 3D face. Since the multiple reference 3D assets have completed the topological unification, when augmenting the data, the present application can randomly weight two or more of the multiple reference 3D assets to mix (because the topology has been unified) to obtain at least one new reference 3D asset, so as to achieve the purpose of enriching the training samples.

[0252] For example, random weights may be applied to perform interpolation or other processing on the reference 3D assets of young people and the reference 3D assets of elderly people to obtain the reference 3D assets of middle-aged people.

[0253] It should be understood that the above Figure 1cIn the embodiment, the geometric data of each 3D face expressed in a triangular mesh is converted into a position map in S1015, S1021, and S1031. In other embodiments, the geometric data of the 3D face can also be converted from a triangular mesh into a position map in the operation of unified topology in S1041, or after the operation of data augmentation in S1043, so as to obtain the final reference three-dimensional asset and its label information for model training. This application does not restrict the timing of converting the geometric data of the 3D face from the triangular mesh to the position map, as long as it is generated Figure 1c Before the final reference 3D asset used for model training, the conversion of geometric data is completed.

[0254] In the embodiment of the present application, it is considered that there may be missing parts of the 3D face in the multi-source asset data, for example, the face data in the multi-source asset data has mosaic or empty areas. Then the present application can perform texture repair on the missing parts of the face area of ​​the multi-source asset data, and perform geometric cleaning on the redundant parts, so that the geometry and material mapping of the face in the training data used to train the 3D asset generation network 101 are more complete, so that the 3D asset generation network 101 trained according to the training data can output a complete 3D asset representation of the 3D face. In addition, it is considered that there may be differences in the image style of the 3D face in the multi-source asset data, such as differences in the image style of the brightness and hue of the image; for example, some face pictures have eyelashes, and some face data do not have eyelashes; for example, there are differences in hair between different face data. Then the present application can perform a unified style operation on the multi-source asset data, so that the 3D image corresponding to the 3D asset representation generated each time by the 3D asset generation network 101 trained using the training data is of a unified style. In addition, the present application can also perform data augmentation on the training data to increase the data size of the reference three-dimensional assets and their label information used to train the 3D asset generation network 101, so that the 3D asset generation network 101 is easier to converge after training based on rich training samples, and can generate more accurate 3D asset representations about faces. Moreover, when preprocessing multi-source asset data, the present application can also convert the 3D asset representation that expresses geometric information in a triangular mesh into a position map that expresses the geometric information. Among them, the geometric data represented by the triangular mesh cannot be directly used by the neural network (such as the 3D asset generation network 101), and the converted position map is a two-dimensional expansion of the triangular mesh under a fixed topology. The position map stores the coordinates of each vertex in the geometric data in space. The information expressed by the triangular mesh and the position map is essentially the same, and both can accurately represent the geometric information. However, the position map is in an image format and can be used directly by a neural network (e.g., the 3D asset generation network 101). Thus, the present application can convert multi-source asset data into a 3D asset representation that can be directly used by the 3D asset generation network 101, thereby facilitating the use of the position map to train the 3D asset generation network 101.

[0255] Example 4

[0256] Combined with Figure 1a , optionally in combination with Figure 1b , Figure 1c , Figure 1d A schematic diagram showing an implementation process of training a neural network in the face generation system 100 shown in FIG. 2 of the present application is shown.

[0257] The neural network in the face generation system 100 may include: Figure 1dThe 3D asset generation network 101 shown may optionally further include a super-resolution network 103 .

[0258] Combine the following Figure 1d , to introduce the training process of the 3D asset generation network 101 and the super-resolution network 103.

[0259] like Figure 1d As shown, the process may include the following steps:

[0260] S301a, inputting the reference 3D asset and its tag information into the 3D asset generation network 101 of the digital human.

[0261] The reference 3D asset and its label information can be Figure 1a and Figure 1c The training data introduced in, for example, the reference three-dimensional asset can be a training sample converted from multi-source asset data.

[0262] The reference three-dimensional asset is a 3D asset representation of a human face, specifically including a reference position map of the 3D human face and a reference material map of the 3D human face.

[0263] In one example, the 3D asset generation network 101 may be a pre-trained GAN network.

[0264] For example, first, the GAN network can be pre-trained using hundreds of millions of texts and / or pictures, so that after training, it can generate images that match the semantics of the texts and pictures.

[0265] In this Example 4, the 3D asset generation network 101 is a GAN network as an example for illustration. When the 3D asset generation network 101 is other neural networks, the implementation principle of the method is the same and will not be repeated here.

[0266] In one possible implementation, Figure 1a and Figure 1c As shown, the reference 3D asset in S301a is generated by preprocessing multi-source asset data. When using the reference 3D asset to train the 3D asset generation network 101, the first reference 3D asset obtained by preprocessing a large number (e.g., 100,000, not specifically limited) of 2D face images (or text) can be used to continue to optimize the training of the pre-trained GAN network. Then, part of the second reference 3D asset obtained by preprocessing medium-quality 3D face data is used to continue to optimize the training of the GAN network. Finally, part of the second reference 3D asset obtained by preprocessing a small amount of high-quality 3D face data is used to continue to optimize the training of the GAN network to fine-tune the network parameters in the GAN network.

[0267] In other implementations, when training the GAN network using reference three-dimensional assets obtained by preprocessing multi-source asset data, the order of the various reference three-dimensional assets used may also be other orders, which is not limited here.

[0268] Optionally, S301b , random noise is input into the 3D asset generation network 101 of the digital human.

[0269] When training the 3D asset generation network 101 , by inputting random noise, the trained 3D asset generation network 101 can generate different 3D asset representations of faces based on the same user input information, thereby improving the diversity of the output results of the 3D asset generation network 101 .

[0270] After S301a and S301b, S302 is executed.

[0271] S302 , the 3D asset generation network 101 generates and outputs a 3D asset representation (a position map and a material map of a 3D face) of a first resolution according to the tag information and random noise.

[0272] After S302 , the network parameters of the 3D asset generation network 101 may be optimized through S105 .

[0273] Optionally, S105 , the 3D asset generation network 101 is optimized according to the reference three-dimensional asset and the 3D asset representation of the first resolution obtained in S302 .

[0274] The implementation principle of S105 in this example 4 is the same as the implementation principle of S105 in example 1, and will not be repeated here.

[0275] In a possible implementation, after S302, S303a may be executed.

[0276] S303a, the differentiable rendering network 102 may receive the 3D asset representation of the first resolution output by the 3D asset generation network 101, and perform differentiable rendering on the 3D asset representation of the first resolution under randomly set camera perspective parameters and lighting parameters regarding the digital human, so as to obtain a first 2D face image (an example of a 2D face image obtained by rendering the 3D asset representation).

[0277] In a possible implementation, after S303a, S304a may be executed.

[0278] S304a, based on the first 2D face image (an example of the first image), the massive 2D face images (an example of the multiple second reference images) and the differentiable rendering network 102, the 3D asset generation network 101 is updated using a back propagation algorithm to optimize the 3D asset generation network 101.

[0279] In this embodiment, the massive 2D face images in the multi-source asset data (as the first reference image) are the same as the second reference image used to train the 3D asset generation network 101 in combination with the network 102 for differentiable rendering, that is, the first reference image and the second reference image in this embodiment are the same, but in other embodiments, the second reference image may not be the massive 2D face images in the multi-source asset data, but a 2D face image obtained by other means that is different from the massive 2D face images. The second reference image may be a reference two-dimensional image of any virtual image, and the virtual image is not limited to the virtual image to be generated as described in the above-mentioned label information.

[0280] In the related art, the 3D asset generation network 101 is optimized by only calculating the loss for the reference three-dimensional asset and the 3D asset representation of the first resolution obtained by the 3D asset generation network 101 based on the label information. This loss is a 3D loss, which will result in that the 2D error cannot be directly transmitted back to the 3D asset generation network 101 for training optimization. To this end, the embodiment of the present application can obtain a large number of 2D face images as the second reference image, and the second reference image has rich 2D information about the face; and the 3D asset generation network 101 is updated using the back propagation algorithm using the micro-rendering network 102 and the second reference image with rich 2D information and the first image output by the 3D asset generation network 101. In this way, the image realism of the 3D asset representation output by the 3D asset generation network 101 can be improved based on the massive 2D face images by means of the micro-rendering network 102.

[0281] In some embodiments, when executing S304a, a loss calculation may be performed based on the first image and the plurality of second reference images to obtain a loss function; the loss function is passed by the differentiable rendering network 102 to the 3D asset generation network 101 to update the 3D asset generation network 101 using a back propagation algorithm.

[0282] like Figure 1dAs shown, at S302, the 3D asset generation network 101 can output the 3D asset representation of the first resolution generated based on the label information to the micro-rendering network 102; the micro-rendering network 102 can positively render each 3D asset representation received from the 3D asset production network 101 to positively obtain multiple first 2D face images; then, based on the multiple first 2D face images and the massive 2D face images that provide rich 2D information of the face, a loss function (such as GAN loss) is calculated. Then, the loss function can be passed to the 3D asset generation network 101 through the micro-rendering network 102, so as to use the loss function to update the 3D asset generation network 101 using the back propagation algorithm. In this way, the gradient corresponding to the loss function can be transmitted back to the 3D asset generation network 101 via the differentiable rendering network 102 to learn the network weights, so as to achieve iterative optimization of the 3D asset generation network 101 until convergence, so that the 3D asset representation generated by the trained 3D asset generation network 101 can have a stronger sense of reality than the corresponding 2D image.

[0283] In the process of S304a, a large number of real 2D face images (or texts) can be used as supervision data, and the differentiable rendering network 102 can be used to realize the transmission of the loss function to optimize the network parameters in the 3D asset generation network 101, so that the face image corresponding to the 3D asset representation output by the optimized 3D asset generation network 101 (for example, the 3D asset representation is rendered as a 2D image) is closer to the real face.

[0284] Optionally, after S303a, S304b may also be executed.

[0285] S304b, calculating the text consistency loss according to the first 2D face image obtained in S303a and the label information of the reference 3D asset in S301a, and optimizing the 3D asset generation network 101 based on the text consistency loss.

[0286] Specifically, the label information of the reference three-dimensional asset can be input into a text encoder to be converted into a latent vector 1 to obtain a first feature of the virtual image to be generated described in the label information; and the first 2D face image output by the differentiable rendering network 102 is input into an image encoder to be converted into a latent vector 2 to obtain a second feature of the virtual image to be generated described in the first 2D face image; then, the loss between latent vector 1 and latent vector 2 is calculated (for example, text consistency such as cosine distance, which is not specifically limited); finally, based on the text consistency loss, the 3D asset generation network 101 is updated using a back propagation algorithm to optimize network parameters until it converges.

[0287] In the process of S304b, the label information of the reference 3D asset in S301a can be used as supervision data, and the text consistency loss between the supervision data and the 2D face image rendered by the micro-rendering network 102 is calculated to optimize the network parameters in the 3D asset generation network 101, so that the 3D face corresponding to the 3D asset representation output by the optimized converged 3D asset generation network 101 can better meet the needs of the digital human face generated by the user. In other words, after the 3D asset generation network 101 is optimized and converged through S304b, when the 3D asset generation network 101 is used to generate a 3D asset representation according to the user input information, the 3D face corresponding to the 3D asset representation output by the 3D asset generation network 101 can better meet the needs expressed by the user input information.

[0288] It should be understood that the label information of the reference 3D asset used in S304b may be the label information of the reference 3D asset in S301a transmitted through the 3D asset generation network 101 and the micro-rendering network 102. Alternatively, the label information of the reference 3D asset used in S304b may also be externally input into the training system separately for use in the calculation of the text consistency loss in S304b, which is not limited here.

[0289] In addition, the module for executing the loss calculation of S304a and S304b above may also be implemented as a neural network, which is not limited here.

[0290] This application does not limit the execution order between S304a and S304b, and they are both executed after S303a.

[0291] Continue to refer to Figure 1d After S302, the present application may not only execute S303a to optimize the 3D asset generation network 101 through the micro-renderable network 102, but also, optionally, execute S303c to train the super-resolution network 103 through the material map in the reference three-dimensional asset of S301a, which may be specifically implemented through S301c, S303c, and S304c.

[0292] Continue to refer to Figure 1d The 3D asset representation of the first resolution obtained through S302 can be input by the 3D asset generation network 101 into the super-resolution network 103 as a training sample in the training data of the super-resolution network 103 .

[0293] In addition, if Figure 1d As shown, the training data of the super-resolution network 103 may also include supervision data of the 3D asset representation at the first resolution, such as a large number of 2D face images, 2D face images of 3D face data, etc.

[0294] Then, S303c, the super-resolution network 103 may perform super-resolution processing on the received 3D asset representation of the first resolution to obtain a 3D asset representation of the second resolution, so as to increase the resolution of the 3D asset representation.

[0295] For example, a 3D asset of the first resolution is represented as a position map and a material map with a resolution of 1000*1000. Then, through super-resolution, the super-resolution network 103 can super-resolve the position map and the material map with a resolution of 1000*1000 into position maps and material maps with a resolution of 4000*4000 (an example of the second resolution).

[0296] The second resolution is higher than the first resolution, and the present application does not limit the size of the two resolutions. In addition, the present application does not limit whether the resolution size of the position map and the material map in the 3D asset representation before or after super-resolution is the same.

[0297] Finally, after S303c, S304c may be executed.

[0298] S304c, calculating the loss between the supervision data and the 3D asset representation at the second resolution output by the super-resolution network 103, and optimizing the super-resolution network 103 according to the loss.

[0299] Among them, Figure 1d As shown, the supervision data used to train the super-resolution network 103 may include but is not limited to at least one of the following: a large number of 2D face images (examples of first reference images) in the multi-source asset data, and 2D face images (also referred to as third reference images) converted (for example, by rendering) from 3D face data (examples of first reference three-dimensional data) in the multi-source asset data.

[0300] It should be understood that the resolution of each picture in the supervision data is the resolution of the 3D asset representation of the face that the face generation system 100 of the present application is intended to generate.

[0301] In this example 4, the super-resolution network 103 and the 3D asset generation network 101 are two separate but interconnected networks, wherein the output result of the 3D asset generation network 101 can be used as the input of the super-resolution network 103. In other embodiments, the super-resolution network 103 can also be incorporated into the 3D asset generation network 101, so as to generate a high-resolution 3D asset representation of the 3D face of a digital human through a complete network.

[0302] In the embodiment of the present application, when training the 3D asset generation network 101 in the face generation system 100, the 3D asset representation output by the 3D asset generation network 101 can be rendered into a 2D face image through the pre-trained differentiable rendering network 102, and then the real massive 2D face images are used as supervision data to optimize the network parameters in the 3D asset generation network 101. On the one hand, since the image quality rendered by the differentiable rendering network 102 is high and has a high differentiable property, the 3D asset generation network 101 trained with this is easier to converge. On the other hand, since the real massive 2D face images can provide rich supervision signals, the 3D face corresponding to the 3D asset representation output by the optimized 3D asset generation network 101 is closer to the real face. Therefore, the 3D asset generation network 101 that converges after training can output high-quality 3D asset representations of faces. On the other hand, the present application may also use the label information of the reference three-dimensional asset as the supervision data in the training data of the 3D asset generation network 101, and optimize the network parameters in the 3D asset generation network 101 by calculating the text consistency loss between the supervision data and the 2D face image rendered by the micro-rendering network 102 (obtained by rendering the 3D asset representation output by the 3D asset generation network 101), so that the 3D face corresponding to the 3D asset representation output by the optimized converged 3D asset generation network 101 can better meet the user's desired digital human face generation needs.

[0303] In the embodiment of the present application, the face generation system 100 may include not only a 3D asset generation network 101, but also a super-resolution network 103. By training the super-resolution network 103, the resolution of the 3D asset representation output by the trained 3D asset generation network 101 of the present application can be improved, thereby combining with the super-resolution network to obtain high-quality 3D assets of digital humans.

[0304] After training the neural network in the 3D digital human face generation system 100 through any one of the above examples 1 to 4, the present application can use the face generation system 100 to generate a 3D asset representation (position map and material map) of the digital human face that conforms to the user input information (at least one form such as text, picture, or voice).

[0305] The following describes the use of the face generation system 100 of the present application in combination with different examples to implement the face generation method of the present application.

[0306] Example 5

[0307] Figure 2a A schematic diagram showing an application scenario of the face generation system 100 is shown.

[0308] like Figure 2a As shown, the face generation system 100 may include a 3D asset generation network 101 for a digital human that converges after training, and may optionally include a conversion module 105 .

[0309] Combine the following Figure 2a The following is an introduction to the use process of the face generation system 100 of the present application, which may include the following steps:

[0310] S201, the digital human 3D asset generation network 101 receives user input information.

[0311] The first user input information is used to describe the characteristics of the target virtual image to be generated.

[0312] Taking the target virtual image as the face of a digital person as an example, the user input information is information instructing to generate the face of the digital person. The user input information may be at least one of text, picture, voice, etc.

[0313] The characteristics of the target virtual image may be contour characteristics of a face, characteristics of the eyes (eg, color of the eyeballs), etc., and are not limited here.

[0314] For example, the user input information is the following text: "a middle-aged male with yellow skin and wrinkles and freckles on his face, with red curly hair and thin eyebrows". In other words, the user expects that the face generation system 100 can generate a face image of a digital person that meets the facial features of the above text.

[0315] S202 : The 3D asset generation network 101 generates a first 3D asset representation (also referred to as a third three-dimensional asset) of the target virtual image according to the user input information.

[0316] For example, the user input information may be converted into a latent vector and then input into the 3D asset generation network 101 .

[0317] The first 3D asset representation may include a position map and a material map of the 3D human face.

[0318] Then, the third three-dimensional asset may be rendered to obtain a second image of the target virtual image, wherein the second image is a two-dimensional image of the target virtual image. In this embodiment, this may be achieved through S206, S207, and S208. In other embodiments, the purpose of rendering the third three-dimensional asset into a 2D image may also be achieved through other methods, which are not limited here.

[0319] S206 , the conversion module 105 may convert the first 3D asset representation into a second 3D asset representation.

[0320] The conversion module 105 may convert the position map in the first 3D asset representation into a triangle mesh, thereby obtaining the second 3D asset representation.

[0321] In this way, the second 3D asset representation in which the geometric information is represented by a triangular mesh can be compatible with a traditional rendering pipeline, so as to render a 2D face image of the 3D digital human.

[0322] Optionally, S207 , the animation binding and driving module may bind the animation signal to the second 3D asset representation to obtain a 3D asset representation bound with the animation signal.

[0323] When a dynamic 3D digital human face image needs to be rendered, an animation signal can be bound to the 3D asset representation.

[0324] For the specific implementation process of S207, reference may be made to the prior art and will not be elaborated here.

[0325] Optionally, S208, a high-quality head rendering module renders the 3D asset representation bound with the animation signal to obtain a 2D image of the face of the 3D digital human or a 2D image sequence of the face (eg, 2D video).

[0326] Among them, the high-quality head rendering module may include a traditional rendering pipeline (such as a CG rendering pipeline) and a traditional neural renderer.

[0327] Since the data for rendering input into the high-quality head rendering module is a 3D asset representation, it can be compatible with the traditional rendering pipeline. In this way, the 3D asset representation generated by the face generation system 100 of the present application can be compatible with the existing game rendering pipeline and video rendering pipeline, and can be adapted to any image rendering scene, enriching the application scenario of the face generation system 100 of the present application.

[0328] Finally, the 2D image or 2D video rendered by the high-quality head rendering module can be output for display.

[0329] For example, the output 2D image is a face image of a middle-aged male with yellow skin, wrinkles and freckles, red curly hair, and thin eyebrows.

[0330] Optionally, in some embodiments, the face generation system 100 may also include: Figure 2a At least one of the animation binding and driving modules and the high-quality human head rendering module shown.

[0331] In addition, the 3D asset generation network 101 and the conversion module 105 in the face generation system 100 may also be combined or further split, and this application does not impose any limitation on this.

[0332] In the embodiment of the present application, the 3D asset generation network 101 in the face generation system 100 is optimized by the differentiable rendering network 102 during training, for example, by Figure 1d The network model is optimized in S304a shown, and the training data may include a 3D asset representation obtained by converting 2D data and 3D data, so that the image quality of the 3D asset representation output by the 3D asset generation network 101 of the present application is high; in addition, when training the 3D asset generation network 101, the network 102 with differentiable rendering is optimized, for example, by Figure 1d The S304b shown performs network model optimization so that the 3D asset generation network 101 in the face generation system 100 can output a 3D asset representation that is semantically consistent with the user input information. In this way, the face generation system 100 only needs to receive guidance information for generating a face in any modality (text, image, or voice, etc.) (such as the user input information in Example 5), and can quickly (for example, in minutes or milliseconds, etc.) generate a 3D asset representation of a digital human with high quality and semantically consistent with the guidance information. For example, if the 3D asset generation network 101 is a neural network model trained based on GAN, the 3D asset generation network 101 can generate a 3D asset representation of a digital human with high quality and semantically consistent with the guidance information within milliseconds.

[0333] Example 6

[0334] Figure 2b A schematic diagram showing the use process of the face generation system 100 is shown.

[0335] Most of the contents of Example 6 and Example 5 are the same, except that in Example 6, the face generation system 100 may also include a super-resolution network 103 and a head accessory generation and registration module 104.

[0336] In other embodiments, the difference between the face generation system 100 and Example 5 may also be that the face generation system 100 also includes at least one of a super-resolution network 103 and a head accessory generation and registration module 104.

[0337] The super-resolution network 103 is used to improve the resolution of the 3D asset representation output by the 3D asset generation network 101. The super-resolution network 103 is the trained super-resolution network 103 in Example 4. The training process of the super-resolution network 103 can refer to the relevant introduction in the above Example 4, which will not be repeated here.

[0338] The head accessory generation and registration module 104 is used to add accessories (such as eyebrows, hair, nose, etc.) to the 3D asset representation of the human face obtained through the 3D asset generation network 101. The added accessories are accessories of the human head, thereby obtaining a 3D asset representation of a complete human head of the 3D digital human.

[0339] like Figure 2b As shown, the process may include the following steps:

[0340] S201, the digital human 3D asset generation network 101 receives user input information.

[0341] Here, S201 is the same as S201 in Example 5, and will not be described in detail here.

[0342] The first user input information is used to describe the features of the target virtual image (e.g., the face of a digital person) to be generated. For example, the user input information is the following text: "a middle-aged male with yellow skin, wrinkles and freckles on his face, with red curly hair and thin eyebrows."

[0343] S202 , the 3D asset generation network 101 may generate a first 3D asset representation that complies with the user input information according to the user input information.

[0344] Here, S202 is the same as S202 in Example 5, and will not be described in detail here.

[0345] In a possible implementation, S203 , the super-resolution network 103 may perform super-resolution processing on the first 3D asset representation from the 3D asset generation network 101 to obtain a third 3D asset representation.

[0346] The image resolution of the third 3D asset representation is higher than the image resolution of the first 3D asset representation, thereby facilitating improving the image quality of the 3D asset representation obtained by the face generation system 100 .

[0347] In a possible implementation, S204, the head accessory generation and registration module 104 may add a 3D asset representation of the accessory missing from the head (the example of the fourth three-dimensional asset described below) to the third 3D asset representation, thereby obtaining a complete fourth 3D asset representation of the 3D head.

[0348] In this embodiment, based on a preset three-dimensional asset library of the virtual image, a fourth three-dimensional asset of at least one accessory that matches the characteristics of the target virtual image can be determined, the preset three-dimensional asset library includes three-dimensional data of the attachment, the three-dimensional data includes geometric data and material parameters, and the fourth three-dimensional asset includes a fourth position map and a fourth material map; then, the fourth three-dimensional asset of the at least one attachment is added to the third three-dimensional asset.

[0349] For example, the first 3D asset representation of a human face obtained by the 3D asset generation network 101 only includes the geometric texture of the human face (expressing the external contour of the human face) and the material texture (expressing the skin color, skin texture, etc.), but does not include the geometric and material information of the facial features, eyebrows, hair and other accessories in the human head. Correspondingly, the third 3D asset representation obtained after super-resolution also lacks the geometric and material information of the facial features, eyebrows, hair and other accessories in the human head.

[0350] Then, in order to render a 2D image or a 2D image sequence of a complete head of a 3D digital human, the head attachment generation and registration module 104 can be used to search the preset 3D asset library for 3D data of at least one attachment that matches the characteristics of the target virtual image described by the user input information. Then, based on the 3D data of the at least one attachment found, a fourth 3D asset of the at least one attachment is obtained to be added to the third 3D asset. In this way, the 3D asset representation (e.g., Figure 2b The third 3D asset representation shown in the figure) is added with the 3D asset representation (including position map and material map) of the missing accessories (such as hair and eyebrows), thereby obtaining a 3D asset representation of a complete human head of the 3D digital human (such as the fourth 3D asset representation described below).

[0351] The detailed process of adding and configuring the attachments by the head attachment generation and registration module 104 will be introduced in Example 7.

[0352] In other embodiments, S204 may also be executed before S203, so that the 3D asset representation of the missing accessories is first added to obtain a complete 3D asset representation of the 3D human head, and then the 3D asset representation is super-resolution processed to improve the resolution of the image and improve the image quality of the rendered image.

[0353] After S204 , in S205 , the conversion module 105 may convert the fourth 3D asset representation into a second 3D asset representation.

[0354] The geometric data in the second 3D asset representation is represented by a triangle mesh.

[0355] The operations after S205 are the same as those in Example 5 and will not be described in detail here.

[0356] In the embodiment of the present application, the 3D asset generation network 101 can be used to generate a high-quality 3D asset representation of the 3D face of a digital human that conforms to its semantics for user guidance information of any modality; in order to improve the image quality, the super-resolution network 103 can also be used to super-resolution the 3D asset representation, so that the resolution of the image (such as material map) in the generated 3D asset representation can reach 2 kilobytes (2KB) to 8KB (not specifically limited, and can also be other high-resolution numerical ranges); in addition, in order to render an image of a complete human head, the head accessory generation and alignment module 104 can be used to add a 3D asset representation of the accessories missing from the human head to the super-resolution 3D asset representation, thereby obtaining a complete 3D asset representation of the 3D human head, and then the complete 3D asset representation can be used to render an image or video of the complete head of the 3D digital human.

[0357] Example 7

[0358] Combined with Figure 2b , Figure 2c The figure shows a process diagram of the method of adding and configuring accessories implemented by the head accessory generation and registration module 104 in the face generation system 100 .

[0359] In the introduction Figure 2c Before that, the attachment asset library (an example of a preset three-dimensional asset library for a virtual image) pre-configured by the head attachment generation and registration module 104 of the present application is first introduced.

[0360] The accessory asset library may include 3D data of the corresponding accessory of the avatar (here, a human face), and the 3D data includes geometric data and material parameters (which may be a material map or information other than a map that can describe the material) of the corresponding accessory. One or more fixed values ​​of the material parameters of the corresponding accessory may be preconfigured in the accessory asset library.

[0361] For example, the attachments in the attachment asset library may include, but are not limited to, at least one of: eyes, hair, eyebrows, eyelashes, and mouth.

[0362] For example, the geometric data of the hair attachment may include geometric information of hair of various shapes and structures, such as geometric data of long wavy hair, geometric data of straight hair, etc. The material parameters of the hair attachment may include but are not limited to at least one of the following: color, roughness, density and other parameters, and any material parameter may be assigned one or more fixed values ​​in advance. In this way, the hair asset library may include multiple hair assets, each of which may include geometric data and material parameters, and different hair assets may have differences in at least one of geometry and material.

[0363] The composition of the attachment asset library for other attachments, such as eyes, eyebrows, eyelashes, mouth, etc., is similar and will not be repeated here.

[0364] In some embodiments, when determining the fourth three-dimensional asset of at least one accessory that matches the characteristics of the target virtual image based on the preset three-dimensional asset library of the virtual image, the three-dimensional data of at least one accessory that matches the characteristics of the target virtual image may be determined based on the preset three-dimensional asset library of the virtual image; the three-dimensional data of the at least one accessory may be rendered by a micro-rendering module to obtain a third image of the at least one accessory; a third feature describing the target virtual image may be extracted from the first user input information; a fourth feature describing the target virtual image may be extracted from the third image; and a loss between the third feature and the fourth feature may be calculated to optimize the material parameters in the three-dimensional data of the at least one accessory to obtain the optimized three-dimensional data of the at least one accessory. The optimized three-dimensional data of the at least one accessory may be converted into the third three-dimensional asset of the at least one accessory that matches the characteristics described by the first user input information.

[0365] In an embodiment of the present application, the three-dimensional data of at least one attachment matching the feature of the target virtual image described by the user input information can be first searched in the preset three-dimensional asset library. The at least one attachment is an attachment about the human face that is missing in the third three-dimensional asset, such as eyebrows, hair, and other attachments. Considering that the three-dimensional data of the attachment matching the feature in the preset three-dimensional asset library is not the best match for the feature indicated by the user input information, the three-dimensional data of the at least one attachment found can be micro-rendered (can be any micro-rendering module) to obtain a third image of the at least one attachment. Then, the loss between the feature of the third image and the feature of the user input information can be used to optimize the three-dimensional data of the at least one attachment found, and finally the optimized three-dimensional data of the at least one attachment is converted into a three-dimensional asset. The three-dimensional asset of the attachment matching the user input information can be obtained.

[0366] Combine the following Figure 2c To describe the specific process of implementing the above attachment addition.

[0367] like Figure 2c As shown, the process may include the following steps:

[0368] S401: Convert the user input text into a vector of at least one attachment to be added.

[0369] In Example 7, the above Figure 2a , Figure 2bThe user input information shown is user input text, which is "a middle-aged male with yellow skin, wrinkles and freckles on his face, with red curly hair and thin eyebrows." The text "a middle-aged male with yellow skin, wrinkles and freckles on his face" describes the characteristics of a human face, not the characteristics of the appendages of a human head. Figure 2c It mainly shows the process of adding missing attachments, so Figure 2c The user input text shown only shows the text information "red curly hair, thin eyebrows", but does not show the text "middle-aged male with yellow skin and wrinkles and freckles on his face" also included in the user input text.

[0370] Among them, the present application can convert the text describing the human head accessories in the user input information into a vector, so that only the text "red curly hair, thin and long eyebrows" is converted into a vector to obtain the latent vector b1 about the text "red curly hair" and the latent vector b2 about the text "thin and long eyebrows".

[0371] For example, when converting text into latent vectors, it can be achieved through a contrastive language-image pre-training (CLIP) model based on contrastive learning, or through other methods, which are not limited here.

[0372] The at least one attachment is an attachment that matches the characteristics of the target avatar, that is, the user expects the generated digital person to have hair and eyebrows, and the hair is red curly hair and the eyebrows are thin and long.

[0373] S501, based on the randomly sampled viewing angle and light source, the parameters in the hair asset library and the eyebrow asset library are rendered to obtain a 2D image library of hair accessories and a 2D image library of eyebrow accessories.

[0374] As described above, the hair asset library may include multiple hair assets, each of which may include hair geometry parameters and hair material parameters (such as color, roughness, density, etc.), and each hair material parameter has an initial fixed value. The hair geometry parameters may be used to represent the geometric shape of the hair.

[0375] Then, each hair asset in the hair asset library may be rendered based on the randomly sampled perspective and light source, thereby obtaining a 2D image library of hair accessories, wherein the 2D image library includes multiple 2D images of hair, and the multiple 2D images of hair may have at least one difference in geometry and material.

[0376] Similarly, the eyebrow asset library may include multiple eyebrow assets, each eyebrow asset may include eyebrow geometry parameters and eyebrow material parameters, and each eyebrow material parameter may have multiple fixed values. Among them, the eyebrow geometry parameters can be used to represent the geometric shape of the eyebrow.

[0377] Then, each eyebrow asset in the eyebrow asset library can be rendered based on the randomly sampled view angle and light source, thereby obtaining a 2D image library of eyebrow attachments, which includes multiple 2D images of eyebrows, and there may be at least one difference in geometry and material between the multiple 2D images of eyebrows. Among them, for multiple eyebrow assets with the same geometric parameters but different fixed values ​​assigned to the material parameters, they can be rendered into multiple 2D images of eyebrows, and the eyebrow shapes of the multiple 2D images of eyebrows are the same, but the material parameters such as the color and density of the eyebrows are different.

[0378] The asset library for other types of attachments in the attachment asset library is handled in the same way and will not be described here.

[0379] In this way, a 2D image of the corresponding attachment obtained by rendering all the attachment assets in the preset attachment asset library can be obtained.

[0380] S502, convert each 2D picture in the 2D picture library of hair accessories and the 2D picture library of eyebrow accessories into a vector to obtain a latent vector set A1 corresponding to the 2D picture library of hair accessories and a latent vector set A2 corresponding to the 2D picture library of eyebrow accessories.

[0381] Of course, in order to illustrate the latent vector sets of different attachments, they are divided into latent vector set A1 and latent vector set A2 for separate description. However, in other embodiments, they can also be converted into one latent vector set, which includes latent vectors corresponding to various attachment asset libraries.

[0382] S502 can also realize the conversion from image to latent vector through the CLIP model, which is not limited or elaborated here.

[0383] Among them, S501 and S502 can be completed in advance before receiving user input information.

[0384] After S401 and S502 , S601 may be executed.

[0385] S601, for each latent vector of the attachment to be added obtained through S401, search for the latent vector that best matches the latent vector in the latent vector set obtained through S502, thereby obtaining a candidate set of latent vectors for each required attachment (an example of three-dimensional data of at least one attachment that matches the features of the target virtual image).

[0386] like Figure 2c As shown, S601 can be used to search for one or more latent vectors that best match the latent vector b1 of the hair accessory (e.g., have the smallest vector distance) in the latent vector set A1 related to the hair accessory to obtain a candidate set K1 of latent vectors for the hair accessory.

[0387] The candidate set K1 may include k1 latent vectors from the latent vector set A1.

[0388] Similarly, S601 can be used to search for one or more latent vectors that best match the latent vector b2 of the eyebrow attachment (for example, with the smallest vector distance) in the latent vector set A2 related to the hair attachment to obtain a candidate set K2 of latent vectors for the eyebrow attachment.

[0389] The candidate set K2 may include k2 latent vectors from the latent vector set A2.

[0390] The values ​​of k1 and k2 can be any positive integers.

[0391] In a possible implementation manner A, after S601 , S701 may be executed.

[0392] S701, select an optimal latent vector about hair attachment from the candidate set K1, and determine the optimal hair asset (including position map and material map) corresponding to the optimal latent vector; and select an optimal latent vector about eyebrow attachment from the candidate set K2, and determine the optimal eyebrow asset (including position map and material map) corresponding to the optimal latent vector.

[0393] Let’s take determining the optimal hair asset as an example. The method for determining the optimal eyebrow asset is similar and will not be elaborated here.

[0394] The method of selecting the optimal latent vector from the candidate set K1 can be implemented in any way, and there is no limitation here. For example, the optimal latent vector is a latent vector with the smallest vector distance to the latent vector b1. Alternatively, the 2D picture of the hair attachment corresponding to each latent vector in the candidate set K1 can be output for display, and the user can select an optimal 2D picture that best meets the hair requirements (here, red curly hair) of the generated 3D digital person from the 2D pictures of the output hair attachments, and then the hair asset corresponding to the optimal 2D picture can be determined as the optimal hair asset.

[0395] In addition, as described above, the optimal hair asset determined by the present application is the position map and material map of the hair. However, the geometric parameters of each accessory asset in the preset accessory asset library are not necessarily the position map, and its material parameters are not necessarily the material map. Then, when determining the optimal hair asset, the geometric parameters of the hair asset corresponding to the optimal latent vector of the hair in the hair asset library can be converted into the position map, and the material parameters of the hair asset corresponding to the optimal latent vector of the hair in the hair asset library can be converted into the material map, so as to obtain the optimal hair asset represented by the 3D asset.

[0396] Similarly, the optimal eyebrow asset represented by a 3D asset can be obtained.

[0397] After S701 , S702 may be executed.

[0398] S702 , adding the optimal hair asset and the optimal eyebrow asset to the third 3D asset representation through a registration algorithm to obtain a fourth 3D asset representation.

[0399] like Figure 2b As shown, the third 3D asset representation is a 3D asset representation of a 3D face that has been super-resolution processed but lacks accessories (such as hair and eyebrows). Then, the automatic registration algorithm can be used to determine the positions of the hair and eyebrows to be added in the 3D model of the face corresponding to the third 3D asset representation. Then, based on the positions, the optimal hair asset and the optimal eyebrow asset represented by the 3D asset can be added to the third 3D asset representation to obtain a fourth 3D asset representation including a complete face.

[0400] The fourth 3D asset representation is a 3D asset representation of a complete human head, which may include complete human head accessories such as face, hair, eyebrows, eyes, eyelashes, nose, mouth, ears, etc. Then, a complete 2D picture or 2D video of a human head can be rendered.

[0401] In this way, for any attachment that needs to be added in the user input text, the present application can search through the attachment and find the optimal attachment asset in the preset attachment asset library that best matches the user input text (including the guidance information of the attachment to be added) in terms of geometric parameters and material parameters, and convert the optimal attachment asset into a 3D asset representation to obtain a 3D asset representation of the attachment to be added that matches the user's semantics, thereby improving the efficiency of searching for the attachment to be added.

[0402] In another possible implementation B, when searching for the optimal accessory asset (such as the hair asset, eyebrow asset, etc.), for the material parameters of the accessory, only the vector distance calculation in S601 is used to determine the optimal accessory asset that matches the user input text, and the accuracy is relatively low. In order to improve the accuracy of the determined optimal accessory asset and the matching degree with the user guidance information, the material parameters of the accessory can be optimized through S602 to S605, so that the optimized material parameters can match the user input text, thereby obtaining the optimal hair accessory and the optimal eyebrow accessory. Of course, in other embodiments, the geometric parameters of the accessory can also be optimized according to the principles of S602 to S605 below, so that the optimized geometric parameters can match the user input text, thereby obtaining the optimal hair accessory and the optimal eyebrow accessory.

[0403] The following takes the determination of the optimal hair attachment as an example to illustrate the process of optimizing the material parameters from S602 to S605. Figure 2c As shown, the processing process of each latent vector in the candidate set K1 in the following S602 to S605 is the same. The processing process is explained below by taking a latent vector p in the candidate set K1 as an example.

[0404] After S601 , S605 may be executed.

[0405] S605, calculating the loss between the user input text and the latent vector (eg, latent vector p) in the candidate set K1.

[0406] For example, “red curly hair” in the user input text can be converted into a latent vector (an example of the third feature), and then the loss (e.g., vector distance) between it and the latent vector p is calculated.

[0407] After S605 , S602 is executed.

[0408] S602, optimizing material parameters of the hair attachment based on the loss obtained in S605.

[0409] Among them, the hair asset corresponding to the latent vector p in the hair asset library is close to "red curly hair". For example, the geometric parameters of the hair asset corresponding to the latent vector p correspond to the geometric information of "curly hair", and the material parameters of the hair asset corresponding to the latent vector p in the hair asset library are pre-configured with an initial fixed value, such as the color is reddish brown.

[0410] Then the initial fixed value of the color of the hair asset (hereinafter referred to as the target hair asset) corresponding to the latent vector p in the hair asset library is reddish brown (color is a type of material parameter). Although it is close to the red color expected by the user, it is not completely consistent. Therefore, the parameter value of the material parameter (such as color parameter) of the target hair asset can continue to be optimized until it is very close to the parameter value of "red" expected by the user, or is the parameter value of "red".

[0411] The optimization goal of S602 is to minimize the loss. The smaller the loss, the closer the material parameters of the optimized target hair asset are to the user input text.

[0412] like Figure 2c As shown, after S602, S603 may be executed.

[0413] S603 , performing microscopic rendering of the hair accessory based on the optimized material parameters of the hair accessory, thereby obtaining a 2D image of the hair accessory.

[0414] For example, under randomly sampled perspectives and light sources, differentiable rendering may be performed based on geometric parameters of the target hair asset (e.g., a 3D model of hair corresponding to the geometric parameters) and optimized material parameters to obtain a 2D image of the hair accessory.

[0415] Among them, the module used to perform the differentiable rendering operation of S603 can be a differentiable renderer in the prior art, or it can be a differentiable rendering network for accessories in the face (such as hair accessories) trained according to the principle of the training method of the differentiable rendering network 102 of the present application, so as to render a 2D picture of the corresponding attachment.

[0416] After S603, S604, the 2D image of the hair attachment is converted into a vector to obtain a latent vector of the hair attachment.

[0417] Among them, the conversion from image to latent vector can be realized through the CLIP model, which is not limited or elaborated here.

[0418] After S604, S605 is executed again, but when S605 is executed again, the loss between the latent vector output by S604 and the user input text (such as "red curly hair") is calculated.

[0419] Then, when it is determined that the loss has not reached the minimum value or has not converged, S602, S603, S604, and S605 are continuously executed in a loop based on the loss until at least one of the conditions of the loss being minimized or convergence is met. At this time, the material parameters of the target hair accessory after multiple iterations of optimization are the material parameters of the optimal hair asset.

[0420] When at least one of the conditions of the loss being minimized or convergence in S605 is met, S606 may be executed.

[0421] S606: Convert the geometric parameters of the target hair attachment (the hair asset corresponding to the latent vector p) and the optimized (eg, with the minimum loss) material parameters into a 3D asset representation (position map and material map) to obtain an optimal hair asset.

[0422] In this way, for a latent vector p in the candidate set K1, a corresponding optimal hair asset is obtained; when the candidate set K1 includes multiple latent vectors, a corresponding optimal hair asset can be obtained according to the same process. Then, the user or the system can select an optimal hair asset from the multiple optimal hair assets for attachment registration.

[0423] Similarly, in a similar manner, the optimal eyebrow asset can be obtained.

[0424] Finally, through the above S702, the optimal hair asset and the optimal eyebrow asset can be added to the third 3D asset representation with incomplete attachments through attachment registration to obtain a 3D asset representation of a complete human head.

[0425] In this way, for any attachment that needs to be added in the user input text, the present application can search through the attachment and find the candidate attachment asset in the preset attachment asset library that best matches the user input text (including the guidance information of the attachment to be added) in terms of geometric parameters and material parameters. Then, the candidate attachment asset is optimized by optimizing the material parameters to obtain the optimal attachment asset, so that the material of the optimal attachment asset is more consistent with the user semantics, thereby improving the accuracy of the added attachments.

[0426] Of course, in some embodiments, the above implementation method A may be used for some of the missing multiple attachments to obtain the optimal attachment asset by searching the asset library, and the above implementation method B may be used for the other attachments to obtain the optimal attachment asset by searching the asset library and optimizing the material parameters. The specific solution can be flexibly selected and used according to the needs, and is not limited here.

[0427] In this example 7, the 3D asset representation generated by the 3D face generation network 101 that conforms to the semantics of the user input information is used to obtain a high-quality 3D asset representation of a 3D human head that conforms to the user's semantics through attachment search, material parameter optimization, and attachment alignment, so as to facilitate high-quality rendering of the 3D digital human.

[0428] Example 8

[0429] Please refer to Figure 2d, this Example 8 introduces the usage process of another face generation system 100.

[0430] like Figure 2d As shown, the face generation system 100 in this Example 8 may include not only a 3D asset generation network 101 for a digital human, but also a differentiable rendering network 102 trained through Example 2.

[0431] In other words, different from the above Figure 2a (Example 5) and Figure 2b (Example 6) When the 3D asset generation network 101 is applied, the first 3D asset representation output by the 3D asset generation network 101 can also be optimized with the help of the micro-rendering network 102, so that the optimized first 3D asset representation is more consistent with the semantics of the user input information.

[0432] In a possible implementation A, the 3D asset generation network 101 in Example 8 can be implemented by the above Examples 1 to 4 ( Figure 1a to Figure 1d ) is obtained by the training method of any example in .

[0433] In another possible implementation B, when training the 3D asset generation network 101 in Example 8, the 3D asset generation network 101 may not be trained and optimized by the differentiable rendering network 102, but by, for example Figure 1a S105 or Figure 1d The training optimization is performed in S105 shown in the figure to obtain the 3D asset generation network 101 of the digital human in the face generation system 100 in Example 8. For example, the 3D asset generation network 101 is Stable Diffusion.

[0434] like Figure 2d As shown, the use process of the face generation system 100 may include the following steps:

[0435] S201, the face generation system 100 receives user input information.

[0436] For the user input information in S201, please refer to the relevant introduction of S201 in Example 5, which will not be repeated here.

[0437] S202 , the 3D asset generation network 101 may generate a first 3D asset representation that complies with the user input information according to the user input information.

[0438] For the introduction of S202, please refer to the introduction of S202 in the above example 5. The implementation principles of the two are the same and will not be repeated here.

[0439] In some embodiments, the 3D asset generation network 101 may be trained by the training method described in any of Examples 1 to 4 above.

[0440] In this embodiment, the first 3D asset representation output by the 3D asset generation network 101 may not be further optimized, so that the 3D asset representation of the digital human can be quickly output within milliseconds.

[0441] In other embodiments, referring to Figure 1a , when the 3D asset generation network 101 is trained, only S105 can exist in the two optimization processes of S104 and S105, and the network 102 with differentiable rendering can be not used to optimize the model. For example, if the 3D asset generation network 101 is Stable Diffusion, then when the 3D asset generation network 101 is used, Figure 2d The micro-rendering network 102 is used to optimize the output result of the 3D asset generation network 101 (such as S802 and S803 described below) to improve the image quality.

[0442] In a possible implementation, in Example 5 Figure 2a The process shown in the figure can render a 2D image or a 2D video of a 3D digital human. However, the user is not satisfied with the rendering effect of the human face in the 2D image or the 2D video. For example, the user inputs text that he wants to get a face of a "middle-aged male with yellow skin and wrinkles and freckles on his face". However, the rendered 2D picture (for example, Figure 2d The first 3D asset output by the 3D asset generation network 101 represents the rendered image) in which the male face has darker skin and fewer freckles. Figure 2d to generate an optimized first 3D asset representation.

[0443] In one possible implementation, Figure 2d As shown, after S202, S801 may be executed.

[0444] S801 , calculating the loss between the first 3D asset representation and the user input information, and optimizing the aforementioned first 3D asset representation (eg, the first 3D asset representation obtained in S202 ) based on the loss.

[0445] Among them, the user input information can be converted into a latent vector for calculating the loss. The specific conversion method can be referred to the introduction above and will not be repeated here.

[0446] Wherein, the first 3D asset representation may include a position map and a material map. Then, when optimizing the first 3D asset representation, the position map may be optimized based on the loss between the position map (e.g., converted from an image to a latent vector) and the user input text (e.g., converted from text to a latent vector); and the material map may be optimized based on the loss between the material map (e.g., converted from an image to a latent vector) and the user input text (e.g., converted from text to a latent vector) until the loss is minimized. In this way, the user input text may be used to optimize the first 3D asset representation, so that the optimized first 3D asset representation is closer to the semantics of the user input text, and the image rendered by the optimized first 3D asset representation is closer to the semantics of the user input text, so as to improve the image quality rendered by the optimized first 3D asset representation. Moreover, a high-quality 3D asset representation can be output within minutes.

[0447] In one possible implementation, Figure 2d As shown, after S202, S802 and S803 may also be executed.

[0448] S802, the differentiable rendering network 102 may perform differentiable rendering (for example, under randomly set lighting and camera perspective) on the first 3D asset representation (which may be before or after optimization in S801) to obtain a first 2D face image, and optionally may also obtain a normal vector of the face.

[0449] S803: The loss between the first 2D face image (optionally including the normal vector) output by the differentiable rendering network 102 and the user input text may be calculated, and based on the loss, the first 3D asset representation input to the differentiable rendering network 102 may be optimized.

[0450] Among them, the user input information and the first 2D face image can be converted into latent vectors for calculating the loss. The specific conversion method can be referred to the introduction above and will not be repeated here.

[0451] Similar to S801 above, the first 3D asset representation to be optimized may include a position map and a material map. Therefore, when optimizing the first 3D asset representation based on the loss obtained in S803, the position map and the material map may be optimized separately until the loss is minimized.

[0452] In a possible implementation, the neural network model of the 2D Vincent graph in the prior art may be used to execute S803.

[0453] Among them, the neural network model of the 2D text map is a model that can output an image that conforms to the semantics of any text. Since the model has been pre-trained with relatively rich training samples, the model can output any image that conforms to the text, and the diversity of the output images is stronger.

[0454] Specifically, the first 2D face image (optionally including a normal vector) is input into the neural network model of the 2D Wensheng graph by the micro-renderable network 102, and the user input text in S201 is input into the neural network model of the 2D Wensheng graph. Then, the neural network model of the 2D Wensheng graph can generate a target image matching the user input text, calculate the loss between the target image and the first 2D face image, and reversely optimize the first 3D asset representation based on the loss.

[0455] In this embodiment, the neural network model of the 2D text image is used to calculate the loss between the first 2D face image (optionally including the normal vector) and the user input text to optimize the first 3D asset representation, thereby improving the diversity of the generated digital human face. In addition, a high-quality 3D asset representation can be output within minutes.

[0456] This application does not impose any limitation on the execution order of the optimization process of S801 and the optimization processes of S802 and S803.

[0457] Then, the first 3D asset representation optimized through the above two optimization processes has higher image quality when rendered as a 2D image, and the image content is more consistent with the semantics of the user input information.

[0458] It should be understood that the Figure 2d The use process of the face generation system 100 can also be Figure 2a , Figure 2b , Figure 2c They are combined to form new embodiments. For example, Figure 2d The first 3D asset representation after optimization in S801 and S803 can be input into Figure 2b The super-resolution network 103 shown in the figure performs operations such as super-resolution and attachment addition.

[0459] Example 9

[0460] Please refer to Figure 2eThe difference between Example 9 and Examples 5 to 8 is that the face generation system 100 of the present application supports users to input editing information twice, so that the 3D asset representation of a 3D face will not be regenerated based on the editing information input twice, so as to avoid the difference between the regenerated 3D asset representation and the previous 3D asset representation being too large (for example, the 3D assets of two completely different people). Instead, the semantic difference between the two most recent user input information is combined, and based on the first 3D asset representation generated last time, the first 3D asset representation is adjusted to obtain a fifth 3D asset representation, and the difference between the first 3D asset representation and the fifth 3D asset representation matches the semantic difference between the two most recent user input information. In this way, the face generation system 100 of the present application can support users to edit face generation multiple times.

[0461] Combined with Figure 2a (optionally combined with Figure 2b to Figure 2d ), such as Figure 2e As shown, the process may include the following steps:

[0462] S201, the digital human 3D asset generation network 101 receives user input information.

[0463] The specific implementation of S201 may refer to the introduction of S201 in Example 5, the principle is the same and will not be repeated here.

[0464] For example, the user input information is the following text: "a middle-aged male with yellow skin, wrinkles and freckles on his face, red curly hair, and thin eyebrows".

[0465] S202 , the 3D asset generation network 101 may generate a first 3D asset representation that complies with the user input information according to the user input information.

[0466] The specific implementation of S202 may refer to the introduction of S202 in Example 5, and the principle is the same, which will not be repeated here.

[0467] S206 , the conversion module 105 may convert the first 3D asset representation into a second 3D asset representation.

[0468] The specific implementation of S206 may refer to the introduction of S206 in Example 5, the principle is the same and will not be repeated here.

[0469] After S206, refer to Figure 2a , through operations such as S207 and S208, a face image or a video of a 3D digital human rendered from the 3D asset representation generated by the face generation system 100 of the present application can be output.

[0470] So, for example Figure 2b The second 3D asset generated by the face generation system 100 shown in the figure indicates that the rendered 2D face image may be a face picture of a middle-aged male with wrinkles and freckles, yellow skin, red curly hair, and thin eyebrows.

[0471] Then, the user needs to make some minor adjustments to the face image generated this time, such as removing freckles from the face, adding pale lips, changing the skin color to white, and at least one other minor adjustment.

[0472] Then the user can input the editing information so that the system 100 of the present application can execute S902.

[0473] S902, the digital human 3D asset generation network 101 receives the editing information of the user input information.

[0474] The editing information can be used to provide feature editing information of the target virtual image. The feature editing information is information about changes in the features of the target virtual image compared to those described by the user input information in S201.

[0475] For example, if the text of the editing information of the user input information is "Please remove the freckles on the face", then the feature editing information may be "Remove the freckles on the face".

[0476] Or the text of the editing information of the user input information is at least one text such as "please adjust the lip color to a pale color" or "please change the skin color to white".

[0477] Alternatively, the editing information is information with more accurate semantics than the last user input information (information in S201). For example, if the last user input information causes the image corresponding to the 3D asset representation output by the 3D asset generation network 101 to not quite meet the user's needs, then the user may input editing information that better meets his or her needs than the last user input information, so that the 3D asset generation network 101 can output a 3D asset representation of a 2D image of a digital human face that meets the user's needs.

[0478] Then, the face generation system 100 of the present application can recognize that the user input received in S902 is related to the previous user input (e.g., the user input information in S201) based on the semantics of the two previous user inputs, and thus will not treat the edit information in S902 as a new user input information to perform the above Figure 2a , Figure 2b , Figure 2c , Figure 2dAny process in the process will not generate a 3D asset representation of the 3D digital human that only matches the editing information according to the editing information.

[0479] In contrast, the 3D asset generation network 101 in the face generation system 100 of the present application may execute S903 .

[0480] S903, the 3D asset generation network 101 may adjust the first 3D asset representation generated based on the previous user input information according to the feature difference between the two most recent user input information (the user input in S201 and S902) regarding the target virtual image, and generate a fifth 3D asset representation of the target virtual image that matches the feature editing information.

[0481] Among them, the fifth 3D asset representation may include a position map and a material map.

[0482] Among them, Figure 2e As shown, the first 3D asset representation previously generated by the 3D asset generation network 101 can be reused as the input of the 3D asset generation network 101 to obtain the fifth 3D asset representation.

[0483] In other embodiments, the first 3D asset representation based on S903 may also be input by the user to the face generation system 100. For example, the face generation system 100 of the present application may output the first 3D asset representation obtained by the 3D asset generation network 101 based on the user input information in S201 from the system 100, so that the user can obtain the previous first 3D asset representation. When the user is not satisfied with the rendered 2D image corresponding to the previous first 3D asset representation, the user may not only input the editing information in S902 into the system 100, but also input the first 3D asset representation obtained the most recently into the system, so that the 3D asset generation network 101 in the system 100 executes S903.

[0484] The difference between the first 3D asset representation and the fifth 3D asset representation matches the feature difference about the target virtual image between the two most recent user input information.

[0485] The difference between the two most recent user input information can be determined by converting the user input information into latent vectors and then calculating the distance between the latent vectors.

[0486] In addition, the difference between the first 3D asset representation and the fifth 3D asset representation can be represented by the difference between the position maps and the difference between the material maps. The difference between the position maps and the difference between the material maps can also be determined by converting the image into a latent vector and determining the corresponding vector distance to determine the corresponding difference.

[0487] For example, if the text of the editing information in S902 is “Please remove the freckles on your face”, the fifth 3D asset representation generated in S903 (of course, the position map needs to be converted into a triangle mesh) is rendered into a 2D face image, and the 2D face image is represented by Image 1.

[0488] The second 3D asset representation corresponding to the first 3D asset representation generated in S202 is rendered into a 2D face image, and the 2D face image is represented by Image 2.

[0489] The picture 1 is a facial image of a middle-aged male A with wrinkles and freckles on his face and yellow skin. In the picture 1, the male A has red curly hair and thin eyebrows.

[0490] The difference between the middle-aged man B in Picture 2 and the middle-aged man A in Picture 1 is that the middle-aged man B has no freckles, but the other features of his head are exactly the same as those in Picture 1.

[0491] Exemplarily, the 3D asset generation network 101 may include a module for secondary editing of data, for example, the module may be a delta denoising score (DDS) algorithm module, or other modules that can be used for secondary editing, which are not limited here.

[0492] In a possible implementation, when the user is satisfied with the 2D face image rendered based on the fifth 3D asset representation, the system 100 can also use the fifth 3D asset representation as supervision data to calculate the loss (also called error) between the fifth 3D asset representation and the first 3D asset representation generated previously, so as to continue to optimize the 3D asset generation network 101 and realize the reverse propagation of the network. In this way, the user inputs the editing information twice, which can also facilitate the iterative update of the 3D asset generation network 101, so that it can output 3D assets of 3D digital humans that better meet the needs of users.

[0493] In the embodiment of the present application, the user's secondary creation and editing of the 3D digital human can form a data closed loop, and the editing information input by the user for the second time can be used to optimize the 3D asset representation previously output by the 3D asset generation network 101, so as to quickly obtain the 3D asset representation of the digital human that meets the user's needs.

[0494] It should be understood that Figure 2a , Figure 2b , Figure 2c , Figure 2e The system shown is only an example, and the system of the present application may have more or fewer components than those shown in the figure, may combine two or more components, or may have a different component configuration. Figure 2a , Figure 2b , Figure 2c , Figure 2e The various components shown in the EMBODIMENTS 2000 may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.

[0495] Example 10

[0496] Please refer to Figure 2f , this example 10 introduces another use (also described as reasoning) process of the face generation system 100.

[0497] Figure 2f The reasoning process of the face generation system 100 shown can be similar to the above Figure 2a , Figure 2b , Figure 2c , Figure 2e Any one of the embodiments can be combined to generate a 3D digital human.

[0498] like Figure 2f As shown, the face generation system 100 may include a 3D asset generation network 101 for a digital human.

[0499] In a possible implementation, the 3D asset generation network 101 may include a control network 106 , a Stable Diffusion, and a graph generation network.

[0500] In a possible implementation, the Stable Diffusion may not be trained, the graph generation network may be trained, and the control network 106 may be trained.

[0501] The graph generation network and the control network 106 may be trained separately or jointly to obtain the Figure 2f A face generation system 100 is shown.

[0502] In another possible implementation, the Stable Diffusion, the graph network, and the control network 106 are all trained.

[0503] Among them, Figure 2f The training process of Stable Diffusion shown above can be Figure 1aThe 3D asset generation network 101 (e.g., Stable Diffusion as shown in FIG. Figure 1a The training process of the 3D asset generation network 101 shown in FIG. Figure 2f When training with Stable Diffusion, refer to Figure 1a , Stable Diffusion as Figure 1a The 3D asset generation network 101 shown in FIG. 1 may involve the following steps during the training process: Figure 1a The optimization of S105 is shown, but the optimization of S104 is not involved.

[0504] The above-mentioned graph-generated network is a network that generates images from images, and specifically can be a network model that generates material maps from color maps, that is, the input is a color map and the output is a material map. For example, the graph-generated network is an Albedo2PBR network. Among them, Albedo is a material map used to record color information, and PBR (Physically Based Rendering) is a physically based material map, wherein PBR may include but is not limited to normal maps, specular maps, roughness maps, and displacement maps. The Albedo2PBR network can receive Albedo to generate PBR.

[0505] The control network 106 may be a 3D-aware control network, such as a 3D-aware ControlNet, wherein ControlNet is a neural network architecture that can enhance a pre-trained image diffusion model through task-specific conditions.

[0506] Combine the following Figure 2f , taking Stable Diffusion without training, and the control network 106 and the graph generation network after training as an example, the application process of the face generation system 100 of the present application is explained.

[0507] like Figure 2f As shown, the process may include the following steps:

[0508] S201, receiving user input information (at least one of text and picture).

[0509] The user input information is information indicating the generation of a digital human face, and the user input information may be at least one of text, image, voice, etc.

[0510] The user input information is used to describe the characteristics of the target virtual image to be generated (here, the head of a digital person).

[0511] For example, the user input information is the following text: "a middle-aged male with yellow skin and wrinkles and freckles on his face, with red curly hair and thin eyebrows". In other words, the user expects that the face generation system 100 can generate a face image of a digital person that meets the facial features of the above text.

[0512] For ease of explanation, the following description takes the case where the user input information is text as an example. When the user input information is voice, the voice can be converted into text for similar processing.

[0513] S2001, generating geometric data of a digital human based on the user input information.

[0514] The geometric data is geometric data that matches the head features of the digital human described by the above user input information.

[0515] The geometric data may be a triangular mesh. The geometric data is also referred to as rough geometry.

[0516] S2001 can be implemented through any network that obtains geometric data of the digital human head through user input, and there is no limitation here.

[0517] In a possible implementation, when implementing S2001, the present application may construct a 3D head candidate library, which may include not only candidate geometric data (such as geometric data of a head), but also at least one of the following three types of data associated with the geometric data:

[0518] A color real head image (referred to as reference portrait), attribute values ​​of various attributes of the head in the reference portrait, and a rendered image of the geometric data, wherein the rendered image is a 2D image.

[0519] For example, the reference portrait is a real color image of a human head with associated geometric data, such as an image taken by a mobile phone. In addition, the attributes associated with the reference portrait may include attribute values ​​of accessories such as face shape and eyebrows in the reference portrait, such as round face shape, thick eyebrows, black eyebrows, etc. The rendered image is an image (such as a gray image) obtained by rendering the geometric data. The rendered image can only express the candidate geometric data on a 2D image, and does not have attribute values ​​of various accessories in the human head.

[0520] Then, when obtaining geometric data matching the user input information (such as text) from the candidate library, the similarities between the user input information and the above three types of data can be calculated respectively. Specifically, the similarity between the user input information and the reference portrait can be calculated; and the similarity between the user input information and the rendered image can be calculated; and the similarity between the user input information and the attribute values ​​of various attributes of the head in the reference portrait can be calculated; then, the three similarities are weighted to obtain the similarity between the user input information and the candidate geometric data (i.e., the geometric data associated with the above three types of data), so that the candidate geometric data with the highest similarity to the user input information can be screened from the candidate library as the geometric data that meets the user input obtained in S2001.

[0521] Specifically, when calculating the similarity between the user input information and any of the above three types of data, the attribute values ​​of each attachment (such as wrinkles on the face, yellow skin, red curly hair, thin eyebrows, etc.) can be extracted from the user input information. For example, the attribute values ​​can be extracted from the user input information based on the dialogue large model (Large Language Model). Then, the similarity between the attribute values ​​of each attachment extracted from the user input information and each of the above three types of data is calculated. For example, when calculating the similarity between the attribute values ​​of each attachment extracted from the user input information and the attribute values ​​of each attribute of the head in the reference portrait, the similarity can be determined by calculating the distance between the attribute values ​​of the same attribute, wherein the greater the distance, the smaller the similarity. For example, when calculating the similarity between the user input information and the reference portrait, the user input information and the reference portrait can be converted into vectors, and the similarity is calculated in the vector space; similarly, when calculating the similarity between the user input information and the rendered image, the user input information and the rendered image can be converted into vectors, and the similarity is calculated in the vector space.

[0522] When the user input information is a picture, the principle of the process of obtaining the geometric data of the candidate with the highest similarity screened out from the candidate library is the same, which will not be repeated here.

[0523] In this way, when using the user input information to generate the geometric data of the human face (or human head) that matches it, the rendered image of the geometric data of the human face, the real color image with the geometric data, and at least one of the attribute values ​​of each human face attribute in the color image can be combined to find the geometric data of the human face that best matches the user input information. In this way, the rich attribute information of the human face (or human head) can be combined to guide the generation of the geometric data of the human face (or human head), so that the geometric data about the human face (or human head) obtained in S2001 is more accurate.

[0524] S2002, rendering the geometric data to obtain a texture map.

[0525] The geometric data of the human head is rendered to obtain two texture maps of the geometric data of the human head in the texture space.

[0526] The geometric data may include, but is not limited to, position information of key feature points of a human head, and may also include direction information of each point of the human head.

[0527] For example, the texture map obtained by rendering the geometric data may include a landmark map of a key point position of the face and a geometry normal map of the face.

[0528] Among them, the key point position map of the face can describe the position information of the key points of the face, and the geometric normal map of the face can describe the normal direction information of each point on the entire face.

[0529] The two texture maps describe the geometric data of the human head in texture space.

[0530] S202, the digital human 3D asset generation network 101 generates a first 3D asset representation of the target virtual image (here, the head of the digital human) that matches the features of the target virtual image to be generated based on the user input information and the texture map.

[0531] Optionally, random noise may be input to the 3D asset generation network 101 to generate the first 3D asset representation.

[0532] The user input information input into the 3D asset generation network 101 in S202 is the user input information input into the face generation system 100 in S201 .

[0533] The texture map may be used as a 3D control signal (eg, a ControlNet signal) to guide the 3D asset generation network 101 of the digital human to generate a first 3D asset representation of the digital human that matches the user input information.

[0534] In the embodiment of the present application, the texture map provides 3D geometric prior information that conforms to the user input information, which describes the geometric data of the digital human head that matches the user input information in the texture space. Then, the 3D asset generation network 101 can generate the 3D assets of the digital human under the guidance of the 3D control signal and in combination with the user input information. For example, the 3D asset generation network 101 may include a 2D Wensheng map large model (such as Stable Diffusion). The 2D Wensheng map large model has the characteristic of strong diversity of generated results. Then, combined with the above-mentioned prior 3D geometric information, robust and rich 3D digital human generation can be achieved.

[0535] In one possible implementation, Figure 2f As shown, the 3D asset generation network 101 may include a control network 106, a Stable Diffusion, and a graph generation network. During inference, the texture map obtained in S2002 may be input to the control network 106, which performs forward calculation on the texture map to obtain a latent vector and output it to the Stable Diffusion.

[0536] Among them, the control network 106 is connected to the Stable Diffusion, and the output result of the control network 106 can be used as the input of the Stable Diffusion. The output result of the control network 106 is a latent vector.

[0537] Stable Diffusion can receive user input information in S201, and perform forward calculation under the guidance of the latent vector to obtain a color map and output the color map to the graph generation network.

[0538] The graph-to-graph network can then perform a forward pass on the color map to generate the PBR.

[0539] Among them, PBR may include but is not limited to material maps such as normal maps, specular maps, roughness maps, and displacement maps.

[0540] For example, the displacement map can be understood as the difference between the position map of the 3D digital human and the rough geometry (ie, the geometric data).

[0541] Then the face generation system 100 can obtain the position map of the 3D digital human based on the above-mentioned geometric data and the displacement map.

[0542] The above-mentioned normal map, specular map, roughness map and color map output by Stable Diffusion can constitute the material map of the 3D digital human, thereby obtaining the first 3D asset representation in S202.

[0543] Finally, the face generation system 100 can output the color map of the digital person (from Stable Diffusion), the position map of the digital person, and the material map other than the color map (all from the image generation network) as the first 3D asset representation. The color map is one type of material map.

[0544] certainly, Figure 2f The outputted first 3D asset representation shown may also be obtained by, for example Figure 2b The super-resolution, attachment generation and configuration, conversion and other processes shown can ultimately render a 2D image or 2D video of the 3D digital human. The specific process is consistent with the similar scheme of the aforementioned embodiment, and will not be repeated here.

[0545] In the embodiment of the present application, the face generation system 100 can generate geometric data of the digital human for the user input information, and then render the geometric data to the texture space to obtain a texture map; then use the control network to generate a latent vector corresponding to the texture map to guide Stable Diffusion to generate a color map that matches the user input information; finally, the image generation network generates PBR for the color map, so that the face generation system can obtain a position map and a material map of the 3D digital human that matches the user input information. In this process, Stable Diffusion can generate a color map under the guidance of a 3D control signal (such as a ControlNet signal, the above two texture maps). There is no need to use a large number of samples to train Stable Diffusion. Only a small number of samples are needed to train the control network to achieve the effect of generating 3D assets of digital humans of the same quality. Of course, the more training data there is, the higher the quality of the 3D asset representation of the digital human is, so that the quality of the rendered 2D picture or 2D video is also high. For example, the control network can be trained within 1 hour, and the face generation system 100 can complete the generation of a 2D image or 2D video of a digital human within 3 seconds.

[0546] The following introduces Figure 2f The training process of each network in the face generation system 100 is shown.

[0547] The training data of this training process is Figure 1a and Figure 1cThe generation process of the training data shown is similar and will not be repeated here. The training data may include training samples and their label information, wherein the training samples may include reference position maps and reference material maps. The reference material maps may include reference color maps and other material maps except the reference color maps.

[0548] The network to be trained may include but is not limited to: Figure 2f The control network 106 shown is a graph network.

[0549] When training the graph-generated network, the reference color map can be input into the graph-generated network for forward calculation according to the conventional network training method to obtain the predicted PBR; then, the predicted PBR and the reference PBR are compared to calculate the loss, and the loss is used to update the graph-generated network using the back propagation algorithm to optimize the graph-generated network until it converges. The reference color map and the reference PBR are both from, for example Figure 1a The training data shown

[0550] When training the control network 106, for example, the training of the image-generating network is completed in advance according to the above process, and then the control network 106 is trained. The training samples may include a reference position map and a reference material map. The reference material map may include a reference color map and other material maps except the reference color map. The training process is similar to Figure 2f The reasoning process shown is similar, and the training of the control network 106 can be specifically performed by implementing the following method:

[0551] The label information and its reference color map are input into the face generation system 100. After processing the label information at S2001, geometric data can be obtained. After rendering the geometric data at S2002, a texture map can be obtained. Then, the control network 106 can generate a corresponding latent vector for the texture map at S2003. Then, Stable Diffusion generates a predicted color map based on the label information under the guidance of the latent vector. Then, the loss of the predicted color map and the reference color map is calculated, and the control network 106 is updated using a back propagation algorithm using the loss to optimize the network parameters of the control network.

[0552] Optionally, the predicted color map output by Stable Diffusion can be input into the graph-generated network, so as to obtain the predicted position map and the predicted material map other than the color map based on the PBR generated by the graph-generated network, so that the predicted 3D asset representation (specifically including the predicted position map and the predicted material map) can be obtained; then, the loss can be calculated between the predicted 3D asset representation and the reference position map and the reference material map, and then the loss can be used to update the control network 106 using the back propagation algorithm, thereby optimizing the control network 106.

[0553] In the above embodiment, the graph generation network is first trained separately, and then the control network is trained to obtain Figure 2f In other embodiments, the graph network and the control network may be jointly trained to obtain the following: Figure 2f The graph generation network and control network 106 shown in the figure have similar principles and will not be repeated here.

[0554] In this embodiment, Stable Diffusion does not need to be trained. Only a small amount of 3D asset training samples are needed to train the control network. This can generate 3D assets of digital humans with the same quality as those obtained in Example 8, greatly reducing the number of training samples and training time.

[0555] Of course, in some embodiments, the Stable Diffusion in the 3D asset generation network 101 can also be trained in combination with the quality requirements for the output digital human, for example, by Figure 1a The S105 shown is optimized, optionally, it can also be combined with Figure 1a The differentiable rendering network shown is used to perform the optimization of S104, and no limitation is made here.

[0556] In a possible implementation, an embodiment of the present application provides a data processing device. Figure 3 The schematic diagram of the structure of the data processing device 300 is shown in FIG. Figure 3The device 300 may include: a first acquisition module 301, used to acquire a feature guidance label, wherein the feature guidance label is used to describe the features of the virtual image to be generated; a control module 302, used to input the feature guidance label into the virtual image generation network; the virtual image generation network 303, used to perform forward calculation based on the feature guidance label to output a first three-dimensional asset of the virtual image to a differentiable rendering module; the differentiable rendering module 304, used to render the first three-dimensional asset to output a first image, wherein the first image is a two-dimensional image of the virtual image; a second acquisition module 305, used to acquire a plurality of second reference images, wherein the plurality of second reference images are reference two-dimensional images of the virtual image; a first training module 306, used to update the virtual image generation network using a back propagation algorithm based on the first image, the plurality of second reference images and the differentiable rendering module.

[0557] In a possible implementation, the device is arranged on a cloud data center, and a cloud management platform is used to manage an infrastructure for providing cloud services, wherein the infrastructure includes a plurality of cloud data centers arranged in different regions, and each region is provided with at least one cloud data center.

[0558] In one possible implementation, the first training module 306 is specifically used to: perform loss calculation based on the first image and the multiple second reference images to obtain a loss function; and the differentiable rendering module passes the loss function to the virtual image generation network to update the virtual image generation network using a back-propagation algorithm.

[0559] In one possible implementation, the first three-dimensional asset includes a first position map and a first material map, and the device 300 also includes: a third acquisition module, used to acquire at least one first reference image, wherein the at least one first reference image is a two-dimensional image that provides a reference for the virtual image to be generated; a first conversion module, used to convert the at least one first reference image into at least one first reference three-dimensional asset, wherein the first reference three-dimensional asset includes a first reference position map and a first reference material map, and the feature guidance label corresponds to the at least one first reference image, or the feature guidance label corresponds to the first reference position map and the first reference material map; a first loss calculation module, used to perform loss calculation on the first position map using at least one first reference position map, and perform loss calculation on the first material map using at least one first reference material map, so as to train the virtual image generation network.

[0560] In a possible implementation, the first three-dimensional asset includes a first position map and a first material map, and the device 300 also includes: a fourth acquisition module, used to acquire at least one first reference three-dimensional data, wherein the at least one first reference three-dimensional data is three-dimensional data that provides reference for the virtual image to be generated; a second conversion module, used to convert the at least one first reference three-dimensional data into at least one second reference three-dimensional asset, wherein the second reference three-dimensional asset includes a second reference position map and a second reference material map, and the feature guidance label corresponds to the at least one first reference image, or the feature guidance label corresponds to the second reference position map and the second reference material map; a second loss calculation module, used to use at least one second reference position map to perform loss calculation on the first position map, and use at least one second reference material map to perform loss calculation on the first material map, so as to train the virtual image generation network.

[0561] In a possible implementation, the device 300 also includes a super-resolution network connected to the virtual image generation network 303, and the virtual image generation network 303 is also used to output the first three-dimensional asset to the super-resolution network, wherein the resolution of the first three-dimensional asset is a first resolution; the super-resolution network is used to perform super-resolution processing on the first three-dimensional asset to obtain a second three-dimensional asset with a second resolution, wherein the second resolution is higher than the first resolution; the device 300 also includes: a second training module, used to update the super-resolution network using a back propagation algorithm based on at least one of the first reference image and the third reference image, and the second three-dimensional asset, wherein the third reference image is a two-dimensional image obtained by rendering the first reference three-dimensional data and providing a reference for the virtual image to be generated.

[0562] In a possible implementation, the device 300 further includes: a first extraction module, used to extract a first feature describing the virtual image to be generated from the first image; a second extraction module, used to extract a second feature describing the virtual image to be generated from the feature guidance label; and a third loss calculation module, used to calculate the loss between the first feature and the second feature, so as to update the virtual image generation network using a back-propagation algorithm.

[0563] In a possible implementation, the first reference three-dimensional data includes a first triangular mesh providing a geometric reference for the virtual image to be generated and a first material map providing a material reference for the virtual image to be generated, and the second conversion module is specifically used to: convert at least one of the first triangular meshes in the at least one first reference three-dimensional data into at least one second reference position map; convert at least one of the first material maps in the at least one first reference three-dimensional data into at least one second reference material map.

[0564] In a possible implementation, the device 300 further includes: a style unification processing module, configured to perform a style unification operation on the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset to obtain the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset after style unification, wherein the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset after style unification have the same accessory objects described in each of the virtual images to be generated, or have the same image parameters of the features described in each of the accessory objects.

[0565] In a possible implementation, the device 300 further includes: a blending processing module for blending at least two of the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset using blending weights to add at least one third reference three-dimensional asset, wherein the third reference three-dimensional asset includes a third reference position map and a third reference material map.

[0566] In a possible implementation, the first reference three-dimensional data also includes a second triangular mesh providing a geometric reference for objects other than the virtual image to be generated, and a second material map providing a material reference for the objects other than the virtual image to be generated, and the device 300 also includes: a first cleaning module, used to delete the second triangular mesh and the second material map of the object in the first reference three-dimensional data to obtain the cleaned first reference three-dimensional data; a first repair module, used to add a third triangular mesh and a third material map of missing attachments to the cleaned first reference three-dimensional data, wherein the missing attachments are attachments that are missing in the cleaned first reference three-dimensional data and can provide references for the virtual image to be generated; wherein the first triangular mesh includes the third triangular mesh, and the first material map includes the third material map.

[0567] In a possible implementation, the first conversion module is specifically used to: convert the at least one first reference image into second reference three-dimensional data, wherein the second reference three-dimensional data is three-dimensional data that provides a reference for the virtual image to be generated; add third reference three-dimensional data of missing attachments to the second reference three-dimensional data, wherein the missing attachments are attachments that are missing in the second reference three-dimensional data and can provide a reference for the virtual image to be generated; and convert the second reference three-dimensional data after adding the third reference three-dimensional data into a first reference three-dimensional asset.

[0568] In a possible implementation, the device 300 also includes: a first receiving module, used to receive first user input information, wherein the first user input information is used to describe the characteristics of the target virtual image to be generated; the virtual image generation network 303 is also used to generate a third three-dimensional asset of the target virtual image based on the first user input information, the third three-dimensional asset including a third position map and a third material map; a first rendering module, used to render the third three-dimensional asset to obtain a second image of the target virtual image, wherein the second image is a two-dimensional image of the target virtual image.

[0569] In a possible implementation, the device 300 also includes: a generation module, which is used to generate geometric data that matches the characteristics of the target virtual image to be generated based on the first user input information; a second rendering module, which is used to render the geometric data to obtain a texture map; the virtual image generation network 303 is specifically used to generate a third three-dimensional asset of the target virtual image that matches the characteristics of the target virtual image to be generated based on the first user input information and the texture map.

[0570] In a possible implementation, the device 300 also includes: a determination module, used to determine, based on a preset three-dimensional asset library of the virtual image, a fourth three-dimensional asset of at least one attachment that matches the characteristics of the target virtual image, the preset three-dimensional asset library including three-dimensional data of the attachment, the three-dimensional data including geometric data and material parameters, and the fourth three-dimensional asset including a fourth position map and a fourth material map; an attachment adding module, used to add the fourth three-dimensional asset of the at least one attachment to the third three-dimensional asset.

[0571] In a possible implementation, the determination module is specifically used to: determine the three-dimensional data of at least one accessory that matches the characteristics of the target virtual image based on a preset three-dimensional asset library of the virtual image; render the three-dimensional data of the at least one accessory through a micro-rendering module to obtain a third image of the at least one accessory; extract a third feature that describes the target virtual image from the first user input information; extract a fourth feature that describes the target virtual image from the third image; calculate the loss between the third feature and the fourth feature to optimize the material parameters in the three-dimensional data of the at least one accessory to obtain the optimized three-dimensional data of the at least one accessory; and convert the optimized three-dimensional data of the at least one accessory into a third three-dimensional asset of the at least one accessory that matches the characteristics described by the first user input information.

[0572] In a possible implementation, the device 300 also includes: a second receiving module, used to receive second user input information, wherein the second user input information is used to provide feature editing information of the target virtual image; the virtual image generation network 303 is also used to optimize the third three-dimensional asset based on the feature difference between the second user input information and the first user input information about the target virtual image, and generate a fourth three-dimensional asset of the target virtual image that matches the feature editing information, wherein the fourth three-dimensional asset includes a fourth position map and a fourth material map, and the difference between the fourth three-dimensional asset and the third three-dimensional asset matches the feature difference.

[0573] Among them, the above modules and networks can be implemented by software or by hardware. Among them, the modules and networks are taken as an example of software functional units, and the above modules and networks may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the first training module 306 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.

[0574] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0575] As an example of a hardware functional unit, the module may include at least one computing device, such as a server, etc. Alternatively, the module may also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0576] The multiple computing devices included in the above data processing device can be distributed in the same region or in different regions. The multiple computing devices included in the above data processing device can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the above data processing device can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0577] It should be noted that, in other embodiments, the above modules and networks can be used to execute corresponding steps in the above data processing method to realize all functions of the data processing device.

[0578] The present application also provides a computing device 900. Figure 4 As shown, the computing device 900 includes: a bus 902, a processor 904, a memory 906, and a communication interface 909. The processor 904, the memory 906, and the communication interface 909 communicate with each other through the bus 902. The computing device 900 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 900.

[0579] The bus 902 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 The bus 902 may include a path for transmitting information between various components of the computing device 900 (eg, the memory 906, the processor 904, and the communication interface 909).

[0580] The processor 904 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0581] The memory 906 may include a volatile memory, such as a random access memory (RAM). The processor 904 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0582] The memory 906 stores executable program codes, and the processor 904 executes the executable program codes to respectively implement the functions of the first acquisition module, the control module, the virtual image generation network, the differentiable rendering module, the second acquisition module, the first training module, and other modules in the above-mentioned device 300, thereby implementing the data processing method. That is, the memory 906 stores instructions for executing the data processing method.

[0583] The communication interface 909 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 900 and other devices or a communication network.

[0584] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0585] like Figure 5 As shown, the computing device cluster includes at least one computing device 1000. The memory 1006 in one or more computing devices 1000 in the computing device cluster may store the same instructions for executing the data processing method.

[0586] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 1000 may jointly execute instructions for executing the data processing method.

[0587] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the data processing device. That is, the instructions stored in the memory 1006 in different computing devices 1000 can implement the functions of one or more modules in the above-mentioned device 300, such as the first acquisition module, the control module, the virtual image generation network, the differentiable rendering module, the second acquisition module, and the first training module.

[0588] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Figure 6 A possible implementation is shown. Figure 6 As shown, two computing devices 1100A and 1100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1106 in the computing device 1100A stores instructions for executing the functions of the first acquisition module, the control module, the virtual image generation network, and the micro-rendering module. At the same time, the memory 1106 in the computing device 1100B stores instructions for executing the functions of the second acquisition module and the first training module.

[0589] It should be understood that Figure 6 The functions of the computing device 1100A shown in FIG. 1100A may also be completed by multiple computing devices 1100. Similarly, the functions of the computing device 1100B may also be completed by multiple computing devices 1100.

[0590] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Figure 4 and Figure 6 The connection mode of the computing device cluster is different in that the memory 1106 in one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the data processing method.

[0591] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 1100 may jointly execute instructions for executing the data processing method.

[0592] It should be noted that the memory 1106 in different computing devices 1100 in the computing device cluster may store different instructions for executing part of the functions of the data processing device. That is, the instructions stored in the memory 1106 in different computing devices 1100 may implement the functions of one or more devices in the data processing device.

[0593] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the data processing method in the above embodiment.

[0594] The present application embodiment also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the data processing method in the above embodiment.

[0595] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A data processing method, characterized in that: The method comprises: Acquire a feature guidance tag, where the feature guidance tag is used to describe the features of the virtual image to be generated; Inputting the feature guidance label into a virtual image generation network, and the virtual image generation network performs forward calculation to output a first three-dimensional asset of the virtual image to a differentiable rendering module; Rendering the first three-dimensional asset by the micro-rendering module to output a first image, wherein the first image is a two-dimensional image of the avatar; Acquire a plurality of second reference images, wherein the plurality of second reference images are reference two-dimensional images of the virtual image; Based on the first image, the multiple second reference images and the differentiable rendering module, the virtual image generation network is updated using a back-propagation algorithm.

2. The method according to claim 1, characterized in that The updating of the virtual image generation network by using a back propagation algorithm based on the first image, the plurality of second reference images and the differentiable rendering module comprises: Performing loss calculation based on the first image and the plurality of second reference images to obtain a loss function; The loss function is passed to the virtual image generation network by the differentiable rendering module to update the virtual image generation network using a back propagation algorithm.

3. The method according to claim 1 or 2, characterized in that: The first three-dimensional asset includes a first position map and a first material map, and the method further includes: Acquire at least one first reference image, where the at least one first reference image is a two-dimensional image that provides a reference for the virtual image to be generated; Converting the at least one first reference image into at least one first reference three-dimensional asset, wherein the first reference three-dimensional asset includes a first reference position map and a first reference material map, and the feature guidance tag corresponds to the at least one first reference image, or the feature guidance tag corresponds to the first reference position map and the first reference material map; A loss calculation is performed on the first position map using at least one of the first reference position maps, and a loss calculation is performed on the first material map using at least one of the first reference material maps to train the virtual image generation network.

4. The method according to any one of claims 1 to 3, characterized in that The first three-dimensional asset includes a first position map and a first material map, and the method further includes: Acquire at least one first reference three-dimensional data, where the at least one first reference three-dimensional data is three-dimensional data providing a reference for the virtual image to be generated; Converting the at least one first reference three-dimensional data into at least one second reference three-dimensional asset, wherein the second reference three-dimensional asset includes a second reference position map and a second reference material map, and the feature guidance label corresponds to the at least one first reference image, or the feature guidance label corresponds to the second reference position map and the second reference material map; The loss of the first position map is calculated using at least one of the second reference position maps, and the loss of the first material map is calculated using at least one of the second reference material maps to train the virtual image generation network.

5. The method according to claim 3 or 4, characterized in that: The virtual image generation network is connected to the super-resolution network, and the method further includes: Outputting the first three-dimensional asset to the super-resolution network by the avatar generation network, wherein the resolution of the first three-dimensional asset is a first resolution; The super-resolution network performs super-resolution processing on the first three-dimensional asset to obtain a second three-dimensional asset with a second resolution, wherein the second resolution is higher than the first resolution; Based on at least one of the first reference image and the third reference image, and the second three-dimensional asset, the super-resolution network is updated using a back-propagation algorithm, wherein the third reference image is a two-dimensional image rendered by the first reference three-dimensional data and provides a reference for the virtual image to be generated.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Extracting a first feature describing a virtual image to be generated from the first image; Extracting a second feature describing the virtual image to be generated from the feature guidance label; The loss between the first feature and the second feature is calculated to update the virtual image generation network using a back-propagation algorithm.

7. The method according to claim 4, characterized in that The first reference three-dimensional data includes a first triangular mesh providing a geometric reference for the virtual image to be generated and a first material map providing a material reference for the virtual image to be generated, and the converting the at least one first reference three-dimensional data into at least one second reference three-dimensional asset includes: Converting at least one of the first triangular meshes in the at least one first reference three-dimensional data into at least one second reference position map; At least one of the first texture maps in the at least one first reference three-dimensional data is converted into at least one second reference texture map.

8. The method according to claim 4, characterized in that Before using at least one of the first reference position maps to perform loss calculation on the first position map, and using at least one of the first reference material maps to perform loss calculation on the first material map, so as to train the virtual image generation network, before using at least one of the second reference position maps to perform loss calculation on the first position map, and using at least one of the second reference material maps to perform loss calculation on the first material map, so as to train the virtual image generation network, the method further includes: A style unification operation is performed on the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset to obtain the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset after style unification, wherein the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset after style unification have the same accessory objects described for the virtual image to be generated, or the same image parameters of the features described for the accessory objects.

9. The method according to claim 4, characterized in that Before using at least one of the first reference position maps to perform loss calculation on the first position map, and using at least one of the first reference material maps to perform loss calculation on the first material map, so as to train the virtual image generation network, before using at least one of the second reference position maps to perform loss calculation on the first position map, and using at least one of the second reference material maps to perform loss calculation on the first material map, so as to train the virtual image generation network, the method further includes: At least two of the at least one first reference 3D asset and the at least one second reference 3D asset are blended using blending weights to add at least one third reference 3D asset, wherein the third reference 3D asset includes a third reference position map and a third reference material map.

10. The method according to claim 7, characterized in that The first reference three-dimensional data further includes a second triangular mesh providing a geometric reference for an object other than the virtual image to be generated, and a second material map providing a material reference for the object other than the virtual image to be generated, and the method further includes: Deleting the second triangular mesh and the second material map of the object in the first reference three-dimensional data to obtain cleaned first reference three-dimensional data; Adding a third triangular mesh and a third material map of a missing attachment to the cleaned first reference three-dimensional data, wherein the missing attachment is an attachment that is missing from the cleaned first reference three-dimensional data and can provide a reference for the virtual image to be generated; The first triangular mesh includes the third triangular mesh, and the first texture map includes the third texture map.

11. The method according to claim 3, characterized in that The converting the at least one first reference image into at least one first reference three-dimensional asset comprises: Converting the at least one first reference image into second reference three-dimensional data, wherein the second reference three-dimensional data is three-dimensional data providing a reference for the virtual image to be generated; adding third reference three-dimensional data of a missing attachment to the second reference three-dimensional data, wherein the missing attachment is an attachment that is missing from the second reference three-dimensional data and can provide a reference for the virtual image to be generated; The second reference three-dimensional data after adding the third reference three-dimensional data is converted into a first reference three-dimensional asset.

12. The method according to any one of claims 1 to 11, characterized in that After the virtual image generation network is updated using a back propagation algorithm based on the first image, the plurality of second reference images and the differentiable rendering module, the method further includes: Receiving first user input information, wherein the first user input information is used to describe the characteristics of a target virtual image to be generated; The avatar generation network generates a third three-dimensional asset of the target avatar based on the first user input information, wherein the third three-dimensional asset includes a third position map and a third material map; The third three-dimensional asset is rendered to obtain a second image of the target virtual image, wherein the second image is a two-dimensional image of the target virtual image.

13. The method according to claim 12, characterized in that The method further comprises: Based on the first user input information, generating geometric data matching the features of the target virtual image to be generated; Rendering the geometric data to obtain a texture map; The step of generating, by the avatar generation network based on the first user input information, a third three-dimensional asset of the target avatar comprises: The virtual image generation network generates a third three-dimensional asset of the target virtual image that matches the features of the target virtual image to be generated based on the first user input information and the texture map.

14. The method according to claim 12 or 13, characterized in that Before rendering the third three-dimensional asset to obtain the second image of the target virtual image, the method further includes: determining, based on a preset 3D asset library of the avatar, a fourth 3D asset of at least one attachment that matches the characteristics of the target avatar, the preset 3D asset library comprising 3D data of the attachment, the 3D data comprising geometric data and material parameters, the fourth 3D asset comprising a fourth position map and a fourth material map; A fourth three-dimensional asset of the at least one attachment is added to the third three-dimensional asset.

15. The method according to claim 14, characterized in that The step of determining, based on the preset three-dimensional asset library of the virtual image, a fourth three-dimensional asset of at least one accessory that matches the feature of the target virtual image comprises: Determining, based on a preset three-dimensional asset library of the avatar, three-dimensional data of at least one accessory that matches the characteristics of the target avatar; The three-dimensional data of the at least one attachment is rendered by a micro-rendering module to obtain a third image of the at least one attachment; extracting a third feature describing the target virtual image from the first user input information; extracting a fourth feature describing the target virtual image from the third image; Calculating a loss between the third feature and the fourth feature to optimize a material parameter in the three-dimensional data of the at least one accessory to obtain optimized three-dimensional data of the at least one accessory; The optimized three-dimensional data of the at least one accessory is converted into a third three-dimensional asset of the at least one accessory that matches the features described by the first user input information.

16. The method according to any one of claims 12 to 15, characterized in that After the avatar generation network generates the third three-dimensional asset of the target avatar based on the first user input information, the method further includes: receiving second user input information, wherein the second user input information is used to provide feature editing information of the target virtual image; The virtual image generation network optimizes the third three-dimensional asset based on the feature difference between the second user input information and the first user input information regarding the target virtual image, and generates a fourth three-dimensional asset of the target virtual image that matches the feature editing information, wherein the fourth three-dimensional asset includes a fourth position map and a fourth material map, and the difference between the fourth three-dimensional asset and the third three-dimensional asset matches the feature difference.

17. A data processing device, characterized in that: The device comprises: A first acquisition module is used to acquire a feature guidance label, wherein the feature guidance label is used to describe the features of the virtual image to be generated; A control module, used for inputting the feature guidance label into a virtual image generation network; The avatar generation network is configured to perform forward computation based on the feature guidance label to output a first three-dimensional asset of the avatar to a differentiable rendering module; The micro-rendering module is configured to render the first three-dimensional asset to output a first image, wherein the first image is a two-dimensional image of the avatar; A second acquisition module, used to acquire a plurality of second reference images, wherein the plurality of second reference images are reference two-dimensional images of the virtual image; The first training module is used to update the virtual image generation network using a back propagation algorithm based on the first image, the multiple second reference images and the differentiable rendering module.

18. The device according to claim 17, characterized in that The first training module is specifically used for: Performing loss calculation based on the first image and the plurality of second reference images to obtain a loss function; The loss function is passed to the virtual image generation network by the differentiable rendering module to update the virtual image generation network using a back propagation algorithm.

19. The device according to claim 17 or 18, characterized in that The first three-dimensional asset includes a first position map and a first material map, and the apparatus further includes: A third acquisition module, used to acquire at least one first reference image, where the at least one first reference image is a two-dimensional image providing a reference for the virtual image to be generated; A first conversion module, configured to convert the at least one first reference image into at least one first reference three-dimensional asset, wherein the first reference three-dimensional asset includes a first reference position map and a first reference material map, and the feature guidance tag corresponds to the at least one first reference image, or the feature guidance tag corresponds to the first reference position map and the first reference material map; The first loss calculation module is used to use at least one of the first reference position maps to perform loss calculation on the first position map, and use at least one of the first reference material maps to perform loss calculation on the first material map, so as to train the virtual image generation network.

20. The device according to any one of claims 17 to 19, characterized in that The first three-dimensional asset includes a first position map and a first material map, and the apparatus further includes: A fourth acquisition module, used to acquire at least one first reference three-dimensional data, where the at least one first reference three-dimensional data is three-dimensional data providing reference for the virtual image to be generated; A second conversion module, configured to convert the at least one first reference three-dimensional data into at least one second reference three-dimensional asset, wherein the second reference three-dimensional asset includes a second reference position map and a second reference material map, and the feature guidance label corresponds to the at least one first reference image, or the feature guidance label corresponds to the second reference position map and the second reference material map; The second loss calculation module is used to use at least one of the second reference position maps to perform loss calculation on the first position map, and use at least one of the second reference material maps to perform loss calculation on the first material map, so as to train the virtual image generation network.

21. The device according to claim 19 or 20, characterized in that The apparatus further comprises a super-resolution network connected to the avatar generation network; The avatar generation network is further configured to output the first three-dimensional asset to the super-resolution network, wherein the resolution of the first three-dimensional asset is a first resolution; The super-resolution network is used to perform super-resolution processing on the first three-dimensional asset to obtain a second three-dimensional asset with a second resolution, wherein the second resolution is higher than the first resolution; The device also includes: The second training module is used to update the super-resolution network using a back-propagation algorithm based on at least one of the first reference image and the third reference image, and the second three-dimensional asset, wherein the third reference image is a two-dimensional image rendered by the first reference three-dimensional data and provides a reference for the virtual image to be generated.

22. The device according to any one of claims 17 to 21, characterized in that The device also includes: A first extraction module, configured to extract, from the first image, a first feature describing a virtual image to be generated; A second extraction module, used for extracting a second feature describing the virtual image to be generated from the feature guidance label; The third loss calculation module is used to calculate the loss between the first feature and the second feature to update the virtual image generation network using a back propagation algorithm.

23. The device according to claim 20, characterized in that The first reference three-dimensional data includes a first triangular mesh providing a geometric reference for the virtual image to be generated and a first material map providing a material reference for the virtual image to be generated. The second conversion module is specifically used to: Converting at least one of the first triangular meshes in the at least one first reference three-dimensional data into at least one second reference position map; At least one of the first texture maps in the at least one first reference three-dimensional data is converted into at least one second reference texture map.

24. The device according to claim 20, characterized in that The device also includes: A style unification processing module is used to perform a style unification operation on the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset to obtain the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset after style unification, wherein the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset after style unification have the same accessory objects described in the virtual image to be generated, or the same image parameters of the features described in the accessory objects.

25. The device according to claim 20, characterized in that The device also includes: A blending processing module is used to blend at least two reference three-dimensional assets among the at least one first reference three-dimensional asset and the at least one second reference three-dimensional asset using blending weights to add at least one third reference three-dimensional asset, wherein the third reference three-dimensional asset includes a third reference position map and a third reference material map.

26. The device according to claim 23, characterized in that The first reference three-dimensional data further includes a second triangular mesh providing a geometric reference for an object other than the virtual image to be generated, and a second material map providing a material reference for the object other than the virtual image to be generated, and the device further includes: A first cleaning module, used for deleting the second triangular mesh and the second material map of the object in the first reference three-dimensional data to obtain cleaned first reference three-dimensional data; A first repair module, configured to add a third triangular mesh and a third material map of a missing attachment to the cleaned first reference three-dimensional data, wherein the missing attachment is an attachment that is missing from the cleaned first reference three-dimensional data and can provide a reference for the virtual image to be generated; The first triangular mesh includes the third triangular mesh, and the first texture map includes the third texture map.

27. The device according to claim 19, characterized in that The first conversion module is specifically used to: Converting the at least one first reference image into second reference three-dimensional data, wherein the second reference three-dimensional data is three-dimensional data providing a reference for the virtual image to be generated; adding third reference three-dimensional data of a missing attachment to the second reference three-dimensional data, wherein the missing attachment is an attachment that is missing from the second reference three-dimensional data and can provide a reference for the virtual image to be generated; The second reference three-dimensional data after adding the third reference three-dimensional data is converted into a first reference three-dimensional asset.

28. The device according to any one of claims 17 to 27, characterized in that The device also includes: A first receiving module, configured to receive first user input information, wherein the first user input information is used to describe the characteristics of a target virtual image to be generated; The avatar generation network is further used to generate a third three-dimensional asset of the target avatar based on the first user input information, wherein the third three-dimensional asset includes a third position map and a third material map; The first rendering module is used to render the third three-dimensional asset to obtain a second image of the target virtual image, wherein the second image is a two-dimensional image of the target virtual image.

29. The device according to claim 28, characterized in that The device also includes: A generating module, configured to generate geometric data matching the features of the target virtual image to be generated based on the first user input information; A second rendering module, used for rendering the geometric data to obtain a texture map; The virtual image generation network is specifically used to generate a third three-dimensional asset of the target virtual image that matches the features of the target virtual image to be generated based on the first user input information and the texture map.

30. The device according to claim 28 or 29, characterized in that The device also includes: a determination module, configured to determine, based on a preset three-dimensional asset library of the avatar, a fourth three-dimensional asset of at least one attachment that matches the characteristics of the target avatar, wherein the preset three-dimensional asset library includes three-dimensional data of the attachment, the three-dimensional data includes geometric data and material parameters, and the fourth three-dimensional asset includes a fourth position map and a fourth material map; The attachment adding module is used to add a fourth three-dimensional asset of the at least one attachment to the third three-dimensional asset.

31. The device according to claim 30, characterized in that The determining module is specifically used for: Determining, based on a preset three-dimensional asset library of the avatar, three-dimensional data of at least one accessory that matches the characteristics of the target avatar; Rendering the three-dimensional data of the at least one attachment through a micro-rendering module to obtain a third image of the at least one attachment; extracting a third feature describing the target virtual image from the first user input information; extracting a fourth feature describing the target virtual image from the third image; Calculating a loss between the third feature and the fourth feature to optimize a material parameter in the three-dimensional data of the at least one accessory to obtain optimized three-dimensional data of the at least one accessory; The optimized three-dimensional data of the at least one accessory is converted into a third three-dimensional asset of the at least one accessory that matches the features described by the first user input information.

32. The device according to any one of claims 28 to 31, characterized in that The device also includes: A second receiving module, configured to receive second user input information, wherein the second user input information is used to provide feature editing information of the target virtual image; The virtual image generation network is further used to optimize the third three-dimensional asset based on the feature difference between the second user input information and the first user input information about the target virtual image, and generate a fourth three-dimensional asset of the target virtual image that matches the feature editing information, wherein the fourth three-dimensional asset includes a fourth position map and a fourth material map, and the difference between the fourth three-dimensional asset and the third three-dimensional asset matches the feature difference.

33. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 16.

34. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 16.

35. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 16.