Construction method and device of texture data set, equipment, medium and program product
UV texture images are generated and rendered through text UV texture image models, which solves the problem of low precision in texture data sets caused by insufficient three-dimensional information in the prior art, and achieves fast and efficient generation of UV texture images and texture data sets, which is suitable for a variety of application scenarios.
Patent Information
- Application Number
- CN202410141278.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-08-01
AI Technical Summary
When building UV texture data sets, the lack of relying on three-dimensional information leads to low accuracy of the texture data set. Especially when there is insufficient three-dimensional scan and image or video, it is difficult to generate high-quality texture data sets.
By obtaining the UV description text of the target character, using the text UV texture image model to generate UV texture images, and render it, building a texture data set, getting rid of the dependence on three-dimensional data, using the text UV texture image model (T2UV) to generate UV texture images, and combining with the 3D rendering application to generate rendered images.
It realizes the rapid generation of UV texture images and texture data sets without three-dimensional data, improves generation speed and accuracy, reduces processing costs, is suitable for application scenarios such as digital entertainment, AR, VR, MR, etc., and improves the accuracy of neural network models.
Smart Images

Figure CN120411269A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and particularly to a method, device, equipment, medium and program product for constructing a texture data set. Background Art
[0002] A UV texture image can map a two-dimensional image onto the surface of a three-dimensional model, and is usually used to add textures to the three-dimensional model. The texture can be at least one of skin texture, graphics, patterns, colors, shadows, and lighting. UV texture images have been widely used in application scenarios such as digital entertainment, Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR).
[0003] In related technologies, the methods for generating a texture data set based on UV texture images include two categories: constructing a texture data set based on three-dimensional scanning. For example, through a three-dimensional scanning device, the three-dimensional information of multiple real three-dimensional objects is scanned, and the UV texture images of the multiple real three-dimensional objects are generated according to the three-dimensional information, and the texture data set is obtained by rendering the UV texture images. Or, constructing a texture data set based on image or video reconstruction. For example, according to the three-dimensional information of multiple real three-dimensional objects in an image or video, a UV texture image is reconstructed, and the texture data set is obtained by rendering the UV texture image.
[0004] However, related technologies rely on three-dimensional information. In the case where the three-dimensional information included in three-dimensional scanning, images or videos is insufficient, the accuracy of the constructed texture data set is not high. Summary of the Invention
[0005] The present application provides a method, device, equipment, medium and program product for constructing a texture data set. The technical solutions are as follows:
[0006] On the one hand, a method for constructing a texture data set is provided. The method includes:
[0007] Obtain a UV description text corresponding to a target character, where the UV description text is used to describe the content of the UV texture image corresponding to the target character;
[0008] Input the UV description text into a text UV texture image model to generate a UV texture image corresponding to the UV description text;
[0009] Render the UV texture image to obtain a rendered image corresponding to the target character. The rendered image includes the texture information of the target character, and the rendered image is used to construct the texture data set.
[0010] On the other hand, a device for constructing a texture dataset is provided, and the device includes:
[0011] An acquisition module, configured to acquire a UV description text corresponding to a target character, where the UV description text is used to describe the content of a UV texture image corresponding to the target character;
[0012] A processing module, configured to input the UV description text into a text UV texture image model to generate a UV texture image corresponding to the UV description text;
[0013] A rendering module, configured to render the UV texture image to obtain a rendered image corresponding to the target character, where the rendered image includes texture information of the target character, and the rendered image is used to construct the texture dataset.
[0014] On the other hand, a computer device is provided, and the computer device includes: a processor and a memory, where the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the above-mentioned method for constructing a texture dataset.
[0015] On the other hand, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the above-mentioned method for constructing a texture dataset.
[0016] On the other hand, a computer program product is provided, and the computer program product includes computer instructions, where the computer instructions are stored in a computer-readable storage medium, and a processor obtains the computer instructions from the computer-readable storage medium, so that the processor loads and executes to implement the above-mentioned method for constructing a texture dataset.
[0017] The beneficial effects brought by the technical solution provided in the embodiments of the present application at least include:
[0018] The computer device obtains the UV description text corresponding to the target character, and the UV description text is used to describe the content of the UV texture image corresponding to the target character; the UV description text is input into the text UV texture image model to generate a UV texture image corresponding to the UV description text; the UV texture image is rendered to obtain a rendered image corresponding to the target character, and the rendered image contains the texture information of the target character, and the rendered image is used to construct a texture dataset. Since the text UV texture image model can generate a corresponding UV texture image based on the UV description text, this solution can quickly generate a UV texture image without obtaining the three-dimensional data of the target character, getting rid of the dependence on the three-dimensional data of the target character. Since the three-dimensional data is difficult to obtain, costly, and time-consuming and laborious to process, and this solution does not require the use of three-dimensional data, thereby improving the generation speed of the UV texture image and reducing the processing cost. Further, by rendering the UV texture image, a texture dataset can be constructed, and while improving the generation speed of the UV texture image, the construction speed of the texture dataset is also improved. The texture dataset constructed by this solution can be applied in application scenarios such as digital entertainment, AR, VR, and MR to construct a more realistic humanoid target character in the above application scenarios. It can also be applied in the training process of some neural network models related to data processing of texture information to improve the accuracy of such neural network models. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 Shows a block diagram of the structure of a computer system provided by an exemplary embodiment;
[0021] Figure 2 Shows a schematic diagram of a method for constructing a texture dataset provided by an exemplary embodiment;
[0022] Figure 3 Shows a flowchart of a method for constructing a texture dataset provided by an exemplary embodiment;
[0023] Figure 4 Shows a schematic diagram of a text UV texture image model provided by an exemplary embodiment;
[0024] Figure 5 Shows a schematic diagram of training a text UV texture image model provided by an exemplary embodiment;
[0025] Figure 6The flowchart of the construction method of the texture data set provided by an exemplary embodiment is shown;
[0026] Figure 7 The schematic diagram of obtaining the real UV texture image provided by an exemplary embodiment is shown;
[0027] Figure 8 The schematic diagram of obtaining a plurality of view images provided by an exemplary embodiment is shown;
[0028] Figure 9 The flowchart of the construction method of the texture data set provided by an exemplary embodiment is shown;
[0029] Figure 10 The schematic diagram of obtaining the UV description text provided by an exemplary embodiment is shown;
[0030] Figure 11 The schematic diagram of obtaining the UV texture image provided by an exemplary embodiment is shown;
[0031] Figure 12 The schematic diagram of the construction method of the texture data set provided by an exemplary embodiment is shown;
[0032] Figure 13 The schematic diagram of rendering the UV texture image provided by an exemplary embodiment is shown;
[0033] Figure 14 The schematic diagram of the construction method of the texture data set provided by an exemplary embodiment is shown;
[0034] Figure 15 The schematic diagram of the construction method of the texture data set provided by an exemplary embodiment is shown;
[0035] Figure 16 The block diagram of the construction device of the texture data set provided by an exemplary embodiment is shown;
[0036] Figure 17 The structural block diagram of a computer device provided by an exemplary embodiment is shown. Detailed implementation manners
[0037] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0038] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0039] The terms used in this application are for the purpose of describing particular embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0040] It should be understood that although the terms first, second, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first parameter may also be referred to as the second parameter, and similarly, the second parameter may also be referred to as the first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0041] It should be noted that before and during the collection of relevant data of the user in this application (for example, sample video data related to the sample character, sample text, real UV texture image, bone image, several view images, global rotation information, joint pose information, and UV description text, UV texture image related to the target character, etc.), a prompt interface, pop-up window or voice prompt information can be displayed. The prompt interface, pop-up window or voice prompt information is used to prompt the user that their relevant data is being collected at present, so that this application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the confirmation operation of the user on the prompt interface or pop-up window. Otherwise (that is, when the confirmation operation of the user on the prompt interface or pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are ended, that is, the relevant data of the user is not obtained. In other words, all user data collected in this application is collected with the consent and authorization of the user, and the collection, use and processing of the relevant user data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0042] First, a brief introduction to the nouns involved in the embodiments of this application is given:
[0043] Character: A general term for movable objects. Optionally, a character can be an movable object in the real world, such as a real human body. Alternatively, a character can be an AI-powered digital movable object constructed in the real world, such as a virtual human body. Alternatively, a character can be an movable object in the virtual world, such as a virtual person or virtual animal. A character is a three-dimensional model with its own unique shape and volume, occupying a portion of space in the real or virtual world.
[0044] Texture: refers to the pattern wrapped around the surface of a 3D model, such as at least one of wood grain, cloth grain, skin, graphics, pattern, color, shadow, and lighting.
[0045] UV: Also known as mapping coordinates, they are used to position textures. UVs allow you to map points in the 3D model space to pixels on the texture image.
[0046] UV texture images, also known as UV maps, are images that include UV and texture information. UV texture images can map a 2D image onto the surface of a 3D model and are commonly used to add texture to 3D models. UV texture images are widely used in digital entertainment, augmented reality (AR), virtual reality (VR), and mixed reality (MR).
[0047] The Text-to-UV Module (T2UV) is a neural network model that generates a corresponding UV texture image based on UV description text. The UV description text is used to describe the content of the UV texture image of the target character, including at least one of the following information: the target character's name, address, appearance, gender, clothing, hairstyle, age, category, and personality.
[0048] Skinned Multi-Person Linear Model (SMPL): is a parametric model used to generate the surface shape and posture of the human body. The SMPL model takes into account skeleton points and skin information and can generate a highly realistic human surface model based on a small number of posture parameters and shape parameters.
[0049] Skinned Multi-Person Linear-X Model (SMPL eXpressive): A parametric model based on the SMPL model with the addition of face and hands.
[0050] Texture dataset: Also known as the ArTicuLated humAntextureS (ATLAS), a high-resolution 3D human texture dataset, is constructed by the method of this embodiment and contains 1100 high-fidelity human texture data, which are called rendered images in this embodiment. Each human texture data has corresponding texture information, which includes at least one of the following: UV description text, SMPL-X animation sequence, background segmentation mask, and synthetic rendering video.
[0051] Artificial Intelligence (AI): It is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0052] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operating / interactive systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the base model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0053] Figure 1 The block diagram of a computer system 100 provided by an exemplary embodiment of the present application is shown. The computer system 100 can be implemented as the system architecture of the method for constructing a texture dataset. The computer system 100 includes: a terminal 120 and a server 140.
[0054] The terminal 120 may be an electronic device such as a mobile phone, a tablet computer, an in-vehicle terminal (car computer), a wearable device, a PC (Personal Computer), an unattended reservation terminal, etc. A text UV texture image model may be stored in the terminal 120. The text UV texture image model is used to generate a UV texture image corresponding to the UV description text according to the UV description text. A client for running a target application may also be installed and run in the terminal 120. The target application may be an application for rendering a UV texture image or other applications providing a rendering function. This application is not limited in this application. In addition, the form of the target application is not limited in this application, including but not limited to an App (Application) installed in the terminal 120, a mini-program, etc., and may also be in the form of a web page.
[0055] The server 140 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, a cloud database, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, a content delivery network (Content Delivery Network, CDN), and a cloud server for basic cloud computing services such as a big data and artificial intelligence platform. A text UV texture image model may be stored in the server 140. The text UV texture image model is used to generate a UV texture image corresponding to the UV description text according to the UV description text. The server 140 may be a background server of the target application in the terminal 120 and is used to provide background services for the client of the target application.
[0056] Among them, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, and can form a resource pool, which can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system support, which can only be achieved through cloud computing.
[0057] In some embodiments, the server 140 can also be implemented as a node in a blockchain system. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Essentially, a blockchain is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer.
[0058] Communication can be carried out between the terminal 120 and the server 140 through a network, such as a wired or wireless network.
[0059] In the method for constructing the texture dataset provided in the embodiments of the present application, the execution subject of each step can be a computer device, and a computer device refers to an electronic device with data computing, processing, and storage capabilities. Taking Figure 1 the illustrated implementation environment as an example, the method for constructing the texture dataset can be executed by the terminal 120 (for example, executed by the client of the target application installed and running in the terminal 120), or can be executed by the server 140, or can be executed by the interaction and cooperation between the terminal 120 and the server 140. The present application does not make any limitation in this regard.
[0060] Those skilled in the art can know that the number of terminals 120 can be more or less. For example, there can be only one terminal 120, or dozens or hundreds of terminals 120, or a larger number. The embodiments of the present application do not limit the number and device type of the terminals 120.
[0061] Next, taking the three-dimensional model as a three-dimensional human body model and constructing a three-dimensional human body texture dataset as an example, a brief introduction to the related technologies involved in the embodiments of the present application is given:
[0062] In the process of generating a three-dimensional human body model, human body texture is a key component. The realism of the texture directly determines the realism of the generated model. However, three-dimensional human body models have always been relatively complex and difficult digital assets in the industrial production field. Obtaining high-quality human body textures is a challenging and time-consuming task, and texture datasets are very scarce.
[0063] In related technologies, the methods related to the construction of texture datasets are mainly divided into two categories: constructing texture datasets based on three-dimensional scanning, or constructing texture datasets based on image or video reconstruction.
[0064] 1. Construction of texture dataset based on 3D scanning: Although the 3D data obtained by 3D scanning devices has high precision, the acquisition process is difficult and time-consuming. For researchers, there are two ways to obtain 3D data: one is to purchase a 3D scanning device (e.g., digital single-lens reflex camera structure group, camera array) and then perform scanning, and the other is to directly purchase digital human assets. However, 3D scanning highly depends on hardware devices. At the same time, the 3D data obtained by 3D scanning usually has the characteristics of excessive vertices and unstructured meshes. Without additional processing of this 3D data, it is difficult to directly obtain a human texture image with a reasonable UV layout from this 3D data.
[0065] 2. Construction of texture dataset based on image or video reconstruction: Use a neural network model and 3D prior data to reconstruct a 3D human body from an image or video and extract the texture. However, the videos are usually real human A-Pose rotation videos, which usually do not include any 3D information. The work of reconstructing a 3D human body from a single image has the problem of difficult image acquisition. Only a very small number of image datasets contain UV textures, and the 3D information in most images is very insufficient.
[0066] Based on this, the embodiment of the present application provides a method for constructing a texture dataset. When constructing the texture dataset, this method can get rid of the dependence on 3D data and directly convert the production task of 3D human body texture into the generation task of ordinary 2D images. Moreover, since the rendered images are synthesized in the rendering application of the computer device, corresponding 3D human body shapes, UV textures, segmentation maps, and UV text descriptions can be obtained for each frame of the rendered image. The finally constructed texture dataset can be fully applicable to most graphics rendering applications in the computer device and is also suitable for application in the industrial production process.
[0067] Figure 2 The figure shows a schematic diagram of the method for constructing a texture dataset provided by an exemplary embodiment of the present application. This method is executed by a computer device, and the computer device is Figure 1 the shown terminal 120 and / or server 140. Please refer to Figure 2 , and take the computer device being the server 140 as an example for illustration. The server 140 stores a text UV texture image model (T2UV). The steps executed by the server 140 are briefly described as follows:
[0068] Step 1: The server 140 obtains a UV description text 141 corresponding to the target character, and this UV description text 141 is used to describe the content of the UV texture image corresponding to the target character.
[0069] Optionally, the UV description text 141 can be sent to the server 140 through the terminal 120. The UV description text 141 includes at least one of the following information: the name, residence, appearance, gender, clothing, hairstyle, age, category, and personality of the target character.
[0070] Reference Figure 2 Taking the example of, the target character UV description text 141 includes the following content: "1. A humanoid in a standing form, 2 meters tall, with a muscular build; 2. The hair is golden, with a goatee, the face shape is oval, and the eyebrows are black; 3. Wearing metal armor, there are circular protective decorations at the joints of the armor, and wearing a belt."
[0071] Step 2: The server 140 inputs the UV description text 141 into the text UV texture image model 200 to generate a UV texture image 142 corresponding to the UV description text 141.
[0072] Optionally, the text UV texture image model (T2UV) 200 is used to generate a corresponding UV texture image 142 according to the UV description text 141. The server 140 inputs the UV description text 141 into the text UV texture image model 200, and through the text UV texture image model 200, a UV texture image 142 corresponding to the UV description text 141 can be generated.
[0073] It should also be noted that the text UV texture image model 200 can be trained by the server 140 before executing Step 2, or it can be obtained from another server ( Figure 2 (not shown) a pre-trained text UV texture image model 200. The model structure and training process of the text UV texture image model 200 will be described in detail in the subsequent embodiments.
[0074] Step 3: The server 140 renders the UV texture image 142 to obtain a rendered image 143 corresponding to the target character. The rendered image 143 contains the texture information of the target character, and the rendered image 143 is used to construct a texture dataset.
[0075] Optionally, the server 140 installs and runs a 3D rendering application, which is used to provide a rendering function. The server 140 renders the UV texture image 142 through the 3D rendering program to obtain a rendered image 143 corresponding to the target character. The rendered image 143 contains the texture information of the target character, and the rendered image 143 is used to construct a texture dataset. The texture information includes at least one of the following: UV description text, SMPL-X animation sequence, background segmentation mask, and composite rendering video.
[0076] In summary, in the method provided by the embodiment of the present application, a computer device obtains a UV description text corresponding to a target role, where the UV description text is used to describe the content of the UV texture image corresponding to the target role; inputs the UV description text into a text-UV texture image model to generate a UV texture image corresponding to the UV description text; renders the UV texture image to obtain a rendered image corresponding to the target role, where the rendered image includes the texture information of the target role, and the rendered image is used to construct a texture dataset. Since the text-UV texture image model can generate a corresponding UV texture image based on the UV description text, this solution can quickly generate a UV texture image without obtaining the three-dimensional data of the target role, getting rid of the dependence on the three-dimensional data of the target role. Since it is difficult to obtain three-dimensional data, the cost is high, and the processing process is time-consuming and laborious, and this solution does not require the use of three-dimensional data, thereby improving the generation speed of the UV texture image and reducing the processing cost. Further, by rendering the UV texture image, a texture dataset can be constructed, which improves the construction speed of the texture dataset while improving the generation speed of the UV texture image.
[0077] Next, a method for constructing a texture dataset provided by the embodiment of the present application will be introduced.
[0078] Figure 3 FIG. shows a flowchart of a method for constructing a texture dataset provided by an exemplary embodiment of the present application. Taking this method as being applied to a computer device, the computer device may be Figure 1 illustrated by the terminal 120 and / or the server 140 shown, and the method includes step 220, step 240, and step 260:
[0079] Step 220, obtain a UV description text corresponding to a target role, where the UV description text is used to describe the content of the UV texture image corresponding to the target role.
[0080] A role is a general term for an active object. Optionally, the role may be an active object in the real world. For example, the role is a real human body. Or, the role may also be an AI digital active object constructed in the real world. For example, the role is a virtual digital human body. Or, the role may also be an active object in the virtual world. For example, the role is a virtual character or a virtual animal. The role is a three-dimensional model with its own shape and volume, occupying a part of the space in the real world or the virtual world. In one example, the target role is a real human body or a digital human body, and the UV texture image corresponding to the target role is a human body UV texture image.
[0081] The UV description text is text used to describe the content of the UV texture image corresponding to the target role, including at least one of the following information: the name, residence, appearance, gender, clothing, hairstyle, age, category, and personality of the target role.
[0082] In some embodiments, the UV description text can be text with any sentence structure, and the any sentence structure can be any one of common simple sentence structures, compound sentence structures, and complex sentence structures. Alternatively, the UV description text can also be text with a fixed sentence structure.
[0083] As an example, the fixed sentence structure can be any one of the following fixed sentence structures: Fixed sentence structure 1: (residence)(appearance)
gender
clothing
hair style
age
name
clothing
name
clothing
hair style
category
[0084] Specifically, the computer device obtains the UV description text corresponding to the target character, and the UV description text is used to describe the content of the UV texture image corresponding to the target character. Next, by inputting the UV description text into the text UV texture image model, the computer device can obtain the UV texture image.
[0085] Step 240: Input the UV description text into the text UV texture image model to generate a UV texture image corresponding to the UV description text.
[0086] The text UV texture image model (T2UV) is a pre-trained neural network model for generating a corresponding UV texture image according to the UV description text. The UV texture image can describe at least one of the skin texture, graphics, patterns, colors, shadows, and lighting on the surface of the three-dimensional model of the target character. When mapping the UV texture image to the surface of the three-dimensional model of the target character, the target character becomes more realistic.
[0087] The model structure and training process of the text UV texture image model will be described in detail in subsequent embodiments. In this step, the computer device can directly use the trained text UV texture image model. It should be noted that the text UV texture image model can be trained by the computer device before executing this step, or the computer device can obtain the pre-trained text UV texture image model from other computer devices before executing this step.
[0088] Specifically, the computer device inputs the UV description text into the text UV texture image model, and through the text UV texture image model, generates a UV texture image corresponding to the UV description text.
[0089] Step 260: Render the UV texture image to obtain a rendered image corresponding to the target character. The rendered image contains the texture information of the target character, and the rendered image is used to construct a texture dataset.
[0090] Rendering refers to the process of combining the UV texture image with rendering information such as a 3D model, light source, material, and camera to generate a final 2D rendered image. Rendering enables the UV texture image to be presented, showing a more realistic target character under near-natural conditions.
[0091] Exemplarily, the computer device is installed with and runs a 3D rendering application, which is used to provide a rendering function. For example, the 3D rendering program can be the Unreal Engine. Specifically, the computer device renders the UV texture image through the 3D rendering program to obtain a rendered image corresponding to the target character.
[0092] Optionally, the rendered image contains the texture information of the target character, and the texture information includes at least one of: a UV description text, an SMPL-X animation sequence, a background segmentation mask, and a composite rendered video. The rendered image is used to construct a texture dataset.
[0093] Exemplarily, the texture dataset can be applied in application scenarios such as digital entertainment, AR, VR, and MR to construct a more realistic humanoid target character in the above application scenarios. In other examples, the texture dataset can be applied in the training process of some neural network models related to data processing of texture information to improve the accuracy of such neural network models. For example, the neural network model is an Image-To-UV Module (I2UV), which is a neural network model that specifically generates a corresponding UV texture image according to a 2D image. Specifically, the texture dataset can be used as sample data in the training process of the Image-To-UV Module to improve the accuracy of the Image-To-UV Module.
[0094] In some embodiments, for a given image containing a target character, by using the method provided in the embodiments of the present application, the texture information corresponding to the given image can also be extracted. Then, based on the texture information of the given image, a virtual image, a cartoon image, or an animated image based on the target character can be further generated, and an animation can be generated based on the virtual image, the cartoon image, or the animated image.
[0095] In summary, for the method provided in the embodiment of the present application, the computer device obtains the UV description text corresponding to the target character, and the UV description text is used to describe the content of the UV texture image corresponding to the target character; inputs the UV description text into the text-UV texture image model to generate a UV texture image corresponding to the UV description text; renders the UV texture image to obtain a rendered image corresponding to the target character, where the rendered image includes the texture information of the target character, and the rendered image is used to construct a texture dataset. Since the text-UV texture image model can generate a corresponding UV texture image based on the UV description text, this solution can quickly generate a UV texture image without obtaining the three-dimensional data of the target character, getting rid of the dependence on the three-dimensional data of the target character. Since it is difficult to obtain three-dimensional data, the cost is high, and the processing process is time-consuming and laborious, and this solution does not require the use of three-dimensional data, thereby improving the generation speed of the UV texture image and reducing the processing cost. Further, by rendering the UV texture image, a texture dataset can be constructed, which improves the construction speed of the texture dataset while improving the generation speed of the UV texture image.
[0096] · Training and Use of Text-UV Texture Image Model (T2UV)
[0097] 1. Model Structure
[0098] Figure 4 The figure shows a schematic diagram of a text-UV texture image model 200 provided by an exemplary embodiment of the present application. The text-UV texture image model 200 includes: a text encoder 10 and an image generation network 50. The output end of the text encoder 10 is connected to the input end of the image generation network 50. The text encoder 10 is used to encode the sample text into sample text features, and the image generation network 50 is used to map to obtain the predicted noise of the predicted UV texture image corresponding to the sample text and obtain the predicted UV texture image based on the predicted noise.
[0099] Optionally, the type of the image generation network 50 can be at least one of a convolutional neural network (CNN), a generative adversarial network (GANs), a variational auto-encoder (VAEs), a diffusion model (DM), and a latent diffusion model (LDM). In this embodiment, the image generation network 50 can be specifically set as a U-Net network, which is a type of LDM. The U-Net network can decode and generate a corresponding predicted UV texture image according to the input features.
[0100] In some embodiments, the image generation network 50 of the text UV texture image model 200 is obtained by adding trainable model parameters 51 to a pre-trained image generation network. The trainable model parameters 51 refer to the newly added model parameters that need to be trained during the training phase of the text UV texture image model 200. Moreover, during the training phase of the text UV texture image model 200, all other model parameters in the pre-trained image generation network except for the trainable model parameters remain unchanged.
[0101] Optionally, in the image generation network 50, the trainable model parameters 51 are used as a parallel bypass branch and added to the original feature processing layer of the pre-trained image generation network. The input features are respectively input into the parallel bypass branch and the original feature processing layer, and the two processing results are added during output. Alternatively, the newly added processing layer (LoRA layer) includes the trainable parameters 51, and the newly added processing layer is added to the original feature processing layer of the pre-trained image generation network. Accordingly, during the training phase of the text UV texture image model 200, by using sample texts to train the trainable model parameters 51, not only can the texture adaptive fine-tuning of the pre-trained image generation network be achieved, but also the training speed can be increased and the computational amount can be reduced.
[0102] Exemplarily, the weight matrix of the model parameters of the image generation network 50 has the following representation:
[0103] w φunet +ΔW = W φunet +BA
[0104] Wherein, the weight matrix of the model parameters φ t-enc of the text encoder 10 is represented as W φt-enc , the weight matrix of the model parameters φ unet of the pre-trained image generation network is represented as W φunet ∈R d×k , the weight matrix of the newly added trainable model parameters is represented as ΔW, and ΔW = BA. B ∈ R d×r , A ∈ R r×k , r << min(d, k). The trainable model parameters are in the weight matrix A and the weight matrix B. The weight matrix W φunet of the pre-trained image generation network has a shape of d × k, the weight matrix A has a shape of r × k, the weight matrix B has a shape of d × r, r is the rank, and the rank r is a preset value and is much smaller than the minimum of d and k. For example, the preset value can be optionally set to one of 1, 2, 4, 8.
[0105] During the training process, for the input feature s input to the image generation network 50, The forward propagation formula in the image generation network 50 is represented as follows:
[0106]
[0107] Among them, the initialization method of the trainable model parameters is as follows: randomly Gaussian initialize A, initialize B to zero, and at the beginning of training, ΔW = BA = 0; through scaling ΔW and W φtenc-unet s, α is a constant in r. Through appropriate scaling initialization, the method of adjusting the constant α is basically the same as the method of adjusting the learning rate. During the training process, this scaling method is beneficial to reducing the number of times of readjusting hyperparameters and can improve the training speed.
[0108] In this embodiment, the text UV texture image model includes a text encoder and an image generation network. Among them, the image generation network is obtained by adding trainable model parameters to a pre-trained image generation network. The number of parameters of the trainable model parameters is much smaller than the model parameters of the pre-trained image generation network. Therefore, during the training process of the text UV texture image model, only the model parameters of the text encoder and the trainable model parameters of the image generation network need to be adjusted, which can achieve adaptive fine-tuning of the model parameters, reduce the amount of training data, and improve the training speed of the text UV texture image model.
[0109] 2. Training process of the text UV texture image model (T2UV)
[0110] In some embodiments, the text UV texture image model is trained by a computer device. Figure 6 The text UV texture image model provided by an exemplary embodiment of the present application is shown. The text UV texture image model is obtained through the following training method, including: step 320, step 340, and step 360:
[0111] Step 320, obtain a sample text, the sample text corresponds to a real UV texture image of a sample role, the sample text is used to describe the content of the real UV texture image, and the real UV texture image is drawn based on a plurality of view images.
[0112] The sample text refers to the UV description text used in the training stage of the text UV texture image model.
[0113] The sample role refers to the sample role described by the sample text.
[0114] The real UV texture image is the UV texture image of the sample role. In the training stage of the text UV texture image model, the real UV texture image is used as the UV texture image truth value of the sample text.
[0115] The real UV texture image is drawn based on several view images of the sample character. Among them, the view image is an image that can reflect the shape and geometric structure of the sample character. For example, when an observer observes the same spatial geometric body from three different angles: above, left, and front, the drawn images are called the top view, left view, and front view respectively.
[0116] A single view image can only reflect the shape and combined structure of one orientation of the sample character, and cannot completely represent the overall shape and geometric structure of the sample character. By combining several view images of the sample character, the overall shape and geometric structure of the sample character can be completely represented. In this embodiment, several view images are images of the sample character from different perspectives. Taking the front view angle of the sample character as 0°, the perspectives include at least one of the following: 0, ±45°, ±90°, ±135°, 180°.
[0117] Specifically, the computer device obtains a sample text, and the sample text corresponds to the real UV texture image of the sample character. The sample text is used to describe the content of the real UV texture image, and the real UV texture image is drawn based on several view images.
[0118] Step 340: Input the sample text into the text UV texture image model to generate a predicted UV texture image corresponding to the sample text.
[0119] The text UV texture image model in this step refers to the text UV texture image model to be trained. Specifically, the computer device inputs the sample text into the text UV texture image model to generate a predicted UV texture image corresponding to the sample text.
[0120] Step 360: Use reducing the error between the real UV texture image and the predicted UV texture image as the training objective to train the model parameters of the text UV texture image model.
[0121] Specifically, the computer device uses reducing the error between the real UV texture image and the predicted UV texture image as the training objective to train the model parameters of the text UV texture image model. Among them, the model parameters include the model parameters of the text encoder 10 and the trainable model parameters in the image generation network 50.
[0122] In this embodiment, using reducing the error between the real UV texture image corresponding to the sample text and the predicted UV texture image as the training objective to train the model parameters of the text UV texture image model can make the accuracy of the text UV texture image model higher, improve the accuracy of the UV texture image generated by the text UV texture image model, and is beneficial to the subsequent construction of the texture dataset.
[0123] 2.1. Drawing the real UV texture image
[0124] In some embodiments, the sample text corresponds to a real UV texture image, which is drawn based on a plurality of view images. Next, the drawing process of the real UV texture image will be described. Optionally, the real UV texture image is obtained through step 322 and step 324:
[0125] Step 322, obtain a plurality of view images of the sample character; the plurality of view images are images of the sample character from different perspectives.
[0126] The plurality of view images are images of the sample character from different perspectives. Taking the front view angle of the sample character as 0°, the perspectives include at least one of the following: 0, ±45°, ±90°, ±135°, 180°.
[0127] Specifically, the computer device obtains a plurality of view images of the sample character.
[0128] Step 324, draw the real UV texture image of the sample character according to the plurality of view images.
[0129] Specifically, the computer device performs texture projection on the plurality of view images according to the plurality of view images, and draws the real UV texture image of the sample character.
[0130] As an example, Figure 7 shows a schematic diagram of obtaining a real UV texture image provided by an exemplary embodiment of the present application. Refer to Figure 7 the perspective schematic diagram 4200 in. By setting cameras at different perspectives (8 shooting angles) respectively, the sample character is photographed to obtain a plurality of view images 4220 of the sample character. The computer device obtains the plurality of view images 4220 of the sample character, and through texture projection on the plurality of view images 4220, the real UV texture image 4410 of the sample character can be drawn.
[0131] In this embodiment, the plurality of view images of the sample character can relatively completely represent the overall shape and geometric structure of the sample character, and can ensure that there are no gaps in the drawn real UV texture image as much as possible, which can improve the accuracy of the drawn real UV texture image, thereby improving the accuracy of the trained text UV texture image model.
[0132] In some embodiments, a plurality of view images include two types, one being a plurality of view images of a real human body, and the other being a plurality of synthesized view images of a digital human body. Correspondingly, there are two ways to obtain a plurality of view images, which are introduced separately below. In actual use, one of the ways can be used, or the two ways can be combined and used.
[0133] · Method 1 for obtaining a plurality of view images
[0134] In some embodiments, a plurality of view images refer to a plurality of view images of a real human body. Then the sample role includes a first real sample role, and the first real sample role is a real human body. In this embodiment, a plurality of view images of the first real sample role can be obtained through sample video data. Specifically, step 322 can be optionally implemented as steps 322-1, 322-2, and 322-3:
[0135] Step 322-1: Obtain sample video data including the first real sample role.
[0136] The first real sample role refers to a real human body.
[0137] Sample video data refers to publicly available video data containing a first real sample character. Optionally, the source of the sample video data includes at least one of the following: People-Snapshot dataset (Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Video based reconstruction of 3D people models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8387–8397, 2018.) and iPER dataset (Wen Liu, Zhixin Piao, Jie Min, Wenhan Luo, Lin Ma, and Shenghua Gao. Liquid warping gan: A unified framework for human motion imitation, appearance transfer and novel view synthesis. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 5904–5913, 2019.)
[0138] Specifically, the computer device obtains publicly available sample video data containing a first real sample character.
[0139] Step 322-2: Determine the multi-view information of the first real sample character in the sample video data, where the multi-view information includes at least one of global rotation information and joint pose information.
[0140] Multi-view information refers to image information for capturing the first real sample character from different shooting angles. The computer device can determine several view images of the first real sample character through the multi-view information.
[0141] Optionally, the multi-view information includes at least one of the global rotation information and joint pose information of the first real sample character. Among them, the global rotation information is used to characterize the global rotation angle of the first real sample character, and the joint pose information is used to characterize the poses of the respective joints of the first real sample character.
[0142] In some embodiments, the computer device uses an action capture algorithm (Carry Location Information in Full Frames, CLIFF) to estimate the multi-view information of the first real sample character in at least a part of the sample video frames in the sample video data. Among them, the action capture algorithm is a method for estimating human body postures and shapes by combining full-frame location information. It mainly splices the position of the CLIFF model detection box with the feature vector of the cropped image to provide more global information for the CLIFF model, and improves the estimation accuracy of the global rotation angle and joint postures of the first real sample character.
[0143] In some other embodiments, due to the diversity of the sources of the sample video data, before performing this step, the computer device can also perform background segmentation on each sample video frame in the sample video data, and then perform the step of determining the multi-view information of the first real sample character in the sample video data to improve the estimation accuracy.
[0144] Step 322-3: Reconstruct a plurality of view images of the first real sample character according to the multi-view information of the first real sample character.
[0145] Specifically, the computer device adds the multi-view information of the first real sample character to the initial blank image to obtain a reconstructed view image, and reconstructs a plurality of view images of the first real sample character with the goal of minimizing the difference between the sample video frames in the sample video data and this reconstructed view image.
[0146] · Method 2 for obtaining a plurality of view images
[0147] In some embodiments, the plurality of view images refer to a plurality of view images of a digital human body. Then the sample character includes a second virtual sample character, and the second virtual sample character is a digital human body. There is no publicly available dataset for the digital human body. In this embodiment, a plurality of view images of the second virtual sample character are generated through a series of neural network models. Specifically, step 322 can be optionally implemented as steps 322-4 and 322-5:
[0148] Step 322-4: Input the image of the second virtual sample character into a pose control model to generate a plurality of pose images of the second virtual sample character; the plurality of pose images are bone images of the second virtual sample character with the same pose at different shooting angles, and the pose control model is used to generate the bone images of the second virtual sample character.
[0149] The second virtual sample character refers to a digital human body.
[0150] The image of the second virtual sample character refers to any two-dimensional image containing the second virtual sample character.
[0151] A skeletal image refers to an image that contains the skeleton of a second virtual sample character.
[0152] A pose image refers to a skeletal image of a second virtual sample character maintaining a certain pose.
[0153] The pose control model (DWpose) is used to generate the skeletal images of the second virtual sample character. This pose control model is a pre-trained open-source neural network model for achieving effective full-body pose estimation through a two-stage distillation method, for generating the skeletal images of the second virtual sample character, and then generating a number of pose images of the second virtual sample character maintaining the same pose.
[0154] Specifically, the computer device inputs the image of the second virtual sample character into the pose control model to generate a number of pose images of the second virtual sample character. It should be noted that the image of the second virtual sample character can include: the second virtual sample character maintaining any pose, and the image at any shooting angle.
[0155] In some embodiments, taking the front of the second virtual sample character being 0° as an example, the shooting angles corresponding to a number of pose images include at least one of the following: 0, ±45°, ±90°, ±135°, 180°.
[0156] Step 322-5, input a number of pose images and the description texts respectively corresponding to the number of pose images into the image processing model. Based on the image processing model, convert the number of pose images into a number of view images of the second virtual sample character; the image processing model is used to generate multi-view images of the second virtual sample character; wherein, the number of view images corresponds one-to-one with the number of pose images, and the description text includes at least one of the front description text and the back description text.
[0157] The description text is the text used to describe the content of a number of view images of the second virtual sample character. Each view image among the number of view images can correspond to a kind of description text. The description text includes at least one of the front description text (T pos ) and the back description text (T nrg ). The front description text is the text used to describe the content of a number of view images of the second virtual sample character expected to be generated, specifically including at least one of the following: the identity description text (T id ), the pose description text (T pose ), the background description text (T bg)). The reverse description text is the text for describing the content of several view images of the second virtual sample character that is not expected to be generated. For a view image, its positive description text can be used as the reverse description text of another view image. For example, when generating the back view image of the second virtual sample character based on the back pose image, the positive description text can be: "Back, rear side"; the reverse description text can be: "Face, front side".
[0158] In some embodiments, the image processing model includes an image control model and an image generation model. Among them, the image control model (ControlNet) is used to generate an image under specific control conditions (in this embodiment, it is the description text), and is specifically used to refine and control the image elements in several view images of the first virtual sample character. Each image element includes at least one of the first virtual sample character pose and the image structure. The image structure can specifically include background, style, lines, colors, etc. The image generation model (Stable Diffusion) is a type of implicit diffusion model (LDM), and is specifically used to generate several view images that conform to the description text according to the description texts respectively corresponding to several pose images. In this embodiment, by inputting several pose images and the description texts respectively corresponding to the several pose images into the image control model and the image generation model, and combining the image control model and the image generation model, several view images of the second virtual sample character that meet the description text are jointly generated.
[0159] Specifically, input several pose images and the description texts respectively corresponding to the several pose images into the image processing model. Based on the image processing model, convert the several pose images into several view images of the second virtual sample character. Among them, the several view images correspond one-to-one with the several pose images.
[0160] As an example, Figure 8A schematic diagram of obtaining several view images provided by an exemplary embodiment of the present application is shown. The computer device inputs the image of the second virtual sample character into the pose control model DWpose to generate several pose images 4210 of the second virtual sample character. The several pose images 4210 are 8 in number and are skeletal images of the second virtual sample character with the same pose at different shooting angles (8). The computer device inputs the several pose images 4210, the positive description text 4211 and the negative description text 4212 respectively corresponding to each pose image of the several pose images 4210 into the image control model ControlNet and the image generation model Stable Diffusion. Based on the image control model ControlNet and the image generation model Stable Diffusion, the several pose images 4210 are converted into several view images 4220 of the second virtual sample character. Among them, the several view images 4220 correspond to the several pose images 4210 one by one.
[0161] In the above embodiment, multiple methods for obtaining several view images of the sample character are provided. In practical applications, one or a combination of multiple methods can be selected to improve the flexibility of data processing, ensure that relatively accurate and sufficient view images can be obtained, and facilitate the drawing of real UV texture images.
[0162] · Texture Projection of Several View Images
[0163] In some embodiments, after obtaining several view images of the sample character, the real UV texture image of the sample character can be drawn. Specifically, step 324 can be optionally implemented as step 324-1 and step 324-2:
[0164] Step 324-1, extract the UV texture from each of the several view images.
[0165] Specifically, the computer device extracts the UV texture from each of the several view images through texture sampling, or through an application program with texture extraction function installed and running, or through the DensePose method. Among them, the DensePose method is a method that maps 2D image pixels to the human 3D surface to achieve efficient pose estimation and is used to obtain the SMPL-UV space texture.
[0166] Step 324-2, fuse each UV texture to draw the real UV texture image of the sample character.
[0167] Specifically, the computer device fuses and renders the UV textures separately extracted from each view image to render a real UV texture image of the sample character.
[0168] In this embodiment, by separately extracting UV textures from each of several view images and performing fusion rendering, a real UV texture image can be generated from the several view images, ensuring the accuracy of the real UV texture image. Since there are multiple view images, gaps are minimized in the real UV texture image as much as possible to improve the accuracy of the real UV texture image.
[0169] 2.2. Detailed training steps
[0170] In some embodiments, after obtaining a sample text and determining the real UV texture image corresponding to the sample text, the sample text can be used to train a text UV texture image model. By minimizing the difference between the real UV texture image and the predicted UV texture image output by the text UV texture image model, the training of the text UV texture image model is achieved.
[0171] Figure 5 FIG. shows a schematic diagram of training a text UV texture image provided by an exemplary embodiment of the present application. The text UV texture image model 200 includes a text encoder 10 and an image generation network 50. The text encoder 10 is used to encode the sample text 11 into sample text features. The image generation network 50 is used to map and obtain the predicted noise of the predicted UV texture image corresponding to the sample text 11 and obtain the predicted UV texture image based on the predicted noise. The image generation network 50 is obtained by adding trainable model parameters 51 to a pre-trained image generation network. Specifically, step 360 is implemented as step 361:
[0172] Step 361, with reducing the error between the real UV texture image and the predicted UV texture image as the training objective, trains the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0173] Exemplarily, the computer device only needs to train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network. Then, the computer device uses minimizing the error between the real UV texture image and the predicted UV texture image as the training objective to train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0174] In this embodiment, the computer device only needs to train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network, which can improve the training speed of the text UV texture image model, reduce the computational amount and memory occupation.
[0175] In some embodiments, the training phase of the text UV texture image model includes a first sub-phase and a second sub-phase. The predicted UV texture image includes a first sub-predicted UV texture image of the first sub-phase and a second sub-predicted UV texture image of the second sub-phase.
[0176] In both the first sub-phase and the second sub-phase, it is necessary to train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network. After the training in the first sub-phase is completed, each first sub-predicted UV texture image output by the text UV texture image model is called a generated UV texture image (GT UV). In order to improve the consistency of text and image, an alignment enhancement strategy is further adopted. In the training of the second sub-phase, combined with the above-mentioned generated UV texture images, the training of the second sub-phase is continued.
[0177] Specifically, step 361 can be implemented as step 420, step 440, step 460, and step 480:
[0178] Step 420: Taking reducing the error between the real UV texture image and the first sub-predicted UV texture image as the training objective, train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0179] The first sub-predicted UV texture image is a predicted UV texture image generated by the text UV texture image model according to the sample text in the first sub-training phase.
[0180] Specifically, the computer device takes minimizing the error between the real UV texture image and the first sub-predicted UV texture image as the training objective, and trains the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0181] In some embodiments, taking a sample text as an example, a sample text can correspond to several first sub-predicted UV texture images. Optionally, the computer device can also screen several first sub-predicted UV texture images according to the metric scores respectively corresponding to the several first sub-predicted UV texture images, and use the first sub-predicted UV texture image with the highest metric score as the final first sub-predicted UV texture image corresponding to the sample text. Among them, the metric score (CLIPScore) is determined based on the similarity between the sample text and the first sub-predicted UV texture image. The magnitude of the similarity is positively correlated with the magnitude of the metric score. The greater the similarity, the greater the value of the metric score.
[0182] In some embodiments, there are several ways to determine whether the training of the first sub-training stage has ended: 1. Determine based on the experience of the developer. When the developer observes that the first sub-predicted UV texture image is close to the real UV texture image, it is determined that the training of the first sub-stage has ended; 2. Determine based on the training duration. When the preset training duration is reached, it is determined that the training of the first sub-stage has ended; 3. Determine based on the number of iterations. When the preset number of iterations is reached, it is determined that the first sub-training stage has ended; 4. Determine based on the sample test text in the sample test set. The following embodiments will describe the fourth determination method in detail.
[0183] Exemplarily, during the training of the first sub-stage, the computer device inputs the sample test text into the text UV texture image model to generate a test UV texture image corresponding to the sample test text; calculates the metric score corresponding to the test UV texture image, and the metric score is determined based on the similarity between the test UV texture image and the sample test text; when the metric score is greater than the set threshold, it is determined that the training of the first sub-stage of the text UV texture image model has ended.
[0184] The sample test text is used to evaluate the prediction performance of the text UV texture image model in the first sub-training stage. The test UV texture image refers to the predicted UV texture image corresponding to the sample test text generated by the text UV texture image model. The metric score (CLIP Score) is a quantification of the above prediction performance. The metric score is determined based on the similarity between the test UV texture image and the sample test text. The magnitude of the similarity has a positive correlation with the magnitude of the metric score. The greater the similarity, the greater the value of the metric score, and the better the prediction performance of the text UV texture image model. In another example, when there is also a corresponding real UV texture image for the sample test text, the metric score can also be determined based on the similarity between the test UV texture image and the real UV texture image corresponding to the sample test text.
[0185] The set threshold is a preset numerical value of the metric score for determining the end of the training of the first sub-stage. When the metric score is greater than the set threshold, the computer device can determine that the training of the first sub-stage of the text UV texture image model has ended.
[0186] Step 440, after the training of the first sub-stage ends, input the sample text corresponding to the sample role into the text UV texture image model after the training of the first sub-stage ends to generate a generated UV texture image corresponding to the sample text.
[0187] The generated UV texture image refers to the predicted UV texture image corresponding to the sample text generated by the text UV texture image model after the training of the first sub-stage ends.
[0188] Exemplarily, after the training in the first sub-phase ends, the computer device inputs the sample text corresponding to the sample character into the text UV texture image model after the training in the first sub-phase ends, and generates a generated UV texture image (GT UV) corresponding to the sample text. It should be noted that the above sample text can be the sample text that has been used, or some other sample text that has not been used.
[0189] In some embodiments, for each sample text, the text UV texture image model can output several generated UV texture images. Then the computer device can screen out a generated UV texture image with the highest metric score from the several generated UV texture images corresponding to each sample text, so that the finally trained text UV texture image model has the best text consistency.
[0190] Step 460: Use the generated UV texture image as the real UV texture image in the second sub-phase.
[0191] Specifically, the computer device uses the generated UV texture image as the real UV texture image in the second sub-phase. Add the sample text and its generated UV texture image back to the sample dataset, and retrain the text UV texture image model.
[0192] Step 480: Take reducing the error between the generated UV texture image and the second sub-predicted UV texture image as the training objective, and train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0193] The second sub-predicted UV texture image is the predicted UV texture image generated by the text UV texture image model according to the sample text in the second sub-training phase.
[0194] Specifically, the computer device takes minimizing the error between the generated UV texture image and the second sub-predicted UV texture image as the training objective, and trains the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0195] In some embodiments, in the first sub-training phase of the text UV texture image model, the computer device trains the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network based on the training loss. Then step 420 can be specifically implemented as step 422 and step 424:
[0196] Step 422: Determine the training loss of the text UV texture image model based on the noise corresponding to the real UV texture image corresponding to the sample text and the predicted noise corresponding to the first sub-predicted UV texture image.
[0197] In some embodiments, the text UV texture image model is a sub-model within the UV texture image model. The UV texture image model further includes a noisy image encoder (Stable Diffusion Encoder, SD Encoder). The output end of the noisy image encoder is connected to the input end of the image generation network.
[0198] The noisy image encoder is used to encode a real UV texture image into image encoding features, and add noise to the image encoding features to obtain noisy image encoding features. That is, the noisy image encoding features are obtained by encoding a real UV texture image into image encoding features and adding noise to the image encoding features. It should also be noted that the amount of this noise has a positive correlation with the change in the time dimension. For example, as the training progresses, more and more noise can be added.
[0199] The noise corresponding to the real UV texture image is the noise added by the noisy image encoder. The predicted noise corresponding to the first sub-predicted UV texture image is obtained through mapping by the image generation network. Specifically, the computer device, based on the image generation network, maps the sample text features corresponding to the sample text, the noisy image encoding features corresponding to the real UV texture image corresponding to the sample text, and the number of noise addition times corresponding to the noisy image encoding features, to obtain the predicted noise corresponding to the first sub-predicted UV texture image.
[0200] Optionally, the number of noise addition times refers to the number of times the noise in the noisy image encoding features is increased. The number of noise addition times is associated with the time step (t), and the amount of noise increased in each time step is the same. In one example, the value of the number of noise addition times is related to the number of time steps. For example, if the current time step is the 1st time step in the time dimension, the corresponding number of noise addition times at this time is 1 time; if the current time step is the 2nd time step in the time dimension, the corresponding number of noise addition times at this time is 2 times.
[0201] Exemplarily, the computer device subtracts the predicted noise corresponding to the first sub-predicted UV texture image from the noise corresponding to the real UV texture image to obtain a noise difference. Based on the norm corresponding to this noise difference, the training loss of the text UV texture image model is determined. Then the training loss L is represented as follows:
[0202]
[0203] Among them, the real UV texture image is represented as x, the sample text is represented as c, the noisy image encoder is represented as E, the noise is represented as ε, and the predicted noise is represented as φ unet (z t ,t,φ t-enc(c)), the number of noise addition is denoted as t, the image encoding feature is denoted as z, and the noisy image encoding feature is denoted as z t , the model parameters of the text encoder are denoted as φ t-enc , the model parameters of the image generation network are denoted as φ unet , denotes the square of the L2 norm.
[0204] Step 424: Based on the training loss, train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0205] Specifically, the computer device trains the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network based on the training loss, with the goal of minimizing the training loss. When the training loss reaches the minimum, the first sub-training stage of the text UV texture image model is completed. And / or, during the training of the first sub-stage, when the metric score is greater than the set threshold, the training of the first sub-stage of the text UV texture image model ends. Next, the computer device can perform the second sub-training stage on the text UV texture image model.
[0206] In some embodiments, in the second sub-training stage of the text UV texture image model, the computer device trains the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network based on the training loss. Then step 480 can be specifically implemented as step 482 and step 484:
[0207] Step 482: Based on the noise corresponding to the generated UV texture image and the predicted noise corresponding to the second sub-predicted UV texture image, determine the training loss of the text UV texture image model.
[0208] In some embodiments, the text UV texture image model is a sub-model of the UV texture image model. The UV texture image model further includes a noisy image encoder (Stable Diffusion Encoder, SD Encoder). The output end of the noisy image encoder is connected to the input end of the image generation network.
[0209] The noisy image encoder is used to encode the generated UV texture image into an image encoding feature, and add noise to the image encoding feature to obtain a noisy image encoding feature. That is, the noisy image encoding feature is obtained by encoding the generated UV texture image into an image encoding feature and adding noise to the image encoding feature. It should also be noted that the amount of this noise is positively correlated with the change in the time dimension. For example, as the training progresses, more and more noise can be added.
[0210] The noise corresponding to the generated UV texture image is the noise added by the noisy image encoder. The noise corresponding to the second sub-predicted UV texture image is obtained by mapping through the image generation network. Specifically, the computer device, based on the image generation network, maps the sample text features corresponding to the sample text, generates the noisy image encoding features corresponding to the UV texture image, and the number of noise addition times corresponding to the noisy image encoding features, to obtain the predicted noise corresponding to the second sub-predicted UV texture image.
[0211] Optionally, the number of noise addition times refers to the number of times the noise in the noisy image encoding features is increased. The number of noise addition times is associated with the time step (t), and the amount of noise added at each time step is the same. In one example, the value of the number of noise addition times is related to the number of time steps. For example, if the current time step is the 1st time step in the time dimension, then the corresponding number of noise addition times is 1; if the current time step is the 2nd time step in the time dimension, then the corresponding number of noise addition times is 2.
[0212] In some embodiments, the computer device subtracts the predicted noise corresponding to the second sub-predicted UV texture image from the noise corresponding to the generated UV texture image to obtain a noise difference. Based on the norm corresponding to the noise difference, the training loss of the text UV texture image model is determined. The training loss L1 is represented as follows:
[0213]
[0214] where the generated UV texture image is represented as x, the sample text is represented as c, the noisy image encoder is represented as E, the noise is represented as ε, and the predicted noise is represented as φ unet (z t ,t,φ t-enc (c)), the number of noise addition times is represented as t, the image encoding features are represented as z, the noisy image encoding features are represented as z t , the model parameters of the text encoder are represented as φ t-enc , the model parameters of the image generation network are represented as φ unet , represents the square of the L2 norm.
[0215] Step 484, based on the training loss, train the model parameters of the text encoder and the trainable model parameters of the image generation network of the text UV texture image model.
[0216] Specifically, based on the training loss, the computer device can optimize the denoising loss of the model parameters of the text UV texture image model. Taking minimizing the training loss as the training objective, the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network are trained. When the training loss reaches the minimum, the training of the text UV texture image model is completed.
[0217] As an example, in combination with Figure 5 , the overall training process of the text UV texture image model will be described. Specifically, the computer device inputs the sample text 11 into the text UV texture image model 200. The sample text is encoded into sample text features by the text encoder 10, and the first sub-predicted UV texture image corresponding to the sample text 11 is generated by the image generation network 50. Taking reducing the error between the real UV texture image corresponding to the sample text 11 and the first sub-predicted UV texture image as the training objective, the model parameters of the text encoder 10 of the text UV texture image model 200 and the trainable model parameters 51 of the image generation network 50 are trained; after the training in the first sub-stage is completed, the sample text 11 is input into the text UV texture image model 200 after the training in the first sub-stage is completed, and the generated UV texture image (GT UV image) 21 corresponding to the sample text 11 is generated; the generated UV texture image 21 is encoded into image encoding features 23 by the SD image encoder 20, and noise 22 is added to the image encoding features 23 to obtain noisy image encoding features 24. Based on the noise 22 corresponding to the generated UV texture image 21 and the predicted noise 12 corresponding to the second sub-predicted UV texture image, the training loss L1 of the text UV texture image model 200 is determined; by minimizing the training loss L1, the model parameters of the text encoder 10 of the text UV texture image model 200 and the trainable model parameters 51 of the image generation network 50 are trained.
[0218] In the above embodiments, detailed training steps of the text UV texture image model are provided. Among them, through the training of the text UV texture image model in the first sub-stage, a text UV texture image model with relatively high accuracy can already be obtained. By retraining the text UV texture image model in the second sub-stage, the text UV texture image model can have better text consistency, further improving the accuracy of the text UV texture image model. This is further conducive to improving the accuracy of the subsequent generated UV texture images.
[0219] 3. Usage process of the text UV texture image model (T2UV)
[0220] After the training of the text UV texture image model is completed, the trained text UV texture image model can be directly used. By inputting the UV description text of the target character into the text UV texture image model, the corresponding UV texture image can be generated. Figure 9 FIG. shows a flowchart of a method for constructing a texture data set provided by an exemplary embodiment of the present application. The text UV texture image model includes a text encoder and an image generation network. Then, step 240 is specifically implemented as step 620 and step 640:
[0221] Step 620, encoding the UV description text into text features through the text encoder.
[0222] Step 640, generating the UV texture image corresponding to the UV description text through the image generation network.
[0223] The text features are the features obtained after the UV description text is processed by the text encoder.
[0224] The UV texture image is an image generated after the text features are processed by the image generation network.
[0225] Specifically, the computer device inputs the UV description text of the target character into the text UV texture image model. The text UV texture image model encodes the UV description text into text features through the text encoder, inputs the text features into the image generation network, and generates the UV texture image corresponding to the UV description text through the image generation network.
[0226] In this embodiment, the computer device directly inputs the UV description text into the text UV texture image model, and the UV texture image corresponding to the UV description text can be obtained, which improves the data processing speed. Moreover, this process does not require any three-dimensional data or three-dimensional information, and can get rid of the dependence of the UV texture image generation process on three-dimensional data or three-dimensional information.
[0227] · Acquisition of UV description text
[0228] In some embodiments, step 220 is implemented as step 221:
[0229] Step 221, inputting at least one piece of identity description information corresponding to the target character into the large language model, and generating the UV description text corresponding to the target character based on the large language model.
[0230] Since the texture data set is determined based on a large number of UV texture images, and each UV texture image needs to be generated based on the UV description text, a large number of UV description texts are required. In this embodiment, a large language model (or a general language model) is used to quickly obtain the UV description text.
[0231] Specifically, the computer device inputs at least one piece of identity description information corresponding to the target role into the large language model, and generates a UV description text corresponding to the target role based on the large language model. Among them, the identity description information includes at least one of the name, residence, appearance, gender, clothing, hairstyle, age, category, and personality of the target role.
[0232] As an example, Figure 10 FIG. shows a schematic diagram of obtaining a UV description text provided by an exemplary embodiment of the present application. According to the type of the target role, the computer device can input different identity description information into the large language model 600, so as to generate a UV description text 610 corresponding to the target role based on the large language model 600. Among them, the target role is set to 4 categories, and each category has corresponding identity description information and a UV description text with a fixed sentence structure. The identity description information of the first category of target roles includes: residence, appearance, gender, clothing, hairstyle, age, and the UV description text 1 is the fixed sentence structure 1: (residence)(appearance)
gender
clothing
hairstyle
age
name
clothing
name
clothing
hairstyle
category
[0233] As an example, Figure 11 FIG. shows a schematic diagram of obtaining a UV texture image provided by an exemplary embodiment of the present application. The computer device obtains 1100 UV description texts 610 through the large language model 600. These 1100 UV description texts 610 are respectively: 259 UV description texts 1; 421 UV description texts 2; 355 UV description texts 3; 87 UV description texts 4. Then, the computer device can input these 1100 UV description texts 610 into the text UV texture image model (T2UV) 200 respectively, and generate 1100 UV texture images 611 through the text UV texture image model 200.
[0234] In this embodiment, a large number of UV description texts can be obtained through a large language model. Since the large language model is a relatively mature natural language processing model, it can improve the text quality and acquisition speed of the UV description texts.
[0235] · Screening of UV texture images
[0236] In some embodiments, the UV texture images corresponding to a UV description text include several UV texture images. Then, before step 260, the several UV texture images corresponding to the UV description text can also be screened to determine one UV texture image corresponding to each UV description text for subsequent rendering. Then the method may further optionally include step 710 and step 712:
[0237] Step 710, determine the UV texture images that meet the screening conditions from the several UV texture images.
[0238] Step 712, use the UV texture images that meet the screening conditions as the UV texture images corresponding to the UV description text; wherein, the screening conditions include at least one of the index score corresponding to the UV texture image being greater than the threshold and the index score corresponding to the UV texture image ranking among the top N, N>0, and the index score is determined based on the similarity between the UV texture image and the UV description text.
[0239] The screening conditions are used to characterize the conditions that the UV texture images finally required for rendering need to meet. The screening conditions can be set based on the index score (CLIP Score), and this index score is determined based on the similarity between several UV texture images corresponding to the UV description text and the UV description text.
[0240] The screening conditions include at least one of the index score corresponding to the UV texture image being greater than the threshold and the index score corresponding to the UV texture image ranking among the top N, N>0. The threshold can be set to the maximum value of the index scores corresponding to each UV texture image, and N can be set to 1.
[0241] Specifically, the computer device determines the UV texture images that meet the screening conditions from the several UV texture images, and uses the UV texture images that meet the screening conditions as the UV texture images corresponding to the UV description text. Accordingly, the UV texture image with the best effect can be selected for subsequent rendering.
[0242] In this embodiment, by screening the UV texture images, the UV texture image with the best semantic consistency for each UV description text can be determined, improving the accuracy of the UV texture images. In addition, since the UV texture images are screened, the amount of rendering data can also be reduced, improving the accuracy of the rendered images.
[0243] · Rendering of UV texture images
[0244] Figure 12 The figure shows a flowchart of a method for constructing a texture data set provided by an exemplary embodiment of the present application. In some embodiments, in order to simulate a real image, a computer device needs to render a UV texture image into a final two-dimensional rendered image. Step 260 is specifically implemented as step 720 and step 740:
[0245] Step 720: Obtain at least one piece of rendering information corresponding to the UV texture image.
[0246] Step 740: Based on at least one piece of rendering information, render the UV texture image to obtain a rendered image corresponding to the target character; wherein, the at least one piece of rendering information includes at least one of: an animation sequence for capturing the movement of the target character, lighting information, material information, camera information, and background image.
[0247] Rendering information is the information required to render a UV texture image into a corresponding two-dimensional rendered image.
[0248] Specifically, the computer device installs and runs at least one 3D rendering program. The computer device obtains at least one piece of rendering information corresponding to the UV texture image, and based on the at least one piece of rendering information, renders the UV texture image through the 3D rendering application program to obtain a rendered image corresponding to the target character.
[0249] Optionally, the at least one piece of rendering information includes at least one of: an animation sequence for capturing the movement of the target character, lighting information (HDR Lighting), material information (Material), camera information (Camera), and background image (Background).
[0250] Specifically, the animation sequence is an SMPL-X animation sequence, and the motion rate corresponding to the animation sequence is the same as the rendering rate; the light source in the light information is a High Dynamic Range Imaging (HDR) image, and the dielectric specular reflection in the material information is a first preset value, the roughness is a second preset value, the gloss color is a third preset value, the varnish roughness is a fourth preset value, and the transmission refractive index is a fifth preset value; each camera in the camera information is used to follow the movement of a preset joint in the animation sequence, and each camera is located at a preset position, the focal length is a sixth preset value, the rendering resolution is a seventh preset value, the number of samples per pixel is an eighth preset value, and the frame rate is a ninth preset value; the background image includes at least one of a natural scene, a city street, an indoor scene, an abstract texture, and a solid color image. In one example, the rendering rate is 24 frames per second (FPS), the first preset value is 0.1, the second preset value is 0.6, the third preset value is 0.5, the fourth preset value is 0.03, the fifth preset value is 1.45, the preset joint can be any joint of the whole body, in this embodiment, it is set as the pelvis, the preset position is 5 meters (m) in front of the grid, the sixth preset value is 80 millimeters (mm), the seventh preset value is 1024*1024, the eighth preset value is 64, and the ninth preset value is 24 FPS.
[0251] As an example, Figure 13 FIG. shows a schematic diagram of a rendered UV texture image provided by an exemplary embodiment of the present application. In order to describe a real human body image in a natural scene, the computer device synthesizes the final rendered image 717 (which can also be called a rendered frame) using the generated UV texture image 611, the animation sequence 712, and the background image 716 of the Blender rendering pipeline. The computer device sets the light information 713, the camera information 714, and the material information 715 during rendering. Among them, the material information 715 sets the PBR material shader, the light information 713 sets the high dynamic range imaging HDR image illumination, and the camera information 714 sets a camera constrained by the SMPL-X skeleton.
[0252] Specifically, for the animation sequence 712, in this embodiment, the SMPL-X animation sequence is added as a three-dimensional theme. Using a subset ACCAD of the motion capture dataset AMASS, 30 different motion sequences are used. These 20 motion sequences include the basic shapes of both females and males. The motion rate corresponding to the motion sequence is set to 24 FPS, the same as the rendering frame rate, thereby generating a minimum of 72 rendering frames and a maximum of 400 rendering frames. For the lighting information 713, the image-based lighting (IBL) process can simulate real scenes and ensure uniform lighting. In this embodiment, an HDR image is used as the "sunlight" light source. For the material information 715, in this embodiment, the bidirectional reflectance distribution function (BSDF), also known as the PBR material shader, is used. To obtain a realistic human-like material, the dielectric specular reflection is set to 0.1, the roughness is increased to 0.6, the gloss color is set to 0.5, the varnish roughness is set to 0.03, and the transmission refractive index (IOR) is set to 1.45. The Alpha channel remains 1, and the remaining channels are all 0. To further improve the rendering efficiency, EEVEE rendering can be used. For the camera information 714, since each SMPL-X animation sequence has a global transformation. To more completely capture each human body movement, each camera is set to follow the movement of the pelvis joint in the SMPL-X animation series. Each camera is located 5 m in front of the mesh, and the focal length is 80 mm. The rendering resolution is 1024*1024, the number of samples per pixel is 64, and the frame rate is 24 FPS. For the background image 716, 100 images from Pexels are used, including: natural scenes, city streets, indoor scenes, abstract textures, and solid color images. Using post-processing techniques, the Alpha channels of the background image 716 and the UV texture image 611 are calculated to synthesize a rendered image 717.
[0253] In this embodiment, by combining at least one rendering information of the UV texture image for rendering, each rendered image has a corresponding UV description text, SMPL-X animation sequence, background segmentation mask, and synthesized rendered video, improving the richness of the texture information corresponding to the rendered image. It can also make the rendered image more realistic and improve the accuracy of the rendered image. Since the texture dataset is constructed from the rendering data, the accuracy of the texture dataset can be improved.
[0254] Next, with reference to specific schematic diagrams, the method for constructing the texture dataset provided by the embodiments of the present application will be described as a whole. Figure 14 、 Figure 15 FIG. shows a schematic diagram of the method for constructing the texture dataset provided by an exemplary embodiment of the present application. The method mainly includes three processes: (a) efficient texture adaptive fine-tuning; (b) diverse texture generation; (c) synthesized rendering.
[0255] (a) Efficient texture adaptive fine-tuning
[0256] Please refer to Figure 14 In part (a) of, to simplify the process of obtaining the UV texture image corresponding to the target character, in this embodiment, a text UV texture image model (T2UV) 200 is pre-trained. T2UV 200 can generate the corresponding UV texture image according to the input UV description text.
[0257] The training phase of T2UV 200 requires the use of 91 real UV texture images. In this embodiment, a number of view images of the first real sample character (real human body) are obtained by video back-projection, and a number of view images of the second virtual sample character (digital human body) are obtained by texture projection drawing technology, and then the corresponding real UV texture images of the number of view images are drawn respectively.
[0258] · The steps for the computer device to obtain the real UV texture image are briefly described as follows:
[0259] 1. Video back-projection method: Obtain sample video data containing the first real sample character from the People-Snapshot dataset and the iPER dataset; use the CLIFF method to estimate the multi-view information of the first real sample character in the sample video frames of the sample video data, and the multi-view information includes global rotation information and joint pose information; according to the multi-view information, reconstruct a number of view images of the first real sample character.
[0260] 2. Texture projection drawing technology: Input the image of the second virtual sample character into the pose control model DWpose to generate a number of pose images of the second virtual sample character; the number of pose images are the bone images of the second virtual sample character with the same pose at different shooting angles, and the pose control model is used to generate the bone images of the second virtual sample character; input the number of pose images, the positive description text and the negative description text corresponding to the number of pose images respectively into the image control model ControlNet and the image generation model LDM to convert the number of pose images into a number of view images of the second virtual sample character.
[0261] 3. Refer to Figure 14 In part (a) of, the computer device extracts and fuses the UV texture from each of the number of view images 61, and draws the real UV texture image 62.
[0262] · The steps for the computer device to train T2UV are briefly described as follows:
[0263] The schematic diagram of the training process of T2UV 200 can be referred to Figure 5 . Continue to refer to Figure 14In part (a), the computer device obtains the sample text 63, and the sample text 63 corresponds to the real UV texture image 62. The computer device inputs the sample text 63 into T2UV200 to generate the predicted UV texture image corresponding to the sample text 63. Taking reducing the error between the real UV texture image 63 and the predicted UV texture image as the training objective, the model parameters of the text encoder of T2UV200 and the trainable model parameters in the image generation network are trained. After the training of T2UV200 is completed, the computer device can continue to execute (b) diverse texture generation.
[0264] (b) Diverse texture generation
[0265] Please refer to Figure 14 part (b) in. The computer device can use the trained T2UV200 to generate various UV texture images. The computer device obtains 1100 UV description texts 64 through the large language model. In this embodiment, the UV description texts are pre-divided into 4 categories. These 1100 UV description texts 610 are respectively: 259 UV description texts 1; 421 UV description texts 2; 355 UV description texts 3; 87 UV description texts 4. Then, the computer device can input these 1100 UV description texts 64 into T2UV200 respectively, and 1100 UV texture images 65 can be generated.
[0266] (c) Composite rendering
[0267] Please refer to Figure 15 , in order to simulate real human images in natural scenes, in this embodiment, the generated UV texture image 65, the SMPL-X animation sequence 67 and the background image 66 of the Blender rendering pipeline are used to synthesize the final rendered image (rendering frame) 71 in the ATLAS application. In order to improve the authenticity, the material information 70 is set as the PBR human body material shader, the lighting information 68 is set as the high dynamic range imaging HDR image illumination, and the camera information 69 is set to use a camera constrained by the SMPL-X skeleton.
[0268] · Animation sequence 67: The SMPL-X animation sequence is added as the three-dimensional theme. Using a subset ACCAD of the human motion capture dataset AMASS, 30 different motion sequences are used. These 20 motion sequences include the basic shapes of women and men, and the motion rate corresponding to the motion sequence is set to 24 FPS, which is the same as the rendering frame rate, so as to generate a minimum of 72 rendering frames and a maximum of 400 rendering frames.
[0269] · Lighting information 68 and material information 70: The image-based lighting (IBL) process can simulate real scenes and ensure uniform lighting. In this embodiment, an HDR image is used as the "sunlight" light source. For the material information 70, the bidirectional reflectance distribution function (BSDF), also known as the PBR material shader, is used in this embodiment. To obtain a realistic human-like material, the dielectric specular reflection is set to 0.1, the roughness is increased to 0.6, the gloss color is set to 0.5, the varnish roughness is set to 0.03, and the transmission refractive index (IOR) is set to 1.45. The Alpha channel is kept as 1, and the remaining channels are all 0. To further improve the rendering efficiency, EEVEE rendering can be used.
[0270] · Camera information 69: Since each SMPL-X animation sequence has a global transformation. To more completely capture each human movement, each camera is set to follow the movement of the pelvis joint in the SMPL-X animation series. Each camera is located 5m in front of the mesh, and the focal length is 80mm. The rendering resolution is 1024*1024, the number of samples per pixel is 64, and the frame rate is 24FPS. For the background image 66, 100 images from Pexels are used, including: natural scenes, city streets, indoor scenes, abstract textures, and solid color images. Using post-processing techniques, the Alpha channels of the background image 66 and the UV texture image 65 are calculated to synthesize a rendered image 71. The rendered image 71 has corresponding UV description text, SMPL-X animation sequence, background segmentation mask, and synthesized rendered video. This rendered video 71 is used to generate a texture dataset, which is referred to as the multi-modal 3D human dataset ATLAS in this embodiment.
[0271] Table 1 shows the comparison of the texture dataset constructed in this embodiment with the 3D human datasets in the related art in terms of data type and quality. It can be seen from Table 1 that the texture dataset constructed in this embodiment has a larger data volume, higher image resolution, and better UV texture quality.
[0272] Table 1
[0273]
[0274]
[0275] In summary, the method provided by the embodiments of the present application can construct a multi-modal three-dimensional human dataset ATLAS, which is the first and largest high-resolution (1024*1024) three-dimensional human texture dataset in existence, filling the gap in high-quality three-dimensional human texture datasets and improving the generation efficiency of generating three-dimensional human models to a certain extent. At the same time, the ATLAS dataset contains 1100 high-fidelity human textures, each texture is paired with a UV description text, an SMPL-X animation sequence, and a synthetic rendering frame with a background segmentation mask, generating 7.5 million synthetic rendering frames. The ATLAS dataset has the characteristics of multi-modal, effectively solving the problem that the texture dataset lacks three-dimensional information or UV texture images, and is a texture dataset with high-quality texture images and rich texture information.
[0276] Figure 16 FIG. shows a block diagram of a texture dataset construction device 800 provided by an exemplary embodiment of the present application. The texture dataset construction device 800 includes:
[0277] An acquisition module 810, configured to acquire a UV description text corresponding to a target character, where the UV description text is used to describe the content of the UV texture image corresponding to the target character;
[0278] A processing module 820, configured to input the UV description text into a text UV texture image model to generate a UV texture image corresponding to the UV description text;
[0279] A rendering module 830, configured to render the UV texture image to obtain a rendering image corresponding to the target character, where the rendering image includes the texture information of the target character, and the rendering image is used to construct the texture dataset.
[0280] In some embodiments, the processing module 820 is configured to:
[0281] Acquire a sample text, where the sample text corresponds to a real UV texture image of a sample character, the sample text is used to describe the content of the real UV texture image, and the real UV texture image is drawn based on a plurality of view images;
[0282] Input the sample text into the text UV texture image model to generate a predicted UV texture image corresponding to the sample text;
[0283] Use reducing the error between the real UV texture image and the predicted UV texture image as a training objective to train the model parameters of the text UV texture image model.
[0284] In some embodiments, the acquisition module 810 is configured to:
[0285] Obtain the plurality of view images of the sample character; the plurality of view images are images of the sample character from different perspectives;
[0286] Draw the true UV texture image of the sample character according to the plurality of view images.
[0287] In some embodiments, the sample character includes a first true sample character;
[0288] In some embodiments, the obtaining module 810 is configured to:
[0289] Obtain sample video data including the first true sample character;
[0290] Determine the multi-view information of the first true sample character in the sample video data, where the multi-view information includes at least one of global rotation information and joint pose information;
[0291] Reconstruct the plurality of view images of the first true sample character according to the multi-view information of the first true sample character.
[0292] In some embodiments, the sample character includes a second virtual sample character;
[0293] In some embodiments, the obtaining module 810 is configured to:
[0294] Input the image of the second virtual sample character into a pose control model to generate a plurality of pose images of the second virtual sample character; the plurality of pose images are skeletal images of the second virtual sample character with the same pose at different shooting angles, and the pose control model is used to generate the skeletal images of the second virtual sample character;
[0295] Input the plurality of pose images and the description texts respectively corresponding to the plurality of pose images into an image processing model, and based on the image processing model, convert the plurality of pose images into a plurality of view images of the second virtual sample character; the image processing model is used to generate the multi-view images of the second virtual sample character;
[0296] Wherein, the plurality of view images correspond to the plurality of pose images one by one, and the description text includes at least one of a front description text and a back description text.
[0297] In some embodiments, the obtaining module 810 is configured to:
[0298] Extract UV textures from each of the plurality of view images respectively;
[0299] Fuse each of the UV textures to draw the true UV texture image of the sample character.
[0300] In some embodiments, the text UV texture image model includes a text encoder and an image generation network. The text encoder is used to encode the sample text into sample text features, and the image generation network is used to generate the predicted UV texture image corresponding to the sample text and map to obtain the predicted noise corresponding to the predicted UV texture image. The image generation network is obtained by adding trainable model parameters to a pre-trained image generation network.
[0301] In some embodiments, the processing module 820 is configured to:
[0302] Taking reducing the error between the true UV texture image and the predicted UV texture image as the training objective, train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0303] In some embodiments, the training stage of the text UV texture image model includes a first sub-stage and a second sub-stage. The predicted UV texture image includes a first sub-predicted UV texture image in the first sub-stage and a second sub-predicted UV texture image in the second sub-stage.
[0304] In some embodiments, the processing module 820 is configured to:
[0305] Taking reducing the error between the true UV texture image and the first sub-predicted UV texture image as the training objective, train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0306] After the training in the first sub-stage ends, input the sample text corresponding to the sample character into the text UV texture image model after the training in the first sub-stage ends to generate the generated UV texture image corresponding to the sample text.
[0307] Use the generated UV texture image corresponding to the sample text as the true UV texture image in the second sub-stage.
[0308] Taking reducing the error between the generated UV texture image and the second sub-predicted UV texture image as the training objective, train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0309] In some embodiments, the processing module 820 is configured to:
[0310] During the training of the first sub-phase, the sample test text is input into the text UV texture image model to generate a test UV texture image corresponding to the sample test text;
[0311] Calculate the metric score corresponding to the test UV texture image, where the metric score is determined based on the similarity between the test UV texture image and the sample test text;
[0312] When the metric score is greater than the set threshold, it is determined that the training of the first sub-phase of the text UV texture image model ends.
[0313] In some embodiments, the processing module 820 is configured to:
[0314] Determine the training loss of the text UV texture image model based on the noise corresponding to the generated UV texture image and the predicted noise corresponding to the second sub-predicted UV texture image;
[0315] Based on the training loss, train the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
[0316] In some embodiments, the processing module 820 is configured to:
[0317] Based on the image generation network, map the sample text features corresponding to the sample text, the noisy image encoding features corresponding to the generated UV texture image, and the number of noise addition times corresponding to the noisy image encoding features to obtain the predicted noise corresponding to the second sub-predicted UV texture image;
[0318] Wherein, the noisy image encoding feature is obtained by encoding the generated UV texture image into an image encoding feature and adding the noise to the image encoding feature.
[0319] In some embodiments, the text UV texture image model includes a text encoder and an image generation network; the processing module 820 is configured to:
[0320] Encode the UV description text into text features through the text encoder;
[0321] Generate the UV texture image corresponding to the UV description text through the image generation network.
[0322] In some embodiments, the obtaining module 810 is configured to:
[0323] Input at least one piece of identity description information corresponding to the target role into the large language model, and generate the UV description text corresponding to the target role based on the large language model.
[0324] In some embodiments, the UV texture image corresponding to the UV description text includes a plurality of UV texture images; the processing module 820 is further configured to:
[0325] Determine a UV texture image that meets the screening conditions from the plurality of UV texture images;
[0326] Use the UV texture image that meets the screening conditions as the UV texture image corresponding to the UV description text;
[0327] Wherein, the screening conditions include at least one of the index score corresponding to the UV texture image being greater than a threshold value and the index score corresponding to the UV texture image ranking among the top N, where N>0, and the index score is determined based on the similarity between the UV texture image and the UV description text.
[0328] In some embodiments, the rendering module 830 is configured to:
[0329] Obtain at least one piece of rendering information corresponding to the UV texture image;
[0330] Render the UV texture image based on the at least one piece of rendering information to obtain the rendering image corresponding to the target character;
[0331] Wherein, the at least one piece of rendering information includes at least one of an animation sequence for capturing the movement of the target character, lighting information, material information, camera information, and background image.
[0332] It should be noted that the specific limitations in the embodiments of the construction device 800 of one or more texture data sets provided above can be referred to the limitations on the construction method of the texture data set in the above text, and will not be elaborated here. Each module of the above device can be implemented in whole or in part by software, hardware, and their combination. Each module can be embedded in the processor of the computer device in hardware form or be independent, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0333] The embodiments of the present application further provide a computer device, which includes: a processor and a memory, and a computer program is stored in the memory; the processor is configured to execute the computer program in the memory to implement the texture data set construction method provided in each of the above method embodiments.
[0334] Exemplarily, Figure 17 is the structural block diagram of the computer device 1000 provided by an exemplary embodiment of the present application. Optionally, the computer device 1000 is a server 1000.
[0335] Generally, the server 1000 includes a processor 1001 and a memory 1002.
[0336] The processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1001 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1001 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.
[0337] The memory 1002 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1002 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1002 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1001 to implement the method for constructing a texture data set provided in the method embodiments of the present application.
[0338] In some embodiments, the server 1000 may further optionally include: an input interface 1003 and an output interface 1004. The processor 1001, the memory 1002, the input interface 1003, and the output interface 1004 may be connected via a bus or signal lines. Each peripheral device may be connected to the input interface 1003 and the output interface 1004 via a bus, signal lines, or a circuit board. The input interface 1003 and the output interface 1004 may be used to connect at least one peripheral device related to input / output (I / O) to the processor 1001 and the memory 1002. In some embodiments, the processor 1001, the memory 1002, the input interface 1003, and the output interface 1004 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, the memory 1002, the input interface 1003, and the output interface 1004 may be implemented on a separate chip or circuit board, and the embodiments of the present application do not limit this.
[0339] Those skilled in the art can understand that Figure 17 the structure shown in does not constitute a limitation on the computer device 1000, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.
[0340] In an exemplary embodiment, the present application provides a chip, which includes programmable logic circuits and / or program instructions. When the chip runs on a computer device, it is used to implement the method for constructing a texture data set provided by the above method embodiment.
[0341] The present application provides a computer-readable storage medium, which stores a computer program. The computer program is loaded and executed by a processor to implement the method for constructing a texture data set provided by the above method embodiment.
[0342] The present application provides a computer program product or a computer program, which includes computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the processor of the computer device is loaded and executed to implement the method for constructing a texture data set provided by the above method embodiment.
[0343] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0344] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned computer-readable storage medium can be a read-only memory, a disk, an optical disc, etc.
[0345] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0346] The above are only alternative embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for constructing a texture data set, characterized in that, The method includes: Obtaining a UV description text corresponding to a target character, where the UV description text is used to describe the content of the UV texture image corresponding to the target character; Inputting the UV description text into a text UV texture image model to generate a UV texture image corresponding to the UV description text; Rendering the UV texture image to obtain a rendered image corresponding to the target character, where the rendered image includes the texture information of the target character, and the rendered image is used to construct the texture dataset.
2. The method according to claim 1, characterized in that The text UV texture image model is obtained through the following training method: Obtaining a sample text, where the sample text corresponds to a real UV texture image of a sample character, the sample text is used to describe the content of the real UV texture image, and the real UV texture image is drawn based on a plurality of view images; Inputting the sample text into the text UV texture image model to generate a predicted UV texture image corresponding to the sample text; Using reducing the error between the real UV texture image and the predicted UV texture image as the training objective to train the model parameters of the text UV texture image model.
3. The method according to claim 2, wherein The method further includes: Obtaining the plurality of view images of the sample character; the plurality of view images are images of the sample character from different perspectives; Drawing the real UV texture image of the sample character according to the plurality of view images.
4. The method according to claim 3, wherein The sample character includes a first real sample character; the obtaining the plurality of view images of the sample character includes: Obtaining sample video data including the first real sample character; Determining multi-view information of the first real sample character in the sample video data, where the multi-view information includes at least one of global rotation information and joint pose information; Reconstructing the plurality of view images of the first real sample character according to the multi-view information of the first real sample character.
5. The method according to claim 3, characterized in that, The sample character includes a second virtual sample character; The obtaining the plurality of view images of the sample character includes: Inputting an image of the second virtual sample character into a pose control model to generate a plurality of pose images of the second virtual sample character; the plurality of pose images are skeletal images of the second virtual sample character with the same pose from different shooting angles, and the pose control model is used to generate the skeletal images of the second virtual sample character; Inputting the plurality of pose images and description texts respectively corresponding to the plurality of pose images into an image processing model, and based on the image processing model, converting the plurality of pose images into a plurality of view images of the second virtual sample character; the image processing model is used to generate the multi-view images of the second virtual sample character; Wherein, the plurality of view images correspond one by one to the plurality of pose images, and the description text includes at least one of a front description text and a back description text.
6. The method according to claim 3, characterized in that, The drawing the real UV texture image of the sample character according to the plurality of view images includes: Extract UV textures from each of the several view images respectively; Fuse the UV textures and draw the true UV texture image of the sample character.
7. The method according to any one of claims 2 to 6, characterized in that The text UV texture image model includes a text encoder and an image generation network. The text encoder is used to encode the sample text into sample text features, and the image generation network is used to map to obtain the predicted noise of the predicted UV texture image corresponding to the sample text and obtain the predicted UV texture image based on the predicted noise. The image generation network is obtained by adding trainable model parameters to a pre-trained image generation network; Using reducing the error between the true UV texture image and the predicted UV texture image as the training objective, training the model parameters of the text UV texture image model includes: Using reducing the error between the true UV texture image and the predicted UV texture image as the training objective, training the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
8. The method according to claim 7, wherein The training stage of the text UV texture image model includes a first sub-stage and a second sub-stage. The predicted UV texture image includes a first sub-predicted UV texture image in the first sub-stage and a second sub-predicted UV texture image in the second sub-stage; Using reducing the error between the true UV texture image and the predicted UV texture image as the training objective, training the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network includes: Using reducing the error between the true UV texture image and the first sub-predicted UV texture image as the training objective, training the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network; After the training of the first sub-stage ends, input the sample text corresponding to the sample character into the text UV texture image model after the training of the first sub-stage ends to generate the generated UV texture image corresponding to the sample text; Use the generated UV texture image corresponding to the sample text as the true UV texture image of the second sub-stage; Using reducing the error between the generated UV texture image and the second sub-predicted UV texture image as the training objective, training the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network.
9. The method according to claim 8, wherein The method further includes: During the training of the first sub-stage, input the sample test text into the text UV texture image model to generate the test UV texture image corresponding to the sample test text; Calculate the metric score corresponding to the test UV texture image. The metric score is determined based on the similarity between the test UV texture image and the sample test text; When the metric score is greater than the set threshold, determine that the training of the first sub-stage of the text UV texture image model ends.
10. The method according to claim 8, characterized in that Using the reduction of the error between the generated UV texture image and the second sub-predicted UV texture image as the training objective, training the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network, includes: Determining the training loss of the text UV texture image model based on the noise corresponding to the generated UV texture image and the predicted noise corresponding to the second sub-predicted UV texture image; Training the model parameters of the text encoder of the text UV texture image model and the trainable model parameters of the image generation network based on the training loss.
11. The method according to claim 10, characterized in that, The method further includes: Based on the image generation network, mapping the sample text features corresponding to the sample text, the noisy image encoding features corresponding to the generated UV texture image, and the number of noise addition times corresponding to the noisy image encoding features, to obtain the predicted noise corresponding to the second sub-predicted UV texture image; Wherein, the noisy image encoding features are obtained by encoding the generated UV texture image into image encoding features and adding the noise to the image encoding features.
12. The method according to any one of claims 1 to 11, characterized in that, The text UV texture image model includes a text encoder and an image generation network; The step of inputting the UV description text into the text UV texture image model to generate the UV texture image corresponding to the UV description text, includes: Encoding the UV description text into text features through the text encoder; Generating the UV texture image corresponding to the UV description text through the image generation network.
13. The method according to any one of claims 1 to 11, characterized in that, The step of obtaining the UV description text corresponding to the target character, includes: Inputting at least one identity description information corresponding to the target character into a large language model, and generating the UV description text corresponding to the target character based on the large language model.
14. The method according to any one of claims 1 to 11, characterized in that, The UV texture image corresponding to the UV description text includes a plurality of UV texture images; the method further includes: Determining the UV texture image that meets the screening conditions from the plurality of UV texture images; Using the UV texture image that meets the screening conditions as the UV texture image corresponding to the UV description text; Wherein, the screening conditions include at least one of the index score corresponding to the UV texture image being greater than a threshold value, and the index score corresponding to the UV texture image ranking among the top N, N>0, and the index score is determined based on the similarity between the UV texture image and the UV description text.
15. The method according to any one of claims 1 to 11, characterized in that The step of rendering the UV texture image to obtain the rendered image corresponding to the target character, includes: Obtaining at least one rendering information corresponding to the UV texture image; Rendering the UV texture image based on the at least one rendering information to obtain the rendered image corresponding to the target character; Wherein, the at least one rendering information includes at least one of an animation sequence for capturing the movement of the target character, lighting information, material information, camera information, and background image.
16. An apparatus for constructing a texture data set, characterized in that The device includes: An acquisition module, configured to acquire a UV description text corresponding to a target character, where the UV description text is used to describe the content of the UV texture image corresponding to the target character; A processing module, configured to input the UV description text into a text UV texture image model to generate a UV texture image corresponding to the UV description text; A rendering module, configured to render the UV texture image to obtain a rendered image corresponding to the target character, where the rendered image includes texture information of the target character, and the rendered image is used to construct the texture dataset.
17. A computer device, characterized in that, The computer device includes: a processor and a memory, where the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the method for constructing a texture dataset according to any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the method for constructing a texture dataset according to any one of claims 1 to 15.
19. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, and the processor obtains the computer instructions from the computer-readable storage medium, so that the processor loads and executes to implement the method for constructing a texture dataset according to any one of claims 1 to 15.